{"id":514,"date":"2023-02-16T20:14:41","date_gmt":"2023-02-16T18:14:41","guid":{"rendered":"https:\/\/thomaskosch.com\/?p=514"},"modified":"2023-10-24T10:38:58","modified_gmt":"2023-10-24T08:38:58","slug":"using-smartphones-to-sense-human-emotions-in-the-wild","status":"publish","type":"post","link":"https:\/\/thomaskosch.com\/index.php\/2023\/02\/16\/using-smartphones-to-sense-human-emotions-in-the-wild\/","title":{"rendered":"Using Smartphones to Sense Emotions In-the-Wild"},"content":{"rendered":"\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"485\" src=\"https:\/\/thomaskosch.com\/wp-content\/uploads\/2023\/02\/1-tl4xRJFdQLwO2JXHitJmwA-1024x485.png\" alt=\"\" class=\"wp-image-510\" srcset=\"https:\/\/thomaskosch.com\/wp-content\/uploads\/2023\/02\/1-tl4xRJFdQLwO2JXHitJmwA-1024x485.png 1024w, https:\/\/thomaskosch.com\/wp-content\/uploads\/2023\/02\/1-tl4xRJFdQLwO2JXHitJmwA-300x142.png 300w, https:\/\/thomaskosch.com\/wp-content\/uploads\/2023\/02\/1-tl4xRJFdQLwO2JXHitJmwA-768x364.png 768w, https:\/\/thomaskosch.com\/wp-content\/uploads\/2023\/02\/1-tl4xRJFdQLwO2JXHitJmwA.png 1400w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption>Emotion sensing via contextual data only. We present a new virtual emotion sensor embedded into a smartphone app that fuses various contextual information like vehicle- and traffic dynamics, road characterization, environmental weather, and in-vehicle context.<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">Imagine\n it\u2019s a rainy morning, and you are spending your time in a traffic jam \non Autobahn 8 near Stuttgart. The whole working week is only about to \nstart, with many appointments scheduled back-to-back. You\u2019re already \nlate for your first appointment with your manager\u2019s manager. How do you \nthink this situation would make you feel? Most likely not happy and \nquite stressed. From reading this anecdote, how do you come to this \nconclusion? It\u2019s all about the context \u2014 the information that describes \nthe situation you are currently experiencing.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This\n example made us curious: Why do so many companies and individuals try \nto infer emotions from observing the driver with cameras or microphones?\n Our studies found that most people look neutral and drive without \nemotional expressiveness. Car manufacturers use in-car microphones and \ncomputer vision to detect emotions by analyzing the driver\u2019s voice and \nfacial expressions. However, most people only talk while driving, \nespecially when driving alone. Filming the driver and the car interiors \nposes privacy threats and provide little utility for emotion detection. \nHence, we asked ourselves how to exploit driver context for emotion \nrecognition.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">We  proposed a novel approach to detect emotions in the wild. We use a  smartphone, which is a device that almost everyone has access to sensors  such as GPS, camera, and microphone to derive contextual features. The  context can give insights into a person\u2019s emotional state. The following  section will present the various context streams available for emotion  prediction.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>GPS<\/strong>: We use a smartphone\u2019s GPS  sensor to determine travel speed and acceleration. This information  helps to determine if and how the user is driving, walking, or using  public transport. Furthermore, via reverse-geocoding, the weather  (degree of sunlight, raininess, temperature outside) and traffic flow  (onset of traffic jam, reduced speed on road segment vs. average speed)  of the location via Google or Microsoft APIs give an environmental  (driving) context. OpenStreetMaps also allows for a better understanding  of the static road context, such as road type, speed limits, traffic  signs, and the number of available lanes. At last, the GPS location can  be used to get a satellite image of the surroundings and deduct the  average greenness. In total, a GPS position enables to proxy of the  person\u2019s activity, current weather conditions, surrounding traffic flow,  road types, and greenness.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Camera:<\/strong>\n Usually, the camera assesses the driver\u2019s facial expressions and \nderives its emotion. As the driver mostly looks unexpressed, facial \nexpressions mostly return neutral emotional states, although other \nemotions may be present. Yet, we can access the backward-camera frames \nand analyze the visual driving scene: the number of cars, trucks, and \nbicycles are extracted via computer vision approaches (e.g., VGG-16 \ntrained on ImageNet). This visual scene analysis provides a more direct \nsignal of how complex the scene in the wild for the person is, which \ndirectly relates to emotional feelings. A smartphone\u2019s camera can also \ndetermine the weather conditions outside. For instance, the driver might\n feel sad if it\u2019s raining, but if the sun is shining, they might feel \nhappy.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Voice:<\/strong>\n A smartphone\u2019s microphone can capture the driver\u2019s voice and detect \nchanges in tone or pitch. Mostly, people do not speak while driving, \nlimiting the performance of dynamic analysis of the signals. This \ninformation could be used to determine if the driver feels happy, sad, \nangry, or any other emotion. In our case, we use the microphone only for\n getting ground-truth labels.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Static Smartphone<\/strong>:\n We are capturing the daytime using the smartphone timestamps, providing\n the system information whether it is early in the morning (and people \nare potentially moodier).<\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"509\" src=\"https:\/\/thomaskosch.com\/wp-content\/uploads\/2023\/02\/0-dnVC5urr6TeRoDBE-1024x509.png\" alt=\"\" class=\"wp-image-512\" srcset=\"https:\/\/thomaskosch.com\/wp-content\/uploads\/2023\/02\/0-dnVC5urr6TeRoDBE-1024x509.png 1024w, https:\/\/thomaskosch.com\/wp-content\/uploads\/2023\/02\/0-dnVC5urr6TeRoDBE-300x149.png 300w, https:\/\/thomaskosch.com\/wp-content\/uploads\/2023\/02\/0-dnVC5urr6TeRoDBE-768x382.png 768w, https:\/\/thomaskosch.com\/wp-content\/uploads\/2023\/02\/0-dnVC5urr6TeRoDBE.png 1396w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption>We record contextual data (e.g., weather, road type, traffic flow) and the driver\u2019s facial expressions while driving. We fuse the collected data and use it as input for a machine learning predictor that predicts the driver\u2019s emotions. The audio stream is used to detect the baseline emotion in our study experiment. Facial expressions can be included as a feature in our system based on individual privacy policies and is therefore depicted as a dashed line. The audio stream input analyzed via a speech-to-text engine is used to extract the label for our system and is not included as an input feature to the machine learning engine.<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">In\n a user study, we validated whether contextual signals can predict \nemotions in the wild. We provided 27 participants with our smartphone \napp gathering real-time contextual signals and asked them to provide \ntheir current emotional state every 60s verbally. The verbally provided \nemotion represents the ground-truth label of emotions in the wild.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">We  train a random forest classification model to predict emotions for all  participants. We investigate how powerful each feature was for creating a  classification model. Overall, we detect the highest feature importance  of vehicle trajectory variables (e.g., vehicle speed). This might be  because \u2018happy\u2019 emotions are often reported in unhindered speed  scenarios. Similarly, higher negative emotions (i.e., \u2018anger\u2019 and  \u2018fear\u2019) are frequently observed during unforeseen traffic incidents  (e.g., increased traffic densities or red light series) that require  high cognitive demands of the driver.<\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"425\" src=\"https:\/\/thomaskosch.com\/wp-content\/uploads\/2023\/02\/0-ceKxIKf2CSASmyL0-1024x425.png\" alt=\"\" class=\"wp-image-511\" srcset=\"https:\/\/thomaskosch.com\/wp-content\/uploads\/2023\/02\/0-ceKxIKf2CSASmyL0-1024x425.png 1024w, https:\/\/thomaskosch.com\/wp-content\/uploads\/2023\/02\/0-ceKxIKf2CSASmyL0-300x125.png 300w, https:\/\/thomaskosch.com\/wp-content\/uploads\/2023\/02\/0-ceKxIKf2CSASmyL0-768x319.png 768w, https:\/\/thomaskosch.com\/wp-content\/uploads\/2023\/02\/0-ceKxIKf2CSASmyL0.png 1396w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption>Feature importances measured by the mean decrease of Gini-impurity for the Leave-One-Participant-Out cross-validation.<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">Our analysis shows that contextual emotion recognition is significantly more robust than facial recognition, leading to an overall improvement of 7% using a leave-one-participant-out cross-validation. Overall, our contextual features make it possible to predict emotions in real time for unseen roads and drivers. The overall accuracy of emotions is 71.70%. In other words, it is 29% better than relying on the \u2018facial expression\u2019 engine alone. We validated the facial expression engine using other common facial expression classifier systems (AWS, EmoPy), but the results remained the same.<\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"238\" src=\"https:\/\/thomaskosch.com\/wp-content\/uploads\/2023\/02\/0-QqF4uwwdtNu1IUrQ-1024x238.png\" alt=\"\" class=\"wp-image-513\" srcset=\"https:\/\/thomaskosch.com\/wp-content\/uploads\/2023\/02\/0-QqF4uwwdtNu1IUrQ-1024x238.png 1024w, https:\/\/thomaskosch.com\/wp-content\/uploads\/2023\/02\/0-QqF4uwwdtNu1IUrQ-300x70.png 300w, https:\/\/thomaskosch.com\/wp-content\/uploads\/2023\/02\/0-QqF4uwwdtNu1IUrQ-768x178.png 768w, https:\/\/thomaskosch.com\/wp-content\/uploads\/2023\/02\/0-QqF4uwwdtNu1IUrQ.png 1396w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption>Accuracy, precision, recall, and weighted F1 scores of the global 10-fold cross-validation on unseen consecutive driving segments, aggregate of participant-dependent Leave-One-of-10-Road-Segments-Out cross-validation, and Leave-One-Participant-Out cross-validation.<br><\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">Empathic\n interface designers can tailor the driver experience by changing the \nmusic playing or showing more relevant advertisements. In conclusion, \nusing smartphones to sense human emotions is a relatively new but \npromising field. With advancements in machine learning and artificial \nintelligence, this technology will likely become more accurate and \nwidely used. Whether it\u2019s to improve the driver experience or to \nunderstand human behavior better, using smartphones to sense human \nemotions will significantly impact the future.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Authors: David Bethge (LMU Munich), Thomas Kosch (HU Berlin), Tobias Grosse-Puppendahl (Dr. Ing. h. c. F. Porsche AG). This post is based on a joint post from <a href=\"https:\/\/medium.com\/@d.bethge\/using-smartphones-to-sense-human-emotions-in-the-wild-4d5f7e03d091\">Medium.<\/a><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">References:<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">[1] David  Bethge, Thomas Kosch, Tobias Grosse-Puppendahl, Lewis L. Chuang,  Mohamed Kari, Alexander Jagaciak, and Albrecht Schmidt. 2021. VEmotion:  Using Driving Context for Indirect Emotion Prediction in Real-Time. In  The 34th Annual ACM Symposium on User Interface Software and Technology  (UIST \u201821). Association for Computing Machinery, New York, NY, USA,  638\u2013651.<a rel=\"noreferrer noopener\" href=\"https:\/\/doi.org\/10.1145\/3472749.3474775\" target=\"_blank\"> https:\/\/doi.org\/10.1145\/3472749.3474775<\/a><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">[2] David  Bethge, Luis Falconeri Coelho, Thomas Kosch, Satiyabooshan  Murugaboopathy, Ulrich von Zadow, Albrecht Schmidt, and Tobias  Grosse-Puppendahl. 2023. Technical Design Space Analysis for Unobtrusive  Driver Emotion Assessment Using Multi-Domain Context. Proc. ACM  Interact. Mob. Wearable Ubiquitous Technol. 6, 4, Article 159 (December  2022), 30 pages.<a rel=\"noreferrer noopener\" href=\"https:\/\/doi.org\/10.1145\/3569466\" target=\"_blank\"> https:\/\/doi.org\/10.1145\/3569466<\/a><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">[3]  Bethge, D., Patsch, C., Hallgarten, P., &amp; Kosch, T. (2023, April).  Interpretable Time-Dependent Convolutional Emotion Recognition with  Contextual Data Streams. In Extended Abstracts of the 2023 CHI Conference on Human Factors in Computing Systems. <a href=\"https:\/\/doi.org\/10.1145\/3544549.3585672\">https:\/\/doi.org\/10.1145\/3544549.3585672<\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Imagine it\u2019s a rainy morning, and you are spending your time in a traffic jam on Autobahn 8 near Stuttgart. The whole working week is only about to start, with many appointments scheduled back-to-back. You\u2019re already late for your first appointment with your manager\u2019s manager. How do you think this situation would make you feel? [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":510,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[1],"tags":[],"class_list":["post-514","post","type-post","status-publish","format-standard","has-post-thumbnail","category-uncategorized","czr-hentry"],"_links":{"self":[{"href":"https:\/\/thomaskosch.com\/index.php\/wp-json\/wp\/v2\/posts\/514","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/thomaskosch.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/thomaskosch.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/thomaskosch.com\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/thomaskosch.com\/index.php\/wp-json\/wp\/v2\/comments?post=514"}],"version-history":[{"count":4,"href":"https:\/\/thomaskosch.com\/index.php\/wp-json\/wp\/v2\/posts\/514\/revisions"}],"predecessor-version":[{"id":542,"href":"https:\/\/thomaskosch.com\/index.php\/wp-json\/wp\/v2\/posts\/514\/revisions\/542"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/thomaskosch.com\/index.php\/wp-json\/wp\/v2\/media\/510"}],"wp:attachment":[{"href":"https:\/\/thomaskosch.com\/index.php\/wp-json\/wp\/v2\/media?parent=514"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/thomaskosch.com\/index.php\/wp-json\/wp\/v2\/categories?post=514"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/thomaskosch.com\/index.php\/wp-json\/wp\/v2\/tags?post=514"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}