MULTIMEDIA MESSAGING APPARATUS AND METHOD FOR TRANSMITTING MULTIMEDIA MESSAGES - Patent application
The multimedia messaging system addresses the challenge of conveying sender emotions by combining multimedia and text content with graphical representations, allowing for accurate emotional expression in multimedia messages.
Patent Information
- Application Number
- JP2025529917
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-09-29
- Filing Date
- 2023-11-17
- Publication Date
- 2026-01-14
AI Technical Summary
Existing multimedia messaging systems fail to accurately convey the sender's emotions, making it difficult for recipients to understand the emotional level intended by the sender.
A multimedia messaging device and method that combines multimedia content, corresponding text content, and a graphical representation of the sender's mood, allowing for the sequential display of multiple graphical representations based on the multimedia and text content, and enabling the transmission of multimedia messages via gestures.
Enables recipients to easily recognize the sender's sentiment by incorporating graphical representations of emotions, enhancing emotional expression in multimedia messages.
Smart Images

Figure 2026501071000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of and priority to U.S. Patent Application Nos. 18 / 478,382, 18 / 478,393, and 18 / 478,436, all filed on September 29, 2023, and U.S. Provisional Patent Application Nos. 63 / 427,466, 63 / 427,467, and 63 / 427,468, all filed on November 23, 2022, the contents of all of which are incorporated herein by reference in their entireties.
[0002] TECHNICAL FIELD The present disclosure relates generally to a multimedia messaging device and method for transmitting multimedia messages, and more particularly to a multimedia messaging device and method for transmitting multimedia messages with emotions. [Background technology]
[0003] Messaging applications are a common medium for communication between users. Various systems and applications can utilize multimedia messaging to send rich media to users. Multimedia messages may include video and / or audio data along with text messages. Generally, users can infer the sender's emotions from multimedia and text messages. However, multimedia messages may not truly convey the sender's emotions to other users because the sender's emotions and feelings are subjective and self-contained phenomenal experiences. Even if the receiving user understands the general emotion of the multimedia message, it may be difficult for the receiving user to understand the sender's emotional level. Summary of the Invention
[0004] The present disclosure relates to a multimedia messaging device and method for sending a multimedia message with a graphical representation of a sender's mood or emotion. In particular, the present disclosure relates to a multimedia messaging device and method for combining multimedia content, corresponding text content, and a graphical representation of a sender's mood and sequentially displaying multiple graphical representations based on the multimedia content and the corresponding text content. Furthermore, the present disclosure relates to a multimedia messaging device and method for sending a multimedia message based on a sender's gesture.
[0005] According to aspects of the present disclosure, a messaging device for sending multimedia messages via gestures includes a processor and a memory coupled to the processor, the memory including stored instructions that, when executed by the processor, cause the messaging device to receive sensor data from one or more sensors, analyze the received sensor data to identify gestures, perform a search within a list of predefined gestures based on the identified gestures, select a multimedia message based on the searched predefined gestures, and send the selected multimedia message via a messaging application.
[0006] In some embodiments, the list of predefined gestures is pre-stored.
[0007] In some embodiments, the messaging device further includes a touchscreen that functions as one or more sensors, the touchscreen generating sensor data based on a user's movements across the touchscreen.
[0008] In some embodiments, the one or more sensors include a motion sensor, the sensor being worn by a user, the motion sensor including at least one of a gyroscope, an accelerometer, a magnetometer, a radar, or a lidar.
[0009] In some embodiments, the sensor data is collected within a predetermined time period.
[0010] In some embodiments, a relational database between the list of predefined gestures and the multimedia messages is stored in the memory, and the selected multimedia message is associated with the retrieved predefined gesture based on the relational database.
[0011] According to aspects of the present disclosure, a messaging method for sending a multimedia message via a gesture includes receiving sensor data from one or more sensors, analyzing the received sensor data to identify a gesture, performing a search in a list of predefined gestures based on the identified gesture, selecting a multimedia message based on the searched predefined gesture, and sending the selected multimedia message via a messaging application.
[0012] In some embodiments, the list of predefined gestures is pre-stored.
[0013] In some embodiments, the messaging device further includes a touchscreen that functions as one or more sensors, the touchscreen generating sensor data based on a user's movements across the touchscreen.
[0014] In some embodiments, the one or more sensors include a motion sensor, the sensor being worn by a user, the motion sensor including at least one of a gyroscope, an accelerometer, a magnetometer, a radar, or a lidar.
[0015] In some embodiments, the sensor data is collected within a predetermined time period.
[0016] In some embodiments, a relational database between the list of predefined gestures and the multimedia messages is stored in the memory, and the selected multimedia message is associated with the retrieved predefined gesture based on the relational database.
[0017] According to aspects of the present disclosure, a non-transitory computer-readable storage medium includes stored instructions that, when executed by a computer, cause the computer to perform a messaging method for transmitting a multimedia message with an emotional expression. The method includes receiving sensor data from one or more sensors, analyzing the received sensor data to identify a gesture, performing a search in a list of predefined gestures based on the identified gesture, selecting a multimedia message based on the searched predefined gesture, and sending the selected multimedia message via a messaging application.
[0018] In some embodiments, the messaging device trains the first predetermined gesture in a list of predetermined gestures.
[0019] In some embodiments, training the first predetermined gesture in the list of predetermined gestures includes initially tapping the device on the user's chest.
[0020] In some embodiments, training a first predetermined gesture in the list of predetermined gestures includes one or more subsequent taps of the device on the user's chest to complete training of the first predetermined gesture.
[0021] In some aspects, tapping the device on the user's chest upon completion of training the first predetermined gesture includes selecting a multimedia message associated with the first predetermined gesture and sending the selected multimedia message via a messaging application. [Brief explanation of the drawings]
[0022] A detailed description of aspects of the present disclosure will be made with reference to the accompanying drawings, in which like numerals indicate corresponding parts throughout the drawings.
[0023] [Figure 1] FIG. 1 illustrates a block diagram of a multimedia messaging system for transmitting multimedia messages, according to one or more aspects. [Figure 2] FIG. 2 illustrates a graphical representation of a mobile device screen showing an exchange of multimedia messages, according to one or more aspects. [Figure 3] FIG. 3 illustrates a graphical flow representation of a mobile device screen for forming a flattened multimedia message, according to one or more aspects. [Figure 4] FIG. 4 illustrates a block diagram of layers in a flattened multimedia message, according to one or more aspects. [Figure 5] FIG. 5 illustrates a graphical representation of a flattened multimedia message, according to one or more aspects. [Figure 6] FIG. 6 illustrates a flowchart of a method for forming a flattened multimedia message, according to one or more aspects. [Figure 7] FIG. 7 illustrates a graphical representation of emotion categories, according to one or more embodiments. [Figure 8] FIG. 8 illustrates a block diagram of a multimedia messaging server, according to one or more aspects. [Figure 9] FIG. 9 illustrates a flowchart of a method for sequencing and displaying graphical representations of emotions, according to one or more embodiments. [Figure 10] FIG. 10 illustrates a block diagram for gesture recognition, according to one or more aspects. [Figure 11]FIG. 11 illustrates a flowchart for recognizing gestures for sending multimedia messages, according to one or more aspects. [Figure 12] FIG. 12 illustrates a block diagram of a computing device, according to one or more aspects. [Figure 13] FIG. 13 illustrates a flowchart of an exemplary method for training a gesture recognition module to identify gestures, according to one or more aspects. [Figure 14A] 14A-14E show graphical representations of mobile device screens illustrating exemplary heart bump training gestures, according to one or more embodiments. [Figure 14B] 14A-14E show graphical representations of mobile device screens illustrating exemplary heart bump training gestures, according to one or more embodiments. [Figure 14C] 14A-14E show graphical representations of mobile device screens illustrating exemplary heart bump training gestures, according to one or more embodiments. [Figure 14D] 14A-14E show graphical representations of mobile device screens illustrating exemplary heart bump training gestures, according to one or more embodiments. [Figure 14E] 14A-14E show graphical representations of mobile device screens illustrating exemplary heart bump training gestures, according to one or more embodiments. [Figure 15] FIG. 15 illustrates a graphical representation of a mobile device screen showing an exemplary heart bump gesture collection, according to one or more aspects. [Figure 16] FIG. 16 illustrates a graphical representation of a mobile device screen illustrating an exemplary heart bump gesture message sending operation, according to one or more aspects. DETAILED DESCRIPTION OF THE INVENTION
[0024] The present disclosure provides a multimedia messaging server, a computing device, and a method for flattening layers of content to generate a multimedia message, sequencing and displaying graphic representations of emotions, and transmitting the multimedia message based on gestures. The multimedia message incorporating the sender's sentiment is transmitted so that a recipient of the multimedia message can easily recognize the sender's sentiment. The sender's sentiment or emotion is represented by a corresponding graphic representation.
[0025] Furthermore, the sentiments or emotions are classified into different categories and levels. After analyzing the audiovisual content and the corresponding textual content, one emotion category is brought to the forefront, and other categories are subsequently presented based on their distance from that one category.
[0026] Additionally, sensor data is collected and analyzed to identify gestures, and flattened multimedia messages correlated with the identified gestures can be automatically sent to other users.
[0027] 1 shows a block diagram of a multimedia messaging system 100 for sending and receiving multimedia messages between users 140a-140n according to an embodiment of the present disclosure. The multimedia messaging system 100 includes an audiovisual content server 110, a text content server 120, and a multimedia messaging server 135. In one embodiment, the multimedia messaging system 100 may optionally include a social media platform 130. In this disclosure, users 140a-140n are referred to collectively and users 140 are referred to individually.
[0028] Users 140a-140n communicate with each other using a social media platform 130 by using their own computing devices 145a-145n, such as mobile devices or computers. The computing devices 145a-145n are used collectively, and the computing devices 145 are used individually, as are users 140a-140n and users 140. The social media platform 130 may be YouTube®, Facebook®, Twitter®, Tik Tok®, Pinterest®, Snapchat®, LinkedIn®, etc. The social media platform may also be a virtual reality (VR), mixed reality (MR), augmented reality (AR), and metaverse platform. In response to shared text and multimedia messages, users 140 may send graphic representations (e.g., emojis, graphic images, GIFs, animated GIFs, etc.) to indicate their emotions / sentiments / feelings. However, no social media platform 130 exists that sends multimedia messages incorporating the sender's sentiments. The multimedia messaging system 100, in accordance with aspects of the present disclosure, can incorporate the sentiment of the sender into a multimedia message.
[0029] In one aspect, the multimedia messaging server 135 may be incorporated within the social media platform 130 to enable users 140 to send multimedia messages along with sender sentiments. In another aspect, the multimedia messaging server 135 may be a standalone server for sharing multimedia messages along with sender sentiments among users 140a-140n.
[0030] The audiovisual content server 110 can be a music server containing song audio and music videos, an audio recording server storing recordings of phone conversations and meetings, a video server storing movie files or any video files, or any audiovisual server containing audiovisual content, each of which may be a digital file and may have information about its duration.
[0031] Audiovisual content server 110 categorizes audiovisual content into different genre categories, so that when user 140 searches for a particular genre, audiovisual content server 110 may send a list of audiovisual content for the searched genre to user 140. Similarly, audiovisual content server 110 may categorize audiovisual content into emotion / mood categories, so that audiovisual content server 110 can send a list of audiovisual content with a particular mood in response to a request from user 140 via multimedia messaging server 135.
[0032] Based on the audiovisual content stored on the audiovisual content server 110, the text content server 120 may have corresponding text content. In other words, if the audiovisual content server 110 has a song, the text content server 120 has the lyrics for that song, and if the audiovisual content server 110 has a recording of a phone conversation, the text content server 120 has the transcribed text of that phone conversation. The text content may be stored in "lrc" or "sct" format. The "lrc" format is a computer file format that synchronizes song lyrics with an audio file, such as MP3, Vorbis, or MIDI, while the "sct" format is a common subtitle file format for audiovisual content. These examples of text content format types are provided for illustrative purposes only and are not limiting. The list of format types may include other formats as would be readily understood by one of ordinary skill in the art.
[0033] These formats generally include two parts: one is a timestamp including a start time and an end time, and the other is a text information within the start time and end time. Therefore, if text content for a specific period (e.g., 00:01:30 to 00:03:00) is requested, the corresponding text content can be easily extracted and obtained from the text content based on the timestamp. In one aspect, the multimedia messaging server 135 may be able to extract and obtain a portion of the text content, or the text content server 120 may provide such a portion upon request.
[0034] In one aspect, if the selected audiovisual content is a recording of a conference and there is no textual content for the selected audiovisual content, the multimedia messaging server 135 may perform a transcription operation to transcribe the recording or may contact a transcription server to receive the transcribed text information.
[0035] After receiving the audiovisual content from the audiovisual content server 110 and the textual content from the textual content server 120, the multimedia messaging server 135 may combine the two to generate a flattened multimedia message that may include a graphical representation of the sender's emotions. In one aspect, the flattened multimedia message may be inseparable into multimedia content and textual content and may be stored locally on the sender's computing device 145. In another aspect, the multimedia messaging server 135 may store the flattened multimedia message in its memory under the sender's account.
[0036] 2 illustrates a screenshot 200 of a mobile device according to an embodiment of the present disclosure. The screenshot 200 illustrates a messaging interface 210 of a social media platform (e.g., the social media platform 130 in FIG. 1 ) or a multimedia messaging server (e.g., the multimedia messaging server 135 in FIG. 1 ) as a standalone system. Since the description of the messaging interface 210 of the multimedia messaging server is substantially similar to the description of the messaging interface 210 of the social media platform, the following description will be given of the social media platform, and the description of the multimedia messaging server will be omitted, and reference can be made to the following.
[0037] The messaging interface 210 includes a message chain 220, which may include multimedia messages, text messages, and emojis. The message chain 220 is scrollable, allowing a user to track messages 220 by scrolling up and down. As shown, the message chain 220 may include a multimedia message 225 that is sent via a social media platform but not by a multimedia messaging server. The multimedia message 225 includes a video portion and a subtitle portion overlaid on the video portion. The subtitle portion may display the song title, artist name, and / or lyrics. However, the subtitle portion is separable from the video. Furthermore, the multimedia message 225 does not inseparably incorporate a graphical representation of the sender's emotions. Incorporating the sender's emotions into a multimedia message is described below with reference to FIGS. 3-5.
[0038] The messaging interface 210 includes an input section 230. A user can take a photo, upload a saved image, or enter a text message in a text field. The messaging interface 210 also includes an additional input section 240 that allows a user to send various types of messages. One of the icons in the additional input section 240 is an icon for a multimedia messaging server (e.g., multimedia messaging server 135 of FIG. 1). Through this icon, a user can request multimedia content with a specific sentiment, combine layers of content to generate a flattened multimedia message with the embedded sentiment, and send the multimedia message to other users.
[0039] When the multimedia messaging server icon is selected or clicked, another window may pop up or be displayed within the messaging interface 210. For example, as shown in FIG. 3, window 300 may be displayed on the screen of the user's computing device. Window 300 may include four icons in the top region: a multimedia icon 302, a text content icon 304, a sentiment icon 306, and a combination icon 308. When an icon is selected, the selected icon may be inverted or have a graphic effect applied to it to make it stand out from the other icons. Window 300 also includes a text entry section 310 where the user enters text input.
[0040] If a user wants to send a sad multimedia message, the user selects the multimedia icon 302 and enters the user's feelings in the text input section 310. The multimedia messaging server sends the user's feelings as search terms to the audiovisual content server. The user may submit any search terms, such as "party," "Halloween," "Christmas," "graduation," "wedding," "funeral," etc. Search terms may also include lyrics, artist names, emotions, titles, and any other words or phrases. In response, the audiovisual content server sends a list of audiovisual content (e.g., songs, videos, movies, etc.) based on the search terms. If the user presses the Enter key without entering a search term in the text input section 310, the audiovisual content server may return a list of any audiovisual content.
[0041] In some aspects, the multimedia messaging server may use artificial intelligence to train a mood model so that the multimedia messaging server transmits additional information along with the search terms to the audiovisual content server. The mood model may be trained to analyze information (e.g., message exchange history) to further define the user's current mood, and the multimedia messaging server may transmit the search terms along with the current user mood to the audiovisual content server so that the user can receive a focused or more refined list of audiovisual content from the audiovisual content server. The mood model may further analyze weather conditions, the user's activity on the Internet, or other activity stored on the user's computing device to refine the user's current mood. For example, if a user searches the Internet for "Halloween" or "costume," the mood model may identify the user's mood as expectant or happy. If a user searches the Internet for terms related to "funeral," the mood model may identify the user's mood as sad or grief-stricken.
[0042] After the multimedia messaging server displays the list in window 300, the user selects one audiovisual content from the displayed list, and the selected audiovisual content is displayed in video section 312 in window 300. If the user wants a portion of the selected audiovisual content, the user specifies a period starting from a start time and ending at an end time within the selected audiovisual content. The multimedia messaging server may store the specified portion in temporary storage.
[0043] In some aspects, the multimedia messaging server may provide visual effects (hereinafter "lenses") for selected audiovisual content. For example, lenses may allow users to customize audiovisual content and save the customization settings to their collection. Lenses may be designed, preset, and stored on the multimedia messaging server. Lenses may be filters that adjust the visual attributes of the audiovisual content. Key lens adjustments include exposure, contrast, highlights, shadows, saturation, color temperature, tint, tone, and sharpness. Lenses may not be editable by the user.
[0044] After completing the selection of audiovisual content, the user selects or clicks the text content icon 304. The multimedia messaging server then automatically sends a request for text content corresponding to the selected audiovisual content to the text content server and automatically receives the corresponding text content. If a portion of the audiovisual content selected by the user is selected, the multimedia messaging server may extract the corresponding text content from the retrieved text content based on the start time and end time. The multimedia messaging server stores the text content in temporary storage and indicates to the user that the text information has been retrieved.
[0045] Returning now to window 300 of Figure 3, the user selects emotion icon 306. Window 300 shows emotion section 314, which displays graphical representations of emotions. When the user selects one of the graphical representations 322 displayed in emotion section 314, the selected graphical representation is inverted or has a graphical effect applied to it to distinguish it from unselected graphical representations.
[0046] When the user is satisfied with the selected graphical representation 322, the user clicks, drags, and drops the selected graphical representation 322 into the video section 312. The user selects or clicks the combine icon 308. The multimedia messaging server then inseparably combines the selected audiovisual content, textual content, and graphical representation 322 to generate a flattened multimedia message 330, which is then saved in the multimedia messaging server's permanent storage and / or locally on the user's computing device storage.
[0047] In some embodiments, a user can determine the location of text content in the audiovisual content when the audiovisual content and the text content are combined. The audiovisual content can be divided into three regions: top, center, and bottom. Each of the three regions can be further divided into three sub-regions: left, center, and right. Once one of the three regions or nine sub-regions is selected, text information can be incorporated into the selected region or sub-region within the flattened multimedia message.
[0048] In some aspects, the multimedia messaging server may automatically change the position of text information within the flattened multimedia message by training an artificial intelligence or machine learning algorithm. In particular, the artificial intelligence or machine learning algorithm may be trained to distinguish between foreground and background and place text content in the background so that the foreground is not obscured by the text or machine learning. Furthermore, the artificial intelligence or machine learning algorithm may change the color of the text content so that the color of the text content is substantially different from or stands out from the color of the background.
[0049] Referring to FIG. 4, a flattened multimedia message is shown with three content layers. The first layer is audiovisual content, the second layer is a graphic representation of a user's feelings or emotions, and the third layer is text content corresponding to the audiovisual content. Within the flattened multimedia message, the second layer is layered or superimposed on the first layer, and the third layer is layered or superimposed on the second layer, in turn. This means that the second layer may obscure the first layer, and the third layer may obscure the first and second layers. Because the flattened multimedia message has only one layer, the first, second, and third layers cannot be separated from the flattened multimedia message after it is flattened. In other words, the flattened multimedia message has only one layer.
[0050] A lens is not part of the three layers, but is a graphic effect applied to the first layer, the audiovisual content. For example, a lens may be a filter configured to modify the exposure, contrast, highlights, shadows, saturation, color temperature, tint, tone, and sharpness of the first layer.
[0051] When the flattened multimedia message is saved, information about the flattened multimedia message, the three layers, and the start and end times of the selected portions of audiovisual content saved may be saved in a database on the multimedia messaging server or on the user's computing device. In one aspect, all information may be saved as metadata in the header of the digital file of the flattened multimedia message.
[0052] When a user searches for audiovisual content, the multimedia messaging server, or the user's computing device, may first perform a local search within the previously stored flattened multimedia messages by comparing the search terms with a database or metadata of previously stored flattened multimedia messages. If there is no match, the multimedia messaging server sends the search terms to the audiovisual server.
[0053] Referring now to FIG. 5, a flattened multimedia message 500 according to an embodiment of the present disclosure is illustrated. The flattened multimedia message 500 incorporates a first layer 510 (i.e., a music video for Billie Eilish's "Copycat"), a second layer 520 (i.e., a graphic representation of the user's feelings or emotions), and a third layer 530 (i.e., text content or lyrics corresponding to the displayed music video). The graphic representation 520 is displayed as a colored circle, the shape of which may change. Furthermore, the graphic representation 520 may be an emoji, a GIF, an animated GIF, or any other graphic format. The second layer remains on top of the first layer for the entire playback time of the first layer. That is, when other users view any portion of the flattened multimedia message 500, they can easily understand the sending user's feelings based on the second layer 520 without guessing or assumptions.
[0054] 6 illustrates a method 600 for combining layers of content to generate a flattened multimedia message according to an aspect of the present disclosure. Method 600 begins in step 605 with a multimedia messaging server displaying a list of audiovisual content retrieved from an audiovisual content server. A user may submit a search term to the multimedia messaging server, which provides a list of audiovisual content based on the search term, which the multimedia messaging server displays on the screen of the user's computing device.
[0055] In step 610, the user reviews the displayed audiovisual content and submits a selection of one audiovisual content, and the multimedia messaging server receives the selection of the audiovisual content from the user. In some embodiments, the user may specify a start time and an end time within the selected audiovisual content to use a portion of the selected audiovisual content. The multimedia messaging server may store the selected audiovisual content or the selected portion in temporary storage.
[0056] In step 615, the user may select one or more lenses, which are filters for modifying the selected audiovisual content. In response to the selection, the multimedia messaging server applies the selected lenses to the selected audiovisual content. For example, the lenses may be filters for modifying the exposure, contrast, highlights, shadows, saturation, color temperature, tint, tone, and sharpness of the selected audiovisual content. After applying the lenses, the selected audiovisual content may become the first layer of the flattened multimedia message.
[0057] In step 620, the multimedia messaging server may search for text content corresponding to the selected audiovisual content in a text content server without receiving input from the user. In response, the multimedia messaging server receives the corresponding text content from the text content server in step 625. If the user selects a portion of the selected audiovisual content, the multimedia messaging server may extract the corresponding text content from the retrieved text content based on the selected portion.
[0058] In step 630, it is determined whether a user sentiment is to be added to the flattened multimedia message. If it is determined that a user sentiment is not to be added, in step 645, the selected audiovisual content and the obtained text content are combined to generate a flattened multimedia message.
[0059] If it is determined that a user sentiment is to be added, the multimedia messaging server receives the graphical representation of the user sentiment in step 635, and a flattened multimedia message is generated in step 640, in which the selected audiovisual content, the selected graphical representation, and the retrieved text content are sequentially combined. The flattened multimedia message has only one layer, and the selected audiovisual content, the retrieved text content, and the graphical representation are inseparable so that they cannot be extracted from the flattened multimedia message.
[0060] After generating the flattened multimedia message in steps 640 and 645, the flattened multimedia message is transmitted to other users by the multimedia messaging server. Method 600 may be repeated whenever the user wishes to send or share the flattened multimedia message with other users.
[0061] Referring now to FIG. 7 , a circular representation 700 of emotions is shown in accordance with an embodiment of the present disclosure. The circular representation 700 is borrowed from Robert Plutchik's Wheel of Emotions. There are eight primary emotion categories 710-780 in the circular representation 700. The primary emotions include, in clockwise circular order, joy, trust, fear, surprise, sadness, anticipation, anger, and disgust. Each primary emotion has an opposite primary emotion, i.e., one primary emotion is located opposite another primary emotion in the circular representation 700. For example, joy is the opposite of sadness, fear is the opposite of anger, anticipation is the opposite of surprise, and disgust is the opposite of trust. These emotion categories are provided for illustrative purposes only, and emotions may be categorized differently.
[0062] Each emotion category includes three levels. For example, emotion category 710 includes calm, joy, and ecstasy. Calm is the lowest level, joy is the primary emotion, and ecstasy is the highest level. Similarly, in category 720, acceptance is the lowest level, trust is the primary emotion, and admiration is the highest level. In category 730, anxiety is the lowest level, fear is the primary emotion, and terror is the highest level. In category 740, daze is the lowest level, surprise is the primary emotion, and amazement is the lowest level. In category 750, melancholy is the lowest level, sadness is the primary emotion, and grief is the highest level. In category 760, disgust is the lowest level, disgust is the primary emotion, and strong disgust is the highest level. In category 770, annoyance is the lowest level, anger is the primary emotion, and rage is the highest level. In category 780, interest is the lowest level, anticipation is the primary emotion, and vigilance is the highest level.
[0063] The circular representation 700 further includes combination categories 790 in which each emotion is combined with two adjacent primary emotions. For example, love is a combination of joy and trust. Submission is a combination of trust and fear. Awe is a combination of fear and surprise. Rejection is a combination of surprise and sadness. Regret is a combination of sadness and disgust. Contempt is a combination of disgust and anger. Aggression is a combination of anger and anticipation. Optimism is a combination of anticipation and joy. Thus, there are 32 emotions in the circular representation 700.
[0064] In some embodiments, colors may be assigned to primary emotions in eight categories 710-780. For example, yellow may be assigned to category 710 joy. Light green may be assigned to category 720 confidence. Blue may be assigned to category 750 sadness. Red may be assigned to category 770 anger. The level of each category may be assigned by varying the shade, brightness, shading, saturation, or tone of the color. Returning now to FIG. 5 , within the flattened multimedia message 500, the graphic representation of the emotion or sentiment in the second layer 520 may be displayed in color as described with respect to the circular representation 700. Also, as described, different emotions within the same category may be displayed in the same color with different shades, brightness, shading, saturation, or tone.
[0065] In some aspects, emotions may be represented by emojis, graphic images, gifs, animated gifs, or any other graphic representation. For example, a smiley face emoji may be used for joy, and the same color scheme may be applied to the smiley face emoji. In other words, yellow may be the color of the smiley face emoji, and calm or ecstasy may be represented by smiley face emojis with different shades, brightness, shades, saturations, or tones. The above color schemes are provided as examples, and other types of color schemes may also be used to represent different categories and levels of emotions.
[0066] Considering the circular representation of emotions 700, there are 32 emotions, and similarly, there are 32 graphical representations of the 32 emotions. That is, displaying 32 graphical representations may confuse a user in selecting one that appropriately represents the user's emotion or mood. Returning now to FIG. 3 , a graphical representation of one category having a high priority may be displayed at the top of the mood section 314, and another category having a lower priority may be displayed at the bottom of the mood section 314. The order of display of categories from top to bottom may be determined based on distance from the category with the highest priority.
[0067] 8, a multimedia messaging server 800 according to an aspect of the present disclosure is illustrated. The multimedia messaging server 800 may include a sentiment analysis module 810 for performing sentiment analysis. In particular, when the multimedia messaging server 800 receives audiovisual content from an audiovisual content server (e.g., 120 in FIG. 1 ) and obtains text content from a text content server (e.g., 130 in FIG. 1 ), the sentiment analysis module 810 performs sentiment analysis on the received audiovisual content and text content.
[0068] In some embodiments, the sentiment analysis module 810 may classify each scene of the audiovisual content into one of the nine categories 710-790 of FIG. 7 and determine which category is more dominant than the other categories. In a similar manner, the sentiment analysis module 810 may classify each term or phrase of the textual content into one of the nine categories 710-790 and determine which category is more dominant than the other categories. By comparing the two sentiment analysis results, the sentiment analysis module 810 may identify the more dominant category and assign the highest priority to the identified category. The multimedia messaging server 800 then displays the identified category in the front of a sentiment section of the user's computing device (e.g., sentiment section 314 of FIG. 3 ) so that the user can easily select an appropriate graphical representation.
[0069] Additionally, the sentiment analysis module 810 may employ artificial intelligence or machine learning algorithms when identifying a dominant category of emotion based on contextual factors along with the audiovisual content and corresponding textual content. Contextual factors such as location, temperature / weather, social context of proximity of others, social context of the message, phrases in the conversation history, frequency of the emotion, the user's affinity group, the user's regional group, and the total number of users worldwide may be used by the sentiment analysis module 810 to identify a dominant category of emotion. In particular, the contextual factors may be used to calculate a weight for each category, and the sentiment analysis module 810 uses the weighting calculation to identify the dominant category and displays the dominant category first.
[0070] Additionally, the sentiment analysis module 810 may display other, less prioritized categories based on their distance from the dominant category. For example, if category 710 in Figure 7 is identified as the dominant category, the graphical representation of category 710 may be displayed first, categories 720 and 780 may be displayed second, categories 730 and 770 may be displayed third, categories 740 and 760 may be displayed fourth, and category 750 may be displayed at the bottom. A graphical representation of combined category 790 may also be displayed based on its distance from the dominant category.
[0071] If two adjacent categories are equally dominant, the combination of the two adjacent categories may be displayed first, and other categories may be displayed based on their distance from the combination. For example, if categories 730 and 740 are equally dominant, the graphical representation of awe may be displayed first, categories 730 and 740 may be displayed second, and other categories may be displayed based on their distance from the combination.
[0072] 9 illustrates a method 900 for sequentially displaying graphic representations of emotions according to an aspect of the present disclosure. Method 900 prioritizes and displays categories of emotions to facilitate a user's selection of a graphic representation that best expresses their current state of mind. Method 900 begins in step 910 by receiving audiovisual content selected by a user. The list of audiovisual content may be received from an audiovisual content server upon providing a search term from the user. The multimedia messaging server displays the list of audiovisual content, and the user selects one from the list.
[0073] The multimedia messaging server may automatically retrieve text content corresponding to the selected audiovisual content from a text content server in step 920 .
[0074] The sentiment analysis module of the multimedia messaging server may perform sentiment analysis on the selected audiovisual content and the retrieved textual content in step 930. Based on the sentiment analysis, the sentiment analysis module may identify a dominant category of sentiment.
[0075] In step 940, the graphical representation corresponding to the dominant category may be displayed first, and the graphical representations of other categories may be displayed next based on their distance from the dominant category. In this way, relevant emotional categories may be brought to the forefront for the user.
[0076] 10 , a block diagram of a gesture recognition system 1000 for transmitting a multimedia message based on a gesture is shown, in accordance with an aspect of the present disclosure. The gesture recognition system 1000 may receive sensor data of a user, analyze the sensor data to identify a gesture, and transmit a multimedia message according to the identified gesture.
[0077] The gesture recognition system 1000 includes a data collection module 1010, a collection gesture database 1020, a gesture recognition module 1030, a learning gesture database 1040, a gesture learning module 1050, a feeling database 1060, and a multimedia message database 1070. The data collection module 1010 may accept sensors worn by a user. The sensors may include accelerometers, gyroscopes, magnetometers, radar, lidar, microphones, cameras, or other sensors. If the user is making movements on a touchscreen, the sensors may further include a touchscreen. The accelerometers and gyroscopes may generate and provide data related to acceleration, velocity, and their location. By integrating the acceleration and velocity given an initial position, the movement of a user wearing the sensors can be identified. The magnetometer may measure the direction, strength, or relative change of a magnetic field at a specific location. Radar and lidar may be used to determine distance from the user.
[0078] A single sensor may not provide enough data to track a user's movements to identify the user's gestures. However, when sensor data from one sensor is combined with sensor data from other sensors, the gesture is more likely to be accurately and reliably identified. The sensor data may be collected for a predetermined period of time. In one aspect, if there is no substantial change in the sensor data, the data collection module 1010 may ignore the sensor data. If there is a substantial change in the sensor data, the data collection module 1010 may begin collecting sensor data for a predetermined period of time.
[0079] The data collection module 1010 receives sensor data from each sensor and organizes the sensor data based on the sensor type. The organized sensor data is stored in the collected gesture database 1020.
[0080] The cleaned sensor data is then provided to the gesture recognition module 1030. Within the gesture recognition module 1030, the cleaned sensor data is preprocessed to remove noise and outliers. The sensor data may be used to identify an initial pose. For example, camera data including images captured by a camera undergoes image analysis to identify the user's initial pose. Then, sensor data from the accelerometer, gyroscope, magnetometer, radar, and lidar are integrated to track each of the user's body parts from the initial pose to generate a series of movement segments of a gesture.
[0081] The gesture learning module 1050 employs artificial intelligence or machine learning algorithms to train gesture models based on the collected sensor data and sequences of gesture segments. The artificial intelligence or machine learning algorithms may be trained by supervised, unsupervised, semi-supervised, or reinforcement learning methods, or any combination thereof.
[0082] Gestures may be identified by tracking the movements of body parts. For example, rolling dice may be identified by tracking the movements of the arms, hands, and fingers, and bowling may be identified by tracking the movements of the legs and arms. Similarly, uncorking a champagne bottle may be identified by tracking the arms, hands, and fingers. The gesture learning module 1050 analyzes the sequence of segments of a gesture to identify the body parts and track their movement and direction. Once a gesture is identified, the learning gesture database 1040 stores the gesture along with the sequence of segments.
[0083] Because the purpose of the gesture is to send a multimedia message, the identified gesture may be connected or correlated to a multimedia message that is a flattened multimedia message and that is previously stored in the multimedia message database 1070. In other words, when an identified gesture is detected or identified, the corresponding flattened multimedia message stored in the multimedia message database 1070 may be automatically sent.
[0084] If the flattened multimedia message is not correlated with a gesture, the user needs to search for and select audiovisual content from the audiovisual content server and select a graphic representation from the emotion database 1060. The identified gesture corresponds to one or more emotions or feelings stored in the feeling database 1060. Data from the microphone may be used to determine the user's feeling. Search terms corresponding to the feeling are sent to the audiovisual content server. As described above, the user selects one from a list of audiovisual content, and text content corresponding to the selected audiovisual content may be automatically retrieved from the text content server. The selected audiovisual content, the graphic representation of the user's feelings, and the text content are then combined to generate the flattened multimedia message. A correlation is made between the identified gesture and the flattened multimedia message. Based on this correlation, each time an identified gesture is detected, a corresponding flattened multimedia message is sent to other users. Such correlations may be stored in a database, such as a relational database.
[0085] Examples of predetermined gestures include uncorking a champagne bottle and rolling dice. Uncorking a champagne bottle may be correlated with a flattened multimedia message having a feeling of party or joy. Rolling dice may send any multimedia message. An example of a predetermined gesture on a touchscreen is drawing a heart, which may be correlated with a heartwarming multimedia message. Also, flipping a coin in the air may be correlated with any one of the flattened multimedia messages.
[0086] When a user wears a head-mounted display to play in VR, AR, MR, or the Metaverse, a gesture may be performed via a series of keyboard strokes or mouse movements. In this case, the gesture may be recognized via the movements shown in the VR, AR, MR, or Metaverse. In other words, the gesture may be recognized via image processing. In this example, the data collection module 1010 may acquire video from the VR, AR, MR, or Metaverse, and the gesture recognition module 1030 may perform image processing to identify body parts, detect movements or movements, and recognize the gesture.
[0087] Where a user's movements in VR, AR, MR, or the Metaverse correspond to the user's movements in the real world, gestures may be recognized based on sensor data from sensors placed on the user's body.
[0088] 11 , a method 1100 for sending a multimedia message based on a gesture is shown according to an aspect of the present disclosure. A user moves a body part, and one or more sensors on the body part generate sensor data. At step 1110, a multimedia messaging device receives the sensor data. If the user's gesture is made via a series of keyboard strokes, mouse movements, or other input device, the sensor may be a keyboard, a mouse, or any other input device, and the sensor data is data from the keyboard and mouse.
[0089] In step 1120, a gesture recognition module of the multimedia messaging device analyzes the sensor data to identify the gesture. In some embodiments, all sensor data may be jointly analyzed by an artificial intelligence or machine learning algorithm. A camera may be used to capture the initial pose, or a magnetometer, lidar, and / or radar may be used to estimate the initial pose by estimating the distance of each body part from a reference position. In the case of VR, AR, MR, or the Metaverse, the analysis may be performed on data from a keyboard, mouse, or any other input device, or, in the absence of sensor data, on video data from the VR, AR, MR, or Metaverse.
[0090] In various embodiments, the gestures may be learned through a training method described below in Figure 13. Briefly, the training method may teach a gesture recognition module of a multimedia messaging device one or more movements that identify a gesture.
[0091] Once the gesture is identified, a search is performed in a database storing pre-saved multimedia messages based on the identified gesture in step 1130. In one aspect, a relational database may be used in the search.
[0092] At step 1140, it is determined whether a search result is found in the list. If it is determined that the identified gesture is found in the relational database, at step 1170, a corresponding multimedia message may be found and selected based on the correlation.
[0093] If it is determined that the identified gesture is not found in the list, a selection of audiovisual content and a graphical representation of the user's emotion may be made, and an association between the identified gesture and both the selected audiovisual content and the graphical representation may be made in step 1150. The identified gesture is then saved as a predefined gesture in a list of predefined gestures, and the associated correspondence of the identified gesture is saved in a relational database in step 1160 for later retrieval.
[0094] After steps 1160 and 1170, the selected multimedia message is sent to other users via a messaging application in step 1180.
[0095] 12, a block diagram of a computing device 1200 is shown, which may represent any device for sending and receiving multimedia messages according to aspects of the present disclosure. The computing device 1200 may include, by way of non-limiting example, a server computer, a desktop computer, a laptop computer, a notebook computer, a subnotebook computer, a netbook computer, a netpad computer, a set-top computer, a handheld computer, an Internet appliance, a mobile smartphone, a tablet computer, a personal digital assistant, a video game console, an embedded computer, a cloud server, and the like. Those skilled in the art will recognize that many smartphones are suitable for use with the multimedia messaging system described herein. Suitable tablet computers include those in booklet, slate, and convertible configurations known to those skilled in the art.
[0096] In some embodiments, computing device 1200 includes an operating system configured to execute executable instructions. An operating system is software (including programs and data) that, for example, manages the device's hardware and provides a server for the execution of applications. Those skilled in the art will recognize that suitable server operating systems include, by way of non-limiting example, FreeBSD, OpenBSD, NetBSD®, Linux, Apple® Mac OS X Server®, Oracle® Solaris®, Windows Server®, Novell® NetWare®, iOS®, Android®, and the like. Those skilled in the art will recognize that suitable personal computer operating systems include, by way of non-limiting example, UNIX-like operating systems such as Microsoft® Windows®, Apple® Mac OS X®, UNIX®, and GNU / Linux®. In some embodiments, the operating system is provided by cloud computing. Those skilled in the art will also recognize that suitable mobile smartphone operating systems include, by way of non-limiting example, Nokia® Symbian® OS, Apple® iOS®, Research In Motion® BlackBerry® OS®, Google® Android®, Microsoft® Windows Phone® OS, Microsoft® Windows Mobile® OS, Linux®, and Palm® WebOS®.
[0097] In some embodiments, computing device 1200 may include storage 1210. Storage 1210 is one or more physical devices used to temporarily or permanently store data or programs. In some embodiments, storage 1210 is volatile memory and may require power to maintain stored information. In some embodiments, storage 1210 is nonvolatile memory and may retain stored information even when computing device 1200 is powered off. In some embodiments, nonvolatile memory includes flash memory. In some embodiments, nonvolatile memory includes dynamic random access memory (DRAM). In some embodiments, nonvolatile memory includes ferroelectric random access memory (FRAM). In some embodiments, nonvolatile memory includes phase change random access memory (PRAM). In some embodiments, storage 1210 includes, by way of non-limiting example, a CD-ROM, a DVD, a flash memory device, a magnetic disk drive, a magnetic tape drive, an optical disk drive, and cloud-based storage. In some embodiments, storage 1210 may be a combination of devices such as those disclosed herein.
[0098] Computing device 1200 further includes processor 1220, expansion unit 1230, display 1240, input device 1250, and network card 1260. Processor 1220 is the brain of computing device 1200. Processor 1220 executes instructions that implement the tasks or functions of a program. When a user runs a program, processor 1220 reads the program stored in storage 1210, loads the program into RAM, and executes the instructions defined by the program.
[0099] Processor 1220 may include a microprocessor, central processing unit (CPU), application specific integrated circuit (ASIC), math co-processor, or graphics processor, each of which is electronic circuitry within a computer that executes the instructions of a computer program by performing basic arithmetic, logic, control, and input / output (I / O) operations specified by the instructions.
[0100] In some embodiments, expansion unit 1230 may include several ports, such as one or more Universal Serial Bus (USB), IEEE 1394 ports, parallel ports, and / or expansion slots, such as Peripheral Component Interconnect (PCI) and PCI Express (PCIe). Expansion unit 1230 is not limited to this list and may include other slots or ports that can be used for any suitable purpose. Expansion unit 1230 may be used to install hardware or add additional functionality to the computer that may further the computer's purpose. For example, a USB port may be used to add additional storage to the computer, and / or an IEEE 1394 port may be used to receive video / still image data.
[0101] In some embodiments, display 1240 may be a cathode ray tube (CRT), liquid crystal display (LCD), or light emitting diode (LED). In some embodiments, display 1240 may be a thin film transistor liquid crystal display (TFT-LCD). In some embodiments, display 1240 may be an organic light emitting diode (OLED) display. In various embodiments, the OLED display is a passive matrix OLED (PMOLED) or an active matrix OLED (AMOLED) display. In some embodiments, display 1240 may be a plasma display. In some embodiments, display 1240 may be a video projector. In some embodiments, the display may be interactive (e.g., having a touch screen or sensors such as a camera, 3D sensor, LiDAR, radar) capable of detecting user interactions / gestures / responses, etc. Further, in some embodiments, display 1240 may be a combination of devices such as those disclosed herein.
[0102] A user may enter and / or modify data via input device 1250, which may include a keyboard, a mouse, a virtual keyboard, or any other device through which a user can enter data. Display 1240 displays data on the screen of display 1240. Display 1240 may be a touch screen, and display 1240 may be used as an input device.
[0103] The network card 1260 is used to communicate with other computing devices via wireless or wired connections. Multimedia messages may be exchanged or relayed between users via the network card 1260.
[0104] The computing device 1200 may further include a graphics processing unit (GPU) 1270, which generally accelerates graphics rendering. However, because the GPU 1270 can process a large amount of data simultaneously in parallel, the GPU 1270 may also be used for machine learning systems and algorithms. The GPU 1270 may cooperate with the processor 1220 for artificial intelligence to create, execute, and enhance gesture learning algorithms. In some embodiments, the GPU 1270 may include multiple GPUs to further enhance processing power.
[0105] 13-14E, an example technique for training a gesture recognition module is described. For example, FIG. 13 illustrates a flowchart of an example method 1300 for training a gesture recognition module to identify gestures, according to one or more aspects.
[0106] For example, in step 1305, a training sequence is initiated. In various embodiments, the training sequence may include training a gesture recognition module (e.g., gesture recognition module 1030) to recognize one or more movements or operations associated with a gesture. In an exemplary embodiment, a message may be sent by tapping a user's device (e.g., computing device 145) on the user's heart or other body area, a gesture that may be defined as a "heartbump gesture."
[0107] 14A-14E illustrate graphical representations of mobile device screens showing exemplary heart bump training gestures, according to one or more aspects. As shown in FIG. 14A, screen 1400A shows a list of available gestures. For purposes of explanation, the heart bump gesture is described here, but other gestures may be available in the "Available Gestures" area shown on screen 1400A. A selection button 1410 allows or disallows the heart bump gesture. In various embodiments, selection button 1410 may be a slider-type selector, where a user slides an on-screen button to turn a feature (e.g., the heart bump gesture) on (allow) or off (disallow).
[0108] A "Training Calibration" operation 1420 may be launched by clicking on the text to begin training the heart bump gesture. Thus, a user may begin training by selecting the "Training Calibration" operation 1420 (step 1305 of method 1300). Thus, once training has begun, a user may perform the training operation in step 1310.
[0109] For example, to train the device 145 to recognize a heart bump gesture, the user holds the phone up to the area of the heart and taps lightly on the chest one or more times. FIG. 14B shows an exemplary screen 1400B in which training begins and the user is instructed to tap the device (e.g., the phone) on the chest. In various embodiments, an exemplary number of times may be three. That is, the user may be asked to tap the phone on the chest three times to train the heart bump gesture. Thus, in step 1315, if the training operation has not been performed a sufficient number of times to learn the gesture, the method returns to step 1310.
[0110] Thus, after the first tap, the screen appears as screen 1400C of FIG. 14C, with the first checkbox highlighted. After the second tap, the screen appears as screen 1400D of FIG. 14D, with the second checkbox highlighted. After the third tap, the screen appears as screen 1400E of FIG. 14E, with three checkboxes highlighted, indicating that the heart bump gesture has been set (i.e., trained). Thus, method 1300 proceeds to step 1320, where gesture training is complete.
[0111] 15 illustrates a graphical representation of a mobile device screen 1500 showing an exemplary heart bump gesture collection, according to one or more aspects. Screen 1500 shows a collection addition portion 1510 and a collection portion 1520. Collection portion 1520 may, in various exemplary embodiments, include media selected for transmission to a recipient during a heart bump gesture operation, as described below.
[0112] 16 illustrates a graphical representation of a mobile device screen 1600 illustrating an exemplary heart bump gesture message sending operation, according to one or more aspects. As shown in FIG. 16, a recipient area 1610 may be included in the screen 1600, where, for example, a recipient's phone number may be entered for receiving the heart bump gesture.
[0113] 13, once device 145 has been trained for the heart bump gesture, a user may enter the recipient's phone number in recipient area 1610. When the user taps the trained device 145 on their chest, the gesture is recognized as a heart bump gesture and media from heart bump gesture collection 1620 may be selected and sent to the recipient.
[0114] Any method, program, algorithm, or code described herein may be converted into or expressed in a programming language or computer program. The terms "programming language" and "computer program," as used herein, include any language used to specify instructions to a computer, including (but not limited to) the following languages and their derivatives: Assembler, Basic, batch files, BCPL, C, C+, C++, C#, Delphi, Fortran, Java, JavaScript, machine code, operating system command languages, Pascal, Perl, PL1, scripting languages, Visual Basic, metalanguages that specify the program itself, and all first-, second-, third-, fourth-, fifth-, or later-generation computer languages. Also included are databases, other data schemas, and any other metalanguages. There is no distinction between interpreted and compiled languages. There is no distinction between compiled and source versions of a program. Thus, a reference to a program, where a programming language may exist in multiple states (e.g., source, compiled, object, or linked), is a reference to all such states. A reference to a program may encompass the actual instructions and / or the intent of those instructions.
[0115] In one or more examples, the described techniques may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored as one or more instructions or code on a computer-readable medium and executed by a hardware-based processing unit. The computer-readable medium may include non-transitory computer-readable media, corresponding to tangible media such as data storage media (e.g., RAM, ROM, EEPROM, flash memory, or any other medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer).
[0116] The instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), GPUs, or other equivalent integrated or discrete logic circuitry. Accordingly, the term "processor," as used herein, may refer to any of the foregoing structures or any other physical structure suitable for implementing the described techniques. Also, the techniques may be implemented entirely in one or more circuits or logic elements.
[0117] It should be understood that the various aspects disclosed herein may be combined in different combinations than those specifically presented in the description and accompanying drawings. It should also be understood that, in some instances, certain actions or events of any process or method described herein may be performed in a different order, added, combined, or omitted entirely (e.g., not all described actions or events are necessary to implement a technique). Furthermore, while certain aspects of the present disclosure are described for clarity as being performed by a single module or unit, it should be understood that the techniques of the present disclosure may also be implemented by a combination of units or modules associated with, for example, the servers and computing devices described above.
[0118] In various embodiments, a computing device (e.g., a mobile phone, a wearable, a television with a camera, an AR / VR input device) capable of recognizing user movement data using an accelerometer, a camera, a radar, or any combination thereof may perform the above method.
[0119] In various embodiments, the above methods may be performed by software that recognizes gestures using artificial intelligence (AI) or other methods. In various embodiments, the gestures may include explicit gestures that a user intentionally performs (e.g., gestures that mimic normal movements such as throwing a ball or a heart bump). In various embodiments, the gestures may include iconic explicit gestures (e.g., a triangle, a heart shape, etc.). In various embodiments, the gestures may include implicit gestures and movements that people normally perform, such as walking, jumping, running, dancing, etc. In various embodiments, the gestures may include user-created gestures.
[0120] In various embodiments, the system allows the user to associate gestures with a set of predefined emotional states (e.g., sad, happy, etc.) defined in the previous patent.
[0121] In various embodiments, the system allows for the association of an emotional state with a particular feel or group of feels that reflects the user's emotional state and can be suggested to the user when a gesture is triggered.
[0122] In various embodiments, the dialogue system triggers the sending of a feel to a particular person or group of people or public service when a gesture is recognized.
[0123] In various embodiments, the dialogue system measures the emotional state from previous feels sent using other methods (e.g., text) and updates the suggested feel when a gesture is triggered.
[0124] In various embodiments, the dialogue system recognizes the user's overall context (e.g., geolocation at school, in the car, at the gym, and activities such as dancing) and updates associations between gestures and emotions and feels based on the context.
[0125] In various embodiments, the dialogue system associates movements and gestures with audio input (e.g., commands, song text, etc.) to help select a feel and simplify the transmission of a feel.
[0126] In various embodiments, the multimedia messages may be selected randomly or via any desired selection algorithm.
[0127] While the above description refers to particular aspects of the present disclosure, it should be understood that many modifications may be made without departing from the spirit thereof. Additional steps and changes may be made to the sequence of the algorithm while carrying out the key teachings of the disclosure. Accordingly, the appended claims are intended to cover such modifications as fall within the true scope and spirit of the disclosure. The aspects of the present disclosure are therefore to be considered in all respects as illustrative and not restrictive, the scope of the disclosure being indicated by the appended claims, rather than the foregoing description. Unless the context dictates otherwise, any aspect disclosed herein may be combined with any other aspect or aspects disclosed herein. All changes that come within the meaning and range of equivalency of the claims are intended to be embraced therein.
Claims
1. 1. A messaging device for sending multimedia messages via gestures, comprising: a processor; a memory coupled to the processor and having instructions stored therein; Including, The instructions, when executed by the processor, cause the messaging device to: receiving sensor data from one or more sensors; analyzing the received sensor data to identify a gesture; performing a search in a list of predefined gestures based on the identified gesture; selecting a multimedia message based on the retrieved predetermined gesture; sending the selected multimedia message via a messaging application; A messaging device that performs the following.
2. 2. The messaging device of claim 1, wherein the list of predetermined gestures is pre-stored.
3. The messaging device of claim 1 further comprising a touchscreen that functions as the one or more sensors.
4. The messaging device of claim 3 , wherein the touchscreen generates the sensor data based on a user's movements on the touchscreen.
5. The messaging device of claim 1 , wherein the one or more sensors include a motion sensor.
6. The messaging device of claim 5 , wherein the sensor is worn by a user.
7. The messaging device of claim 5 , wherein the motion sensor includes at least one of a gyroscope, an accelerometer, a magnetometer, a radar, or a lidar.
8. The messaging device of claim 1 , wherein the sensor data is collected within a predetermined time period.
9. 2. The messaging device of claim 1, wherein a relational database between the list of predetermined gestures and multimedia messages is stored in the memory.
10. 10. The messaging device of claim 9, wherein the selected multimedia message is associated with the retrieved predetermined gesture based on the relational database.
11. The messaging device of claim 1 , further comprising training a first predetermined gesture in the list of predetermined gestures.
12. 12. The messaging device of claim 11, wherein training the first predetermined gesture in the list of predetermined gestures includes initially tapping the device on the user's chest.
13. 13. The messaging device of claim 12, wherein training the first predetermined gesture in the list of predetermined gestures includes one or more subsequent taps of the device on the user's chest to complete training of the first predetermined gesture.
14. tapping the device on the user's chest upon completion of training the first predetermined gesture selects a multimedia message associated with the first predetermined gesture; sending the selected multimedia message via the messaging application; 14. The messaging device of claim 13, comprising:
15. 1. A messaging method for sending multimedia messages via gestures, comprising: receiving sensor data from one or more sensors; analyzing the received sensor data to identify a gesture; performing a search in a list of predefined gestures based on the identified gesture; selecting a multimedia message based on the retrieved predetermined gesture; sending the selected multimedia message via a messaging application; A messaging method, including:
16. The messaging method of claim 15, wherein the list of predefined gestures is pre-stored.
17. The messaging method of claim 15 , wherein the one or more sensors include a touch screen.
18. 18. The messaging method of claim 17, further comprising the touch screen sensing the sensor data based on a user's movement on the touch sensor.
19. The messaging method of claim 15, wherein the sensor is a motion sensor.
20. 20. The messaging method of claim 19, wherein the motion sensor includes at least one of a gyroscope, an accelerometer, a magnetometer, a radar, or a lidar.
21. The messaging method of claim 15 , wherein the sensor data is collected within a predetermined time period.
22. 16. The messaging method of claim 15, wherein a database of relationships between the list of predefined gestures and multimedia messages is stored in a memory.
23. 23. The messaging method of claim 22, wherein the selected multimedia message is associated with the retrieved predetermined gesture based on the relational database.
24. The messaging method of claim 15 , further comprising training a first predefined gesture in the list of predefined gestures.
25. 25. The messaging method of claim 24, wherein training the first predetermined gesture in the list of predetermined gestures includes initially tapping a device on the user's chest.
26. 26. The messaging method of claim 25, wherein training the first predetermined gesture in the list of predetermined gestures includes one or more subsequent taps on the user's chest to complete training of the first predetermined gesture.
27. tapping the device against the user's chest upon completion of training the first predetermined gesture; selecting a multimedia message associated with the first predetermined gesture; sending the selected multimedia message via the messaging application; 27. The messaging method of claim 26, comprising:
28. A non-transitory computer-readable storage medium having instructions stored thereon, comprising: The instructions, when executed by a computer, cause the computer to perform a messaging method for sending a multimedia message, the messaging method comprising: receiving sensor data from one or more sensors; analyzing the received sensor data to identify a gesture; performing a search in a list of predefined gestures based on the identified gesture; selecting a multimedia message based on the retrieved predetermined gesture; sending the selected multimedia message via a messaging application; 1. A non-transitory computer-readable storage medium comprising:
Citation Information
Patent Citations
Music data generation system, music data generation server system, and music data generation method
JP2004226671A
Haptic communication system and method in portable computing devices
JP2013511897A
Method, system, and computer program for expressing emotion in dialog message by use of gestures
JP2021103520A
Music / video messaging system and method
US20110066940A1