Method and system for generating emotion based multimedia content
Patent Information
- Application Number
- KR1020200170569
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2020-12-08
- Publication Date
- 2026-08-03
- Estimated Expiration
- 2040-12-08
Smart Images

Figure 112020132989312-PAT00029_ABST
Abstract
Description
Technology Field
[0001] The present disclosure relates to a method and system for generating emotion-based multimedia content, and specifically to a method and system for generating emotion-based multimedia content capable of expressing a user's emotional state on an instant messaging application. Background Technology
[0002] Due to the proliferation of mobile devices such as smartphones and the development of the Internet, instant messaging services using mobile devices are widely used. Users of instant messaging services can naturally communicate and chat with each other in their daily lives. However, for languages with a large number of characters (e.g., Chinese), there is a problem in that it is difficult to input desired messages as text.
[0003] Meanwhile, instant messaging services provide voice messaging services that allow users to transmit their voices through chat rooms. Consequently, users can easily exchange conversations with others using voice messaging services without having to type text. However, there is a problem in that these voice messaging services merely transmit the user's recorded voice and cannot provide visual effects. The problem to be solved
[0004] The present disclosure provides a method for generating emotion-based multimedia content to solve the above-mentioned problem, a computer program stored on a recording medium, and a system (device). means of solving the problem
[0005] The present disclosure may be implemented in various ways, including a method, a system (device), or a computer program stored on a readable storage medium.
[0006] According to one embodiment of the present disclosure, a method for generating emotion-based multimedia content performed by at least one processor of a user terminal comprises the steps of receiving voice data in which a user’s voice is recorded, receiving a selection from a user for one of a plurality of characters, and transmitting to another user multimedia content generated based on the voice data, the user’s emotional state detected from the voice data, and the selected character.
[0007] According to one embodiment of the present disclosure, the generated multimedia content is transmitted to another user through a chat room on an instant messaging application.
[0008] According to one embodiment of the present disclosure, a selected character is associated with a plurality of animated graphic objects expressing different emotional states. Multimedia content includes voice data and animated graphic objects expressing the user's emotional state.
[0009] According to one embodiment of the present disclosure, the action of a selected character included in multimedia content is determined based on the emotional state of the user.
[0010] According to one embodiment of the present disclosure, voice data includes a first time interval having a signal strength below a predetermined threshold and a second time interval having a signal strength above a predetermined threshold. Multimedia content includes voice data and an animated graphic object representing the user's emotional state. The animated graphic object within the multimedia content is maintained in a static state during the first time interval and is played during the second time interval.
[0011] According to one embodiment of the present disclosure, voice data includes a first time interval associated with a first emotional state and a second time interval associated with a second emotional state. A selected character is associated with a first animated graphic object representing the first emotional state and a second animated graphic object representing the second emotional state. Multimedia content plays the first animated graphic object and voice data together during the first time interval and plays the second animated graphic object and voice data together during the second time interval.
[0012] According to one embodiment of the present disclosure, the emotional state of a user is detected based on the audio frequency characteristics of voice data.
[0013] According to one embodiment of the present disclosure, the emotional state of a user is detected based on a string detected from voice data.
[0014] According to one embodiment of the present disclosure, the method further includes the step of displaying a plurality of characters on a display. Each of the plurality of characters has an animated graphic object that expresses the emotional state of a user detected from voice data.
[0015] According to one embodiment of the present disclosure, a plurality of characters are arranged and displayed based on the user's past usage history.
[0016] According to one embodiment of the present disclosure, a plurality of characters are recommendations for characters frequently used by other users in relation to the user's emotional state.
[0017] A computer program stored on a computer-readable recording medium is provided to execute the above-described method according to one embodiment of the present disclosure on a computer.
[0018] An information processing system according to one embodiment of the present disclosure includes a communication module, a memory, and at least one processor connected to the memory and configured to execute at least one computer-readable program included in the memory. The at least one program includes instructions for receiving voice data recording the voice of a first user from a first user terminal, detecting the emotional state of the first user from the received voice data, receiving a selection of one of a plurality of characters from the first user terminal, and generating multimedia content based on the voice data, the emotional state of the first user, and the selected character.
[0019] According to one embodiment of the present disclosure, at least one program further includes instructions for transmitting multimedia content to a second user terminal associated with a second user included in a chat room on the same instant messaging application as the first user.
[0020] According to one embodiment of the present disclosure, a selected character is associated with a plurality of animated graphic objects expressing different emotional states. Multimedia content includes voice data and animated graphic objects associated with the emotional state of a first user.
[0021] According to one embodiment of the present disclosure, the action of a selected character included in multimedia content is determined based on the emotional state of a first user.
[0022] According to one embodiment of the present disclosure, voice data includes a first time interval having a signal strength below a predetermined threshold and a second time interval having a signal strength above a predetermined threshold. Multimedia content includes voice data and an animated graphic object associated with the emotional state of a first user. The animated graphic object within the multimedia content is maintained in a static state during the first time interval and is played during the second time interval.
[0023] According to one embodiment of the present disclosure, voice data includes a first time interval associated with a first emotional state and a second time interval associated with a second emotional state. A selected character is associated with a first animated graphic object representing the first emotional state and a second animated graphic object representing the second emotional state. Multimedia content plays the first animated graphic object and voice data together during the first time interval and plays the second animated graphic object and voice data together during the second time interval.
[0024] According to one embodiment of the present disclosure, detecting the emotional state of a first user from received voice data includes detecting the emotional state of the first user regardless of language and content by analyzing the audio frequency characteristics of the voice data to detect the emotional state of the first user.
[0025] According to one embodiment of the present disclosure, at least one program further includes instructions for detecting a string from voice data through voice recognition. The emotional state of the first user is detected based on the detected string. Effects of the invention
[0026] In various embodiments of the present disclosure, a user can effectively express their current emotional / mood state by transmitting multimedia content that combines voice data and a character representing their emotional state, without simply transmitting a voice message to another user.
[0027] In various embodiments of the present disclosure, emotion-based multimedia content can be generated such that a character moves according to the actual user's voice, thereby allowing the user's emotional state to be intuitively expressed to other users.
[0028] In various embodiments of the present disclosure, even when voice data of multiple emotional states is included, the user can create multimedia content that effectively reflects their emotional state and transmit the created multimedia content to other users.
[0029] The effects of the present disclosure are not limited to those mentioned above, and other unmentioned effects will be clearly understood by a person skilled in the art to which the present disclosure pertains from the description in the claims (“person skilled in the art”). Brief explanation of the drawing
[0030] Embodiments of the present disclosure will be described with reference to the accompanying drawings described below, wherein similar reference numerals indicate similar elements, but are not limited thereto. FIG. 1 is a diagram illustrating an example in which emotion-based multimedia content is provided through an instant messaging application operating on a user terminal according to one embodiment of the present disclosure. FIG. 2 is a schematic diagram showing a configuration in which an information processing system is connected to communicate with a plurality of user terminals to provide an emotion-based multimedia content creation service according to one embodiment of the present disclosure. FIG. 3 is a block diagram showing the internal configuration of a user terminal and an information processing system according to one embodiment of the present disclosure. FIG. 4 is a drawing showing an example of multimedia content being generated according to one embodiment of the present disclosure. FIG. 5 is an exemplary drawing showing animated graphic objects associated with various emotional states included in a character according to one embodiment of the present disclosure. FIG. 6 is an exemplary drawing illustrating the operation of a graphic object according to a section of voice data according to one embodiment of the present disclosure. FIG. 7 is an exemplary drawing illustrating the operation of an animated graphic object according to a section of voice data including two emotional states, according to one embodiment of the present disclosure. FIG. 8 is an exemplary drawing showing a chat room interface on an instant messaging application in which multimedia content is transmitted according to one embodiment of the present disclosure. FIG. 9 is a flowchart illustrating a method for transmitting multimedia content according to one embodiment of the present disclosure. FIG. 10 is a flowchart illustrating a method for creating multimedia content according to one embodiment of the present disclosure. Specific details for implementing the invention
[0031] Hereinafter, specific details for implementing the present disclosure will be described in detail with reference to the attached drawings. However, in the following description, specific descriptions regarding widely known functions or configurations will be omitted if there is a risk that the gist of the present disclosure may be unnecessarily obscured.
[0032] In the attached drawings, identical or corresponding components are assigned the same reference numerals. Additionally, in the description of the following embodiments, the description of identical or corresponding components may be omitted. However, even if a description of a component is omitted, it is not intended that such component is not included in any embodiment.
[0033] The advantages and features of the disclosed embodiments and the methods for achieving them will become clear by referring to the embodiments described below in conjunction with the accompanying drawings. However, the present disclosure is not limited to the embodiments disclosed below but may be implemented in various different forms, and the embodiments provided are merely to make the present disclosure complete and to fully inform those skilled in the art of the scope of the invention.
[0034] The terms used in this specification will be briefly explained, and the disclosed embodiments will be described in detail. The terms used in this specification have been selected to be as generally used as possible, taking into account their functions in this disclosure; however, these terms may vary depending on the intent of those skilled in the art, case law, the emergence of new technologies, etc. Additionally, in specific cases, terms may be arbitrarily selected by the applicant, and in such cases, their meanings will be described in detail in the relevant description of the invention. Therefore, the terms used in this disclosure should be defined not merely by their names, but based on their meanings and the content throughout this disclosure.
[0035] In this specification, singular expressions include plural expressions unless the context clearly specifies them as singular. Additionally, plural expressions include singular expressions unless the context clearly specifies them as plural. Throughout the specification, when a part is described as including a certain component, this means that, unless specifically stated otherwise, it does not exclude other components but may include additional components.
[0036] Additionally, the terms 'module' or 'part' as used in the specification refer to software or hardware components, and the 'module' or 'part' performs certain roles. However, the meaning of 'module' or 'part' is not limited to software or hardware. The 'module' or 'part' may be configured to reside in an addressable storage medium or configured to run on one or more processors. Thus, as an example, the 'module' or 'part' may include components such as software components, object-oriented software components, class components, and task components, and at least one of processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, or variables. The components and the functions provided within the 'module' or 'part' may be combined into a smaller number of components and 'modules' or 'parts', or further separated into additional components and 'modules' or 'parts'.
[0037] According to one embodiment of the present disclosure, a ‘module’ or ‘part’ may be implemented as a processor and memory. The term ‘processor’ should be broadly interpreted to include a general-purpose processor, a central processing unit (CPU), a microprocessor, a digital signal processor (DSP), a controller, a microcontroller, a state machine, etc. In some environments, the term ‘processor’ may refer to an application-specific integrated circuit (ASIC), a programmable logic device (PLD), a field programmable gate array (FPGA), etc. The term ‘processor’ may also refer to a combination of processing devices, such as, for example, a combination of a DSP and a microprocessor, a combination of multiple microprocessors, a combination of one or more microprocessors combined with a DSP core, or any other combination of such configurations. Additionally, the term ‘memory’ should be broadly interpreted to include any electronic component capable of storing electronic information. 'Memory' may refer to various types of processor-readable media, such as Random Access Memory (RAM), Read-Only Memory (ROM), Non-Volatile Random Access Memory (NVRAM), Programmable Read-Only Memory (PROM), Erasable-Programmable Read-Only Memory (EPROM), Electrically Erasable PROM (EEPROM), Flash Memory, Magnetic or Optical Data Storage Devices, Registers, etc. If a processor can read information from memory and / or write information to memory, the memory is said to be in an electronic communication state with the processor. Memory integrated into a processor is in an electronic communication state with the processor.
[0038] In the present disclosure, a "chat room" may refer to a virtual space or group in which one or more users (or user accounts) can participate, which can be created in an instant messaging application installed on a computing device. For example, one or more user accounts may participate in or be included in a chat room to exchange various types of messages, files, etc. with one another. Additionally, the chat room may be provided with VoIP (Voice over Internet Protocol) voice call functions, VoIP video call functions, live broadcasting functions (VoIP real-time video transmission functions), and multimedia content creation functions, so that voice calls, video calls, video streaming, and multimedia content transmission between user accounts can be performed.
[0039] In the present disclosure, 'user' may refer to a user utilizing an instant messaging application or a user account of an instant messaging application. Here, a user account may represent an account created and utilized by a user within the instant messaging application or data related thereto.
[0040] FIG. 1 is a diagram illustrating an example in which emotion-based multimedia content (132) is provided through an instant messaging application operating on a user terminal (120) according to one embodiment of the present disclosure. A user (110) can exchange messages with other users who are subscribed to the instant messaging application using the user terminal (120). In one embodiment, the user (110) can exchange text messages, voice messages, video messages, multimedia content (132), etc., with other users through the instant messaging application.
[0041] A user (110) can send messages or data to another user through a chat room interface (130). In one embodiment, the user (110) can send voice data containing the user's voice (112) to another user through the chat room interface (130). For example, the user (110) can record the user's voice (112) by selecting a recording button provided on an instant messaging application, such as by touch input, and send the recorded voice data to another user. In this case, the user (110) can send the recorded voice data to another user by combining a character (or sticker, emoji, etc.) with the recorded voice data.
[0042] A user (110) can transmit multimedia content (132), created by combining the user's voice (112) and a character, to another user through a chat room interface (130). Here, the character is used to visually convey the user's (110) emotions or feelings and may include emoticons, emojis, stickers, images, videos, etc. For example, to create multimedia content (132), the user (110) may select a character provided by default in the application or purchase and use a new character from a store, etc. In this case, there may be a dedicated character for creating multimedia content (132), but is not limited thereto, and existing emoticons, etc. may also be used to create multimedia content (132).
[0043] A user (110) can transmit emotion-based multimedia content (132) that represents their emotions to another user. In this case, to generate emotion-based multimedia content (132), the emotional state of the user (110) can be extracted / detected from voice data in which the user's voice (112) is recorded. For example, the emotional state of the user (110) can be detected based on the audio frequency characteristics of the voice data. In another example, the emotional state of the user (110) can be detected based on a string detected from the voice data. In this case, the multimedia content (132) can be generated by combining voice data with a motion character that can represent the detected emotional state of the user (110), and can be transmitted to another user.
[0044] In FIG. 1, it is illustrated that a user (110) uses a user terminal (120) to record the user's voice (112) in real time, but this is not limited thereto. For example, voice data stored in advance in the user terminal (120) may be used to create multimedia content (132), or voice data received from another computing device may be used to create multimedia content (132). With such a configuration, the user (110) can effectively express their current emotional / mood state by transmitting multimedia content (132) combined with voice data and a character representing their emotional state, rather than simply transmitting a voice message to another user.
[0045] FIG. 2 is a schematic diagram showing a configuration in which an information processing system (230) is connected to communicate with a plurality of user terminals (210_1, 210_2, 210_3) to provide an emotion-based multimedia content creation service according to one embodiment of the present disclosure. The information processing system (230) may include system(s) capable of providing an instant messaging service including an emotion-based multimedia content creation service through a network (220). In one embodiment, the information processing system (230) may include one or more server devices and / or databases capable of storing, providing, and executing computer-executable programs (e.g., downloadable applications) and data related to the instant messaging service and the creation of emotion-based multimedia content, or one or more distributed computing devices and / or distributed databases based on a cloud computing service. For example, the information processing system (230) may include separate systems (e.g., servers) for providing an emotion-based multimedia content creation service and / or an instant messaging service.
[0046] The instant messaging service provided by the information processing system (230) can be provided to the user through an instant messaging application installed on each of the multiple user terminals (210_1, 210_2, 210_3). For example, the instant messaging service may include a text messaging service between users of the instant messaging application, a voice messaging service, a video call service, a voice call service, a video streaming service, an emotion-based multimedia content creation / provision service, etc.
[0047] Multiple user terminals (210_1, 210_2, 210_3) can communicate with an information processing system (230) through a network (220). The network (220) can be configured to enable communication between the multiple user terminals (210_1, 210_2, 210_3) and the information processing system (230). Depending on the installation environment, the network (220) may be configured as a wired network such as Ethernet, Power Line Communication, telephone line communication devices and RS-serial communication, a mobile communication network, a Wireless LAN (WLAN), Wi-Fi, Bluetooth and ZigBee, or a combination thereof. The communication method is not limited and may include not only communication methods utilizing communication networks that the network (220) may include (e.g., mobile communication network, wired internet, wireless internet, broadcasting network, satellite network, etc.) but also short-range wireless communication between user terminals (210_1, 210_2, 210_3).
[0048] In FIG. 2, a mobile phone terminal (210_1), a tablet terminal (210_2), and a PC terminal (210_3) are illustrated as examples of user terminals, but are not limited thereto. The user terminals (210_1, 210_2, 210_3) may be any computing device capable of wired and / or wireless communication and capable of installing and running instant messaging applications, etc. For example, user terminals may include smartphones, mobile phones, navigation systems, computers, laptops, digital broadcasting terminals, PDAs (Personal Digital Assistants), PMPs (Portable Multimedia Players), tablet PCs, game consoles, wearable devices, IoT (Internet of Things) devices, VR (Virtual Reality) devices, AR (Augmented Reality) devices, etc. Additionally, FIG. 2 illustrates three user terminals (210_1, 210_2, 210_3) communicating with an information processing system (230) through a network (220), but is not limited thereto, and may be configured so that a different number of user terminals communicate with an information processing system (230) through a network (220).
[0049] In one embodiment, the information processing system (230) may receive voice data containing a user's voice recorded from a user terminal (210_1, 210_2, 210_3). Additionally, the information processing system (230) may receive a selection of one of a plurality of characters from the user terminal (210_1, 210_2, 210_3). In this case, the information processing system (230) may receive voice data and a selection of characters through an instant messaging application installed on the user terminal (210_1, 210_2, 210_3). Afterward, the information processing system (230) may detect the user's emotional state from the received voice data, generate multimedia content based on the voice data, the user's emotional state, and the selected character, and provide the generated multimedia content to other users.
[0050] FIG. 3 is a block diagram showing the internal configuration of a user terminal (210) and an information processing system (230) according to one embodiment of the present disclosure. The user terminal (210) may refer to any computing device capable of running instant messaging applications, etc., and capable of wired / wireless communication, and may include, for example, the mobile phone terminal (210_1), tablet terminal (210_2), PC terminal (210_3) of FIG. 2. As illustrated, the user terminal (210) may include a memory (312), a processor (314), a communication module (316), and an input / output interface (318). Similarly, the information processing system (230) may include a memory (332), a processor (334), a communication module (336), and an input / output interface (338). As illustrated in FIG. 3, the user terminal (210) and the information processing system (230) may be configured to communicate information and / or data through the network (220) using their respective communication modules (316, 336). Additionally, the input / output device (320) may be configured to input information and / or data to the user terminal (210) or output information and / or data generated from the user terminal (210) through the input / output interface (318).
[0051] The memory (312, 332) may include any non-transient computer-readable recording medium. According to one embodiment, the memory (312, 332) may include a permanent mass storage device such as random access memory (RAM), read-only memory (ROM), disk drive, solid state drive (SSD), flash memory, etc. As another example, a permanent mass storage device such as ROM, SSD, flash memory, disk drive, etc. may be included in the user terminal (210) or information processing system (230) as a separate permanent storage device distinct from the memory. Additionally, the memory (312, 332) may store an operating system and at least one program code (e.g., code for an instant messaging application installed and running on the user terminal (210)).
[0052] These software components may be loaded from a computer-readable recording medium separate from memory (312, 332). This separate computer-readable recording medium may include a recording medium that can be directly connected to the user terminal (210) and the information processing system (230), for example, a computer-readable recording medium such as a floppy drive, disk, tape, DVD / CD-ROM drive, or memory card. As another example, the software components may be loaded into memory (312, 332) via a communication module (316, 336) rather than a computer-readable recording medium. For example, at least one program may be loaded into memory (312, 332) based on a computer program (e.g., an application providing an instant messaging service or an emotion-based multimedia content creation / provision service) that is installed by files provided through a network (220) by developers or a file distribution system distributing installation files of the application.
[0053] The processor (314, 334) may be configured to process instructions of a computer program by performing basic arithmetic, logic, and input / output operations. Instructions may be provided to the processor (314, 334) by memory (312, 332) or a communication module (316, 336). For example, the processor (314, 334) may be configured to execute instructions received according to program code stored in a recording device such as memory (312, 332).
[0054] The communication module (316, 336) may provide a configuration or function for the user terminal (210) and the information processing system (230) to communicate with each other via the network (220), and may provide a configuration or function for the user terminal (210) and / or the information processing system (230) to communicate with another user terminal or another system (e.g., a separate cloud system). For example, a request or data (e.g., voice data recording the user's voice, selection of one of a plurality of characters, etc.) generated by the processor (314) of the user terminal (210) according to program code stored in a recording device such as memory (312) may be transmitted to the information processing system (230) via the network (220) under the control of the communication module (316). Conversely, control signals or commands provided under the control of the processor (334) of the information processing system (230) may be received by the user terminal (210) through the communication module (336) and the network (220) via the communication module (316) of the user terminal (210). For example, the user terminal (210) may receive multimedia content generated based on the user's emotional state and selected character from the information processing system (230).
[0055] The input / output interface (318) may be a means for interfacing with an input / output device (320). As an example, the input device may include a device such as a camera including an audio sensor and / or an image sensor, a keyboard, a microphone, a mouse, etc., and the output device may include a device such as a display, a speaker, a haptic feedback device, etc. As another example, the input / output interface (318) may be a means for interfacing with a device in which the configuration or function for performing input and output is integrated into one, such as a touchscreen, etc.
[0056] In FIG. 3, the input / output device (320) is depicted as not being included in the user terminal (210), but is not limited thereto and may be configured as a single device with the user terminal (210). Additionally, the input / output interface (338) of the information processing system (230) may be a means for interfacing with a device (not shown) for input or output that may be connected to the information processing system (230) or included in the information processing system (230). In FIG. 3, the input / output interface (318, 338) is depicted as an element configured separately from the processor (314, 334), but is not limited thereto and may be configured so that the input / output interface (318, 338) is included in the processor (314, 334).
[0057] The user terminal (210) and the information processing system (230) may include more components than those of FIG. 3. However, it is not necessary to clearly illustrate most of the prior art components. In one embodiment, the user terminal (210) may be implemented to include at least some of the input / output devices (320) described above. Additionally, the user terminal (210) may further include other components such as a transceiver, a GPS (Global Positioning System) module, a camera, various sensors, a database, etc. For example, if the user terminal (210) is a smartphone, it may include components that are generally included in a smartphone, and may be implemented to include various components such as an accelerometer, a gyroscope, a microphone module, a camera module, various physical buttons, buttons using a touch panel, input / output ports, and a vibrator for vibration.
[0058] According to one embodiment, the processor (314) of the user terminal (210) may be configured to operate an instant messaging application or a web browser application that provides an instant messaging service including an emotion-based multimedia content creation service. At this time, program code associated with the application may be loaded into the memory (312) of the user terminal (210). While the application is running, the processor (314) of the user terminal (210) may receive information and / or data provided from an input / output device (320) through an input / output interface (318) or receive information and / or data from an information processing system (230) through a communication module (316), and may process the received information and / or data and store it in the memory (312). Additionally, such information and / or data may be provided to the information processing system (230) through the communication module (316).
[0059] While the instant messaging application is in operation, the processor (314) may receive voice data, text, images, videos, etc., that are input or selected through an input device such as a touch screen, keyboard, audio sensor and / or image sensor, camera, microphone, etc., connected to an input / output interface (318), and may store the received voice data, text, images and / or videos, etc. in memory (312) or provide them to an information processing system (230) through a communication module (316) and a network (220). In one embodiment, the processor (314) may receive voice data recording the user's voice and the user's selection of one of a plurality of characters through an input device, and provide the corresponding data / request to an information processing system (230) through a network (220) and a communication module (316).
[0060] The processor (334) of the information processing system (230) may be configured to manage, process, and / or store information and / or data received from a plurality of user terminals and / or a plurality of external systems. In one embodiment, the processor (334) may store, process, and transmit voice data, selection information regarding a character, etc., received from the user terminal (210). For example, the processor (334) may detect the user's emotional state from the received voice data and generate multimedia content based on the voice data, the user's emotional state, and the selected character. The multimedia content generated in this way may be provided to other users through an instant messaging application, etc.
[0061] FIG. 4 is a diagram illustrating an example of multimedia content being generated according to an embodiment of the present disclosure. A user can transmit emotion-based multimedia content to another user through a chat room on an instant messaging application. In one embodiment, the user can transmit emotion-based multimedia content expressing their own emotions through a first operation (410), a second operation (420), and a third operation (430).
[0062] The first operation (410) indicates that a recording icon (412) capable of recording the user's voice is displayed on a display (e.g., a chat room interface). In one embodiment, the user can receive an interface for recording the user's voice by selecting the recording icon (412) displayed on the display via touch input or the like. Although the recording icon (412) is shown as being displayed to the right of the message input window in FIG. 4, it is not limited thereto, and the recording icon (412) may be displayed at any location on the chat room interface.
[0063] The second operation (420) indicates that, in response to the user selecting the recording icon (412), a plurality of characters and a recording button (424) are displayed on the display. Here, each of the plurality of characters may be a sticker from one of the sticker sets. In another example, each of the plurality of characters may be an image and may be associated with a specific theme. In one embodiment, all characters owned by the user may be displayed on the display. Alternatively, among the characters owned by the user, characters having animated graphic objects that express the user's emotional state detected from the user's voice data may be displayed on the display.
[0064] The user can perform voice recording by selecting the recording button (424) via touch input or the like. Additionally, the user can determine a character for creating multimedia content by selecting one of a plurality of characters displayed on the display via touch input or the like. For example, the user can determine the “Bunny” character (422) as the character for creating multimedia content by selecting the “Bunny” character (422) displayed on the display via touch input or the like. In this case, the user can perform voice recording after selecting the character first, or select the character after completing the voice recording.
[0065] In one embodiment, a plurality of characters displayed on the display may be arranged based on the user's past usage history. For example, a plurality of characters may be arranged and displayed in the order of the user's most recent use. In another example, a plurality of characters may be arranged and displayed in order of the user's past usage frequency. In yet another example, a plurality of characters may be arranged and displayed by comprehensively considering the user's past usage frequency, past usage times, etc. In yet another example, if the user performs a voice recording before selecting a character, the characters may be arranged and displayed in the order of the user's recent use in relation to the emotional state detected in the voice recording.
[0066] In another embodiment, a plurality of characters may be recommendations for characters frequently used by other users in relation to the user's emotional state. That is, among the characters owned by the user, characters frequently used by other users of the instant messaging application may be displayed on the display. For example, a plurality of characters may be displayed in order based on the number of past uses, past usage times, etc., of other users. In another example, if the user performs a voice recording before selecting a character, the characters may be displayed in order of frequent use by other users in relation to the emotional state detected in the voice recording.
[0067] The third operation (430) indicates the process of performing a voice recording of the user when the user selects the recording button (424) via touch input or the like. When the user performs a voice recording, the audio frequency of the voice, the recording time, etc., may be displayed in real time. After the voice recording is completed, the user may complete the voice recording by selecting the recording completion button (432) displayed on the display via touch input or the like. When the recording completion button (432) is selected by the user, the voice data containing the user's recorded voice may be transmitted to a system (information processing system) for creating emotion-based multimedia content. In this case, information about the character selected by the user may also be transmitted to the system.
[0068] In FIG. 4, four characters are shown displayed on the display, but this is not limited thereto, and any number of characters may be displayed on the display. Also, in FIG. 4, the input for selecting a character and the input for selecting the recording button (424) are described as being distinct, but this is not limited thereto, and when the user selects one character, voice recording may start automatically. Conversely, when the user selects the voice recording button (424) or the recording complete button (432), the character most preferred by the user or other users for the detected emotional state may be automatically selected.
[0069] FIG. 5 is an exemplary drawing illustrating animated graphic objects (510, 520, 530, 540) associated with various emotional states included in a character (500) according to one embodiment of the present disclosure. The character (500) may be associated with a plurality of animated graphic objects (510, 520, 530, 540) expressing different emotional states. As illustrated, the character (500) may include a graphic object (510) expressing Sadness, a graphic object (520) expressing Happiness, a graphic object (530) expressing Anger, a graphic object (540) expressing Neutral, etc. Additionally, the character (500) may further include a graphic object expressing Fear, a graphic object expressing Contempt, a graphic object expressing Disgust, a graphic object expressing Surprise, etc.
[0070] An animated graphic object (510, 520, 530, 540) may include multiple preset images or videos representing the movement of the character (500). For example, an animated graphic object (510) expressing sadness may include five preset images (510_1, 510_2, 510_3, 510_4, 510_5) in which the appearance of the character (500) (e.g., direction of gaze, shape of mouth, body movement, etc.) is slightly changed. In one embodiment, when multimedia content is created using a graphic object (510) that expresses sadness, the multimedia content may be configured such that preset images (510_1, 510_2, 510_3, 510_4, 510_5) included in the graphic object (510) are played repeatedly at predetermined time intervals (e.g., 0.1 seconds), or may be configured to be played at time intervals determined by the user's voice.
[0071] In FIG. 5, a graphic object (510) is depicted as including five preset images (510_1, 510_2, 510_3, 510_4, 510_5), but is not limited thereto, and each graphic object may include a different number of preset images. Additionally, in FIG. 5, the order in which preset images are played is described as being predetermined, but is not limited thereto, and after determining the user's mouth shape based on the user's voice data, preset images may be played to display a mouth shape similar to the determined mouth shape.
[0072] FIG. 6 is an exemplary diagram illustrating the operation of a graphic object according to a section of voice data (610) according to an embodiment of the present disclosure. A user can transmit multimedia content generated based on voice data, the user's emotional state detected from the voice data, and a selected character to another user. For example, multimedia content may be generated by combining the user's voice and an animated graphic object of a character associated with the user's emotional state. In this case, the animated graphic object may include a plurality of preset images or videos.
[0073] In one embodiment, voice data (610) that records the user's voice has a section having a signal strength below a predetermined threshold ( , ) and a section having a signal strength greater than a predetermined threshold ( It may include ). That is, voice data (610) may be divided into sections containing the user's voice at a volume greater than a certain level and sections not containing the user's voice. Here, a predetermined threshold value is a criterion for determining whether the user's voice is included, and may be applied equally to all users using the emotion-based multimedia content creation service or applied differently to each user.
[0074] Sections having signal strength below a predetermined threshold ( , ) may correspond, for example, to the interval before the user presses the record button and performs voice recording, the interval before the user completes the voice recording and presses the recording complete button, the interval during which the user does not speak while recording, etc. In addition, a interval having a signal strength greater than or equal to a predetermined threshold ( ) may be, for example, a section containing the user's voice.
[0075] In one embodiment, a section having a signal strength below a predetermined threshold ( , In ), animated graphic objects included in multimedia content may be maintained in a static state. For example, in a section having a signal strength below a predetermined threshold ( , In ), one preset image (620_1) included in the animated graphic object can be continuously displayed. That is, in the section of voice data (610) ( , Determining that the user's voice is not included in ), the section of the multimedia content ( , One preset image (620_1) included in the animated graphic object can be continuously displayed during the animation.
[0076] In one embodiment, a section having a signal strength greater than or equal to a predetermined threshold ( In ), animated graphic objects included in multimedia content can be played. For example, the section ( In ), multiple preset images (620_1, 620_2, 620_3, 620_4, 620_5) included in the object can be repeatedly displayed at predetermined time intervals. That is, a section ( of voice data (610) Determining that the user's voice is included in ), the segment of the multimedia content ( Animated graphic objects can be played during )
[0077] In FIG. 6, the voice data (610) has a signal strength greater than one predetermined threshold ( ) and a section having signal strengths below two predetermined thresholds ( , It has been described as including, but is not limited to. For example, voice data may include intervals having a signal strength greater than or equal to an arbitrary number of predetermined thresholds and intervals having a signal strength less than an arbitrary number of predetermined thresholds. With such a configuration, emotion-based multimedia content can be generated so that a character moves according to the actual user's voice, thereby allowing the user's emotional state to be intuitively expressed to other users.
[0078] FIG. 7 is an exemplary diagram illustrating the operation of animated graphic objects (710_1, 720_1) according to intervals of voice data (700) including two emotional states, in accordance with one embodiment of the present disclosure. A user can transmit multimedia content generated based on voice data (700), the user's emotional state detected from the voice data, and a selected character to another user. In this case, the user can transmit voice data (700) containing the user's voice and selection information regarding the character to an information processing system for generating / providing multimedia content in order to generate multimedia content.
[0079] The information processing system can detect the user's emotional state from the received voice data (700). In one embodiment, the information processing system can detect the user's emotional state regardless of language and content by analyzing the audio frequency characteristics of the voice data (700) to detect the user's emotional state. Additionally or alternatively, the information processing system can detect strings from the voice data (700) through voice recognition and detect the user's emotional state based on the detected strings. That is, the information processing system can convert the voice contained in the voice data (700) into text using voice recognition technology (e.g., STT (Speech-To-Text) technology, etc.). Then, the information processing system can detect words and / or strings, etc., representing the user's emotions from the converted text and detect the user's emotional state based on the detected words and / or strings.
[0080] In one embodiment, two or more emotional states may be detected for each time interval of a single voice data (700). In the illustrated example, the voice data (700) may include two different emotional states of a user and may include a time interval associated with each emotional state. For example, the voice data (700) may include a time interval associated with a neutral emotional state ( Time intervals associated with the emotional states of ) and Happiness ( It may include ).
[0081] The information processing system can generate multimedia content by combining an animated graphic object of a character selected by a user with voice data (700). In one embodiment, the voice data (700) is a segment associated with two emotional states ( If ) is included, the information processing system can generate multimedia content so that an animated graphic object associated with the emotional state detected in each interval is played. For example, the information processing system may include a time interval ( During ) an animated graphic object (710_1) expressing a neutral emotional state is played, and a time interval ( Multimedia content can be created so that an animated graphic object (720_1) expressing an emotional state of happiness is played during ). In this case, the multimedia content is a time interval ( During ) the animated graphic object (710_1) and voice data (700) are played together, and the time interval ( During this time, animated graphic objects (720_1) and voice data (700) can be played together.
[0082] In FIG. 7, voice data (700) is depicted as including segments associated with two emotional states, but is not limited thereto and may include segments associated with three or more emotional states. Additionally, in FIG. 7, one segment associated with one emotional state is depicted thereto, but is not limited thereto and may include two or more segments separated from each other associated with one emotional state. With such a configuration, even when voices of multiple emotional states are included in one voice data (700), the user can create multimedia content that effectively reflects their emotional state and transmit the created multimedia content to other users.
[0083] FIG. 8 is an exemplary drawing illustrating a chat room interface (800) on an instant messaging application in which multimedia content (810, 820) is transmitted according to one embodiment of the present disclosure. As illustrated, a user can transmit multimedia content (810, 820) generated based on voice data recorded from a user's voice, an emotional state detected from the voice data, and a selected character to another user. For example, the generated multimedia content (810, 820) can be transmitted to another user through a chat room on an instant messaging application.
[0084] In one embodiment, users included in the chat room and other users can select and play multimedia content (810, 820) transmitted through the chat room using touch input or the like. For example, when another user included in the chat room selects multimedia content (810, 820), voice recorded by the user who transmitted the multimedia content (810, 820) is output, and a character expressing the user's emotional state can move on the display. The multimedia content (810, 820) can be created using different characters and can be used to visually express different emotional states of the user.
[0085] In one embodiment, the user can select a share button (812, 822) via touch input or the like to share multimedia content (810, 820) sent to another user through a chat room with another chat room / user on an instant messaging application or to share it with another application such as a text application. Another user who has received multimedia content (810, 820) from the user can also share the multimedia content (810, 820) with other users by using a similar share button.
[0086] In FIG. 8, it is described that when users included in a chat room select multimedia content (810, 820) by touch input or the like, the multimedia content (810, 820) is played, but this is not limited thereto. For example, when a user of the chat room enters the chat room, the multimedia content (810, 820) may be played automatically. In another example, when a user of the chat room moves to the location of previously transmitted multimedia content (810, 820) by scroll input or the like, the multimedia content (810, 820) may be played automatically.
[0087] FIG. 9 is a flowchart illustrating a multimedia content transmission method (900) according to one embodiment of the present disclosure. The multimedia content transmission method (900) may be performed by a user terminal (e.g., at least one processor of the user terminal). The multimedia content transmission method (900) may be initiated by the processor receiving voice data in which the user's voice is recorded (S910). For example, the processor may receive voice data through an instant messaging application installed on the user terminal.
[0088] Additionally, the processor may display multiple characters on the display. Then, the processor may receive a selection from the user for one of the multiple characters (S920). Here, the character may be associated with multiple animated graphic objects representing different emotional states. In one embodiment, the multiple characters may each have an animated graphic object representing the user's emotional state detected from voice data. For example, the multiple characters may be displayed in order based on the user's past usage history. Additionally, or alternatively, the multiple characters may be recommendations for characters frequently used by other users in relation to the user's emotional state.
[0089] The processor can transmit multimedia content generated based on voice data, the user's emotional state detected from the voice data, and a selected character to another user (S930). For example, the multimedia content can be transmitted to another user through a chat room on an instant messaging application. Here, the multimedia content may include voice data and an animated graphic object representing the user's emotional state. In this case, the action of the selected character included in the multimedia content can be determined based on the user's emotional state.
[0090] FIG. 10 is a flowchart illustrating a method for generating multimedia content (1000) according to an embodiment of the present disclosure. The method for generating multimedia content (1000) may be performed by an information processing system (e.g., at least one processor of the information processing system). The method for generating multimedia content (1000) may be initiated by the processor receiving voice data in which the voice of a first user is recorded from a first user terminal (S1010). In this case, the processor may detect the emotional state of the first user from the received voice data (S1020). For example, the processor may detect the emotional state of the first user regardless of language and content by analyzing the audio frequency characteristics of the voice data to detect the emotional state of the first user. Additionally or alternatively, the processor may detect a string from the voice data through voice recognition and detect the emotional state of the first user based on the detected string.
[0091] Additionally, the processor may receive a selection of one of a plurality of characters from a first user terminal (S1030). For example, the character may be associated with a plurality of animated graphic objects expressing different emotional states. Then, the processor may generate multimedia content based on voice data, the first user's emotional state, and the selected character (S1040). Here, the multimedia content may include voice data and animated graphic objects expressing the first user's emotional state. For example, the processor may determine the action of the selected character included in the multimedia content based on the user's emotional state. Then, the processor may transmit the generated multimedia content to a second user terminal associated with a second user who is in the same chat room as the first user (S1050).
[0092] The method described above may be provided as a computer program stored on a computer-readable recording medium for execution on a computer. The medium may continuously store a program executable by a computer, or temporarily store it for execution or download. Additionally, the medium may be various recording or storage means in the form of a single or multiple hardware components, and may not be limited to a medium directly connected to a computer system but may exist distributed over a network. Examples of media may include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and media configured to store program instructions, including ROM, RAM, and flash memory. Furthermore, other examples of media may include recording or storage media managed by app stores that distribute applications or sites and servers that supply or distribute various other software.
[0093] The methods, operations, or techniques of the present disclosure may be implemented by various means. For example, these techniques may be implemented in hardware, firmware, software, or a combination thereof. Those skilled in the art will understand that the various exemplary logical blocks, modules, circuits, and algorithmic steps described in connection with the disclosure herein may be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate such interchangeability between hardware and software, various exemplary components, blocks, modules, circuits, and steps have been generally described above in terms of their functional aspects. Whether such functions are implemented in hardware or in software depends on the design requirements imposed on the specific application and the overall system. Those skilled in the art may implement the functions described in various ways for each specific application, but such implementations should not be construed as departing from the scope of the present disclosure.
[0094] In a hardware implementation, the processing units used to perform the techniques may be implemented in one or more ASICs, DSPs, digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, electronic devices, other electronic units designed to perform the functions described in this disclosure, computers, or a combination thereof.
[0095] Accordingly, the various exemplary logic blocks, modules, and circuits described in connection with the present disclosure may be implemented or performed by any combination of general-purpose processors, DSPs, ASICs, FPGAs or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or those designed to perform the functions described herein. A general-purpose processor may be a microprocessor, but alternatively, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented by a combination of computing devices, for example, a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors coupled with a DSP core, or any other combination of configurations.
[0096] In firmware and / or software implementations, techniques may be implemented as instructions stored on a computer-readable medium such as random access memory (RAM), read-only memory (ROM), non-volatile random access memory (NVRAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable PROM (EEPROM), flash memory, compact disc (CD), magnetic or optical data storage devices, etc. The instructions may be executable by one or more processors, and the processor(s) may be enabled to perform specific aspects of the functions described in this disclosure.
[0097] Although the embodiments described above have been described as utilizing aspects of the subject matter disclosed herein in one or more standalone computer systems, the present disclosure is not limited thereto and may be implemented in conjunction with any computing environment, such as a network or a distributed computing environment. Furthermore, aspects of the subject matter in the present disclosure may be implemented in a plurality of processing chips or devices, and storage may be similarly affected across a plurality of devices. Such devices may include PCs, network servers, and portable devices.
[0098] Although the present disclosure has been described in relation to some embodiments, various modifications and changes may be made without departing from the scope of the present disclosure as understood by a person skilled in the art to which the invention of the present disclosure pertains. Furthermore, such modifications and changes should be considered to fall within the scope of the claims appended to this specification. Explanation of the symbols
[0099] 110: User 120: User terminal 130: Chat room interface 132: Multimedia content
Claims
Claim 1 A method for generating emotion-based multimedia content performed by at least one processor of a user terminal, comprising: receiving voice data in which a user’s voice is recorded; detecting the user’s emotional state based on the voice data; receiving a selection of one of a plurality of characters from the user; and transmitting to another user multimedia content including the voice data and an animated graphic object of the selected character representing the user’s emotional state, wherein the animated graphic object includes a plurality of preset images, the voice data includes a first time interval having a signal strength below a predetermined threshold and a second time interval having a signal strength above the predetermined threshold, the first time interval includes a period during which the user did not speak during recording, and the animated graphic object of the selected character within the multimedia content is maintained in a static state by continuously displaying one of the plurality of preset images during the first time interval, and is played by changing and displaying the plurality of preset images at predetermined time intervals during the second time interval. Claim 2 A method for generating emotion-based multimedia content according to claim 1, wherein the generated multimedia content is transmitted to the other user through a chat room on an instant messaging application. Claim 3 A method for generating emotion-based multimedia content according to claim 1, wherein the selected character is associated with a plurality of animated graphic objects expressing different emotional states. Claim 4 A method for generating emotion-based multimedia content according to claim 1, wherein the action of the selected character included in the multimedia content is determined based on the emotional state of the user. Claim 5 delete Claim 6 A method for generating emotion-based multimedia content according to claim 1, wherein the voice data includes a third time interval associated with a first emotional state and a fourth time interval associated with a second emotional state, the selected character is associated with a first animated graphic object representing the first emotional state and a second animated graphic object representing the second emotional state, the multimedia content plays the first animated graphic object and the voice data together during the third time interval and plays the second animated graphic object and the voice data together during the fourth time interval, the first emotional state and the second emotional state are different from each other, and the first animated graphic object representing the first emotional state and the second animated graphic object representing the second emotional state are different from each other. Claim 7 A method for generating emotion-based multimedia content according to claim 1, wherein the emotional state of the user is detected based on the audio frequency characteristics of the voice data. Claim 8 A method for generating emotion-based multimedia content according to claim 1, wherein the emotional state of the user is detected based on a string detected from the voice data. Claim 9 A method for generating emotion-based multimedia content according to claim 1, further comprising the step of displaying the plurality of characters on a display, wherein each of the plurality of characters has an animated graphic object expressing the emotional state of the user detected from the voice data. Claim 10 In claim 9, a method for generating emotion-based multimedia content, wherein the plurality of characters are arranged and displayed based on the user's past usage history. Claim 11 In claim 9, the plurality of characters are recommenders for characters frequently used by other users in relation to the user's emotional state, a method for generating emotion-based multimedia content. Claim 12 A computer program stored on a computer-readable recording medium for executing a method according to any one of paragraphs 1 through 4 and paragraphs 6 through 11 on a computer. Claim 13 As an information processing system, communication module; memory; and includes at least one processor connected to the memory and configured to execute at least one computer-readable program included in the memory, wherein the at least one program includes instructions for receiving voice data recording the voice of a first user from a first user terminal, detecting the emotional state of the first user from the received voice data, receiving a selection of one of a plurality of characters from the first user terminal, and generating multimedia content including the voice data and an animated graphic object of the selected character expressing the emotional state of the first user, wherein the animated graphic object includes a plurality of preset images, wherein the voice data includes a first time interval having a signal strength below a predetermined threshold and a second time interval having a signal strength above the predetermined threshold, wherein the first time interval includes a period during which the user did not speak during recording, wherein the multimedia content includes the voice data and an animated graphic object associated with the emotional state of the first user, wherein the animated graphic object within the multimedia content is maintained in a stationary state by continuously displaying one of the plurality of preset images during the first time interval, and wherein the plurality of preset images during the second time interval a predetermined time An information processing system that is played by changing and displaying at intervals. Claim 14 In paragraph 13, the information processing system further comprises at least one program for transmitting the multimedia content to a second user terminal associated with a second user included in a chat room on the same instant messaging application as the first user. Claim 15 In claim 13, the selected character is associated with a plurality of animated graphic objects expressing different emotional states, and the multimedia content includes animated graphic objects associated with the voice data and the emotional state of the first user, an information processing system. Claim 16 In paragraph 13, an information processing system in which the action of the selected character included in the multimedia content is determined based on the emotional state of the first user. Claim 17 delete Claim 18 In claim 13, the voice data includes a third time interval associated with a first emotional state and a fourth time interval associated with a second emotional state, the selected character is associated with a first animated graphic object representing the first emotional state and a second animated graphic object representing the second emotional state, the multimedia content plays the first animated graphic object and the voice data together during the third time interval and plays the second animated graphic object and the voice data together during the fourth time interval, the first emotional state and the second emotional state are different from each other, and the first animated graphic object representing the first emotional state and the second animated graphic object representing the second emotional state are different from each other, an information processing system. Claim 19 An information processing system according to claim 13, wherein detecting the emotional state of the first user from the received voice data includes detecting the emotional state of the first user regardless of language and content by analyzing the audio frequency characteristics of the voice data. Claim 20 In paragraph 13, the above-mentioned at least one program further comprises instructions for detecting a string from the voice data through voice recognition, and the emotional state of the first user is detected based on the detected string, an information processing system.