Generation system and generation method

The generation system adjusts speech bubble shapes and sound effects based on emotional estimation to enhance the expressiveness of chat services, addressing the lack of emotional flexibility in existing avatar systems.

JP2025167110APending Publication Date: 2025-11-07NTT DOCOMO INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024071418
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-04-25
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Chat services using avatars lack the flexibility to reflect users' emotional expressions, with uniform speech bubble expressions failing to convey emotional nuances.

Method used

A generation system that includes a reception unit for user requests, an image generation unit, a sound effect generation unit, a feeling estimation unit, and a shape determination unit to dynamically adjust speech bubble shapes based on estimated emotions, synchronizing image and sound effects with emotional expressions.

Benefits of technology

Enables output that flexibly reflects users' emotional expressions through dynamically shaped speech bubbles and matching sound effects, enhancing the emotional expressiveness of chat services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025167110000001_ABST
    Figure 2025167110000001_ABST
Patent Text Reader

Abstract

To create output for chat service with user's feeling expression flexibly reflected.SOLUTION: A processing device 20 comprises: a reception unit 21 which accepts a generation request including request details related to image generation from a user; an image generation unit 22 which generates an image on the basis of the generation request; an effect sound generation unit 23 which generates an effect sound based upon the image generated by the image generation unit 22; a feeling estimation unit 24 which estimates a feeling based upon the effect sound generated by the effect sound generation unit or the request details; a shape determination unit 25 which determines a shape of a speech bubble displayed in the circumference of the image generated by the image generation unit 22 based upon a result of the feeling estimation by the feeling estimation unit 24; and an output unit 26 which displays a display image including the speech bubble and the image and also outputs the effect sound in time with the display of the display image.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a production system and a production method. [Background technology]

[0002] Patent Document 1 discloses a technique for changing the shape of a message balloon in a chat service using an avatar after a predetermined time has elapsed. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2004-206544 Summary of the Invention [Problem to be solved by the invention]

[0004] In chat services using avatars, expressions such as speech bubbles are uniform, and it is not possible to flexibly reflect the user's emotional expressions.

[0005] The present disclosure has been made in consideration of the above-mentioned circumstances, and aims to provide a generation system and generation method that can produce output in a chat service that flexibly reflects a user's emotional expression. [Means for solving the problem]

[0006] A generation system according to one aspect of the present disclosure includes: a reception unit that receives a generation request from a user, the generation request including requested content related to image generation; an image generation unit that generates an image based on the generation request; a sound effect generation unit that generates a sound effect based on the image generated by the image generation unit; a feeling estimation unit that performs feeling estimation based on the sound effect generated by the sound effect generation unit or the requested content; a shape determination unit that determines the shape of speech bubbles to be displayed around the image generated by the image generation unit based on a result of the feeling estimation by the feeling estimation unit; and an output unit that displays a display image including the speech bubbles and the image, and outputs sound effects in synchronization with the timing at which the display image is displayed.

[0007] In a generation system according to the present disclosure, a sound effect is generated based on an image generated in response to a user request, emotion estimation is performed based on the sound effect or the user's request, and the shape of a speech bubble is determined based on the emotion estimation result. Then, the speech bubble, image, and sound effect are output together. With this configuration, the shape of the speech bubble is determined from an emotion that is appropriately estimated taking into account the user's request, so that the shape of the speech bubble can reflect the user's emotion. Then, an image is displayed inside such a speech bubble, and a sound effect that matches the image is output, thereby enabling an output that flexibly reflects the user's emotional expression. [Effects of the Invention]

[0008] According to the present disclosure, in a chat service, it is possible to provide output that flexibly reflects the emotional expression of the user. [Brief explanation of the drawings]

[0009] [Figure 1] FIG. 1 is a diagram illustrating an overview of an avatar dialogue system according to this embodiment. [Figure 2] FIG. 2 is a diagram showing the device configuration of the avatar dialogue system. [Figure 3] FIG. 3 is a diagram showing an example of output. [Figure 4] FIG. 4 is a diagram showing an example of output. [Figure 5] FIG. 5 is a flowchart showing the processing executed by the avatar dialogue system. [Figure 6] FIG. 6 is a diagram illustrating an example of a hardware configuration of the avatar dialogue system. DETAILED DESCRIPTION OF THE INVENTION

[0010] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings. In the description of the drawings, the same or similar elements are designated by the same reference numerals, and redundant description will be omitted.

[0011] FIG. 1 is a diagram illustrating an overview of the avatar dialogue system according to this embodiment. The avatar dialogue system is a system that uses an avatar to engage in dialogue with a user. An avatar is a character or icon that moves in a digital space. In this system, the avatar may be a character that represents the user himself or herself, or a character that represents a person with whom the user is having a dialogue. In the example shown in FIGS. 1(a) and 1(b), the avatar responds to a user's inquiry.

[0012] An avatar dialogue system may be used, for example, in an avatar customer service system that aims to support users in both work and everyday life using AI that can understand the human mind and converse appropriately with literacy, the situation, and the topic. An avatar dialogue system may be built, for example, by combining CX analysis technology, non-verbal information recognition technology, and responses using large-scale language models (LLMs). CX analysis technology, for example, predicts a user's current state and future behavior based on user attribute or behavioral data, and provides responses tailored to the user. Non-verbal information recognition technology, for example, detects user anxieties from the user's speech and video, and provides responses that are sensitive to the user's feelings. Responses using LLMs, for example, use language models to expand the range of available dialogues and respond to a wide range of questions, including casual conversations and general questions.

[0013] Here, the expressions of the avatars need to be made in real time in accordance with the input from the user, so they tend to be simple and lack expressiveness.

[0014] The avatar dialogue system according to this embodiment realizes a variety of expressions for avatars, and specifically, provides output that flexibly reflects the user's emotional expression. For example, in the example shown in FIG. 1(a), in response to a user's request to see a cute cat, an image of a cat is displayed and a cat's meowing sound is output as audio. Furthermore, the user's emotion is estimated based on the user's request (here, "I want to see a cute cat"), and a speech bubble (avatar speech bubble) shaped in accordance with the estimated emotion is displayed. Here, a speech bubble shaped in a soft image is displayed in response to the request "I want to see a cute cat." Furthermore, in the example shown in FIG. 1(b), in response to a user's request to see a cool lion, an image of a lion is displayed and a lion's roar is output as audio. Furthermore, in response to the user's request (here, "I want to see a cool lion"), the user's emotion is estimated based on the user's request (here, "I want to see a cool lion"), and a speech bubble (avatar speech bubble) shaped in accordance with the estimated emotion is displayed. Here, a speech bubble shaped in a violent image is displayed in response to the request "I want to see a cool lion." In this way, a speech bubble shape that reflects the user's emotions is adopted, an image is displayed inside the speech bubble, and sound effects that match the image are output, thereby realizing a variety of outputs that flexibly reflect the user's emotional expression.

[0015] 2 is a diagram showing the device configuration of the avatar dialogue system. The avatar dialogue system includes a terminal 10, a processing device 20 (generation system), an image generation server 30, a sound effect generation server 40, and a shape determination server 50, which are configured to be able to communicate with each other via a network including a wireless communication network and a fixed communication network. Note that the configuration of the avatar dialogue system is not limited to the above, and for example, all processing may be performed by the terminal 10, or all processing may be performed by the terminal 10 and the processing device 20, or the system may be configured to further include other devices.

[0016] The terminal 10 is a device used by a user who utilizes the avatar dialogue system. The terminal 10 is, for example, a personal computer, a smartphone, a tablet terminal, a feature phone, a server device, a game console, or the like. Note that while only one terminal 10 is illustrated in FIG. 2, the avatar dialogue system may include any number of terminals 10, two or more.

[0017] The processing device 20 is a device that receives a generation request from a user via the terminal 10, generates images, sound effects, and speech bubbles in cooperation with each server, and transmits (outputs) the generated information to the terminal 10. The processing device 20 is configured to include, as functional components, a reception unit 21, an image generation unit 22, a sound effect generation unit 23, a feeling estimation unit 24, a shape determination unit 25, and an output unit 26.

[0018] The reception unit 21 receives a generation request from a user via the terminal 10, the generation request including a request for image generation. The request for image generation is information related to an image desired by the user, such as, for example, "I want to see a cool lion." The reception unit 21 may receive a generation request in the form of at least one of voice, text, and image. That is, the reception unit 21 may receive a generation request in the form of, for example, a user's speech (voice), characters (text) input by the user, or an image input by the user. The generation request received by the reception unit 21 may be output by the output unit 26 and displayed on the terminal 10. In this case, if the generation request is voice, a text transcribed from the voice may be displayed on the terminal 10; if the generation request is text, the text may be displayed; and if the generation request is an image, the image may be displayed. The reception unit 21 outputs the received generation request to the image generation unit 22.

[0019] The image generation unit 22 generates an image based on the generation request. The image generation unit 22 may generate an image based on the request content included in the generation request and user information. Here, the user information may be, for example, at least one of the user's location information, the user's service usage history, the user's attribute information, and the user's hobbies and preferences. Each piece of user information may be acquired by the terminal 10 and periodically transmitted to the processing device 20, or may be transmitted from the terminal 10 to the processing device 20 as information included in the generation request, or may be acquired from another server or the like and stored in the processing device 20. The user's location information is, for example, information indicating the user's latitude and longitude. The user's service usage history is information about services used by the user and may include online purchase history and website browsing history. The user's attribute information is information such as the user's gender, age, family composition, occupation, educational background, and life stage. The user's hobbies and preferences are information indicating the user's interests, the user's favorite color, the user's current emotions, the user's personality information, the user's behavioral history, and the like.

[0020] The image generation unit 22 generates a prompt including, for example, the request content included in the generation request and each user information, and transmits the prompt to the image generation server 30. The prompt expresses, in text, for example, the command to be executed by the generative AI model, the task to be executed by the generative AI model, the background and context (e.g., role, condition) to be considered by the generative AI model, the question to be answered by the generative AI model, and the output format of the response information from the generative AI model. The prompt may also include input information that is the target of the command and task to be executed by the generative AI model. Examples of such input information include data files with file names that include a predetermined extension, such as text data, image data, application-related data, audio data, video data, and still image data. Application-related data includes document data, table data, graph data, and other data that can be processed by a predefined application program.

[0021] The image generation server 30 stores a generative AI model that generates an image based on a prompt sent from the image generation unit 22, for example. The generative AI model may be an interactive AI that includes, for example, a large language model (LLM) and a user interface (UI) for dialogue with a user, enabling text-based or voice-based chat with the user. Examples of such generative AI models include ChatGPT, GPT (registered trademark)-3.5, GPT-4V, and PaLM2. The image generation server 30 transmits the generated image to the image generation unit 22. The image generation unit 22 outputs the image to the sound effect generation unit 23. Note that the image generation unit 22 may generate images by itself by having the functions of the image generation server 30.

[0022] The sound effect generation unit 23 generates sound effects based on the image generated by the image generation unit 22. The sound effect generation unit 23 generates music (sound effects) that evokes the image. The sound effect generation unit 23 may, for example, generate a corpus of words that describe the characteristics of the image from the image and transmit the corpus of words to the sound effect generation server 40. In this case, the sound effect generation server 40 stores, for example, a machine learning model that generates sound effects from a corpus of words, and generates sound effects that match the characteristics of the image by inputting the corpus of words into the machine learning model. The sound effect generation server 40 transmits the sound effects to the sound effect generation unit 23. The sound effect generation unit 23 outputs the sound effects to the emotion estimation unit 24. Note that the sound effect generation unit 23 may generate sound effects by itself by having the functions of the sound effect generation server 40.

[0023] The emotion estimation unit 24 estimates the emotion of the user based on the sound effects generated by the sound effect generation unit 23. The emotion estimation unit 24 may estimate the emotion of the user from the request content included in the generation request accepted by the acceptance unit 21. The emotion estimation unit 24 outputs the result of the estimation of the emotion of the user to the shape determination unit 25.

[0024] The shape determination unit 25 determines the shape of a speech bubble to be displayed around the image generated by the image generation unit 22 based on the result of emotion estimation by the emotion estimation unit 24. The shape determination unit 25 may determine the shape of the speech bubble based on both the result of emotion estimation and the image. The shape determination unit 25 may, for example, transmit the result of emotion estimation (and the image) to the shape determination server 50. In this case, the shape determination server 50 stores a machine learning model using, for example, a convolutional neural network, and inputs the image related to the result of emotion estimation (and the image generated by the image generation unit 22) to the machine learning model, thereby selecting the shape of the speech bubble that best matches the input content. Note that information obtained by visualizing a sound effect using a mel spectrogram may also be input to the machine learning model of the shape determination server 50. The shape determination server 50 transmits the selected shape of the speech bubble to the shape determination unit 25. The shape determination unit 25 determines the shape of the speech bubble based on the received information.

[0025] The shape determination unit 25 may dynamically change the shape of the speech bubble in response to a change in the result of emotion estimation by the emotion estimation unit 24. That is, the shape determination unit 25 may dynamically change the shape of the speech bubble by changing the information to be sent to the shape determination server 50 each time the result of emotion estimation changes. The shape determination unit 25 may further dynamically change the shape of the speech bubble in consideration of the content of previous and subsequent dialogues in the avatar dialogue system. The shape determination unit 25 outputs information indicating the determined shape of the speech bubble to the output unit 26.

[0026] The output unit 26 displays a display image including a speech bubble and an image on the terminal 10, and outputs sound effects to the terminal 10 in synchronization with the timing at which the display image is displayed. The output unit 26 displays a display image in which an image is displayed along the frame of the speech bubble (at the maximum size that fits within the frame). The output unit 26 displays the display image so that it does not overlap with avatars or other objects. The output unit 26 may display the display image including a speech bubble and an image as a 2D image, as shown in FIG. 3. The output unit 26 may display the display image including a speech bubble and an image as a 3D image, as shown in FIG. 4.

[0027] Next, the processing executed by the avatar dialogue system will be described with reference to Fig. 5. Fig. 5 is a flowchart showing the processing executed by the processing device 20 of the avatar dialogue system.

[0028] As shown in FIG. 5, first, the processing device 20 receives a generation request from the user via the terminal 10 (step S1).

[0029] Next, the processing device 20 generates an image based on the generation request (step S2).

[0030] Next, the processing device 20 generates sound effects based on the generated image (step S3).

[0031] Next, the processing device 20 estimates emotions based on the generated sound effects (step S4).

[0032] Next, the processing device 20 determines the shape of the speech bubble to be displayed around the generated image based on the result of emotion estimation (step S5).

[0033] Finally, the processing device 20 outputs (displays) the display image including the speech bubble and the image to the terminal 10, and also outputs a sound effect to the terminal 10 in synchronization with the timing at which the display image is displayed (step S6).

[0034] Next, the effects of the processing device 20 according to this embodiment will be described.

[0035] The processing device 20 according to this embodiment includes: a reception unit 21 that receives a generation request including a request related to image generation from a user; an image generation unit 22 that generates an image based on the generation request; a sound effect generation unit 23 that generates a sound effect based on the image generated by the image generation unit 22; a feeling estimation unit 24 that performs feeling estimation based on the sound effect generated by the sound effect generation unit or the request content; a shape determination unit 25 that determines the shape of speech bubbles to be displayed around the image generated by the image generation unit 22 based on the result of feeling estimation by the feeling estimation unit 24; and an output unit 26 that displays a display image including the speech bubbles and the image and outputs sound effects in synchronization with the timing at which the display image is displayed.

[0036] In the processing device 20 according to this embodiment, sound effects are generated based on images generated in response to a user request, emotion estimation is performed based on the sound effects or the user's request, and the shape of the speech bubble is determined based on the emotion estimation result. Then, the speech bubble, image, and sound effect are output together. With this configuration, the shape of the speech bubble is determined based on an emotion that is appropriately estimated taking into account the user's request, so that the shape of the speech bubble can reflect the user's emotion. Then, an image is displayed inside such a speech bubble, and sound effects that match the image are output, thereby enabling output that flexibly reflects the user's emotional expression.

[0037] The receiving unit 21 may receive a generation request in the form of at least one of a voice, a text, and an image. With this configuration, a generation request including the user's request content can be received easily and appropriately.

[0038] The image generation unit 22 may generate an image based on the request content included in the generation request and at least one of the user's location information, service usage history, attribute information, and hobbies and preferences. This allows the image generation unit 22 to generate an image that better meets the user's needs by taking into account the user's request content as well as information related to the user.

[0039] The shape determination unit 25 may determine the shape of the speech bubble based on the emotion estimation result and the image. By taking the image into consideration in addition to the emotion estimation result, the shape of the speech bubble can be determined to better reflect the user's request.

[0040] The shape determination unit 25 may dynamically change the shape of the speech bubble in response to changes in the result of emotion estimation by the emotion estimation unit 24. This allows the shape of the speech bubble to be changed in accordance with changes in the user's emotion, and enables output that more flexibly reflects the user's emotional expression.

[0041] The output unit 26 may display a display image in which an image is displayed along the frame of the speech bubble. This improves the consistency between the speech bubble and the image, and allows the display to be natural to the user.

[0042] The output unit 26 may display the display image as a 3D image, thereby enabling a more diverse image display.

[0043] The generation system and generation method of the present disclosure have the following configuration.

[0044] [1] a receiving unit that receives a generation request including a request for image generation from a user; an image generation unit that generates an image based on the generation request; a sound effect generation unit that generates a sound effect based on the image generated by the image generation unit; an emotion estimation unit that performs emotion estimation based on the sound effect generated by the sound effect generation unit or the request content; a shape determination unit that determines a shape of a speech bubble to be displayed around the image generated by the image generation unit based on a result of emotion estimation by the emotion estimation unit; an output unit that displays a display image including the speech bubble and the image, and outputs the sound effect in synchronization with the timing at which the display image is displayed.

[0045] [2] The generation system according to [1], wherein the reception unit receives the generation request by at least one of voice, text, and image.

[0046] [3] The generation system according to [1] or [2], wherein the image generation unit generates the image based on the request content included in the generation request and at least one of the user's location information, service usage history, attribute information, and hobbies and preferences.

[0047] [4] The generation system according to any one of [1] to [3], wherein the shape determination unit determines a shape of the speech bubble based on the result of the emotion estimation and the image.

[0048] [5] The generation system according to any one of [1] to [4], wherein the shape determination unit dynamically changes the shape of the speech bubble in response to changes in a result of emotion estimation by the emotion estimation unit.

[0049] [6] The generation system according to any one of [1] to [5], wherein the output unit displays the display image in which the image is displayed along a frame of the speech bubble.

[0050] [7] The generation system according to any one of [1] to [6], wherein the output unit displays the display image as a 3D image.

[0051] [8] A production method executed by a production system, comprising: receiving a generation request from a user, the generation request including a request for image generation; generating an image based on the generation request; generating a sound effect based on the generated image; Estimating an emotion based on the generated sound effect or the request content; determining a shape of a speech bubble to be displayed around the image based on a result of emotion estimation; displaying a display image including the speech bubble and the image, and outputting the sound effect in synchronization with the timing at which the display image is displayed.

[0052] The block diagrams used to explain the above embodiments show functional blocks. These functional blocks (components) are realized by any combination of hardware and / or software. Furthermore, the method for realizing each functional block is not particularly limited. That is, each functional block may be realized using a single device that is physically or logically coupled, or may be realized using two or more physically or logically separated devices that are connected directly or indirectly (for example, by wire, wirelessly, etc.) and these multiple devices. The functional block may also be realized by combining the single device or multiple devices with software.

[0053] Functions include, but are not limited to, judgment, determination, judgment, calculation, computation, processing, derivation, investigation, search, confirmation, reception, transmission, output, access, resolution, selection, election, establishment, comparison, assumption, expectation, consideration, broadcasting, notifying, communicating, forwarding, configuring, reconfiguring, allocation, mapping, and assignment. For example, a functional block (component) that performs transmission is called a transmitting unit or transmitter. As mentioned above, there are no particular limitations on how these functions are implemented.

[0054] For example, the processing device 20 constituting the avatar dialogue system according to an embodiment of the present disclosure may function as a computer that performs processing of the control method of the present disclosure. FIG. 6 is a diagram illustrating an example of the hardware configuration of the processing device 20 according to this embodiment. The processing device 20 described above may be physically configured as a computer including a processor 1001, a memory 1002, a storage 1003, a communication device 1004, an input device 1005, an output device 1006, a bus 1007, and the like. Note that the processing device 20 may be configured as a computer including at least one processor such as a CPU or a GPU, or may be configured as a computer including multiple processors or may be configured to include multiple computer devices. The terminal 10 may also have a similar hardware configuration.

[0055] In the following description, the term "apparatus" can be interpreted as a circuit, a device, a unit, etc. The hardware configuration of the processing device 20 may be configured to include one or more of the apparatuses shown in the drawings, or may be configured to exclude some of the apparatuses.

[0056] Each function in the processing device 20 is realized by loading predetermined software (programs) onto hardware such as the processor 1001 and memory 1002, causing the processor 1001 to perform calculations, control communication via the communication device 1004, and control at least one of reading and writing data in the memory 1002 and storage 1003.

[0057] The processor 1001 controls the entire computer by running, for example, an operating system. The processor 1001 may be configured by a central processing unit (CPU) including an interface with peripheral devices, a control device, an arithmetic unit, a register, etc. For example, the above-mentioned reception unit 21, image generation unit 22, sound effect generation unit 23, emotion estimation unit 24, shape determination unit 25, output unit 26, etc. may be realized by the processor 1001.

[0058] The processor 1001 also reads programs (program codes), software modules, data, etc. from at least one of the storage 1003 and the communication device 1004 into the memory 1002 and executes various processes in accordance with the programs. The programs used are those that cause a computer to execute at least some of the operations described in the above-described embodiments. For example, the reception unit 21, the image generation unit 22, the sound effect generation unit 23, the emotion estimation unit 24, the shape determination unit 25, and the output unit 26 may be implemented by a control program stored in the memory 1002 and running on the processor 1001, and similarly may be implemented for other functional blocks. While the above-described various processes have been described as being executed by one processor 1001, they may also be executed simultaneously or sequentially by two or more processors 1001. The processor 1001 may be implemented by one or more chips. The programs may also be transmitted from a network via a telecommunications line.

[0059] The memory 1002 is a computer-readable recording medium and may be configured, for example, by at least one of a read-only memory (ROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), a random access memory (RAM), etc. The memory 1002 may also be called a register, a cache, a main memory (primary storage device), etc. The memory 1002 can store executable programs (program codes), software modules, etc. for implementing a control method according to an embodiment of the present disclosure.

[0060] Storage 1003 is a computer-readable recording medium, and may be, for example, at least one of an optical disk such as a CD-ROM (Compact Disc ROM), a hard disk drive, a flexible disk, a magneto-optical disk (e.g., a compact disk, a digital versatile disk, a Blu-ray disc), a smart card, a flash memory (e.g., a card, a stick, a key drive), a floppy disk, a magnetic strip, etc. Storage 1003 may also be referred to as an auxiliary storage device. The above-mentioned storage medium may be, for example, a database, a server, or other appropriate medium including at least one of memory 1002 and storage 1003.

[0061] The communication device 1004 is hardware (transmission / reception device) for communicating between computers via at least one of a wired network and a wireless network, and is also referred to as, for example, a network device, a network controller, a network card, or a communication module. The communication device 1004 may be configured to include a high-frequency switch, a duplexer, a filter, a frequency synthesizer, etc. to realize at least one of frequency division duplex (FDD) and time division duplex (TDD). For example, the above-mentioned reception unit 21 and the like may be realized by the communication device 1004.

[0062] The input device 1005 is an input device (for example, a keyboard, a mouse, a microphone, a switch, a button, a sensor, etc.) that receives input from the outside. The output device 1006 is an output device (for example, a display, a speaker, an LED lamp, etc.) that outputs to the outside. The input device 1005 and the output device 1006 may be integrated into one device (for example, a touch panel).

[0063] Furthermore, each device, such as the processor 1001 and the memory 1002, is connected by a bus 1007 for communicating information. The bus 1007 may be configured using a single bus, or may be configured using different buses between each device.

[0064] Furthermore, the processing device 20 may be configured to include hardware such as a microprocessor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a programmable logic device (PLD), or a field programmable gate array (FPGA), and some or all of the functional blocks may be realized by the hardware. For example, the processor 1001 may be implemented using at least one of these pieces of hardware.

[0065] The notification of information is not limited to the aspects / embodiments described in the present disclosure and may be performed using other methods. For example, the notification of information may be performed by physical layer signaling (e.g., Downlink Control Information (DCI), Uplink Control Information (UCI)), higher layer signaling (e.g., Radio Resource Control (RRC) signaling, Medium Access Control (MAC) signaling, broadcast information (Master Information Block (MIB), System Information Block (SIB))), other signals, or a combination thereof. Furthermore, the RRC signaling may be referred to as an RRC message, and may be, for example, an RRC Connection Setup message, an RRC Connection Reconfiguration message, or the like.

[0066] The order of the procedures, sequences, flowcharts, etc. of each aspect / embodiment described in this disclosure may be changed unless it is consistent. For example, the methods described in this disclosure present elements of various steps using an example order, and are not limited to the particular order presented.

[0067] Input and output information may be stored in a specific location (for example, memory) or may be managed using a management table. Input and output information may be overwritten, updated, or added to. Output information may be deleted. Input information may be sent to another device.

[0068] The determination may be made based on a value represented by one bit (0 or 1), a Boolean value (true or false), or a numerical comparison (e.g., comparison with a predetermined value).

[0069] Each aspect / embodiment described in this disclosure may be used alone, in combination, or switched depending on the implementation. Furthermore, notification of predetermined information (e.g., notification that "X is true") is not limited to being done explicitly, but may be done implicitly (e.g., by not notifying the predetermined information).

[0070] Although the present disclosure has been described in detail above, it is clear to those skilled in the art that the present disclosure is not limited to the embodiments described herein. The present disclosure can be implemented in modified and altered forms without departing from the spirit and scope of the present disclosure as defined by the claims. Therefore, the description of the present disclosure is intended to be illustrative and does not have any limiting meaning on the present disclosure.

[0071] Software shall be construed broadly to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, threads of execution, procedures, functions, etc., whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise.

[0072] Software, instructions, information, etc. may also be transmitted or received over a transmission medium. For example, if software is transmitted from a website, server, or other remote source using wired technologies (such as coaxial cable, fiber optic cable, twisted pair, Digital Subscriber Line (DSL)), and / or wireless technologies (such as infrared, microwave), these wired and / or wireless technologies are included within the definition of transmission media.

[0073] The information, signals, etc. described in this disclosure may be represented using any of a variety of different technologies. For example, data, instructions, commands, information, signals, bits, symbols, chips, etc. that may be referred to throughout the above description may be represented by voltages, currents, electromagnetic waves, magnetic fields or magnetic particles, optical fields or photons, or any combination thereof.

[0074] Note that terms explained in this disclosure and terms necessary for understanding this disclosure may be replaced with terms having the same or similar meanings. For example, at least one of a channel and a symbol may be a signal (signaling). Furthermore, a signal may be a message. Furthermore, a component carrier (CC) may be called a carrier frequency, a cell, a frequency carrier, etc.

[0075] Furthermore, the information, parameters, etc. described in the present disclosure may be expressed using absolute values, relative values ​​from a predetermined value, or other corresponding information. For example, a radio resource may be indicated by an index.

[0076] The names used for the above-described parameters are not intended to be limiting in any way. Furthermore, the mathematical expressions using these parameters may differ from those explicitly disclosed in this disclosure. The various channels (e.g., PUCCH, PDCCH, etc.) and information elements may be identified by any suitable names, and therefore the various names assigned to these various channels and information elements are not intended to be limiting in any way.

[0077] In this disclosure, the terms "Mobile Station (MS)," "user terminal," "User Equipment (UE)," "terminal," etc. may be used interchangeably.

[0078] A mobile station may also be referred to by those skilled in the art as a subscriber station, mobile unit, subscriber unit, wireless unit, remote unit, mobile device, wireless device, wireless communication device, remote device, mobile subscriber station, access terminal, mobile terminal, wireless terminal, remote terminal, handset, user agent, mobile client, client, or some other suitable terminology.

[0079] As used in this disclosure, the terms "determining" and "determining" may encompass a wide variety of actions. "Determining" and "determining" may include, for example, judging, calculating, computing, processing, deriving, investigating, looking up, searching, inquiring (e.g., searching in a table, database, or other data structure), ascertaining, and the like. "Determining" and "determining" may also include receiving (e.g., receiving information), transmitting (e.g., sending information), input, output, accessing (e.g., accessing data in memory), and the like. Furthermore, "judgment" and "decision" can include regarding resolving, selecting, choosing, establishing, comparing, etc. as having been "judged" or "decided." In other words, "judgment" and "decision" can include regarding some action as having been "judged" or "decided." Furthermore, "judgment (decision)" can be interpreted as "assuming," "expecting," "considering," etc.

[0080] The terms "connected," "coupled," or any variation thereof, refer to any direct or indirect connection or coupling between two or more elements, and may include the presence of one or more intermediate elements between two elements that are "connected" or "coupled" to each other. The coupling or connection between elements may be physical, logical, or a combination thereof. For example, "connected" may be read as "access." As used in this disclosure, two elements may be considered to be "connected" or "coupled" to each other using one or more wires, cables, and / or printed electrical connections, as well as electromagnetic energy having wavelengths in the radio frequency range, microwave range, and optical (both visible and invisible) range, as some non-limiting and non-exhaustive examples.

[0081] As used in this disclosure, the phrase "based on" does not mean "based only on," unless expressly stated otherwise. In other words, the phrase "based on" means both "based only on" and "based at least on."

[0082] Any reference to an element using a designation such as "first," "second," etc., used in this disclosure does not generally limit the quantity or order of those elements. These designations may be used in this disclosure as a convenient way to distinguish between two or more elements. Thus, a reference to a first and a second element does not imply that only two elements may be employed or that the first element must in some way precede the second element.

[0083] When used in this disclosure, the terms "include," "including," and variations thereof are intended to be inclusive, similar to the term "comprising." Furthermore, when used in this disclosure, the term "or" is not intended to be an exclusive or.

[0084] In this disclosure, where articles are added by translation, such as a, an, and the in English, the disclosure may include that the nouns following these articles are in the plural form.

[0085] In the present disclosure, the term "A and B are different" may mean "A and B are different from each other." The term may also mean "A and B are each different from C." Terms such as "separate" and "coupled" may also be interpreted in the same way as "different." [Explanation of symbols]

[0086] 20...processing device (generation system), 21...reception unit, 22...image generation unit, 23...sound effect generation unit, 24...emotion estimation unit, 25...shape determination unit, 26...output unit.

Claims

1. a receiving unit that receives a generation request including a request for image generation from a user; an image generation unit that generates an image based on the generation request; a sound effect generation unit that generates a sound effect based on the image generated by the image generation unit; an emotion estimation unit that performs emotion estimation based on the sound effect generated by the sound effect generation unit or the request content; a shape determination unit that determines a shape of a speech bubble to be displayed around the image generated by the image generation unit based on a result of emotion estimation by the emotion estimation unit; an output unit that displays a display image including the speech bubble and the image, and outputs the sound effect in synchronization with the timing at which the display image is displayed.

2. The generation system according to claim 1 , wherein the reception unit receives the generation request in the form of at least one of a voice, a text, and an image.

3. The generation system according to claim 1 , wherein the image generation unit generates the image based on the request content included in the generation request and at least one of the user's location information, service usage history, attribute information, and hobbies and preferences.

4. The generation system according to claim 1 , wherein the shape determination unit determines a shape of the speech bubble based on the result of the emotion estimation and the image.

5. The generation system according to claim 1 , wherein the shape determination unit dynamically changes the shape of the speech bubble in response to a change in a result of emotion estimation by the emotion estimation unit.

6. The generation system according to claim 1 , wherein the output unit displays the display image in which the image is displayed along a frame of the speech bubble.

7. The generation system according to claim 1 , wherein the output unit displays the display image as a 3D image.

8. A production method executed by a production system, comprising: receiving a generation request from a user, the generation request including a request for image generation; generating an image based on the generation request; generating a sound effect based on the generated image; Estimating an emotion based on the generated sound effect or the request content; determining a shape of a speech bubble to be displayed around the image based on a result of emotion estimation; displaying a display image including the speech bubble and the image, and outputting the sound effect in synchronization with the timing at which the display image is displayed.

Citation Information

Patent Citations

  • Information processing system, information processing device and method, recording medium, and program

    JP2004206544A