Information processing system, information processing method, and program

The information processing system enhances AI-based interactions by integrating voice and image inputs to generate synchronized voice and animation responses, addressing the limitations of existing systems in providing immersive user experiences.

JP2025128858AInactive Publication Date: 2025-09-03DIGITAL HUMAN CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
JP2024025822
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-22
Publication Date
2025-09-03
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing AI-based response systems lack efficiency in providing comprehensive and immersive user interactions, particularly in handling voice and image inputs.

Method used

An information processing system that integrates voice and image inputs to an artificial intelligence module, generating synthesized voice and animation responses, enhancing user interaction through an avatar interface.

Benefits of technology

Improves usability by providing immersive and responsive interactions, allowing for quicker and more accurate processing of user inputs, including voice and image data, through an avatar-based conversation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025128858000001_ABST
    Figure 2025128858000001_ABST
Patent Text Reader

Abstract

To provide an information processing system, an information processing method, and a program that improve the usability of a response system that uses AI to respond to user utterances.SOLUTION: In an information processing system in which a server and an information processing device are connected via a telecommunications line (network), an information processing method by a control unit of the server includes, as reception steps, step S001 of receiving image information related to at least one image captured by an imaging device and step S002 of receiving voice information corresponding to voice uttered by a user, and as an input step, step S003 of inputting the voice information and the image information into an artificial intelligence module, and as an output step, step S004 of outputting synthesized voice data in response to a response from the artificial intelligence module.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing system, an information processing method, and a program. [Background technology]

[0002] Patent Document 1 is a document related to efficiently obtaining a reply message that a user expects when the user sends a chat message related to a question. When a multi-cloud chat service providing device 10 disclosed in Patent Document 1 receives a question message from a user terminal 20, it transmits the question message to an AI chat cloud service system 30 and receives a reply message to the question message from the AI ​​chat cloud service system 30. Furthermore, when the multi-cloud chat service providing device 10 determines that a predetermined condition is satisfied, it transmits the question message to an operator terminal 40A operated by an operator and receives a reply message to the question message from the operator terminal 40A. The multi-cloud chat service providing device 10 then returns the reply message received from the AI ​​chat cloud service system 30 or the operator terminal 40A to the user terminal 20. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Publication No. 2019-128737 Summary of the Invention [Problem to be solved by the invention]

[0004] However, there is still room for improvement in the technology of AI-based response systems that respond to user utterances. [Means for solving the problem]

[0005] According to one aspect of the present invention, there is provided an information processing system. The information processing system includes at least one processor configured to execute the following steps by reading a program: a receiving step for receiving voice information corresponding to a voice uttered by a user and image information relating to at least one image captured by an imaging device; an input step for inputting the voice information and image information into an artificial intelligence module; and an output step for outputting synthesized voice data in response to a response from the artificial intelligence module.

[0006] According to one aspect of the present invention, a more useful information processing system can be provided as a response system technology using AI. [Brief explanation of the drawings]

[0007] [Figure 1] 1 is a configuration diagram illustrating an information processing system 1. FIG. [Figure 2] FIG. 2 is a block diagram showing the hardware configuration of the server 2. [Figure 3] FIG. 2 is a block diagram showing a hardware configuration of an information processing device 3. [Figure 4] FIG. 2 is a diagram showing an outline of processing executed by the information processing system 1. [Figure 5] FIG. 2 is a diagram illustrating an example of a usage mode of the information processing system 1. [Figure 6] 2 is an activity diagram showing an example of the flow of processing executed by the information processing system 1. FIG. [Figure 7] FIG. 10 is a diagram illustrating an example of a relationship between a captured image and a user's speech. [Figure 8] FIG. 10 is a diagram showing an example of a screen 8 for accepting user input information. [Figure 9] 1 is a diagram showing an example of an image 9 including a user Y together with an object 91 and an object 92. FIG. [Figure 10] FIG. 10 shows an example of dividing a reply 10 into multiple segments. [Figure 11]10A and 10B are diagrams illustrating the relationship between generation and output of video data corresponding to each segment. DETAILED DESCRIPTION OF THE INVENTION

[0008] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The present invention will be described below with reference to the accompanying drawings. Various features shown in the following embodiments can be combined with each other.

[0009] Incidentally, the program for realizing the software appearing in this embodiment may be provided as a non-transitory computer-readable medium, or may be provided so that it can be downloaded from an external server, or may be provided so that the program is started on an external computer and its functions are realized on a client terminal (so-called cloud computing).

[0010] In this embodiment, the term "unit" may also include, for example, a combination of hardware resources implemented by a circuit in the broad sense and software information processing that can be specifically realized by these hardware resources. In addition, various types of information are handled in this embodiment, and this information may be represented by, for example, physical values ​​of signal values ​​representing voltages and currents, high and low signal values ​​as a binary bit set consisting of 0 or 1, or quantum superposition (so-called quantum bits), and communication and calculations may be performed on a circuit in the broad sense.

[0011] In addition, a circuit in the broad sense is a circuit realized by at least appropriately combining a circuit, circuitry, a processor, a memory, etc. That is, it includes an application specific integrated circuit (ASIC), a programmable logic device (e.g., a simple programmable logic device (SPLD), a complex programmable logic device (CPLD), and a field programmable gate array (FPGA)), etc.

[0012] 1. Hardware Configuration This section explains the hardware configuration.

[0013] <Information Processing System 1> FIG. 1 is a configuration diagram showing an information processing system 1. The information processing system 1 includes a server 2 and an information processing device 3. The server 2 and the information processing device 3 are configured to be able to communicate with each other via a telecommunications line (network). Here, the system exemplified as the information processing system 1 is made up of one or more devices or components. Therefore, it should be noted that the information processing system 1 includes either the server 2 alone or both the server 2 and the information processing device 3. More specifically, the information processing system 1 may include an element selected from the group consisting of the server 2 and the information processing device 3. The unselected element may not be included in the information processing system 1, but may be electrically connected to the selected element as an external element. These components will be described below.

[0014] <Server 2> 2 is a block diagram showing the hardware configuration of server 2. Server 2 includes a communication unit 21, a storage unit 22, and a control unit 23, and these components are electrically connected via a communication bus 20 inside server 2. Each component will be further described below.

[0015] The communication unit 21 is preferably a wired communication means such as USB, IEEE1394, Thunderbolt (registered trademark), wired LAN network communication, etc., but may also include wireless LAN network communication, mobile communication such as 3G / LTE / 5G, BLUETOOTH (registered trademark) communication, etc. as needed. In other words, it is more preferable to implement it as a collection of multiple communication means. In other words, the server 2 may communicate various information from the outside via the communication unit 21 and the network.

[0016] The memory unit 22 stores various pieces of information defined above. This can be implemented, for example, as a storage device such as a solid state drive (SSD) that stores various programs and the like related to the server 2 executed by the control unit 23, or as a memory such as a random access memory (RAM) that stores temporarily required information (arguments, arrays, etc.) related to the program operations. The memory unit 22 stores various programs, variables, etc. related to the server 2 executed by the control unit 23.

[0017] The control unit 23 processes and controls the overall operations related to the server 2. The control unit 23 is, for example, a central processing unit (CPU) not shown. The control unit 23 realizes various functions related to the server 2 by reading out predetermined programs stored in the storage unit 22. In other words, information processing by software stored in the storage unit 22 is specifically realized by the control unit 23, which is an example of hardware, and each step related to each function described below can be executed. This will be described in further detail in the next section. Note that the control unit 23 is not limited to being single, and multiple control units 23 may be provided for each function. A combination of these may also be used.

[0018] <Information processing device 3> 3 is a block diagram showing the hardware configuration of the information processing device 3. The information processing device 3 includes a communication unit 31, a storage unit 32, a control unit 33, a display unit 34, an input unit 35, and an output unit 36, and these components are electrically connected via a communication bus 30 inside the information processing device 3. Each component will be further described. The description of the communication unit 31, the storage unit 32, and the control unit 33 will be omitted because they are the same as the description of each unit in the server 2.

[0019] The display unit 34 may be included in the housing of the information processing device 3 or may be externally attached. The display unit 34 displays a graphical user interface (GUI) screen that can be operated by the user. This is preferably implemented by selectively using display devices such as a CRT display, a liquid crystal display, an organic EL display, and a plasma display depending on the type of the information processing device 3.

[0020] The input unit 35 may be included in the housing of the information processing device 3 or may be externally attached. The input unit 35 receives voice uttered by the user and receives input of images captured by the imaging device. The input unit 35 may be configured, for example, with a sound collector such as a microphone, collect external sound, and output a sound signal indicating the collected sound. The sound signal is transferred as a command signal to the control unit 33 via the communication bus 30, and the control unit 33 may execute predetermined control or calculation as necessary. The input unit 35 also supplies images input from the imaging device to the control unit 33. The information received by the input unit 35 is not limited to the above-mentioned voice and images. Specifically, the input unit 35 may be integrated with the display unit 34 and configured to receive input from the user via a touch panel, switch buttons, a mouse, a QWERTY keyboard, or the like. The input unit 35 receives, for example, information other than voice and images from the user.

[0021] The output unit 36 ​​may be included in or external to the housing of the information processing device 3. The output unit 36 ​​is configured by, for example, a speaker or the like, and outputs voice, signal sounds, etc. generated by the information processing system 1.

[0022] The information processing device 3 may be a smartphone, a tablet terminal, a personal computer, a wearable device, or the like.

[0023] 2. Functional configuration of Server 2 The control unit 23 is configured to execute, for example, the following steps: The following steps can be optionally omitted.

[0024] The control unit 23 is configured to receive information from the information processing device 3 or another device as a receiving step. The control unit 23 is also configured to receive various pieces of information by reading various pieces of information stored in a storage area, which is at least a part of the memory unit 22, and writing the read information to a working area, which is at least a part of the memory unit 22. The storage area is, for example, an area of ​​the memory unit 22 implemented as a storage device such as an SSD. The working area is, for example, an area implemented as a memory such as a RAM. For example, the receiving step may be a step of receiving audio information corresponding to a voice uttered by a user and image information related to at least one image captured by an imaging device. The receiving step may further be a step of receiving user input information other than the audio information and the image information based on a user's terminal operation.

[0025] In the input step, the control unit 23 inputs various information to the artificial intelligence module. The information may be voice information corresponding to a voice uttered by the user, image information relating to at least one image captured by an imaging device, user input information other than the voice information and image information received based on the user's terminal operation, etc. The control unit 23 may generate a prompt for the artificial intelligence module based on the information and reference information previously stored in the storage unit 22, and input the prompt together with the information to the artificial intelligence module.

[0026] The control unit 23 outputs various information in the output step. The information can be output to the user via the display unit 34 or another device. In such a case, for example, the control unit 23 may control the display unit 34 to display visual information such as a screen, an image including a still image or a video, an icon, or a message. For example, the control unit 23 outputs voice data synthesized in response to a response from the artificial intelligence module in the output step. Furthermore, the output step may be a step of outputting the voice data synthesized in response to the response from the artificial intelligence module in synchronization with an animation of an avatar responding.

[0027] In the generation step, the control unit 23 generates various information based on the response received from the artificial intelligence module via the communication unit 21 and the network. The information is response data including an animation of an avatar speaking and a synthetic voice corresponding to the response. In the generation step, the control unit 23 is configured to be able to generate synthetic voices and animations. For example, in the generation step, the control unit 23 may generate an animation of an avatar responding based on the response from the artificial intelligence module.

[0028] As a display control step, the control unit 23 controls the display unit 34 to display visual information such as a screen, an image including a still image or a video, an icon, a message, etc. The control unit 23 may generate only rendering information for displaying the visual information on the display unit 34. As a display control step, the control unit 23 may display an avatar in a manner that is visible to the user.

[0029] 3. Information processing flow This section describes the flow of an information processing method executed by the information processing system 1. As shown below, the information processing method includes each step executed by the information processing system. The information processing program of this embodiment causes a computer to execute each step of the information processing system. Note that the order of the processes can be changed as appropriate, multiple processes may be executed simultaneously, or some processes may be omitted.

[0030] 3.1 Overview 4 is a diagram showing an outline of the processing executed by the information processing system 1. In this processing, first, the control unit 23 receives an image captured by an imaging device as a receiving step (step S001). In parallel with step S001, the control unit 23 receives voice information corresponding to a voice uttered by a user as a receiving step (step S002). As an input step, the control unit 23 inputs the voice information and image information to an artificial intelligence module (step S003). Subsequently, as an output step, the control unit 23 outputs voice data synthesized based on a response from the artificial intelligence module (step S004).

[0031] In summary, an information processing system according to one embodiment includes at least one processor, and the processor reads a program to perform the following steps. In a receiving step, the control unit 23 receives audio information corresponding to a voice uttered by a user and image information related to at least one image captured by an imaging device. In an input step, the control unit 23 inputs the audio information and image information to an artificial intelligence module. In an output step, the control unit 23 outputs synthesized audio data in response to a response from the artificial intelligence module. With this configuration, a response can be output by audio data in response to the audio information and image information received from the user, improving usability.

[0032] 3.2 Specific examples Hereinafter, details of the above information processing will be described as an example with reference to FIGS. 5 to 9. FIG. 5 is a diagram illustrating an example of how the information processing system 1 is used. FIG. 5 illustrates an information processing device 3 and a user Y. An avatar AV is displayed on the display unit 34 of the information processing device 3, and the user Y is speaking a voice 51. The control unit 33 of the information processing device 3 receives the voice 51 spoken by the user Y via the input unit 35. An imaging device 351 is also connected to the information processing device 3 and captures an image of the speaking user Y. The control unit 33 receives an image captured by the imaging device 351 via the input unit 35. In this way, the information processing system 1 is, for example, a system in which the avatar AV responds to the user Y by voice. By using the avatar AV in this way, it is possible to provide the user Y with a sense of immersion that is not possible with a conversation based solely on text. Note that the information processing system 1 of this embodiment only needs to be configured to receive image information related to an image captured by the imaging device 351. Therefore, the imaging device 351 does not necessarily need to be connected to the information processing device 3.

[0033] FIG. 6 is an activity diagram showing an example of the flow of processing executed by the information processing system 1. This example of the flow may fall within the scope defined in the above-mentioned overview. The following description will be given along with each activity in this activity diagram. Note that the information processing may include any exception handling not shown. Exception handling includes the interruption of the information processing or the omission of each process. Selections or inputs made in the information processing may be based on user operation or may be made automatically without user operation.

[0034] First, the control unit 23 displays an avatar on the display unit 34 (activity A101). Here, the avatar may be a human figure, or an anthropomorphized animal or inanimate object. Preferably, the avatar is a realistic 3D model created to look exactly like a human and move in a human-like manner. A more human-like figure can enhance the user Y's sense of immersion and encourage deeper dialogue.

[0035] Next, the control unit 23 executes the processing of the activity A102 and the processing of the activity A104 in parallel.

[0036] In activity A102, user Y is imaged by imaging device 351 (activity A102). Specifically, imaging device 351 images user Y while he is speaking to the avatar AV, and control unit 33 receives the image captured by imaging device 351 via input unit 35. Imaging device 351 may capture multiple images (imaging device 351 may capture images continuously). Control unit 33 may temporarily store the captured image in storage unit 32.

[0037] Next, as a receiving step, the control unit 23 receives image information regarding at least one image captured by the imaging device from the control unit 33 via the network (activity A103). The image information may be the image data itself, or image data converted to a predetermined file size or file format. The image information may also include a time code and other information. The time code is information for identifying a time within user Y's speech. The imaging device may capture an image once while user Y is speaking, and the control unit 23 may receive one image. Alternatively, the imaging device may capture images multiple times while user Y is speaking, and the control unit 23 may receive multiple images. Alternatively, the imaging device may capture a video continuously rather than capturing still images, and extract and receive still images from the video. With this configuration, while the control unit 23 is receiving the user's voice, at least one image can be automatically received from among the images captured by the imaging device. Conventionally, when inputting an image to an artificial intelligence module, a user had to capture an image, select the captured image, and input it to the artificial intelligence module. Compared to the conventional method, it is possible to input images quickly and without the user having to do much work.

[0038] Furthermore, in one preferred embodiment, the imaging device 351 continuously captures images, and the control unit 23 receives, as image information, image information relating to the image captured at the final timing of the user Y's utterance. According to this configuration, the control unit 23 can automatically receive image information relating to the image captured at the final timing of the user's utterance and input it to the artificial intelligence module.

[0039] The following describes the acceptance of image information related to an image captured at the end of a user's speech, using FIG. 7 . FIG. 7 is a diagram illustrating an example of the relationship between captured images and the user's speech. While the user is speaking, the imaging device captures images at a predetermined interval. The predetermined interval can be set to any interval. For example, if the imaging device is set to capture images every 15 seconds, images 71, 72, 73, and 74 are captured every 15 seconds while the user is speaking. When the control unit 23 determines that the user's speech has ended, the control unit 23 accepts image information related to the last captured image at the end of the speech. That is, in the case of FIG. 7 , the control unit 23 accepts image 74. Furthermore, compared to accepting all images captured by the imaging device and inputting all images into the artificial intelligence module, this allows for faster processing due to the smaller data size. The reason for accepting an image captured at the end of a user's speech is that it is likely to clearly represent the content of the user's speech. Of course, the system may be configured to accept images captured at other times.

[0040] In activity A104, the control unit 23 receives, as a receiving step, audio information corresponding to the voice uttered by the user Y via the input unit 35. The audio information may be the audio data itself corresponding to the voice uttered by the user Y, or audio data converted into a predetermined file format. The audio information may include a time code. By including a common time code in the image information and the audio information, the timing of the audio information and the image information can be associated with each other.

[0041] Incidentally, the timing of the last moment of the user's utterance may be, for example, when the voice from user Y is interrupted for a predetermined period of time. Specifically, for example, when the voice from user Y is interrupted for 1 second or more, 2 seconds or more, 3 seconds or more, 4 seconds or more, or 5 seconds or more, the control unit 23 may determine that the voice of user Y has been interrupted. With this configuration, the voice of user Y can be quickly accepted as voice information at the last moment of the user's utterance.

[0042] Furthermore, in a more typical embodiment, the control unit 23 may convert the speech of user Y into text in real time (activity A105). That is, the control unit 23 may input speech information corresponding to user Y's speech into a speech recognition model and convert the speech information into text data. The speech recognition model may be configured with elements such as a language model or an acoustic model, or may use a model that has been trained in advance using speech information and text data. In the case of converting the speech into text in real time, the control unit 23 may also determine the timing of the end of the user's speech based on the context, punctuation marks, etc. Examples of punctuation marks include periods, commas, periods, semicolons, colons, hyphens, parentheses, etc. Note that if the artificial intelligence module described below is configured to accept speech data as is, converting the speech of user Y into text (including determining whether there is a break in the speech, etc.) may be omitted. For example, the control unit 23 may input the speech itself to an artificial intelligence module having an internal speech recognition function.

[0043] Furthermore, the control unit 23 may execute the processing of activity A106 in parallel with the processing of activity A102 and the processing of activity A104. In activity A106, the control unit 23 may accept, as a receiving step, user input information other than audio information and image information based on a terminal operation by the user.

[0044] FIG. 8 is a diagram showing an example of a screen 8 for accepting user input information. The acceptance screen 8 includes an area 81, a button 82, and a button 83. The area 81 is configured to be able to accept various files. When user Y operates the terminal to drag and drop various files into the area 81, the control unit 23 accepts the various files. The button 82 is a button for accepting an instruction to start accepting operations for selecting various files. When the control unit 23 accepts the pressing of the button 82 by user Y, it presents a screen for selecting various files to user Y. The button 83 is a button for accepting the registration of files. When user Y presses the button 82, selects various files, and then presses the button 83, the control unit 23 can accept the selected various files.

[0045] The user input information may include, for example, image data, text data, and data acquired by sensors such as temperature and humidity sensors. Various file formats include PDF, JPG, PNG, TXT, ZIP, and PSD. While activity A103 accepts image information related to an image captured by an imaging device, the control unit 23 can also accept image data selected by the user as user input information. This configuration allows, for example, image data to be received along with a user's voice request such as "I'd like to hear your opinion about this photo" and a response to that request. It is also possible to receive text data along with a user's voice request such as "I'd like your thoughts after reading this article" and respond with a user's thoughts about the article, or to respond to a user's voice request such as "It's cold, isn't it?" based on information acquired by a temperature sensor.

[0046] Next, the control unit 23 inputs the voice information and image information to the artificial intelligence module as an input step (activity A107). Furthermore, when the control unit 23 receives user input information different from the voice information and image information, it may input the voice information, image information, and user input information to the artificial intelligence module. The artificial intelligence module is a module that outputs a response to the voice information, etc., upon receiving input of voice information, etc. The artificial intelligence module is preferably a module that outputs a response in natural language in response to input of voice information, etc. Furthermore, the artificial intelligence module is preferably a module that understands the content of user Y's utterance and responds according to that content. In other words, the artificial intelligence module is preferably a conversational artificial intelligence module.

[0047] More typically, the artificial intelligence module may have a large-scale language model. That is, the artificial intelligence module may respond to input items based on the large-scale language model. Note that the large-scale language model is a deep learning model that pre-trains a language model that models human spoken language based on its occurrence probability from a huge amount of data. Note that the models that the artificial intelligence module can have are not limited to those described above.

[0048] Furthermore, the artificial intelligence module is a module that can accept input of not only text information but also image information. The artificial intelligence module is preferably a module that can also accept input of voice data and files of various formats. In other words, the artificial intelligence module is preferably a multimodal-compatible type artificial intelligence module. A multimodal-compatible type artificial intelligence module is an artificial intelligence module that integrates and processes data of multiple different formats, such as text, voice, image, and video. By using a multimodal-compatible type artificial intelligence module, it is possible, for example, to associate images and voices and then respond. Therefore, the artificial intelligence module is preferably a multimodal-compatible conversational artificial intelligence module.

[0049] Furthermore, in a more typical embodiment, when the control unit 23 inputs voice information and image information to the artificial intelligence module as an input step, the control unit 23 may input a generated prompt, voice information, and image information based on pre-stored reference information and voice information. The prompt may include, for example, requirements for a response. This configuration makes it possible to control the quality of the response. The prompt may also include information associating the voice information with the image information. This configuration clarifies the requirements for the response required regarding the image information, thereby further improving the quality of the response. Furthermore, when the control unit 23 inputs voice information, image information, and user input information to the artificial intelligence module as an input step, the prompt may also include requirements regarding the user input information. This configuration makes it possible to control the quality of the response regarding the user input information.

[0050] Next, the control unit 23 receives a response from the artificial intelligence module (activity A108).

[0051] Next, the control unit 23 generates reply data for the user Y in response to the received reply (activity A109).

[0052] Next, the control unit 23 outputs response data to the user Y (activity A110). Here, the response data includes at least voice data synthesized in response to the response from the artificial intelligence module. In a more preferred embodiment, the response data is output in synchronization with an animation of an avatar AV speaking in response to the response from the artificial intelligence module and a synthesized voice corresponding to the response.

[0053] More preferably, the animation of the avatar AV's speech may include at least one of mouth movement, body movement, and facial expression of the avatar AV in accordance with the speech of the avatar AV. This configuration further improves the realism of the animation of the avatar AV's speech and further enhances the user's sense of immersion.

[0054] 4. Variations Furthermore, the following aspects may be adopted: The above-described information processing aspects are merely examples, and the present invention is not limited to these, and can be modified as appropriate within the scope of the technical concept of the invention.

[0055] In the above embodiment, it has been described that an avatar is displayed on the display unit 34 in activity A101, but the avatar does not have to be displayed on the display unit 34 at the stage of activity A101. The control unit 23 may output an avatar in response to an utterance from the user.

[0056] In the above embodiment, as a preferred aspect, in activity A103, image information about an image captured at the final timing of the user's utterance is described. However, if the user's voice contains information about a shooting time, the control unit 23 may accept, as a receiving step, image information about an image captured at a timing corresponding to the information. If the user's voice contains information about a captured image, the control unit 23 may accept, as a receiving step, image information about an image captured at a timing corresponding to the information. That is, if the user's utterance contains information about an image, such as "How does my expression look right now?", the control unit 23 may accept image information captured at the timing closest to that information. With this configuration, if the user's utterance contains information about a shooting time, image information about an image captured at a timing corresponding to the information can be input to the artificial intelligence module.

[0057] In the above embodiment, as an example, image information regarding an image captured of a user is received. However, the image information may also include information regarding objects other than the user. Here, a case where the image information is an image including an object other than the user will be described with reference to FIG. 9 . FIG. 9 is a diagram illustrating an example of an image 9 including user Y as well as objects 91 and 92. Object 91 is a hat worn by user Y, and object 92 is a cat photographed with user Y. With this configuration, image information including an object other than the user is input to the artificial intelligence module. For example, when user Y utters, "Does this hat look good on me?", the control unit 23 inputs voice information corresponding to "Does this hat look good on me?" and image information including object 91 to the artificial intelligence module. The artificial intelligence module outputs a response based on the information about object 91. Furthermore, when user Y utters, "This is my pet," the control unit 23 inputs voice information corresponding to "This is my pet" and image information including object 92 to the artificial intelligence module. The artificial intelligence module can output a response based on the information about object 92.

[0058] In the activity A108 of the above embodiment, the control unit 23 may divide the response from the artificial intelligence module into multiple segments before accepting the response. By dividing the response into multiple segments at this stage, the time required for processing such as generating response data can be shortened.

[0059] In the activity A107 of the above embodiment, the control unit 23 may include a requirement regarding the language to be output by the AI ​​module when inputting the voice information and image information to the AI ​​module as an input step. The requirement regarding the language may be specified, for example, by a language tag. In this manner, a response can be obtained in a language different from the language input to the AI ​​module.

[0060] Here, when the artificial intelligence module outputs a response in a streaming format, the control unit 23 may divide the response transmitted from the artificial intelligence module in a streaming format into multiple segments and accept the divided response. To put it another way, outputting a response in a streaming format means that the artificial intelligence module outputs a response to an input prompt in succession in stages. In such a case, the control unit 23 may divide the response that is output in stages into multiple segments and accept the divided response. In this manner, the control unit 23 can accept the first segment of the response earlier than when all responses have been output.

[0061] The segments may be formed by dividing a reply based on punctuation marks. Punctuation marks include periods, commas, periods, semicolons, colons, hyphens, parentheses, etc. Punctuation marks are used to organize sentences in an easy-to-understand manner and make them easier to read. In this manner, the control unit 23 can divide the reply into segments for each clearly organized unit. The control unit 23 may accept a reply by dividing it into multiple segments depending on the language of the reply. In other words, the control unit 23 may divide the reply using punctuation marks according to the language of the reply.

[0062] Next, the control unit 23 generates reply data corresponding to the reply for each segment as a generation step. Here, the control unit 23 may generate reply data corresponding to the received segments in order as the generation step. With this configuration, reply data is generated in order from the received segments without waiting for the last reply from the artificial intelligence module, thereby shortening the time required to generate reply data corresponding to the first segment and improving usability.

[0063] FIG. 10 is a diagram showing an example of dividing a reply 10 into multiple segments. The reply 10 includes a segment S1, a segment S2, a segment S3, and a segment Sn. The reply 10 is an example of a reply transmitted in streaming format from an artificial intelligence module. The segments S1 to Sn are obtained by dividing the reply 10 by periods.

[0064] FIG. 11 is a diagram illustrating the relationship between the generation and output of video data corresponding to each segment. The explanatory diagram shown in FIG. 11 illustrates the timing at which activities A108 to A110 are executed for each of segments S1 to Sn. First, upon receiving segment S1 from the artificial intelligence module, the control unit 23 generates video data AD1 corresponding to segment S1 and outputs the video data AD1 corresponding to segment S1. Here, the control unit 23 receives segment S2 following segment S1, generates video data AD1 corresponding to segment S1, and then generates video data AD2 corresponding to segment S2. Furthermore, the control unit 23 generates video data AD2 corresponding to segment S2, and then generates video data AD3 corresponding to segment S3. Thus, as a generation step, the control unit 23 generates video data ADn-1 corresponding to segment Sn-1, and then generates video data ADn corresponding to segment Sn. Meanwhile, as a generation step, the control unit 23 outputs video data AD1 corresponding to segment S1 in parallel with generating video data AD2 corresponding to segment S2.

[0065] However, the time required to generate the video data AD1 corresponding to the segment S1 and the time required to output the video data AD1 corresponding to the segment S1 are not necessarily the same.

[0066] If the time required for generation is longer than the time required for output, generation of video data AD2 will not be completed when video data AD1 is output, resulting in a time gap between the output of video data AD1 and the start of output of video data AD2. Therefore, it is preferable that the control unit 23 generates the nth video data in the generation step so that the nth video data corresponding to the nth segment can be output consecutively with the n-1th video data corresponding to the n-1th segment. This configuration allows video data corresponding to multiple segments to be output continuously and without interruption. Because there is no interruption in the video data, user Y can converse with the avatar without stress.

[0067] On the other hand, if the time required for generation is shorter than the time required for output, when generation of video data AD2 is completed, output of video data AD1 has not yet finished, and video data AD2 cannot be output immediately. Therefore, when control unit 23 generates the nth video data in the generation step, if output of the (n-1)th video data has not yet been completed, the processor queues output of the nth video data in the output step. With this configuration, if the nth video data is generated before output of the (n-1)th video data, it can be queued and then output in order.

[0068] 11, if the time required for generation is shorter than the time required for output, queuing Q2 occurs after generating video data AD2 corresponding to segment S2 and before outputting the video data AD2; queuing Q3 occurs after generating video data AD3 corresponding to segment S3 and before outputting the video data AD3; and queuing Qn occurs after generating video data ADn corresponding to segment Sn and before outputting the video data ADn. In such a case, the number of video data being queued may increase. Therefore, the control unit 23 may adjust the generation speed of generating video data in the generation step so as to optimize the number of video data being queued. With this configuration, video data is generated in accordance with the output speed, thereby optimizing resource utilization.

[0069] The overall configuration shown in Fig. 1 is an example and is not limited to this. For example, the server 2 may be distributed across two or more devices, or may be replaced by a cloud computing system. Furthermore, all processing may be performed by the server 2, or all processing may be performed by the information processing device 3. An application may be installed on the information processing device 3, and the information processing device 3 and the server 2 may work together to execute the processing described above.

[0070] The server 2 may be an on-premise server or a cloud server. The cloud server 2 may provide the above functions and processes in the form of, for example, SaaS (Software as a Service) or cloud computing.

[0071] In the above embodiment, the server 2 performs various storage and control operations, but multiple external devices may be used instead of the server 2. That is, various information and programs may be distributed and stored in multiple external devices using block chain technology or the like.

[0072] It may be provided in the following manner.

[0073] (1) An information processing system comprising at least one processor, the processor being configured to execute the following steps by reading a program: a reception step receiving audio information corresponding to a voice uttered by a user and image information relating to at least one image captured by an imaging device; an input step inputting the audio information and the image information into an artificial intelligence module; and an output step outputting synthesized audio data in response to a response from the artificial intelligence module.

[0074] According to this configuration, a response can be output by the artificial intelligence module in response to the voice information and image information received from the user, and the response can be output as voice data, thereby improving usability.

[0075] (2) In the information processing system described in (1) above, the generation step further generates an animation of an avatar responding based on the response from the artificial intelligence module, and the output step outputs the voice data and the animation in synchronization.

[0076] With this configuration, the user can feel as if he or she is having a conversation with an avatar, improving usability.

[0077] (3) In the information processing system described in (2) above, the animation includes at least one of the avatar's mouth movements, body movements, and facial expressions in accordance with the avatar's speech.

[0078] This configuration further improves the realism of the animation of the avatar speaking, further increasing the user's sense of immersion.

[0079] (4) In the information processing system described in any one of (1) to (3) above, the imaging device continuously captures images, and in the receiving step, image information regarding the image captured at the last moment of the user's speech is received as the image information.

[0080] According to this configuration, image information relating to the image captured at the final timing of the user's utterance is automatically input to the artificial intelligence module.

[0081] (5) In the information processing system according to any one of (1) to (4) above, the image information includes information about an object other than the user.

[0082] With this configuration, image information about an object related to a user's utterance can be input to the artificial intelligence module.

[0083] (6) In the information processing system described in any one of (1) to (5) above, in the reception step, if the voice uttered by the user contains information regarding time, the system receives image information regarding an image captured at a timing corresponding to the information.

[0084] With this configuration, if the user's speech contains information regarding the shooting time, image information regarding the image captured at the timing corresponding to that information can be input to the artificial intelligence module.

[0085] (7) In the information processing system described in any one of (1) to (6) above, in the reception step, user input information different from the voice information and the image information is received based on the user's terminal operation, and in the input step, the voice information, the image information, and the user input information are input into the artificial intelligence module.

[0086] According to this configuration, input information input by the user can be input to the artificial intelligence module together with audio information and image information.

[0087] (8) An information processing method, comprising the steps of the information processing system according to any one of (1) to (7) above.

[0088] With this configuration, a response can be output by the artificial intelligence module in response to the voice and image received from the user, and the response can be output as voice data, improving usability.

[0089] (9) A program that causes at least one computer to execute each step in the information processing system according to any one of (1) to (7) above.

[0090] With this configuration, a response can be output by the artificial intelligence module in response to the voice and image received from the user, and the response can be output as voice data, improving usability. Of course, this is not the case.

[0091] Finally, while various embodiments of the present invention have been described, these are presented by way of example only and are not intended to limit the scope of the invention. The novel embodiments may be embodied in various other forms, and various omissions, substitutions, and modifications may be made without departing from the spirit of the invention. Such embodiments and modifications are intended to be included within the scope and spirit of the invention, as well as within the scope of the inventions and their equivalents as defined in the accompanying claims. [Explanation of symbols]

[0092] 1: Information processing system 2: Server 20: Communication bus 21: Communications Department 22: Storage section 23: Control section 3: Information processing equipment 30: Communication bus 31: Communications Department 32: Storage section 33: Control section 34: Display section 35: Input section 351: Imaging device 36: Output section 51: Audio 71: Image 72: Image 73: Image 74: Image 8: Reception screen 81 :Area 82: Button 83: Button 9: Image 91: Object 92: Object AD: Video data AD1: Video data AD2: Video data AD3: Video data ADn: Video data AV:Avatar Q2: Queuing Q3: Queuing Qn: Queuing S1: Segment S2: Segment S3: Segment Sn: Segment

Claims

1. An information processing system, At least one processor is provided, the processor being configured to execute the following steps by reading a program: In the receiving step, voice information corresponding to a voice uttered by the user and image information relating to at least one image captured by the imaging device are received; In the input step, the voice information and the image information are input to an artificial intelligence module; In the output step, the system outputs synthesized voice data in response to the response from the artificial intelligence module.

2. 2. The information processing system according to claim 1, Furthermore, in the generating step, an animation of an avatar making a response is generated based on the response from the artificial intelligence module; In the output step, the audio data and the animation are output in synchronization with each other.

3. 3. The information processing system according to claim 2, The system, wherein the animation includes at least one of mouth movements, body movements, and facial expressions of the avatar in sync with the avatar's speech.

4. 2. The information processing system according to claim 1, the imaging device continuously captures images; In the receiving step, the system receives, as the image information, image information relating to an image captured at the final timing of the user's utterance.

5. 2. The information processing system according to claim 1, The system, wherein the image information includes information about an object other than the user.

6. 2. The information processing system according to claim 1, In the receiving step, if the voice uttered by the user includes information about time, the system receives image information about an image captured at a timing corresponding to the information.

7. 2. The information processing system according to claim 1, In the receiving step, user input information different from the voice information and the image information is received based on a terminal operation by the user; In the input step, the voice information, the image information, and the user input information are input to the artificial intelligence module.

8. An information processing method, comprising: A method comprising the steps of the information processing system according to any one of claims 1 to 7.

9. A program, A program that causes at least one computer to execute each step in the information processing system according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method and system for emotional conversations between humans and machines

    JP2019012255A

  • Peripheral device, communication system, and communication program

    JP2019139653A

  • Program, device, and method for interacting with user in accordance with multimodal information around the user

    JP2022056638A

  • Impression evaluation method and impression evaluation system

    JP2023038870A

  • Human-computer interaction method, device, system, electronic device, computer-readable medium, and program

    JP2023552854A