COMMUNICATION SYSTEM AND SERVICE PROCESSING METHOD FOR COMMUNICATION SYSTEM

The communication system addresses the limitations of text-based interactions by integrating language, auditory, and visual elements, ensuring seamless and continuous communication without user account processing.

JP7827339B1Active Publication Date: 2026-03-10LIPRONEXT INC
View PDF 8 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing communication systems focus on unambiguous text-based interactions, often failing to provide intuitive visual and auditory responses, and require user account processing, leading to disrupted network access when response processing is needed.

Method used

A communication system that integrates language, auditory, and visual elements by generating language information and screen settings in response to user input, allowing for seamless interaction through a data terminal with a server device.

Benefits of technology

Enables a comprehensive response service combining language, auditory, and visual elements, providing intuitive and continuous communication without requiring user account processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007827339000001_ABST
    Figure 0007827339000001_ABST
Patent Text Reader

Abstract

To provide flexible support for response services that go beyond responses based on natural language processing and combine auditory and visual elements. [Solution] In a communication system in which a data terminal 1 connected by a user to a network 21 can communicate with a server device 100 that provides a specified communication service in response to the launch of a voice-enabled application, the server device 100 is characterized by a configuration in which it generates response information from language information input at the data terminal 1 and screen information linked to the response information.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a communication system and a service processing method for a communication system in which a data terminal, which a user operates to connect to a network, can communicate with a server device that provides a predetermined communication service while connected to the network via a predetermined communication medium. [Background technology]

[0002] Systems have been put into practical use that enable two-way communication by operating a data terminal to connect to a website and receive some kind of response from the server, allowing for free consultation, contracts, and inquiries based on mutual understanding.

[0003] Specifically, Patent Document 1 below discloses that "to facilitate user interaction with the device and facilitate more effective use of local and / or remote services, the intelligent automated assistant system 1002 engages the user through integrated dialogue using natural language dialogs and invokes external services when appropriate to obtain information or perform various actions. The system is implemented using any of a number of different platforms, such as the web, email, smartphones, etc., or any combination thereof. In one embodiment, the system employs additional functionality powered by external services with which the system can interact, based on a set of interrelated domains and tasks."

[0004] Furthermore, the following Patent Document 2 discloses that "as part of an interaction session between a user and an automated assistant, an implementation receives a stream of audio data capturing an utterance including an assistant query, determines a set of assistant outputs each predicted to be responsive to the assistant query based on processing the stream of audio data, processes the assistant outputs and the context of the interaction session using the large-scale language model output to generate a set of modified assistant outputs, and a given modified assistant output from among the set of modified assistant outputs can be presented to the user in response to the utterance. In some implementations, the LLM output can be generated in an offline manner for use in an online manner. In an additional or alternative implementation, the LLM output can be generated in an online manner as the utterance is received." [Prior art documents] [Patent documents]

[0005] [Patent Document 1] Japanese Patent Application Publication No. 2024-122975 [Patent Document 2] Special Publication No. 2024-521053 Summary of the Invention [Problem to be solved by the invention]

[0006] However, in automatic response processing, systems that use natural language processing focus on unambiguous text-based interactions, and users often expect intuitive, specifically visual, responses.

[0007] Here, "univocal" means a response that focuses only on the text provided by the user. If the response is the obvious "That's a problem. Is there anything else we can help you with...", the user responding will end up giving an average response and may not be able to find an effective solution.

[0008] Additionally, while there are systems that support visual responses, it is difficult to obtain the intended answer without switching the user interface provided by the server device to the user terminal or operating a button, and improvements in this area are desired.

[0009] Furthermore, in the above-mentioned systems, response processing is performed via a specific user interface provided by the server device, which requires some kind of user account processing and pre-processing to enable smooth communication services to begin.However, users who dislike being tied down by the server device often find that their network access is cut off at the time the response processing is required.

[0010] The present invention has been made to solve the above-mentioned problems, and aims to provide a communication system and a service processing method for a communication system that can provide a service that goes beyond responses based on natural language processing and freely supports response services that combine language elements, auditory elements, and visual elements by receiving language information generated by a server device in accordance with the operation of inputting the language information intended by a user operating a data terminal, and screen setting information to be output together with the language information. [Means for solving the problem]

[0011] The communication system of the present invention that achieves the above object has the following configuration.

[0012] The present invention is a communication system in which a data terminal having a function of inputting and outputting language information can communicate with a server device that provides a predetermined communication service via a predetermined communication medium, and the data terminal has an input unit that inputs the predetermined language information and an output unit that outputs the predetermined language information. and, The language information input from the input unit is transferred to the server device, and generated by the server device. language Information and the foregoing language information Output of Displayed in conjunction with will beScreen Information and generated by the server device. an acquisition means for acquiring operation instruction information; language a process of outputting information from the output unit, and a process of outputting information based on the operation instruction information acquired by the acquisition means. generating screen information for controlling the behavior and background of an avatar displayed on the display unit; and a control means for controlling the processing of outputting the information processed by the data terminal to a display unit. The aforementioned Get language information ,before A response setting for the recorded language information, a motion determination setting for determining the motion of the avatar to be displayed on the display unit, and a background setting for determining the background of the avatar to be displayed on the display unit. Ruta Meno back a language processing means for performing the answer setting, the action determination setting, the background determination setting, and the background determination setting set by the language processing means; Before and receiving a combination of the language information as a prompt, and displaying the language information to be responded to on the data terminal and the display unit. The screen information to be displayed is Avatar behavior settings and Background settings and the operation instruction information including and a generating means for generating the The operation instruction information includes operation settings and background settings of the avatar displayed on the display unit. It is characterized by: [Effects of the Invention]

[0014] According to the present invention, by receiving language information generated by a server device in response to the operation of inputting the language information intended by the user operating the data terminal, as well as screen setting information to be output together with the language information, it is possible to provide a service that goes beyond responses based on natural language processing and freely supports response services that combine language elements, auditory elements, and visual elements. [Brief explanation of the drawings]

[0015] The drawings illustrate particular embodiments of the present invention, including essential features of the invention as well as alternative and preferred embodiments. [Figure 1] 1 is a block diagram illustrating the configuration of a communication system according to an embodiment of the present invention. [Figure 2] 1 is a block diagram illustrating the configuration of a communication system according to an embodiment of the present invention. [Figure 3] FIG. 3 is a block diagram illustrating the configuration of the server device shown in FIG. 2. [Figure 4] 2A is a block diagram illustrating the configuration of the data terminal shown in FIG. 1, and FIG. 2B is a diagram illustrating the configuration of a program expanded in RAM shown in FIG. [Figure 5] 3A and 3B are diagrams illustrating a sequence of a series of communication service processes in the communication system according to the present embodiment. [Figure 6] 3A and 3B are diagrams illustrating a sequence of a series of communication service processes in the communication system according to the present embodiment. [Figure 7] 3A and 3B are diagrams illustrating a sequence of a series of communication service processes in the communication system according to the present embodiment. [Figure 8] FIG. 2 is a diagram showing a setting sequence for avatar processing on the server device side shown in FIG. [Figure 9] 4 is a flowchart showing a data processing procedure on the data terminal side in the communication system according to the embodiment. [Figure 10] 4 is a flowchart showing a data processing procedure on the data terminal side in the communication system according to the embodiment. [Figure 11] 10 is a flowchart showing a data processing procedure on the server device side in the communication system according to the embodiment. [Figure 12] 10 is a flowchart showing a data processing procedure on the server device side in the communication system according to the embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0016] Next, the best mode for carrying out the present invention will be described with reference to the drawings.

[0017] <System configuration description> [First embodiment] 1 and 2 are block diagrams illustrating the configuration of a communication system according to this embodiment. In this example, the configuration, operation, and effects of the communication system will be described in detail, based on a communication system in which a data terminal 1, which a user operates to connect to a network 21, and a server device 100, which provides a predetermined communication service while connected to the network 21, can communicate with each other via a predetermined communication medium.

[0018] 1 and 2 show an example of inputting linguistic information in which a text question converted into text information is sent to server device 100. However, the scope of the present invention also includes a configuration in which linguistic information is input as voice, and the voice information is then converted into digital information, which is then directly transferred to server device 100. The scope of the present invention also includes a configuration in which a text response (text data) as output information generated by server device 100 is converted into voice data, and a voice response is made from data terminal 1.

[0019] Therefore, this system is exemplified as a system in which a server device 100 having an output function that outputs to a data terminal 1 while coordinating a first generation process that acquires question information as linguistic information input at a data terminal 1, performs language processing on the question information based on natural language, and generates output information to be output at the data terminal 1, and a second generation process that generates operation instruction information that specifies screen information to be displayed at the data terminal 1 in conjunction with the output information generated by the first generation process.

[0020] Similarly, the data terminal 1 is assumed to be a system configured with a function to control a process of outputting the acquired answer information from a speaker as a predetermined output unit, and a process of outputting screen information generated based on the acquired operation instruction information to a display unit (corresponding to the display screen 13 described later), as well as an acquisition function of transferring voice information input from a predetermined input unit, for example, to the server device 100, and acquiring answer information as output information generated by the server device 100 and operation instruction information for screen information to be displayed in conjunction with the answer information.

[0021] That is, in this embodiment, the term "linguistic information" includes voice information and text information, and the language processing is not limited to a system adapted to processing only voice data or only text data.

[0022] Therefore, in the description of this specification or the drawings, the terms question information (question data) and response information (answer data) are to be interpreted as synonymous with linguistic information.

[0023] In addition, the data terminal 1 is assumed to have a function to make an avatar appear on the display screen 13 and move it around the display screen 13 based on operation instruction information obtained from the server device 100, and a function to display various background screens on the display screen 13 based on response information obtained from the server device 100.

[0024] Furthermore, the communication environment between the data terminal 1 and the server device 100 to which the present invention is applied is not limited to a system adapted to a specific OS environment, assuming an environment in which various OSs communicate using a predetermined protocol via their own browsers.

[0025] In Figure 1, reference numeral 1 denotes a data terminal equipped with a voice input unit 2, which inputs voice information generated by the user to a voice determination unit 4. The voice input unit 2, which may be configured, for example, as a microphone, has the function of converting analog voice signals into digital signals at a predetermined sampling frequency, and is configured to output, for example, 8-bit digital voice information. Communication on the data terminal 1 can include chat formats, voice responses, responses by displaying an avatar, and screen responses using a background screen displayed on the display unit, and it is also possible to combine these communication functions in response to the site the user accesses.

[0026] The voice determination unit 4 determines whether or not it is receiving voice input from the voice input unit 2, and outputs the received voice input to the voice-to-text converter STT 5. STT stands for Speech to Text, and performs the process of converting voice input into text in real time. Specifically, it uses voice recognition technology to convert the speaker's voice into text data. Note that some systems include multiple language models and are equipped with a function that can handle simultaneous voice input in languages ​​other than Japanese.

[0027] The voice determination unit 4 receives the question information converted into text by the voice-to-text conversion STT 5, and outputs it to the communication interface unit 8 with a header attached.

[0028] The communication interface unit 8 receives the question information from the voice determination unit 4 via the network 21 and outputs the packetized question information to the natural language processing unit 101 of the server device 100 .

[0029] The natural language processing unit 101 functions as a means for performing language processing based on a large language model. Specifically, this refers to a general term for technologies in the field known as natural language processing (NLP), which are constructed using massive amounts of text data and advanced deep learning technologies.

[0030] On the other hand, the communication interface unit 8 receives the answer information received from the natural language generation unit 102 and the operation instructions for the data terminal 1 in response to the question information from the natural language processing unit 101 on the server device 100 side via the network 21, outputs the answer information to the voice speech unit 6, and passes the operation instructions to the operation judgment unit 9 on the data terminal 1 side, which outputs them to the avatar and environment operation unit 10.

[0031] The avatar and environment operation unit 10 comprises an avatar operation unit 10-1 that processes the avatar's operations, a background operation unit 10-2 that processes the background of the screen, and an emotion operation unit 10-3 that processes the avatar's emotions.

[0032] Reference numeral 11 denotes a movement processing unit that comprehensively processes the movements processed by avatar movement unit 10-1, which processes the movements of the avatar, background movement unit 10-2, which processes the background of the screen, and emotion movement unit 10-3, which processes the emotions of the avatar. Reference numeral 12 denotes an avatar processing unit that controls the process of generating screen information including the avatar processed by movement processing unit 11 and outputting it to display screen 13. Here, display screen 13 corresponds to the display screen of display 117 shown in FIG. 3, which will be described later.

[0033] In addition, by providing a display control unit in place of or together with the avatar processing unit 12, which displays various background images on the display screen 13 in accordance with the answer information, it is possible to provide a highly impressive communication tool service by allowing the user to see visual information linked to the background image that cannot be expressed by the movements of the avatar displayed on the screen alone.

[0034] Furthermore, the communication interface unit 8 receives answer information from the natural language generation unit 102 via the natural language processing unit 101 on the server device 100 side, and outputs it to the voice utterance unit 6. After outputting the answer information to the text-to-speech unit 7, the voice utterance unit 6 receives the voice output uttered by the text-to-speech unit 7, and outputs the voice to the voice output unit 3 provided in the data terminal 1.

[0035] Here, the natural language processing corresponds to natural language processing using a large-scale language model, can adapt to the language environment selected by the user, and can transfer the language information selected by the user to the server device 100.

[0036] In the server device 100 shown in FIG. 2, the natural language processing unit 101 outputs the received question information, answer setting, and action instruction setting as a response sentence creation request to the natural language generation unit 102 based on the answer setting by the answer content setting unit 103 and the action instruction setting by the action instruction content setting unit 104. Here, the natural language generation unit 102 is assumed to be configured with various generation AIs, such as OpenAI (registered trademark). Furthermore, various versions of the natural language generation unit 102 exist, and it is assumed that the latest functions are not guaranteed depending on the version. In this embodiment, it is assumed to be configured with the latest version of the generation AI.

[0037] The natural language generation unit 102 generates answer information, for example composed of text, and operational instructions for the data terminal 1 in response to the question information, in order to respond to the data terminal 1 based on the question information, answer settings, and operational instruction settings received from the natural language processing unit 101.

[0038] The natural language processing unit 101 receives answer information from the natural language generation unit 102 via the network 21 and responds to the communication interface unit 8 of the data terminal 1 with operational instruction information for the data terminal 1 in response to the question information at an appropriate timing using a predetermined protocol.

[0039] Reference numeral 105 denotes a conversation database, which stores the conversation history (including questions and answers) exchanged between the data terminal 1 and the server device 100 according to a predetermined data structure.

[0040] A conversation history analysis unit 106 acquires conversation histories in chronological order from the conversation database 105 and executes predetermined analysis processing.

[0041] A prompt setting database 107 is made up of an answer content setting database 107-1, an avatar action setting database 107-2, and a background action setting database 107-3, and refers to the answer content setting database 107-1, the avatar action setting database 107-2, and the background action setting database 107-3 as needed to present necessary action setting information (setting information for issuing action instructions in response to question information) related to the screen display on the data terminal 1 to the natural language processing unit 101. Note that in the answer sentence generation sequence described below, the prompt setting database 107 may be described as the entity that controls each database.

[0042] As a result, the natural language processing unit 101 generates answer information and operation instructions for the question information based on the information provided from the prompt setting database 107, and outputs them to the communication interface unit 8 of the requesting data terminal 1 via the network 21. Note that one aspect of this embodiment is a configuration in which a separate database is employed that stores settings such as judgment criteria and knowledge for determining what emotional operation instructions should be given in response to user input.

[0043] The avatar and environment operation unit 10 plays image information (including still images, illustrations, and videos) displayed on the display screen 13 of the data terminal 1 in accordance with the operation instructions generated by the natural language generation unit 102, and also makes a registered avatar appear on the display screen 13 as a concierge (or attendant) and performs processing to explain the contents of the display screen 13.

[0044] In this system, the avatar and environment operation unit 10 is configured to be able to control the movement of the avatar and the output of sound in both real space and virtual space.

[0045] This allows for the avatar's movement speed and audio playback speed to be freely controlled.

[0046] [Example of Hardware Configuration of Server Device 100] Fig. 3 is a block diagram illustrating the configuration of the server device 100 shown in Fig. 2. In Fig. 3, the CPU 113 can also be configured as a processing device including a DSP (Digital Signal Processor), a CPU (Central Processing Unit), and a GPU (Graphics Processing Unit).

[0047] The DSP and GPU operate under the control of the CPU 113 and are responsible for executing image processing. Here, a processing device including a DSP, CPU 113, and GPU is given as an example of a processor, but this is merely an example, and the processor may be one or more CPUs and one or more GPUs, one or more CPUs and DSPs with integrated GPU functions, one or more CPUs and DSPs without integrated GPU functions, or may be equipped with a TPU (Tensor Processing Unit).

[0048] The RAM 116 is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0049] The RAM 116 may be, for example, a dynamic random access memory (DRAM) or a static random access memory (SRAM).

[0050] The communication unit 111 is an interface including a communication processor, an antenna, etc., and is connected to the internal bus 112. The communication unit 111 controls communication between the data terminal 1 and the server device 100, etc. The communication standard applied to the communication unit 111 may be, for example, a wireless communication standard including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc., or a wired communication standard including Ethernet (registered trademark), Fast Ethernet (registered trademark), Gigabit Ethernet (registered trademark), etc.

[0051] An external memory 118 stores web pages (HTML files) to be presented on the platform. The platform of this system is configured to be viewable on an account set by the server device 100, and is configured to be operable on a so-called cloud.

[0052] Reference numeral 114 denotes a setting database, which functions as an image database that stores large-scale models of background image patterns. Specifically, the setting database 114 corresponds to an integrated version of the prompt setting database 107 of the server device 100 shown in FIG. 2, and is configured to be able to update the setting data as needed. Furthermore, the setting data is managed by the CPU 113 with version information individually assigned. Reference numeral 115 denotes a keyboard, which is used to specify icons and dashboards displayed on the display 117, and to input text and commands.

[0053] An AI support unit 119 cooperates with an AI server (not shown) to support background image generation processing and edited image processing in synchronization with requests (including prompts) generated from voice responses from the data terminal 1. The AI ​​support unit 119 includes the natural language generation unit 102 shown in Fig. 2, and generates answer texts shown in the flowchart described below, and generates operational instructions for question information (operation instructions for determining display operations such as background processing on the screen, the timing of avatar appearance, and the form of the avatar to be displayed).

[0054] In this embodiment, a communication system is provided in which a data terminal 1, which a user operates to connect to the network 21, and a server device 100, which provides a predetermined communication service while an application is running, can communicate with each other via a predetermined communication medium (network 21), and the server device 100 is configured to include a natural language processing unit 101 that processes question information acquired from a communication interface unit 8 provided in the data terminal 1, and a natural language generation unit 102 that performs a generation AI function of generating answer information to the question information and operation instruction information for the question information, which are sent back to the data terminal 1, based on the question information acquired by the natural language processing unit 101 and multiple pieces of setting information generated by the natural language processing unit 101. Note that a configuration in which the server device 100 and the data terminal 1 communicate via a browser without going through a platform is also included in one aspect of this embodiment.

[0055] The RAM 116 is loaded with a natural language processing program 101P, a natural language generation program 102P, a natural language response content setting program 103P, an action instruction content setting program 104P, and a conversation history analysis program 106P, and the CPU 113 comprehensively controls the execution state of various functional processes in accordance with the procedures based on the flowcharts described below. Note that although each step is executed in chronological order, this embodiment also includes a mode in which the order is changed as appropriate in the program configuration and program control process.

[0056] The stable diffusion function used by the server device 100 will be described below.

[0057] (Learning model) Large-scale neural network models used for text generation are typically Transformer models such as GPT-3 and GPT-4.

[0058] (Hardware) In addition to the CPU 113, this system may require hardware resources such as a high-performance GPU (Graphics Processing Unit) or TPU (Tensor Processing Unit), which allows for efficient execution of large-scale models.

[0059] [Data terminal 1 hardware configuration example] FIG. 4 is a block diagram illustrating the configuration of data terminal 1 shown in FIG.

[0060] In FIG. 4, 301 denotes a CPU, which starts up a BIOS stored in a ROM 302 and controls an I / O device.

[0061] Reference numeral 303 denotes a RAM, which loads applications stored in an external memory (not shown) into the RAM 303 and executes various data processing. Reference numeral 304 denotes a communication unit, which is connected to the network 21 and is configured to be able to communicate bidirectionally with the server device 100. Reference numeral 311 denotes a display with a touch panel function, which displays input information. Reference numeral 303-3 denotes a UI control unit, which controls the rendering of a UI screen to be displayed on the display 311, and in this embodiment, controls the display of avatars and background screens based on operation instruction information acquired from the server device 100 in conjunction with the control of displaying output information generated by the server device 100 in response to input language information.

[0062] The display 311 includes a microphone 311M for audio input and a speaker 311S for audio output. The microphone 311M and the speaker 311S may be configured as separate devices from the display 311, or may be configured as Bluetooth (registered trademark) earphones with a microphone, etc.

[0063] Specifically, a voice judgment program 4P, a voice-to-text STT program 5P, a voice speech program 6P, a text-to-voice program 7P, a communication program 8P, and an action judgment program 9P are deployed in RAM 303, and CPU 301 executes each step according to the flowchart described below to comprehensively control the execution status of various functional processes.

[0064] Although the steps are executed in chronological order, this embodiment also includes a case in which the order is changed as appropriate in the course of program configuration and program control.

[0065] This example is a communication system in which a data terminal 1, which a user operates to connect to the network 21, can communicate with a server device 100 that provides a predetermined communication service while connected to the network 21 via a predetermined communication medium (network 21).The data terminal 1 is characterized by having a predetermined voice input unit 2 and a predetermined voice output unit 3, a voice judgment unit 4 that judges voice information input from the predetermined voice input unit 2 and generates question information, a communication interface unit 8 that connects the question information generated by the voice judgment unit 4 to a natural language processing unit 101 provided in the server device 100 via a predetermined API and receives answer information and predetermined operation instruction information processed by the server device 100, and a voice speaking unit 6 that performs a voice generation function of generating voice information from the answer information received by the communication interface unit 8 from the server device 100 and outputting it from the predetermined voice output unit 3.

[0066] 5 to 7 are diagrams illustrating a sequence of a series of communication service processes in the communication system according to this embodiment. Note that each of Messages 1 to 33 indicates an instruction.

[0067] In addition, the reference numerals at the top are the same as those in the block diagrams shown in FIGS. 1 and 2, and are used to denote components that perform the same functions.

[0068] Furthermore, due to limitations on the size of application drawings, Figures 5 to 7 are shown in multiple drawings, but the alphabets A to G are added to the branching points of each drawing to indicate that they are connected.

[0069] 5 and 7 correspond to a series of data processing on the data terminal 1 side, and Fig. 6 corresponds to a series of data processing on the server device 100 side. In particular, "par" in the figures indicates parallel processing, and "alt" indicates alternative processing.

[0070] First, as shown in Figure 5, a user operating the data terminal 1 speaks a question into the voice input unit 2 (Message 1), the question (text) is received by the voice judgment unit 4 (Message 2), question information is generated by the voice-to-text conversion STT5, and the question information (question) generated by the communication interface unit 8 is handed over to the natural language processing unit 101 of the server device 100 (Message 3).

[0071] Next, the natural language processing unit 101 acquires the question sent from the communication interface unit 8 of the data terminal 1 via the network 21 (Message 4).

[0072] Next, the natural language processing unit 101 sets an answer (Message 5) by referring to the prompt setting database 107. Similarly, the natural language processing unit 101 acquires an action determination setting by referring to the prompt setting database 107 (Message 6), and sets the action determination setting by referring to the prompt setting database 107 (Message 7).

[0073] Similarly, the natural language processing unit 101 acquires background determination settings by referring to the prompt setting database 107 (Message 8), and performs background determination settings by referring to the prompt setting database 107 (Message 9).

[0074] Next, the natural language processing unit 101 transmits the question sentence (question information) and the above-mentioned answer setting to the natural language generation unit 102 (Message 10). In response to this, the natural language generation unit 102 generates answer information (for example, an answer sentence based on a text format) and responds to the natural language processing unit 101 (Message 11).

[0075] Next, the natural language processing unit 101 transmits the question sentence (question information) and the above-mentioned action determination setting to the natural language generation unit 102 (Message 12). In response to this, the natural language generation unit 102 generates an action determination result and responds to the natural language processing unit 101 (Message 13).

[0076] Next, the natural language processing unit 101 transmits the question sentence (question information) and the above-mentioned background determination setting to the natural language generation unit 102 (Message 14). In response to this, the natural language generation unit 102 generates a background determination result and responds to the natural language processing unit 101 (Message 15).

[0077] Next, the natural language processing unit 101 responds to the data terminal 1 via the network 21 with the answer information generated by the natural language generation unit 102, the action determination result, and the background determination result (Message 16).

[0078] Next, avatar processing unit 12 instructs avatar operating unit 10-1 to perform an action, for example, a bow, based on the action determination result returned from server device 100 (Message 17).

[0079] As a result, the avatar processing unit 12 displays, for example, an animation of a bow on the display screen 13 (Message 18).

[0080] On the other hand, when the avatar processing unit 12 instructs the background operation unit 10-2 to display, for example, a specific background based on the operation judgment result received from the server device 100 (Message 19), the avatar processing unit 12 displays the specific background on the display screen 13 (Message 20).

[0081] Similarly, when the avatar processing unit 12 instructs the voice utterance unit 6 to utter an answer sentence based on the answer sentence returned from the server device 100 (Message 21), the answer sentence is reproduced from the voice output unit 3 (Message 22).

[0082] As a result, compared to when a simple text response is obtained in response to question information input by voice from the voice input unit 2, the user operating the data terminal 1 can obtain a response result that combines screen information (AI-generated screen or screen information related to avatar display) generated by the natural language generation unit 102 and voice information (AI-generated answer sentence), and the user can perceive a response that is both visually and audibly satisfying, as image processing by the display screen 13 is executed in parallel with the voice corresponding to the answer sentence. This concludes the explanation of the series of processes that occur when the natural language generation unit 102 independently determines that the voice input by the user in response to question information received by the server device 100 is a short conversation such as a greeting.

[0083] Next, a series of processes will be described below when it is determined that the voice input by the user is a long conversation such as a greeting.

[0084] Next, avatar processing unit 12 instructs avatar operating unit 10-1 to perform, for example, a backchannel action (Message 23) based on the action determination result returned from server device 100. As a result, an animation of the backchannel action is displayed on display screen 13 (Message 24).

[0085] Next, based on the operation determination result returned from server device 100, avatar processing unit 12 instructs background operation unit 10-2 to display, for example, a specific background (Message 25), and displays the specific background on display screen 13 (Message 26). Here, when the user inquires about property information, display screen 13 displays the floor plan of the selected document and the state of each room, allowing the user to visually confirm the property as if viewing it.

[0086] The contents of Message 25 and Message 20 may be different or partially the same. The criteria for distinguishing between different and partially the same may be based on the user's gender, age, tone of voice, rhythm, etc.

[0087] Next, based on the reply sentence received from the server device 100, the avatar processing unit 12 instructs the voice utterance unit 6 to utter a reply sentence (Message 27), and then reproduces the reply sentence (standard response sentence) from the voice output unit 3 (Message 28).

[0088] In addition, instead of Message 27, the avatar processing unit 12 instructs the avatar processing unit 12 to prepare to speak the answer sentence based on the answer sentence returned from the server device 100 (Message 29), and when preparation for playing the answer sentence is complete, it instructs the avatar processing unit 12 to that effect (Message 30).

[0089] Next, the voice utterance unit 6 reproduces the answer sentence from the voice output unit 3 (Message 31).

[0090] Similarly, avatar processing unit 12 instructs avatar operation unit 10-1 to perform an operation in accordance with the answer received from server device 100 (Message 32). Next, avatar operation unit 10-1 displays the answering operation of the avatar in accordance with the answer on display screen 13 in various display modes such as video and illustrations (Message 33).

[0091] As a result, the user operating the data terminal 1 can perceive a response that is visually, audibly, and even familiarly satisfying to the user, as a combination of the screen information, voice information, and avatar actions generated by the natural language generation unit 102, compared to when a simple text response is obtained in response to question information input by voice.

[0092] [Avatar processing setting sequence] FIG. 8 is a diagram showing a setting sequence of the avatar processing on the server device 100 side shown in FIG.

[0093] As shown in FIG. 8, in this embodiment, answer setting, action determination setting, and background determination setting are processed in parallel, so that the natural language generation unit 102 can respond to voice input by the user without any time lag in answer processing.

[0094] [Processing on Data Terminal 1 side] 9 and 10 are flowcharts showing the data processing procedure on the data terminal 1 side in the communication system according to this embodiment. Note that (1) to (16) indicate each step, which is realized by the CPU 301 executing a program stored in the ROM 302 and an external memory (not shown).

[0095] First, when the CPU 301 determines that the application started by the user operating the data terminal 1 is a voice-enabled application (1), when the CPU 301 determines that it is checking the voice input spoken by the user from the microphone 311M (2), and further when the CPU 301 determines that the voice input is temporarily interrupted because the user is taking a breath or thinking (voice input is intermittent) (3), the CPU 301 activates the voice judgment program 4P and the voice-to-text STT program 5P expanded in the RAM 303 to convert the voice data spoken by the user that has been input up to that point into text, and executes the process of converting the voice input into text (4).

[0096] Here, the text conversion process corresponds to the process of generating question information and additional information (which may include expressive elements (fast speech, language, etc.)) to be transferred to the server device 100 based on the voice generated by the user.

[0097] Next, the CPU 301 starts the voice judgment program 4P and the voice-to-text STT program 5P expanded in the RAM 303, and when it acquires the question information to be transferred to the server device 100, it transfers the question information to the natural language processing unit 101 of the server device 100 via the network 21 (5).

[0098] Next, the CPU 301 determines whether or not response information such as a question and answer sentence generated by the composite response processing consisting of the natural language processing and the natural language generation processing of the server device 100 has been received via the network 21 (6).

[0099] Here, if the CPU 301 determines that it has received response information from the server device 100, it further determines whether the received response information represents answer information (7). If the CPU 301 determines that it has received answer information, it executes the voice speech program 6P and the text-to-speech program 7P to generate an audio output for voice output, and outputs the answer information aloud from the speaker 311S (8).

[0100] Next, CPU 301 determines whether the user of data terminal 1 is performing an operation to disconnect from network 21 (12), and if it determines that an operation to disconnect from network 21 is being performed, ends the processing.

[0101] On the other hand, if the CPU 301 determines in step (7) that the response information is not answer information, it determines (9) whether the response information received from the server device 100 is an operation instruction in response to the question information, and if the CPU 301 determines that the response information is not an operation instruction in response to the question information, it returns to step (6).

[0102] On the other hand, in step (9), if the CPU 301 determines that the response information received from the server device 100 is an operation instruction in response to the question information, the CPU 301 further determines whether the received operation instruction is related to background processing or avatar appearance processing (10).

[0103] If the CPU 301 determines that the received operation instruction does not relate to background processing or avatar appearance processing, the CPU 301 proceeds to step (12).

[0104] On the other hand, if the CPU 301 determines in step (10) that the received action instruction is related to background processing or avatar appearance processing, the CPU 301 executes an analysis process to determine the action content that accompanies the performance based on the action instruction (11).

[0105] Next, CPU 301 executes action determination program 9P to determine whether the action instruction is related to making an avatar appear (13). If action determination program 9P determines that the action instruction is related to making an avatar appear, CPU 301 further activates avatar and environment action program 10P to perform an avatar appearance effect on data terminal 1 (14), and returns to step (2).

[0106] On the other hand, if the CPU 301 determines in step (13) that the content of the operation instruction does not relate to making an avatar appear, the CPU 301 determines whether the content of the operation instruction relates to adjusting the background of the display screen 13 (15), and if the CPU 301 determines that the content relates to adjusting the background, it performs calculations for the background image of the display screen 13 (16) and returns to step (2).

[0107] On the other hand, if the CPU 301 determines in step (15) that the content does not relate to background adjustment, the process returns to step (2).

[0108] In this embodiment, the background image includes a still image, a moving image, and a combination thereof, and includes image processing in real space and image processing in virtual space.

[0109] [Processing on the Server Device 100 Side] 11 and 12 are flowcharts showing the processing on the server device 100 side in the communication system according to this embodiment. Note that (21) to (33) indicate each step, which is realized by the CPU 113 loading a control program stored in the external memory 118 into the RAM 116 and executing it.

[0110] First, the CPU 113 determines whether the data terminal 1 is requesting an AI response to question information generated from the user's voice spoken into the microphone 311M (21). Note that this determination process can be skipped depending on the settings on the server device 100 side.

[0111] Here, if the CPU 113 determines that the request received from the data terminal 1 is requesting an AI response to the question information, the CPU 113 compares the setting data (including answer content setting data and operation instruction content setting data) accumulated in the setting database 114 with the setting data (including answer content setting data and operation instruction content setting data) stored in the external memory 118, and determines whether the setting data stored in the external memory 118 is the latest based on the assigned version information (22).

[0112] Here, if the CPU 113 determines that the setting data stored in the external memory 118 is the latest, the CPU 113 executes a process of reading the setting data stored in the setting database 114 and updating the setting data stored in the external memory 118 (23).

[0113] Next, if the CPU 113 determines that it has acquired answer content setting data from the answer content setting unit 103 (24), and further determines that it has acquired an operation instruction file related to the display screen 13 from the prompt setting database 107 (25), the CPU 113 executes a process of analyzing the question information received from the data terminal 1 (26).

[0114] Next, the CPU 113 executes a process (27) for creating optimal question information based on the operation instruction information acquired in step (25). Furthermore, the CPU 113 acquires operation instruction content setting data from the operation instruction content setting unit 104, and generates a process (28) for creating optimal operation instruction settings.

[0115] Next, the CPU 113 executes a process (29) to generate an optimum response setting to be sent to the data terminal 1 based on the results of the generation processes in steps (27) and (28).

[0116] Next, the CPU 113 activates the natural language generation unit 102 to generate an action instruction (avatar+background) in response to the natural language+question information (30).

[0117] Then, the CPU 113 determines whether the word-of-mouth index for similar responses exceeds a predetermined threshold (whether the satisfaction rate for the response exceeds, for example, 60%) (31).

[0118] Here, if the CPU 113 determines that the answer satisfaction level exceeds, for example, 60%, the CPU 113 responds to the data terminal 1 with the generated answer information, whether or not an avatar will appear, and background settings for the screen to be displayed (operation instructions for the question information) (32).

[0119] Next, CPU 113 determines whether a predetermined time has elapsed since the start of processing the previous request received from data terminal 1 (33). If CPU 113 determines that the predetermined time has elapsed, it terminates the processing, and if it determines that the predetermined time has not elapsed, it executes a process returning to step (21).

[0120] [Effects of the first embodiment] According to this embodiment, the user operating the data terminal 1 can listen to the audio output of answer information adapted to the question information input by voice, and can also visually confirm on the screen the response actions of an avatar that appears in conjunction with the audio output and the status of the background image. This makes it possible to freely provide a response service that combines natural language elements, auditory elements, and visual elements, going beyond audio output responses based on natural language processing.

[0121] Second Embodiment In the above embodiment, a system has been described in which the natural language generation unit 102 comprehensively processes the expected response processing, but it is also possible to construct a system in which multiple engines work together to process information individually for the fields (cultural, everyday, scientific, educational, etc.) and attributes (race, age, gender, cultural background, language used in everyday conversation, etc.) that the natural language generation unit 102 handles.

[0122] [Effects of the second embodiment] According to this embodiment, by using the most suitable natural language generation unit 102 in accordance with the question information received from the data terminal 1 and freely organizing combinations of natural language generation units 102, it is possible to freely support a system that can respond to specialized responses that are difficult to handle with a standardized response.

[0123] Third Embodiment In the above embodiment, a system has been described in which the natural language generation unit 102 comprehensively processes all response processing on one server device 100. However, it is also possible to configure a system in which different natural language generation units 102 are placed on multiple server devices 100, and multiple engines that individually process responses according to the field (cultural, everyday, scientific, educational, etc.) and attributes (race, age, gender, cultural background, language used in everyday conversation, etc.) to be responded to are linked together. For example, the technical terms (including abbreviations) to be used must be adjusted depending on whether the user making an inquiry about science and technology is an elementary school student or a university student.

[0124] [Effects of the third embodiment] According to this embodiment, by using the most suitable server device 100 depending on the question information received from the data terminal 1, it is possible to freely support a rapid system that can respond to specialized responses that are difficult to handle with a standardized response.

[0125] [Fourth embodiment] In the above embodiment, a system was described that controls screen display when performing voice input and voice output in a communication system in which a data terminal 1 operated by a user to connect to a network 21 via a specified communication medium and a server device 100 that provides a specified communication service in response to the launch of a voice-enabled application can communicate.However, the data terminal 1 may also be configured to have a function for determining whether a setting has been made to enable the AI-assisted mode, and a function for, if it is determined that a setting has been made to enable the AI-assisted mode, to acquire the AI-assisted service provided by the server device 100 in conjunction with the launch of the voice-enabled application and expand the communication tool function.

[0126] In this case, the AI ​​assistance mode can be configured to conform to the installed operating system.

[0127] [Effects of the fourth embodiment] According to this embodiment, the voice response function is automatically expanded by the generation AI that is optimal for the operating system installed in the data terminal 1, thereby significantly improving usability.

[0128] The disclosure of the present invention described above can be summarized at least as follows.

[0129] (1) A communication system in which a data terminal having a function for inputting and outputting language information can communicate with a server device that provides a predetermined communication service via a predetermined communication medium, wherein the data terminal has an input unit for inputting the predetermined language information and an output unit for outputting the predetermined language information. and, The language information input from the input unit is transferred to the server device, and generated by the server device. language Information and the foregoing language information Output of Displayed in conjunction with will be Screen Information and generated by the server device. an acquisition means for acquiring operation instruction information; languagea process of outputting information from the output unit, and a process of outputting information based on the operation instruction information acquired by the acquisition means. generating screen information for controlling the behavior and background of an avatar displayed on the display unit; and a control means for controlling the processing of outputting the information processed by the data terminal to a display unit. The aforementioned Get language information ,before A response setting for the recorded language information, a motion determination setting for determining the motion of the avatar to be displayed on the display unit, and a background setting for determining the background of the avatar to be displayed on the display unit. Ruta Meno back a language processing means for performing the answer setting, the action determination setting, the background determination setting, and the background determination setting set by the language processing means; Before and receiving a combination of the language information as a prompt, and displaying the language information to be responded to on the data terminal and the display unit. The screen information to be displayed is Avatar behavior settings and Background settings and the operation instruction information including and a generating means for generating the The operation instruction information includes operation settings and background settings of the avatar displayed on the display unit. It is characterized by:

[0132] (2) The control means language The device is characterized by controlling at least one of facial expressions, gestures, hand movements and / or three-dimensional movement of an avatar that appears on a screen displayed in the background of the display unit in synchronization with the output of information.

[0133] (3) The control means language The present invention is characterized in that the mode of the background image to be displayed on the display unit is switched and controlled in synchronization with the output of information.

[0134] (4) A service processing method for a communication system in which a data terminal having a function of inputting and outputting language information and a server device that provides a predetermined communication service can communicate with each other via a predetermined communication medium, wherein the data terminal has an input unit that inputs the predetermined language information and an output unit that outputs the predetermined language information, and transfers the language information input from the input unit to the server device and generates a communication service. language Information and the foregoing language information Output of Displayed in conjunction with will be Screen Information and generated by the server device. an acquisition step of acquiring operation instruction information; language a process of outputting information from the output unit; Steps Based on the operation instruction information acquired by generating screen information for controlling the behavior and background of an avatar displayed on the display unit; and a control step of controlling a process of outputting the processed information to a display unit, The aforementioned Get language information ,before Language information versus an action determination setting for determining the action of the avatar to be displayed on the display unit; and an action determination setting for determining the background of the avatar to be displayed on the display unit. Ruta Meno back a language processing step for performing a response setting, an action determination setting, a background determination setting, and a background judgment setting set by the language processing step; Before and receiving a combination of the language information as a prompt, and displaying the language information to be responded to on the data terminal and the display unit. The screen information to be displayed is Avatar behavior settings and Background settings and the operation instruction information including and a generating step of generating The operation instruction information includes operation settings and background settings of the avatar displayed on the display unit. It is characterized by: [Industrial Applicability]

[0136] In this embodiment, a case has been described where a system is configured with one data terminal and one server device in a one-to-one correspondence, but the present invention is naturally also applicable to a system in which a plurality of data terminals communicate with a server device.

[0137] Furthermore, although this embodiment assumes one language model, it is also applicable to a system that simultaneously targets multiple language models.

[0138] The AI ​​processing power in the current generative AI market is based on the assumption that the fields in which it can be introduced will be used in closed network environments. However, as generative AI development environments are established, it is expected that generative AI collaboration interfaces will be built that will enable collaboration with existing generative AI systems in order to evolve generative AI created by individuals. [Explanation of symbols]

[0139] 1 data terminal 100 Server device

Claims

1. A communication system in which a data terminal having a function of inputting and outputting language information and a server device providing a predetermined communication service can communicate with each other via a predetermined communication medium, The data terminal an input unit for inputting predetermined language information; an output unit that outputs predetermined language information; an acquisition means for transferring language information input from the input unit to the server device, and acquiring language information generated by the server device and operation instruction information that is screen information displayed in conjunction with the output of the language information and that is generated by the server device; a control means for controlling a process of outputting the language information acquired by the acquisition means from the output unit, and a process of generating screen information for controlling the action and background of an avatar displayed on a display unit based on the action instruction information acquired by the acquisition means, and outputting the screen information to the display unit; The server device a language processing means for acquiring the language information processed by the data terminal and performing a response setting for the language information, an action determination setting for determining an action of an avatar to be displayed on the display unit, and a background determination setting for determining a background of the avatar to be displayed on the display unit; a generating means for receiving a combination of the answer setting, the action determination setting, the background determination setting and the language information set by the language processing means as a prompt, and generating language information to be responded to by the data terminal and the action instruction information, which is screen information to be displayed on the display unit and includes an action setting and a background setting of an avatar; The communication system is characterized in that the operation instruction information includes operation settings and background settings of an avatar displayed on the display unit.

2. The communication system described in claim 1, characterized in that the control means controls at least one of facial expressions, gestures, hand movements and / or three-dimensional movement of an avatar that appears on a screen displayed in the background of the display unit in synchronization with the output of the linguistic information.

3. 2. The communication system according to claim 1, wherein the control means controls switching of the form of the background image to be displayed on the display unit in synchronization with the output of the linguistic information.

4. A service processing method for a communication system in which a data terminal having a function of inputting and outputting language information and a server device providing a predetermined communication service can communicate with each other via a predetermined communication medium, comprising: The data terminal An input unit that inputs predetermined language information and an output unit that outputs the predetermined language information, an acquisition step of transferring language information input from the input unit to the server device, and acquiring language information generated by the server device and operation instruction information that is screen information displayed in conjunction with the output of the language information and that is generated by the server device; a control step of controlling a process of outputting the language information acquired in the acquisition step from the output unit, and a process of generating screen information for controlling the action and background of an avatar displayed on a display unit based on the action instruction information acquired in the acquisition step, and outputting the screen information to the display unit; The server device a language processing step of acquiring the language information processed by the data terminal and performing a response setting for the language information, an action determination setting for determining an action of an avatar to be displayed on the display unit, and a background determination setting for determining a background of the avatar to be displayed on the display unit; a generating step of receiving a combination of the answer setting, the action determination setting, the background determination setting, and the language information set by the language processing step as a prompt, and generating language information to be responded to by the data terminal and the action instruction information, which is screen information to be displayed on the display unit and includes an action setting and a background setting of an avatar; The service processing method for a communication system, wherein the operation instruction information includes operation settings and background settings of an avatar displayed on the display unit.

Citation Information

Patent Citations

  • Interaction system and program

    JP2024176062A

  • Data processing apparatus, data processing method, and data processing program

    JP2025019910A

  • Action control system

    JP2025022833A

  • System

    JP2025044251A

  • Avatar Creation System and Communication System

    JP7577392B1