Service information providing system
The service information provision system addresses the limitation of one-to-one relationships in generative AI by using multiple media generation models to present output results in diverse formats, enhancing user experience through varied and sensory responses.
Patent Information
- Application Number
- JP2024074137
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-04-30
- Publication Date
- 2025-11-12
AI Technical Summary
Existing generative AI systems, such as ChatGPT and smart speakers, provide one-to-one relationships between questions and answers, limiting the presentation of results to a single medium, which does not facilitate diverse perspectives or sensory experiences.
A service information provision system utilizing multiple media generation models (language, image, audio, video, and music) on the server side to generate and present output results in two or more types of media from a single input, allowing for diverse perspectives and sensory experiences.
Enables users to receive output results that are more profound and sensory by providing responses from different perspectives across multiple media formats, enhancing the user experience in generating ideas and inspirational hints.
Smart Images

Figure 2025169106000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a system for providing service information to users via a network. [Background technology]
[0002] In recent years, many systems that utilize artificial intelligence have been provided, such as language model chat systems that support multiple languages such as English and Japanese (e.g., non-patent document 1) and image generation systems that can generate images by entering keywords (e.g., non-patent document 2).These systems are now accessible from browsers and are becoming available to general users as well.
[0003] Providing generative AI via browsers has created an environment that makes it easy for general users to use. Following this trend, enterprises and companies are also creating an environment in which generative AI can be used for internal and commercial purposes. Patent Document 1 discloses technology that uses a generative AI language model to input a spoken conversation and accurately generate a summary. This technology converts the audio data, which is the audio source of the spoken conversation, into text, and when the same string of characters, such as "yes," is converted into text, there are cases where the person is angry and cases where they are not angry. This technology discloses a method for generating a summary accurately by analyzing the emotional state of the voice data.
[0004] Patent Document 2 discloses a technology that generates images with a novel style by inputting keywords for the image a user wants to create into an image generation model of a generation AI.The method shows a method for generating a novel style of art by using N-dimensional feature data of an image generated from keywords input by the user and N-dimensional feature data of accumulated images, and by taking into account feature data in sparse areas in the distribution of certain feature data.
[0005] Furthermore, Patent Document 3 shows that as a method for efficiently improving the recognition accuracy of an automatic learning device for recognition results in a voice chatbot, the AI chatbot repeats the recognition results of the AI chatbot to determine whether the recognized content is correct or incorrect based on the speaker's response, and learns based on the difference in the recognition results of things that were initially judged as "incorrect" but later judged as "correct," thereby improving recognition accuracy. [Prior art documents] [Patent documents]
[0006] [Patent Document 1] Patent No. 6513869 [Patent Document 2] Patent No. 7270894 [Patent Document 3] Patent Publication No. 2021-56392 [Non-patent literature]
[0007] [Non-Patent Document 1] OpenAI ChatGPT Searched on April 12, 2024 Internet URL: https: / / chat.openai.com / [Non-patent document 2] Stability AI SDXL Searched on April 12, 2024. Internet URL: https: / / ja.stability.ai / stable-diffusion [Non-patent document 3] OpenAI Sora Searched on April 12, 2024 Internet URL: https: / / www.ciciai.com / Summary of the Invention [Problem to be solved by the invention]
[0008] However, with systems like ChatGPT, when a question is entered as text in a browser, the generated answer is added to the question. Also, with AI chatbots equipped with AI and communication functions known as smart speakers, the speaker can converse with the chatbot using its microphone and speaker, and the speaker can respond to the question. Such smart speakers can handle everything from general questions like checking the weather or stock prices to tasks like listening to the news, setting a timer, or playing music. In ChatGPT, these conversations are handled by text media in response to text input in the browser, while in the case of smart speakers, responses are generated by voice in response to the speaker's speech (voice). Here, the dialogue is text or voice, and the questions and answers are the same text or voice media, creating a one-to-one relationship in which one answer is displayed for each question.
[0009] Unlike ChatGPT and smart speakers, the technologies disclosed in the above patent documents convert the media of questions and answers from voice to text, text to image, etc., such as converting voice to text and creating a summary as shown in Patent Document 1, and generating a novel style image from keywords entered by the user as shown in Patent Document 2. However, the output results for those questions are in a one-to-one relationship, the same as ChatGPT and smart speakers.
[0010] ChatGPT and smart speakers use generative AI models to generate answers to questions. Examples include questions like "Please explain dementia," "What's the weather like today?", "Summarize the following sentence," "You are a professional writer. Please create a catchy slogan for the following product that resonates with your target audience," and "Classify the text as neutral, negative, or positive. Text: This food was okay." For image generation models, detailed questions are used to obtain a single answer, such as "Children playing in the park in the evening," "Top quality, beautiful woman, long hair, brown hair, wearing a camisole and skirt, shy smile, sunlight lighting, cityscape background." Again, the relationship between question and answer is one-to-one: in language generation models, input text is output text, and in image generation models, input text is output image.
[0011] In addition, in the case of a language generation model, a request such as "Write five poems with a spring theme" or an image generation model, "You are an excellent designer. Please create five packaging design ideas for new product A targeted at XX" will result in five responses (five poems and five designs), which is a one-to-five relationship in terms of the number of responses to the number of inputs. However, in terms of the type of input and output media, there is a one-to-one relationship between input text and output text, and between input text and output image in the image generation model.
[0012] These generative AIs are used as tools to achieve goals such as answering questions with a single purpose, such as the weather or the time, creating package design ideas, determining whether an input sentence is negative, positive, or neutral, generating a desired image, or writing a poem.The relationship between input and output is always one-to-one, and they do not provide results from different perspectives by presenting them across multiple media in response to a user's search for ideas or inspirational hints. [Means for solving the problem]
[0013] In order to solve the above problem, the present invention provides a service information provision system that is characterized by having at least two or more types of media generation models among a plurality of media generation models, such as a language generation model capable of generating language, an image generation model capable of generating images, an audio generation model capable of generating sound, a music generation model capable of generating music, and a video generation model capable of generating video, on the server side that generates the output results, and presenting the user with generation results obtained from two or more types of media generation models from input information of one type of media sent from the user, thereby presenting the user with results from multiple perspectives. [Effects of the Invention]
[0014] According to the service information providing system of the present invention, by presenting the user with output results in two or more types of media and generative models as response results from input data in a single medium from the user, the user can obtain output results that are more sensory and profound by providing results from different perspectives. [Brief explanation of the drawings]
[0015] [Figure 1] FIG. 1 is a diagram illustrating a system configuration according to an embodiment of the present invention. [Figure 2] FIG. 2 is a diagram showing a system configuration according to an embodiment of the present invention, which is configured differently from that shown in FIG. [Figure 3] FIG. 10 is a diagram showing an input screen of a user terminal according to an embodiment of the present invention. [Figure 4] FIG. 10 is a diagram showing an output screen of a user terminal according to an embodiment of the present invention. [Figure 5] FIG. 5 is a diagram showing an output screen of an embodiment having a different configuration from that of FIG. 4. [Figure 6] FIG. 7 is a diagram showing an output screen of an embodiment having a different configuration from those of FIGS. 5 and 6. [Figure 7] FIG. 1 is a diagram showing an outline of the overall processing flow of an embodiment of the present invention. [Figure 8] FIG. 8 is a diagram showing an outline of a processing flow according to an embodiment having a different configuration from that of FIG. 7. [Figure 9] FIG. 8 is a diagram showing details of the processing flow on the server side in FIG. 7. [Figure 10] FIG. 1 is a diagram illustrating an embodiment of the present invention. [Figure 11] FIG. 1 is a diagram illustrating an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0016] An embodiment of the present invention will be described with reference to the drawings. Note that the embodiment described below is an example of a means for realizing the present invention, and the configuration of the device to which the present invention is applied will vary depending on the service information provided to the user, and is not necessarily limited to the embodiment described below.
[0017] (System Configuration Overview of This Embodiment) 1 and 2 show an example of the system configuration of a service information providing system according to this embodiment. In Fig. 1, three types of components, namely, a user terminal 1, a server 2, and four generative models 3, are connected via a network. In contrast, in Fig. 2, the server 2 and generative models 3, which were separate and independent in Fig. 1, are configured such that the generative models 3 are included within the server. In these configurations, whether the server 2 and the generative models 3 are included in the same system or not, processing is performed via data communication, so there is a difference in configuration but no difference in functionality.
[0018] (Configuration of input screen of user terminal in this embodiment) 3 and 4 show examples of the configuration of the data input screen 5 of the user terminal 1 of the service information providing system according to this embodiment. FIG. 3 shows an example of the data input screen 5 of the user terminal when the data to be transmitted to the server 2 is text, and FIG. 4 shows an example of the data input screen 5 of the user terminal when the data to be transmitted to the server 2 is sound, image, or video. In FIG. 3, the data input screen 5 of the user terminal is provided with a text input field 6 and a send button 7. When input into the text data input field 6 is completed, the input data can be transmitted to the server 2 by clicking or tapping the send button 7. Also, in FIG. 4, the input data to be transmitted, such as a sound file, image file, or video file, can be selected and then transmitted to the server 2 by clicking or tapping the send button 7. A configuration of at least these elements is required to transmit data from the user terminal 1 to the server 2.
[0019] (Configuration of the output screen of the user terminal in this embodiment) 5 and 6 show examples of the configuration of the display screen 9 of the output result of the user terminal of the service information providing system according to this embodiment. Fig. 5 shows an example in which the media of the generated result returned from the server 2 is text media and image media. Fig. 6 shows an example in which the media of the generated result returned from the server 2 is text media and either audio media or video media.
[0020] FIG. 5 shows an example of a method for presenting an output result when the generation result is text media and image media. Generated text 10, which is the generation result for text media, and generated image 11, which is the generation result for image media, are displayed on output result display screen 9 of the user terminal. FIG. 6 also shows an example of a method for presenting an output result when the generation result is two media, that is, text media and either audio media or video media. Generated text 10, which is the generation result for text, and a generation result play button 12 for playing the generation result for audio media or video media are provided on output result display screen 9 of the user terminal. Generated text 10, which is the generation result for text media, is already displayed on output result display screen 9 of the user terminal, but the generation result for audio media or video media is not displayed. By operating generation result play button 12, the sound of the generation result for audio media is output from a speaker, and for the generation result for video media, a video media playback screen is displayed and a generated video (not shown) is played. The generated video may be a static generated video 13, such as generated image 11 shown on display screen 9 of the output result of the user terminal in Fig. 5, and may be played back by clicking or tapping on the screen of generated video 13. At least such a configuration of elements is required to display and check the output result obtained from server 2 on user terminal 1.
[0021] Note that the generated text 10, generated image 11, generated result play button 12, etc. do not need to be arranged in the same location or size as in the example of this embodiment as long as they are displayed on the output result display screen 9 of the user terminal, and even if the generated text 10, generated image 11, generated result play button 12, etc. are not displayed, they will still meet the requirements as long as they can be displayed or played by some means such as audio. Here, we have shown a method in which three media are displayed or played as output results, but three or more media can also be presented in the same way.
[0022] (Overview of the overall processing flow of the service information provision system) The overall processing flow when using the service information providing system according to the present invention will be described below. Figures 7 and 8 are flowcharts showing an outline of the overall processing flow in the service information providing system according to the present invention.
[0023] 7 and 8 are flowcharts showing an example of an outline of the processing flow of the entire system in the service information providing system according to the present invention. The flow shown in Fig. 7 is configured to reduce the amount of processing at the user terminal 1. The processing 15 on the server side in Fig. 7 will be explained using a diagram that combines the processing flow and configuration in Fig. 8. The focus is on the processing at server 2, which increases the amount of processing at server 2.
[0024] When using a service provided on a website or an app installed on a smartphone, not limited to the present invention, the user must first open the website where the service is provided or run the app, and then open the website or run the app on the user terminal where the service information providing system can be used. This will display a data entry screen 5 on the user terminal. The processing flow from this state will now be described.
[0025] First, the user inputs data into the text input field 6 displayed on the data input screen 5 (step S10). After inputting the data, the user clicks the send button 7 on the data input screen 5 (step S11). This operation causes the data entered into the text input field 6 to be sent to the server 2 (step S12). This is the process 14 on the user terminal.
[0026] From here, processing 15 on the server side begins. The server 2 receives the data 16 sent from the user terminal 1 (step S13). Next, using the received data from the user terminal 1 and the base 18 of input data for the generative model stored in the storage device 17, input data 19 for the generative model that reflects the user's requests is created (step S14). Note that the input data 19 for the generative model can only be created by using both the base 18 of input data for the generative model and the data 16 received from the user terminal 1. As shown in FIG. 9, the media to be generated differs depending on the generative model used, and therefore the way the input data is written differs. Therefore, it is necessary to create a base 18 of input data for each generative model. Here, three types of input data 19 for the generative model are illustrated to illustrate the creation of bases 18 for multiple input data. As for the base generative models for the data 16 input data received from the user terminal 1, there are many generative models available, such as large-scale language models for language generation, models such as diffusion models and deep generative models for image generation, video generation models for video generation, and diffusion models, which are also used for image generation, and music generation models for music generation, for example. FIG. 9 shows the use of three types of generative models: generative model a, generative model b, and generative model c. For example, these generative models may be different models, such as a language generation model, an image generation model, and a video generation model. Now that the input data for the generative models is prepared, the generative models are executed 20 using the respective generative model input data 19 (step S15). Execution results 21 are obtained from the execution 20 of the generative models. In FIG. 9, execution 20 is performed using three types of generative models, so three types of execution results 21 are obtained (step S16). Now that the execution results 21 have been obtained, these three types of execution results 21 are sent 22 to the user terminal 1 (step S17).
[0027] From here on, the user terminal performs processing 14, receiving the execution results sent from the server 2 (step S18). Finally, the received execution results are presented to the user terminal using a presentation method corresponding to each media (step S19). This completes the processing requested by the user.
[0028] Fig. 8 is a flowchart showing an outline of the processing flow of the entire system in the service information providing system according to the present invention. The flow shown in Fig. 8 is configured to reduce the processing in the user terminal 1 shown in Fig. 7.
[0029] The differences between Fig. 7 and Fig. 8 will be explained. In Fig. 8, step S15 in Fig. 7, which was processed on the server side 15, is processed in step S52 of the user terminal shown in Fig. 8. Accordingly, base 18 of the input data stored in storage device 17 shown in Fig. 9 is stored in user terminal 1.
[0030] First, the user inputs data into the text input field 6 displayed on the data input screen 5 (step S50). After completing the data input, the user clicks the send button 7 on the data input screen 5 (step S51). Up to this point, the process is the same as the processing flow in FIG. 7, but before the data is sent, input data 19 for each generative model is created (step S52). Once the creation of the input data 19 for the generative model is complete, the input data 19 for the generative model is sent to the server 2 (step S52). Up to this point, the processing 14 on the user terminal is completed.
[0031] From this point on, processing 15 takes place on the server side. The server 2 receives the data sent from the user terminal 1 (step S54). Next, since the input data 19 for the generative model is completed, this is used to execute the generative model 20 (step S55). The subsequent processing flow is the same as in Figure 7, with the execution results being obtained (step S56) and the output results being sent to the user terminal (step S57).
[0032] From this point on, processing 14 is performed on the user terminal. The execution results sent from the server are received (step S58). Finally, the received execution results are presented on the user terminal using a presentation method corresponding to each media (step S59). This completes the processing requested by the user.
[0033] (Application example of embodiment) An application example of the embodiment is shown using FIGS. 10 and 11. FIG. 10 shows an example in which the user's input media is speech and the generated output media are two types of media: text and music. FIG. 10 is composed of three diagrams: FIG. 10(a) shows a data input screen 5' on the user terminal, FIG. 10(b) shows the contents of processing 15' on the server side, and FIG. 10(c) shows a display screen 9' of the output results on the user terminal. FIG. 11 shows an example in which the user's input media is text and the generated output media are three types of media: text, images, and music. Like FIG. 10, FIG. 11 is also composed of three diagrams: FIG. 11(a) shows a data input screen 5'' on the user terminal, FIG. 11(b) shows the contents of processing 15'' on the server side, and FIG. 11(c) shows a display screen 9'' of the output results on the user terminal.
[0034] An application example of the embodiment shown in FIG. 10 will be described in which the user's input media is voice and the generated output media are two types of media: text and music. First, the process up to the time the user sends data will be described using FIG. 10(a). FIG. 10(a) shows an input screen 5' of the user terminal, on which a recording start button 23, a recording stop button 24, and a send button 7' are arranged. Since the user's input media is voice, the process begins by recording the voice. When preparations for recording are complete, the user taps or clicks the recording start button 23, and when recording is complete, the user taps or clicks the recording stop button 24 to perform the recording operation. Note that if the recording is not successful, the user may record again. When recording is complete, the user taps or clicks the send button 7', and the input data is sent to the server 2.
[0035] The processing after receiving data 16' from the user terminal 1 will be described using Figure 10(b). Because the received data is audio media data, the audio data is converted to text 25 to create input data 19 for the generation model. Using the base 18a' of language generation model input data and the base 18b' of music generation model input data stored in the storage device 17' and the converted text input data, input data a19a' for the language generation model and input data b19b' for the music generation model are created, respectively, and are executed 20a in language generation model a and executed 20b in music generation model b, respectively. By executing each generation model, it is possible to obtain an execution result a21a', which is the result of execution in the language model, and an execution result b21b', which is the result of execution in the music model. The obtained execution results are then sent 22' to the user terminal. At this point, processing 15 on the server side ends.
[0036] Next, processing 14 on the user terminal will be explained using Figure 10(c). First, the execution results sent from server 2 are received, and the received execution results are presented on the user terminal using a presentation method corresponding to each media. Here, since there are two types of media presented to the user, text and music, the display screen 9' of the output results on the user terminal in Figure 10(c) is provided with generated text 10' and a music play button 26, which is an operation button for listening to the generated music. The text media results are presented on the screen, and the music media results can be heard on the user terminal 1 by tapping the music play button 26. This completes all processing.
[0037] An application example of the embodiment shown in FIG. 11 will be described in which the user's input media is text and the generated output media are three types of media: text, images, and music. First, the process up to the point where the user sends data will be described with reference to FIG. 11(a). FIG. 11(a) shows an input screen 5'' of the user terminal, which has an input field 6'' for entering text and a send button 7''. Since the user's input media is text, the entered data is sent from the user terminal 1 to the server 2 by tapping or clicking the send button 7'' when the user has finished entering the text.
[0038] The processing performed after receiving 16'' of data from the user terminal 1 will be described with reference to FIG. 11(b). Using the received text data and the input data for three types of generation models, namely, base 18a'' of language generation model input data, base 18b'' of image generation model input data, and base 18c'' of music generation model input data stored in storage device 17'', and the received text data, input data a19a'' for the language generation model, input data b19b'' for the image generation model, and input data c19c'' for the music generation model are created, and the created data are used to perform execution 20a'' on language generation model a, execution 20b'' on image generation model b, and execution 20c'' on music generation model c. Execution results a21a'', b21b'', and c21'', which are the execution results of each generation model, can be obtained. These execution results are then sent 22'' to the user terminal. At this point, processing 15 on the server side ends.
[0039] Next, processing 14 on the user terminal will be explained using Figure 11(c). First, the execution results sent from server 2 are received, and the received execution results are presented on the user terminal using a presentation method corresponding to each media. Here, three types of media are presented to the user: text, image, and music. The output result display screen 9'' on the user terminal in Figure 10(c) is provided with generated text 10'', generated image 11'', and a music play button 26'' which is an operation button for listening to the generated music. The text media output result 10'' and the image media output result 11'' are displayed on the screen. For music media results, tapping the music play button 26 at the bottom of the screen will start playback, allowing you to check the generated music media results. This completes all processing.
[0040] Here, data input and output at the user terminal 1 in the embodiment shown in FIG. 11 will be described with reference to FIGS. 11(a) and 11(c). Here, it is assumed that the user is using the service information providing system with the thought, "What state am I in?" For data input at the user terminal 1, as shown in FIG. 11(a), an input field 6'' for inputting text is provided on the input screen 5'', and the user enters text into this input field 6''. In FIG. 11(a), instructions for the input field are written as "Please enter about three words to describe how you are feeling right now." For example, words such as "sad, blue, lonely" are entered.
[0041] Regarding data output on the user device 1, as shown in Figure 11(c), the output screen 9'' displays generated text 10'', generated image 11'', and music play button 26''. These represent the locations and means for presenting the execution results of the language generation model, image generation model, and music generation model, respectively. The execution results of the language generation model are displayed as generated text 10'', the execution results of the image generation model are displayed as generated image 11'', and the execution results of the music generation model are generated by tapping the music play button 26''. An example of generated text 10'' might include a psychological analysis such as "You seem to be feeling disappointment, frustration, or loss..." The generated image 11'' might depict a picture mainly in deep blue or green, and the generated music might be a song in a minor key with a slow tempo. Presenting the results to the user on different devices, such as text, images, and music, allows for different perspectives and provides information that resonates with the senses. [Industrial Applicability]
[0042] In the service information providing system of the present invention, when the purpose of using the generation AI is not to obtain specific results such as creating a summary of a text or preparing presentation materials, but rather to obtain ideas, the relationship between the user's inquiry and the response from the generation AI is usually one-to-one. However, when searching for ideas, information from different perspectives is often useful. By using the service information providing system of the present invention, it is possible to obtain responses in a one-to-many relationship, and thus the service providing system can be used advantageously by obtaining responses in multiple media that provide the user with idea hints from different perspectives. [Explanation of symbols]
[0043] 1. User terminal 2 Server 3 Server and Generative Model 4. Generative Model 5, 5', 5'' data entry screen 6, 6'' text field 7, 7', 7'' Send button 8 File selection button 9, 9', 9'' output result display screen 10, 10', 10'' generated text 11, 11'' generated image 12 Playback button for generated results 13. Generated Images 14 Processing on the user terminal 15, 15', 15'' Server-side processing 16, 16', 16'' Receiving data from user terminal 17, 17', 17'' storage device 18, 18a, 18a', 18a'', 18b, 18b', 18b'', 18c, 18c'' Input data base 19, 19a, 19a', 19a'', 19b, 19b', 19b'', 19c, 19c'' Input data for generative models 20, 20a, 20a', 20a'', 20b, 20b', 20b'', 20c, 20c'' Running the generative model 21, 21a, 21a', 21a'', 21b, 21b', 21b'', 21c, 21c'' generated results 22, 22', 22'' Sending execution results to the user terminal 23 Recording start button 24 Recording stop button 25 Converting audio data to text 26, 26'' music play button 30 Network
Claims
1. a server having a receiving function for connecting to a user terminal via a network and receiving data, a transmitting function for transmitting data to the user terminal, and an information processing function; the server has at least two or more types of generation models, such as a language generation model for language generation, an image generation model for image generation, a sound generation model for sound generation, a music generation model for music generation, and a video generation model for video generation; The user terminal or the server has an information processing function for creating input data for a generative model from one type of data, such as text, image, sound, music, video, etc., input to the user terminal and a stored input base for the generative model, The system has a function of transmitting input data for each of two or more generation models among the language generation model, the image generation model, the sound generation model, the music generation model, and the video generation model as a question to each generation model, receiving answers from each generation model, transmitting the results to a user terminal, and presenting the generation results to the user terminal, This service information provision system is characterized by helping users to imagine and come up with ideas by providing results from multiple different perspectives when they utilize generative models to seek ideas and inspiration.
2. The service information providing system of claim 1, characterized in that the user terminal has an information processing function for creating input data for a generative model from user input data processed by the server and a stored input base for the generative model, and information processing is performed using the generative model stored in the user terminal.
3. 3. The service information providing system according to claim 1, wherein the input data from the user is one of text, image and voice, and the content does not include any clear instruction.
4. A service information providing system as described in claim 1 or 2, characterized in that the content of the user's input data is a plurality of words expressing emotions or sensations, and these words alone are insufficient as input data for the generative model.
5. The service information providing system according to claim 1, characterized in that the content transmitted from the user terminal to the server is a plurality of words expressing emotions or sensations, thereby facilitating the creation of input templates for different generation models.
6. The system includes a screen having an input field in which text can be input to a user terminal and a button for transmitting the input text when the user has completed the text input, and the server receives and processes the text data transmitted from the user terminal, the server has the language generation model that understands a language and generates a language, and the image generation model that understands a language and generates an image from the understood language, The service information providing system of claim 1, characterized in that input data for the language generation model is created from the text data entered by the user and the generation model input material stored on the server, and the input data is used to execute the language generation model to obtain a generation result, and input data is created from the text data entered by the user and the image generation model input material stored on the server to create input data and the image generation model to obtain a generation result, which are returned to the user's terminal and presented on the user's terminal.
Citation Information
Patent Citations
Automatic learning device and method for recognition result of voice chatbot, and computer program and recording medium
JP2021056392A
Dialogue summary generation device, dialogue summary generation method, and program
JP6513869B1
Identifying digital data with a new style
JP7270894B1