Terminal device and prediction method
By acquiring and learning the text information and emotion tags of users' speech on terminal devices, artificial intelligence is used to infer the content that users are going to say, solving the problem of tedious manual annotation and improving the accuracy and efficiency of emotion inference.
Patent Information
- Application Number
- CN202510601865.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-05-15
- Filing Date
- 2025-05-12
- Publication Date
- 2025-11-18
AI Technical Summary
In existing technologies, when using large-scale language models (LLM) to infer the content of a user's speech, manually labeling sentiment data is cumbersome and makes it difficult to accumulate a sufficient amount of data, resulting in poor inference results.
By launching a conversational application on the terminal device, the system continuously acquires the text information of the user's speech and stores relevant data when posting emotion tags. Artificial intelligence is then used to learn from this data to infer the content that the user wants to say, reducing the burden of manual annotation.
It improves the accuracy and efficiency of emotion inference, reduces the hassle of users manually inputting large amounts of data, and improves the inference technology for speech content.
Smart Images

Figure CN120977340A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to a terminal device and a prediction method. BACKGROUND
[0002] In the past, a technology of estimating the utterance content of a user according to an emotion has been known. For example, Patent Literature 1 discloses a technology of analyzing the biological information of a user and the SNS (Social Network Service) posting content and the like to estimate the emotion of the user.
[0003] PRIOR ART DOCUMENTS
[0004] PATENT LITERATURE
[0005] Patent Literature 1: Japanese Patent Application Publication No. 2023-142137
[0006] In recent years, artificial intelligence (AI) using a large-scale language model (LLM) is being actively utilized effectively. The LLM refers to a language model constructed by a large amount of data and a deep learning technique. For example, a user accumulates own utterance content in which own emotion is manually labeled as data. Then, the user causes the LLM to perform deep learning on the accumulated data. However, the manual labeling is troublesome. Therefore, it is difficult to accumulate a sufficient amount of data by the manual labeling. Therefore, there is room for improvement in the technology of estimating the utterance content of a user according to an emotion. SUMMARY
[0007] The present disclosure accomplished in view of the circumstances aims at improving the technology of estimating the utterance content of a user according to an emotion.
[0008] The terminal device of one embodiment of the present disclosure includes a database and a control portion that uses artificial intelligence capable of searching the database, wherein the control portion continuously acquires text information spoken by a user in a conversation using a conversation application program started on the terminal device, stores reference data including text information spoken by the user and a tag indicating the user's emotion in a predetermined period before and after the tag is posted whenever the tag is posted by the user in the database, extracts reference data including a tag corresponding to an emotion exposed by the user from the database using the artificial intelligence when the emotion exposed by the user is sensed in a case where the amount of the stored reference data exceeds a threshold, and learns the extracted reference data using the artificial intelligence, thereby inferring text information to be spoken by the user.
[0009] The prediction method of one embodiment of the present disclosure performs the following actions by a terminal device: continuously acquiring text information spoken by a user in a conversation using a conversation application program started on the terminal device; storing reference data including text information spoken by the user and a tag indicating the user's emotion in a predetermined period before and after the tag is posted whenever the tag is posted by the user in a database; extracting reference data including a tag corresponding to an emotion exposed by the user from the database using artificial intelligence when the emotion exposed by the user is sensed in a case where the amount of the stored reference data exceeds a threshold; and learning the extracted reference data using the artificial intelligence, thereby inferring text information to be spoken by the user.
[0010] Effects of Invention
[0011] According to one embodiment of the present disclosure, a technology of inferring a user's speech content from an emotion is improved. BRIEF DESCRIPTION OF DRAWINGS
[0012] Figure 1 FIG. 1 is a block diagram illustrating a schematic configuration example of a system of one embodiment of the present disclosure.
[0013] Figure 2 FIG. 2 is a flowchart illustrating an action example of a terminal device.
[0014] Figure 3 FIG. 3 is a schematic diagram illustrating a large-scale language model in which RAG is installed.
[0015] DETAILED DESCRIPTION
[0016] 1: system; 2: network; 3: user; 10: terminal device; 11: communication section; 12: input section; 13: output section; 14: storage section; 14A, 22A: database (DB); 15: control section; 20: server; 21: communication section; 22: storage section; 23: control section. DETAILED DESCRIPTION
[0017] Hereinafter, an embodiment of the present disclosure will be described.
[0018] (SUMMARY OF EMBODIMENT)
[0019] REFERENCE Figure 1 A summary of the system 1 of the embodiment of the present disclosure will be described. The system 1 is provided with a terminal device 10 and a server 20. The terminal device 10 and the server 20 are communicably connected with a network 2 including the Internet and a mobile communication network, and the like, for example.
[0020] The terminal device 10 is a personal computer (PC) or the like held by a user 3. However, the terminal device 10 is not limited to a personal computer, and can be any information processing terminal. The terminal device 10 is capable of communicating with the server 20 via the network 2.
[0021] The server 20 is a computer owned by an operator that provides an AI service, for example. The server 20 is capable of communicating with the terminal device 10 via the network 2.
[0022] First, a summary of the present embodiment will be described, and details will be described later. The terminal device 10 continuously acquires text information uttered by the user 3 in a session in which a session application program started on the device is used, and stores reference data including text information uttered by the user 3 and a tag indicating the emotion of the user 3 in a prescribed period before and after the time when the tag is posted by the user 3 every time the tag indicating the emotion of the user 3 is posted, and in a case where the amount of the stored reference data exceeds a threshold value, extracts reference data including a tag corresponding to the emotion exposed by the user 3 from the database 14A using artificial intelligence when the emotion exposed by the user 3 is sensed, and learns the extracted reference data using artificial intelligence, thereby inferring text information to be uttered by the user 3.
[0023] Thus, according to this embodiment, in order to infer the text information that user 3 intends to speak, it is necessary to publish a tag indicating user 3's emotion during a conversation using a chat application. Therefore, it is possible to prevent others from arbitrarily inferring the text information that user 3 intends to speak. Furthermore, the actions performed by user 3 are limited to publishing tags indicating user 3's emotion during a conversation using a chat application. Therefore, the hassle of user 3 manually inputting large amounts of reference data is reduced. Therefore, the technology for inferring the content of a user's speech based on emotion is improved.
[0024] Next, the components of System 1 will be described in detail.
[0025] (Composition of the terminal device)
[0026] like Figure 1 As shown, the terminal device 10 includes a communication unit 11, an input unit 12, an output unit 13, a storage unit 14, and a control unit 15.
[0027] The communication unit 11 is configured to include at least one communication module that can be connected to the network 2. The communication module may be, for example, a communication module corresponding to mobile communication standards such as LTE (Long Term Evolution), 4G (4th Generation), or 5G (5th Generation), a wired LAN (Local Area Network) standard, or a wireless LAN standard. Alternatively, the communication module may be a communication module corresponding to a short-range wireless communication standard such as Bluetooth (registered trademark). However, the communication module is not limited to these. The communication module can correspond to any communication standard. In this embodiment, the terminal device 10 communicates with the server 20 via the communication unit 11 and the network 2.
[0028] The input unit 12 is configured to include at least one input interface that allows input from the user 3. The input interface may be an input screen integrated with the display for the user 3 to input text data. Alternatively, the input interface may be a camera that captures an image of the user 3's face. Furthermore, the input interface may be a microphone that captures the user 3's speech. However, the input interface is not limited to these.
[0029] Output unit 13 is capable of outputting data. Output unit 13 is configured to include at least one output interface capable of outputting data. The output interface is, for example, a speaker and a display. The display is, for example, an LCD (Liquid Crystal Display) or an organic EL (Electro-Luminescent) display. However, the output interface is not limited to these.
[0030] Storage unit 14 includes one or more memories. These memories may be, for example, semiconductor memories, magnetic memories, or optical memories, but are not limited to these. Each memory included in storage unit 14 may function as a main storage device, an auxiliary storage device, or a cache memory. Storage unit 14 stores any information used for the operation of terminal device 10. For example, storage unit 14 may store system programs, application programs, embedded software, session applications, and AI engines. Furthermore, storage unit 14 includes a database 14A that stores reference data for AI learning. Alternatively, the information stored in storage unit 14 may be updated using information obtained from network 2 via communication unit 11.
[0031] The control unit 15 includes one or more processors, one or more programmable circuits, one or more dedicated circuits, or combinations thereof. The processor may be a general-purpose processor such as a CPU (Central Processing Unit) or a GPU (Graphics Processing Unit), or a dedicated processor for specific processing. However, the processor is not limited to these. The programmable circuit may be, for example, a FPGA (Field-Programmable Gate Array). However, the programmable circuit is not limited to FPGAs. The dedicated circuit may be, for example, an ASIC (Application Specific Integrated Circuit). However, the dedicated circuit is not limited to ASICs. The control unit 15 performs information processing related to the operation of the terminal device 10.
[0032] (Server configuration)
[0033] like Figure 1 As shown, server 20 includes a communication unit 21, a storage unit 22, and a control unit 23.
[0034] The communication unit 21 includes one or more communication interfaces connected to the network 2. These communication interfaces may correspond to, for example, mobile communication standards, wired LAN standards, or wireless LAN standards. However, the communication interface is not limited to these and may correspond to any communication standard. In this embodiment, the server 20 communicates with the terminal device 10 via the communication unit 21 and the network 2.
[0035] The storage section 22 includes one or more memories. Each memory included in the storage section 22 can function as a main storage device, an auxiliary storage device, or a cache memory, for example. The storage section 22 stores arbitrary information for the operation of the server 20. For example, the storage section 22 can store a system program, a database, an application program, an AI engine, and the like. Also, the storage section 22 can have a database 22A that stores reference data for AI learning and information on words of natural language processed by humans, and the like. It can also be that the information stored in the storage section 22 is updated using information acquired from the network 2 via the communication section 21, for example.
[0036] The control section 23 includes one or more processors, one or more programmable circuits, one or more dedicated circuits, or a combination thereof. The control section 23 performs information processing related to the operation of the server 20.
[0037] (Action flow of the terminal device)
[0038] Figure 2 is a flowchart showing an example of the operation of the terminal device 10. Referring to Figure 2 , the operation of the terminal device 10 of the present embodiment will be described. The present operation is related to the estimation of the text information ci (i is 1 to n) to be spoken by the user 3.
[0039] S101: The control section 15 continuously acquires the text information ci spoken by the user 3 in a conversation using a conversation application program started on the terminal device 10.
[0040] The conversation application program is LINE (registered trademark), for example. However, the conversation application program is not limited to LINE. The text information ci is a word, a phrase, or an article. However, the text information ci is not limited thereto.
[0041] S102: The control section 15 acquires the label emi (i is 1 to n) indicating the emotion of the user 3 posted by the user 3.
[0042] Among the classification methods of emotions, for example, there is the basic emotion theory that asserts that there are six basic emotions of joy, fear, surprise, disgust, anger, and sadness. Among other classification methods of emotions, there is Russell's circumplex model. Russell's circumplex model has two axes that indicate the degree of pleasure on the vertical axis and the degree of arousal on the horizontal axis. Russell's circumplex model indicates nineteen human emotions of surprise, happiness, boredom, melancholy, anger, fear, and the like on the circumference. In the present disclosure, the emotion of the user 3 is classified into the six basic emotions of joy, fear, surprise, disgust, anger, and sadness. However, the emotion of the user 3 is not limited thereto, and can be the nineteen emotions based on Russell's circumplex model.
[0043] The tag emi representing the emotion of the user 3 is, for example, a sticker used on LINE. In the present disclosure, the control section 15 classifies a tag corresponding to the emotion of joy as eml, a tag corresponding to the emotion of fear as em2, a tag corresponding to the emotion of surprise as em3, a tag corresponding to the emotion of disgust as em4, a tag corresponding to the emotion of anger as em5, and a tag corresponding to the emotion of sadness as em6. In this case, when the user 3 has any one of the six emotions in the conversation, the user 3 posts a sticker corresponding to any one of the tags eml to em6.
[0044] S103: The control section 15 generates reference data di (i is 1 to n) containing the text information ci uttered by the user 3 in the prescribed period T before and after the time ti (i is 1 to n) at which the tag emi is posted and the tag emi representing the emotion of the user 3, in the prescribed period T before and after the time ti (i is 1 to n) at which the tag emi is posted.
[0045] The prescribed period T is, for example, set to a period T (ti ± 180 seconds) of three minutes (180 seconds) before and after the time ti at which the tag emi is posted. However, the prescribed period T is not limited to three minutes and can be defined arbitrarily.
[0046] S104: The control section 15 stores the generated reference data di in the database 14A in order.
[0047] S105: The control section 15 confirms whether the amount of the stored reference data di exceeds the threshold value a. In the case where the amount of the reference data di exceeds the threshold value a, the control section 15 moves to S106 and performs the information processing. In the case where the amount of the reference data di is less than or equal to the threshold value a, the control section 15 moves to S101 and performs the information processing.
[0048] The control section 15 performs the generation of the reference data di until the amount of the reference data di referred to when learning using artificial intelligence (AI) in the database 14A reaches a required amount (threshold value a).
[0049] S106: The control section 15 senses the emotion exposed by the user 3.
[0050] The control section 15 can sense the emotion exposed by the user 3 by receiving the posting of the tag emi by the user 3 in the conversation using the conversation application. In addition, the control section 15 can also receive the exposure of the emotion based on the voice of the user 3 collected by a microphone in a web conference. Furthermore, the control section 15 can also receive the exposure of the emotion based on the expression of the user 3 captured by a camera in the web conference. However, the method of sensing the emotion exposed by the user 3 is not limited thereto.
[0051] S107: The control section 15 extracts, using artificial intelligence, reference data di containing a label emi corresponding to the emotion exposed by the user 3 from the database 14A.
[0052] The control section 15 selects a label emi corresponding to the emotion exposed by the user 3 from the label emi, the label em2, the label em3, the label em4, the label em5, and the label em6 corresponding to the emotions of joy, fear, surprise, disgust, anger, and sadness, respectively. As shown in Figure 3 In a case where the emotion exposed by the user 3 is an emotion of anger, the control section 15 selects the label em5 corresponding to the emotion of anger, and extracts, from the database 14A, reference data di containing the label em5.
[0053] S108: The control section 15 learns the extracted reference data di using artificial intelligence, thereby inferring the character information ci to be spoken by the user 3.
[0054] Artificial intelligence (AI) is a large-scale language model (LLM) installed with RAG (Retrieval Augmented Generation). The LLM is an artificial intelligence included in a part of AI, which can understand human language and conduct a conversation. The LLM learns a huge amount of character information (text data), generates a human who processes such natural language, and conducts understanding. The LLM has an advantage of being able to conduct a natural conversation as if having a conversation with a human. On the other hand, the LLM has a disadvantage of being able to process only learned information, i.e., being unable to process closed information.
[0055] RAG combines retrieval of external information in generation of character information (text data) by the LLM, thereby compensating for the disadvantage of the LLM of being unable to process closed information. RAG is translated as Retrieval Augmented Generation.
[0056] Figure 3 is a schematic diagram illustrating a large-scale language model installed with RAG. As shown in Figure 3 The RAG is able to perform retrieval of the database 14A, and the RAG conducts information processing of the following (i) to (vi).
[0057] (i) The RAG receives, from the control section 15, the label em5 indicating the emotion of anger issued by the user 3.
[0058] (ii) The RAG accesses the database 14A, and retrieves reference data containing the label em5.
[0059] (iii) The RAG extracts, from the database 14A, reference data d3 containing the label em5. InFigure 3 In the example of FIG. 10, only the reference data d3 is extracted. However, in practice, a large number of 1000 pieces, 10000 pieces, and so on of the reference data can be extracted.
[0060] (iv) The RAG passes the extracted reference data d3 to the LLM.
[0061] (v) The RAG receives the character information c8 estimated by the LLM.
[0062] (vi) The RAG presents the character information c8 to the control section 15.
[0063] The RAG can also be configured to perform a search on other databases (for example, the database 22A shown in FIG. 12) that incorporate human-processed natural language vocabularies, in addition to the search on the database 14A. Figure 1 In the example of FIG. 10, the location of the other database that incorporates human-processed natural language vocabularies is set to the server 20. However, the location of the other database can be any location that can communicate via the network 2. Figure 1
[0064] The control section 15 can also acquire an LLM engine that installs the RAG from the server 20 owned by an operator that provides an AI service via the network 2. The engine is a collection of programs and algorithms that are the core of artificial intelligence. On the other hand, the control section 15 can also use a cloud service of the LLM that installs the RAG provided by the server 20 via the network 2. The cloud service refers to a service that provides desktop virtualization, shared disks, and other hardware and infrastructure functions via the Internet.
[0065] S109: The control section 15 confirms with the user 3 whether the estimation result of the character information ci that the user 3 wants to speak is correct or not. In the case where the estimation result is correct, the control section 15 moves to S110 and performs information processing. In the case where the estimation result is not correct, the control section 15 moves to S111 and performs information processing.
[0066] S110: In the case where the estimation result is correct, the control section 15 further stores the reference data di that includes the estimated character information ci and the label emi corresponding to the emotion expressed by the user 3 in the database 14A.
[0067] The reference data that includes the character information ci determined by the user 3 as the correct estimation result is further stored in the database 14A. By using this highly accurate reference data for learning by the LLM, the accuracy of the estimation of the character information that the user 3 wants to speak while having a specific emotion will improve.
[0068] S111: In a case where the result of the presumption is incorrect, the control section 15 deletes the presumed character information ci.
[0069] S112: The control section 15 determines whether or not to continue the information processing. In a case of continuation, the control section 15 moves to S106, and executes the information processing. In a case of non-continuation, the control section 15 ends the information processing.
[0070] Note that the control section 15 can also accumulate, in the database 14A, reference data di' including, in addition to the character information uttered by the user 3 and the tag emi indicating the emotion of the user 3 within the prescribed period T, data of the number of conversation partners with whom the user 3 converses within the prescribed period T or data of attributes of the conversation partners. In this way, it is possible to presume the character information ci that the user 3 intends to utter while having a specific emotion, in accordance with each conversation situation in which the number of conversation partners or the attributes of the conversation partners (family, friends, superiors in a company, customers, etc.) are different.
[0071] As described above, the terminal device 10 of the present embodiment continuously acquires the character information uttered by the user 3 in a conversation in which a conversation application program started on the device is used, stores, every time a tag indicating the emotion of the user 3 is posted, reference data including the character information uttered by the user 3 and the tag indicating the emotion of the user 3 within a prescribed period before and after the time at which the tag is posted in the database 14A, and in a case where the amount of the stored reference data exceeds a threshold value, extracts, using artificial intelligence, reference data including a tag corresponding to the emotion exposed by the user 3 from the database 14A at the time when the emotion exposed by the user 3 is sensed, and learns the extracted reference data using the artificial intelligence, thereby presuming the character information that the user 3 intends to utter.
[0072] According to this configuration, in order to presume the character information that the user 3 intends to utter, the posting of a tag indicating the emotion of the user 3 is required in a conversation in which a conversation application program is used. Therefore, it is possible to avoid the character information that the user 3 intends to utter from being presumed arbitrarily by others. Further, the operation performed by the user 3 is only the posting of a tag indicating the emotion of the user 3 in a conversation in which a conversation application program is used. Therefore, the user 3 is relieved of the trouble of manually inputting a large amount of reference data. Thus, the technology of presuming the utterance content of the user in accordance with the emotion is improved.
[0073] The present disclosure is described based on the drawings and the embodiments, but note that the skilled person can make various modifications and changes based on the present disclosure. Therefore, note that these modifications and changes are included in the scope of the present disclosure. For example, the functions and the like included in each constituent or each step and the like can be reconfigured in a manner that does not contradict in logic, a plurality of constituents or steps and the like can be combined into one, or a plurality of constituents or steps and the like can be divided.
[0074] For example, in the above-described embodiments, it can also be an embodiment in which the configuration and the operation of the terminal device 10 are dispersed among a plurality of computers that can communicate with each other. Further, for example, it can also be an embodiment in which a part or all of the constituent elements of the terminal device 10 are provided to the server 20. For example, it can also be an embodiment in which the terminal device 10 entrusts the estimation of the text information in which the user 3 is to speak, which is performed by the control section 15, to the server 20 on the basis of the transmission of the reference data to the server 20.
[0075] Further, for example, it can also be an embodiment in which a general-purpose computer functions as the terminal device 10 of the above-described embodiments. Specifically, a program that describes the processing content of each function of the terminal device 10 of the above-described embodiments is stored in the memory of the general-purpose computer, and the program is read out and executed by the processor. Therefore, the present disclosure can also be realized as a program executable by the processor or a non-transitory computer-readable medium that stores the program.
Claims
1. A terminal device comprising: Database; and The control unit uses artificial intelligence capable of retrieving data from the database. The control unit continuously acquires text information spoken by the user during a session using a conversational application launched on the terminal device. Whenever the user posts a tag indicating the user's emotion, it stores reference data in the database containing text information spoken by the user during a predetermined period before and after the time the tag was posted, along with the tag indicating the user's emotion. If the amount of stored reference data exceeds a threshold, when the user's expressed emotion is sensed, the control unit uses artificial intelligence to extract reference data containing tags corresponding to the user's expressed emotion from the database. The control unit then uses artificial intelligence to learn from the extracted reference data, thereby inferring the text information the user is about to speak.
2. The terminal device according to claim 1, wherein, The artificial intelligence mentioned is a large-scale language model equipped with retrieval-enhanced RAG generation. The RAG can retrieve data from the database, extract reference data stored in the database, and pass the extracted reference data to the large-scale language model.
3. The terminal device according to claim 1, wherein, The control unit confirms with the user whether the estimated text information the user wants to speak is correct. If the estimated text information is correct, the control unit further stores reference data, including the estimated text information and the tag corresponding to the emotion expressed by the user, in the database.
4. The terminal device according to claim 1, wherein, In addition to the text information spoken by the user during the specified period and the tags indicating the user's emotions, the reference data also includes data on the number of conversation partners the user had during the specified period or data on the attributes of the conversation partners.
5. A prediction method, wherein the following actions are performed via a terminal device: Continuously acquire text information spoken by the user in a session using a session application launched on the terminal device; Whenever a user posts a tag indicating the user's emotion, the database stores the text information spoken by the user during a specified period before and after the time the tag was posted, along with reference data for the tag indicating the user's emotion. If the amount of stored reference data exceeds a threshold, when an emotion expressed by the user is sensed, artificial intelligence is used to extract reference data from the database containing tags corresponding to the emotion expressed by the user. as well as The artificial intelligence is used to learn from the extracted reference data, thereby inferring the text information that the user wants to say.
Citation Information
Patent Citations
Emotion adjustment support device, emotion adjustment support method, program, and recording medium
JP2023142137A