Question and answer processing method, apparatus, device, storage medium, and program product

By using voice emotion recognition and knowledge base rewriting technologies, the answer format is adjusted according to the user's emotional state, solving the problem of poor user experience in intelligent question-answering systems and achieving a higher quality interactive experience.

CN118692450BActive Publication Date: 2026-07-31BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING BAIDU NETCOM SCI & TECH CO LTD
Filing Date
2024-06-19
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing intelligent question-answering systems struggle to adjust the content and format of answers based on the user's emotional state during interaction, resulting in a poor user experience.

Method used

By using voice emotion recognition technology to determine the user's interaction state, and using a pre-set knowledge base to rewrite the semantics of the answer into standard-length standard answer information, the needs of users in different emotional states can be adapted.

Benefits of technology

It improved the user's interactive experience, making it easier for users to accept and understand the answers, and thus improving the quality of responses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118692450B_ABST
    Figure CN118692450B_ABST
Patent Text Reader

Abstract

This disclosure provides a question-and-answer processing method, apparatus, electronic device, computer-readable storage medium, and computer program product, relating to artificial intelligence technologies such as speech recognition, natural language processing, and intelligent question answering. One specific implementation of the method includes: determining the semantic meaning of the answer to the question information based on the question information provided by the target user in voice form; determining the current interaction state of the target user based at least on the tone recognition result of the question information; in response to the interaction state being a first interaction state, rewriting the semantic meaning of the answer into standard-length standard answer information using a preset knowledge base; and playing the standard answer information to the target user. Therefore, in the process of finding and providing answers to a user's question, the content and format of the answer can be adjusted based on the user's interaction state to help the user more easily accept and understand the answer, improving the quality of interaction and response with the user, and enhancing the user's interactive experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, specifically to the field of artificial intelligence technologies such as speech recognition, natural language processing, and intelligent question answering, and particularly to question answering methods, devices, electronic devices, computer-readable storage media, and computer program products. Background Technology

[0002] With the development of computer technology, intelligent question answering (QA) technology has emerged to facilitate users' access to knowledge and answers to their questions. QA is an artificial intelligence technology designed to enable computer systems to understand natural language questions posed by humans and provide answers in a precise and accurate manner.

[0003] Correspondingly, to make this technology more convenient for users, an increasing number of service providers are choosing to deploy related applications on terminal devices such as smart speakers and smartphones. This allows users to interact with these devices via voice and more easily access intelligent question-and-answer services. Therefore, in this process, how to provide higher-quality services and improve the user's interactive experience is a matter of concern and urgent need. Summary of the Invention

[0004] This disclosure provides a question-and-answer processing method, apparatus, electronic device, computer-readable storage medium, and computer program product.

[0005] In a first aspect, embodiments of this disclosure propose a question-and-answer processing method, comprising: determining the answer semantics of the question information based on the question information provided by the target user in the form of voice; determining the current interaction state of the target user based at least on the tone recognition result of the question information; in response to the interaction state being a first interaction state, rewriting the answer semantics into standard length standard answer information using a preset knowledge base; and playing the standard answer information to the target user.

[0006] Secondly, embodiments of this disclosure propose a question-and-answer processing apparatus, comprising: an answer semantic determination unit configured to determine the answer semantic of a question based on question information provided by a target user in the form of voice; an interaction state determination unit configured to determine the current interaction state of the target user based at least on tone recognition results for the question information; a first answer rewriting unit configured to rewrite the answer semantic into standard-length standard answer information using a preset knowledge base in response to the interaction state being a first interaction state; and a first answer sending unit configured to play the standard answer information to the target user.

[0007] Thirdly, embodiments of this disclosure provide an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to implement the question-and-answer processing method as described in any implementation of the first aspect.

[0008] Fourthly, embodiments of this disclosure provide a non-transitory computer-readable storage medium storing computer instructions that enable a computer to implement the question-and-answer processing method as described in any implementation of the first aspect.

[0009] Fifthly, embodiments of this disclosure provide a computer program product including a computer program that, when executed by a processor, can implement the question-and-answer processing method as described in any implementation of the first aspect.

[0010] The question-and-answer processing method, apparatus, electronic device, computer-readable storage medium, and computer program product provided in this disclosure first determine the answer semantics of the question information based on the question information provided by the target user in the form of voice. Then, based at least on the tone recognition result of the question information, the current interaction state of the target user is determined. Next, if the interaction state is a first interaction state, the answer semantics are rewritten into standard-length standard answer information using a preset knowledge base. Finally, the standard answer information is played to the target user.

[0011] In the process of finding and providing answers to user questions, this disclosure can adjust the content and format of the answers based on the user's interaction status to help users more easily accept and understand the answers, improve the quality of interaction and responses with users, and enhance the user's interactive experience.

[0012] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0013] Other features, objects, and advantages of this disclosure will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:

[0014] Figure 1 This is an exemplary system architecture to which this disclosure can be applied;

[0015] Figure 2 A flowchart of a question-and-answer processing procedure provided in this embodiment of the disclosure;

[0016] Figure 3 A flowchart illustrating a process for determining the semantics of an answer, provided as an embodiment of this disclosure;

[0017] Figure 4 This is a flowchart illustrating a question-and-answer processing method in an application scenario provided by an embodiment of this disclosure;

[0018] Figure 5 A structural block diagram of a question-and-answer processing device provided in an embodiment of this disclosure;

[0019] Figure 6 This is a schematic diagram of the structure of an electronic device suitable for performing a question-and-answer processing method, provided as an embodiment of the present disclosure. Detailed Implementation

[0020] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding; these should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description. It should be noted that, unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.

[0021] Furthermore, the acquisition, storage, use, processing, transportation, provision, and disclosure of user personal information (such as images containing facial objects as discussed later in this disclosure) involved in the technical solutions disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0022] Figure 1 An exemplary system architecture 100 is shown, to which embodiments of the question-answering processing methods, apparatus, electronic devices, and computer-readable storage media of this disclosure may be applied.

[0023] like Figure 1 As shown, system architecture 100 may include terminal devices 101, 102, and 103, a network 104, and a server 105. Network 104 serves as the medium for providing communication links between terminal devices 101, 102, and 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.

[0024] User 106 can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various applications for enabling information communication between the terminal devices 101, 102, and 103 and server 105 can be installed. These applications include question-and-answer applications, semantic interaction applications, and instant messaging applications.

[0025] Terminal devices 101, 102, and 103 and server 105 can be either hardware or software. When terminal devices 101, 102, and 103 are hardware, they can be various electronic devices with displays, including but not limited to smartphones, tablets, smart speakers, laptops, and desktop computers. When terminal devices 101, 102, and 103 are software, they can be installed in the aforementioned electronic devices, and can be implemented as multiple software programs or software modules, or as a single software program or software module; no specific limitation is made here. When server 105 is hardware, it can be implemented as a distributed server cluster composed of multiple servers, or as a single server. When server 105 is software, it can be implemented as multiple software programs or software modules, or as a single software program or software module; no specific limitation is made here.

[0026] Server 105 can provide various services through its built-in applications. Taking a question-and-answer application as an example, when running this application, server 105 can achieve the following: First, user 106 can interact with terminal devices 101, 102, and 103 via voice to provide question information. Then, server 105 can obtain this question information from terminal devices 101, 102, and 103 via network 104, and determine the semantic meaning of the answer based on the question information provided by the target user via voice. Next, server 105 determines the current interaction state of the target user based at least on the tone recognition result of the question information. Then, in response to the interaction state being the first interaction state, server 105 rewrites the answer semantics into a standard-length standard answer information using a preset knowledge base. Finally, server 105 can use terminal devices 101, 102, and 103 to play the standard answer information to the target user.

[0027] Since finding the semantic meaning of the answer corresponding to the question information and rewriting the semantic meaning of the answer may require a lot of computing resources and strong computing power, the question-and-answer processing methods provided in the subsequent embodiments of this disclosure are generally executed by a server 105 with strong computing power and abundant computing resources. Accordingly, the question-and-answer processing device is generally also set in the server 105. However, it should also be noted that when the terminal devices 101, 102, and 103 also have sufficient computing power and computing resources, in this case, for example, for the purpose of user convenience, the terminal devices 101, 102, and 103 can also complete the above-mentioned calculations performed by the server 105 through the question-and-answer application installed on them, and thus output the same results as the server 105. Especially when there are multiple terminal devices with different computing capabilities at the same time, but the question-and-answer application determines that the terminal device has strong computing power and abundant computing resources, it can let the terminal device perform the above calculations, thereby appropriately reducing the computing pressure on the server 105. Accordingly, the question-and-answer processing device can also be set in the terminal devices 101, 102, and 103. In this case, the exemplary system architecture 100 may also exclude the server 105 and the network 104.

[0028] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0029] Please refer to Figure 2 , Figure 2 A flowchart of a question-and-answer processing procedure provided in this disclosure embodiment is provided, wherein process 200 includes the following steps:

[0030] Step 201: Determine the semantic meaning of the answer to the question based on the question information provided by the target user in the form of voice;

[0031] This step is intended for the entity executing the question-and-answer processing method (e.g., Figure 1 After obtaining the question information provided by the target user (e.g., user 106 mentioned above) via voice, the server 105 determines the semantic meaning of the answer corresponding to the question information. For example, after obtaining and determining the question information, the executing entity can generate the answer corresponding to the question information using online databases, pre-trained answer matching models, etc. For example, the executing entity can generate a "reference answer" by aggregating and combining content from multiple platforms and databases. Then, based on the semantic analysis of such a "reference answer," the "answer semantics" corresponding to the question information is extracted.

[0032] In some embodiments, answer semantics can generally be understood as the "simplest" form that expresses the actual content contained in the "answer." This "simplest" form is at least readable by the executing entity and can be used to convert it into text content that can be read and understood by the "user" (e.g., by mapping the answer semantics based on a pre-set mapping relationship to obtain a text form of the answer that can be read by the user, or by adding words, phrases, or sentences to aid understanding, thereby adjusting the answer semantics into a text form that can be read by the user). Of course, in some scenarios, such as when the final answer is too short, the answer semantics can also be directly read by the user; this disclosure is not intended to limit this.

[0033] For example, a user might ask about tomorrow's weather. The executing entity retrieves the first answer provided by database A, "Tomorrow's weather will be sunny," and the second answer provided by platform B, "Tomorrow's wind direction will be southeast, and the wind force will be level 3," and so on. Then, the executing entity summarizes these "answers" to obtain the semantics "sunny, southeast level 3 wind."

[0034] It should be noted that the "at least part of the answer" used to generate the answer semantics can be obtained directly from the local storage device by the aforementioned executing entity, or it can be obtained from a non-local storage device (e.g., Figure 1 The data is obtained from the terminal devices 101, 102, and 103 shown. The local storage device can be a data storage module located within the aforementioned execution entity, such as a server hard drive. In this case, the "at least partial answer" can be quickly read locally. The non-local storage device can also be any other electronic device configured to store data, such as some user terminals. In this case, the aforementioned execution entity can obtain the required "at least partial answer" by sending an acquisition command to the electronic device.

[0035] Step 202: Determine the current interaction state of the target user based at least on the tone recognition results for the question information;

[0036] Building upon step 201, this step aims to have the aforementioned executing entity determine the target user's current interaction state, at least based on the tone recognition results of the question information. For example, the executing entity can determine the user's interaction state by parsing the speech segment (or audio segment) corresponding to the question information. For instance, the executing entity can determine the user's current interaction state based on the parsing results of the speech rate, volume, tone, etc., of the speech segment.

[0037] For example, in scenarios where only one dimension—speech rate, volume, or tone—is considered, the executing entity can choose to set multiple numerical segments based on the corresponding indicator to represent different interaction states. For instance, when the executing entity determines that the speech rate is in the first numerical segment, it determines that the user is in the first interaction state. When the executing entity determines that the speech rate is in the second numerical segment, it determines that the user is in the second interaction state. The numerical value corresponding to the second segment is higher than that of the first segment; that is, the "speech rate" analysis result falling into the second numerical segment has a higher and faster speech rate compared to the "speech rate" analysis result falling into the first numerical segment.

[0038] Accordingly, the correspondence between speech rate and state can be determined in advance based on data analysis of a large number of relevant users. For example, the question-and-answer processing method provided in this disclosure can be understood as a "personalized service" (e.g., it can be understood as a service that adjusts the content and form of the answer and the interactive form of providing the answer according to the user's tone and emotion). In such a scenario, such "relevant users" can be understood as personalized users or users with special needs who expect to use such a service. Accordingly, when providing services to these personalized users or users with special needs, personalized users who expect to use such a function to achieve voice interaction and / or "other users" who expect or exist to provide reference for such a function can be used as reference users (or, the users being referenced). Then, through group analysis of a large number of reference users, the correspondence between the speech rate indicator and the interaction state can be determined. Similarly, the subsequent implementing entity can determine whether a user belongs to the "personalized user or special user" through user settings or through user analysis, and then provide such "personalized service".

[0039] For example, interaction states can be categorized or classified in terms of, for instance, emotional states. For instance, the first interaction state might correspond to a general emotional state of the referenced user group, such as a stable emotional state; the second interaction state might correspond to emotional states such as "anxious," "irritable," "frustrated," or "angry"; and the third interaction state might correspond to emotional states such as "doubtful," "cautious," "taciturn," "melancholy," or "frustrated." This allows the subsequent implementing entity to adjust its response strategy to the user based on the classification of interaction states.

[0040] Speech Emotion Recognition (SER) identifies the user's (or speaker's) interactive state, or emotional state, by analyzing the acoustic features of a speech segment (these features are independent of the speech's content and language). Users can express different emotions by adjusting the movements of their vocal organs to alter the acoustic features of the speech signal. Accordingly, discrete and continuous emotion description models are typically used to process the user's recorded speech content to identify their "interactive state."

[0041] In some embodiments, if the interaction state is determined jointly based on two or more of speech rate, volume, and tone, the executing entity can generate a score corresponding to each indicator. Then, the interaction state is determined based on the sum of the scores, or a weighted sum. For example, the summed score, from low to high, can correspond to scenarios where the user gradually changes from cautious, doubtful, or depressed states to emotionally stable states, and further to "impatient" states.

[0042] For example, in this process, as the speaking speed increases (leading to an increase in the score corresponding to speaking speed), the volume increases (leading to an increase in the score corresponding to volume), and the proportion of stressed characters in the tone increases (leading to an increase in the score corresponding to volume), a user may be determined to have a higher sum of scores. For example, for a first question based on a slower speaking speed, lower volume, and a smaller proportion of stressed characters, a second question with a faster speaking speed, higher volume, and a higher proportion of stressed characters may have a higher sum of scores.

[0043] Step 203: In response to the interaction state being the first interaction state, rewrite the semantics of the answer into standard-length standard answer information using a preset knowledge base;

[0044] Building upon step 201, this step aims to have the executing entity, upon determining that the target user is currently in the first interactive state, rewrite the answer semantics into standard-length standard answer information using a preset knowledge base. For example, the executing entity can segment the answer semantics into words, extract the corresponding entity words from the preset knowledge base, and replace them to construct standard-length standard answer information. In some embodiments, multiple standard-length (reference) standard answer information can also be directly configured in the preset knowledge base. Then, the executing entity can determine the standard answer information based on the comparison results between these (reference) standard answer information and the answer semantics (e.g., the standard answer information in the knowledge base with the highest similarity among all comparison results, where the similarity exceeds a preset similarity threshold).

[0045] In some optional implementations of this embodiment, the executing entity can also rewrite the semantics of the answer using a pre-trained rewriting model to obtain "standard answer information." Accordingly, this "standard answer information" can be considered as the response content that meets general reading requirements, enabling the user to clearly understand it. Correspondingly, its length can also be considered as the "standard length" of the answer information provided to the user (it is usually not the minimum length that meets basic clarity requirements, but it is usually a "normal, general" length that meets reading clarity and comfort without easily causing ambiguity).

[0046] Step 204: Play the standard answer information to the target user.

[0047] This step aims to have the executing entity play the standard answer information obtained based on step 203 above to the target user. For example, if the executing entity is a server, it can push the standard answer information back to the "terminal device" that collected the target user's question information, and use the "terminal device" to play the standard answer information to reply to the target user.

[0048] The question-and-answer processing method provided in this embodiment can adjust the content and format of the answer based on the user's interaction state during the process of finding and providing answers to user questions, so as to help users more easily accept and understand the answer, improve the quality of interaction and response to users, and enhance the user's interactive experience.

[0049] In some embodiments, if the interaction state is determined to be a second interaction state, which, for example, could be the aforementioned "impatient" emotional state, the executing entity may, in response to the interaction state being the second interaction state, first rewrite the answer semantics into standard-length standard answer information using a preset knowledge base, as discussed above.

[0050] Then, the executing entity selects to generate a first rewritten answer information with a length shorter than the standard answer information based on the deletion and / or replacement of at least one non-entity word and / or entity words below the preset entity level in the standard answer information.

[0051] Specifically, the executing entity can choose to perform word segmentation on the standard answer, and then determine whether each word segment is an entity based on the word segmentation result, and the entity level corresponding to the word segment that is an entity (wherein, the entity that has a greater impact on the semantic understanding of the answer can have a higher entity level. For example, in the weather example above, since the target user has explicitly specified the date as "tomorrow", the entity level of "tomorrow" used to represent the date will be lower than the entity level of "sunny" which represents the weather result. Accordingly, "tomorrow" can be deleted because it is lower than the preset entity level).

[0052] Then, the executing entity can choose to delete non-entity words and / or entity words below a preset entity level, or update the expression of non-entity words and entity words (e.g., using shorter words with similar meanings as replacements). For example, in a shopping scenario, if the question is related to price, such as the price of product A, and the standard answer is "1 product A, X1 yuan; 3 product A, X2 yuan," the executing entity can, for example, delete the non-entity word "one" to obtain a first rewritten answer of "X1 yuan; 3 products, X2 yuan." Thus, by shortening the "standard length," a rewritten solution that is easier for the target user to read quickly can be obtained.

[0053] Accordingly, the implementing entity can provide and play the first rewritten answer information to the target user based on the question information. This facilitates the target user's faster information acquisition, avoids generating more negative emotions due to potentially redundant content, and improves the target user's interactive experience.

[0054] In some optional implementations of this embodiment, when the target user may remain in the target interaction state for an extended period (e.g., an "anxious" interaction state), the executing entity may choose to proactively intervene or provide reassurance to offer care to the target user. Specifically, the executing entity may also play a preset reassurance voice message to the target user in response to determining that the number of times the target user is in the target interaction state within a preset time period equals a preset threshold.

[0055] For ease of understanding, let's continue with the example based on the scenario described above, where the second interaction state represents the emotional state of "impatience." Each time the executing entity detects that the target user is in the second interaction state, it can use the current time as a base point (or the end point of a preset duration) to find the number of times the target user has been in the second interaction state within that preset historical time period. If the number of times the target user is in the second interaction state equals a preset threshold (typically determined based on a pre-determined number of times "intervention" is needed), a preset soothing voice message is played to the target user. In some embodiments, the soothing voice message can be provided by the target user through interaction (e.g., voice content provided by the target user that they believe may have a "soothing" effect). In some embodiments, the soothing voice message can also be pre-configured based on the service provider of the question-and-answer processing method.

[0056] Therefore, when the target user may be in a state of "impatience" for a long time, the implementing entity can actively and proactively "intervene" to soothe them, enriching the interactive functions while improving the user's interactive experience and providing care for the user.

[0057] In some embodiments, if the interaction state is determined to be a third interaction state, which, for example, could be the "cautious" emotional state described above, the executing entity may, in response to the interaction state being a third interaction state, first rewrite the answer semantics into standard-length standard answer information using a preset knowledge base, as discussed above.

[0058] Then, the executing entity generates a second rewritten answer information longer than the first answer information by inserting at least one phrase into the standard answer information. In some embodiments, the phrase can be a further detailed explanation of a specific entity word in the standard answer information. For example, the executing entity can analyze words in the standard answer information that require further explanation based on historical question-and-answer results from a large number of historical users. For instance, the executing entity can identify words that are "easy" to be followed up on by analyzing words in second-round questioning and follow-up dialogues from a large number of historical users. Then, if such words exist in the standard answer information, the executing entity proactively inserts the further explanation of the corresponding word as phrase content into the standard answer information. In some embodiments, the phrase content can also be words or sentences with a soothing or calming tone.

[0059] Accordingly, the implementing entity can enhance the quality of the response by inserting these phrases into the standard answer information, making it at least "more approachable" in tone. This aims to alleviate the target user's caution and tension. Correspondingly, the implementing entity can provide and play the second rewritten answer to the target user based on the question information. Thus, by providing richer, more "warm" (rewritten) answer information, the user's "tension and caution" are alleviated, improving the user's interactive experience.

[0060] In some optional implementations of this embodiment, to determine the script content more conveniently and efficiently, a script database can also be maintained and set up in advance, so that the executing entity can use the script database to determine the script content. For example, for the "soothing" function, some words and sentences with soothing functions can be collected and set up in advance to obtain the script database.

[0061] Then, the executing entity can select to determine the actual script content to be used based on the similarity matching results between the local and target user's historical dialogue content and the reference script content in the preset script database. For example, the executing entity can compare each reference script content in the script database with the historical dialogue content provided by the target user (to improve comparison efficiency, this historical dialogue can also be based on dialogue content provided by the user that they believe has a "soothing" effect). Then, the reference script content in the script database that has the highest similarity to the target dialogue content provided by the target user in the historical dialogue content, exceeding the preset similarity threshold, is determined as the target reference script content to be used, and is used as the script content when inserting standard answer information.

[0062] Therefore, by using this method, we can find "script content" that better meets the emotional value needs of the target users and is easier for them to understand, based on their language habits, thereby improving the quality of script content selection.

[0063] It should be understood that the "emotional states" corresponding to the first, second, and third interaction states mentioned above are merely examples to distinguish the states of the three different interactions. For ease of understanding, in this text, the first interaction state is used as an example to refer to the "normal emotional state of emotional stability" discussed above, the second interaction state as an example to refer to the "impatient" emotional state discussed above, and the third interaction state as an example to refer to the "doubtful" and "cautious" emotional states discussed above. In some scenarios, this correspondence can also be regarded as the "default" configuration of the executing subject.

[0064] Correspondingly, such a "default" configuration can be adjusted based on user needs and actual circumstances. For example, in some scenarios, users may prefer to obtain a more comprehensive "second rewritten answer" when they are identified as being in an "impatient" state. Accordingly, in some embodiments, the executing entity can also allow the target user to adjust the actual "meaning" of the first, second, and third interaction states based on the way they communicate and interact with the target user. In short, the target user can adjust the executing entity's criteria for determining and identifying the first, second, and third interaction states through interaction with the executing entity.

[0065] For example, if a target user prefers to receive standard-length answer information in a "frustrated" emotional state, they can choose to adjust the actual meaning of the first interaction state to the "frustrated" emotional state, or in other words, define the first interaction state as the state in which the executor determines they are "frustrated." Similarly, if a target user expects to receive more detailed and / or reassuring rewritten answer information in a "stable, ordinary emotional state," they can adjust the third interaction state to a "stable, ordinary" emotional state.

[0066] Accordingly, in some embodiments, as discussed above, the "boundaries" for the division of the first, second, and third interaction states can also be adaptively adjusted based on the needs of the actual scenario. Correspondingly, the "specific state" corresponding to each interaction state can also be divided according to different criteria. For example, the aforementioned "cautious" state can be replaced by the "frustrated" state. For instance, when the third interaction state is "frustrated," the executing entity can also provide "reassurance" to the user by adding dialogue, hoping to help them escape the "frustrated" interaction state.

[0067] Furthermore, it should be understood that if, in actual use, the executing entity adjusts the various interaction states according to the target user's instructions, so that the content of the first, second, and third interaction states differs from the content executored above, then adaptive adjustments will also be made to the implementation process of certain embodiments. For example, regarding the case executored above where the "second interaction state" is used as the "target interaction state," if the user adjusts the first interaction state to "impatient," the executing entity can adaptively determine the "target interaction state" as the first interaction state. For example, the executing entity can determine the "target interaction state" based on factors such as the sum of the scores, parameters, and emotion classification results provided when the target user sets the "interaction state." For example, the executing entity can determine all interaction states considered to belong to "negative emotions" as the target interaction states.

[0068] In some embodiments, the executing entity may also select an interaction state belonging to the target interaction state based on the settings of the target user.

[0069] In some embodiments, in determining the interaction state of the target user, the executing entity may also choose to combine the facial action recognition results and the tone recognition results of the target user in providing the question information to determine the current interaction state of the target user.

[0070] Specifically, the executing entity can use Face Reader (FER) technology to recognize the facial movements of the target user while they are providing information, thus obtaining a recognition result. For example, the executing entity can determine the first sub-interaction state based on the changes in the movements of key points on the face, such as the eyes, eyebrows, mouth, and facial muscles. This first sub-interaction state is then combined with the second sub-interaction state determined from the speech, as discussed above, to obtain the final recognition result—the target user's current interaction state. Therefore, by using both facial movement and speech recognition paths, the target user's interaction state can be more accurately determined.

[0071] In some embodiments, during the process of playing an answer (e.g., standard answer information, first rewritten answer information, or second rewritten answer information) to the user, the voice playback parameters used during playback can be further adjusted based on the determined interaction state. Voice playback parameters include at least one of the following: volume parameters, speech rate parameters, and tone parameters. For example, the voice playback parameters can also be determined by configuring them, as described above, either by the user or by the service provider of the question-and-answer service.

[0072] In some embodiments, the logic of volume, speech rate, and tone parameters used in determining voice playback parameters for an interaction state can be reversed compared to when determining the interaction state. For example, for an interaction state determined to be "impatient," the corresponding voice playback parameters can be determined by using lower volume, slower speech rate, and a calmer tone to play the answer information (e.g., rewriting the answer information). This further caters to the target user's emotional state during playback, enhancing the user experience.

[0073] Accordingly, after determining the voice playback parameters, the executing entity can play the corresponding answer information (e.g., standard answer information, first rewritten answer information, or second rewritten answer information) to the target user based on the voice playback parameters.

[0074] Furthermore, to enhance the usage scenarios and capabilities of the executor, it can be configured to provide unique responses to specific questions. For example, it can be configured to respond to inquiries about health status (e.g., such responses could be based on the analysis of the user's physical indicators).

[0075] Please refer to Figure 3 , Figure 3 The flowchart provided in this disclosure describes a process for determining the semantics of an answer, which can be applied to scenarios where the question information is a query about the health status of a target user. Figure 3 This includes process 300, which can serve as an alternative to or replacement for step 201 in scenarios where the question information is a query about the health status of the target user. Process 300 specifically includes the following steps:

[0076] Step 301: In response to receiving the question information provided by the target user in the form of voice, send a request to the target user to obtain the target user's physical indicators;

[0077] Specifically, upon receiving a question from the target user via voice, the executing entity responds by requesting the target user to provide their health metrics. For example, the executing entity can use voice interaction or a display screen to provide the user with the query information in order to request the target user's health metrics.

[0078] In some embodiments, the acquisition request issued by the executing entity may also explicitly include the specific details of the various physical indicators that are expected to be acquired, so as to enable the target user to make more precise authorizations and protect the target user's privacy.

[0079] Step 302: In response to receiving an authorization instruction from the target user in response to the request, obtain a list of body indicators based on the preset first communication path;

[0080] Specifically, if the target user confirms authorization, they can return an authorization instruction to the executing entity (e.g., an authorization instruction in the form of a voice, or a click on the "confirmation control" provided in the query information on the display). Accordingly, in response to receiving the authorization instruction from the target user in response to the acquisition request, the executing entity obtains the list of body indicators based on a preset first communication path.

[0081] The list of physical indicators may include at least one physical indicator specific to the target user (e.g., blood pressure, pulse, uric acid, etc.). The first communication path may be a terminal device for storing the target user's physical indicators; for example, it may be a storage device of a medical institution, a terminal device used by the target user, etc. In some embodiments, if the list of physical indicators is stored in the local storage of the executing entity, the first communication path may also be an internal path of the executing entity, so that the executing entity can directly access the "list of physical indicators" locally.

[0082] It should be understood that the list of physical indicators can be stored separately in the aforementioned entities in multiple parts, and correspondingly, the implementing entity can obtain the complete "list of physical indicators" by communicating with these entities respectively.

[0083] Step 303: Based on the evaluation results of each physical indicator in the physical indicator list, determine the semantic meaning of the answer to the question information.

[0084] Specifically, after obtaining the "body indicator list" based on step 302 above, the executing entity can determine the evaluation result of each body indicator based on the numerical relationship between each body indicator and its corresponding preset reference threshold. For example, if blood pressure is higher than the corresponding blood pressure threshold, the executing entity can determine the evaluation result for the body indicator "blood pressure" (e.g., unhealthy blood pressure). Then, the executing entity can comprehensively determine the answer semantics of the question information based on the evaluation results of each body indicator. For example, if there are 10 body indicators, and the executing entity determines that 9 of them are in an "unhealthy" state (i.e., these 9 do not meet the requirements of their respective thresholds), then the executing entity can determine the answer semantics of the question information as "the user is currently in an unhealthy state." It should be understood that the descriptions of 9 and 10 above are for illustrative purposes only and are not intended to be limiting.

[0085] In some embodiments, the implementing entity can further refine and enrich the semantics of the answer based on specific indicators of an unhealthy state. For example, "The user is currently in an unhealthy state caused by excessive blood pressure," etc.

[0086] Therefore, with the user's authorization, health monitoring services can be provided based on the user's physical indicators to meet the user's care needs.

[0087] In some optional implementations of this embodiment, when the executing entity determines that the answer semantics indicate that the target user is currently in a target health state, the executing entity can respond by communicating the target health state to the target user device based on a preset second communication path. The target user device can be a pre-recorded other user with a caregiving or assistance relationship with the target user. For example, the target device can be a communication device of an institution such as a medical institution or a rescue institution, or a device used by a relative of the target user, etc.

[0088] This allows these entities to understand the status of target users in a timely and synchronous manner, enabling personalized and special user groups, such as those with personalized needs that may require more effort to help and care, to receive more adequate "care".

[0089] Accordingly, the target health status can also be configured based on the needs of the target user and / or these receiving entities. For example, the target health status can be determined to be "in an unhealthy state" so that when it is determined that a target user who needs more effort to be helped and cared for, for example, is in an unhealthy state, the situation can be synchronized to the various entities more quickly so that the various entities can provide assistance to the target user as early and timely as possible.

[0090] Building upon any of the above embodiments, the interaction effect and quality can be improved by introducing a Generative Language Model (GLM). For example, GLM can be used to improve processing efficiency and quality in determining the interaction state, determining the semantics of the answer, and modifying the semantics of the answer.

[0091] Generative Language Models (GLMs) are a type of Large Language Model (LLM). LLMs are artificial intelligence models designed to understand and generate human language. Furthermore, generative language models can perform processing operations based on their understanding to obtain corresponding results. Conversely, other types of LLMs can be used to equivalently implement the processing procedures of GLMs; this discussion focuses solely on GLMs.

[0092] LLMs are characterized by their large scale, typically including a large number of parameters to help them learn complex patterns in language data. These models are often based on deep learning architectures, such as transformers, which helps them deliver better processing performance on a variety of NLP tasks.

[0093] In embodiments of this disclosure, the executing entity may also optionally acquire a generative language model. The generative language model is at least configured to rewrite the semantics of an answer into standard-length standard answer information using a pre-defined knowledge base. For example, the generative language model may be pre-trained to enable it to rewrite the semantics of an answer into standard-length standard answer information using a pre-defined knowledge base. In this case, the executing entity can instruct the rewriting of the semantics of an answer into standard-length standard answer information using a pre-configured guide word, guide tag, etc. For example, in a scenario where the entity is instructed to rewrite the semantics of an answer into standard-length standard answer information using a pre-defined knowledge base, the guide word could be, for example: "Based on the pre-defined knowledge base, rewrite 'XXX' into an answer to the 'question information'."

[0094] Similarly, generative language models can omit "guide words" through default configuration. For example, for the purpose of rewriting the semantics of an answer into a standard-length standard answer using a pre-defined knowledge base, the generative language model can, based on its default configuration, naturally understand the need to utilize the pre-defined knowledge base to rewrite the semantics of the answer into a standard-length standard answer. Thus, through default configuration, the generative language model can stably and purposefully utilize the pre-defined knowledge base to rewrite the semantics of the answer into a standard-length standard answer. Therefore, generative language models can be used to utilize the pre-defined knowledge base to rewrite the semantics of the answer into a standard-length standard answer more efficiently and with higher quality. Typically, the "standard" corresponding to this "standard answer" can also be "learned" through pre-training of the GLM.

[0095] Accordingly, in some embodiments, the executing entity may similarly utilize GLM to determine the target user's interaction state, generate first rewritten answer information, second rewritten answer information, and determine the speech playback parameters for the standard answer information. For example, the executing entity can instruct GLM to perform data analysis on a large number of users to determine the correspondence between speech rate and state.

[0096] For example, in the process of providing the above-mentioned question-and-answer processing methods primarily for "personalized users," one can choose to acquire voice emotion feature modeling data and facial emotion tagging data by focusing on the daily lives and interests of personalized users and their corresponding reference users, through the internet, books, and manual surveys. Regarding the dialogue content, one can obtain user-referenced dialogue content from publicly accessible channels (e.g., social media, user forums that gather reference users, health and wellness websites, and other internet platforms) to analyze the dialogue content that "personalized users" prefer to use, and so on.

[0097] Then, based on the data cleaning (e.g., using hash algorithms or similarity comparison techniques to identify and remove duplicate data records; for example, using methods such as mean imputation, multiple imputation, or machine learning-based predictive models to fill in missing values, data distribution, and correlation analysis; for example, using statistical methods such as IQR or Z-score to identify outliers and correcting or removing them according to the actual situation), the GLM is trained on the cleaned data to enable the GLM to have the processing capabilities in the corresponding scenarios and to provide data processing capabilities that meet the needs (e.g., to determine the wording content that is more suitable for personalized users).

[0098] To enhance understanding, this disclosure also provides a specific implementation scheme based on a particular application scenario. Please refer to the example below. Figure 4The process shown is 400. For ease of understanding, it will be partially combined... Figure 1 The exemplary system architecture 100 shown is illustrated.

[0099] In process 400, for example, user 106 with the aforementioned personalized needs can interact with terminal device 103 via voice. Accordingly, terminal device 103 can obtain processing capability support (e.g., question-and-answer processing capability support) based on communication with server 105.

[0100] As shown in process 400, user 106 can first execute S401 to interact with terminal device 103 and provide problem information 410 in the form of voice.

[0101] Then, the terminal device 103 can execute S402 to provide the received question information 410 to the server 105, so as to use the processing capability of the server 105 to process the question information 410 and provide question and answer services to the user 106.

[0102] After receiving the question information 410, server 105 can execute S403 to use GLM to determine the answer semantics 415 corresponding to the question information 410. For example, as explained above, server 105 can construct the guiding word "Please determine the answer semantics corresponding to question information 410" based on the obtained question information 410 to instruct GLM to generate answer semantics 415.

[0103] Next, server 105 continues to execute S404 to use GLM to determine the current interaction state 420 of user 106 (for example, it could be one of the first, second, or third interaction states discussed above). Server 105 can similarly instruct GLM to determine the current interaction state 420 of user 106 based on the voice segment corresponding to the question information 410 by constructing a guide word, which will not be repeated here.

[0104] After obtaining the interaction state 420, the executing entity can rewrite the answer semantics 415 based on the selection execution S405 to obtain the rewritten answer information 430. In this process, as discussed above, the executing entity can determine, based on the rewriting strategy corresponding to the interaction state (e.g., ultimately rewritten to standard-length standard answer information, ultimately rewritten to first rewritten answer information, or ultimately rewritten to second rewritten answer information), to rewrite the answer semantics 415 into one of the above three options based on the determined interaction state 420 to obtain the rewritten answer information 430.

[0105] For ease of understanding, for example, the standard answer information of standard length obtained in the first interaction state can be exemplified as "YYYYY", the first rewritten answer information can be exemplified as "YYYY", and the second rewritten answer information can be exemplified as "YYYYYYYY".

[0106] For example, in process 400, if the example interaction state 420 is determined to be "second interaction state", then the corresponding rewritten answer information 430 can be in the form of "YYYY".

[0107] Next, server 105 can execute S406 to return the rewritten answer information 430 to terminal device 103.

[0108] Finally, the terminal device 103 plays the rewritten answer information 430 to the user 106 to "answer" the question information 410 provided by the user.

[0109] Further reference Figure 5 As an implementation of the methods shown in the above figures, this disclosure provides an embodiment of a question-and-answer processing device, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.

[0110] like Figure 5 As shown, the question-and-answer processing device 500 of this embodiment may include: an answer semantic determination unit 501, an interaction state determination unit 502, a first answer rewriting unit 503, and a first answer sending unit 504. The answer semantic determination unit 501 is configured to determine the answer semantics of the question information based on the question information provided by the target user in voice form; the interaction state determination unit 502 is configured to determine the current interaction state of the target user based at least on the tone recognition result of the question information; the first answer rewriting unit 503 is configured to rewrite the answer semantics into standard-length standard answer information using a preset knowledge base in response to the interaction state being a first interaction state; and the first answer sending unit 504 is configured to play the standard answer information to the target user.

[0111] In this embodiment, the specific processing of the answer semantic determination unit 501, the interaction state determination unit 502, the first answer rewriting unit 503, and the first answer sending unit 504 in the question-and-answer processing device 500, and the resulting technical effects, can be found in reference to [reference needed]. Figure 2 The relevant descriptions of steps 201-204 in the corresponding embodiments will not be repeated here.

[0112] In some optional implementations of this embodiment, the device 500 further includes: a second answer rewriting unit configured to, in response to an interaction state of a second interaction state, rewrite the answer semantics into standard-length standard answer information using a preset knowledge base; a third answer rewriting unit configured to generate first rewritten answer information based on the deletion and / or replacement of at least one non-entity word and / or entity words lower than a preset entity level in the standard answer information, wherein the length of the first rewritten answer information is less than that of the standard answer information; and a second answer sending unit configured to play the first rewritten answer information to a target user.

[0113] In some optional implementations of this embodiment, the device 500 further includes: a soothing voice playback unit, configured to play preset soothing voice information to the target user in response to determining that the number of times the target user is in the target interaction state within a preset time period is equal to a preset quantity threshold.

[0114] In some optional implementations of this embodiment, the device 500 further includes: a fourth answer rewriting unit configured to, in response to an interaction state of a third interaction state, rewrite the answer semantics into standard-length standard answer information using a preset knowledge base; a fifth answer rewriting unit configured to generate second rewritten answer information by inserting at least one phrase into the standard answer information, wherein the length of the second rewritten answer information is greater than that of the first answer information; and a third answer sending unit configured to play the second rewritten answer information to a target user.

[0115] In some optional implementations of this embodiment, the dialogue content is determined based on the similarity matching results between the local and target user's historical dialogue content and reference dialogue content in a preset dialogue database. The dialogue content is the target reference dialogue content that is recorded in the dialogue database and whose similarity to the target dialogue content provided by the target user in the historical dialogue content is greater than or equal to a preset similarity threshold.

[0116] In some optional implementations of this embodiment, the interaction state determination unit 502 is further configured to determine the current interaction state of the target user based on the facial action recognition results during the process of the target user providing question information and the tone recognition results for the question information.

[0117] In some optional implementations of this embodiment, the first answer sending unit 504 includes: a voice parameter determination subunit, configured to determine voice playback parameters for the standard answer information based on the interaction state, wherein the voice playback parameters include at least one of the following: volume parameter, speech rate parameter, and tone parameter; and an answer sending subunit, configured to play the standard answer information to the target user based on the voice playback parameters.

[0118] In some optional implementations of this embodiment, the question information is a health status inquiry of the target user. The answer semantic determination unit 501 includes: an indicator request acquisition subunit, configured to, in response to receiving the question information provided by the target user in the form of voice, send an acquisition request to the target user to acquire the target user's body indicators; an indicator list acquisition subunit, configured to, in response to receiving an authorization instruction returned by the target user in response to the acquisition request, acquire a body indicator list based on a preset first communication path; and an answer semantic determination subunit, configured to, based on the evaluation results of each body indicator in the body indicator list, determine the answer semantics of the question information, wherein the evaluation results are determined based on the numerical relationship between each body indicator and its corresponding preset reference threshold.

[0119] In some optional implementations of this embodiment, the device 500 further includes: a health status communication unit, configured to communicate the target health status to the target user device based on a preset second communication path in response to an answer semantic indication that the target user is currently in a target health status.

[0120] In some optional implementations of this embodiment, the standard answer information is obtained by rewriting the answer semantics using a generative language model.

[0121] This embodiment exists as a device embodiment corresponding to the above method embodiment. The question-and-answer processing device provided in this embodiment can adjust the content and form of the answer based on the user's interaction state during the process of finding and providing answers to the user's questions, so as to help the user more easily accept and understand the answer, improve the quality of interaction and response with the user, and enhance the user's interactive experience.

[0122] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0123] Figure 6 A schematic block diagram of an example electronic device 600 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0124] like Figure 6As shown, device 600 includes a computing unit 601, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 602 or a computer program loaded from storage unit 608 into random access memory (RAM) 603. RAM 603 may also store various programs and data required for the operation of device 600. The computing unit 601, ROM 602, and RAM 603 are interconnected via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.

[0125] Multiple components in device 600 are connected to I / O interface 605, including: input unit 606, such as keyboard, mouse, etc.; output unit 607, such as various types of monitors, speakers, etc.; storage unit 608, such as disk, optical disk, etc.; and communication unit 609, such as network card, modem, wireless transceiver, etc. Communication unit 609 allows device 600 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0126] The computing unit 601 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above, such as question-and-answer processing methods. For example, in some embodiments, the question-and-answer processing method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed on device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by the computing unit 601, one or more steps of the question-and-answer processing method described above may be performed. Alternatively, in other embodiments, the computing unit 601 may be configured to perform question-and-answer processing methods by any other suitable means (e.g., by means of firmware).

[0127] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0128] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0129] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0130] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0131] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0132] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is established by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, also known as cloud computing servers or cloud hosts, which are hosting products within the cloud computing service ecosystem to address the management difficulties and weak business scalability inherent in traditional physical hosts and Virtual Private Servers (VPS) services. Servers can also be categorized as distributed system servers or servers incorporating blockchain technology.

[0133] According to the technical solution of this disclosure, in the process of finding and providing answers to user questions, the content and format of the answers can be adjusted based on the user's interaction state to help the user more easily accept and understand the answers, improve the quality of interaction and response to the user, and enhance the user's interactive experience.

[0134] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution provided in this disclosure can be achieved, and this is not limited herein.

[0135] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A question-and-answer processing method, comprising: Based on the question information provided by the target user in the form of voice, determine the semantic meaning of the answer to the question information; Based at least on the tone recognition results for the question information, the current interaction state of the target user is determined; In response to the first interaction state, the semantics of the answer are rewritten into standard-length standard answer information using a preset knowledge base; Play the standard answer information to the target user; In response to the interaction state being the second interaction state, the semantics of the answer are rewritten into standard-length standard answer information using a preset knowledge base; Based on the deletion and / or replacement of at least one non-entity word and / or entity word below a preset entity level in the standard answer information, a first rewritten answer information is generated, wherein the length of the first rewritten answer information is less than that of the standard answer information; The first rewritten answer information is played to the target user.

2. The method according to claim 1, further comprising: In response to determining that the number of times the target user is in the target interaction state within a preset time period is equal to a preset quantity threshold, a preset soothing voice message is played to the target user.

3. The method according to claim 1, further comprising: In response to the interaction state being the third interaction state, the semantics of the answer are rewritten into standard-length standard answer information using a preset knowledge base; By inserting at least one phrase into the standard answer information, a second rewritten answer information is generated, wherein the length of the second rewritten answer information is greater than that of the first rewritten answer information; as well as The second rewritten answer information is played to the target user.

4. The method according to claim 3, wherein, The dialogue content is determined based on the similarity matching results between the local and target user's historical dialogue content and reference dialogue content in a preset dialogue database. The dialogue content is the target reference dialogue content that is recorded in the dialogue database and whose similarity with the target dialogue content provided by the target user in the historical dialogue content is greater than or equal to a preset similarity threshold.

5. The method according to claim 1, wherein, Determining the current interaction state of the target user based at least on the tone recognition result of the question information includes: Based on the facial motion recognition results during the process of the target user providing the question information, and the tone recognition results for the question information, the current interaction state of the target user is determined.

6. The method according to claim 1, wherein, Playing the standard answer information to the target user includes: Based on the interaction state, voice playback parameters for the standard answer information are determined, wherein the voice playback parameters include at least one of the following: volume parameter, speech rate parameter, and tone parameter; Based on the aforementioned voice playback parameters, the standard answer information is played to the target user.

7. The method according to claim 1, wherein, The question information is a health status inquiry for the target user. Determining the semantic meaning of the answer to the question information based on the question information provided by the target user in voice format includes: In response to receiving a question from the target user via voice, a request is sent to the target user to obtain the target user's physical indicators; In response to receiving an authorization instruction from the target user in response to the acquisition request, the body indicator list is acquired based on a preset first communication path; Based on the evaluation results of each physical indicator in the list of physical indicators, the semantic meaning of the answer to the question information is determined, wherein the evaluation results are determined based on the numerical relationship between each physical indicator and its corresponding preset reference threshold.

8. The method according to claim 7, further comprising: In response to the semantic indication of the answer indicating that the target user is currently in a target health state, the target health state is communicated to the target user device based on a preset second communication path.

9. The method according to any one of claims 1-8, wherein, The standard answer information is obtained by rewriting the semantics of the answer using a generative language model.

10. A question-and-answer processing apparatus, comprising: The answer semantic determination unit is configured to determine the answer semantics of the question information provided by the target user in the form of voice. The interaction state determination unit is configured to determine the current interaction state of the target user based at least on the tone recognition result of the question information. The first answer rewriting unit is configured to respond to the first interaction state by using a preset knowledge base to rewrite the answer semantics into standard length standard answer information. The first answer sending unit is configured to play the standard answer information to the target user; The second answer rewriting unit is configured to respond to the second interaction state by using a preset knowledge base to rewrite the answer semantics into standard length standard answer information. The third answer rewriting unit is configured to generate first rewritten answer information based on the deletion and / or replacement of at least one non-entity word and / or entity word below a preset entity level in the standard answer information, wherein the length of the first rewritten answer information is less than that of the standard answer information; The second answer sending unit is configured to play the first rewritten answer information to the target user.

11. The apparatus of claim 10, further comprising: The soothing voice playback unit is configured to play a preset soothing voice message to the target user in response to determining that the number of times the target user is in a target interaction state within a preset time period is equal to a preset quantity threshold.

12. The apparatus of claim 10, further comprising: The fourth answer rewriting unit is configured to respond to the interaction state as the third interaction state by rewriting the semantics of the answer into standard-length standard answer information using a preset knowledge base; The fifth answer rewriting unit is configured to generate second rewritten answer information by inserting at least one phrase into the standard answer information, wherein the length of the second rewritten answer information is greater than that of the first rewritten answer information; as well as The third answer sending unit is configured to play the second rewritten answer information to the target user.

13. The apparatus according to claim 12, wherein, The dialogue content is determined based on the similarity matching results between the local and target user's historical dialogue content and reference dialogue content in a preset dialogue database. The dialogue content is the target reference dialogue content that is recorded in the dialogue database and whose similarity with the target dialogue content provided by the target user in the historical dialogue content is greater than or equal to a preset similarity threshold.

14. The apparatus according to claim 10, wherein, The interaction state determination unit is further configured to determine the current interaction state of the target user based on the facial action recognition results during the process of the target user providing the question information, and the tone recognition results for the question information.

15. The apparatus according to claim 10, wherein, The first answer sending unit includes: The voice parameter determination subunit is configured to determine voice playback parameters for the standard answer information based on the interaction state, wherein the voice playback parameters include at least one of the following: volume parameter, speech rate parameter, and tone parameter; The answer sending subunit is configured to play the standard answer information to the target user based on the voice playback parameters.

16. The apparatus according to claim 10, wherein, The question information is a health status inquiry for the target user, and the answer semantic determination unit includes: The indicator request acquisition subunit is configured to, in response to receiving a question provided by the target user in the form of voice, send an acquisition request to the target user to obtain the target user's body indicators; The indicator list acquisition subunit is configured to acquire the body indicator list based on a preset first communication path in response to receiving an authorization instruction from the target user in response to the acquisition request. The answer semantic determination subunit is configured to determine the answer semantic of the question information based on the evaluation results of each physical indicator in the list of physical indicators, wherein the evaluation results are determined based on the numerical relationship between each physical indicator and its corresponding preset reference threshold.

17. The apparatus of claim 16, further comprising: A health status communication unit is configured to, in response to the semantic indication of the answer indicating that the target user is currently in a target health status, communicate the target health status to the target user device based on a preset second communication path.

18. The apparatus according to any one of claims 10-17, wherein, The standard answer information is obtained by rewriting the semantics of the answer using a generative language model.

19. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the question-and-answer processing method according to any one of claims 1-9.

20. A non-transitory computer-readable storage medium storing computer instructions for causing the computer to perform the question-and-answer processing method according to any one of claims 1-9.

21. A computer program product comprising a computer program that, when executed by a processor, implements the question-and-answer processing method according to any one of claims 1-9.