A natural language interpretation method, device and storage medium based on character portrait
By combining interactive scenes and character portrait interpretation models in the speech recognition process, the keywords of the speech data are corrected, and the problem of low accuracy of speech recognition in the prior art is solved, achieving more efficient and more accurate speech interpretation.
Patent Information
- Application Number
- CN202210553460.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-20
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2042-05-20
AI Technical Summary
The existing speech recognition methods and language interpretation methods cannot provide targeted interpretation and recognition based on character portraits, resulting in low accuracy.
By determining the interactive behavior interpretation model corresponding to the interactive scene, and using this model to correct the keywords in the recognition results of the speech data, and interpreting and correcting them in combination with the character portrait information.
It improves the accuracy and efficiency of speech recognition and language interpretation, and enhances the pertinence and accuracy of user interaction behavior.
Smart Images

Figure CN114974253B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech recognition technology, and in particular to a natural language interpretation method, device and storage medium based on character portraits. Background Art
[0002] With the continuous development of computer technology, human-computer interaction methods are becoming increasingly diverse and intelligent. Currently, more and more interactive platforms are adopting voice interaction. Voice interaction can improve user interaction efficiency and enhance user enjoyment, and has become an important human-computer interaction method. For example, self-service voice customer service systems first pose questions to users through voice, and then the user answers by voice. Another example is some navigation systems and shopping systems, which require users to issue voice commands to control the displayed content. In these scenarios, accurate voice recognition of users is required to provide appropriate feedback.
[0003] However, the accuracy of existing speech recognition methods or language interpretation methods is not high, and they cannot perform targeted interpretation and recognition based on character portraits. Summary of the Invention
[0004] Based on the above problems, the present invention proposes a natural language interpretation method, device and storage medium based on character portraits. By determining a first interactive behavior interpretation model corresponding to a first interactive scene and using the first interactive behavior interpretation model to correct the first keyword in the first recognition result of the first voice data, the efficiency of voice recognition / language interpretation can be improved while the accuracy of voice recognition / language interpretation can be improved.
[0005] In view of this, one aspect of the present invention proposes a natural language interpretation method based on character portraits, comprising:
[0006] receiving first voice data;
[0007] Determining a first interaction scenario to which the first voice data belongs;
[0008] Selecting a first interaction behavior explanation model corresponding to the first interaction scenario;
[0009] performing speech recognition on the first speech data to obtain a first recognition result;
[0010] Using the first interactive behavior explanation model, modifying the first keyword that meets the preset condition in the first recognition result;
[0011] Among them, the first interactive behavior explanation model includes the association relationship between interactive scene information and character portrait information.
[0012] Optionally, after the step of determining the first interaction scenario to which the first voice data belongs, the method further includes:
[0013] According to the first interaction scenario, prompting the user to perform a first action and / or prompting the user to speak first text data and / or prompting the user to input first selection data;
[0014] collecting the first action data and / or the first text data and / or the first selection data;
[0015] First key information is extracted from the first action data and / or the first text data and / or the first selection data.
[0016] Optionally, the step of selecting a first interaction behavior explanation model corresponding to the first interaction scenario includes:
[0017] Based on the pre-established correspondence between key information and interaction behavior explanation models, selecting an interaction behavior explanation model that matches the first key information;
[0018] Determine the interaction behavior explanation model that matches the first key information as the first interaction behavior explanation model corresponding to the first interaction scenario.
[0019] Optionally, the step of modifying the first keyword that meets a preset condition in the first recognition result by using the first interaction behavior explanation model includes:
[0020] extracting a first keyword that meets a preset condition from the first recognition result;
[0021] Extracting a first character portrait from the first interactive behavior explanation model;
[0022] The first keyword is modified according to the first character portrait.
[0023] Optionally, the step of determining the first interaction scenario to which the first voice data belongs includes:
[0024] extracting first attribute information from the first voice data;
[0025] The first interaction scenario is determined according to the first attribute information.
[0026] Optionally, the first attribute information includes: a collection tool, a collection method, a collection time, a collection location, the number of people, and a semantic environment of the first voice data.
[0027] Optionally, after the step of modifying the first keyword that meets a preset condition in the first recognition result by using the first interaction behavior explanation model, the method further includes:
[0028] outputting a first modified result of the first keyword in the form of voice or text;
[0029] receiving user feedback on the correction result;
[0030] When the evaluation feedback is a positive value, increasing the priority of the first interaction behavior explanation model corresponding to the first interaction scenario;
[0031] When the evaluation feedback is a negative value, the priority of the first interaction behavior explanation model corresponding to the first interaction scenario is lowered.
[0032] Optionally, before the step of receiving the first voice data, the method further includes:
[0033] Determining a relationship between a first user and a second user from the user group, and generating a first relationship tag using the unique identity tags of the first user and the second user;
[0034] Acquire first interaction behavior data between the first user and the second user;
[0035] Constructing a first persona portrait of the first user, a second persona portrait of the second user, and an interaction behavior database between the first user and the second user based on the first interaction behavior data and the first relationship label;
[0036] Repeat the above operations until all users have established character role portraits according to different roles, and an interaction behavior database has been established between different character role portraits, and key information is extracted from the interaction behavior database;
[0037] Inputting the data in the interactive behavior database into a trained neural network to obtain multiple interactive behavior explanation models;
[0038] A corresponding relationship between the key information and the multiple interaction behavior explanation models is established.
[0039] Another aspect of the present invention provides a natural language interpretation device based on character portraits, comprising: a speech receiving module, a processing module, a speech recognition module, and a result correction module;
[0040] The voice receiving module is used to receive first voice data;
[0041] The processing module is configured to determine a first interaction scenario to which the first voice data belongs, and select a first interaction behavior interpretation model corresponding to the first interaction scenario;
[0042] The speech recognition module is configured to perform speech recognition on the first speech data to obtain a first recognition result;
[0043] The result correction module is configured to correct the first keyword that meets a preset condition in the first recognition result by using the first interaction behavior explanation model;
[0044] Among them, the first interactive behavior explanation model includes the association relationship between interactive scene information and character portrait information.
[0045] The third aspect of the present invention provides a computer-readable storage medium, which stores at least one instruction, at least one program, code set or instruction set. The at least one instruction, at least one program, code set or instruction set is loaded and executed by a processor to implement any of the natural language interpretation methods based on character portraits as described above.
[0046] Using the technical solution of the present invention, a natural language interpretation method based on character portraits includes: receiving first voice data; determining a first interaction scenario to which the first voice data belongs; selecting a first interaction behavior interpretation model corresponding to the first interaction scenario; performing voice recognition on the first voice data to obtain a first recognition result; and using the first interaction behavior interpretation model to correct a first keyword in the first recognition result that meets a preset condition. By determining the first interaction behavior interpretation model corresponding to the first interaction scenario and using the first interaction behavior interpretation model to correct the first keyword in the first recognition result of the first voice data, the efficiency of voice recognition / language interpretation is improved while also improving the accuracy of voice recognition / language interpretation. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 This is a flow chart of a natural language interpretation method based on character portraits provided by one embodiment of the present invention;
[0048] Figure 2 is a flowchart after the step of determining the first interaction scenario to which the first voice data belongs in another embodiment of the present invention;
[0049] Figure 3 is a specific execution flow chart of the step of selecting the first interaction behavior explanation model corresponding to the first interaction scenario in another embodiment of the present invention;
[0050] Figure 4 is a specific execution flow chart of the step of modifying the first keyword that meets the preset condition in the first recognition result by using the first interactive behavior explanation model in another embodiment;
[0051] Figure 5is a flowchart after the step of modifying the first keyword that meets the preset condition in the first recognition result by using the first interactive behavior explanation model in another embodiment;
[0052] Figure 6 is a flow chart of a method for constructing an interactive behavior explanation model in another embodiment;
[0053] Figure 7 It is a schematic block diagram of a natural language interpretation device based on character portrait provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0054] In order to more clearly understand the above-mentioned objects, features and advantages of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that, in the absence of conflict, the embodiments of the present application and the features therein can be combined with each other.
[0055] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Therefore, the scope of protection of the present invention is not limited to the specific embodiments disclosed below.
[0056] The terms "first," "second," and the like in the specification and claims of this application and the accompanying drawings are used to distinguish between different objects, not to describe a particular order. Furthermore, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements is not limited to the listed steps or elements but may optionally include steps or elements not listed, or may optionally include other steps or elements inherent to the process, method, product, or apparatus.
[0057] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0058] Refer to the following Figures 1 to 7 To describe a natural language interpretation method, device and storage medium based on character portrait provided according to some embodiments of the present invention.
[0059] like Figure 1 As shown, one embodiment of the present invention provides a natural language interpretation method based on character portraits, comprising:
[0060] receiving first voice data;
[0061] Determining a first interaction scenario to which the first voice data belongs;
[0062] Selecting a first interaction behavior explanation model corresponding to the first interaction scenario;
[0063] performing speech recognition on the first speech data to obtain a first recognition result;
[0064] Using the first interactive behavior explanation model, modifying the first keyword that meets the preset condition in the first recognition result;
[0065] Among them, the first interactive behavior explanation model includes the association relationship between interactive scene information and character portrait information.
[0066] It is understandable that the natural language interpretation method based on character portraits provided by the embodiment of the present invention can be applied to smart terminals such as smart phones, computers, smart TVs, etc., and can also be used in intercom equipment, robots, access control systems, etc.
[0067] In an embodiment of the present invention, the first voice data can be acquired by a voice acquisition unit (such as a microphone), or by an acquisition from a server or an intelligent terminal via a communication network. During the acquisition of the first voice data, relevant information about the scene in which the voice occurs is simultaneously stored as first attribute information of the first voice data.
[0068] It should be noted that after receiving the first voice data, the first interaction scenario to which the first voice data belongs can be determined based on the first attribute information carried by the first voice data. For example, if the first attribute information of the first voice data is the collection location, the building corresponding to the coordinates (such as home, company, shopping mall, etc.) can be determined based on the collection location coordinates. Assuming the collection location is a company, combined with other first attribute information such as the collection time (such as Monday at 10:00 am) and the number of people (such as 5 people; this can be determined based on voiceprint characteristics), the interaction scenario to which the first voice data belongs can be determined to be a "company meeting." Depending on the actual application scenario, interaction scenarios may include but are not limited to: family chats, work discussions, shopping, and gatherings with friends.
[0069] Furthermore, a first interaction behavior interpretation model corresponding to the first interaction scenario is selected, wherein the first interaction behavior interpretation model includes an association relationship between interaction scenario information and character portrait information, so that the first recognition result can be specifically interpreted or corrected by determining the corresponding character portrait information.
[0070] In an embodiment of the present invention, voice recognition is performed on the first voice data. The voice recognition module can segment the first voice data according to different voiceprints, or segment the first voice data according to a preset time length, or segment the first voice data according to a preset file size. Each segmented voice segment is queued in the order of the time when the voice occurs, and a voice recognition algorithm is used to convert each voice segment into corresponding text information according to the queue sequence; the text information is fused in chronological order and adjusted according to the context to obtain a first recognition result.
[0071] For the first keyword (such as local terms, industry jargon, professional terms, etc.) in the first recognition result that meets the preset conditions (such as the frequency of occurrence and / or the frequency of error is within a preset range), the first interactive behavior interpretation model is used to interpret / correct the first keyword to obtain a first interpretation / correction result, such as using the character portrait information it contains and combining the character's characteristics to interpret / correct the first keyword.
[0072] Using the technical solution of this embodiment, the persona-based natural language interpretation method includes: receiving first voice data; determining a first interaction scenario to which the first voice data belongs; selecting a first interaction behavior interpretation model corresponding to the first interaction scenario; performing voice recognition on the first voice data to obtain a first recognition result; and using the first interaction behavior interpretation model to correct a first keyword in the first recognition result that meets a preset condition. By determining the first interaction behavior interpretation model corresponding to the first interaction scenario and using the first interaction behavior interpretation model to interpret / correct the first keyword in the first recognition result of the first voice data, the efficiency of voice recognition / language interpretation is improved while also improving the accuracy of voice recognition / language interpretation.
[0073] like Figure 2 As shown, in some possible implementations of the present invention, after the step of determining the first interaction scenario to which the first voice data belongs, the method further includes:
[0074] According to the first interaction scenario, prompting the user to perform a first action and / or prompting the user to speak first text data and / or prompting the user to input first selection data;
[0075] collecting the first action data and / or the first text data and / or the first selection data;
[0076] First key information is extracted from the first action data and / or the first text data and / or the first selection data.
[0077] It should be noted that in order to further clarify the characteristics of the user's interactive behavior in the first interactive scenario and to select the optimal interactive behavior explanation model, in an embodiment of the present invention, some interactive events are constructed according to the first interactive scenario to obtain the user's interactive behavior in the first interactive scenario, such as issuing an instruction to prompt the user to perform a first action, and / or providing a first text data and issuing an instruction to prompt the user to speak the first text data, and / or providing some options and issuing an instruction to prompt the user to enter the first selection data, etc. By collecting the first action data and / or the first text data and / or the first selection data, and extracting the first key information therefrom. The first key information can be a specific action, a special tone, a special pronunciation or a special preference, etc., and the embodiments of the present invention are not limited to this.
[0078] It can be understood that, based on the first interaction scenario, the more interaction events are constructed, the more comprehensive the event types are covered, and the more interaction behavior data are obtained, the more accurate the subsequent selection of the interaction behavior explanation model will be.
[0079] like Figure 3 As shown, in some possible implementations of the present invention, the step of selecting a first interaction behavior explanation model corresponding to the first interaction scenario includes:
[0080] Based on the pre-established correspondence between key information and interaction behavior explanation models, selecting an interaction behavior explanation model that matches the first key information;
[0081] Determine the interaction behavior explanation model that matches the first key information as the first interaction behavior explanation model corresponding to the first interaction scenario.
[0082] It is understandable that in an embodiment of the present invention, by analyzing the interactive behavior data under different historical interactive scenarios and processing them through a neural network, a correspondence between key information and the interactive behavior explanation model is established. As described above in explaining the first key information, the key information can be a specific action, a special tone, a special pronunciation, or a special preference, etc. Based on the correspondence between the key information and the interactive behavior explanation model, the interactive behavior explanation model that matches the first key information is selected and used as the first interactive behavior explanation model corresponding to the first interactive scenario. Through this solution, the first interactive behavior explanation model can be selected quickly and accurately, which improves execution efficiency and enhances user experience.
[0083] like Figure 4 As shown, in some possible implementations of the present invention, the step of using the first interactive behavior explanation model to correct the first keyword that meets the preset condition in the first recognition result includes:
[0084] extracting a first keyword that meets a preset condition from the first recognition result;
[0085] Extracting a first character portrait from the first interactive behavior explanation model;
[0086] The first keyword is modified according to the first character portrait.
[0087] It is understandable that in an embodiment of the present invention, the first recognition result can be text information, and a first keyword (such as a term with local characteristics, industry jargon, professional terminology, etc.) that meets a preset condition (such as an occurrence frequency and / or an error frequency within a preset range) is extracted from the first recognition result; then a first character role portrait is extracted from the first interactive behavior explanation model, and the character labels contained in the first character role portrait (such as permanent residence, industry, accent characteristics, gender, character relationships, etc.) are used to comprehensively analyze the first recognition result, and when there is an error, the first keyword is interpreted / corrected to obtain a first interpretation / correction result. In this embodiment, the character role portrait is used to conduct a targeted analysis of the first recognition result and to correct the first keyword, which greatly improves the recognition readiness rate.
[0088] In some possible implementations of the present invention, the step of determining the first interaction scenario to which the first voice data belongs includes:
[0089] extracting first attribute information from the first voice data;
[0090] The first interaction scenario is determined according to the first attribute information.
[0091] It can be understood that, as mentioned above, in the process of collecting the first voice data, the relevant information of the voice occurrence scene is also saved as the first attribute information of the first voice data. Specifically, the sound data and the first attribute information are packaged to form the first voice data, or the data format of the sound data is modified, and a part is added to record the first attribute information to form the first voice data.
[0092] Among them, the first attribute information includes: the collection tool of the first voice data (such as mobile phone, drone, office robot, smart camera, etc.), collection method (such as direct collection through the device, collection through network connection to other devices, etc.), collection time (such as 6 o'clock in the morning, 9 o'clock in the morning, 3 o'clock in the afternoon, 8 o'clock in the evening, etc.), collection location (such as company, home, shopping mall, hospital, school, etc.), number of people and semantic environment (mainly including the preface and context of expression and comprehension).
[0093] In the embodiment of the present invention, by recording relevant information of the speech occurrence scene, an additional reference dimension is provided for subsequent speech recognition, thereby improving the efficiency and accuracy of speech recognition.
[0094] like Figure 5 As shown, in some possible implementations of the present invention, after the step of using the first interactive behavior explanation model to correct the first keyword that meets the preset condition in the first recognition result, the following step is further included:
[0095] outputting a first modified result of the first keyword in the form of voice or text;
[0096] receiving user feedback on the correction result;
[0097] When the evaluation feedback is a positive value, increasing the priority of the first interaction behavior explanation model corresponding to the first interaction scenario;
[0098] When the evaluation feedback is a negative value, the priority of the first interaction behavior explanation model corresponding to the first interaction scenario is lowered.
[0099] It can be understood that in order to improve recognition efficiency, an embodiment of the present invention sets up a feedback mechanism, which outputs the first interpretation / correction result of the first keyword in the form of voice or text, and provides an interactive interface for users to evaluate and feedback on the recognition results; receives user evaluation feedback on the correction result, and adjusts the priority of the first interactive behavior explanation model corresponding to the first interactive scenario according to the positive and negative values represented by the evaluation feedback.
[0100] like Figure 6 As shown, in some possible implementations of the present invention, before the step of receiving the first voice data, the method further includes:
[0101] S1. Determine a relationship between a first user and a second user from a user group, and generate a first relationship tag using the unique identity tags of the first user and the second user respectively;
[0102] It is understood that each user has a unique identity tag, and the relationship between users can be a role relationship such as parent-child, husband-wife, friend-colleague, or other basic relationships. The user's role tag can be constructed based on their unique identity tag using preset rules, such as adding a role field after the unique identity tag. The first relationship tag can be constructed by combining the role tags of the first and second users. In this step, any two different users are randomly selected as the first and second users.
[0103] S2. Acquire first interaction behavior data between the first user and the second user;
[0104] In this step, interactive behaviors include chatting, discussing, teaching, commanding, etc. The first interactive behavior data between the first user and the second user is extracted from the interactive behavior data between multiple users / roles (such as voice, action, text, geographic location, role distance, number of simultaneous participants, background noise, etc.).
[0105] S3. Constructing a first persona portrait of the first user, a second persona portrait of the second user, and an interaction behavior database between the first user and the second user based on the first interaction behavior data and the first relationship label;
[0106] In this step, character portraits are created based on vocabulary, emotion, age, gender, education level, accent, hobbies, etc. Based on the character portraits and necessary technical means (keyword recognition, emotion recognition, and attitude analysis, etc.), a database of interaction behaviors between two characters is established. The interaction behavior database contains the relationship tags of each interaction behavior (as constructed by the first relationship tag mentioned above);
[0107] S4. Repeat the above operations until all users have established character portraits according to different roles, and an interaction behavior database has been established between different character portraits, and key information is extracted from the interaction behavior database;
[0108] In this step, the key information, such as the first key information mentioned above, can be a specific action, a special tone, a special pronunciation or a special preference, etc.
[0109] S5. Inputting the data in the interactive behavior database into a trained neural network to obtain multiple interactive behavior explanation models;
[0110] S6. Establishing a corresponding relationship between the key information and the multiple interaction behavior explanation models.
[0111] This embodiment is based on the role relationship between users and the interactive behavior data as the analysis basis, and uses a pre-trained neural network to generate and construct an interactive behavior explanation model. The execution method is simple, and the generated interactive behavior explanation model can also provide accurate explanation results.
[0112] In some possible implementations of the present invention, a verification mechanism is established by setting multiple interaction behavior interpretation models corresponding to the same interaction scenario. After the first keyword in the first recognition result of the first voice data is interpreted / corrected using the multiple interaction behavior interpretation models, the user selects the most correct interpretation / correction result. Specifically, after the first interaction behavior interpretation model is used to interpret / correct the first keyword in the first recognition result that meets the preset conditions and obtains the first interpretation / correction result, the following is also included:
[0113] Selecting a second interaction behavior explanation model corresponding to the first interaction scenario;
[0114] Using the second interactive behavior interpretation model, interpreting / correcting the first keyword that meets the preset condition in the first recognition result to obtain a second interpretation / correction result;
[0115] The first interpretation / correction result and the second interpretation / correction result are presented to the user, and the user selects the optimal one. Based on the user's selection, the priority levels of the first interaction behavior explanation model and the second interaction behavior explanation model are adjusted.
[0116] like Figure 7 As shown, another embodiment of the present invention provides a natural language interpretation device 700 based on character portrait, comprising: a speech receiving module 701, a processing module 702, a speech recognition module 703 and a result correction module 704;
[0117] The voice receiving module 701 is used to receive first voice data;
[0118] The processing module 702 is configured to determine a first interaction scenario to which the first voice data belongs, and select a first interaction behavior interpretation model corresponding to the first interaction scenario;
[0119] The speech recognition module 703 is configured to perform speech recognition on the first speech data to obtain a first recognition result;
[0120] The result correction module 704 is configured to correct the first keyword that meets a preset condition in the first recognition result by using the first interaction behavior explanation model;
[0121] Among them, the first interactive behavior explanation model includes the association relationship between interactive scene information and character portrait information.
[0122] The operating method of the device provided in this embodiment can be found in the aforementioned method embodiments and will not be described in detail here.
[0123] Figure 7 This is a schematic diagram of the hardware composition of the device in this embodiment. It can be understood that Figure 7 Only a simplified design of the device is shown. In actual applications, the device may also include other necessary components, including but not limited to any number of input / output systems, processors, controllers, memories, etc., and all devices that can implement the natural language interpretation method of the embodiments of the present application are within the scope of protection of this application.
[0124] Another embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores at least one instruction, at least one program, code set or instruction set, and the at least one instruction, the at least one program, the code set or instruction set is loaded and executed by a processor to implement any of the methods described above.
[0125] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.
[0126] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0127] In the several embodiments provided in this application, it should be understood that the disclosed devices can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical or other forms.
[0128] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0129] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0130] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a memory and includes a number of instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the above-mentioned methods of each embodiment of the present application. The aforementioned memory includes: various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.
[0131] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a program. The program can be stored in a computer-readable memory, and the memory can include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0132] The above is a detailed introduction to the embodiments of the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core idea of the present application. At the same time, for those skilled in the art, according to the idea of the present application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
[0133] Although the present invention is disclosed above, it is not limited thereto. Any person skilled in the art may readily conceive of variations or substitutions, and may make various modifications and alterations without departing from the spirit and scope of the present invention. Combinations of the above-described functions and implementation steps, including software and hardware implementations, are all within the scope of protection of the present invention.
Claims
1. A natural language interpretation method based on character portrait, characterized in that: include: receiving first voice data; Determining a first interaction scenario to which the first voice data belongs; Selecting a first interaction behavior explanation model corresponding to the first interaction scenario; performing speech recognition on the first speech data to obtain a first recognition result; Using the first interactive behavior explanation model, modifying the first keyword that meets the preset condition in the first recognition result; The first interactive behavior explanation model includes the association between interactive scene information and character portrait information; The natural language interpretation method further comprises: before the step of receiving the first voice data, Determining a relationship between a first user and a second user from the user group, and generating a first relationship tag using the unique identity tags of the first user and the second user; Acquire first interaction behavior data between the first user and the second user; Constructing a first persona portrait of the first user, a second persona portrait of the second user, and an interaction behavior database between the first user and the second user based on the first interaction behavior data and the first relationship label; Repeat the above operations until all users have established character role portraits according to different roles, and an interaction behavior database has been established between different character role portraits, and key information is extracted from the interaction behavior database; Inputting the data in the interactive behavior database into a trained neural network to obtain multiple interactive behavior explanation models; Establishing a corresponding relationship between the key information and the multiple interaction behavior explanation models; The step of modifying the first keyword that meets the preset condition in the first recognition result by using the first interactive behavior explanation model includes: extracting a first keyword that meets a preset condition from the first recognition result; Extracting a first character portrait from the first interactive behavior explanation model; The first keyword is modified according to the first character portrait.
2. The natural language interpretation method based on character portrait according to claim 1, characterized in that: After the step of determining the first interaction scenario to which the first voice data belongs, the method further includes: According to the first interaction scenario, prompting the user to perform a first action and / or prompting the user to speak first text data and / or prompting the user to input first selection data; collecting the first action data and / or the first text data and / or the first selection data; First key information is extracted from the first action data and / or the first text data and / or the first selection data.
3. The natural language interpretation method based on character portrait according to claim 2, characterized in that: The step of selecting a first interaction behavior explanation model corresponding to the first interaction scenario includes: Based on the pre-established correspondence between key information and interaction behavior explanation models, selecting an interaction behavior explanation model that matches the first key information; Determine the interaction behavior explanation model that matches the first key information as the first interaction behavior explanation model corresponding to the first interaction scenario.
4. The natural language interpretation method based on character portrait according to claim 3, characterized in that: The step of determining the first interaction scenario to which the first voice data belongs includes: extracting first attribute information from the first voice data; The first interaction scenario is determined according to the first attribute information.
5. The natural language interpretation method based on character portrait according to claim 4, characterized in that: The first attribute information includes: the collection tool, collection method, collection time, collection location, number of people and semantic environment of the first voice data.
6. The natural language interpretation method based on character portrait according to claim 5, characterized in that: After the step of modifying the first keyword that meets the preset condition in the first recognition result by using the first interaction behavior explanation model, the method further includes: outputting a first modified result of the first keyword in the form of voice or text; receiving user feedback on the correction result; When the evaluation feedback is a positive value, increasing the priority of the first interaction behavior explanation model corresponding to the first interaction scenario; When the evaluation feedback is a negative value, the priority of the first interaction behavior explanation model corresponding to the first interaction scenario is lowered.
7. A natural language interpretation device based on character portrait, characterized in that: include: Speech receiving module, processing module, speech recognition module and result correction module; The voice receiving module is used to receive first voice data; The processing module is configured to determine a first interaction scenario to which the first voice data belongs, and select a first interaction behavior interpretation model corresponding to the first interaction scenario; The speech recognition module is configured to perform speech recognition on the first speech data to obtain a first recognition result; The result correction module is configured to correct the first keyword that meets a preset condition in the first recognition result by using the first interaction behavior explanation model; The first interactive behavior explanation model includes the association between interactive scene information and character portrait information; The natural language interpretation device is further configured to, before receiving the first voice data, Determining a relationship between a first user and a second user from the user group, and generating a first relationship tag using the unique identity tags of the first user and the second user; Acquire first interaction behavior data between the first user and the second user; Constructing a first persona portrait of the first user, a second persona portrait of the second user, and an interaction behavior database between the first user and the second user based on the first interaction behavior data and the first relationship label; Repeat the above operations until all users have established character role portraits according to different roles, and an interaction behavior database has been established between different character role portraits, and key information is extracted from the interaction behavior database; Inputting the data in the interactive behavior database into a trained neural network to obtain multiple interactive behavior explanation models; Establishing a corresponding relationship between the key information and the multiple interaction behavior explanation models; The step of modifying the first keyword that meets the preset condition in the first recognition result by using the first interactive behavior explanation model includes: extracting a first keyword that meets a preset condition from the first recognition result; Extracting a first character portrait from the first interactive behavior explanation model; The first keyword is modified according to the first character portrait.
8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set are loaded and executed by the processor to implement the natural language interpretation method based on character portrait as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Voice recommendation language display method, device and system and electronic equipment
CN112927686A