A personalized voice interaction method and system
By performing feature recognition on user behavior data, a response text matching the total score of the user behavior data is generated, which solves the problem of the lack of logic in existing voice interaction systems and realizes personalized voice interaction.
Patent Information
- Application Number
- CN202210763766.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-30
- Publication Date
- 2025-12-19
- Estimated Expiration
- 2042-06-30
AI Technical Summary
Existing voice interaction systems lack logical structure, making it difficult to engage with users specifically and unable to generate personalized responses based on user behavior data.
By performing feature recognition on user behavior data, a response content matching the user's total behavior data score is generated. The feature recognition model generated using the NLG algorithm is used to generate the response text matching the user's total behavior data score.
It enables voice interaction with users, and the generated responses are tailored to the user's personalized characteristics and are logical.
Smart Images

Figure CN115188376B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of natural language processing, and in particular to a personalized voice interaction method and system. BACKGROUND
[0002] With the continuous popularity of voice interaction technology, current cars are usually equipped with a voice interaction system, which can respond to the collected voice data of the user and realize voice interaction with the user. The existing voice interaction system usually uses a general corpus, and when receiving the voice data of the user, a statement is randomly selected from the general corpus to respond, which lacks logic and thus is difficult to conduct targeted voice interaction with the user. SUMMARY
[0003] The present application provides a personalized voice interaction method and system to solve the problem that the existing voice interaction system is difficult to conduct targeted voice interaction with the user. By performing feature recognition on the behavior data of the user, the personalized features and the total behavior data score of the user are obtained, and then based on the text data in the voice data of the user and the personalized features, a response text matching the total behavior data score of the user is generated, and the response text is converted into audio data, so that the user receives the response text in the form of audio, thereby realizing voice interaction with the user, and the response content is consistent with the personalized features of the user and has logic.
[0004] To solve the above technical problem, the first aspect of the embodiment of the present application provides a personalized voice interaction method, comprising the following steps:
[0005] In response to a voice interaction instruction of a user, behavior data of the user is collected; wherein the behavior data at least includes voice data;
[0006] The behavior data is input into a preset feature recognition model for feature recognition, and based on a preset score value corresponding to each user behavior, personalized features and a total behavior data score of the user are obtained;
[0007] Based on a preset text generation model, text data in the voice data is extracted, and based on the text data and the personalized features, a response text matching the total behavior data score is generated based on feature tags and score tags of each text in a preset corpus, and the response text is converted into audio data.
[0008] As a preferred scheme, the behavior data is input into a preset feature recognition model for feature recognition, and based on a preset score value corresponding to each user behavior, personalized features and a total behavior data score of the user are obtained, which specifically comprises the following steps:
[0009] input the behavior data into the feature recognition model for feature recognition to obtain the personalized feature of the user;
[0010] Based on the preset score value corresponding to each user behavior, the score value of each behavior data is obtained, and according to the score value of each behavior data, the total score value of the behavior data of the user is obtained according to the preset scoring rule.
[0011] As a preferred solution, the response text matched to the total score value of the behavior data is generated based on the feature mark and score mark of each text in the preset corpus according to the text data and the personalized feature, and specifically includes the following steps:
[0012] Based on the feature mark and score mark of each text in the preset corpus, a plurality of texts matched to the text data and the personalized feature are obtained in the preset corpus by using an NLG algorithm;
[0013] According to the total score value of the behavior data and the score mark of the plurality of texts, the plurality of texts are screened to obtain a plurality of screened texts; wherein the score mark of the plurality of screened texts matches the total score value of the behavior data;
[0014] The response text is generated according to the plurality of screened texts.
[0015] As a preferred solution, the method specifically obtains the feature recognition model by the following steps:
[0016] The behavior data with personalized feature mark and score value mark is composed into a training set, and the convolutional neural network is trained by using the training set to obtain the feature recognition model.
[0017] As a preferred solution, the behavior data of the user is collected in response to the voice interaction instruction of the user, and specifically includes the following steps:
[0018] The voice data of the user is collected by the voice acquisition module in response to the voice interaction instruction of the user.
[0019] As a preferred solution, the behavior data further includes image data and central control configuration data.
[0020] As a preferred solution, the behavior data of the user is collected in response to the voice interaction instruction of the user, and specifically includes the following steps:
[0021] The image data of the user is collected by the image acquisition module in response to the voice interaction instruction of the user;
[0022] The central control configuration data of the user is collected by the central control module.
[0023] As a preferred solution, the personalized features at least include age, gender, time, emotional features, preference features and scene environment.
[0024] As a preferred solution, the method further comprises the following steps:
[0025] The personalized features of the user and the behavior data score total value are transmitted to a preset database, so that the personalized features and the behavior data score total value are stored in the database.
[0026] A second aspect of the embodiment of the present application provides a personalized voice interaction system, comprising:
[0027] A behavior data collection module is configured to collect behavior data of a user in response to a voice interaction instruction of the user, wherein the behavior data at least includes voice data;
[0028] A personalized feature identification module is configured to input the behavior data into a preset feature identification model for feature identification, and obtain personalized features of the user and a behavior data score total value based on a preset score value corresponding to each user behavior.
[0029] A response text generation module is configured to extract text data in the voice data based on a preset text generation model, generate a response text matched with the behavior data score total value based on feature labels and score labels of each text in a preset corpus according to the text data and the personalized features, and convert the response text into audio data.
[0030] Compared with the prior art, the embodiment of the present application has the beneficial effects that the personalized features of the user and the behavior data score total value are obtained by feature identification of the behavior data of the user, and then the response text matched with the behavior data score total value of the user can be generated based on the text data in the voice data of the user and the personalized features, and the response text is converted into audio data, so that the user receives the response text in the form of audio, thereby realizing voice interaction with the user, and the response content is consistent with the personalized features of the user and has logic. BRIEF DESCRIPTION OF DRAWINGS
[0031] Figure 1 is a flowchart of the personalized voice interaction method provided by the embodiment of the present application;
[0032] Figure 2 is a structural schematic diagram of the personalized voice interaction system provided by the embodiment of the present application. DETAILED DESCRIPTION
[0033] With reference to the drawings of the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described, obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0034] Referring to Figure 1 The first aspect of the embodiments of the present application provides a personalized voice interaction method, comprising the following steps S1 to S3:
[0035] Step S1, in response to a voice interaction instruction of a user, collecting behavior data of the user; wherein the behavior data at least includes voice data;
[0036] Step S2, inputting the behavior data into a preset feature recognition model for feature recognition, obtaining personalized features of the user and a total score value of the behavior data based on a preset score value corresponding to each user behavior;
[0037] Step S3, based on a preset text generation model, extracting text data in the voice data, generating a response text matched with the total score value of the behavior data according to the text data and the personalized features based on feature labels and score labels of each text in a preset corpus, and converting the response text into audio data.
[0038] In the embodiment, in response to a voice interaction instruction of a user, the behavior data of the user is collected by an information collection module provided in a vehicle, wherein the behavior data of the user at least includes voice data of the user.
[0039] Further, different behaviors of the user represent different personalized features, for example, if the user indicates that he likes to listen to rock music, the personalized features of the user may be a rock music lover and a bold person, therefore, in order to generate a response text that matches the personalized features of the user as much as possible, the behavior data is input into a preset feature recognition model for feature recognition, the personalized features of the user and a total score value of the behavior data are obtained based on a preset score value corresponding to each user behavior, the total score value of the behavior data can represent the current behavior of the user with a quantitative value, which can be used as a basis for judging whether the response text matches the personalized features of the user in the subsequent process of generating the response text.
[0040] Further, the embodiment extracts text data in the voice data based on a preset text generation model, generates response text matching the behavior data score total value of the user based on the feature labels and score labels of each text in the preset corpus according to the text data and the personalized features of the user, and converts the response text into audio data, so that the user receives the response text in the form of audio, thereby realizing voice interaction with the user. It can be understood that the same personalized feature may correspond to multiple texts, but each text has different score labels. At this time, in order to select the text that best matches the personalized features of the user, the text needs to be screened based on the behavior data score total value, so that the score of all screened texts is the score closest to the behavior data score total value, i.e., matching the behavior data score total value, and the screened texts are organized in language, thereby generating response text matching the behavior data score total value of the user.
[0041] The personalized voice interaction method provided by the embodiment can obtain the personalized features and the behavior data score total value of the user by feature recognition of the behavior data of the user, and then generate response text matching the behavior data score total value of the user based on the text data in the voice data and the personalized features of the user, convert the response text into audio data, so that the user receives the response text in the form of audio, thereby realizing voice interaction with the user, and the response content matches the personalized features of the user and has logicality.
[0042] As a preferred solution, the behavior data is input into a preset feature recognition model for feature recognition, and the personalized features and the behavior data score total value of the user are obtained based on a preset score value corresponding to each user behavior. Specifically, the following steps are included:
[0043] The behavior data is input into the feature recognition model for feature recognition, and the personalized features of the user are obtained.
[0044] Based on the preset score value corresponding to each user behavior, the score value of each behavior data is obtained, and the behavior data score total value of the user is obtained according to the score value of each behavior data and a preset scoring rule.
[0045] In the embodiment, based on the preset score value corresponding to each user behavior, the score value of each behavior data can be obtained, and the behavior data score total value of the user is obtained according to the score value of each behavior data and the following expression as a scoring rule:
[0046]
[0047] Wherein, S represents the behavior data score total value, S0 represents the preset initial behavior data score value, N represents the number of behavior data, S1, S2, …, S N represent the score value of the i-th behavior data, 0 < i ≤ N.
[0048] The obtained behavior data score total value corresponds to the digital portrait of the user, forming the specific identity ID of the user.
[0049] As a preferred solution, the response text matched to the behavior data score total value is generated based on the feature label and score label possessed by each text in the preset corpus according to the text data and the personalized features, specifically comprising the following steps:
[0050] Based on the feature label and score label possessed by each text in the preset corpus, a plurality of texts matched to the text data and the personalized features are obtained in the preset corpus by using the NLG algorithm;
[0051] According to the behavior data score total value and the score label of the plurality of texts, the plurality of texts are screened to obtain a plurality of screened texts; wherein the score of the plurality of screened texts matches the behavior data score total value.
[0052] The response text is generated according to the plurality of screened texts.
[0053] It is worth noting that in this embodiment, the working principle of the NLG algorithm is: input abstract propositions, then perform semantic analysis and syntax analysis on the input natural language, combine the personalized features identified by the feature recognition model, perform behavior data score matching, organize language according to the text that best matches the behavior data score total value of the user, and then generate a response text that best fits the personality of the user.
[0054] The NLG algorithm uses the TextRank algorithm, which is a graph-based ranking algorithm for keyword extraction and document summarization, improved from the PageRank algorithm of Google's web importance ranking algorithm. It can extract keywords, keyword groups and key sentences from a given text by using the co-occurrence information (semantics) between words in a document. The text generated by the TextRank algorithm does not have the feature attributes of the user, and the texts in the corpus need to be manually labeled with feature labels and score labels in advance. After manual labeling, the behavior data score matching can be performed in combination with the dynamic personalized features of the user to screen the most suitable response text.
[0055] The basic idea of the TextRank algorithm is to regard a document as a network of words, and the links in the network represent the semantic relationship between the words. The algorithm mainly includes keyword extraction, key phrase extraction, and key sentence extraction.
[0056] Keyword extraction refers to a process of determining some terms from a text that can describe the meaning of the document. For keyword extraction, the text units used to construct the vertex set can be one or more words in a sentence; and the edges are constructed according to the relationship between the words (such as: appearing in a box at the same time). According to the needs of the task, the vertex set can be optimized using syntactic filters. The main role of the syntactic filters is to filter out words of a certain part of speech or several parts of speech as the vertex set.
[0057] After keyword extraction, N keywords can be obtained, and adjacent keywords in the original text form a key phrase.
[0058] The sentence extraction task is mainly aimed at the automatic summary scenario, and each sentence is taken as a vertex. The similarity between two sentences is calculated according to the content repetition degree, and the similarity is used as the connection. Since the similarities between different sentences are inconsistent, a weighted graph is constructed in this scenario, with the similarity size as the edge weight.
[0059] It is worth noting that the text generation model of the embodiment of the present application is based on the NLG algorithm, and the NLG algorithm is used to extract keywords in the voice data, thereby forming text data.
[0060] In the embodiment, based on the feature tags and score tags of each text in the preset corpus, the NLG algorithm is used to obtain, in the preset corpus, a plurality of texts matching the text data extracted by the text generation model using the NLG algorithm and matching the personalized features; then, according to the behavior data score total value and the score tags of the plurality of texts, the plurality of texts are screened to obtain a plurality of screened texts, and the total score values of the screened texts match the behavior data score total value; and the response text is generated according to the plurality of screened texts.
[0061] It is worth noting that the screened texts are calculated according to the same scoring rules as the behavior data to ensure that the finally generated response text is as close as possible to the personalized features of the user.
[0062] As a preferred solution, the feature recognition model is obtained by the following steps:
[0063] The preset behavior data with the personalized feature label and the score value label is composed into a training set, and the convolutional neural network is trained by using the training set, so as to obtain the feature recognition model.
[0064] It is worth noting that, due to the distortion of data such as voice data and image data during the driving of the vehicle, in order to improve the stability and accuracy of feature recognition, the embodiment adopts a convolutional neural network composed of two convolutional layers, two pooling layers and three fully connected layers, the number of neurons of the three fully connected layers is 128, 32 and 1 respectively, the first two layers use a Relu activation function, and the last layer outputs a similarity value of a state.
[0065] As a preferred solution, the behavior data of the user is collected in response to the voice interaction instruction of the user, and the method specifically comprises the following steps:
[0066] The voice data of the user is collected by the voice acquisition module in response to the voice interaction instruction of the user.
[0067] As one of the optional embodiments, the voice acquisition module is a front-mounted microphone or a rear-mounted microphone arranged in the vehicle, and the voice data of the user is collected through the front-mounted microphone or the rear-mounted microphone.
[0068] As a preferred solution, the behavior data further comprises image data and central control configuration data.
[0069] It is worth noting that, since the central control module of the vehicle is a module for controlling the air conditioner, the audio and other comfort and entertainment devices of the vehicle, the behavior data of the user in the entertainment and learning aspects can be obtained by collecting the central control configuration data, for example, the user can control the audio of the vehicle to play the music he likes through the central control module, and the music style preferred by the user can be obtained by collecting the central control configuration data as one of the personalized features of the user.
[0070] As a preferred solution, the behavior data of the user is collected in response to the voice interaction instruction of the user, and the method specifically further comprises the following steps:
[0071] The image data of the user is collected by the image acquisition module in response to the voice interaction instruction of the user.
[0072] The central control configuration data of the user is collected by the central control module.
[0073] As one of the optional embodiments, the image acquisition module is a front-mounted camera or a rear-mounted camera arranged in the vehicle, and the image data of the user can be collected by controlling the shooting angle of the front-mounted camera or the rear-mounted camera.
[0074] As a preferred solution, the personalized features at least include age, gender, time, emotional features, preference features and scene environment.
[0075] As a preferred solution, the method further comprises the following steps:
[0076] The personalized features and the behavior data score total value of the user are transmitted to a preset database, so that the personalized features and the behavior data score total value are stored in the database.
[0077] It is worth noting that the personalized features and the behavior data score total value stored in the database can be used for the next training of the feature recognition model, and through a large number of training, the recognition accuracy of the feature recognition model can be continuously improved.
[0078] Referring to Figure 2 , the second aspect of the embodiment of the present application provides a personalized voice interaction system, comprising:
[0079] The behavior data collection module 201 is configured to collect behavior data of a user in response to a voice interaction instruction of the user, wherein the behavior data at least includes voice data;
[0080] The personalized feature recognition module 202 is configured to input the behavior data into a preset feature recognition model for feature recognition, obtain personalized features of the user and a behavior data score total value of the user based on a preset score value corresponding to each user behavior.
[0081] The response text generation module 203 is configured to extract text data in the voice data based on a preset text generation model, generate a response text matched with the behavior data score total value according to the text data and the personalized features based on feature labels and score labels of each text in a preset corpus, and convert the response text into audio data.
[0082] As a preferred solution, the personalized feature recognition module 202 is configured to input the behavior data into a preset feature recognition model for feature recognition, obtain personalized features of the user and a behavior data score total value of the user based on a preset score value corresponding to each user behavior, and specifically comprises:
[0083] The behavior data is input into the feature recognition model for feature recognition to obtain personalized features of the user.
[0084] Based on a preset score value corresponding to each user behavior, a score value of each behavior data is obtained, and a behavior data score total value of the user is obtained according to the score value of each behavior data and a preset scoring rule.
[0085] As a preferred solution, the response text generation module 203 is configured to generate, according to the text data and the personalized features, a response text matching the behavior data score total value based on feature labels and score labels of each text in a preset corpus, specifically comprising:
[0086] Based on the feature labels and score labels of each text in the preset corpus, a plurality of texts matching the text data and the personalized features are obtained from the preset corpus by using an NLG algorithm;
[0087] According to the behavior data score total value and the score labels of the plurality of texts, the plurality of texts are filtered to obtain a plurality of filtered texts; wherein the score labels of the plurality of filtered texts match the behavior data score total value;
[0088] The response text is generated according to the plurality of filtered texts.
[0089] As a preferred solution, the personalized feature identification module 202 is further configured to obtain the feature identification model by the following steps:
[0090] A training set with preset behavior data having personalized feature labels and score value labels is formed, and the training set is used to train a convolutional neural network to obtain the feature identification model.
[0091] As a preferred solution, the behavior data collection module 201 is configured to collect behavior data of a user in response to a voice interaction instruction of the user, specifically comprising:
[0092] In response to the voice interaction instruction of the user, the voice data of the user is collected by a voice acquisition module.
[0093] As a preferred solution, the behavior data further comprises image data and central control configuration data.
[0094] As a preferred solution, the behavior data collection module 201 is configured to collect behavior data of a user in response to a voice interaction instruction of the user, specifically further comprising:
[0095] In response to the voice interaction instruction of the user, the image data of the user is collected by an image acquisition module;
[0096] The central control configuration data of the user is collected by a central control module.
[0097] As a preferred solution, the personalized features at least include age, gender, time, emotional features, preference features, and scene environment.
[0098] As a preferred solution, the personalized feature identification module 202 is further configured to:
[0099] The personalized features and the behavior data score total value of the user are transmitted to a preset database 204, so that the personalized features and the behavior data score total value are stored in the database 204.
[0100] As a preferred solution, the system further comprises a control module 205, configured to:
[0101] receive the voice interaction instruction of the user, and send the voice interaction instruction to the behavior data acquisition module 201;
[0102] send the acquired behavior data to the personalized feature identification module 202.
[0103] The personalized voice interaction system provided by the embodiment of the present application can obtain the personalized features and the behavior data score total value of the user by performing feature identification on the behavior data of the user, and can further generate a response text matched with the behavior data score total value of the user based on the text data in the voice data of the user and the personalized features, and convert the response text into audio data, so that the user receives the response text in the form of audio, thereby realizing voice interaction with the user, and the response content is consistent with the personalized features of the user and has logic.
[0104] The above describes the preferred embodiments of the present application, and it should be noted that, for those skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, and these improvements and refinements are also considered within the protection scope of the present application.
Claims
1. A method of personalized voice interaction, characterized by, The method comprises the following steps: In response to the voice interaction instruction of the user, the behavior data of the user is collected; wherein the behavior data at least comprises voice data; The behavior data is input into a preset feature recognition model for feature recognition, and based on the preset score value corresponding to each user behavior, the personalized feature of the user and the behavior data score total value are obtained; Based on the preset text generation model, the text data in the voice data is extracted, and based on the feature mark and score mark of each text in the preset corpus, the response text matching the behavior data score total value is generated according to the text data and the personalized feature, and the response text is converted into audio data; The method obtains the feature recognition model through the following steps: The behavior data with personalized feature mark and score value mark is composed into a training set, and the training set is used to train the convolutional neural network to obtain the feature recognition model.
2. The personalized voice interaction method of claim 1, wherein, The behavior data is input into the feature recognition model for feature recognition, and the personalized feature of the user is obtained; Based on the preset score value corresponding to each user behavior, the score value of each behavior data is obtained, and the behavior data score total value of the user is obtained according to the score value of each behavior data and the preset score rule. The text data and the personalized feature are used to generate the response text matching the behavior data score total value based on the feature mark and score mark of each text in the preset corpus, and the response text is converted into audio data.
3. The personalized voice interaction method of claim 2, wherein, Based on the feature mark and score mark of each text in the preset corpus, NLG algorithm is used to obtain a plurality of texts matching the text data and the personalized feature in the preset corpus; According to the behavior data score total value and the score mark of the plurality of texts, the plurality of texts are screened to obtain a plurality of screened texts; wherein the score of the plurality of screened texts matches the behavior data score total value; The response text is generated according to the plurality of screened texts. In response to the voice interaction instruction of the user, the behavior data of the user is collected, which comprises the following steps:
4. The personalized voice interaction method of claim 1, wherein, In response to the voice interaction instruction of the user, the voice data of the user is collected through the voice acquisition module. The behavior data further comprises image data and central control configuration data.
5. The personalized voice interaction method of claim 1, wherein, In response to the voice interaction instruction of the user, the behavior data of the user is collected, which further comprises the following steps:
6. The personalized voice interaction method of claim 5, wherein, In response to the voice interaction instruction of the user, the image data of the user is collected through the image acquisition module; The central control configuration data of the user is collected through the central control module. The personalized feature at least comprises age, gender, time, emotional feature, preference feature and scene environment.
7. The personalized voice interaction method of claim 1, wherein, The method further comprises the following steps:
8. The personalized voice interaction method of claim 1, wherein, The personalized feature and the behavior data score total value of the user are transmitted to a preset database, so that the personalized feature and the behavior data score total value are stored in the database.
9. A personalized voice interaction system, characterized by The application comprises: a behavior data collection module, configured to collect behavior data of a user in response to a voice interaction instruction of the user, wherein the behavior data at least comprises voice data; a personalized feature identification module, configured to input the behavior data into a preset feature identification model for feature identification, and obtain a personalized feature and a behavior data score total value of the user based on a preset score value corresponding to each user behavior; a response text generation module, configured to extract text data in the voice data based on a preset text generation model, generate a response text matched with the behavior data score total value according to the text data and the personalized feature based on feature labels and score labels of each text in a preset corpus, and convert the response text into audio data; wherein the personalized feature identification module is further configured to obtain the feature identification model by the following steps: a training set with preset behavior data having personalized feature labels and score value labels is formed, and a convolutional neural network is trained by using the training set to obtain the feature identification model.
Citation Information
Patent Citations
Voice interaction method based on emotion engine technology, intelligent terminal and storage medium
CN111368609A
User personalized preference prediction method based on multi-angle non-transfer preference relationship
CN111723290A
Voice interaction method of water heater and water heater
CN114678025A