Information collection method and device, electronic equipment and storage medium

By calculating the similarity between the current text information and the previous text information in the vehicle system and combining it with a sentiment analysis model, the system automatically collects voice and text information that the vehicle system cannot respond to correctly, solving the problem of low recognition efficiency in the vehicle system and achieving efficient text information collection and sentiment judgment.

CN116030814BActive Publication Date: 2026-04-21IFLYTEK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
IFLYTEK CO LTD
Filing Date
2022-12-16
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

In existing technologies, vehicle infotainment systems cannot efficiently identify and record repetitive and incorrectly responding voice and text information, resulting in low efficiency for manual analysis.

Method used

By calculating the similarity score between the current text information and the previous text information, and when the similarity score reaches a certain threshold, the text information is sent to the vehicle system. Combined with the sentiment analysis model to judge the user's emotional changes, the system automatically collects similar text information that the vehicle system cannot respond to correctly.

Benefits of technology

It improves the efficiency of acquiring similar text information that the vehicle's infotainment system cannot respond to correctly, and improves the accuracy of recognition and the accuracy of user emotion judgment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116030814B_ABST
    Figure CN116030814B_ABST
Patent Text Reader

Abstract

The application provides an information collection method and device, electronic equipment and a storage medium, and relates to the technical field of voice processing. The information collection method comprises the following steps: converting obtained current voice information into current text information; in the case that the current text information and previous text information are not recognized, determining a similarity score of the current text information and the previous text information; the previous text information is text information corresponding to previous voice information; a time interval between the current voice information and the previous voice information is less than or equal to a first preset value; and when it is determined that the similarity score is greater than or equal to a second preset value, the current text information is sent to a car machine system. The application realizes automatic collection of similar text information that cannot be correctly responded by the car machine, thereby improving the efficiency of obtaining similar text information that cannot be correctly responded by the car machine.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of speech processing technology, and in particular to an information collection method, apparatus, electronic device, and storage medium. Background Technology

[0002] In-vehicle infotainment systems (IVS) refer to the in-vehicle infotainment products installed inside vehicles. Functionally, IVS enables information communication between people and vehicles, and between vehicles and the outside world (vehicle-to-vehicle communication). Specifically, IVS uses a human-machine interaction system to convert user input speech into text, performs semantic understanding, and thus enables information communication.

[0003] In related technologies, the vehicle system uploads all text information corresponding to the user's voice input to the server. Semantic researchers retrieve all text information within a preset time period from the server, and then manually analyze all text information within the preset time period to find repeated text information that the vehicle system cannot respond to correctly multiple times. The repeated text information that the vehicle system cannot respond to correctly multiple times is fed back to the vehicle system, so that the vehicle system can record and optimize the repeated text information that the vehicle system cannot respond to correctly multiple times.

[0004] However, the aforementioned technologies require semantic researchers to manually analyze massive amounts of text information to extract repetitive text information that the vehicle system cannot respond to correctly multiple times, thus reducing the efficiency of obtaining repetitive text information that the vehicle system cannot respond to correctly multiple times. Summary of the Invention

[0005] To address the problems existing in the prior art, embodiments of the present invention provide an information collection method, apparatus, electronic device, and storage medium.

[0006] This invention provides an information collection method, comprising:

[0007] Convert the acquired current voice information into current text information;

[0008] If neither the current text information nor the previous text information is recognized, a similarity score between the current text information and the previous text information is determined; the previous text information is the text information corresponding to the previous voice information; the time interval between the current voice information and the previous voice information is less than or equal to a first preset value.

[0009] When the similarity score is determined to be greater than or equal to the second preset value, the current text information is sent to the vehicle system.

[0010] According to an information collection method provided by the present invention, the step of sending the current text information to the vehicle system when it is determined that the similarity score is greater than or equal to a second preset value includes:

[0011] When the similarity score is determined to be greater than or equal to the second preset value, the number of repetitions of similar texts is updated;

[0012] After a preset duration, when the number of repetitions is greater than or equal to a third preset value and less than or equal to a fourth preset value, it is determined whether the user's emotion has changed to a negative emotion based on the current text information and the previous text information.

[0013] When it is determined that the user's emotion has turned into a negative emotion, the current text information is sent to the vehicle system.

[0014] According to an information collection method provided by the present invention, determining whether a user's emotion has shifted to a negative emotion based on the current text information and the previous text information includes:

[0015] Obtain the sentiment score corresponding to the current text information and the sentiment score corresponding to the previous text information;

[0016] Based on the emotion score corresponding to the current text information and the emotion score corresponding to the previous text information, it is determined whether the user's emotion has changed to a negative emotion.

[0017] According to an information collection method provided by the present invention, determining whether a user's emotion has shifted to a negative emotion based on the emotion score corresponding to the current text information and the emotion score corresponding to the previous text information includes:

[0018] When the difference between the emotion score corresponding to the current text information and the emotion score corresponding to the previous text information is less than or equal to a fifth preset value, the user's emotion is determined to be converted into a negative emotion.

[0019] According to an information collection method provided by the present invention, the method further includes:

[0020] When the number of repetitions exceeds the fourth preset value, the current text information is sent to the vehicle system.

[0021] According to an information collection method provided by the present invention, before obtaining the sentiment score corresponding to the current text information, the method further includes:

[0022] The current text information and the previous text information are input into the sentiment analysis model to obtain the sentiment score corresponding to the current text information output by the sentiment analysis model.

[0023] The sentiment analysis model is trained based on multiple sample pairs and the sentiment labels corresponding to the two text samples in each sample pair; the sentiment labels include one of the following: positive sentiment label, no sentiment label, and negative sentiment label.

[0024] According to an information collection method provided by the present invention, determining the similarity score between the current text information and the previous text information includes:

[0025] Extract a first set of key corpora from the current text information, and extract a second set of key corpora from the previous text information;

[0026] Determine the similarity score between the first key corpus set and the second key corpus set.

[0027] The present invention also provides an information collection device, comprising:

[0028] The conversion unit is used to convert the acquired current speech information into current text information;

[0029] The determining unit is configured to determine a similarity score between the current text information and the previous text information when it is determined that there is no target text information that matches the current text information; the previous text information is the text information corresponding to the previous voice information; the time interval between the current voice information and the previous voice information is less than or equal to a first preset value;

[0030] The sending unit is used to send the current text information to the vehicle system when it is determined that the similarity score is greater than or equal to a second preset value.

[0031] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the information collection method as described above.

[0032] The present invention also provides an electronic device, including a microphone, a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the microphone is used to acquire current voice information;

[0033] The processor performs the conversion of the acquired current voice information into current text information;

[0034] If it is determined that there is no target text information that matches the current text information, the similarity score between the current text information and the previous text information is determined; the previous text information is the text information corresponding to the previous voice information; the time interval between the current voice information and the previous voice information is less than or equal to a first preset value;

[0035] When the similarity score is determined to be greater than or equal to the second preset value, the current text information is sent to the vehicle system.

[0036] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the information collection method as described above.

[0037] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the information collection method as described above.

[0038] The information collection method, apparatus, electronic device, and storage medium provided by this invention determine the similarity score between the current text information and the previous text information when neither the current text information nor the previous text information is recognized. When the similarity score is greater than or equal to a second preset value, it indicates that the current text information and the previous text information are similar text information. At this time, the current text information is sent to the vehicle system, realizing the automatic collection of similar text information that the vehicle system cannot respond to correctly, thereby improving the efficiency of obtaining similar text information that the vehicle system cannot respond to correctly. Attached Figure Description

[0039] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0040] Figure 1 This is one of the flowcharts illustrating the information collection method provided in this embodiment of the invention;

[0041] Figure 2 This is a second schematic flowchart of the information collection method provided in the embodiments of the present invention;

[0042] Figure 3 This is the third flowchart illustrating the information collection method provided in this embodiment of the invention;

[0043] Figure 4 This is a schematic diagram illustrating how the text filtering model provided in this embodiment of the invention filters text information.

[0044] Figure 5 This is a schematic diagram of the structure of the information collection device provided in an embodiment of the present invention;

[0045] Figure 6 This is one of the schematic diagrams of the physical structure of the electronic device provided by the present invention;

[0046] Figure 7 This is the second schematic diagram of the physical structure of the electronic device provided by the present invention. Detailed Implementation

[0047] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0048] To accurately capture unsupported content in user voice messages, research revealed that when electronic devices such as in-vehicle systems fail to provide the expected results, users will input similar voice messages to request again or express dissatisfaction. Therefore, this invention focuses on repetitive voice messages within a short period, especially those accompanied by a shift to negative emotions. Since users may describe similar voice messages, it is necessary to obtain the "key corpus" from the corresponding text information. An attention mechanism is used to extract auxiliary words (including prefixes and modal particles) from the text information. These auxiliary words are then removed, and the similarity score between the text information corresponding to consecutively inputted voice messages is calculated. If the similarity score is higher than a second preset value, the consecutively inputted voice messages are considered similar. Simultaneously, an emotion score for each text message can be determined, and based on the emotion score, it can be judged whether the user experienced an emotional shift when inputting similar voice messages.

[0049] The following is combined with Figures 1-4 The information collection method of the present invention is described.

[0050] Figure 1 This is one of the flowcharts illustrating the information collection method provided in this embodiment of the invention, applied to electronic devices such as in-vehicle systems. The following description uses an in-vehicle system as an example. Figure 1 As shown, this information collection method includes the following steps:

[0051] Step 101: Convert the acquired current voice information into current text information.

[0052] For example, after the user wakes up the vehicle's infotainment system, the system can acquire each voice message input by the user through the microphone and convert each voice message into corresponding text information; that is, when the current voice message input by the user is acquired, the current voice message can be converted into the corresponding current text information.

[0053] Step 102: If neither the current text information nor the previous text information is recognized by the vehicle system, determine the similarity score between the current text information and the previous text information; the previous text information is the text information corresponding to the previous voice information; the time interval between the current voice information and the previous voice information is less than or equal to a first preset value.

[0054] The first preset value can be set based on needs. For example, if the first preset value is 2 minutes, the interval between two voice messages is less than or equal to 2 minutes before the user actively exits the interactive interface, and is considered as continuous voice.

[0055] For example, when the vehicle system receives a voice message, it needs to match the corresponding text message with the corpus stored in the database. If the match fails, it means that the vehicle system cannot recognize the text message. If neither the current text message nor the previous text message is recognized by the vehicle system, it is necessary to calculate the similarity score between the current text message and the previous text message. The similarity score is used to automatically determine whether the current text message and the previous text message are similar text messages.

[0056] Step 103: When it is determined that the similarity score is greater than or equal to the second preset value, the current text information is sent to the vehicle system.

[0057] The second preset value can be set based on requirements; for example, the second preset value is 0.9.

[0058] For example, when obtaining a similarity score, the similarity score is compared with a second preset value. If the similarity score is greater than or equal to the second preset value, it means that the current text information is similar to the previous text information. Since both the current text information and the previous text information are information that the vehicle system cannot recognize, the current text information is sent to the vehicle system. If the similarity score is less than the second preset value, it means that the current text information is not similar to the previous text information, and the information collection process ends at this time.

[0059] It should be noted that the semantic parsing results corresponding to the current text information can also be sent to the vehicle system. This allows semantic researchers to make subsequent optimizations and improvements based on the current text information and the corresponding semantic parsing results, enabling the vehicle system to recognize the current text information and improve the accuracy of the vehicle system's semantic understanding. The semantic parsing results are the information obtained by parsing the current text information. The semantic parsing results can include fields related to the object of control, the control action, and the text meaning. For example, if the current text information is "open the car window," the field value of the object of control could be "vehicle control," the field value of the control action could be "open," and the field value of the text meaning could be "the car window needs to be opened."

[0060] It should be noted that the process of obtaining the semantic parsing results corresponding to the current text information can refer to existing technologies, and will not be elaborated here.

[0061] The information collection method provided by this invention determines the similarity score between the current text information and the previous text information when neither the current text information nor the previous text information is recognized by the vehicle system. When the similarity score is greater than or equal to a second preset value, it indicates that the current text information and the previous text information are similar text information. At this time, the current text information is sent to the vehicle system, realizing the automatic collection of similar text information that the vehicle system cannot respond to correctly, thereby improving the efficiency of obtaining similar text information that the vehicle system cannot respond to correctly.

[0062] In one embodiment, Figure 2 This is a second schematic flowchart of the information collection method provided in this embodiment of the invention, as shown below. Figure 2 As shown, step 103 above can be implemented through the following steps:

[0063] Step 1031: When it is determined that the similarity score is greater than or equal to the second preset value, update the number of repetitions of similar texts.

[0064] For example, when the similarity score between the current text information and the previous text information is determined to be greater than or equal to the second preset value, the repetition count of similar texts i = i + 1 is updated, where the initial value of i is 1; when the similarity score between the current text information and the previous text information is less than the second preset value, the count of i is reset.

[0065] Step 1032: After a preset duration, when the number of repetitions is greater than or equal to a third preset value and less than or equal to a fourth preset value, determine whether the user's emotion has changed to a negative emotion based on the current text information and the previous text information.

[0066] The third and fourth preset values ​​can be set according to requirements; for example, the third preset value is 2 and the fourth preset value is 3.

[0067] For example, after the preset duration for indicating continuous dialogue ends, the number of repetitions is compared with the third and fourth preset values. If it is determined that the number of repetitions is greater than or equal to the third preset value and less than or equal to the fourth preset value, it is necessary to determine whether the user's emotion has changed when inputting the current voice information and has changed to a negative emotion. That is, it is necessary to determine whether the user's emotion has changed to a negative emotion based on the current text information and the previous text information. For example, if the number of repetitions is 2, it is determined that the number of repetitive voice information input by the user within the preset duration is 2.

[0068] Step 1033: When it is determined that the user's emotion has changed to a negative emotion, the current text information is sent to the vehicle system.

[0069] For example, when it is determined that the user's emotion has turned into a negative emotion, it means that the user is dissatisfied with the unexpected feedback from the vehicle system to the previous voice information and has a change in emotion. After the change in emotion, the user inputs the current voice information, which is similar to the previous voice information, into the vehicle system. At this time, the vehicle system can actively output a prompt to soothe the user's emotions and send the current text information as the vehicle system's rejection data to the vehicle system.

[0070] The information collection method provided in this embodiment of the invention only sends the current text information to the vehicle system when the number of repetitions of similar text is greater than or equal to a third preset value, less than or equal to a fourth preset value, and when it is determined that the user's emotion has changed to a negative emotion. By using the user's emotional change to further determine whether the current text information is the vehicle system's rejected data, the accuracy of collecting text information that the vehicle system cannot recognize is improved.

[0071] In one embodiment, step 1032 described above can be implemented in the following way:

[0072] Obtain the emotion score corresponding to the current text information and the emotion score corresponding to the previous text information; based on the emotion score corresponding to the current text information and the emotion score corresponding to the previous text information, determine whether the user's emotion has changed to a negative emotion.

[0073] For example, the emotion score can be represented numerically. A emotion score of 1 indicates that the user's emotion is positive, meaning the current text contains happy-related language. A emotion score of 0 indicates that the user has no emotion, meaning the current text contains neutral language such as statements of fact or commands. A emotion score of -2 indicates that the user's emotion is negative, meaning the current text contains language such as anger or anxiety. Therefore, the user's emotion can be determined to have shifted to a negative emotion based on the emotion score of the current text and the emotion score of the previous text.

[0074] The information collection method provided in this embodiment of the invention can determine whether a user's emotion has turned into a negative emotion based on the emotion score corresponding to the current text information and the emotion score corresponding to the previous text information. Since the emotion score is directly derived from the analysis of the text information, it can improve the accuracy of the user's emotion judgment.

[0075] In one embodiment, based on the sentiment score corresponding to the current text information and the sentiment score corresponding to the previous text information, it is determined whether the user's sentiment has shifted to a negative sentiment. This can be achieved in the following way:

[0076] When the difference between the emotion score corresponding to the current text information and the emotion score corresponding to the previous text information is less than or equal to a fifth preset value, the user's emotion is determined to be converted into a negative emotion.

[0077] The fifth preset value can be set based on the specific value of the emotion score. For example, when an emotion score of 1 represents positive emotion, an emotion score of 0 represents no emotion, and an emotion score of -2 represents negative emotion, the fifth preset value can be -2.

[0078] For example, when the emotion score corresponding to the current text information and the emotion score corresponding to the previous text information are obtained, the difference between the emotion score corresponding to the current text information and the emotion score corresponding to the previous text information is calculated. If the difference is less than or equal to a fifth preset value, it is determined that the user's emotion has changed to a negative emotion. For example, if the emotion score corresponding to the current text information is -2 and the emotion score corresponding to the previous text information is 0, then the difference -2 is equal to the fifth preset value -2, and it can be determined that the user has changed from no emotion to a negative emotion. For example, if the emotion score corresponding to the current text information is -2 and the emotion score corresponding to the previous text information is 1, then the difference -3 is less than the fifth preset value -2, and it can be determined that the user has changed from a positive emotion to a negative emotion.

[0079] In one embodiment, step 1032 described above can be implemented in the following way:

[0080] Obtain the amplitude information corresponding to the current voice message and the amplitude information corresponding to the previous voice message; based on the amplitude information corresponding to the current voice message and the amplitude information corresponding to the previous voice message, determine whether the user's emotion has changed to a negative emotion.

[0081] For example, when the amplitude information corresponding to the current voice information and the amplitude information corresponding to the previous voice information are obtained, the amplitude information corresponding to the current voice information is compared with the amplitude information corresponding to the previous voice information. If it is determined that the amplitude information corresponding to the current voice information is greater than the amplitude information corresponding to the previous voice information, it means that the volume of the user when inputting the previous voice information is less than the volume when inputting the current voice information. The user's current emotion can be judged from the volume. That is, if it is determined that the amplitude information corresponding to the current voice information is greater than the amplitude information corresponding to the previous voice information, it can be determined that the user's emotion has changed to a negative emotion.

[0082] In one embodiment, Figure 3 This is the third flowchart illustrating the information collection method provided in this embodiment of the invention, as shown below. Figure 3 As shown, step 103 above also includes the following steps:

[0083] Step 1034: When the number of repetitions exceeds the fourth preset value, the current text information is sent to the vehicle system.

[0084] For example, if the number of repetitions exceeds the fourth preset value, it is considered that the user has initiated repeated voice messages multiple times in a short period of time. In this case, it can be determined that the repeated voice message is the voice message that the vehicle system cannot provide the user with a correct response. Without the need for the emotion score corresponding to the current text message, the vehicle system will actively send the current text message corresponding to the current voice message to the vehicle system.

[0085] The information collection method provided in this embodiment of the invention can determine the repeated voice information that the vehicle system cannot respond to correctly when the number of repetitions of similar text is greater than a fourth preset value, without judging the user's emotions. The method requires less computation and thus improves the efficiency of collecting text information that the vehicle system cannot recognize.

[0086] In one embodiment, before obtaining the sentiment score corresponding to the current text information, the information acquisition method further includes the following steps:

[0087] The current text information and the previous text information are input into the sentiment analysis model to obtain the sentiment score corresponding to the current text information output by the sentiment analysis model.

[0088] The sentiment analysis model is trained based on multiple sample pairs and the sentiment labels corresponding to the two text samples in each sample pair. The sentiment labels include one of the following: positive sentiment label, no sentiment label, and negative sentiment label. A positive sentiment label can be 1, a no sentiment label can be 0, and a negative sentiment label can be -2. The sentiment analysis model can be a text-based dialogue sentiment recognition model (hierarchical Gated Recurrent Unit, HiGRU).

[0089] For example, the HiGRU model is modeled using a two-level bidirectional gated recurrent unit (GRU) structure. The lower-level GRU is at the sentence level, obtaining the embedding of a single text based on the text sequence of the corpus. The higher-level GRU is at the dialogue level, obtaining the embedding of text in the context based on the text sequence of the dialogue. The input of the higher-level GRU is the output of the lower-level GRU.

[0090] First, define a set of dialogues. Where L represents the total number of dialogues. This indicates that in each dialogue D i N in i A sequence of texts, u j Made with specific emotions c j Speakers s of ∈C jLet S represent the set of speakers, C represent the set of all emotions, and M represent the set of all emotions. j Indicate u j The number of words, and specific emotions can include three categories of labels: positive emotions, no emotions, and negative emotions. As shown in formulas (1) and (2) below, the corresponding single-word embedding sequences... The word embeddings are fed into a low-level bidirectional GRU to learn individual word embeddings in opposite directions:

[0091]

[0092]

[0093] Two hidden states and Linked to And based on the following formula (3), the tanh activation function on a linear transformation is used to generate w. k Context-dependent word embeddings:

[0094] e c (w k ) = tanh(W w ·h s +b w (3)

[0095] Based on the following formula (4), the context-related word embeddings are subjected to max pooling to obtain the embeddings of individual words:

[0096]

[0097] As shown in formulas (5) and (6) below, for the i-th dialogue The learned embedding sequence of a single utterance It is sent to a higher-level bidirectional GRU to obtain the sequence relations and context relations of the text corpus:

[0098]

[0099]

[0100] To distinguish the hidden state h of low-level GRUs k The hidden state of a high-level GRU is represented as Therefore, the embedding of context-related utterances is obtained through the following formula (7):

[0101] e c (u j ) = tanh(W u ·H s +bu (7)

[0102] in, and d2 represents the dimension of the hidden state in a high-level GRU (dialogue sequence modeling). W represents a d²-dimensional real vector. u and b u These are all parameters of the model to be trained. W u It is a real matrix with d2 rows and 2d2 columns. b u It is a d²-dimensional real vector because user sentiment is judged at the utterance level, and the learned context-dependent utterance embedding e c (u j After being directly fed into the fully connected layer, the sentiment score is determined by the softmax function shown in the following formula (8).

[0103]

[0104] in, W represents the sentiment score. fc and b fc These are all parameters of the model to be trained. W fc It is a real matrix with |C| rows and d2 columns, b fc ∈R |C| b fc It is a real vector of |C| dimensions, where |C| represents the vector length of all sentiment sets.

[0105] It should be noted that the aforementioned dialogue in this invention refers to the previous text information and the current text information.

[0106] In one embodiment, the similarity score between the current text information and the previous text information is calculated as follows:

[0107] Extract the first key corpus set from the current text information and the second key corpus set from the previous text information, and determine the similarity score between the first key corpus set and the second key corpus set.

[0108] In order to better determine the similarity between the current text information and the previous text information, we need to pay more attention to the more critical text content in the text information. Therefore, we need to extract the first key corpus from the current text information and the second key corpus from the previous text information, and calculate the similarity score between the first key corpus and the second key corpus to improve the accuracy of the calculated similarity score.

[0109] Exemplarily, a text filtering model can be used to learn auxiliary words, and the auxiliary words in the text information are filtered by the text filtering model. That is, the current text information is input into the text filtering model, and the auxiliary words in the current text information are filtered by the text filtering model to obtain a first key corpus set output by the text filtering model; the previous text information is input into the text filtering model, and the auxiliary words in the previous text information are filtered by the text filtering model to obtain a second key corpus set output by the text filtering model.

[0110] Among them, the auxiliary words include modal particles, prefix words, etc. For example, common modal particles are "de, gei, ma, me, lai, le, ba, ye, ya, jiushi, ne, a, ha", etc., and prefix words are "qing, qingnin, qingni, mafan, mafanni, mafannin, nengbuneng, kebukeyi, nengbunengni, kebukeyini, banggewang, bangwo, geiwo, tiwo, weiwo, woshiyao, wodiansuan, wo xiang, wo yao, wo xiangyao", etc. Filtering these auxiliary words and / or prefix words from the text information can more clearly understand the core demands of the user. Figure 4 It is a schematic diagram of the text filtering model provided by the embodiments of the present invention for filtering text information, as Figure 4 shown. Taking the text information "Hurry up and open the window for me, ah" as an example, the text information "Hurry up and open the window for me, ah" is input into the text filtering model, and the result output by the text filtering model is the result corresponding to each character. Using 1 to represent the key corpus in the text information and 0 to represent the auxiliary word corpus in the text information, from Figure 4 it can be seen that the key corpus set in the text information "Hurry up and open the window for me, ah" is "open the window".

[0111] Specifically, the specific process for the text filtering model to extract the first key corpus set from the current text information is as follows: converting each text corpus in the current text information into a corresponding first vector, and converting each preset auxiliary word corpus into a corresponding second vector; determining the corresponding attention scores based on the first vector and the second vector; determining the attention distribution based on all the attention scores; determining the attention value of the current text information based on each attention score in the attention distribution and the corresponding second vector; and finally removing the non-key corpus in the current text information based on the attention value to obtain the first key corpus set; the specific process for extracting the second key corpus set from the previous text information is similar to the above specific process for extracting the first key corpus set from the current text information, and the present invention will not elaborate here.

[0112] Among them, the attention distribution is used to indicate the correlation degree between the first vector and the second vector.

[0113] Exemplarily, first, the n text corpora y1, y2,..., y in the current text informationn-1 y n After one-hot encoding, it is converted into the input text vectors e1, e2, ..., e n-1 e n To better perform subsequent attention mechanism calculations, the input text vectors are transformed using a linear transformation H = g(We), where g is the LeakyReLU activation function and W is a linear layer. This ultimately yields linear vectors H1, H2, ..., H1 corresponding to each input text vector. n-1 H n Then determine the vector H corresponding to the current text information. i , i∈(1,2,...,n-1,n). Similarly, the auxiliary word corpus set X=x1,x2,…,x n-1 x m After one-hot encoding and linear transformation H = g(We), the vectors h1, h2, ..., h corresponding to each auxiliary word corpus are finally obtained. n-1 h m Then determine the vector h corresponding to the auxiliary word corpus set. j , j∈(1,2,...,n-1,m), let vector H i sum vector h j The attention score e is obtained by performing a dot product operation using formula (9). ij :

[0114] e ij =score(H i h j )=H i *h j (9)

[0115] Here, * represents the dot product operation.

[0116] Then, the attention distribution α is obtained using the softmax function shown in formula (10). ij α ij Used to indicate the degree of correlation between the first vector and the second vector.

[0117]

[0118] Wherein, exp(e ik ) represents e x The exponential function, e represents Napier's constant 2.7182, i represents the i-th text corpus in the current text information, j represents the j-th auxiliary word prediction, L represents the total number of auxiliary word predictions, e ik =score(H i h k )=H i *hk , which represents the dot product operation between the vectors of the i-th input text corpus and the k-th auxiliary word corpus.

[0119] Finally, based on the following formula (11), the attention distribution α is... ij Each attention score is multiplied by its corresponding h, and then the weighted vectors are summed to obtain the attention value:

[0120]

[0121] The attention value c for each text corpus is obtained by fusing the vector information of the text corpus with the feature information of auxiliary words. i At that time, c i Two linear layers are fed into the text. The outputs of these two linear layers are then fed into a LeakyReLU activation function. The output of the LeakyReLU activation function is then used as the input to a softmax activation function. Finally, the output of the softmax activation function is used as the confidence score T for the corresponding text corpus. The confidence score T ranges from 0 to 1, indicating whether the auxiliary word appears in the text corpus. 1-T represents whether the text corpus is a key corpus. Therefore, the value of 1-T is compared with a threshold. If the value of 1-T is less than the threshold, the text corpus is determined to be key corpus; if the value of 1-T is greater than or equal to the threshold, the text corpus is determined to be non-key corpus and removed from the current text information. This yields the first set of key corpus information excluding non-key corpus information. The same method is used to obtain the second set of key corpus information excluding non-key corpus information from the previous text information.

[0122] When the first and second key corpora are obtained, the cosine similarity method is used to calculate the similarity score between the first and second key corpora. Specifically, the first and second key corpora are segmented using a Long-term Memory (LSTM) model, and all words in the first and second key corpora are listed. Each word is vectorized, and then the cosine value of the angle between the vectors corresponding to the first and second key corpora is calculated using the cosine similarity formula (12).

[0123]

[0124] Where, x i y represents the vector corresponding to the first key corpus set. iLet θ represent the vector corresponding to the second key corpus set, and n represent the total number of all key corpora in the first and second key corpus sets. This can also be understood as the length of the vector corresponding to either the first or second key corpus set. The lengths of the vectors corresponding to the first and second key corpus sets are the same. θ represents the angle between the vectors corresponding to the first and second key corpus sets. The smaller the angle, the more similar the first and second key corpus sets are. Therefore, the closer the cosine of the angle is to 1, the more similar the first and second key corpus sets are. In this invention, the cosine of the angle, cos(θ), is determined as the similarity score.

[0125] The information collection device provided by the present invention is described below. The information collection device described below and the information collection method described above can be referred to in correspondence.

[0126] Figure 5 This is a schematic diagram of the information collection device provided in an embodiment of the present invention, such as... Figure 5 As shown, the information collection device 500 includes: a conversion unit 501, a determination unit 502, and a transmission unit 503; wherein:

[0127] The conversion unit 501 is used to convert the acquired current speech information into current text information;

[0128] The determining unit 502 is used to determine the similarity score between the current text information and the previous text information when it is determined that there is no target text information that matches the current text information; the previous text information is the text information corresponding to the previous voice information; the time interval between the current voice information and the previous voice information is less than or equal to a first preset value;

[0129] The sending unit 503 is used to send the current text information to the vehicle system when it is determined that the similarity score is greater than or equal to a second preset value.

[0130] Based on any of the above embodiments, the sending unit 503 is specifically used for:

[0131] When the similarity score is determined to be greater than or equal to the second preset value, the number of repetitions of similar texts is updated;

[0132] After a preset duration, when the number of repetitions is greater than or equal to a third preset value and less than or equal to a fourth preset value, it is determined whether the user's emotion has changed to a negative emotion based on the current text information and the previous text information.

[0133] When it is determined that the user's emotion has turned into a negative emotion, the current text information is sent to the vehicle system.

[0134] Based on any of the above embodiments, the sending unit 503 is further specifically used for:

[0135] Obtain the sentiment score corresponding to the current text information and the sentiment score corresponding to the previous text information;

[0136] Based on the emotion score corresponding to the current text information and the emotion score corresponding to the previous text information, it is determined whether the user's emotion has changed to a negative emotion.

[0137] Based on any of the above embodiments, the sending unit 503 is further specifically used for:

[0138] When the difference between the emotion score corresponding to the current text information and the emotion score corresponding to the previous text information is less than or equal to a fifth preset value, the user's emotion is determined to be converted into a negative emotion.

[0139] Based on any of the above embodiments, the sending unit 503 is further specifically used for:

[0140] When the number of repetitions exceeds the fourth preset value, the current text information is sent to the vehicle system.

[0141] Based on any of the above embodiments, the information collection device 500 further includes:

[0142] The analysis unit is used to input the current text information and the previous text information into the sentiment analysis model to obtain the sentiment score corresponding to the current text information output by the sentiment analysis model.

[0143] The sentiment analysis model is trained based on multiple sample pairs and the sentiment labels corresponding to the two text samples in each sample pair; the sentiment labels include one of the following: positive sentiment label, no sentiment label, and negative sentiment label.

[0144] Based on any of the above embodiments, the determining unit 502 is specifically used for:

[0145] Extract a first set of key corpora from the current text information, and extract a second set of key corpora from the previous text information;

[0146] Determine the similarity score between the first key corpus set and the second key corpus set.

[0147] The information collection device provided by this invention determines the similarity score between the current text information and the previous text information when neither the current text information nor the previous text information is recognized. When the similarity score is greater than or equal to a second preset value, it indicates that the current text information and the previous text information are similar text information. At this time, the current text information is sent to the vehicle system, realizing the automatic collection of similar text information that the vehicle system cannot respond to correctly, thereby improving the efficiency of obtaining similar text information that the vehicle system cannot respond to correctly.

[0148] Figure 6 This is one of the schematic diagrams of the physical structure of the electronic device provided in the embodiments of the present invention, such as... Figure 6 As shown, the electronic device may include a processor 610, a communications interface 620, a memory 630, and a communication bus 640, wherein the processor 610, the communications interface 620, and the memory 630 communicate with each other via the communication bus 640. The processor 610 can call logical instructions in the memory 630 to execute an information collection method, which includes converting the acquired current voice information into current text information;

[0149] If neither the current text information nor the previous text information is recognized, a similarity score between the current text information and the previous text information is determined; the previous text information is the text information corresponding to the previous voice information; the time interval between the current voice information and the previous voice information is less than or equal to a first preset value.

[0150] When the similarity score is determined to be greater than or equal to the second preset value, the current text information is sent to the vehicle system.

[0151] Furthermore, the logical instructions in the aforementioned memory 630 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0152] Figure 7 This is a second schematic diagram of the physical structure of the electronic device provided in the embodiments of the present invention, as shown below. Figure 7 As shown, the electronic device may include a processor 710, a communications interface 720, a memory 730, and a communication bus 740, and also includes a microphone 750. The processor 710, communications interface 720, memory 730, and microphone 750 communicate with each other via the communication bus 740. The microphone 750 is used to acquire current voice information; the processor 710 can call logical instructions in the memory 730 to execute the conversion of the acquired current voice information into current text information.

[0153] If it is determined that there is no target text information that matches the current text information, the similarity score between the current text information and the previous text information is determined; the previous text information is the text information corresponding to the previous voice information; the time interval between the current voice information and the previous voice information is less than or equal to a first preset value;

[0154] When the similarity score is determined to be greater than or equal to the second preset value, the current text information is sent to the vehicle system.

[0155] Furthermore, the logical instructions in the aforementioned memory 730 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0156] On the other hand, the present invention also provides a computer program product, the computer program product including a computer program, the computer program being able to be stored on a non-transitory computer-readable storage medium, and when the computer program is executed by a processor, the computer is able to execute the information collection method provided by the above methods, the method including: converting the acquired current voice information into current text information;

[0157] If neither the current text information nor the previous text information is recognized, a similarity score between the current text information and the previous text information is determined; the previous text information is the text information corresponding to the previous voice information; the time interval between the current voice information and the previous voice information is less than or equal to a first preset value.

[0158] When the similarity score is determined to be greater than or equal to the second preset value, the current text information is sent to the vehicle system.

[0159] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the information collection methods provided by the above methods, the method comprising: converting acquired current voice information into current text information;

[0160] If neither the current text information nor the previous text information is recognized, a similarity score between the current text information and the previous text information is determined; the previous text information is the text information corresponding to the previous voice information; the time interval between the current voice information and the previous voice information is less than or equal to a first preset value.

[0161] When the similarity score is determined to be greater than or equal to the second preset value, the current text information is sent to the vehicle system.

[0162] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0163] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0164] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An information collection method, characterized in that, include: Convert the acquired current voice information into current text information; If neither the current text information nor the previous text information is recognized, a similarity score between the current text information and the previous text information is determined. The previous text information is the text information corresponding to the previous voice information; the time interval between the current voice information and the previous voice information is less than or equal to a first preset value; the fact that neither the current text information nor the previous text information is recognized means that the current text information fails to match the corpus stored in the database, and the previous text information also fails to match the corpus stored in the database. When the similarity score is determined to be greater than or equal to the second preset value, the current text information is sent to the vehicle system.

2. The information collection method according to claim 1, characterized in that, The step of sending the current text information to the vehicle system when the similarity score is determined to be greater than or equal to a second preset value includes: When the similarity score is determined to be greater than or equal to the second preset value, the number of repetitions of similar texts is updated; After a preset duration, when the number of repetitions is greater than or equal to a third preset value and less than or equal to a fourth preset value, it is determined whether the user's emotion has changed to a negative emotion based on the current text information and the previous text information. When it is determined that the user's emotion has turned into a negative emotion, the current text information is sent to the vehicle system.

3. The information collection method according to claim 2, characterized in that, The step of determining whether the user's emotion has shifted to a negative emotion based on the current text information and the previous text information includes: Obtain the sentiment score corresponding to the current text information and the sentiment score corresponding to the previous text information; Based on the emotion score corresponding to the current text information and the emotion score corresponding to the previous text information, it is determined whether the user's emotion has changed to a negative emotion.

4. The information collection method according to claim 3, characterized in that, The step of determining whether a user's emotion has shifted to a negative emotion based on the emotion score corresponding to the current text information and the emotion score corresponding to the previous text information includes: When the difference between the emotion score corresponding to the current text information and the emotion score corresponding to the previous text information is less than or equal to a fifth preset value, the user's emotion is determined to be converted into a negative emotion.

5. The information collection method according to claim 2, characterized in that, The method further includes: When the number of repetitions exceeds the fourth preset value, the current text information is sent to the vehicle system.

6. The information collection method according to claim 3, characterized in that, Before obtaining the sentiment score corresponding to the current text information, the method further includes: The current text information and the previous text information are input into the sentiment analysis model to obtain the sentiment score corresponding to the current text information output by the sentiment analysis model. The sentiment analysis model is trained based on multiple sample pairs and the sentiment labels corresponding to the two text samples in each sample pair; the sentiment labels include one of the following: positive sentiment label, no sentiment label, and negative sentiment label.

7. The information collection method according to any one of claims 1-6, characterized in that, Determining the similarity score between the current text information and the previous text information includes: Extract a first set of key corpora from the current text information, and extract a second set of key corpora from the previous text information; Determine the similarity score between the first key corpus set and the second key corpus set.

8. An information collection device, characterized in that, include: The conversion unit is used to convert the acquired current speech information into current text information; The determining unit is configured to determine the similarity score between the current text information and the previous text information when neither the current text information nor the previous text information is identified. The previous text information is the text information corresponding to the previous voice information; the time interval between the current voice information and the previous voice information is less than or equal to a first preset value; the fact that neither the current text information nor the previous text information is recognized means that the current text information fails to match the corpus stored in the database, and the previous text information also fails to match the corpus stored in the database. The sending unit is used to send the current text information to the vehicle system when it is determined that the similarity score is greater than or equal to a second preset value.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the information collection method as described in any one of claims 1 to 7.

10. An electronic device, comprising a microphone, and further comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The microphone is used to collect current voice information; The processor performs the conversion of the acquired current voice information into current text information; If neither the current text information nor the previous text information is recognized, a similarity score between the current text information and the previous text information is determined. The previous text information is the text information corresponding to the previous voice information; the time interval between the current voice information and the previous voice information is less than or equal to a first preset value; the fact that neither the current text information nor the previous text information is recognized means that the current text information fails to match the corpus stored in the database, and the previous text information also fails to match the corpus stored in the database. When the similarity score is determined to be greater than or equal to the second preset value, the current text information is sent to the vehicle system.

11. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the information collection method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Information processing method and electronic equipment

    CN105810188A

  • Voice instruction recommendation method and device and electronic equipment

    CN113987130A