Method, device, equipment and storage medium for processing vehicle-machine voice interaction data
By analyzing the interaction response status of the vehicle-machine voice interaction data, and automatically identifying and counting the failed voice of interaction, the problems of high manual research costs and incorrect optimization direction are solved, and development accuracy and user experience are improved.
Patent Information
- Application Number
- CN202211330554.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-26
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2042-10-26
AI Technical Summary
In the prior art, manual research costs are high and the direction of optimization development is incorrect, resulting in poor user experience.
By obtaining the original voice interaction data, analyzing the interaction response status, identifying and counting the interaction failed voice and the number of triggered devices, automated failed voice statistics are realized and manual research is avoided.
It realizes interactive failed voice statistics without manual research, improves the accuracy of development and optimization, and improves the user experience.
Smart Images

Figure CN115662400B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of voice processing technology, and in particular to a method, device, equipment and storage medium for processing vehicle-computer voice interaction data. Background Art
[0002] With the development of automotive electronics, more and more automakers are choosing to equip their vehicles with intelligent voice interaction systems. These systems allow users to interact with their smart cars through voice commands. For example, users can use voice commands to access vehicle controls, map navigation, music, multimedia services, check the weather, or chat.
[0003] When developing intelligent voice interaction systems, technical personnel typically optimize or develop features based on market research. However, manual research consumes significant manpower and time, and the accuracy of research results is low, which can lead to incorrect optimization and development, impacting user experience. Summary of the Invention
[0004] In view of this, the purpose of this application is to propose a method, device, equipment and storage medium for processing vehicle-computer voice interaction data to solve the problems of high manual research costs and wrong optimization and development directions in the existing technology.
[0005] Based on the above objectives, this application provides a method for processing vehicle-to-vehicle voice interaction data, including:
[0006] Acquire original voice interaction data, wherein the original voice interaction data includes: at least one historical voice and an interaction response status corresponding to each historical voice, the interaction response status being used to indicate whether the interaction with the historical voice is successful;
[0007] Determining at least one interaction failed voice from at least one of the historical voices according to the interaction response state corresponding to each of the historical voices in the original voice interaction data;
[0008] For each type of interaction failure voice, the number of voices corresponding to the interaction failure voice and / or the number of triggering devices corresponding to the interaction failure voice are determined to obtain a failure number statistical result.
[0009] Based on the above objectives, the present application also provides a device for processing vehicle-machine voice interaction data, comprising:
[0010] A data acquisition module is used to acquire original voice interaction data, wherein the original voice interaction data includes: at least one historical voice and an interaction response status corresponding to each historical voice, the interaction response status being used to indicate whether the interaction with the historical voice is successful;
[0011] a failed voice determination module, configured to determine at least one failed interaction voice from at least one of the historical voices according to an interaction response state corresponding to each of the historical voices in the original voice interaction data;
[0012] The failure statistics module is used to determine the number of voices corresponding to each type of interaction failure voice and / or the number of triggering devices corresponding to the interaction failure voice, and obtain a failure number statistics result.
[0013] Based on the above purpose, the present application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the program, it implements a method for processing vehicle-machine voice interaction data as provided in any embodiment of the present application.
[0014] Based on the above purpose, the present application also provides a non-transitory computer-readable storage medium, which stores computer instructions, and the computer instructions are used to enable a computer to execute the method for processing vehicle-machine voice interaction data provided in any embodiment of the present application.
[0015] From the above, it can be seen that the method for processing vehicle-computer voice interaction data provided by the present application obtains the original voice interaction data, and determines each failed interaction voice from each historical voice according to the interaction response status corresponding to each historical voice in the original voice interaction data, and then determines the number of voices corresponding to each failed interaction voice and at least one of the number of trigger devices corresponding to each failed interaction voice, thereby realizing the statistics of vehicle-computer voice with failed interactions. The statistical results can be used for function development or function optimization without manual research, solving the problem of manual research consuming a lot of costs. Moreover, by statistics on historical voices with failed interactions, the actual usage of users can be reflected. The development direction or optimization direction determined by the statistical results is highly accurate, avoiding the situation where the optimization development direction is wrong. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in this application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are merely embodiments of this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0017] Figure 1 A flowchart of a method for processing vehicle-machine voice interaction data provided in an embodiment of the present application;
[0018] Figure 2A flowchart of another method for processing vehicle-machine voice interaction data provided in an embodiment of the present application;
[0019] Figure 3 A schematic diagram of a process for processing vehicle-machine voice interaction data provided in an embodiment of the present application;
[0020] Figure 4 A schematic diagram of the structure of a device for processing vehicle-machine voice interaction data provided in an embodiment of the present application;
[0021] Figure 5 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0022] In order to make the objectives, technical solutions and advantages of this application more clear, this application is further described in detail below in combination with specific embodiments and with reference to the accompanying drawings.
[0023] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present application should have the usual meanings understood by people with ordinary skills in the field to which this application belongs. The "first", "second" and similar words used in the embodiments of the present application do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0024] Figure 1 This is a flowchart of a method for processing vehicle-machine voice interaction data provided by an embodiment of the present application. The method can be executed by a vehicle-machine voice interaction data processing device, which can be implemented in software and / or hardware, and can be configured in an electronic device. Figure 1 As shown, the method may specifically include the following steps:
[0025] S110. Acquire original voice interaction data, wherein the original voice interaction data includes: at least one historical voice and an interaction response state corresponding to each historical voice.
[0026] The original voice interaction data may include historical voices initiated by each vehicle user and the interactive response status corresponding to each historical voice.
[0027] For example, the raw voice interaction data can be obtained through a pre-established voice access layer. In a specific embodiment, the voice server can collect the user's raw voice interaction data and store it in an open source stream processing platform (such as Kafka); further, a distributed file system (such as Hadoop), a data warehouse tool (such as HIVE), or a scheduler can read the raw voice interaction data from Kafka to further analyze the raw voice interaction data through a relational database service (RDS) or a relational database management system (such as MySQL), thereby extracting each interaction failure voice from the raw voice interaction data.
[0028] It should be noted that, considering that the voice server stores the original voice interaction data of the test environment and the original voice interaction data of the formal environment, the original voice interaction data of the test environment can be the original voice interaction data generated on the test device, and the original voice interaction data of the formal environment can be the original voice interaction data generated on the actual vehicle computer. In order to avoid the influence of the original voice interaction data of the test environment on the statistical results of the number of failures, all the original voice interaction data can be filtered to extract the original voice interaction data of the formal environment. For example, after reading all the original voice interaction data from Kafka, the original voice interaction data of the test environment can be eliminated through the device identifier of the test device. In this way, the influence of the original voice interaction data generated on the test device on the statistical results can be avoided, so that the statistical results are more in line with the actual usage of users, and the accuracy of the development direction or optimization direction is further improved.
[0029] Specifically, each historical voice may include a successful interaction voice and a failed interaction voice; wherein a successful interaction voice may be a historical voice that provides a response service corresponding to the voice to the user, and a failed interaction voice may be a historical voice that does not provide a response service corresponding to the voice to the user. The interaction response status is used to indicate whether the interaction of the historical voice is successful; the interaction response status may include a recognition operation status and a response service status corresponding to the recognition operation status. The recognition operation status is used to indicate whether the recognition operation corresponding to the historical voice is successful, that is, feedback on whether the recognition operation corresponding to the historical voice is successfully recognized; the response service status is used to indicate whether the response service corresponding to the recognition operation is successful, that is, feedback on whether the response service corresponding to the historical voice is successfully recognized.
[0030] In this embodiment, the process by which the vehicle computer provides voice interaction services to users typically involves the following: the user initiates a voice call, the vehicle computer determines the corresponding recognition operation, and based on the recognition operation, further determines the corresponding response service, executing corresponding control or providing corresponding information based on the determined response service. The recognition operation can be the action intended by the voice recognition, such as planning a route, making a call, or opening the news; the response service can be the vehicle computer's skills triggered by the recognition operation, such as weather, navigation, and vehicle control and equipment.
[0031] For example, if a user asks, "Is it cold today?", the corresponding recognition operation is "Check the weather," and the corresponding response service is "Weather." If a user asks, "Please lower the window," the corresponding recognition operation is "Body control," and the corresponding response service is "Vehicle control and vehicle settings."
[0032] S120: Determine at least one interaction-failed voice from at least one historical voice according to the interaction response state corresponding to each historical voice in the original voice interaction data.
[0033] Specifically, if the recognition operation and the corresponding response service corresponding to the historical voice can be determined, the historical voice can be determined as a successful interaction voice; if the recognition operation or the corresponding response service corresponding to the historical voice cannot be determined, the historical voice can be determined as a failed interaction voice.
[0034] Furthermore, for each historical voice in the original voice interaction data, it can be determined whether a recognition operation is determined according to the interaction response status, or whether a response service is determined according to the interaction response status, and then it can be determined whether the historical voice is an interaction failure voice.
[0035] In a specific embodiment, the interactive response status includes a recognition operation status and a response service status corresponding to the recognition operation status, wherein the recognition operation status is used to indicate whether the recognition operation corresponding to the historical voice is successfully recognized, and the response service status is used to indicate whether the response service corresponding to the recognition operation is successfully recognized. According to the interactive response status corresponding to each historical voice in the original voice interaction data, at least one interaction failure voice is determined from at least one historical voice, including: for each historical voice, if the recognition operation status corresponding to the historical voice is recognized as failure, the historical voice is determined to be an interaction failure voice; and / or, for each historical voice, if the recognition operation status corresponding to the historical voice is recognized as success, and the response service status corresponding to the recognition operation status is failure, the historical voice is determined to be an interaction failure voice.
[0036] That is, a historical voice in which a recognition operation cannot be recognized is determined as an interaction failure voice; and a historical voice in which a recognition operation can be recognized but a response service cannot be recognized is determined as an interaction failure voice.
[0037] Exemplarily, the original voice interaction data includes multiple "topic": "dm.output", wherein each "topic": "dm.output" represents a historical voice. Specifically, each "dm.output" contains "input": "xxxxx", wherein "xxxxx" is the specific content of the historical voice; each "dm.output" also contains "intentName": "xxxxx", wherein "xxxxx" is the corresponding recognition operation; each "dm.output" also contains a skill field, and the content in the skill field is the corresponding response service. Therefore, the recognition operation status and the response service status can be determined based on the content in the operation field and the content in the service field corresponding to each historical voice, and then it can be determined whether the recognition operation corresponding to the historical voice is successfully recognized, and whether the corresponding response service is successfully recognized, and then it can be determined whether the historical voice is an interaction failure voice.
[0038] In addition to judging whether the recognition operation status is successful and judging whether the response service status is successful by the content in the field as mentioned above, the recognition operation status or the response service status can also be judged by the automatic reply content corresponding to the historical voice.
[0039] That is, optionally, if the recognition operation status corresponding to the historical voice is recognized as failure, the historical voice is determined to be an interaction failure voice, including: judging whether the historical reply corresponding to the historical voice is the first preset reply, and if so, determining that the recognition operation status corresponding to the historical voice is failure, and determining the historical voice as an interaction failure voice; correspondingly, if the recognition operation status corresponding to the historical voice is recognized as success, and the response service status corresponding to the recognition operation status is failure, including: judging whether the historical reply corresponding to the historical voice is the second preset reply, and if so, determining that the recognition operation status corresponding to the historical voice is success, and the response service status corresponding to the recognition operation status is failure, and determining the historical voice as an interaction failure voice.
[0040] The first preset response may be an automatic response preset for a voice that cannot be recognized for a corresponding recognition operation, that is, an automatic response preset for a voice that is understood as empty. For example, the first preset response may be: "I don't understand what you said, can you please rephrase it?", "I really don't understand what you said?", or "I haven't learned it yet."
[0041] The second preset response can be a pre-set automatic response for voice messages that fail to recognize a response service. Specifically, if the vehicle computer does not have a response service corresponding to the voice recognition operation, or if the vehicle computer cannot determine the response service corresponding to the recognition operation, then the recognition of the response service corresponding to the voice message fails, and the vehicle computer returns the second preset response. Exemplary second preset responses may include: "Sorry," "Excuse me," or "Error."
[0042] Specifically, the historical responses corresponding to each historical voice can be obtained. For each historical voice, if the historical response corresponding to the historical voice is the first preset response or the second preset response, it can be determined that the historical voice is an interaction failure voice. Among them, the historical voice with the first preset response indicates that the recognition operation status corresponding to the historical voice is failed, that is, the corresponding recognition operation cannot be recognized; the historical voice with the second preset response indicates that the recognition operation status corresponding to the historical voice is successful and the response service status is failed, that is, the recognition operation is successfully recognized but the corresponding response service cannot be recognized.
[0043] By obtaining the historical replies corresponding to each historical voice and judging whether it is an interaction failure voice based on the historical replies, the rapid acquisition of interaction failure voice is achieved. There is no need to perform recognition operations and response service recognition on each historical voice one by one, which improves the efficiency of acquiring interaction failure voice and thus improves the efficiency of statistical analysis of failed voice in the vehicle computer.
[0044] Through the above method, accurate acquisition of interaction failure voice can be achieved, and then the voice statistics of interaction failure can be achieved.
[0045] S130. For each type of interaction failure voice, determine the number of voices corresponding to the interaction failure voice and / or the number of triggering devices corresponding to the interaction failure voice, and obtain a failure count statistical result.
[0046] Specifically, after obtaining all the interaction failure voices, all the interaction failure voices can be statistically analyzed to obtain the number of voices corresponding to each interaction failure voice, or the number of trigger devices corresponding to each interaction failure voice, or the number of voices and the number of trigger devices corresponding to each interaction failure voice.
[0047] The number of voices may be the number of all interaction failure voices under this type of interaction failure voices; the number of triggering devices may be the number of initiating devices of all interaction failure voices under this type of interaction failure voices.
[0048] For example, the number of voices corresponding to "It's raining outside" is 688, and the number of devices triggered is 41; the number of voices corresponding to "Sunrise and sunset mode" is 96, and the number of devices triggered is 8. Optionally, the number of voices and / or devices triggered by various interaction failure voices can be determined using a quantity statistics statement (e.g., "topic:recorder.stream.start").
[0049] Consider that failure statistics may vary significantly across different vehicle types. For example, the interaction failure voices differ significantly between SUVs (which often have failed voices related to initiating music) and electric scooters (which often have failed voices related to initiating vehicle controls), and between trucks (which often have failed voices related to initiating navigation) and sedans (which often have failed voices related to initiating music). Furthermore, failure statistics may also differ significantly across time periods, such as during the day (which often have failed voices related to initiating navigation) and early morning (which often have failed voices related to initiating music).
[0050] Therefore, in order to develop and optimize targeted in-vehicle voice services for different models, or for different time periods, it is also possible to separately count the failed interaction voices under different models and time periods, so as to refine the statistical results of the number of failures for each model and time period, and further provide data support for subsequent targeted service development or optimization analysis.
[0051] In a specific embodiment, for each type of interaction failure voice, the number of voices corresponding to the interaction failure voice and / or the number of trigger devices corresponding to the interaction failure voice are determined to obtain a failure count statistical result, including: determining the trigger device type corresponding to each interaction failure voice, and for each trigger device type, determining the number of voices and / or the number of trigger devices corresponding to various interaction failure voices under the trigger device type, and obtaining a failure count statistical result corresponding to the trigger device type; and / or determining the trigger time corresponding to each interaction failure voice, and for each preset time period, determining the number of voices and / or the number of trigger devices corresponding to various interaction failure voices whose trigger time is within the preset time period, and obtaining a failure count statistical result corresponding to the preset time period.
[0052] The trigger device type may be the device type of the vehicle computer that receives the interaction failure voice. Specifically, based on each interaction failure voice for each trigger device type, the number of voices and / or number of trigger devices corresponding to each interaction failure voice for each trigger device type may be determined to obtain a failure count result for each trigger device type.
[0053] The preset time period may be a pre-set time period for performing statistics on failed voice distinctions, such as 5:00-16:00, 16:00-21:00, 21:00-5:00; or Monday-Friday, Saturday-Sunday, etc.
[0054] Specifically, the preset time period to which each interaction failure voice belongs can be determined based on the trigger time corresponding to each interaction failure voice, and then the number of voices and / or the number of triggering devices corresponding to various interaction failure voices in each preset time period can be determined based on each interaction failure voice in each preset time period to obtain the statistical results of the number of failures corresponding to each preset time period.
[0055] Through the above method, the failure statistics results of each device type or each time period can be separately counted, providing data support for subsequent targeted service development or optimization analysis, and further improving the user experience.
[0056] Furthermore, in this embodiment, the determined failure count statistics can be displayed on the cloud interface, or sent to a business system so that the business system can analyze it, such as a reporting system, a machine learning system, a service recommendation system, a user portrait system, or a service optimization system.
[0057] The method for processing vehicle-computer voice interaction data provided in this embodiment obtains original voice interaction data, and determines each failed interaction voice from each historical voice according to the interaction response status indicating whether the interaction is successful corresponding to each historical voice in the original voice interaction data, and then determines the number of voices corresponding to each failed interaction voice and at least one of the number of trigger devices corresponding to each failed interaction voice, thereby realizing vehicle-computer voice statistics of failed interactions. The statistical results can be used for function development or function optimization without manual research, thus solving the problem that manual research requires a lot of cost. Moreover, by counting the historical voices of failed interactions, the actual user usage can be reflected. The development direction or optimization direction determined by the statistical results is highly accurate, thus avoiding the situation where the optimization development direction is wrong.
[0058] Figure 2 This is a flowchart of another method for processing vehicle-to-vehicle voice interaction data provided by an embodiment of the present application. Based on the above embodiments, optionally, the process of determining the voice to be optimized based on the statistical results of the number of failures is exemplified. Figure 2 As shown, the method may specifically include the following steps:
[0059] S210: Acquire original voice interaction data, wherein the original voice interaction data includes: at least one historical voice and an interaction response state corresponding to each historical voice.
[0060] S220: Determine at least one interaction-failed voice from at least one historical voice according to the interaction response state corresponding to each historical voice in the original voice interaction data.
[0061] S230. Determine the trigger device type corresponding to each interaction failure voice. For each trigger device type, determine the number of voices and / or the number of trigger devices corresponding to various interaction failure voices under the trigger device type, and obtain the failure number statistics corresponding to the trigger device type.
[0062] S240. Sort the various interaction failure voices according to the number of voices and / or the number of triggering devices corresponding to the various interaction failure voices in the failure number statistics corresponding to each trigger device type, and obtain the sorting results corresponding to each trigger device type, wherein the sorting results include: a first sorting result from large to small number or a second sorting result from small to large number.
[0063] Specifically, after determining the failure statistics corresponding to various trigger device types, the failure statistics corresponding to various trigger device types can be sorted separately to sort the interaction failure voices under various trigger device types, and then screen out the voices to be optimized under various trigger device types.
[0064] For example, the various interaction failure voices in the failure count statistics can be sorted in descending order of voice quantity or in descending order of triggering devices to obtain a first sorting result. Alternatively, the various interaction failure voices in the failure count statistics can be sorted in descending order of voice quantity or in descending order of triggering devices to obtain a second sorting result.
[0065] S250. Select the interaction failure voice before the preset threshold in the first sorting result, or select the interaction failure voice after the preset threshold in the second sorting result, as the voice to be optimized corresponding to each trigger device type, and determine the processing data output corresponding to each trigger device type based on each voice to be optimized.
[0066] The preset threshold value may be a pre-set value for filtering out failed interaction voices with a large number of voices or triggering a large number of devices. For example, the preset threshold value may be selected or entered by the user on a preset interface.
[0067] Specifically, for each trigger device type, after obtaining the first sorting result or the second sorting result, the first N (preset threshold) interaction failure voices can be selected from the first sorting result as the voices to be optimized corresponding to the trigger device type, or the last N (preset threshold) interaction failure voices can be selected from the second sorting result as the voices to be optimized corresponding to the trigger device type.
[0068] Furthermore, each voice to be optimized corresponding to the trigger device type can be directly used as the processed data output corresponding to the trigger device type, and the processed data output can be sent to the preset interface of the cloud to display each voice to be optimized on the preset interface; or, the processed data output can be sent to other business systems so that other business systems can develop services or optimize models based on each voice to be optimized.
[0069] In a specific embodiment, the processed data output corresponding to each trigger device type is determined based on each voice to be optimized, including: for each voice to be optimized corresponding to each trigger device type, determining the prediction response service corresponding to the voice to be optimized; determining the online upgrade file corresponding to each prediction response service, and sending the online upgrade file to each vehicle computer corresponding to the trigger device type, so as to add each prediction response service to each vehicle computer corresponding to the trigger device type through the online upgrade file.
[0070] The predicted response service corresponding to the voice to be optimized can be manually determined. Specifically, for each triggering device type, the predicted response service corresponding to each voice to be optimized can be obtained, and then the online upgrade file corresponding to each predicted response service can be obtained; the online upgrade file can be used to load each predicted response service into each vehicle computer corresponding to the triggering device type. Each vehicle computer corresponding to the triggering device type can be a vehicle computer integrated into a vehicle of the same type as the triggering device.
[0071] For example, the online upgrade file corresponding to the triggering device type can be proactively sent to each vehicle computer corresponding to the triggering device type, or each vehicle computer corresponding to the triggering device type can proactively obtain the online upgrade file. Furthermore, after each vehicle computer corresponding to the triggering device type runs the online upgrade file, the online upgrade file can be used to add each prediction response service to the vehicle computer.
[0072] Through the above method, for each trigger device type, the interaction failure voices with a large number of statistical results can be optimized, and the online upgrade file can be output as the processing data of the trigger device type. The service corresponding to the interaction failure voice with a large number of failures can be added to each vehicle computer corresponding to the trigger device type. This allows the corresponding response service to be provided the next time a user of the vehicle computer corresponding to the trigger device type initiates the interaction failure voice. This enables the development of interactive services with high user demand, adds new functions that meet the actual needs of users to the vehicle computer voice interaction service, and further optimizes the user's voice interaction experience. In addition, targeted optimization of vehicle computers with different trigger device types is achieved, which improves the accuracy of service optimization and thus enhances the user experience of users of various vehicle models.
[0073] It should be noted that, in addition to determining the voices to be optimized corresponding to different trigger device types, it is also possible to determine the voices to be optimized corresponding to different preset time periods.
[0074] For example, if the failure number statistics corresponding to each preset time period are obtained, the various interaction failure voices are sorted according to the number of voices and / or the number of triggering devices corresponding to the various interaction failure voices in the failure number statistics corresponding to each preset time period, to obtain the sorting results corresponding to each preset time period; wherein the sorting results include: a first sorting result with the number from most to least or a second sorting result with the number from least to most; selecting the interaction failure voice before the preset threshold in the first sorting result, or selecting the interaction failure voice after the preset threshold in the second sorting result, as the voice to be optimized corresponding to each preset time period; determining the processing data output corresponding to each preset time period based on each voice to be optimized.
[0075] Alternatively, all statistical results of the number of failures may be directly sorted and analyzed to obtain the speech to be optimized in all statistical results of the number of failures.
[0076] For example, according to the number of voices and / or the number of triggering devices corresponding to various interaction failure voices in the failure number statistics results, various interaction failure voices are sorted to obtain a sorting result; wherein the sorting result includes: a first sorting result with the number from large to small or a second sorting result with the number from small to large; an interaction failure voice before a preset threshold in the first sorting result is selected, or an interaction failure voice after a preset threshold in the second sorting result is selected, as the voice to be optimized; and the processing data output is determined according to each voice to be optimized.
[0077] In a specific embodiment, the corresponding processed data output is determined according to each voice to be optimized, including: determining the predicted recognition operation corresponding to each voice to be optimized; outputting the updated operation recognition model according to each voice to be optimized and the predicted recognition operation corresponding to each voice to be optimized; wherein the operation recognition model is used to determine the recognition operation corresponding to the current voice when the user's current voice is obtained.
[0078] The voices to be optimized may include the voices to be optimized corresponding to the trigger device types, or may also include the voices to be optimized obtained by sorting the statistical results of all failures.
[0079] The predicted recognition operation corresponding to the speech to be optimized can be manually determined. The operation recognition model can be a pre-trained neural network model, such as a long short-term memory network model or a convolutional neural network model. Specifically, the predicted recognition operation corresponding to each speech to be optimized can be obtained, and then the pre-trained operation recognition model can be updated based on each speech to be optimized and the predicted recognition operation corresponding to each speech to be optimized, thereby optimizing the operation recognition model.
[0080] It should be noted that the purpose of updating the pre-trained operation recognition model is: since the voice to be optimized may be a voice that cannot recognize the operation, in order to further optimize the cloud-based semantic understanding ability, the interaction failure voices with a large number in the statistical results can be determined as voices containing certain semantic information, and then the operation recognition model is optimized and trained to update the network parameters in the operation recognition model, and then the updated operation recognition model is sent to each car computer, so that each car computer can determine the corresponding recognition operation according to the operation recognition model when receiving the user's current voice.
[0081] Of course, in the above process, the operation recognition model can be directly optimized and trained based on each voice to be optimized and the predicted recognition operation corresponding to each voice to be optimized. The voices that can recognize the operation can also be eliminated from the voices to be optimized, so that the operation recognition model can be optimized and trained only based on the voices to be optimized that cannot recognize the operation.
[0082] Through the above method, voices with training value can be screened out from the failed interaction voices, that is, the voices to be optimized, and then the operation recognition model can be optimized and trained through the voices to be optimized, and the optimized operation recognition model can be output to the user's terminal (such as the car computer), thereby optimizing the semantic understanding ability, avoiding the situation where the semantic understanding of the voice initiated by the user is empty, and further improving the user's car-computer voice interaction experience.
[0083] The method for processing vehicle-machine voice interaction data provided in this embodiment is to sort the statistical results of the number of failures under different trigger device types respectively, and obtain a first sorting result from the largest to the smallest number or a second sorting result from the smallest to the largest number, and then select the first preset threshold number of interaction failure voices from the first sorting result as the voices to be optimized corresponding to the trigger device type, or select the second preset threshold number of interaction failure voices from the second sorting result as the voices to be optimized corresponding to the trigger device type, so as to select the interaction failure voices with a large number of voices or a large number of trigger devices as the voices to be optimized corresponding to the trigger device type, providing data support for the mining of new scenarios for vehicle-machine voice services, so as to facilitate subsequent targeted optimization and improve the user's vehicle-machine voice interaction experience. In addition, by separately determining the voices to be optimized for different trigger device types, differentiated optimization of different trigger device types can be achieved, further improving the optimization accuracy and thus improving the user experience.
[0084] For example, Figure 3The figure shows a schematic diagram of the process for processing in-vehicle voice interaction data. The voice server can collect and store raw voice interaction data from both the test environment and the production environment, and send the raw voice interaction data to Kafka. Kafka then sends the raw voice interaction data to HIVE / Hadoop / the scheduler. HIVE / Hadoop / the scheduler uses RDS / MySQL to identify failed voice interactions from the raw voice interaction data, compile statistics on each failed voice interaction, and then sort the failed voice interactions to determine the voices to be optimized based on the sorted results.
[0085] Furthermore, each voice to be optimized can be displayed. Alternatively, each voice to be optimized can be sent to a reporting system to conduct user behavior research and explore new scenarios for in-vehicle voice services. Alternatively, machine learning can be further performed to optimize the intent recognition model. Alternatively, the data can be sent to a development system to develop corresponding response services for each voice to be optimized.
[0086] It should be noted that the method of the embodiment of the present application can be performed by a single device, such as a computer or server. The method of this embodiment can also be applied in a distributed scenario and performed by multiple devices working together. In such a distributed scenario, one of the multiple devices may only perform one or more steps of the method of the embodiment of the present application, and the multiple devices will interact with each other to complete the method.
[0087] It should be noted that the above description is limited to some embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in an order different from that described in the above embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0088] Based on the same inventive concept, corresponding to any of the above-mentioned embodiment methods, the present application also provides a device for processing vehicle-computer voice interaction data. Figure 4 A schematic diagram of the structure of a device for processing vehicle-machine voice interaction data provided in an embodiment of the present application.
[0089] refer to Figure 4 The vehicle-machine voice interaction data processing device includes a data acquisition module 410, a failed voice determination module 420 and a failure statistics module 430, wherein;
[0090] The data acquisition module 410 is configured to acquire original voice interaction data, wherein the original voice interaction data includes: at least one historical voice and an interaction response status corresponding to each historical voice, wherein the interaction response status indicates whether the interaction with the historical voice is successful;
[0091] a failed speech determination module 420 for determining at least one failed interaction speech from at least one of the historical speech according to the interaction response state corresponding to each of the historical speech in the original voice interaction data;
[0092] The failure statistics module 430 is used to determine, for each type of interaction failure voice, the number of voices corresponding to the interaction failure voice and / or the number of triggering devices corresponding to the interaction failure voice, and obtain a failure number statistics result.
[0093] The device for processing vehicle-machine voice interaction data provided by the present application obtains original voice interaction data, and determines each failed interaction voice from each historical voice according to the interaction response status corresponding to each historical voice in the original voice interaction data, and then determines the number of voices corresponding to each failed interaction voice and at least one of the number of trigger devices corresponding to each failed interaction voice, thereby realizing vehicle-machine voice statistics of failed interactions. The statistical results can be used for function development or function optimization without manual research, thus solving the problem of manual research consuming a lot of costs. Moreover, by counting the historical voices of failed interactions, the actual user usage can be reflected. The development direction or optimization direction determined by the statistical results is highly accurate, thus avoiding the situation where the optimization development direction is wrong.
[0094] Based on the above embodiments, optionally, the interactive response state includes: a recognition operation state and a response service state corresponding to the recognition operation state, wherein the recognition operation state is used to indicate whether the recognition operation corresponding to the historical voice is successfully recognized, and the response service state is used to indicate whether the response service corresponding to the recognition operation is successfully recognized; the failed voice determination module 420 is specifically used to:
[0095] For each historical voice, if the recognition operation status corresponding to the historical voice is recognized as failure, the historical voice is determined to be an interaction failure voice; and / or, for each historical voice, if the recognition operation status corresponding to the historical voice is recognized as success, and the response service status corresponding to the recognition operation status is failure, the historical voice is determined to be an interaction failure voice.
[0096] On the basis of the above-mentioned implementation modes, optionally, the failed voice determination module 420 is further used to determine, for each historical voice, whether the historical reply corresponding to the historical voice is the first preset reply; if so, determine that the recognition operation status corresponding to the historical voice is failed, and determine the historical voice as an interaction failure voice; and / or, for each historical voice, determine whether the historical reply corresponding to the historical voice is the second preset reply; if so, determine that the recognition operation status corresponding to the historical voice is successful, and the response service status corresponding to the recognition operation status is failure, and determine the historical voice as an interaction failure voice.
[0097] On the basis of the above-mentioned embodiments, optionally, the failure statistics module 430 is further used to determine the trigger device type corresponding to each of the interaction failure voices, and for each of the trigger device types, determine the number of voices and / or the number of trigger devices corresponding to the various interaction failure voices under the trigger device type, and obtain the failure number statistics corresponding to the trigger device type; and / or determine the trigger time corresponding to each of the interaction failure voices, and for each preset time period, determine the number of voices and / or the number of trigger devices corresponding to the various interaction failure voices whose trigger time is within the preset time period, and obtain the failure number statistics corresponding to the preset time period.
[0098] On the basis of the above-mentioned embodiments, optionally, the processing device of the vehicle-computer voice interaction data also includes a voice determination module to be optimized, and the voice determination module to be optimized is used to sort the various interaction failure voices according to the number of voices and / or the number of trigger devices corresponding to the various interaction failure voices in the failure number statistics corresponding to each trigger device type, if the failure number statistics corresponding to each trigger device type are obtained, to obtain the sorting results corresponding to each trigger device type; wherein the sorting results include: a first sorting result with a large number to a small number or a second sorting result with a small number to a large number; selecting the interaction failure voice before the preset threshold in the first sorting result, or selecting the interaction failure voice after the preset threshold in the second sorting result, as the voice to be optimized corresponding to each trigger device type; and determining the processing data output corresponding to each trigger device type according to each voice to be optimized.
[0099] On the basis of the above-mentioned embodiments, optionally, the voice determination module to be optimized also includes a service delivery unit, which is used to determine the predicted response service corresponding to each voice to be optimized corresponding to each trigger device type, determine the online upgrade file corresponding to each predicted response service, and deliver the online upgrade file to each vehicle computer corresponding to the trigger device type, so as to add each predicted response service to each vehicle computer corresponding to the trigger device type through the online upgrade file.
[0100] On the basis of the above-mentioned embodiments, optionally, the voice determination module to be optimized also includes a model updating unit, which is used to determine the predicted recognition operation corresponding to each of the voices to be optimized; according to each of the voices to be optimized and the predicted recognition operation corresponding to each of the voices to be optimized, the pre-trained operation recognition model is updated, and the updated operation recognition model is output; wherein, the operation recognition model is used to determine the recognition operation corresponding to the current voice when the user's current voice is acquired.
[0101] For the convenience of description, the above devices are described as being divided into various modules according to their functions. Of course, when implementing this application, the functions of each module can be implemented in the same or multiple software and / or hardware.
[0102] The device of the above embodiment is used to implement the corresponding vehicle-machine voice interaction data processing method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be repeated here.
[0103] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments and methods, the present application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the program, it implements the method for processing vehicle-machine voice interaction data described in any of the above embodiments.
[0104] Figure 5 10 is a schematic diagram showing a more specific hardware structure of an electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other within the device via the bus 1050.
[0105] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0106] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage devices, dynamic storage devices, etc. The memory 1020 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.
[0107] The input / output interface 1030 is used to connect input / output modules to implement information input and output. The input / output modules can be configured as components within the device (not shown in the figure) or can be externally connected to the device to provide corresponding functions. Input devices may include a keyboard, mouse, touch screen, microphone, various sensors, etc., and output devices may include a display, speaker, vibrator, indicator light, etc.
[0108] The communication interface 1040 is used to connect to a communication module (not shown) to enable communication between the device and other devices. The communication module can communicate via a wired method (such as USB, network cable, etc.) or a wireless method (such as mobile network, WiFi, Bluetooth, etc.).
[0109] The bus 1050 comprises a path for transmitting information between the various components of the device (eg, the processor 1010 , the memory 1020 , the input / output interface 1030 , and the communication interface 1040 ).
[0110] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in a specific implementation, the device may also include other components necessary for normal operation. In addition, it will be understood by those skilled in the art that the above device may only include the components necessary to implement the embodiments of this specification, and does not necessarily include all the components shown in the figure.
[0111] The electronic device of the above embodiment is used to implement the corresponding vehicle-machine voice interaction data processing method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be repeated here.
[0112] Based on the same inventive concept, corresponding to any of the above-mentioned embodiment methods, the present application also provides a non-transitory computer-readable storage medium, which stores computer instructions, and the computer instructions are used to enable the computer to execute the method for processing vehicle-machine voice interaction data as described in any of the above embodiments.
[0113] The computer-readable media of this embodiment include permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.
[0114] The computer instructions stored in the storage medium of the above embodiment are used to enable the computer to execute the method for processing vehicle-machine voice interaction data as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0115] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present application (including the claims) is limited to these examples. Within the scope of the present application, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of the different aspects of the embodiments of the present application as described above, which are not provided in detail for the sake of simplicity.
[0116] In addition, for simplicity of description and discussion, and in order not to make the embodiment of the application difficult to understand, the known power supply / ground connection with integrated circuit (IC) chip and other components may or may not be shown in the accompanying drawings provided. In addition, the device can be shown in the form of a block diagram to avoid making the embodiment of the application difficult to understand, and this also takes into account the following fact, that is, the details of the embodiment of these block diagram devices are highly dependent on the platform to be implemented in the embodiment of the application (that is, these details should be fully within the scope of understanding of those skilled in the art). When specific details (for example, circuit) are set forth to describe exemplary embodiments of the application, it will be apparent to those skilled in the art that the embodiment of the application can be implemented without these specific details or when these specific details are changed. Therefore, these descriptions should be considered to be illustrative rather than restrictive.
[0117] Although the present invention has been described in conjunction with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those skilled in the art based on the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may utilize the embodiments discussed.
[0118] The embodiments of the present application are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present application should be included in the scope of protection of this application.
Claims
1. A method for processing vehicle-to-vehicle voice interaction data, characterized in that: include: Obtaining original voice interaction data, wherein the original voice interaction data includes: at least one historical voice, and an interaction response status corresponding to each historical voice, the interaction response status being used to indicate whether the interaction with the historical voice is successful, the interaction response status including: a recognition operation status and a response service status corresponding to the recognition operation status, the recognition operation status being used to indicate whether the recognition operation corresponding to the historical voice is successful, and the response service status being used to indicate whether the response service corresponding to the recognition operation is successful; For each historical voice, if the recognition operation status corresponding to the recognized historical voice is failure, the historical voice is determined to be an interaction failure voice; and / or, for each historical voice, if the recognition operation status corresponding to the recognized historical voice is success and the response service status corresponding to the recognition operation status is failure, the historical voice is determined to be an interaction failure voice; Determine the trigger device type corresponding to each of the failed interaction voices, and for each of the trigger device types, determine the number of voices and / or the number of trigger devices corresponding to the various failed interaction voices under the trigger device type, and obtain the failure number statistics corresponding to the trigger device type, wherein the failure number statistics corresponding to various trigger device types are sorted respectively, and the voices to be optimized under various trigger device types are screened out, and the services corresponding to the failed interaction voices with a large number of failures are added to each car computer corresponding to the trigger device type, and the operation recognition model is optimized and trained according to each voice to be optimized and the predicted recognition operation corresponding to each voice to be optimized.
2. The method according to claim 1, characterized in that The method further comprises: If the failure count statistics corresponding to each trigger device type are obtained, sorting the various interaction failure voices according to the number of voices and / or the number of trigger devices corresponding to the various interaction failure voices in the failure count statistics corresponding to each trigger device type to obtain a sorting result corresponding to each trigger device type; wherein the sorting result includes: a first sorting result from large to small number or a second sorting result from small to large number; The interaction failure voices before the preset threshold in the first sorting result are selected, or the interaction failure voices after the preset threshold in the second sorting result are selected, as the voices to be optimized corresponding to the triggering device types.
3. The method according to claim 2, characterized in that The method further comprises: For each of the voices to be optimized corresponding to each of the trigger device types, determine the predicted response service corresponding to the voice to be optimized, determine the online upgrade file corresponding to each of the predicted response services, and send the online upgrade file to each of the vehicle computers corresponding to the trigger device type, so as to add each of the predicted response services to each of the vehicle computers corresponding to the trigger device type through the online upgrade file.
4. The method according to claim 2, characterized in that The method further comprises: Determining the prediction recognition operation corresponding to each of the to-be-optimized speech; updating a pre-trained operation recognition model according to each of the to-be-optimized voices and the predicted recognition operation corresponding to each of the to-be-optimized voices, and outputting the updated operation recognition model; The operation recognition model is used to determine the recognition operation corresponding to the current voice when the user's current voice is acquired.
5. The method according to claim 1, wherein If the recognition operation status corresponding to the historical voice is recognized as failed, it includes: Determining whether the historical response corresponding to the historical voice is a first preset response, and if so, determining that the recognition operation status corresponding to the historical voice is failed; Correspondingly, if the recognition operation status corresponding to the historical voice is recognized as successful, it includes: It is determined whether the historical reply corresponding to the historical voice is a second preset reply. If so, it is determined that the recognition operation status corresponding to the historical voice is successful.
6. A device for processing vehicle-to-vehicle voice interaction data, characterized in that: include: a data acquisition module, configured to acquire original voice interaction data, wherein the original voice interaction data includes: at least one historical voice and an interaction response status corresponding to each historical voice, the interaction response status being used to indicate whether the interaction with the historical voice was successful, the interaction response status including: a recognition operation status and a response service status corresponding to the recognition operation status, the recognition operation status being used to indicate whether the recognition operation corresponding to the historical voice was successful, and the response service status being used to indicate whether the response service corresponding to the recognition operation was successful; a failed voice determination module, configured to determine at least one failed interaction voice from at least one of the historical voices based on the interaction response status corresponding to each of the historical voices in the original voice interaction data, wherein the failed voice determination module is specifically configured to: for each historical voice, if the recognition operation status corresponding to the historical voice is recognized as failed, determine the historical voice as an failed interaction voice; and / or, for each historical voice, if the recognition operation status corresponding to the historical voice is recognized as successful and the response service status corresponding to the recognition operation status is failed, determine the historical voice as an failed interaction voice; A failure statistics module is used to determine the trigger device type corresponding to each of the interaction failure voices, and for each of the trigger device types, determine the number of voices and / or the number of trigger devices corresponding to various interaction failure voices under the trigger device type, and obtain a failure number statistical result corresponding to the trigger device type; The device also includes a voice determination module to be optimized, which is used to sort the failure number statistics corresponding to various trigger device types, screen out the voices to be optimized under various trigger device types, add the services corresponding to the interaction failure voices with a large number of failures to each car computer corresponding to the trigger device type, and optimize and train the operation recognition model according to each voice to be optimized and the predicted recognition operation corresponding to each voice to be optimized.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method for processing vehicle-machine voice interaction data as described in any one of claims 1 to 5 is implemented.
8. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to enable a computer to execute the method for processing vehicle-machine voice interaction data as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Interaction processing method, device and equipment and audio equipment
CN110347248A
Vehicle voice interaction system and interaction method, and vehicle
CN114512126A