An optimization method for improving the recognition results of page voice intents

By analyzing and generating corrected voice configuration data, combined with real-time voice recognition data, the intelligent voice control system can more accurately identify user intentions and perform operations, solving the problem of low accuracy of voice intention recognition and improving system response speed and user experience.

CN119864022BActive Publication Date: 2025-06-27BEIJING NELDA TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510336432.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-06-27
Estimated Expiration
2045-03-21

AI Technical Summary

Technical Problem

In the intelligent voice control system, the accuracy of voice intention recognition is low, and alias learning and correction cannot be performed through the front-end page, resulting in poor user voice control experience.

Method used

By analyzing system configuration data, preset voice configuration data and historical voice recognition data, the object correction data, voice operation data and correction vector data are determined to generate corrected voice configuration data. Based on real-time voice recognition data and corrected voice configuration data, real-time intent data are determined and real-time operations are performed.

Benefits of technology

The response speed and accuracy of the intelligent voice control system are optimized, the ability to understand users' voice commands is enhanced, the accuracy and user experience of voice control are improved, manual operations are reduced by the maintenance personnel, and operation efficiency is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119864022B_ABST
    Figure CN119864022B_ABST
Patent Text Reader

Abstract

The present invention provides an optimization method for improving the recognition result of page voice intent, belonging to the technical field of data processing, including: Step 1: Determine the system configuration data of the intelligent voice control system and the preset voice configuration data, and obtain the real-time voice recognition data and the historical voice recognition data; Step 2: Determine the object correction data, the voice operation data and the correction vector data, and determine the corrected voice configuration data of the intelligent voice control system; Step 3: Determine the real-time intent data; Step 4: Determine the real-time operation data based on the real-time intent data, and perform real-time response to the voice control system based on the real-time operation data. It can optimize the response speed and accuracy of the intelligent voice control system, enhance the understanding ability of the user's voice commands, improve the accuracy of voice control, improve the user's voice control experience, reduce the manual operation of the maintenance personnel, and improve the operation efficiency of the maintenance personnel.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly to an optimization method for improving the recognition result of page voice intent. Background Art

[0002] Intelligent voice recognition technology has a development history of several decades. Early voice recognition technology was mainly based on limited vocabulary and pattern recognition algorithms, and its effect was limited by hardware performance and the quality of voice data. In recent years, intelligent voice control systems have gradually introduced big data analysis and machine learning technologies to realize functions such as interface jump, interface element control, and information entry through intelligent voice technology, achieving the function control of the entire system. However, there are still disadvantages such as low accuracy of voice intent recognition, inability to perform alias learning and correction through the front-end page.

[0003] Therefore, the present invention provides an optimization method for improving the recognition result of page voice intent. Summary of the Invention

[0004] The present invention provides an optimization method for improving the recognition result of page voice intent. By analyzing the determined system configuration data, preset voice configuration data, and the obtained historical voice recognition data, object correction data, voice operation data, and correction vector data are determined, and the corrected voice configuration data of the intelligent voice control system is determined. According to the real-time voice recognition data and the corrected voice configuration data, real-time intent data is determined, real-time operation data is determined and executed. It can optimize the response speed and accuracy of the intelligent voice control system, enhance the ability to understand user voice commands, improve the accuracy of voice control, improve the user voice control experience, reduce the manual operation of maintenance personnel, and improve the operation efficiency of maintenance personnel.

[0005] The present invention provides an optimization method for improving the recognition result of page voice intent, including:

[0006] Step 1: Determine the system configuration data and preset voice configuration data of the intelligent voice control system, and obtain real-time voice recognition data and historical voice recognition data;

[0007] Step 2: Based on the system configuration data, preset voice configuration data, and historical voice recognition data, determine object correction data, voice operation data, and correction vector data, and determine the corrected voice configuration data of the intelligent voice control system;

[0008] Step 3: Based on the real-time voice recognition data and the corrected voice configuration data of the intelligent voice control system, determine real-time intent data;

[0009] Step 4: Determine real-time operation data based on the real-time intent data, and perform real-time response to the voice control system based on the real-time operation data.

[0010] An optimization method for improving the page voice intent recognition result provided by the present invention determines the system configuration data and preset voice configuration data of the intelligent voice control system, including:

[0011] Obtain the page configuration data of each page in the intelligent voice control system, where the page configuration data at least includes a plurality of page configuration objects, the page configuration object types of each page configuration object, the operation actions of each page configuration object, and a plurality of operation action aliases of each operation action of each page configuration object;

[0012] Based on the page configuration data of all pages in the intelligent voice control system, determine the system configuration data of the intelligent voice control system;

[0013] Determine a plurality of voice mapping instructions based on the system configuration data of the intelligent voice control system, and determine the object mapping set of each page configuration object based on all voice mapping instructions;

[0014] Determine the preset voice configuration data based on the object mapping sets of all page configuration objects.

[0015] An optimization method for improving the page voice intent recognition result provided by the present invention obtains real-time voice recognition data and historical voice recognition data, including:

[0016] Obtain the real-time voice recognition data within the current specified time period;

[0017] Obtain historical voice recognition sub-data within a plurality of historical specified time periods before the current specified time period, where the historical voice recognition sub-data includes a plurality of historical operation instructions, the voice configuration objects of each historical operation instruction, the operation labels of each historical operation instruction, and the operation actions of each historical operation instruction;

[0018] Based on the historical voice recognition sub-data of all historical specified time periods, determine the historical voice recognition data.

[0019] An optimization method for improving the page voice intent recognition result provided by the present invention determines object correction data, voice operation data, and correction vector data based on the system configuration data, preset voice configuration data, and historical voice recognition data, and determines the corrected voice configuration data of the intelligent voice control system, including:

[0020] Extract all page configuration objects in the system configuration data to determine the page configuration object set;

[0021] Extract the voice configuration objects of all historical operation instructions in the historical voice recognition sub-data of all historical specified time periods in the historical voice recognition data to determine the voice configuration object set;

[0022] Determine that the number of page configuration objects in the page configuration object set is the number of clusters K for cluster analysis. Randomly select K voice configuration objects from the voice configuration object set as the initial cluster centers, perform cluster analysis on the voice configuration object set to determine K clusters, the cluster center configuration object of each cluster, and the subset of voice configuration objects of each cluster;

[0023] Extract all operation tags of all voice configuration objects in the subset of voice configuration objects of each cluster from the historical speech recognition sub-data within each specified time period based on the historical speech recognition data, and determine the operation tag set of the cluster center configuration object of each cluster;

[0024] Based on the system configuration data, historical speech recognition data, and the operation tag set of the cluster center configuration object of each cluster, determine the cluster center configuration object matrix of the cluster center configuration object of each cluster within each specified time period, and determine the correction success vector of the cluster center configuration object of each cluster;

[0025] Determine the correction mapping set, operation tag set, and correction success vector of each page configuration object;

[0026] Determine the object correction data based on the correction mapping sets of all page configuration objects in the page configuration object set, determine the voice operation data based on the operation tag sets of all page configuration objects in the page configuration object set, and at the same time, determine the correction vector data based on the correction success vectors of all page configuration objects in the page configuration object set;

[0027] Determine the corrected voice configuration data of the intelligent voice control system based on the object correction data, voice operation data, and correction vector data.

[0028] According to an optimization method for improving the page speech intention recognition result provided by the present invention, determining the cluster center configuration object matrix of the cluster center configuration object of each cluster within each specified time period includes:

[0029] Based on the system configuration data, historical speech recognition data, and the operation tag set of the cluster center configuration object of each cluster, determine the cluster center configuration object matrix of the cluster center configuration object of each cluster within each specified time period;

[0030] ;

[0031] ;

[0032] wherein, represents the cluster center configuration object matrix of the cluster center configuration object of the a-th cluster in the t-th historical specified time period, Indicates that the operation label in the operation label set of the cluster center configuration object of the a-th cluster family in the t-th historical specified time period is an operation success, Indicates the number of operation labels in the operation label set of the cluster center configuration object of the a-th cluster family, Respectively indicate the i-th operation label and the aN2-th operation label in the operation label set of the cluster center configuration object of the a-th cluster family in the t-th historical specified time period, Indicates the operation success rate of the operation label in the operation label set of the cluster center configuration object of the a-th cluster family in the t-th historical specified time period, Indicates the operation failure rate of the i-th operation label and the aN2-th operation label in the operation label set of the cluster center configuration object of the a-th cluster family in the t-th historical specified time period, Indicates the operation label of the L-th historical operation instruction of the j-th voice configuration object in the subset of voice configuration objects of the a-th cluster family in the t-th historical specified time period, a Indicates the number of all historical operation instructions based on the j-th voice configuration object in the subset of voice configuration objects of the a-th cluster family, Indicates the number of voice configuration objects in the subset of voice configuration objects of the a-th cluster family in the t-th historical specified time period.

[0033] According to an optimization method for improving the page voice intent recognition result provided by the present invention, determining the correction success vector of the cluster center configuration object of each cluster family includes: determining the correction success vector of the cluster center configuration object of each cluster family based on the cluster center configuration object matrix of the cluster center configuration object of each cluster family within all specified time periods;

[0034] ;

[0035] ;

[0036] Wherein, Indicates the correction success vector of the cluster center configuration object of the a-th cluster family, Respectively indicate the correction success rates of the 2nd operation label, the i-th operation label, and the aN2-th operation label in the operation label set of the cluster center configuration object of the a-th cluster family, Indicates the period decay factor, Indicates the correction factor, and Nu indicates the number of historical specified time periods.

[0037] According to an optimization method for improving the page voice intent recognition result provided by the present invention, determining the correction mapping set, the operation label set, and the correction success vector of each page configuration object includes:

[0038] Determine the corrected mapping set for each page configuration object based on the object mapping set of each page configuration object in the preset voice configuration data, all page configuration objects in the page configuration object set, the cluster center configuration objects of all clusters, and the sub-set of voice configuration objects;

[0039] ;;

[0040] ;

[0041] wherein, represents the similarity value between the x-th page configuration object in the page configuration object set and the a-th cluster; represents the corrected mapping set of the x-th page configuration object in the page configuration object set; represents the cluster center configuration object of the a-th cluster; represents the j-th voice configuration object in the sub-set of voice configuration objects of the a-th cluster; represents the Euclidean norm of the x-th page configuration object in the page configuration object set; represents the Euclidean norm of the cluster center configuration object of the a-th cluster; represents the Euclidean norm of the j-th voice configuration object in the sub-set of voice configuration objects of the a-th cluster; represents the sub-set of voice configuration objects of the a-th cluster; represents the sub-set of voice configuration objects with the largest similarity value to the x-th page configuration object in the page configuration object set; represents the object mapping set of the x-th page configuration object in the page configuration object set;

[0042] Determine the operation label set for each page configuration object in the page configuration object set based on the operation label set of the cluster corresponding to the sub-set of voice configuration objects with the largest similarity value to each page configuration object in the page configuration object set;

[0043] Determine the corrected success vector for each page configuration object in the page configuration object set based on the corrected success vector of the cluster corresponding to the sub-set of voice configuration objects with the largest similarity value to each page configuration object in the page configuration object set.

[0044] According to an optimization method for improving the page voice intent recognition result provided by the present invention, determine the real-time intent data based on the real-time speech recognition data and the corrected voice configuration data of the intelligent voice control system, including:

[0045] Optimize the intent model based on the correction vector data in the corrected voice configuration data, parse the real-time speech recognition data based on the optimized intent model. If the parsing is successful, determine the predicted intent data, match the predicted intent data with the correction mapping sets of each page configuration object in the corrected voice configuration data, and determine the real-time intent data based on the matching result;

[0046] If the parsing is unsuccessful, perform a hard match on the real-time speech recognition data to determine the real-time intent data.

[0047] According to an optimization method for improving the page speech intent recognition result provided by the present invention, determining the real-time intent data based on the matching result includes:

[0048] If there is a page configuration object in the corrected voice configuration data that highly matches the predicted intent data, determine the real-time intent data of the real-time speech data based on the correction mapping set and the operation label set of the page configuration object that highly matches the predicted intent data;

[0049] If there is no page configuration object in the corrected voice configuration data that highly matches the predicted intent data, re-correct the correction mapping set and the operation label set of the page configuration object with the highest match to the predicted intent data.

[0050] According to an optimization method for improving the page speech intent recognition result provided by the present invention, determining the real-time operation data based on the real-time intent data and performing a real-time response to the voice control system based on the real-time operation data includes:

[0051] Parse the real-time intent data to determine the real-time operation data, where the real-time operation data includes at least one or more real-time operation instructions;

[0052] The voice control system executes each real-time operation instruction in the real-time operation data and determines the operation label of each real-time operation instruction;

[0053] Based on all the real-time operation instructions in the real-time operation data and the operation label of each real-time operation instruction, determine the current voice recognition sub-data for the current specified time period.

[0054] Compared with the prior art, the beneficial effects of the present application are as follows:

[0055] Based on the system configuration data determined through analysis, the preset voice configuration data, and the acquired historical speech recognition data, determine the object correction data, voice operation data, and correction vector data, and determine the corrected voice configuration data of the intelligent voice control system. Based on the real-time speech recognition data and the corrected voice configuration data, determine the real-time intent data, determine the real-time operation data, and execute it. This can optimize the response speed and accuracy of the intelligent voice control system, enhance the ability to understand user voice commands, improve the accuracy of voice control, improve the user voice control experience, reduce the manual operations of maintenance personnel, and improve the operation efficiency of maintenance personnel. Brief Description of the Drawings

[0056] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0057] Figure 1 It is a flowchart showing an optimization method for improving the page speech intent recognition result provided by an embodiment of the present invention. Detailed Embodiments

[0058] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0059] Embodiment 1:

[0060] An embodiment of the present invention provides an optimization method for improving the page speech intent recognition result, as Figure 1 shown, including:

[0061] Step 1: Determine the system configuration data and the preset voice configuration data of the intelligent voice control system, and acquire the real-time speech recognition data and the historical speech recognition data;

[0062] Step 2: Based on the system configuration data, the preset voice configuration data, and the historical speech recognition data, determine the object correction data, the voice operation data, and the correction vector data, and determine the corrected voice configuration data of the intelligent voice control system;

[0063] Step 3: Based on the real-time speech recognition data and the corrected voice configuration data of the intelligent voice control system, determine the real-time intent data;

[0064] Step 4: Determine real-time operation data based on real-time intent data, and perform real-time response to the voice control system based on the real-time operation data.

[0065] In this embodiment, by combining system configuration data, voice configuration data, and historical voice recognition data, object correction data, voice operation data, and correction vector data are generated through analysis, and then corrected voice configuration data is generated.

[0066] In this embodiment, using real-time voice recognition data and corrected voice configuration data, real-time intent data is recognized and determined according to the intent model result, that is, by parsing the voice command, the user's real-time intent is determined.

[0067] In this embodiment, real-time operation data (specific real-time operation instructions prepared by the system according to the intent data) is further determined based on the real-time intent data. Then, the intelligent voice control system will perform real-time response and execute the corresponding operations.

[0068] Beneficial effects of the above technical solution: By analyzing and determining the system configuration data, preset voice configuration data, and the obtained historical voice recognition data, object correction data, voice operation data, and correction vector data are determined, and the corrected voice configuration data of the intelligent voice control system is determined. According to the real-time voice recognition data and the corrected voice configuration data, real-time intent data is determined, real-time operation data is determined and executed. It can optimize the response speed and accuracy of the intelligent voice control system, enhance the ability to understand the user's voice commands, improve the accuracy of voice control, improve the user's voice control experience, reduce the manual operation of the maintenance personnel, and improve the operation efficiency of the maintenance personnel.

[0069] Embodiment 2:

[0070] The embodiment of the present invention provides an optimization method for improving the page voice intent recognition result, which determines the system configuration data and preset voice configuration data of the intelligent voice control system, including:

[0071] Obtain the page configuration data of each page in the intelligent voice control system, where the page configuration data at least includes multiple page configuration objects, the page configuration object types of each page configuration object, the operation actions of each page configuration object, and multiple operation action aliases of each operation action of each page configuration object;

[0072] Based on the page configuration data of all pages in the intelligent voice control system, determine the system configuration data of the intelligent voice control system;

[0073] Based on the system configuration data of the intelligent voice control system, determine multiple voice mapping instructions, and determine the object mapping set of each page configuration object based on all the voice mapping instructions;

[0074] Determine preset voice configuration data based on the object mapping set of all page configuration objects.

[0075] In this embodiment, the configuration data of each page contains multiple page configuration objects, and each object has a type, an operation action, and multiple aliases of the operation action. For example, a page may contain a configuration object of "table lamp", the type of this object is "switch device", the operation action is "turn on", and the aliases of the operation action are "turn on", "light up", etc.

[0076] In this embodiment, by integrating the page configuration data of all pages, the global configuration of the intelligent voice control system is determined. That is, all page object types and action aliases constitute the overall configuration of the system, ensuring that the system can support all voice commands and operations.

[0077] In this embodiment, according to the system configuration data, multiple voice instructions are formulated to map to the corresponding page configuration objects. For example, "turn on the light" can be mapped to the "table lamp" configuration object on a certain page.

[0078] In this embodiment, a voice instruction mapping related to each page configuration object is created to ensure that the voice recognition system can correctly understand different instructions and perform corresponding operations.

[0079] In this embodiment, through the object mapping relationship, a set of preset voice configurations are defined, enabling the voice control system to quickly respond to the user's voice instructions in different scenarios.

[0080] Beneficial effects of the above technical solution: Determining the system configuration data and preset voice configuration data of the intelligent voice control system can provide a data basis for determining the corrected voice configuration data, thereby improving the accuracy and response speed of voice instruction recognition and enhancing the user interaction experience.

[0081] Embodiment 3:

[0082] The embodiment of the present invention provides an optimization method for improving the page voice intent recognition result, which obtains real-time voice recognition data and historical voice recognition data, including:

[0083] Obtain real-time voice recognition data within the current specified time period;

[0084] Obtain historical voice recognition sub-data within multiple historical specified time periods before the current specified time period, where the historical voice recognition sub-data includes multiple historical operation instructions, the voice configuration object of each historical operation instruction, the operation label of each historical operation instruction, and the operation action of each historical operation instruction;

[0085] Based on the historical voice recognition sub-data of all historical specified time periods, determine the historical voice recognition data.

[0086] In this embodiment, the real-time speech recognition data refers to the immediate recognition result of the intelligent speech control system for the user's speech input within the current specified time period, including the speech commands issued by the user and the converted text of the speech.

[0087] In this embodiment, the historical speech recognition sub-data refers to the data of each historical operation instruction stored in the system within a certain historical specified time period in the past. Each historical operation instruction includes: speech configuration object: the specific operation object corresponding to the instruction (such as a device, function, etc.); operation label: marking the execution result of the instruction (such as success, unrecognized speech configuration object, missing speech configuration object, etc.); operation action: the specific action executed by the speech instruction (such as "open", "close", etc.).

[0088] In this embodiment, the historical speech recognition data is formed by integrating the historical speech recognition sub-data within multiple time periods to form a complete historical record.

[0089] Beneficial effects of the above technical solution: Obtaining the real-time speech recognition data and the historical speech recognition data can provide a data basis for determining the corrected speech configuration data, provide personalized speech services, and improve the user interaction experience.

[0090] Embodiment 4:

[0091] The embodiment of the present invention provides an optimization method for improving the speech intent recognition result of a page. Based on the system configuration data, the preset speech configuration data, and the historical speech recognition data, determine the object correction data, the speech operation data, and the correction vector data, and determine the corrected speech configuration data of the intelligent speech control system, including:

[0092] Extract all page configuration objects in the system configuration data to determine the page configuration object set;

[0093] Extract the speech configuration objects of all historical operation instructions in the historical speech recognition sub-data of all historical specified time periods in the historical speech recognition data to determine the speech configuration object set;

[0094] Determine that the number of page configuration objects in the page configuration object set is the number of cluster families K for cluster analysis. Randomly select K speech configuration objects in the speech configuration object set as the initial cluster centers, and perform cluster analysis on the speech configuration object set to determine K cluster families, the cluster center configuration object of each cluster family, and the sub-set of speech configuration objects of each cluster family;

[0095] Extract all operation tags of all voice configuration objects in the voice configuration object subset of each cluster in the historical voice recognition sub-data within each specified time period based on the historical voice recognition data, and determine the operation tag set of the cluster center configuration object of each cluster;

[0096] Based on the system configuration data, historical voice recognition data, and the operation tag set of the cluster center configuration object of each cluster, determine the cluster center configuration object matrix of the cluster center configuration object of each cluster within each specified time period, and determine the correction success vector of the cluster center configuration object of each cluster;

[0097] Determine the correction mapping set, operation tag set, and correction success vector of each page configuration object;

[0098] Based on the correction mapping set of all page configuration objects in the page configuration object set, determine the object correction data. Based on the operation tag set of all page configuration objects in the page configuration object set, determine the voice operation data. At the same time, based on the correction success vector of all page configuration objects in the page configuration object set, determine the correction vector data;

[0099] Based on the object correction data, voice operation data, and correction vector data, determine the corrected voice configuration data of the intelligent voice control system.

[0100] In this embodiment, all page configuration objects are extracted from the system configuration data. Each page configuration object includes information on operation items such as devices and functions.

[0101] In this embodiment, the voice configuration object of each historical operation instruction is extracted from the historical voice recognition data, and the voice configuration object reflects the specific operation items of the voice instruction.

[0102] In this embodiment, the number of configuration objects in the page configuration object is used as the number of clusters (K) for clustering analysis. K voice configuration objects are randomly selected as the initial cluster centers, and then clustering analysis is performed on the voice configuration objects to determine K clusters, the cluster center configuration object of each cluster, and the subset of all voice configuration objects of this cluster.

[0103] In this embodiment, based on the historical voice recognition data, the operation tags of all voice configuration objects in each cluster are extracted. These tags reflect the execution results of different historical operation instructions.

[0104] In this embodiment, based on the system configuration data and historical voice recognition data, the operation tags of the cluster center configuration object of each cluster are analyzed to generate a cluster center configuration object matrix.

[0105] In this embodiment, by integrating the object correction data, voice operation data, and correction vector data, the corrected voice configuration data is finally generated.

[0106] Beneficial effects of the above technical solution: Based on the system configuration data, preset voice configuration data, and historical speech recognition data, determine the object correction data, voice operation data, and correction vector data, and determine the corrected voice configuration data of the intelligent voice control system, which can improve the system's recognition ability for complex voice commands and enhance the speech recognition accuracy and personalized experience.

[0107] Embodiment 5:

[0108] The embodiment of the present invention provides an optimization method for improving the speech intent recognition result of a page, which determines the centroid configuration object matrix of the centroid configuration objects of each cluster in each specified time period, including:

[0109] Based on the system configuration data, historical speech recognition data, and the operation label set of the centroid configuration objects of each cluster, determine the centroid configuration object matrix of the centroid configuration objects of each cluster in each specified time period;

[0110] ;

[0111] ;

[0112] where, represents the centroid configuration object matrix of the centroid configuration object of the a-th cluster in the t-th historical specified time period, represents that the operation label in the operation label set of the centroid configuration object of the a-th cluster in the t-th historical specified time period is operation success, represents the number of operation labels in the operation label set of the centroid configuration object of the a-th cluster, respectively represent the i-th operation label and the aN2-th operation label in the operation label set of the centroid configuration object of the a-th cluster in the t-th historical specified time period, represents the operation success rate of the operation label in the operation label set of the centroid configuration object of the a-th cluster in the t-th historical specified time period, represents the operation failure rate of the i-th operation label and the aN2-th operation label in the operation label set of the centroid configuration object of the a-th cluster in the t-th historical specified time period, represents the operation label of the L-th historical operation instruction of the j-th voice configuration object in the voice configuration object subset of the a-th cluster in the t-th historical specified time period, a represents the number of all historical operation instructions based on the j-th voice configuration object in the voice configuration object subset of the a-th cluster, represents the number of voice configuration objects in the voice configuration object subset of the a-th cluster in the t-th historical specified time period.

[0113] In this embodiment, The operation label of the Lth historical operation instruction of the jth voice configuration object in the subset of voice configuration objects of the ath cluster in the tth historical specified time period And the ith operation label in the set of operation labels of the cluster center configuration object of the ath cluster in the tth historical specified time period When they are the same, the value is 1.

[0114] Beneficial effects of the above technical solution: By determining the cluster center configuration object matrix of the cluster center configuration object of each cluster in each specified time period, it can provide a data basis for determining the correction success vector of the cluster center configuration objects of all clusters and correcting the vector data, thereby improving the success rate of real-time speech recognition data recognition.

[0115] Embodiment 6:

[0116] The embodiment of the present invention provides an optimization method for improving the page voice intent recognition result, which determines the correction success vector of the cluster center configuration object of each cluster, including: determining the correction success vector of the cluster center configuration object of each cluster based on the cluster center configuration object matrix of the cluster center configuration object of each cluster in all specified time periods;

[0117] ;

[0118] ;

[0119] wherein, represents the correction success vector of the cluster center configuration object of the ath cluster, respectively represent the correction success rates of the 2nd operation label, the ith operation label, and the aN2th operation label in the set of operation labels of the cluster center configuration object of the ath cluster, represents the period decay factor, represents the correction factor, and Nu represents the number of historical specified time periods.

[0120] In this embodiment, represents the operation success frequency of the operation label in the set of operation labels of the cluster center configuration object of the ath cluster, that is, the weight of the operation success rate of the operation label in the set of operation labels of the cluster center configuration object of the ath cluster.

[0121] In this embodiment, represents the operation failure frequency of the ith operation label and the aN2th operation label in the set of operation labels of the cluster center configuration object of the ath cluster, that is, the weight of the operation failure rate of the ith operation label and the aN2th operation label in the set of operation labels of the cluster center configuration object of the ath cluster.

[0122] Beneficial effects of the above technical solution: Determining the correction success vector of the cluster center configuration object of each cluster can provide a data basis for determining the correction vector data and improve the success rate of real-time speech recognition data recognition.

[0123] Example 7:

[0124] The embodiment of the present invention provides an optimization method for improving the speech intent recognition result of a page, which determines the correction mapping set, operation label set and correction success vector of each page configuration object, including:

[0125] Based on the object mapping set of each page configuration object in the preset speech configuration data, all page configuration objects in the page configuration object set, the cluster center configuration objects of all clusters, and the speech configuration object subset, determine the correction mapping set of each page configuration object;

[0126] ;;

[0127] ;

[0128] Wherein, represents the similarity value between the x-th page configuration object in the page configuration object set and the a-th cluster; represents the correction mapping set of the x-th page configuration object in the page configuration object set; represents the cluster center configuration object of the a-th cluster; represents the j-th speech configuration object in the speech configuration object subset of the a-th cluster; represents the Euclidean norm of the x-th page configuration object in the page configuration object set; represents the Euclidean norm of the cluster center configuration object of the a-th cluster; represents the Euclidean norm of the j-th speech configuration object in the speech configuration object subset of the a-th cluster; represents the speech configuration object subset of the a-th cluster; represents the speech configuration object subset with the largest similarity value to the x-th page configuration object in the page configuration object set; represents the object mapping set of the x-th page configuration object in the page configuration object set;

[0129] Based on the operation label set of the cluster corresponding to the speech configuration object subset with the largest similarity value to each page configuration object in the page configuration object set, determine the operation label set of each page configuration object in the page configuration object set;

[0130] Determine the correction success vector of each page configuration object in the page configuration object set based on the correction success vector of the cluster family corresponding to the sub - set of voice configuration objects with the largest similarity value to each page configuration object in the page configuration object set.

[0131] In this embodiment, represents the weight of the j - th voice configuration object in the sub - set of voice configuration objects of the a - th cluster family at the t - th historical specified time period.

[0132] In this embodiment, represents the weight of the sub - set of voice configuration objects of the a - th cluster family for all historical specified time periods.

[0133] In this embodiment, the value range of a is from 1 to K.

[0134] In this embodiment, select the sub - set of voice configuration objects most similar to each page configuration object and correspond them to the cluster family. Determine the operation label of this cluster family as the operation label of the corresponding page configuration object.

[0135] In this embodiment, for each page configuration object, based on the sub - set of voice configuration objects most similar to the page configuration object, the system extracts the correction success vector of the corresponding cluster family.

[0136] Beneficial effects of the above - mentioned technical solution: Determining the correction mapping set, operation label set, and correction success vector of each page configuration object can provide a data basis for determining the correction vector data, dynamically adjust and optimize the speech recognition and operation response of the intelligent voice control system, improve the accuracy and efficiency of user interaction, and enhance the personalized experience.

[0137] Embodiment 8:

[0138] The embodiment of the present invention provides an optimization method for improving the page speech intent recognition result. Based on the real - time speech recognition data and the corrected speech configuration data of the intelligent voice control system, determine the real - time intent data, including:

[0139] Optimize the intent model based on the correction vector data in the corrected speech configuration data;

[0140] If the parsing is successful after parsing the real - time speech recognition data based on the optimized intent model, determine the predicted intent data, match the predicted intent data with the correction mapping set of each page configuration object in the corrected speech configuration data, and determine the real - time intent data based on the matching result;

[0141] If the parsing is not successful, perform a hard match on the real - time speech recognition data to determine the real - time intent data.

[0142] In this embodiment, the intent model is the core model used by the intelligent voice control system to parse user voice commands. By modifying the correction vector data in the voice configuration data to optimize the intent model, the system can improve the accuracy of the model by the correction success vectors of all page configuration objects after correction, and the intent model can be optimized through a correction algorithm.

[0143] In this embodiment, after the intent model is optimized, the real-time speech recognition data will be parsed. If the parsing is successful, the predicted intent data will be determined.

[0144] In this embodiment, if the intent parsing fails, the system will adopt a hard matching method, that is, directly match the real-time speech recognition data with the page configuration objects, without relying on the optimized intent model, to ensure that a response can still be given when the intent parsing fails.

[0145] Beneficial effects of the above technical solution: Based on the real-time speech recognition data and the corrected voice configuration data of the intelligent voice control system, the real-time intent data is determined, which can ensure the efficiency and reliability of speech recognition, improve the adaptability and robustness of the intelligent voice control system, and enhance the user experience and accuracy of voice control.

[0146] Embodiment 9:

[0147] The embodiment of the present invention provides an optimization method for improving the page voice intent recognition result. Determining the real-time intent data based on the matching result includes:

[0148] If there are page configuration objects in the corrected voice configuration data that highly match the predicted intent data, based on the correction mapping set and operation label set of the page configuration objects that highly match the predicted intent data, determine the real-time intent data of the real-time voice data;

[0149] If there are no page configuration objects in the corrected voice configuration data that highly match the predicted intent data, re-correct the correction mapping set and operation label set of the page configuration object with the highest match to the predicted intent data.

[0150] In this embodiment, when there are page configuration objects in the corrected voice configuration data that highly match the predicted intent data, the real-time intent data of the real-time voice data is determined according to the correction mapping set and operation label set of these page configuration objects. This means that the system will utilize the matching relationship between page configuration objects and combine the corresponding operation labels to confirm the specific intent of the user.

[0151] In this embodiment, if a page configuration object that highly matches the predicted intent data cannot be found in the corrected voice configuration data, the system will select the page configuration object with the highest matching degree with the predicted intent data, and correct the corrected mapping set and operation label set of this page configuration object again to further optimize the matching result and ensure that the system can recognize and respond to the user's intent.

[0152] Beneficial effects of the above technical solution: Determining the real-time intent data based on the matching result can ensure that the intelligent voice control system can accurately understand and respond to complex user instructions, enhance the intelligence and adaptability of the system, improve the accuracy of voice control, and enhance the user experience.

[0153] Embodiment 10:

[0154] An embodiment of the present invention provides an optimization method for improving the page voice intent recognition result. Based on the real-time intent data, real-time operation data is determined, and based on the real-time operation data, a real-time response is made to the voice control system, including:

[0155] Analyze the real-time intent data to determine the real-time operation data, where the real-time operation data includes at least one or more real-time operation instructions;

[0156] The voice control system executes each real-time operation instruction in the real-time operation data and determines the operation label of each real-time operation instruction;

[0157] Based on all the real-time operation instructions in the real-time operation data and the operation labels of each real-time operation instruction, determine the current voice recognition sub-data for the current specified time period.

[0158] In this embodiment, according to the real-time voice recognition data, the system analyzes and determines the user's intent and converts it into operation data (such as opening an application, adjusting the volume, etc.). The real-time operation data usually contains multiple instructions, and each instruction represents a specific operation, such as "playing music" or "turning off the light".

[0159] In this embodiment, the intelligent voice control system sequentially executes each real-time operation instruction according to the determined real-time operation data. After the execution of each real-time operation instruction, an operation label will be determined to identify the execution result of this operation.

[0160] In this embodiment, all the real-time operation instructions and their labels are summarized to generate the voice recognition sub-data for the current specified time period. Beneficial effects of the above technical solution: Determining the real-time operation data based on the real-time intent data and making a real-time response to the voice control system based on the real-time operation data can improve the real-time response ability of the intelligent voice control system, accurately understand and respond to complex user instructions, and provide a personalized and efficient voice control experience.

[0161] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.

[0162] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0163] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An optimization method for improving the recognition result of page speech intent, characterized in that: include: Step 1: Determine the system configuration data and preset voice configuration data of the intelligent voice control system, and obtain real-time voice recognition data and historical voice recognition data; Step 2: Based on the system configuration data, the preset voice configuration data and the historical voice recognition data, determine the object correction data, the voice operation data and the correction vector data, and determine the corrected voice configuration data of the intelligent voice control system; Step 3: Determine real-time intent data based on the real-time speech recognition data and the modified speech configuration data of the intelligent speech control system; Step 4: determining real-time operation data based on the real-time intention data, and responding to the voice control system in real time based on the real-time operation data; Wherein, based on the system configuration data, the preset voice configuration data and the historical voice recognition data, the object correction data, the voice operation data and the correction vector data are determined, and the corrected voice configuration data of the intelligent voice control system is determined, including: Extract all page configuration objects in the system configuration data to determine a page configuration object set; Extracting the voice configuration objects of all historical operation instructions in all historical voice recognition sub-data of the historical specified time period in the historical voice recognition data to determine a voice configuration object set; Determine the number of page configuration objects in the page configuration object set as the number of clusters K for cluster analysis, randomly select K voice configuration objects in the voice configuration object set as initial cluster centers, perform cluster analysis on the voice configuration object set, and determine K clusters, cluster center configuration objects of each cluster, and a subset of voice configuration objects of each cluster; Based on the historical speech recognition data, all operation labels of all voice configuration objects in the voice configuration object subset of each cluster family in the historical speech recognition sub-data within each specified time period are extracted to determine the operation label set of the cluster core configuration object of each cluster family.

2. The optimization method for improving the recognition result of page speech intent according to claim 1, characterized in that: Determine the system configuration data and preset voice configuration data of the intelligent voice control system, including: Obtaining page configuration data of each page in the intelligent voice control system, wherein the page configuration data at least includes a plurality of page configuration objects, a page configuration object type of each page configuration object, an operation action of each page configuration object, and a plurality of operation action aliases of each operation action of each page configuration object; Determine system configuration data of the intelligent voice control system based on page configuration data of all pages in the intelligent voice control system; Determine a plurality of voice mapping instructions based on system configuration data of the intelligent voice control system, and determine an object mapping set for each page configuration object based on all the voice mapping instructions; The preset voice configuration data is determined based on an object mapping set of all page configuration objects.

3. The optimization method for improving the recognition result of page speech intent according to claim 2, characterized in that: Get real-time and historical speech recognition data, including: Get real-time speech recognition data within the current specified time period; Acquire historical speech recognition sub-data within multiple historical specified time periods before the current specified time period, wherein the historical speech recognition sub-data includes multiple historical operation instructions, the speech configuration object of each historical operation instruction, the operation label of each historical operation instruction, and the operation action of each historical operation instruction; Based on all the historical speech recognition sub-data of the historical designated time period, the historical speech recognition data is determined.

4. The optimization method for improving the recognition result of page speech intent according to claim 3, characterized in that: Based on the system configuration data, the preset voice configuration data and the historical voice recognition data, the object correction data, the voice operation data and the correction vector data are determined, and the corrected voice configuration data of the intelligent voice control system is determined, further comprising: Determine a cluster center configuration object matrix of the cluster center configuration objects of each cluster family within each specified time period based on the system configuration data, the historical speech recognition data and the operation tag set of the cluster center configuration objects of each cluster family, and determine a correction success vector of the cluster center configuration objects of each cluster family; Determine a correction mapping set, an operation label set, and a correction success vector for each page configuration object; Determine object correction data based on a correction mapping set of all page configuration objects in the page configuration object set, determine voice operation data based on an operation tag set of all page configuration objects in the page configuration object set, and determine correction vector data based on correction success vectors of all page configuration objects in the page configuration object set; Corrected voice configuration data of the intelligent voice control system is determined based on the object correction data, the voice operation data, and the correction vector data.

5. The optimization method for improving the recognition result of page speech intent according to claim 4, characterized in that: A cluster center configuration object matrix is ​​determined for each cluster family in each specified time period, including: Determine a cluster core configuration object matrix of the cluster core configuration objects of each cluster family within each specified time period based on the system configuration data, the historical speech recognition data and the operation tag set of the cluster core configuration objects of each cluster family; ; ; in, The cluster center configuration object matrix representing the cluster center configuration object of the a-th cluster family in the t-th historical specified time period, Indicates that the operation label in the operation label set of the cluster center configuration object of the a-th cluster group in the t-th historical specified time period is successful. represents the number of operation tags in the operation tag set of the cluster center configuration object of the a-th cluster family, They respectively represent the ith operation label and the aN2th operation label in the operation label set of the cluster center configuration object of the ath cluster family in the tth historical specified time period, represents the operation success rate of the operation tags in the operation tag set of the cluster center configuration object of the a-th cluster group in the t-th historical specified time period, represents the operation failure rate of the ith operation label and the aN2th operation label in the operation label set of the cluster center configuration object of the ath cluster group in the tth historical specified time period, The operation label of the Lth historical operation instruction of the jth voice configuration object in the ath cluster of the voice configuration object subset in the tth historical specified time period, a represents the number of all historical operation instructions based on the jth voice configuration object in the voice configuration object subset of the ath cluster, Indicates the number of voice configuration objects in the voice configuration object subset of the a-th cluster in the t-th historical specified time period.

6. The optimization method for improving the recognition result of page speech intent according to claim 5, characterized in that: Determine the revised success vector of the cluster center configuration object for each cluster family, including: Determine a revised success vector of the cluster center configuration object of each cluster family based on the cluster center configuration object matrix of the cluster center configuration objects of each cluster family within all specified time periods; ; ; in, represents the corrected success vector of the cluster center configuration object of the a-th cluster family, They represent the correction success rates of the second operation label, the ith operation label, and the aN2th operation label in the operation label set of the cluster center configuration object of the ath cluster family, respectively. represents the period attenuation factor, represents the correction factor, and Nu represents the number of historical specified time periods.

7. The optimization method for improving the recognition result of page speech intent according to claim 6, characterized in that: Determine the correction mapping set, operation label set, and correction success vector for each page configuration object, including: Determine a revised mapping set for each page configuration object based on an object mapping set for each page configuration object in the preset voice configuration data, all page configuration objects in the page configuration object set, cluster core configuration objects of all cluster families, and a voice configuration object sub-set; ; ; in, Represents the similarity value between the xth page configuration object and the ath cluster in the page configuration object set. represents the correction mapping set of the xth page configuration object in the page configuration object set, represents the cluster center configuration object of the a-th cluster family, represents the jth voice configuration object in the ath cluster group's voice configuration object subset, Represents the Euclidean norm of the xth page configuration object in the page configuration object collection, represents the Euclidean norm of the cluster center configuration object of the a-th cluster family, represents the Euclidean norm of the jth voice configuration object in the subset of voice configuration objects of the ath cluster, represents the subset of voice configuration objects of the a-th cluster, represents the voice configuration object sub-set with the largest similarity value to the x-th page configuration object in the page configuration object set, An object mapping collection representing the xth page configuration object in the page configuration object collection; Determine an operation tag set for each page configuration object in the page configuration object set based on an operation tag set of a cluster corresponding to a voice configuration object subset with a maximum similarity value to each page configuration object in the page configuration object set; Based on the correction success vector of the cluster group corresponding to the voice configuration object subset with the largest similarity value to each page configuration object in the page configuration object set, a correction success vector of each page configuration object in the page configuration object set is determined.

8. The optimization method for improving the recognition result of page speech intent according to claim 1, characterized in that: Based on the real-time speech recognition data and the modified speech configuration data of the intelligent speech control system, real-time intent data is determined, including: The intent model is optimized based on the correction vector data in the corrected voice configuration data, and the real-time voice recognition data is analyzed based on the optimized intent model. If the analysis is successful, the predicted intent data is determined, and the predicted intent data is matched with the correction mapping set of each page configuration object in the corrected voice configuration data, and the real-time intent data is determined based on the matching result; If the analysis is unsuccessful, a hard match is performed on the real-time speech recognition data to determine the real-time intent data.

9. The optimization method for improving the recognition result of page speech intent according to claim 8, characterized in that: Determine real-time intent data based on matching results, including: If there is a page configuration object in the modified voice configuration data that highly matches the predicted intent data, determine the real-time intent data of the real-time voice data based on the modified mapping set and the operation label set of the page configuration object that highly matches the predicted intent data; If there is no page configuration object in the revised voice configuration data that is highly matched with the predicted intent data, the revised mapping set and operation label set of the page configuration object that is most matched with the predicted intent data are revised again.

10. The optimization method for improving the recognition result of page speech intent according to claim 1, characterized in that: Determine real-time operation data based on the real-time intent data, and respond to the voice control system in real time based on the real-time operation data, including: Parsing the real-time intention data to determine real-time operation data, wherein the real-time operation data includes at least one or more real-time operation instructions; The voice control system executes each real-time operation instruction in the real-time operation data and determines an operation label of each real-time operation instruction; Based on all the real-time operation instructions in the real-time operation data and the operation label of each real-time operation instruction, current speech recognition sub-data of a current specified time period is determined.

Citation Information

Patent Citations

  • Voice interaction method, device and system

    CN111383631A