Intelligent recognition control method based on multi-voice assistant

By monitoring user voice data in real time and selecting appropriate voice processing strategies based on the number of users, the problem of voice recognition chaos in multi-user environments is solved, and intelligent recognition control with high accuracy and stability is achieved.

CN119580734BActive Publication Date: 2025-05-13北京基智科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510125521.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-27
Publication Date
2025-05-13
Estimated Expiration
2045-01-27

AI Technical Summary

Technical Problem

The existing intelligent recognition and control system is difficult to accurately identify each user's voice command in a multi-user environment, and it is prone to problems of recognition disorder and control errors.

Method used

By monitoring the received user voice data in real time, the voice processing strategy of multi-voice assistants is determined based on the number of users, including the first voice processing strategy (single user) and the second voice processing strategy (multi-user), and audio quality grading and voice assistant allocation strategies are adopted to ensure accurate identification and processing.

Benefits of technology

It realizes accurate recognition and processing of voice commands in a multi-user environment, improves the system's recognition accuracy and stability, ensures fast response and correct control, and avoids confusion and errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119580734B_ABST
    Figure CN119580734B_ABST
Patent Text Reader

Abstract

The present invention proposes an intelligent recognition control method based on multiple voice assistants. The intelligent recognition control method based on multiple voice assistants includes: real-time monitoring of the voice data sent by the received users, and determining the voice processing strategy of the multiple voice assistants according to the number of users who sent the voice data; wherein the voice processing strategy includes a first voice processing strategy and a second voice processing strategy; and the first voice processing strategy refers to a voice processing strategy in which multiple voice assistants collaboratively process the voice data of one user; the second voice processing strategy refers to a voice processing strategy in which multiple voice assistants collaboratively process the voice data of multiple users; performing voice recognition processing on the voice data sent by the user according to the voice processing strategy, and obtaining the voice instructions corresponding to the voice data; and intelligently controlling the target system according to the voice instructions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention proposes an intelligent recognition control method based on multiple voice assistants, belonging to the technical field of voice recognition. Background Art

[0002] With the rapid development of artificial intelligence technology, voice assistants have become an important part of smart home, smart office, smart car system and other fields. These voice assistants control various smart devices by receiving users' voice commands, which greatly improves the user experience and the intelligence level of the system. However, in the existing intelligent recognition control system, when faced with the scenario of multiple users issuing voice commands at the same time, the system often has problems of recognition confusion and control errors.

[0003] When designing traditional single voice assistant systems, the interaction between a single user and the system is mainly considered. For multiple voice signals received at the same time, the system often lacks an effective distinction and processing mechanism. As a result, in a multi-user environment, the system may not be able to accurately recognize the instructions of each user, and may even misidentify and misoperate. For example, in a smart home scenario, if family members send different control instructions to the voice assistant at the same time, the system may confuse these instructions, resulting in incorrect device control. In order to solve this problem, some existing technologies attempt to improve the recognition ability of the system by increasing the complexity of the voice recognition module and the accuracy of the algorithm. However, although these methods have improved the recognition accuracy of single-user instructions to a certain extent, it is still difficult to effectively distinguish and correctly process the instructions of each user when multiple users speak at the same time. In addition, these methods often require more computing resources and time, reducing the real-time performance and response speed of the system.

[0004] Therefore, there is an urgent need for an intelligent recognition and control method that can effectively deal with voice commands issued by multiple users at the same time. This method needs to be able to monitor the received user voice data in real time and intelligently select the voice processing strategy according to the number of users to ensure that the system can accurately recognize and process the commands of each user in a multi-user environment. At the same time, this method also needs to be efficient and real-time to meet the requirements of modern intelligent systems for response speed and stability. Summary of the invention

[0005] The present invention provides an intelligent recognition control method based on a multi-voice assistant to solve the above technical problems existing in the prior art. The technical solutions adopted are as follows:

[0006] An intelligent recognition control method based on a multi-voice assistant, the intelligent recognition control method based on a multi-voice assistant comprising:

[0007] Real-time monitoring of the received voice data sent by the user, and determining the voice processing strategy of the multi-voice assistant according to the number of users sending the voice data; wherein the voice processing strategy includes a first voice processing strategy and a second voice processing strategy; and the first voice processing strategy refers to a voice processing strategy for the multi-voice assistants to collaboratively process the voice data of one user; the second voice processing strategy refers to a voice processing strategy for the multi-voice assistants to collaboratively process the voice data of multiple users;

[0008] Performing voice recognition processing on the voice data sent by the user according to the voice processing strategy to obtain voice instructions corresponding to the voice data;

[0009] The target system is intelligently controlled according to the voice command.

[0010] Furthermore, the received voice data sent by the user is monitored in real time, and the voice processing strategy of the multi-voice assistant is determined according to the number of users sending the voice data, including:

[0011] Real-time monitoring of the voice data received from users;

[0012] Determining the number of users currently sending voice data based on the received voice data sent by the user;

[0013] When the number of users is one user, the first voice processing strategy is retrieved as the voice processing strategy;

[0014] When the number of users is multiple users, the second voice processing strategy is called as the voice processing strategy.

[0015] Further, performing voice recognition processing on the voice data sent by the user according to the voice processing strategy to obtain a voice instruction corresponding to the voice data includes:

[0016] Performing voice recognition processing on the voice data sent by the user according to the first voice processing strategy to obtain a voice instruction corresponding to the voice data;

[0017] or

[0018] Perform voice recognition processing on the voice data sent by the user according to the second voice processing strategy to obtain voice instructions corresponding to the voice data.

[0019] Furthermore, the speech recognition processing process of the first speech processing strategy includes:

[0020] When the number of users currently sending voice data is one user, dividing the voice data into voice segments to obtain a plurality of voice segment data;

[0021] Performing audio quality grading on the current multiple voice segment data to obtain audio quality levels of the multiple voice segment data;

[0022] According to the speech recognition quality of the current multiple voice assistants and the audio quality levels of the multiple voice clips, a target voice assistant corresponding to the voice clip of each audio quality level is obtained;

[0023] The voice segments of each audio quality level are assigned to their corresponding target voice assistants for voice recognition processing to obtain voice commands corresponding to the voice data.

[0024] Further, audio quality grading is performed on the current multiple voice segment data to obtain audio quality levels of the multiple voice segments, including:

[0025] Extracting audio parameters of each speech segment data, wherein the audio parameters include a total harmonic distortion value, a spectrum dynamic range, a spectrum width, a distance between a spectrum center and a frequency center, and a ratio between a byte audio intensity and a background audio intensity;

[0026] Obtaining an audio quality adjustment coefficient using the spectrum dynamic range, the spectrum width, and the distance between the spectrum center and the frequency center;

[0027] The audio quality adjustment coefficient is obtained by the following formula:

[0028]

[0029] Where U represents the audio quality adjustment coefficient; S r Indicates the spectrum dynamic range; S w Indicates the spectrum width; S d Indicates the distance between the spectrum center and the frequency center; λ 01 and λ 02 Respectively represent the weight values ​​corresponding to the spectrum dynamic range and spectrum width;

[0030] The audio quality evaluation coefficient is obtained by combining the audio quality adjustment coefficient with the total harmonic distortion value and the ratio between the byte audio intensity and the background audio intensity;

[0031] The audio quality evaluation coefficient is obtained by the following formula:

[0032]

[0033] Where R represents the audio quality evaluation coefficient; T h Indicates the total harmonic distortion value; B indicates the ratio between the byte audio intensity and the background audio intensity; U indicates the audio quality adjustment coefficient;

[0034] Comparing the audio quality evaluation coefficient with a preset first coefficient threshold and a second coefficient threshold;

[0035] When the audio quality evaluation coefficient exceeds a preset first coefficient threshold, determining the voice segment data whose audio quality evaluation coefficient exceeds the preset first coefficient threshold as a high-quality voice segment;

[0036] When the audio quality evaluation coefficient does not exceed the preset first coefficient threshold, but exceeds the second coefficient threshold, the voice segment data whose audio quality evaluation coefficient does not exceed the preset first coefficient threshold, but exceeds the second coefficient threshold is determined as a medium-quality voice segment;

[0037] When the audio quality evaluation coefficient does not exceed the preset second coefficient threshold, the voice segment data whose audio quality evaluation coefficient does not exceed the preset second coefficient threshold is determined as a low-quality voice segment.

[0038] Furthermore, according to the speech recognition quality of the current multiple voice assistants and the audio quality levels of the multiple voice segments, a target voice assistant corresponding to the voice segment of each audio quality level is obtained, including:

[0039] Extract the recognition accuracy and recognition processing time of each voice assistant;

[0040] Obtaining a recognition evaluation coefficient according to the recognition accuracy and recognition processing time corresponding to each voice assistant;

[0041] The identification evaluation coefficient is obtained by the following formula:

[0042]

[0043] Where J represents the recognition evaluation coefficient; n represents the number of times each voice assistant completes recognition; T i represents the recognition processing time corresponding to the i-th speech recognition of each voice assistant; P zi represents the recognition accuracy rate of each voice assistant’s i-th speech recognition; C i represents the amount of voice data corresponding to the i-th voice recognition of each voice assistant; C d Indicates the preset unit voice data volume; T c Indicates the theoretical recognition processing time per unit of voice data for each voice assistant;

[0044] The voice assistant corresponding to the maximum value of the recognition evaluation coefficient is used as the target voice assistant for the low-quality voice segment;

[0045] The voice assistant whose recognition evaluation coefficient is second only to the maximum recognition evaluation coefficient is used as the target voice assistant for the medium-quality voice segment;

[0046] The voice assistant corresponding to the recognition evaluation coefficient of the target voice assistant whose recognition evaluation coefficient is second only to the recognition evaluation coefficient of the medium-quality voice segment is used as the target voice assistant for the high-quality voice segment.

[0047] Furthermore, the speech recognition processing process of the second speech processing strategy includes:

[0048] When the number of users currently sending voice data is multiple users, and the number of multiple users is less than or equal to the number of voice assistants, the users are assigned voice assistants according to the audio intensity ratio, and the assigned voice assistants are used to perform voice recognition processing on the voice data of the corresponding users;

[0049] When the number of users currently sending voice data is multiple users, and the number of multiple users is greater than the number of voice assistants, multiple voice assistants are controlled to collaboratively perform voice recognition processing through user tags.

[0050] Furthermore, assigning a voice assistant to a user according to the audio intensity ratio includes:

[0051] When the number of users currently sending voice data is multiple users, the voice data of the multiple users are split to obtain voice data corresponding to each user;

[0052] Extracting the audio intensity corresponding to the voice data corresponding to each user;

[0053] Compare the audio intensity corresponding to the voice data corresponding to each user with the overall audio intensity corresponding to multiple users to obtain an audio intensity ratio;

[0054] Obtaining a first audio recognition difficulty coefficient by using the audio intensity ratio in combination with a byte frequency corresponding to the voice data of each user;

[0055] The first audio recognition difficulty coefficient is obtained by the following formula:

[0056]

[0057] Among them, J 01 Indicates the difficulty coefficient of the first audio recognition; P d01 represents the audio intensity ratio corresponding to each user; f represents the byte frequency corresponding to the voice data of each user; f c Indicates the preset byte frequency reference value;

[0058] The voice assistant corresponding to the maximum value of the recognition evaluation coefficient is used as the target voice assistant for the user corresponding to the maximum value of the first audio recognition difficulty coefficient;

[0059] According to the allocation strategy of the first audio recognition difficulty coefficient from high to low corresponding to the recognition evaluation coefficient from low to high, other users except the users with the maximum first audio recognition difficulty coefficient are matched with the remaining voice assistants in turn.

[0060] Furthermore, multiple voice assistants are controlled to collaboratively perform voice recognition processing through user tags, including:

[0061] Perform unique identification processing on the voice data of each user, so that the voice data of each user has a unique identification that is uniquely associated with the user;

[0062] Dividing the voice data of all users to obtain a plurality of voice segment data, and assigning a unique identifier consistent with the voice data to the voice segment data;

[0063] Using the audio intensity ratio corresponding to each voice segment data and the amplitude of the byte audio intensity change in each voice segment; obtaining the second audio recognition difficulty coefficient corresponding to each voice segment data;

[0064] The second audio recognition difficulty coefficient is obtained by the following formula:

[0065]

[0066] Among them, J 02 P represents the difficulty coefficient of the second audio recognition; d02 represents the audio intensity ratio corresponding to each speech segment data; k represents the number of bytes contained in each speech segment data; Q i Indicates the audio intensity corresponding to the i-th byte; Q i+1 Indicates the audio intensity corresponding to the i+1th byte; Q b Indicates the standard deviation of the audio intensity corresponding to k bytes;

[0067] Comparing the second audio recognition difficulty coefficient with a preset recognition difficulty coefficient threshold;

[0068] Assigning the speech segment data whose second audio recognition difficulty coefficient exceeds a preset recognition difficulty coefficient threshold to the voice assistant corresponding to the maximum recognition evaluation coefficient for speech recognition processing, obtaining a speech recognition result corresponding to each speech segment data, and assigning a unique identifier consistent with the speech segment data to the speech recognition result;

[0069] Allocate the voice segment data whose second audio recognition difficulty coefficient does not exceed the preset recognition difficulty coefficient threshold to the remaining voice assistants for voice recognition processing according to the principle of approximate uniform distribution, obtain the voice recognition result corresponding to each voice segment data, and assign a unique identifier consistent with the voice segment data to the voice recognition result;

[0070] The voice results are filtered according to the corresponding identifier of each voice segment data, and the voice results with the same unique identifier are integrated to obtain the recognition result corresponding to the voice data uniquely associated with the user.

[0071] Furthermore, intelligently controlling the target system according to the voice command includes:

[0072] When the number of the user is one, the target system is intelligently controlled according to the speech result recognized by the voice assistant;

[0073] When the number of users is multiple, it is determined whether the voice commands corresponding to the multiple users are for the same controlled target parameters; when there are control commands for the same controlled target parameters among the voice commands corresponding to the multiple users, the voice commands of users with priority permissions are filtered according to the user permission priority to perform intelligent control of the target system.

[0074] Beneficial effects of the present invention:

[0075] The intelligent recognition control method based on multi-voice assistant proposed in the present invention realizes accurate recognition and processing of voice commands in a multi-user environment by real-time monitoring of received user voice data and determining the voice processing strategy of the multi-voice assistant according to the number of users (including the strategy of co-processing the voice data of a single user and the strategy of co-processing the voice data of multiple users). This method not only improves the recognition accuracy and stability of the system, but also ensures that when multiple users speak at the same time, the system can correctly and quickly respond to the commands of each user, avoiding control confusion and errors. BRIEF DESCRIPTION OF THE DRAWINGS

[0076] Figure 1 The present invention is a flowchart of the method. DETAILED DESCRIPTION

[0077] The preferred embodiments of the present invention are described below in conjunction with the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.

[0078] The embodiment of the present invention proposes an intelligent recognition control method based on a multi-voice assistant, such as Figure 1 As shown, the intelligent recognition control method based on multi-voice assistant includes:

[0079] S1. Real-time monitoring of the received voice data sent by the user, and determining the voice processing strategy of the multi-voice assistant according to the number of users sending the voice data; wherein the voice processing strategy includes a first voice processing strategy and a second voice processing strategy; and the first voice processing strategy refers to a voice processing strategy for the multi-voice assistants to collaboratively process the voice data of one user; the second voice processing strategy refers to a voice processing strategy for the multi-voice assistants to collaboratively process the voice data of multiple users;

[0080] The voice data sent by the user is subjected to voice recognition processing according to the voice processing strategy to obtain a voice instruction corresponding to the voice data, including:

[0081] Performing voice recognition processing on the voice data sent by the user according to the first voice processing strategy to obtain a voice instruction corresponding to the voice data;

[0082] or

[0083] Perform voice recognition processing on the voice data sent by the user according to the second voice processing strategy to obtain voice instructions corresponding to the voice data.

[0084] S2. Performing voice recognition processing on the voice data sent by the user according to the voice processing strategy to obtain a voice instruction corresponding to the voice data;

[0085] S3. Intelligently control the target system according to the voice command.

[0086] The working principle of the above technical solution is as follows: the system first monitors the voice data sent by the received user in real time. According to the monitored voice data, the system determines the number of users who sent the voice data. According to the number of users, the system selects the corresponding voice processing strategy. If the number of users is one, the first voice processing strategy (the strategy of multiple voice assistants cooperating to process the voice data of one user) is selected; if the number of users is multiple, the second voice processing strategy (the strategy of multiple voice assistants cooperating to process the voice data of multiple users) is selected. According to the selected voice processing strategy, the system performs voice recognition processing on the received voice data. If the first voice processing strategy is selected, the system will use the collaborative working ability of multiple voice assistants to more accurately and quickly recognize the voice data of one user, thereby improving the accuracy and efficiency of recognition. If the second voice processing strategy is selected, the system will use more complex algorithms and models to distinguish and recognize the voice data of multiple users to ensure that the instructions of each user can be accurately understood and processed. After the voice recognition processing, the system obtains the voice instructions corresponding to the voice data. The system intelligently controls the target system according to the recognized voice instructions. The target system can be a smart home device, a smart office device, a smart car system, etc. The system performs corresponding operations according to the content of the instructions, such as switching devices, adjusting parameters, executing commands, etc.

[0087] The effect of the above technical solution is: the intelligent recognition control method based on multi-voice assistant proposed in this embodiment realizes accurate recognition and processing of voice commands in a multi-user environment by real-time monitoring of received user voice data and determining the voice processing strategy of the multi-voice assistant according to the number of users (including the strategy of co-processing the voice data of a single user and the strategy of co-processing the voice data of multiple users). This method not only improves the recognition accuracy and stability of the system, but also ensures that when multiple users speak at the same time, the system can correctly and quickly respond to the instructions of each user, avoiding control confusion and errors.

[0088] On the other hand, by monitoring the number of users in real time and selecting the corresponding voice processing strategy, the system can more accurately identify the user's voice commands, avoiding recognition confusion and control errors in a multi-user environment. The collaborative work of multiple voice assistants can further improve the accuracy and efficiency of recognition. This method can cope with the scenario where multiple users issue voice commands at the same time, ensuring the stability and reliability of the system. By intelligently selecting the voice processing strategy, the system can reasonably allocate resources and avoid system crashes or response delays caused by insufficient resources. Users do not need to worry about command recognition errors or device misoperation in a multi-user environment, thereby improving the user experience. This method enables users to interact with the intelligent system more naturally and conveniently, improving the availability and ease of use of the system. This method is not only applicable to the field of smart homes, but can also be widely used in smart offices, smart car systems and other fields. With the continuous popularization of smart devices and the continuous development of voice technology, the application prospects of this method will be broader.

[0089] In summary, this technical solution achieves accurate recognition and processing of voice commands in a multi-user environment by monitoring the number of users in real time and selecting corresponding voice processing strategies, thereby improving the system's recognition accuracy, stability and user experience.

[0090] In one embodiment of the present invention, voice data sent by users is monitored in real time, and a voice processing strategy of a multi-voice assistant is determined according to the number of users who sent the voice data, including:

[0091] S101, real-time monitoring of received voice data sent by the user;

[0092] S102, determining the number of users currently sending voice data according to the received voice data sent by the user;

[0093] S103: when the number of users is one, calling the first voice processing strategy as the voice processing strategy;

[0094] S104: When the number of users is multiple, the second voice processing strategy is retrieved as the voice processing strategy.

[0095] The working principle of the above technical solution is as follows: the system continuously monitors and receives voice data from users. This is usually done through a microphone or other audio input device to ensure that the system can capture any voice commands issued by the user. The system analyzes the received voice data to determine the number of users who are currently issuing voice data. This can be achieved through a variety of technologies, such as voiceprint recognition, voice activity detection (VAD), directional microphone arrays, etc. Voiceprint recognition can distinguish the voice characteristics of different users, while VAD is used to detect the presence or absence of voice signals. Directional microphone arrays can determine the number of users based on the direction of the sound source. Based on the determined number of users, the system selects the corresponding voice processing strategy. If the number of users is one, the system calls the first voice processing strategy. This strategy usually focuses on optimizing the voice recognition accuracy of a single user, which may include using more complex voice recognition algorithms, enhancing the quality of voice signals, etc. If the number of users is multiple, the system calls the second voice processing strategy. This strategy needs to handle the separation, recognition and priority sorting of multiple voice signals to ensure that the instructions of each user can be accurately understood and the system can execute the instructions according to priority or time order.

[0096] The effect of the above technical solution is: by selecting an appropriate voice processing strategy according to the number of users, the system can more effectively process the voice commands of a single or multiple users, thereby improving the accuracy of recognition. The system can intelligently adjust its processing resources according to the number of users to avoid wasting resources when a single user is using it, while ensuring sufficient processing power when multiple users are using it. Users do not need to worry about command confusion or system response delays in a multi-user environment, thereby improving user experience satisfaction. This technical solution enables the system to adapt to different user environments, and can maintain efficiency and accuracy whether it is used by a single person or by multiple people at the same time. This method is not only suitable for home environments, but also for a variety of scenarios such as conference rooms, classrooms, and public places, which improves the applicability and market competitiveness of the system.

[0097] In summary, this technical solution achieves efficient speech recognition and processing in a multi-user environment by monitoring the voice data emitted by users in real time and selecting appropriate speech processing strategies according to the number of users, thereby improving the accuracy, flexibility and user experience of the system.

[0098] In one embodiment of the present invention, the speech recognition processing process of the first speech processing strategy includes:

[0099] When the number of users currently sending voice data is one user, dividing the voice data into voice segments to obtain a plurality of voice segment data;

[0100] Performing audio quality grading on the current multiple voice segment data to obtain audio quality levels of the multiple voice segment data;

[0101] According to the speech recognition quality of the current multiple voice assistants and the audio quality levels of the multiple voice clips, a target voice assistant corresponding to the voice clip of each audio quality level is obtained;

[0102] The voice segments of each audio quality level are assigned to their corresponding target voice assistants for voice recognition processing to obtain voice commands corresponding to the voice data.

[0103] The working principle of the above technical solution is as follows: when the system detects that only one user is sending voice data, the voice segmentation process is started. This process divides the continuous voice data into multiple shorter voice segment data, which makes it easier to process each segment separately in the future, improving processing efficiency and accuracy. The system will evaluate the audio quality of these voice segment data to obtain the audio quality level of each segment. The audio quality grading may be based on a variety of factors, such as noise level, signal strength, clarity, etc. Through grading, the system can identify which segments have better sound quality and which may be interfered with or of poor quality. Next, the system will select the most suitable target voice assistant for each segment based on the voice recognition quality of the multiple voice assistants currently available and the audio quality level of each voice segment. This process may involve a comprehensive evaluation of the recognition ability, processing speed, stability, etc. of the voice assistant to ensure that high-quality voice segments are assigned to voice assistants with stronger recognition capabilities, thereby improving the overall recognition accuracy. Finally, the system will assign voice segments of each audio quality level to their corresponding target voice assistants for voice recognition processing. The voice assistant will use its own algorithms and models to convert voice segments into text or commands to obtain the voice instructions issued by the user.

[0104] The effect of the above technical solution is: through audio quality grading and the selection of target voice assistants, the system can allocate high-quality voice clips to voice assistants with stronger recognition capabilities, thereby improving the overall recognition accuracy. At the same time, for clips with poor quality, the system can select voice assistants with stronger noise resistance for processing, further reducing the misrecognition rate. The solution can reasonably allocate resources according to the audio quality level of the voice clip and the recognition ability of the voice assistant. This not only avoids wasting high-quality voice clips on voice assistants with weaker recognition capabilities, but also ensures that each voice assistant can process the most suitable clips for itself, thereby improving overall processing efficiency. By improving recognition accuracy and optimizing resource allocation, the solution can provide users with more accurate and efficient voice recognition services. This can not only enhance users' trust and satisfaction with voice interaction, but also promote the application and development of voice technology in more fields.

[0105] In summary, the speech recognition processing process of the first speech processing strategy achieves efficient and accurate processing of speech data through detailed division, classification, selection and allocation steps, providing users with better speech recognition services.

[0106] In one embodiment of the present invention, audio quality grading is performed on the current plurality of voice segment data to obtain the audio quality levels of the plurality of voice segments, including:

[0107] Extracting audio parameters of each speech segment data, wherein the audio parameters include a total harmonic distortion value, a spectrum dynamic range, a spectrum width, a distance between a spectrum center and a frequency center, and a ratio between a byte audio intensity and a background audio intensity;

[0108] Obtaining an audio quality adjustment coefficient using the spectrum dynamic range, the spectrum width, and the distance between the spectrum center and the frequency center;

[0109] The audio quality adjustment coefficient is obtained by the following formula:

[0110]

[0111] Where U represents the audio quality adjustment coefficient; S r Indicates the spectrum dynamic range; S w Indicates the spectrum width; S d Indicates the distance between the spectrum center and the frequency center; λ 01 and λ 02 Respectively represent the weight values ​​corresponding to the spectrum dynamic range and spectrum width;

[0112] The audio quality evaluation coefficient is obtained by combining the audio quality adjustment coefficient with the total harmonic distortion value and the ratio between the byte audio intensity and the background audio intensity;

[0113] The audio quality evaluation coefficient is obtained by the following formula:

[0114]

[0115] Where R represents the audio quality evaluation coefficient; T h Indicates the total harmonic distortion value; B indicates the ratio between the byte audio intensity and the background audio intensity; U indicates the audio quality adjustment coefficient;

[0116] Comparing the audio quality evaluation coefficient with a preset first coefficient threshold and a second coefficient threshold;

[0117] When the audio quality evaluation coefficient exceeds a preset first coefficient threshold, determining the voice segment data whose audio quality evaluation coefficient exceeds the preset first coefficient threshold as a high-quality voice segment;

[0118] When the audio quality evaluation coefficient does not exceed the preset first coefficient threshold, but exceeds the second coefficient threshold, the voice segment data whose audio quality evaluation coefficient does not exceed the preset first coefficient threshold, but exceeds the second coefficient threshold is determined as a medium-quality voice segment;

[0119] When the audio quality evaluation coefficient does not exceed the preset second coefficient threshold, the voice segment data whose audio quality evaluation coefficient does not exceed the preset second coefficient threshold is determined as a low-quality voice segment.

[0120] The working principle of the above technical solution is: extract key audio parameters from each voice segment data, including total harmonic distortion value, spectrum dynamic range, spectrum width, distance between spectrum center and frequency center, and ratio between byte audio intensity and background audio intensity. Calculate the audio quality adjustment coefficient by a specific formula using the extracted spectrum dynamic range, spectrum width, and (the "Sd" in the original text should be a typo, and it is assumed here that it is a parameter that has been correctly understood or included in the evaluation of Sr, so Sr is used instead). In the formula, λ01 and λ02 represent the weight values ​​corresponding to the spectrum dynamic range and spectrum width, respectively, and these weight values ​​can be adjusted according to actual application scenarios and requirements. Combine the calculated audio quality adjustment coefficient with the total harmonic distortion value and the ratio between byte audio intensity and background audio intensity, and calculate the audio quality evaluation coefficient by another formula. This coefficient combines the influence of multiple audio parameters and can more comprehensively reflect the audio quality of the voice segment. Compare the calculated audio quality evaluation coefficient with the preset first coefficient threshold and second coefficient threshold. According to the comparison result, the voice segment data is judged as high quality, medium quality, or low quality.

[0121] The effect of the above technical solution is: by extracting multiple key audio parameters and comprehensively considering their influence, the technical solution can more accurately evaluate the audio quality of the voice segment. This helps to screen out high-quality voice segments in practical applications and improve the accuracy of voice processing or recognition. The weight value in the technical solution can be adjusted according to the actual application scenario and requirements. This means that in different application scenarios, the process of audio quality evaluation can be optimized according to specific requirements, thereby enhancing the flexibility of audio processing. The technical solution can be applied to the audio quality grading of multiple voice segment data, supporting automation and batch processing. This is very useful in practical applications and can greatly improve processing efficiency and accuracy. By determining the voice segment data into different quality levels, the technical solution can provide a basis for subsequent audio optimization. For example, for low-quality voice segments, corresponding measures can be taken to improve or repair them; for high-quality voice segments, they can be further compressed or encoded to save storage space or transmission bandwidth.

[0122] On the other hand, by extracting multiple audio parameters including total harmonic distortion value, spectrum dynamic range, spectrum width, distance between spectrum center and frequency center, and ratio between byte audio intensity and background audio intensity, the scheme can comprehensively and accurately evaluate the audio quality of speech segments. These parameters cover the distortion degree of audio signals, spectrum characteristics, and contrast relationship between signals and background noise, so that speech segments of different qualities can be distinguished more finely. By using audio quality adjustment coefficients and audio quality evaluation coefficients, the scheme can adaptively grade according to the actual audio characteristics of speech segments. By introducing weight values ​​of parameters such as spectrum dynamic range, spectrum width, and distance between spectrum center and frequency center, the scheme can flexibly adjust the grading standard to meet the needs of different application scenarios. By comparing the audio quality evaluation coefficient with the preset first coefficient threshold and second coefficient threshold, the scheme can efficiently divide speech segments into three levels of high quality, medium quality, and low quality. This grading determination method is concise and clear, easy to implement, and can ensure the accuracy of grading. The technical scheme can quickly grade the audio quality of a large number of speech segments, thereby improving the efficiency of speech processing. This is of great significance for scenarios that require processing large amounts of voice data, such as speech recognition and speech synthesis. By classifying voice segments into different quality levels, the solution can guide subsequent processing flows, such as prioritizing high-quality voice segments, or adopting different processing strategies for voice segments of different quality levels. This helps optimize resource utilization and improve overall processing efficiency. In voice interaction applications, high-quality voice segments can provide clearer and more accurate voice recognition results, thereby improving user experience. This technical solution helps to screen out high-quality voice segments and improve the accuracy and fluency of voice recognition through accurate audio quality grading.

[0123] In summary, the technical effects of the above technical solution in terms of performance indicators are mainly reflected in multi-dimensional audio quality evaluation, adaptive audio quality grading, efficient and accurate grading determination, improving speech processing efficiency, optimizing resource utilization, and enhancing user experience. These technical effects make the solution have broad application prospects and important practical value in the field of speech processing. At the same time, the technical solution realizes the audio quality grading of multiple voice segment data through the steps of extracting audio parameters, calculating audio quality adjustment coefficients and audio quality evaluation coefficients, and determining audio quality levels. This technical solution has significant technical effects in improving the accuracy of audio quality assessment, enhancing the flexibility of audio processing, supporting automation and batch processing, and providing a basis for audio optimization.

[0124] In one embodiment of the present invention, according to the speech recognition quality of the current multiple voice assistants combined with the audio quality levels of multiple voice segments, a target voice assistant corresponding to a voice segment of each audio quality level is obtained, including:

[0125] Extract the recognition accuracy and recognition processing time of each voice assistant;

[0126] Obtaining a recognition evaluation coefficient according to the recognition accuracy and recognition processing time corresponding to each voice assistant;

[0127] The identification evaluation coefficient is obtained by the following formula:

[0128]

[0129] Where J represents the recognition evaluation coefficient; n represents the number of times each voice assistant completes recognition; T i represents the recognition processing time corresponding to the i-th speech recognition of each voice assistant; P zi represents the recognition accuracy rate of each voice assistant’s i-th speech recognition; C i represents the amount of voice data corresponding to the i-th voice recognition of each voice assistant; C d Indicates the preset unit voice data volume; T c Indicates the theoretical recognition processing time per unit of voice data for each voice assistant;

[0130] The voice assistant corresponding to the maximum value of the recognition evaluation coefficient is used as the target voice assistant for the low-quality voice segment;

[0131] The voice assistant whose recognition evaluation coefficient is second only to the maximum recognition evaluation coefficient is used as the target voice assistant for the medium-quality voice segment;

[0132] The voice assistant corresponding to the recognition evaluation coefficient of the target voice assistant whose recognition evaluation coefficient is second only to the recognition evaluation coefficient of the medium-quality voice segment is used as the target voice assistant for the high-quality voice segment.

[0133] The working principle of the above technical solution is: extract the recognition accuracy and recognition processing time of each voice assistant. These two indicators are key parameters for measuring the performance of voice assistants. The recognition accuracy reflects the ability of the voice assistant to correctly recognize voice commands, while the recognition processing time reflects its processing speed. According to the recognition accuracy, recognition processing time and corresponding voice data volume of each voice assistant, the recognition rating coefficient is calculated by the above formula. This formula comprehensively considers the impact of recognition accuracy, processing speed and voice data volume on the performance of the voice assistant. Among them, the theoretical recognition processing time per unit voice data volume is used as a benchmark to evaluate the efficiency of the voice assistant when processing different data volumes. According to the calculated recognition rating coefficient, a target voice assistant is assigned to each voice segment of the audio quality level. Specifically, the voice assistant corresponding to the maximum value of the recognition rating coefficient is assigned to the low-quality voice segment, because the recognition of the low-quality voice segment is more difficult and requires a voice assistant with stronger recognition ability to process. Then, the voice assistant with the recognition rating coefficient second only to the maximum value is assigned to the medium-quality voice segment, and so on until the target voice assistant is assigned to the high-quality voice segment. This allocation method ensures that each voice segment of the audio quality level can get the voice assistant that is most suitable for its processing.

[0134] The effect of the above technical solution is: by assigning the most suitable target voice assistant to each voice clip of audio quality level, the technical solution can make full use of the performance advantages of different voice assistants, thereby improving the accuracy of overall speech recognition. The technical solution intelligently matches the performance indicators of the voice assistant and the audio quality level of the voice clip, avoiding waste and unreasonable allocation of resources. This helps to improve resource utilization and processing efficiency in practical applications. The technical solution can adapt to voice clips of different quality levels and voice assistants with different performance, enhancing the flexibility and adaptability of the system. This enables the system to maintain efficient and stable operation in different application scenarios and conditions. By improving speech recognition accuracy and optimizing resource allocation, the technical solution can provide users with a smoother and more accurate voice interaction experience. This helps to enhance users' trust and satisfaction with voice assistants, thereby promoting the development and application of voice technology.

[0135] On the other hand, by combining the recognition accuracy and recognition processing time of multiple voice assistants, as well as the corresponding audio quality level of voice clips, the solution can more accurately match voice clips of different qualities to the most suitable voice assistant. This helps to improve the accuracy and efficiency of overall voice recognition. The recognition rating coefficient (J) is introduced, which comprehensively considers the recognition accuracy (P zi ), recognition processing time (T i ) and the amount of voice data (C i ) and the theoretical recognition processing time per unit of speech data (T c) to more comprehensively evaluate the performance of voice assistants. This comprehensive evaluation helps select voice assistants with faster processing speed and higher accuracy, and improves recognition efficiency. The solution assigns voice assistants to voice clips of different quality levels according to the recognition rating coefficient. Low-quality voice clips are processed by the voice assistant with the highest recognition rating coefficient, medium-quality ones by the second highest, and high-quality ones by the third highest. This allocation strategy ensures that even in the face of low-quality voice input, the most accurate recognition results can be obtained, while high-quality voice clips can also be processed efficiently. By considering the performance indicators of multiple voice assistants and dynamically allocating them according to the quality of actual voice clips, the solution enhances the robustness of the system. Even if a voice assistant performs poorly in some cases, other voice assistants with better performance can be supplemented in time to ensure the stability and reliability of the overall system. The solution provides data support for the continuous optimization of voice assistant technology by quantitatively evaluating the performance of different voice assistants. Developers can make targeted improvements and optimizations to voice assistants based on the feedback of recognition rating coefficients to further improve their recognition accuracy and processing efficiency.

[0136] In summary, the technical effects of this technical solution in terms of performance indicators are mainly reflected in optimizing speech recognition quality matching, improving recognition efficiency, flexibly responding to speech clips of different quality levels, enhancing system robustness, and promoting continuous optimization of voice assistant technology. At the same time, this technical solution assigns the most suitable target voice assistant to each speech clip of each audio quality level by comprehensively considering the performance indicators of the voice assistant and the audio quality level of the voice clip. This solution has significant technical effects in improving speech recognition accuracy, optimizing resource allocation, enhancing system adaptability, and improving user experience.

[0137] In one embodiment of the present invention, the speech recognition processing process of the second speech processing strategy includes:

[0138] When the number of users currently sending voice data is multiple users, and the number of multiple users is less than or equal to the number of voice assistants, the users are assigned voice assistants according to the audio intensity ratio, and the assigned voice assistants are used to perform voice recognition processing on the voice data of the corresponding users;

[0139] When the number of users currently sending voice data is multiple users, and the number of multiple users is greater than the number of voice assistants, multiple voice assistants are controlled to collaboratively perform voice recognition processing through user tags.

[0140] The working principle of the above technical solution is as follows: first, the system determines the number of users who currently send voice data. Then, the system compares the number of users with the number of available voice assistants.

[0141] Voice Assistant Distribution Strategy:

[0142] When the number of users is less than or equal to the number of voice assistants:

[0143] The system ranks or categorizes users based on the audio intensity ratios.

[0144] Then, each user is assigned a corresponding voice assistant, which is responsible for processing the voice data of its corresponding user and performing voice recognition.

[0145] When the number of users is higher than the number of voice assistants:

[0146] The system cannot assign a voice assistant to each user individually.

[0147] At this time, the system uses user tags (such as user ID, session ID, etc.) to distinguish and track the voice data of different users.

[0148] Multiple voice assistants work together to process the voice data of all users. The system attempts to separate the commands of different users from the mixed voice data and perform accurate voice recognition.

[0149] Speech Recognition Processing:

[0150] Regardless of the distribution strategy adopted, the assigned voice assistant or the collaborative voice assistant will process the user's voice data. This includes steps such as voice signal acquisition, preprocessing, feature extraction, model matching, and decoding output, and finally converting the voice data into text or commands.

[0151] The effect of the above technical solution is: by intelligently allocating voice assistants, the solution ensures the effective use of resources. When the number of users is small, each user can get a dedicated voice assistant and enjoy high-quality services. When the number of users is large, through collaborative work, multiple voice assistants can jointly handle the requests of all users, avoiding idleness and waste of resources. For a single user, a dedicated voice assistant can respond to its request faster and provide more personalized services. In a multi-user scenario, through collaborative processing and the use of user tags, the system can identify the instructions of each user as accurately as possible, reduce misidentification and conflicts, and thus improve the overall user experience. The solution can dynamically adjust the resource allocation strategy according to the number of different users, enhancing the flexibility and adaptability of the system. This enables the system to maintain efficient and stable operation in different application scenarios and conditions to meet the needs of different users. The technical solution demonstrates the application potential and challenges of voice technology in multi-user scenarios. Through continuous optimization and improvement, the solution is expected to promote the development and innovation of voice technology and provide more intelligent and convenient services to more users.

[0152] In summary, this technical solution achieves efficient use of resources, improved user experience, enhanced system flexibility, and promoted development of voice technology by intelligently allocating and utilizing voice assistants for voice recognition processing.

[0153] In one embodiment of the present invention, assigning a voice assistant to a user according to the audio intensity ratio includes:

[0154] When the number of users currently sending voice data is multiple users, the voice data of the multiple users are split to obtain voice data corresponding to each user;

[0155] Extracting the audio intensity corresponding to the voice data corresponding to each user;

[0156] Compare the audio intensity corresponding to the voice data corresponding to each user with the overall audio intensity corresponding to multiple users to obtain an audio intensity ratio;

[0157] Obtaining a first audio recognition difficulty coefficient by using the audio intensity ratio in combination with a byte frequency corresponding to the voice data of each user;

[0158] The first audio recognition difficulty coefficient is obtained by the following formula:

[0159]

[0160] Among them, J 01 Indicates the difficulty coefficient of the first audio recognition; P d01 represents the audio intensity ratio corresponding to each user; f represents the byte frequency corresponding to the voice data of each user; f c Indicates the preset byte frequency reference value;

[0161] The voice assistant corresponding to the maximum value of the recognition evaluation coefficient is used as the target voice assistant for the user corresponding to the maximum value of the first audio recognition difficulty coefficient;

[0162] According to the allocation strategy of the first audio recognition difficulty coefficient from high to low corresponding to the recognition evaluation coefficient from low to high, other users except the users with the maximum first audio recognition difficulty coefficient are matched with the remaining voice assistants in turn.

[0163] The working principle of the above technical solution is: when there are multiple users sending voice data at the same time, the system will first split the voice data to ensure that the independent voice data corresponding to each user can be obtained. Next, the system will extract the audio intensity of each user's voice data. Audio intensity is an indicator to measure the strength of the voice signal, which reflects the energy of the voice signal. Then, the system will compare the audio intensity of each user with the overall audio intensity corresponding to multiple users to calculate the audio intensity ratio of each user. This ratio reflects the proportion of each user's voice signal in the overall voice signal. Next, the system will use the audio intensity ratio and the byte frequency of each user's voice data to calculate the first audio recognition difficulty coefficient. This coefficient is a comprehensive indicator that takes into account the intensity and frequency characteristics of the voice signal and is used to evaluate the difficulty of voice recognition. Specifically, Pd01 in the coefficient calculation formula represents the audio intensity ratio, f represents the byte frequency, and fc represents the preset byte frequency reference value. Through the combination and operation of these parameters, a value reflecting the difficulty of voice recognition can be obtained. Finally, the system will assign voice assistants according to the size of the first audio recognition difficulty coefficient. The voice assistant with the highest recognition rating coefficient will be assigned to the user with the highest first audio recognition difficulty coefficient, that is, the user who is the most difficult to recognize. Then, other users will be matched with the remaining voice assistants in turn according to the strategy of first audio recognition difficulty coefficient from high to low and recognition rating coefficient from low to high.

[0164] The effect of the above technical solution is: by comprehensively considering factors such as audio intensity and byte frequency, the solution can more accurately evaluate the difficulty of speech recognition, thereby assigning a more suitable voice assistant to the user. This helps to improve the accuracy of speech recognition and reduce the misrecognition rate. The solution can dynamically allocate voice assistant resources according to the difficulty of the user's speech recognition. For users with greater difficulty in speech recognition, the system will allocate more resources to ensure the accuracy of recognition; for users with less difficulty, fewer resources will be allocated to achieve optimal allocation of resources. By assigning more suitable voice assistants to users, the solution can enhance the user experience. Users can obtain more accurate and timely speech recognition services, thereby improving their satisfaction and trust in intelligent voice assistants. The solution takes into account the impact of multiple factors on the difficulty of speech recognition, including audio intensity and byte frequency. This enables the system to be more robust and adaptable when faced with speech data in different environments and conditions.

[0165] On the other hand, by splitting the voice data of multiple users and extracting the audio intensity separately, the scheme can accurately capture the voice characteristics of each user. By comprehensively considering the audio intensity ratio and byte frequency, the recognition difficulty of each user's voice can be more accurately evaluated, so that the most suitable voice assistant can be assigned to the user. This precise allocation helps to improve the accuracy and efficiency of voice recognition. The scheme realizes the reasonable allocation of voice assistant resources by calculating the first audio recognition difficulty coefficient and combining it with the recognition assessment coefficient. Voice assistants with higher recognition assessment coefficients are assigned to user voices with greater difficulty, while voice assistants with lower recognition assessment coefficients handle relatively simple tasks. This allocation strategy not only improves the utilization rate of voice assistants, but also ensures the performance optimization of the overall voice recognition system. The scheme can respond flexibly when multiple users send voice data at the same time. By splitting and comparing the voice data of each user, the scheme can dynamically adjust the allocation of voice assistants to adapt to the recognition needs of different user voices. This dynamic adaptability makes the scheme perform well in a multi-user environment. By assigning the most suitable voice assistant to the user, the scheme can reduce voice recognition errors and increase recognition speed, thereby optimizing the user experience. Users can interact with voice assistants more quickly and accurately, improving overall satisfaction. Parameters such as the audio intensity ratio and byte frequency reference value in this solution can be adjusted and optimized according to actual needs. This scalability and customizability enables the solution to adapt to changes in different scenarios and user needs, and has greater flexibility and adaptability. By quantitatively evaluating the difficulty of speech recognition and the performance of voice assistants, the solution provides data support for technology iteration. Developers can continuously optimize and improve speech recognition algorithms and voice assistants based on the evaluation results to improve the performance and accuracy of the overall system.

[0166] In summary, the technical effects of this technical solution in terms of performance indicators are mainly reflected in accurate user voice allocation, efficient resource utilization, dynamic adaptation to multi-user environments, optimized user experience, scalability and customizability, and promotion of technical iteration. These technical effects jointly improve the overall performance and user experience of the speech recognition system. At the same time, this technical solution allocates voice assistants to users by comprehensively considering factors such as audio intensity and byte frequency, thereby achieving optimal resource allocation and improved speech recognition accuracy.

[0167] An embodiment of the present invention controls multiple voice assistants to collaboratively perform voice recognition processing through user tags, including:

[0168] Perform unique identification processing on the voice data of each user, so that the voice data of each user has a unique identification that is uniquely associated with the user;

[0169] Dividing the voice data of all users to obtain a plurality of voice segment data, and assigning a unique identifier consistent with the voice data to the voice segment data;

[0170] Using the audio intensity ratio corresponding to each voice segment data and the amplitude of the byte audio intensity change in each voice segment; obtaining the second audio recognition difficulty coefficient corresponding to each voice segment data;

[0171] The second audio recognition difficulty coefficient is obtained by the following formula:

[0172]

[0173] Among them, J 02 P represents the difficulty coefficient of the second audio recognition; d02 represents the audio intensity ratio corresponding to each speech segment data; k represents the number of bytes contained in each speech segment data; Q i Indicates the audio intensity corresponding to the i-th byte; Q i+1 Indicates the audio intensity corresponding to the i+1th byte; Q b Indicates the standard deviation of the audio intensity corresponding to k bytes;

[0174] Comparing the second audio recognition difficulty coefficient with a preset recognition difficulty coefficient threshold;

[0175] Assigning the speech segment data whose second audio recognition difficulty coefficient exceeds a preset recognition difficulty coefficient threshold to the voice assistant corresponding to the maximum recognition evaluation coefficient for speech recognition processing, obtaining a speech recognition result corresponding to each speech segment data, and assigning a unique identifier consistent with the speech segment data to the speech recognition result;

[0176] Allocate the voice segment data whose second audio recognition difficulty coefficient does not exceed the preset recognition difficulty coefficient threshold to the remaining voice assistants for voice recognition processing according to the principle of approximate uniform distribution, obtain the voice recognition result corresponding to each voice segment data, and assign a unique identifier consistent with the voice segment data to the voice recognition result;

[0177] The voice results are filtered according to the corresponding identifier of each voice segment data, and the voice results with the same unique identifier are integrated to obtain the recognition result corresponding to the voice data uniquely associated with the user.

[0178] The working principle of the above technical solution is: uniquely identify each user's voice data to ensure that each user's voice data has a unique identifier uniquely associated with it. This helps to distinguish the voice data of different users in subsequent processing. The voice data of all users are divided into multiple voice segment data, and these voice segment data are assigned a unique identifier consistent with the original voice data. In this way, each voice segment data can be traced back to its original user. The second audio recognition difficulty coefficient is calculated using the audio intensity ratio and byte audio intensity change amplitude of each voice segment data. This coefficient takes into account the audio intensity distribution and change of the voice segment to evaluate the difficulty of its recognition.

[0179] In the specific calculation, the audio intensity of each byte and the standard deviation of the audio intensity of all bytes are taken into account, and the number of bytes contained in the voice segment is combined. The calculated second audio recognition difficulty coefficient is compared with the preset recognition difficulty coefficient threshold to distinguish between high-difficulty and low-difficulty voice segment data. For voice segment data whose difficulty coefficient exceeds the threshold, it is assigned to the voice assistant corresponding to the maximum value of the recognition assessment coefficient for voice recognition processing. This ensures that high-difficulty voice segments are processed more professionally. For voice segment data whose difficulty coefficient does not exceed the threshold, it is allocated to the remaining voice assistants for voice recognition processing according to the principle of approximate uniform distribution. During the recognition process, the voice recognition result of each voice segment data is assigned a unique identifier consistent with the voice segment data for subsequent integration. Finally, the voice recognition results are screened and integrated according to the unique identifier corresponding to each voice segment data to obtain the final recognition result corresponding to the voice data uniquely associated with the user.

[0180] The effect of the above technical solution is: by calculating the second audio recognition difficulty coefficient, the solution can more accurately evaluate the difficulty of recognizing the voice segment, and allocate voice assistant resources according to the difficulty coefficient. This helps to ensure that difficult voice segments are processed more professionally, thereby improving the accuracy of overall voice recognition. The solution can dynamically adjust the allocation strategy of the voice assistant according to the difficulty of recognizing the voice segment to achieve optimal allocation of resources. For voice segments with greater difficulty, more resources are allocated; for voice segments with less difficulty, fewer resources are allocated. This helps to improve resource utilization efficiency. The solution can handle situations where multiple users send voice data at the same time, and maintain the relevance of voice data through unique identification. This enhances the flexibility of the system, enabling it to work normally in a complex multi-user environment. Through more accurate voice recognition and optimized resource allocation, the solution can provide users with faster and more accurate voice recognition services. This helps to improve user satisfaction and trust in intelligent voice assistants.

[0181] On the other hand, by assigning a unique identifier to each user's voice data and dividing and identifying voice segments based on these identifiers, the solution can ensure efficient collaboration between multiple voice assistants. Each voice segment is accurately assigned to a voice assistant for processing, avoiding duplication of work and waste of resources. The solution introduces a second audio recognition difficulty coefficient, which comprehensively considers multiple factors such as the audio intensity ratio of the voice segment, the amplitude of the byte audio intensity change, and the standard deviation of the audio intensity. This comprehensive evaluation method can more accurately reflect the recognition difficulty of the voice segment, thereby helping to assign the more difficult voice segments to voice assistants with better performance for processing. According to the comparison result of the second audio recognition difficulty coefficient with the preset recognition difficulty coefficient threshold, the solution can intelligently assign voice segments to different voice assistants for processing. For more difficult voice segments, they are assigned to the voice assistant with the highest recognition evaluation coefficient to ensure the accuracy of the recognition results; for less difficult voice segments, they are assigned according to the principle of approximate uniform distribution to make full use of the processing capabilities of all voice assistants. Through accurate difficulty evaluation and optimized resource allocation, the solution can significantly improve the accuracy and efficiency of speech recognition. More difficult voice clips are processed by voice assistants with better performance, reducing the possibility of recognition errors; while less difficult voice clips are shared by multiple voice assistants, speeding up the overall processing speed. Throughout the voice recognition process, the solution always maintains data consistency. Each voice clip and the corresponding recognition result are given the same unique identifier as the original voice data, which makes subsequent data integration and result analysis simpler and more accurate. Because the solution can efficiently process voice data from multiple users and provide accurate recognition results, it can significantly improve the user experience. Users can interact with voice assistants more quickly and accurately, thereby improving overall satisfaction and loyalty.

[0182] In summary, the technical effects of this technical solution in terms of performance indicators are mainly reflected in efficient collaborative processing, accurate recognition difficulty assessment, optimized resource allocation, improved recognition accuracy and efficiency, maintaining data consistency, and enhanced user experience. These technical effects have jointly promoted the further development and application of speech recognition technology. At the same time, this technical solution achieves more accurate speech recognition difficulty assessment and resource allocation optimization by comprehensively considering factors such as the audio intensity ratio of the speech segment and the amplitude of the byte audio intensity change. This helps to improve the accuracy of speech recognition and the efficiency of resource utilization, thereby improving the user experience.

[0183] In one embodiment of the present invention, intelligently controlling a target system according to the voice command includes:

[0184] S301, when the number of the user is one, intelligently controlling the target system according to the voice result recognized by the voice assistant;

[0185] S301. When there are multiple users, determine whether the voice instructions corresponding to the multiple users are for the same controlled target parameters; when there are control instructions for the same controlled target parameters among the voice instructions corresponding to the multiple users, filter the voice instructions of users with priority permissions according to the user permission priority to perform intelligent control on the target system.

[0186] The working principle of the above technical solution is: when the system detects that there is only one user, the user's voice command will be directly used as the basis for controlling the target system. This means that the system will directly analyze the user's voice results and perform corresponding intelligent control on the target system based on the analysis results. When the system detects that there are multiple users, the situation will be relatively complicated. At this time, the system will first analyze the voice commands of each user to determine whether they are for the same controlled target parameters. For example, if two users issue the commands of "turn up the temperature" and "turn off the air conditioner" respectively, and these commands are all for the controlled target of the air conditioner, the system needs further processing. If there are control commands for the same controlled target parameters in the voice commands of multiple users, the system will enter the user authority priority judgment stage. At this stage, the system will filter out a user with priority authority based on the preset user authority rules (such as administrator authority is higher than ordinary user authority). Then, the system will intelligently control the target system according to the voice commands of the user.

[0187] The effect of the above technical solution is: the solution can automatically adjust the control strategy according to the number of different users and the content of the instructions, thereby improving the flexibility of control. Whether it is a single user or multiple users, the system can give a reasonable response. In a multi-user environment, the system can accurately judge the intention of the user's instructions and make decisions based on the user's authority priority. This helps to avoid system confusion caused by user instruction conflicts and enhances the robustness of the system. By intelligently judging user instructions and authority priorities, the solution can provide users with a more accurate and timely control experience. Users do not need to worry about their instructions being ignored or misunderstood, thereby improving user satisfaction with the system. In a multi-user situation, the system can reasonably allocate resources according to user authority priorities. This helps to ensure that the instructions of important users are given priority, thereby improving resource utilization efficiency.

[0188] In summary, this technical solution achieves intelligent control of the target system by comprehensively considering the number of users and the content of the user's voice commands. This improves the flexibility of control, enhances the robustness of the system, improves the user experience, and achieves a reasonable allocation of resources.

[0189] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.

Claims

1. An intelligent recognition control method based on multi-voice assistant, characterized in that: include: Real-time monitoring of the received voice data sent by the user, and determining the voice processing strategy of the multi-voice assistant according to the number of users sending the voice data; wherein the voice processing strategy includes a first voice processing strategy and a second voice processing strategy; and the first voice processing strategy refers to a voice processing strategy for the multi-voice assistants to collaboratively process the voice data of one user; the second voice processing strategy refers to a voice processing strategy for the multi-voice assistants to collaboratively process the voice data of multiple users; Performing voice recognition processing on the voice data sent by the user according to the voice processing strategy to obtain voice instructions corresponding to the voice data; Intelligently controlling the target system according to the voice command; The speech recognition processing process of the second speech processing strategy includes: S1: When the number of the multiple users is less than or equal to the number of voice assistants, the users are assigned voice assistants according to the audio intensity ratio, and the assigned voice assistants are used to perform voice recognition processing on the voice data of the corresponding users; S2: When the number of multiple users is greater than the number of voice assistants, multiple voice assistants are controlled to collaboratively perform voice recognition processing through user tags; In the implementation process of S1, the voice data of multiple users are split to obtain the voice data corresponding to each user; Extracting the audio intensity corresponding to the voice data corresponding to each user; Compare the audio intensity corresponding to the voice data corresponding to each user with the overall audio intensity corresponding to multiple users to obtain an audio intensity ratio; Obtaining a first audio recognition difficulty coefficient by using the audio intensity ratio in combination with a byte frequency corresponding to the voice data of each user; The voice assistant corresponding to the maximum value of the recognition evaluation coefficient is used as the target voice assistant for the user corresponding to the maximum value of the first audio recognition difficulty coefficient; According to the allocation strategy of the first audio recognition difficulty coefficient from high to low corresponding to the recognition evaluation coefficient from low to high, other users except the users with the maximum first audio recognition difficulty coefficient are matched with the remaining voice assistants in turn.

2. The intelligent recognition control method based on multi-voice assistant according to claim 1 is characterized in that: The received voice data sent by the users is monitored in real time, and the voice processing strategy of the multi-voice assistant is determined according to the number of users sending the voice data, including: Real-time monitoring of the voice data received from users; Determining the number of users currently sending voice data based on the received voice data sent by the user; When the number of users is one user, the first voice processing strategy is called as the voice processing strategy; When the number of users is multiple users, the second voice processing strategy is called as the voice processing strategy.

3. The intelligent recognition control method based on multi-voice assistant according to claim 1 is characterized in that: The performing speech recognition processing on the speech data sent by the user according to the speech processing strategy to obtain the speech instruction corresponding to the speech data includes: Performing voice recognition processing on the voice data sent by the user according to the first voice processing strategy to obtain a voice instruction corresponding to the voice data; or Perform voice recognition processing on the voice data sent by the user according to the second voice processing strategy to obtain voice instructions corresponding to the voice data.

4. The intelligent recognition control method based on multi-voice assistant according to claim 2 is characterized in that: The speech recognition processing process of the first speech processing strategy includes: When the number of users currently sending voice data is one user, dividing the voice data into voice segments to obtain a plurality of voice segment data; Performing audio quality grading on the current multiple voice segment data to obtain audio quality levels of the multiple voice segment data; According to the speech recognition quality of the current multiple voice assistants and the audio quality levels of the multiple voice clips, a target voice assistant corresponding to the voice clip of each audio quality level is obtained; The voice segments of each audio quality level are assigned to their corresponding target voice assistants for voice recognition processing to obtain voice commands corresponding to the voice data.

5. The intelligent recognition control method based on multi-voice assistant according to claim 4 is characterized in that: Perform audio quality grading on the current multiple voice segment data to obtain the audio quality levels of the multiple voice segments, including: Extracting audio parameters of each speech segment data, wherein the audio parameters include a total harmonic distortion value, a spectrum dynamic range, a spectrum width, a distance between a spectrum center and a frequency center, and a ratio between a byte audio intensity and a background audio intensity; Obtaining an audio quality adjustment coefficient using the spectrum dynamic range, the spectrum width, and the distance between the spectrum center and the frequency center; The audio quality adjustment coefficient is obtained by the following formula: Where U represents the audio quality adjustment coefficient; S r Indicates the spectrum dynamic range; S w Indicates the spectrum width; S d Indicates the distance between the spectrum center and the frequency center; λ 01 and λ 02 Respectively represent the weight values ​​corresponding to the spectrum dynamic range and spectrum width; The audio quality evaluation coefficient is obtained by combining the audio quality adjustment coefficient with the total harmonic distortion value and the ratio between the byte audio intensity and the background audio intensity; The audio quality evaluation coefficient is obtained by the following formula: Where R represents the audio quality evaluation coefficient; T h Indicates the total harmonic distortion value; B indicates the ratio between the byte audio intensity and the background audio intensity; U indicates the audio quality adjustment coefficient; Comparing the audio quality evaluation coefficient with a preset first coefficient threshold and a second coefficient threshold; When the audio quality evaluation coefficient exceeds a preset first coefficient threshold, determining the voice segment data whose audio quality evaluation coefficient exceeds the preset first coefficient threshold as a high-quality voice segment; When the audio quality evaluation coefficient does not exceed the preset first coefficient threshold, but exceeds the second coefficient threshold, the voice segment data whose audio quality evaluation coefficient does not exceed the preset first coefficient threshold, but exceeds the second coefficient threshold is determined as a medium-quality voice segment; When the audio quality evaluation coefficient does not exceed the preset second coefficient threshold, the voice segment data whose audio quality evaluation coefficient does not exceed the preset second coefficient threshold is determined as a low-quality voice segment.

6. The intelligent recognition control method based on multi-voice assistant according to claim 4 is characterized in that: According to the current speech recognition quality of multiple voice assistants combined with the audio quality levels of multiple voice clips, the target voice assistant corresponding to the voice clip of each audio quality level is obtained, including: Extract the recognition accuracy and recognition processing time of each voice assistant; Obtaining a recognition evaluation coefficient according to the recognition accuracy and recognition processing time corresponding to each voice assistant; The identification evaluation coefficient is obtained by the following formula: Where J represents the recognition evaluation coefficient; n represents the number of times each voice assistant completes recognition; T i represents the recognition processing time corresponding to the i-th speech recognition of each voice assistant; P zi represents the recognition accuracy rate of each voice assistant’s i-th speech recognition; C i represents the amount of voice data corresponding to the i-th voice recognition of each voice assistant; C d Indicates the preset unit voice data volume; T c Indicates the theoretical recognition processing time per unit of voice data for each voice assistant; The voice assistant corresponding to the maximum value of the recognition evaluation coefficient is used as the target voice assistant for the low-quality voice segment; The voice assistant whose recognition evaluation coefficient is second only to the maximum recognition evaluation coefficient is used as the target voice assistant for the medium-quality voice segment; The voice assistant corresponding to the recognition evaluation coefficient of the target voice assistant whose recognition evaluation coefficient is second only to the recognition evaluation coefficient of the medium-quality voice segment is used as the target voice assistant for the high-quality voice segment.

7. The intelligent recognition control method based on multi-voice assistant according to claim 1 is characterized in that: in, The first audio recognition difficulty coefficient is obtained by the following formula: Among them, J 01 Indicates the difficulty coefficient of the first audio recognition; P d01 represents the audio intensity ratio corresponding to each user; f represents the byte frequency corresponding to the voice data of each user; f c Indicates the preset byte frequency reference value.

8. The intelligent recognition control method based on multi-voice assistant according to claim 1, characterized in that: Control multiple voice assistants to collaborate on voice recognition processing through user tags, including: Perform unique identification processing on the voice data of each user, so that the voice data of each user has a unique identification that is uniquely associated with the user; Dividing the voice data of all users to obtain a plurality of voice segment data, and assigning a unique identifier consistent with the voice data to the voice segment data; Using the audio intensity ratio corresponding to each voice segment data and the amplitude of the byte audio intensity change in each voice segment; obtaining the second audio recognition difficulty coefficient corresponding to each voice segment data; The second audio recognition difficulty coefficient is obtained by the following formula: Among them, J 02 P represents the difficulty coefficient of the second audio recognition; d02 represents the audio intensity ratio corresponding to each speech segment data; k represents the number of bytes contained in each speech segment data; Q i Indicates the audio intensity corresponding to the i-th byte; Q i+1 Indicates the audio intensity corresponding to the i+1th byte; Q b Indicates the standard deviation of the audio intensity corresponding to k bytes; Comparing the second audio recognition difficulty coefficient with a preset recognition difficulty coefficient threshold; Assigning the speech segment data whose second audio recognition difficulty coefficient exceeds a preset recognition difficulty coefficient threshold to a voice assistant corresponding to a maximum recognition evaluation coefficient for speech recognition processing, obtaining a speech recognition result corresponding to each speech segment data, and assigning a unique identifier consistent with the speech segment data to the speech recognition result; Allocate the speech segment data whose second audio recognition difficulty coefficient does not exceed the preset recognition difficulty coefficient threshold to the remaining voice assistants for speech recognition processing according to the principle of approximate uniform distribution, obtain the speech recognition result corresponding to each speech segment data, and assign a unique identifier consistent with the speech segment data to the speech recognition result; The voice results are filtered according to the corresponding identifier of each voice segment data, and the voice results with the same unique identifier are integrated to obtain the recognition result corresponding to the voice data uniquely associated with the user.

9. The intelligent recognition control method based on multi-voice assistant according to claim 1, characterized in that: Intelligently control the target system according to the voice command, including: When the number of the user is one, the target system is intelligently controlled according to the speech result recognized by the voice assistant; When the number of users is multiple, it is determined whether the voice commands corresponding to the multiple users are for the same controlled target parameters; when there are control commands for the same controlled target parameters among the voice commands corresponding to the multiple users, the voice commands of users with priority permissions are filtered according to the user permission priority to perform intelligent control of the target system.

Citation Information

Patent Citations

  • Method and device for processing voice information collected by multiple voice assistant devices

    CN107393548A

  • Coordination method, device and system for multiple voice assistants

    CN109712624A