Intelligent recognition and control method based on multiple voice assistants
By monitoring the number of users in real time and selecting appropriate voice processing strategies, the system dynamically matches the voice assistant to process multi-user voice data, solving the problem of chaotic voice recognition in multi-user environments. This achieves efficient and accurate voice command processing, improving system stability and user experience.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- BEIJING JIZHI TECH CO LTD
- Filing Date
- 2025-04-17
- Publication Date
- 2026-07-30
AI Technical Summary
Existing intelligent recognition and control systems struggle to accurately recognize and process each user's voice commands when multiple users issue them simultaneously, leading to recognition confusion and control errors. Furthermore, existing methods increase computational resource consumption and response latency.
The intelligent recognition and control method based on multiple voice assistants dynamically selects voice processing strategies by monitoring the number of users in real time. This includes collaboratively processing the voice data of a single user or multiple users, and matching the most suitable voice assistant for recognition processing using audio quality grading and recognition evaluation coefficients.
It achieves accurate recognition and processing of voice commands in a multi-user environment, improves the system's recognition accuracy and stability, avoids control confusion and errors, ensures rapid response, and enhances user experience.
Smart Images

Figure CN2025089526_30072026_PF_FP_ABST
Abstract
Description
Intelligent Recognition and Control Method Based on Multiple Voice Assistants Technical Field
[0001] This invention proposes an intelligent recognition and control method based on multiple voice assistants, belonging to the field of speech recognition technology. Background Technology
[0002] With the rapid development of artificial intelligence technology, voice assistants have become an important component in fields such as smart homes, smart offices, and smart vehicle systems. These voice assistants control various smart devices by receiving users' voice commands, greatly improving the user experience and the level of system intelligence. However, in existing intelligent recognition and control systems, when faced with scenarios where multiple users simultaneously issue voice commands, the system often experiences recognition confusion and control errors.
[0003] Traditional single-user voice assistant systems are designed primarily for interaction between a single user and the system. They often lack effective mechanisms for distinguishing and processing multiple voice signals received simultaneously. This can lead to inaccurate recognition of each user's commands in multi-user environments, potentially resulting in misrecognition or incorrect operation. For example, in smart home scenarios, if family members simultaneously issue different control commands to the voice assistant, the system may confuse these commands, leading to incorrect device control. To address this issue, some existing technologies attempt to improve the system's recognition capabilities by increasing the complexity of the voice recognition module and the accuracy of the algorithm. However, while these methods improve the accuracy of single-user command recognition to some extent, they still struggle to effectively distinguish and correctly process each user's commands when multiple users speak simultaneously. Furthermore, these methods often consume more computing resources and time, reducing the system's real-time performance and response speed.
[0004] Therefore, there is an urgent need for an intelligent recognition and control method capable of effectively handling simultaneous voice commands from multiple users. This method needs to be able to monitor received user voice data in real time and intelligently select a voice processing strategy based on the number of users to ensure that the system can accurately recognize and process each user's commands in a multi-user environment. Simultaneously, the method also needs to be efficient and real-time to meet the requirements of modern intelligent systems for response speed and stability. Summary of the Invention
[0005] This invention provides an intelligent recognition and control method based on multiple voice assistants to solve the aforementioned technical problems in the prior art. The technical solution adopted is as follows:
[0006] A multi-voice assistant-based intelligent recognition and control method, comprising:
[0007] The system monitors received voice data from users in real time and determines a voice processing strategy for multiple voice assistants based on the number of users sending voice data. The voice processing strategy includes a first voice processing strategy and a second voice processing strategy. The first voice processing strategy refers to a voice processing strategy in which multiple voice assistants collaboratively process the voice data of one user. The second voice processing strategy refers to a voice processing strategy in which multiple voice assistants collaboratively process the voice data of multiple users.
[0008] The voice data emitted by the user is processed by voice recognition according to the voice processing strategy to obtain the voice command corresponding to the voice data.
[0009] The target system is intelligently controlled according to the voice commands.
[0010] Furthermore, the system monitors received voice data from users in real time and determines the voice processing strategy for multiple voice assistants based on the number of users sending voice data, including:
[0011] Real-time monitoring of received voice data from users;
[0012] The number of users currently sending voice data is determined based on the received voice data sent by the users.
[0013] When the number of users is one, the first voice processing strategy is invoked as the voice processing strategy.
[0014] When there are multiple users, the second voice processing strategy is invoked as the voice processing strategy.
[0015] Further, the step of performing speech recognition processing on the voice data emitted by the user according to the speech processing strategy to obtain the voice command corresponding to the voice data includes:
[0016] The voice data emitted by the user is processed by voice recognition according to the first voice processing strategy to obtain the voice command corresponding to the voice data.
[0017] or
[0018] The user's voice data is processed by voice recognition according to the second voice processing strategy to obtain the voice command corresponding to the voice data.
[0019] Furthermore, the speech recognition processing of the first speech processing strategy includes:
[0020] When the number of users currently sending voice data is one user, the voice data is divided into multiple voice segment data.
[0021] The audio quality of multiple speech segments is graded to obtain the audio quality level of each segment.
[0022] Based on the speech recognition quality of multiple voice assistants and the audio quality levels of multiple speech segments, the target voice assistant corresponding to each audio quality level of the speech segment is obtained.
[0023] The speech segments of each audio quality level are assigned to their corresponding target voice assistants for speech recognition processing to obtain the voice commands corresponding to the speech data.
[0024] Furthermore, the audio quality of the current multiple speech segments is graded to obtain the audio quality level of each speech segment, including:
[0025] Extract audio parameters for each speech segment data, wherein the audio parameters include total harmonic distortion value, dynamic range of spectrum, spectral width, distance between the spectral center and the frequency center, and the ratio between byte audio intensity and background audio intensity;
[0026] The audio quality adjustment coefficient is obtained using the aforementioned spectral dynamic range, spectral width, and distance between the spectral center and the frequency center.
[0027] The audio quality adjustment coefficient is obtained using the following formula:
[0028] Where U represents the audio quality adjustment coefficient; S r Indicates the dynamic range of the spectrum; S w Indicates the bandwidth; S d λ represents the distance between the center of the spectrum and the center of the frequency. 01 and λ 02 These represent the weight values corresponding to the dynamic range and bandwidth of the spectrum, respectively.
[0029] The audio quality evaluation coefficient is obtained by combining the audio quality adjustment coefficient with the total harmonic distortion value and the ratio between the byte audio intensity and the background audio intensity.
[0030] The audio quality evaluation coefficient is obtained using the following formula:
[0031] Where R represents the audio quality evaluation coefficient; T h B represents the total harmonic distortion value; B represents the ratio between the byte audio intensity and the background audio intensity; U represents the audio quality adjustment factor.
[0032] The audio quality evaluation coefficient is compared with a preset first coefficient threshold and a second coefficient threshold;
[0033] When the audio quality evaluation coefficient exceeds a preset first coefficient threshold, the audio segment data with the audio quality evaluation coefficient exceeding the preset first coefficient threshold is determined as a high-quality audio segment.
[0034] If the audio quality evaluation coefficient does not exceed the preset first coefficient threshold, but exceeds the second coefficient threshold, then the audio segment data whose audio quality evaluation coefficient does not exceed the preset first coefficient threshold but exceeds the second coefficient threshold is determined as a medium quality audio segment.
[0035] If the audio quality evaluation coefficient does not exceed the preset second coefficient threshold, then the speech segment data whose audio quality evaluation coefficient does not exceed the preset second coefficient threshold is determined as a low-quality speech segment.
[0036] Furthermore, based on the speech recognition quality of the current multiple voice assistants and the audio quality levels of multiple speech segments, the target voice assistant corresponding to each audio quality level of the speech segment is obtained, including:
[0037] Extract the recognition accuracy and recognition processing time for each voice assistant;
[0038] The recognition evaluation coefficient is obtained based on the recognition accuracy and recognition processing time corresponding to each voice assistant;
[0039] The identification rating coefficient is obtained using the following formula:
[0040] Where J represents the recognition rating coefficient; n represents the number of times each voice assistant completes recognition; T i P represents the processing time for the i-th speech recognition attempt by each voice assistant; zi C represents the recognition accuracy for the i-th speech recognition attempt by each voice assistant; i C represents the amount of voice data corresponding to the i-th speech recognition of each voice assistant; d Indicates the preset unit voice data volume; T c This represents the theoretical recognition and processing time for each unit of voice data corresponding to each voice assistant;
[0041] The voice assistant corresponding to the maximum value of the recognition evaluation coefficient is used as the target voice assistant for the low-quality voice segment.
[0042] The voice assistant whose recognition rating coefficient is second only to the maximum recognition rating coefficient is selected as the target voice assistant for the medium-quality voice segment.
[0043] The voice assistant whose recognition rating coefficient is second only to that of the medium-quality speech segment is considered as the target voice assistant for the high-quality speech segment.
[0044] Furthermore, the speech recognition processing of the second speech processing strategy includes:
[0045] When the number of users currently sending voice data is multiple users, and the number of multiple users is less than or equal to the number of voice assistants, then the users are assigned voice assistants according to the audio intensity ratio, and the assigned voice assistants are used to perform voice recognition processing on the voice data of their corresponding users.
[0046] When the number of users currently sending voice data is multiple users, and the number of multiple users is higher than the number of voice assistants, then the multiple voice assistants are controlled to collaboratively perform voice recognition processing through user tags.
[0047] Further, assigning a voice assistant to a user based on the audio intensity ratio includes:
[0048] When there are multiple users currently sending voice data, the voice data of multiple users is split to obtain the voice data corresponding to each user;
[0049] Extract the audio intensity corresponding to the voice data of each user;
[0050] The audio intensity ratio is obtained by comparing the audio intensity of each user's voice data with the overall audio intensity of multiple users.
[0051] The first audio recognition difficulty coefficient is obtained by combining the audio intensity ratio with the byte frequency corresponding to each user's voice data.
[0052] The difficulty coefficient of the first audio recognition is obtained by the following formula:
[0053] Among them, J 01 P represents the difficulty level of the first audio recognition; d01 This represents the audio intensity ratio for each user; f represents the byte frequency of each user's voice data; f c This indicates the preset byte frequency reference value;
[0054] The voice assistant corresponding to the highest recognition evaluation coefficient is taken as the target voice assistant of the user corresponding to the highest first audio recognition difficulty coefficient.
[0055] According to the allocation strategy of matching the recognition evaluation coefficient from low to high based on the first audio recognition difficulty coefficient, users other than those with the highest first audio recognition difficulty coefficient are matched with the remaining voice assistants in turn.
[0056] Furthermore, multiple voice assistants are controlled to collaboratively perform speech recognition processing based on user tags, including:
[0057] Each user's voice data is uniquely identified, so that each user's voice data has a unique identifier that is uniquely associated with the user;
[0058] The voice data of all users is divided into multiple voice segment data, and each voice segment data is assigned a unique identifier that is consistent with the voice data.
[0059] By utilizing the audio intensity ratio corresponding to each speech segment data and the byte audio intensity variation amplitude in each speech segment, the second audio recognition difficulty coefficient corresponding to each speech segment data is obtained.
[0060] The difficulty coefficient for the second audio recognition is obtained using the following formula:
[0061] Among them, J 02 Indicates the difficulty level of the second audio recognition; P d02 Q represents the audio intensity ratio corresponding to each speech segment data; k represents the number of bytes contained in each speech segment data; Q i Q represents the audio intensity corresponding to the i-th byte; i+1 Q represents the audio intensity corresponding to the (i+1)th byte; b This represents the standard deviation of audio intensity corresponding to k bytes;
[0062] The second audio recognition difficulty coefficient is compared with a preset recognition difficulty coefficient threshold;
[0063] The speech segment data whose second audio recognition difficulty coefficient exceeds the preset recognition difficulty coefficient threshold is assigned to the voice assistant corresponding to the maximum recognition evaluation coefficient for speech recognition processing, the speech recognition result corresponding to each speech segment data is obtained, and the speech recognition result is assigned a unique identifier consistent with the speech segment data.
[0064] The speech segment data whose second audio recognition difficulty coefficient does not exceed the preset recognition difficulty coefficient threshold are allocated to the remaining voice assistants for speech recognition processing according to the principle of near-uniform distribution, the speech recognition result corresponding to each speech segment data is obtained, and the speech recognition result is assigned a unique identifier consistent with the speech segment data.
[0065] The speech results are filtered according to the identifier corresponding to each speech segment data, and the speech results with the same unique identifier are integrated to obtain the recognition result corresponding to the speech data uniquely associated with the user.
[0066] Furthermore, intelligent control of the target system according to the voice commands includes:
[0067] When the number of users is one, the target system is intelligently controlled according to the voice result recognized by the voice assistant;
[0068] When there are multiple users, it is determined whether the voice commands corresponding to the multiple users are for the same controlled target parameter; when there are control commands for the same controlled target parameter among the voice commands corresponding to the multiple users, the voice commands of the user with higher permission are selected according to the user permission priority to intelligently control the target system.
[0069] Beneficial effects of this invention:
[0070] The intelligent recognition and control method based on multiple voice assistants proposed in this invention monitors received user voice data in real time and determines the voice processing strategy for multiple voice assistants (including strategies for collaboratively processing voice data of a single user and strategies for collaboratively processing voice data of multiple users) according to the number of users. This achieves accurate recognition and processing of voice commands in a multi-user environment. This method not only improves the system's recognition accuracy and stability but also ensures that the system can correctly and quickly respond to each user's commands when multiple users speak simultaneously, avoiding control confusion and errors. Attached Figure Description
[0071] Figure 1 is a flowchart of the method described in this invention. Detailed Implementation
[0072] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.
[0073] This invention proposes an intelligent recognition and control method based on multiple voice assistants, as shown in Figure 1. The intelligent recognition and control method based on multiple voice assistants includes:
[0074] S1. Real-time monitoring of received voice data from users, and determination of a voice processing strategy for multiple voice assistants based on the number of users sending voice data; wherein, the voice processing strategy includes a first voice processing strategy and a second voice processing strategy; and, the first voice processing strategy refers to a voice processing strategy in which multiple voice assistants collaboratively process the voice data of one user; the second voice processing strategy refers to a voice processing strategy in which multiple voice assistants collaboratively process the voice data of multiple users.
[0075] The process of performing speech recognition processing on the user's voice data according to the aforementioned speech processing strategy to obtain the corresponding voice commands includes:
[0076] The voice data emitted by the user is processed by voice recognition according to the first voice processing strategy to obtain the voice command corresponding to the voice data.
[0077] or
[0078] The user's voice data is processed by voice recognition according to the second voice processing strategy to obtain the voice command corresponding to the voice data.
[0079] S2. Perform speech recognition processing on the voice data issued by the user according to the speech processing strategy to obtain the voice command corresponding to the voice data;
[0080] S3. Perform intelligent control of the target system according to the voice commands.
[0081] The working principle of the above technical solution is as follows: The system first monitors the received voice data from users in real time. Based on the monitored voice data, the system determines the number of users who sent the voice data. Based on the number of users, the system selects an appropriate voice processing strategy. If there is only one user, the first voice processing strategy (a strategy where multiple voice assistants collaboratively process the voice data of one user) is selected; if there are multiple users, the second voice processing strategy (a strategy where multiple voice assistants collaboratively process the voice data of multiple users) is selected. According to the selected voice processing strategy, the system performs voice recognition processing on the received voice data. If the first voice processing strategy is selected, the system utilizes the collaborative capabilities of multiple voice assistants to more accurately and quickly recognize the voice data of a single user, thereby improving the accuracy and efficiency of recognition. If the second voice processing strategy is selected, the system uses more complex algorithms and models to distinguish and recognize the voice data of multiple users, ensuring that each user's instructions are accurately understood and processed. After voice recognition processing, the system obtains the voice commands corresponding to the voice data. Based on the recognized voice commands, the system performs intelligent control of the target system. The target system can be smart home devices, smart office equipment, smart vehicle systems, etc. The system performs corresponding operations based on the content of the instructions, such as switching equipment on and off, adjusting parameters, and executing commands.
[0082] The advantages of the above technical solution are as follows: The intelligent recognition and control method based on multiple voice assistants proposed in this embodiment monitors the received user voice data in real time and determines the voice processing strategy for multiple voice assistants (including strategies for collaboratively processing the voice data of a single user and strategies for collaboratively processing the voice data of multiple users) according to the number of users, thereby achieving accurate recognition and processing of voice commands in a multi-user environment. This method not only improves the recognition accuracy and stability of the system, but also ensures that the system can correctly and quickly respond to the commands of each user when multiple users speak simultaneously, avoiding control confusion and errors.
[0083] On the other hand, by monitoring the number of users in real time and selecting appropriate voice processing strategies, the system can more accurately recognize users' voice commands, avoiding recognition confusion and control errors in multi-user environments. The collaborative work of multiple voice assistants can further improve recognition accuracy and efficiency. This method can handle scenarios where multiple users issue voice commands simultaneously, ensuring system stability and reliability. By intelligently selecting voice processing strategies, the system can rationally allocate resources, avoiding system crashes or response delays due to insufficient resources. Users do not need to worry about command recognition errors or device malfunctions in multi-user environments, thus improving the user experience. This method enables users to interact with intelligent systems more naturally and conveniently, improving system usability and ease of use. This method is not only applicable to the smart home field but can also be widely used in smart offices, smart vehicle systems, and other fields. With the continuous popularization of smart devices and the continuous development of voice technology, the application prospects of this method will be even broader.
[0084] In summary, this technical solution achieves accurate recognition and processing of voice commands in a multi-user environment by monitoring the number of users in real time and selecting appropriate voice processing strategies, thereby improving the system's recognition accuracy, stability, and user experience.
[0085] One embodiment of the present invention involves real-time monitoring of received voice data from users, and determining a voice processing strategy for multiple voice assistants based on the number of users sending voice data, including:
[0086] S101. Real-time monitoring of received voice data from users;
[0087] S102. Determine the number of users currently sending voice data based on the received voice data sent by the users;
[0088] S103. When the number of users is one user, the first voice processing strategy is retrieved as the voice processing strategy.
[0089] S104. When there are multiple users, the second voice processing strategy is invoked as the voice processing strategy.
[0090] The working principle of the above technical solution is as follows: The system continuously listens to and receives voice data from users. This is typically done through a microphone or other audio input device, ensuring that the system can capture any voice commands issued by the user. The system analyzes the received voice data to determine the number of users currently issuing voice data. This can be achieved through various technologies, such as voiceprint recognition, voice activity detection (VAD), and directional microphone arrays. Voiceprint recognition can distinguish the voice characteristics of different users, while VAD is used to detect the presence or absence of voice signals. Directional microphone arrays can determine the number of users based on the direction of the sound source. Based on the determined number of users, the system selects an appropriate voice processing strategy. If there is only one user, the system invokes the first voice processing strategy. This strategy typically focuses on optimizing the voice recognition accuracy of a single user and may include using more complex voice recognition algorithms or enhancing the quality of the voice signal. If there are multiple users, the system invokes the second voice processing strategy. This strategy requires handling the separation, recognition, and prioritization of multiple voice signals to ensure that each user's command is accurately understood and that the system can execute commands according to priority or chronological order.
[0091] The above technical solution achieves the following results: By selecting an appropriate voice processing strategy based on the number of users, the system can more effectively process voice commands from single or multiple users, thereby improving recognition accuracy. The system can intelligently adjust its processing resources according to the number of users, avoiding resource waste when using a single user while ensuring sufficient processing capacity when multiple users are using it. Users do not need to worry about command confusion or system response delays in multi-user environments, thus improving user experience satisfaction. This technical solution enables the system to adapt to different user environments, maintaining high efficiency and accuracy whether used by one person or multiple people simultaneously. This method is not only suitable for home environments but also for various scenarios such as conference rooms, classrooms, and public places, improving the system's applicability and market competitiveness.
[0092] In summary, this technical solution achieves efficient speech recognition and processing in multi-user environments by monitoring users' voice data in real time and selecting appropriate speech processing strategies based on the number of users, thereby improving the system's accuracy, flexibility, and user experience.
[0093] In one embodiment of the present invention, the speech recognition processing of the first speech processing strategy includes:
[0094] When the number of users currently sending voice data is one user, the voice data is divided into multiple voice segment data.
[0095] The audio quality of multiple speech segments is graded to obtain the audio quality level of each segment.
[0096] Based on the speech recognition quality of multiple voice assistants and the audio quality levels of multiple speech segments, the target voice assistant corresponding to each audio quality level of the speech segment is obtained.
[0097] The speech segments of each audio quality level are assigned to their corresponding target voice assistants for speech recognition processing to obtain the voice commands corresponding to the speech data.
[0098] The working principle of the above technical solution is as follows: When the system detects that only one user is emitting voice data, it initiates a voice segmentation process. This process divides continuous voice data into multiple shorter voice segments, facilitating individual processing of each segment and improving processing efficiency and accuracy. The system evaluates the audio quality of these voice segments to obtain an audio quality level for each segment. Audio quality grading may be based on various factors, such as noise level, signal strength, and clarity. Through grading, the system can identify which segments have good sound quality and which may be affected by interference or have poor quality. Next, the system selects the most suitable target voice assistant for each segment based on the speech recognition quality of multiple available voice assistants and the audio quality level of each voice segment. This process may involve a comprehensive evaluation of the voice assistant's recognition ability, processing speed, stability, etc., to ensure that high-quality voice segments are assigned to voice assistants with stronger recognition capabilities, thereby improving overall recognition accuracy. Finally, the system assigns each audio quality level voice segment to its corresponding target voice assistant for speech recognition processing. The voice assistant uses its own algorithms and models to convert the voice segments into text or commands, thereby obtaining the user's voice instructions.
[0099] The above technical solution achieves the following results: By classifying audio quality and selecting the target voice assistant, the system can assign high-quality audio segments to voice assistants with stronger recognition capabilities, thereby improving the overall recognition accuracy. Simultaneously, for segments with lower quality, the system can select voice assistants with stronger noise resistance to process them, further reducing the false recognition rate. This solution can rationally allocate resources based on the audio quality level of the audio segment and the voice assistant's recognition capability. This not only avoids wasting high-quality audio segments on voice assistants with weaker recognition capabilities but also ensures that each voice assistant can handle the segments most suitable for it, thus improving overall processing efficiency. By improving recognition accuracy and optimizing resource allocation, this solution can provide users with more accurate and efficient voice recognition services. This not only enhances user trust and satisfaction with voice interaction but also promotes the application and development of voice technology in more fields.
[0100] In summary, the first speech processing strategy achieves efficient and accurate processing of speech data through meticulous division, grading, selection, and allocation steps, providing users with higher-quality speech recognition services.
[0101] One embodiment of the present invention involves performing audio quality grading on multiple current speech segment data to obtain the audio quality levels of the multiple speech segments, including:
[0102] Extract audio parameters for each speech segment data, wherein the audio parameters include total harmonic distortion value, dynamic range of spectrum, spectral width, distance between the spectral center and the frequency center, and the ratio between byte audio intensity and background audio intensity;
[0103] The audio quality adjustment coefficient is obtained using the aforementioned spectral dynamic range, spectral width, and distance between the spectral center and the frequency center.
[0104] The audio quality adjustment coefficient is obtained using the following formula:
[0105] Where U represents the audio quality adjustment coefficient; S r Indicates the dynamic range of the spectrum; S w Indicates the bandwidth; S d λ represents the distance between the center of the spectrum and the center of the frequency. 01 and λ 02 These represent the weight values corresponding to the dynamic range and bandwidth of the spectrum, respectively.
[0106] The audio quality evaluation coefficient is obtained by combining the audio quality adjustment coefficient with the total harmonic distortion value and the ratio between the byte audio intensity and the background audio intensity.
[0107] The audio quality evaluation coefficient is obtained using the following formula:
[0108] Where R represents the audio quality evaluation coefficient; T h B represents the total harmonic distortion value; B represents the ratio between the byte audio intensity and the background audio intensity; U represents the audio quality adjustment factor.
[0109] The audio quality evaluation coefficient is compared with a preset first coefficient threshold and a second coefficient threshold;
[0110] When the audio quality evaluation coefficient exceeds a preset first coefficient threshold, the audio segment data with the audio quality evaluation coefficient exceeding the preset first coefficient threshold is determined as a high-quality audio segment.
[0111] If the audio quality evaluation coefficient does not exceed the preset first coefficient threshold, but exceeds the second coefficient threshold, then the audio segment data whose audio quality evaluation coefficient does not exceed the preset first coefficient threshold but exceeds the second coefficient threshold is determined as a medium quality audio segment.
[0112] If the audio quality evaluation coefficient does not exceed the preset second coefficient threshold, then the speech segment data whose audio quality evaluation coefficient does not exceed the preset second coefficient threshold is determined as a low-quality speech segment.
[0113] The working principle of the above technical solution is as follows: Key audio parameters are extracted from each speech segment data. These parameters include total harmonic distortion (THD), dynamic range, bandwidth, distance between the spectral center and frequency center, and the ratio of byte audio intensity to background audio intensity. Using the extracted dynamic range, bandwidth, and (the original text likely misused "Sd," but here it's assumed to be a correctly understood parameter or already included in the evaluation of Sr, so Sr is used instead), an audio quality adjustment coefficient is calculated using a specific formula. In the formula, λ01 and λ02 represent the weight values corresponding to the dynamic range and bandwidth, respectively. These weight values can be adjusted according to the actual application scenario and requirements. The calculated audio quality adjustment coefficient is combined with the THD and the ratio of byte audio intensity to background audio intensity to calculate an audio quality evaluation coefficient using another formula. This coefficient integrates the influence of multiple audio parameters, providing a more comprehensive reflection of the speech segment's audio quality. The calculated audio quality evaluation coefficient is compared with preset first and second coefficient thresholds. Based on the comparison results, the speech segment data is classified as high quality, medium quality, or low quality.
[0114] The above technical solution achieves the following results: by extracting multiple key audio parameters and comprehensively considering their impact, this solution can more accurately evaluate the audio quality of speech segments. This helps to filter out high-quality speech segments in practical applications, improving the accuracy of speech processing or recognition. The weight values in this technical solution can be adjusted according to actual application scenarios and needs. This means that in different application scenarios, the audio quality evaluation process can be optimized according to specific requirements, thereby enhancing the flexibility of audio processing. This technical solution can be applied to audio quality grading of multiple speech segment data, supporting automated and batch processing. This is very useful in practical applications, greatly improving processing efficiency and accuracy. By classifying speech segment data into different quality levels, this technical solution can provide a basis for subsequent audio optimization. For example, for low-quality speech segments, corresponding measures can be taken to improve or repair them; for high-quality speech segments, further compression or encoding can be performed to save storage space or transmission bandwidth.
[0115] On the other hand, by extracting multiple audio parameters, including total harmonic distortion (THD), spectral dynamic range, spectral width, distance between the spectral center and the frequency center, and the ratio of byte audio intensity to background audio intensity, this scheme can comprehensively and accurately evaluate the audio quality of speech segments. These parameters cover the distortion of the audio signal, spectral characteristics, and the contrast between the signal and background noise, thus enabling more precise differentiation of speech segments of different qualities. Utilizing audio quality adjustment coefficients and audio quality evaluation coefficients, this scheme can adaptively classify speech segments based on their actual audio characteristics. By introducing weighted values for parameters such as spectral dynamic range, spectral width, and distance between the spectral center and the frequency center, this scheme can flexibly adjust the classification criteria to adapt to the needs of different application scenarios. By comparing the audio quality evaluation coefficient with preset first and second coefficient thresholds, this scheme can efficiently classify speech segments into three levels: high quality, medium quality, and low quality. This classification method is simple, easy to implement, and ensures the accuracy of the classification. This technical solution can quickly classify the audio quality of a large number of speech segments, thereby improving the efficiency of speech processing. This is of great significance for scenarios that require processing large amounts of voice data, such as speech recognition and speech synthesis. By classifying speech segments into different quality levels, this solution can guide subsequent processing flows, such as prioritizing high-quality speech segments or employing different processing strategies for speech segments of different quality levels. This helps optimize resource utilization and improve overall processing efficiency. In voice interaction applications, high-quality speech segments can provide clearer and more accurate speech recognition results, thereby enhancing the user experience. This technical solution, through accurate audio quality grading, helps to filter out high-quality speech segments, improving the accuracy and fluency of speech recognition.
[0116] In summary, the technical benefits of the above-mentioned solution in terms of performance indicators are mainly reflected in multi-dimensional audio quality assessment, adaptive audio quality grading, efficient and accurate grading determination, improved speech processing efficiency, optimized resource utilization, and enhanced user experience. These technical benefits make this solution have broad application prospects and significant practical value in the field of speech processing. Furthermore, this technical solution achieves audio quality grading for multiple speech segments by extracting audio parameters, calculating audio quality adjustment coefficients and audio quality evaluation coefficients, and determining audio quality levels. This technical solution demonstrates significant technical benefits in improving the accuracy of audio quality assessment, enhancing the flexibility of audio processing, supporting automated and batch processing, and providing a basis for audio optimization.
[0117] One embodiment of the present invention involves obtaining the target voice assistant corresponding to each audio quality level of a speech segment based on the speech recognition quality of multiple current voice assistants combined with the audio quality levels of multiple speech segments, including:
[0118] Extract the recognition accuracy and recognition processing time for each voice assistant;
[0119] The recognition evaluation coefficient is obtained based on the recognition accuracy and recognition processing time corresponding to each voice assistant;
[0120] The identification rating coefficient is obtained using the following formula:
[0121] Where J represents the recognition rating coefficient; n represents the number of times each voice assistant completes recognition; T i P represents the processing time for the i-th speech recognition attempt by each voice assistant; zi C represents the recognition accuracy for the i-th speech recognition attempt by each voice assistant; i C represents the amount of voice data corresponding to the i-th speech recognition of each voice assistant; d Indicates the preset unit voice data volume; T c This represents the theoretical recognition and processing time for each unit of voice data corresponding to each voice assistant;
[0122] The voice assistant corresponding to the maximum value of the recognition evaluation coefficient is used as the target voice assistant for the low-quality voice segment.
[0123] The voice assistant whose recognition rating coefficient is second only to the maximum recognition rating coefficient is selected as the target voice assistant for the medium-quality voice segment.
[0124] The voice assistant whose recognition rating coefficient is second only to that of the medium-quality speech segment is considered as the target voice assistant for the high-quality speech segment.
[0125] The working principle of the above technical solution is as follows: The recognition accuracy and processing time of each voice assistant are extracted. These two indicators are key parameters for measuring the performance of a voice assistant. Recognition accuracy reflects the voice assistant's ability to correctly recognize voice commands, while processing time reflects its processing speed. Based on the recognition accuracy, processing time, and corresponding voice data volume of each voice assistant, a recognition evaluation coefficient is calculated using the formula described above. This formula comprehensively considers the impact of recognition accuracy, processing speed, and voice data volume on the performance of the voice assistant. The theoretical recognition processing time per unit volume of voice data serves as a benchmark to evaluate the efficiency of the voice assistant when processing different amounts of data. Based on the calculated recognition evaluation coefficient, a target voice assistant is assigned to each audio quality level of the speech segment. Specifically, the voice assistant with the highest recognition evaluation coefficient is assigned to the low-quality speech segment because low-quality speech segments are more difficult to recognize and require a voice assistant with stronger recognition capabilities. Next, the voice assistant with the second-highest recognition evaluation coefficient is assigned to the medium-quality speech segment, and so on, until a target voice assistant is assigned to the high-quality speech segment. This allocation method ensures that each audio quality level of the speech segment receives the most suitable voice assistant for processing.
[0126] The above technical solution achieves the following results: By assigning the most suitable target voice assistant to each audio quality level of speech segment, this solution fully leverages the performance advantages of different voice assistants, thereby improving the overall accuracy of speech recognition. The solution intelligently matches voice assistant performance metrics with the audio quality level of speech segments, avoiding resource waste and unreasonable allocation. This helps improve resource utilization and processing efficiency in practical applications. The solution adapts to speech segments of different quality levels and voice assistants with varying performance, enhancing the system's flexibility and adaptability. This allows the system to maintain efficient and stable operation under different application scenarios and conditions. By improving speech recognition accuracy and optimizing resource allocation, this solution provides users with a smoother and more accurate voice interaction experience. This helps increase user trust and satisfaction with voice assistants, thereby promoting the development and application of voice technology.
[0127] On the other hand, by combining the recognition accuracy and processing time of multiple voice assistants, as well as the corresponding audio quality level of the speech segments, this scheme can more accurately match speech segments of different qualities to the most suitable voice assistant. This helps improve the overall accuracy and efficiency of speech recognition. A recognition evaluation coefficient (J) is introduced, which comprehensively considers the recognition accuracy (P...). zi ), recognition processing time (T) i ) and voice data volume (C i Theoretical recognition and processing time per unit volume of speech data (T) cThis approach allows for a more comprehensive evaluation of voice assistant performance by comparing different performance metrics. This integrated evaluation helps select voice assistants with faster processing speeds and higher accuracy, improving recognition efficiency. The scheme assigns voice assistants to different quality levels of speech segments based on their recognition rating coefficients. Low-quality speech segments are processed by the voice assistant with the highest rating coefficient, medium-quality segments by the next highest, and high-quality segments by the next lowest. This allocation strategy ensures that even with low-quality speech input, the most accurate recognition results can be obtained, while maintaining efficient processing of high-quality speech segments. By considering the performance metrics of multiple voice assistants and dynamically allocating them according to the quality of the actual speech segments, this scheme enhances the robustness of the system. Even if a voice assistant performs poorly in certain situations, other higher-performing voice assistants can promptly compensate, ensuring the overall stability and reliability of the system. This scheme provides data support for the continuous optimization of voice assistant technology by quantitatively evaluating the performance of different voice assistants. Developers can use feedback from the recognition rating coefficients to make targeted improvements and optimizations to the voice assistants, further enhancing their recognition accuracy and processing efficiency.
[0128] In summary, the technical benefits of this solution in terms of performance metrics are mainly reflected in optimizing speech recognition quality matching, improving recognition efficiency, flexibly handling speech segments of different quality levels, enhancing system robustness, and promoting continuous optimization of voice assistant technology. Furthermore, by comprehensively considering the performance metrics of the voice assistant and the audio quality level of the speech segments, this solution assigns the most suitable target voice assistant to each audio quality level of the speech segment. This solution demonstrates significant technical effectiveness in improving speech recognition accuracy, optimizing resource allocation, enhancing system adaptability, and improving user experience.
[0129] In one embodiment of the present invention, the speech recognition processing of the second speech processing strategy includes:
[0130] When the number of users currently sending voice data is multiple users, and the number of multiple users is less than or equal to the number of voice assistants, then the users are assigned voice assistants according to the audio intensity ratio, and the assigned voice assistants are used to perform voice recognition processing on the voice data of their corresponding users.
[0131] When the number of users currently sending voice data is multiple users, and the number of multiple users is higher than the number of voice assistants, then the multiple voice assistants are controlled to collaboratively perform voice recognition processing through user tags.
[0132] The working principle of the above technical solution is as follows: First, the system determines the number of users currently sending voice data. Then, the system compares the number of users with the number of available voice assistants.
[0133] Voice assistant allocation strategy:
[0134] When the number of users is less than or equal to the number of voice assistants:
[0135] The system sorts or categorizes users based on their audio intensity ratio.
[0136] Then, a corresponding voice assistant is assigned to each user. The assigned voice assistant is responsible for processing the user's voice data and performing speech recognition.
[0137] When the number of users is greater than the number of voice assistants:
[0138] The system cannot assign a voice assistant to each user individually.
[0139] At this point, the system uses user tags (such as user ID, session ID, etc.) to distinguish and track the voice data of different users.
[0140] Multiple voice assistants work together to process the voice data from all users. The system attempts to separate the commands from different users from the mixed voice data and perform accurate speech recognition.
[0141] Speech recognition processing:
[0142] Regardless of the allocation strategy used, the assigned voice assistant or the voice assistant working in collaboration with it will process the user's voice data. This includes steps such as voice signal acquisition, preprocessing, feature extraction, model matching, and decoding output, ultimately converting the voice data into text or commands.
[0143] The above technical solution achieves the following results: By intelligently allocating voice assistants, it ensures efficient resource utilization. When the number of users is small, each user receives a dedicated voice assistant and enjoys high-quality service. When the number of users is large, multiple voice assistants can work together to handle all user requests, avoiding resource idleness and waste. For individual users, a dedicated voice assistant can respond to their requests faster and provide more personalized service. In multi-user scenarios, through collaborative processing and the use of user tags, the system can accurately identify each user's commands, reducing misidentification and conflicts, thereby improving the overall user experience. This solution can dynamically adjust resource allocation strategies according to different user numbers, enhancing the system's flexibility and adaptability. This allows the system to maintain efficient and stable operation under different application scenarios and conditions, meeting the needs of different users. This technical solution demonstrates the application potential and challenges of voice technology in multi-user scenarios. Through continuous optimization and improvement, this solution is expected to promote the development and innovation of voice technology, providing more users with smarter and more convenient services.
[0144] In summary, this technical solution achieves efficient resource utilization, improved user experience, enhanced system flexibility, and promoted development of voice technology by intelligently allocating and utilizing voice assistants for speech recognition processing.
[0145] One embodiment of the present invention involves assigning a voice assistant to a user based on the audio intensity ratio, including:
[0146] When there are multiple users currently sending voice data, the voice data of multiple users is split to obtain the voice data corresponding to each user;
[0147] Extract the audio intensity corresponding to the voice data of each user;
[0148] The audio intensity ratio is obtained by comparing the audio intensity of each user's voice data with the overall audio intensity of multiple users.
[0149] The first audio recognition difficulty coefficient is obtained by combining the audio intensity ratio with the byte frequency corresponding to each user's voice data.
[0150] The difficulty coefficient of the first audio recognition is obtained by the following formula:
[0151] Among them, J 01 P represents the difficulty level of the first audio recognition; d01 This represents the audio intensity ratio for each user; f represents the byte frequency of each user's voice data; f c This indicates the preset byte frequency reference value;
[0152] The voice assistant corresponding to the highest recognition evaluation coefficient is taken as the target voice assistant of the user corresponding to the highest first audio recognition difficulty coefficient.
[0153] According to the allocation strategy of matching the recognition evaluation coefficient from low to high based on the first audio recognition difficulty coefficient, users other than those with the highest first audio recognition difficulty coefficient are matched with the remaining voice assistants in turn.
[0154] The working principle of the above technical solution is as follows: When multiple users simultaneously transmit voice data, the system first splits this voice data to ensure that independent voice data corresponding to each user can be obtained. Next, the system extracts the audio intensity of each user's voice data. Audio intensity is an indicator of the strength of a voice signal, reflecting its energy level. Then, the system compares each user's audio intensity with the overall audio intensity of all users to calculate the audio intensity ratio for each user. This ratio reflects the proportion of each user's voice signal in the overall voice signal. Next, the system uses the audio intensity ratio and the byte frequency of each user's voice data to calculate a first audio recognition difficulty coefficient. This coefficient is a comprehensive indicator that considers the strength and frequency characteristics of the voice signal to assess the difficulty of voice recognition. Specifically, in the coefficient calculation formula, Pd01 represents the audio intensity ratio, f represents the byte frequency, and fc represents a preset byte frequency reference value. Through the combination and calculation of these parameters, a numerical value reflecting the difficulty of voice recognition can be obtained. Finally, the system assigns a voice assistant based on the magnitude of the first audio recognition difficulty coefficient. The voice assistant with the highest recognition rating coefficient will be assigned to the user with the highest first audio recognition difficulty coefficient, i.e., the most difficult user to recognize. Then, following the strategy of matching the first audio recognition difficulty coefficient from high to low and the recognition rating coefficient from low to high, the other users will be matched with the remaining voice assistants in turn.
[0155] The above technical solution achieves the following results: By comprehensively considering factors such as audio intensity and byte frequency, this solution can more accurately assess the difficulty of speech recognition, thereby assigning users a more suitable voice assistant. This helps improve the accuracy of speech recognition and reduce the false recognition rate. The solution can dynamically allocate voice assistant resources based on the user's speech recognition difficulty. For users with higher speech recognition difficulty, the system allocates more resources to ensure recognition accuracy; while for users with lower difficulty, fewer resources are allocated, thus optimizing resource allocation. By assigning users a more suitable voice assistant, this solution enhances the user experience. Users can obtain more accurate and timely speech recognition services, thereby increasing their satisfaction and trust in the intelligent voice assistant. This solution considers the impact of multiple factors on speech recognition difficulty, including audio intensity and byte frequency. This enables the system to exhibit stronger robustness and adaptability when facing speech data under different environments and conditions.
[0156] On the other hand, by splitting the voice data of multiple users and extracting the audio intensity separately, this scheme can accurately capture the voice features of each user. By comprehensively considering the audio intensity ratio and byte frequency, the recognition difficulty of each user's voice can be more accurately assessed, thus assigning the most suitable voice assistant to each user. This precise allocation helps improve the accuracy and efficiency of voice recognition. The scheme achieves a reasonable allocation of voice assistant resources by calculating a first audio recognition difficulty coefficient and combining it with a recognition evaluation coefficient. Voice assistants with higher recognition evaluation coefficients are assigned to more difficult user voices, while voice assistants with lower recognition evaluation coefficients handle relatively simple tasks. This allocation strategy not only improves the utilization rate of voice assistants but also ensures the performance optimization of the overall voice recognition system. The scheme can flexibly handle situations where multiple users simultaneously send voice data. By splitting and comparing the voice data of each user, the scheme can dynamically adjust the allocation of voice assistants to adapt to the recognition needs of different users' voices. This dynamic adaptability makes the scheme perform excellently in multi-user environments. By assigning the most suitable voice assistant to each user, the scheme can reduce voice recognition errors, improve recognition speed, and thus optimize the user experience. Users can interact with voice assistants more quickly and accurately, improving overall satisfaction. The parameters in this solution, such as the audio intensity ratio and byte frequency reference value, can be adjusted and optimized according to actual needs. This scalability and customizability enable the solution to adapt to changes in different scenarios and user needs, providing greater flexibility and adaptability. By quantitatively evaluating the difficulty of speech recognition and the performance of voice assistants, this solution provides data support for technological iteration. Developers can continuously optimize and improve the speech recognition algorithm and voice assistant based on the evaluation results to improve the overall system performance and accuracy.
[0157] In summary, the technical benefits of this solution in terms of performance metrics are mainly reflected in accurate user voice allocation, efficient resource utilization, dynamic adaptation to multi-user environments, optimized user experience, scalability and customizability, and promotion of technological iteration. These technical benefits collectively improve the overall performance and user experience of the speech recognition system. Furthermore, by comprehensively considering factors such as audio intensity and byte frequency to allocate voice assistants to users, this solution achieves optimized resource allocation and improved speech recognition accuracy.
[0158] One embodiment of the present invention controls multiple voice assistants to collaboratively perform speech recognition processing based on user tags, including:
[0159] Each user's voice data is uniquely identified, so that each user's voice data has a unique identifier that is uniquely associated with the user;
[0160] The voice data of all users is divided into multiple voice segment data, and each voice segment data is assigned a unique identifier that is consistent with the voice data.
[0161] By utilizing the audio intensity ratio corresponding to each speech segment data and the byte audio intensity variation amplitude in each speech segment, the second audio recognition difficulty coefficient corresponding to each speech segment data is obtained.
[0162] The difficulty coefficient for the second audio recognition is obtained using the following formula:
[0163] Among them, J 02 Indicates the difficulty level of the second audio recognition; P d02 Q represents the audio intensity ratio corresponding to each speech segment data; k represents the number of bytes contained in each speech segment data; Q i Q represents the audio intensity corresponding to the i-th byte; i+1 Q represents the audio intensity corresponding to the (i+1)th byte; b This represents the standard deviation of audio intensity corresponding to k bytes;
[0164] The second audio recognition difficulty coefficient is compared with a preset recognition difficulty coefficient threshold;
[0165] The speech segment data whose second audio recognition difficulty coefficient exceeds the preset recognition difficulty coefficient threshold is assigned to the voice assistant corresponding to the maximum recognition evaluation coefficient for speech recognition processing, the speech recognition result corresponding to each speech segment data is obtained, and the speech recognition result is assigned a unique identifier consistent with the speech segment data.
[0166] The speech segment data whose second audio recognition difficulty coefficient does not exceed the preset recognition difficulty coefficient threshold are allocated to the remaining voice assistants for speech recognition processing according to the principle of near-uniform distribution, the speech recognition result corresponding to each speech segment data is obtained, and the speech recognition result is assigned a unique identifier consistent with the speech segment data.
[0167] The speech results are filtered according to the identifier corresponding to each speech segment data, and the speech results with the same unique identifier are integrated to obtain the recognition result corresponding to the speech data uniquely associated with the user.
[0168] The working principle of the above technical solution is as follows: Each user's voice data is uniquely identified, ensuring that each user's voice data has a unique identifier associated with it. This helps distinguish voice data from different users in subsequent processing. All users' voice data is divided into multiple voice segment data, and each voice segment data is assigned a unique identifier consistent with the original voice data. In this way, each voice segment data can be traced back to its original user. A second audio recognition difficulty coefficient is calculated using the audio intensity ratio and byte audio intensity variation amplitude of each voice segment data. This coefficient takes into account the audio intensity distribution and variation of the voice segment to assess its recognition difficulty.
[0169] In the specific calculation, the audio intensity of each byte and the standard deviation of the audio intensity of all bytes were considered, along with the number of bytes contained in the speech segment. The calculated second audio recognition difficulty coefficient was compared with a preset recognition difficulty coefficient threshold to distinguish between high-difficulty and low-difficulty speech segment data. For speech segment data with a difficulty coefficient exceeding the threshold, it was assigned to the voice assistant with the highest recognition evaluation coefficient for speech recognition processing. This ensures that high-difficulty speech segments receive more professional processing. For speech segment data with a difficulty coefficient not exceeding the threshold, it was assigned to the remaining voice assistants for speech recognition processing according to the principle of near-uniform distribution. During the recognition process, a unique identifier consistent with the speech segment data was assigned to the speech recognition result of each speech segment data for subsequent integration. Finally, the speech recognition results were filtered and integrated according to the unique identifier corresponding to each speech segment data to obtain the final recognition result corresponding to the speech data uniquely associated with the user.
[0170] The effects of the above technical solution are as follows: By calculating a second audio recognition difficulty coefficient, this solution can more accurately assess the recognition difficulty of speech segments and allocate voice assistant resources accordingly. This helps ensure that highly difficult speech segments receive more professional processing, thereby improving the overall accuracy of speech recognition. The solution can dynamically adjust the voice assistant's allocation strategy based on the recognition difficulty of speech segments, achieving optimized resource allocation. More resources are allocated to more difficult speech segments, while fewer resources are allocated to easier ones. This helps improve resource utilization efficiency. The solution can handle situations where multiple users simultaneously send voice data and maintain the correlation of voice data through unique identifiers. This enhances the system's flexibility, enabling it to function normally in complex multi-user environments. Through more accurate speech recognition and optimized resource allocation, this solution can provide users with faster and more accurate speech recognition services. This helps improve user satisfaction and trust in the intelligent voice assistant.
[0171] On the other hand, by assigning a unique identifier to each user's voice data and segmenting and recognizing voice segments based on these identifiers, this scheme ensures efficient collaboration among multiple voice assistants. Each voice segment is accurately assigned to a voice assistant for processing, avoiding duplication of work and waste of resources. The scheme introduces a second audio recognition difficulty coefficient, which comprehensively considers multiple factors such as the audio intensity ratio of the voice segment, the amplitude of byte audio intensity variation, and the standard deviation of audio intensity. This comprehensive evaluation method can more accurately reflect the recognition difficulty of voice segments, thus helping to assign more difficult voice segments to higher-performing voice assistants. Based on the comparison between the second audio recognition difficulty coefficient and a preset recognition difficulty coefficient threshold, the scheme can intelligently assign voice segments to different voice assistants for processing. For more difficult voice segments, they are assigned to the voice assistant with the highest recognition evaluation coefficient to ensure the accuracy of the recognition results; for less difficult voice segments, they are assigned according to a principle of near-uniform distribution to fully utilize the processing capabilities of all voice assistants. Through precise difficulty assessment and optimized resource allocation, this scheme can significantly improve the accuracy and efficiency of speech recognition. More complex speech segments are handled by higher-performance voice assistants, reducing the possibility of recognition errors; while less complex speech segments are processed by multiple voice assistants, accelerating the overall processing speed. Throughout the speech recognition process, this solution maintains data consistency. Each speech segment and its corresponding recognition result is assigned the same unique identifier as the original speech data, making subsequent data integration and result analysis simpler and more accurate. Because this solution can efficiently process speech data from multiple users and provide accurate recognition results, it significantly improves the user experience. Users can interact with voice assistants more quickly and accurately, thereby increasing overall satisfaction and loyalty.
[0172] In summary, the technical benefits of this solution are mainly reflected in efficient collaborative processing, accurate assessment of recognition difficulty, optimized resource allocation, improved recognition accuracy and efficiency, maintenance of data consistency, and enhanced user experience. These benefits collectively drive the further development and application of speech recognition technology. Furthermore, by comprehensively considering factors such as the audio intensity ratio of speech segments and the amplitude of byte audio intensity variations, this solution achieves more accurate speech recognition difficulty assessment and optimized resource allocation. This helps improve the accuracy of speech recognition and resource utilization efficiency, thereby enhancing the user experience.
[0173] One embodiment of the present invention includes intelligent control of a target system according to the voice command, comprising:
[0174] S301. When the number of users is one, the target system is intelligently controlled according to the voice result recognized by the voice assistant.
[0175] S301. When there are multiple users, it is determined whether the voice commands corresponding to the multiple users are for the same controlled target parameter. If there are control commands for the same controlled target parameter among the voice commands corresponding to the multiple users, the voice commands of the user with higher permission are selected according to the user permission priority to intelligently control the target system.
[0176] The working principle of the above technical solution is as follows: When the system detects only one user, that user's voice commands will directly serve as the basis for controlling the target system. This means the system will directly parse the user's voice output and perform corresponding intelligent control on the target system based on the parsing results. When the system detects multiple users, the situation becomes more complex. In this case, the system will first analyze each user's voice commands to determine whether they are directed at the same controlled target parameter. For example, if two users issue commands such as "increase the temperature" and "turn off the air conditioner," and these commands are both directed at the air conditioner as the controlled target, the system needs further processing. If multiple users' voice commands contain control commands for the same controlled target parameter, the system will enter the user permission priority judgment stage. At this stage, the system will select a user with priority based on preset user permission rules (such as administrator privileges being higher than ordinary user privileges). Then, the system will perform intelligent control on the target system based on that user's voice commands.
[0177] The above technical solution offers the following advantages: It automatically adjusts control strategies based on the number of users and the content of their commands, thereby improving control flexibility. Whether dealing with a single user or multiple users, the system provides appropriate responses. In a multi-user environment, the system accurately determines the intent of user commands and makes decisions based on user permission priorities. This helps avoid system chaos caused by conflicting user commands, enhancing system robustness. By intelligently judging user commands and permission priorities, this solution provides users with a more accurate and timely control experience. Users do not need to worry about their commands being ignored or misunderstood, thus increasing user satisfaction with the system. In multi-user scenarios, the system can rationally allocate resources based on user permission priorities. This helps ensure that commands from important users are processed first, thereby improving resource utilization efficiency.
[0178] In summary, this technical solution achieves intelligent control of the target system by comprehensively considering the number of users and the content of user voice commands. This improves control flexibility, enhances system robustness, improves user experience, and enables rational allocation of resources.
[0179] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
Claims
1. A smart recognition and control method based on multiple voice assistants, characterized in that, The intelligent recognition and control method based on multiple voice assistants includes: The system monitors received voice data from users in real time and determines a voice processing strategy for multiple voice assistants based on the number of users sending voice data. The voice processing strategy includes a first voice processing strategy and a second voice processing strategy. The first voice processing strategy refers to a voice processing strategy in which multiple voice assistants collaboratively process the voice data of one user. The second voice processing strategy refers to a voice processing strategy in which multiple voice assistants collaboratively process the voice data of multiple users. The voice data emitted by the user is processed by voice recognition according to the voice processing strategy to obtain the voice command corresponding to the voice data. The target system is intelligently controlled according to the voice commands.
2. The intelligent recognition and control method based on multiple voice assistants according to claim 1, characterized in that, Real-time monitoring of received voice data from users; determination of voice processing strategies for multiple voice assistants based on the number of users sending voice data, including: Real-time monitoring of received voice data from users; The number of users currently sending voice data is determined based on the received voice data sent by the users. When the number of users is one, the first voice processing strategy is invoked as the voice processing strategy. When there are multiple users, the second voice processing strategy is invoked as the voice processing strategy.
3. The intelligent recognition and control method based on multiple voice assistants according to claim 1, characterized in that, The step of performing speech recognition processing on the voice data emitted by the user according to the speech processing strategy to obtain the voice command corresponding to the voice data includes: The voice data emitted by the user is processed by voice recognition according to the first voice processing strategy to obtain the voice command corresponding to the voice data. or The user's voice data is processed by voice recognition according to the second voice processing strategy to obtain the voice command corresponding to the voice data.
4. The intelligent recognition and control method based on multiple voice assistants according to claim 2, characterized in that, The speech recognition process of the first speech processing strategy includes: When the number of users currently sending voice data is one user, the voice data is divided into multiple voice segment data. The audio quality of multiple speech segments is graded to obtain the audio quality level of each segment. Based on the speech recognition quality of multiple voice assistants and the audio quality levels of multiple speech segments, the target voice assistant corresponding to each audio quality level of the speech segment is obtained. The speech segments of each audio quality level are assigned to their corresponding target voice assistants for speech recognition processing to obtain the voice commands corresponding to the speech data.
5. The intelligent recognition and control method based on multiple voice assistants according to claim 4, characterized in that, The audio quality of multiple speech segments is graded to obtain their respective audio quality levels, including: Extract audio parameters for each speech segment data, wherein the audio parameters include total harmonic distortion value, dynamic range of spectrum, spectral width, distance between the spectral center and the frequency center, and the ratio between byte audio intensity and background audio intensity; The audio quality adjustment coefficient is obtained using the aforementioned spectral dynamic range, spectral width, and distance between the spectral center and the frequency center. The audio quality adjustment coefficient is obtained using the following formula: Where U represents the audio quality adjustment coefficient; S r Indicates the dynamic range of the spectrum; S w Indicates the bandwidth; S d λ represents the distance between the center of the spectrum and the center of the frequency. 01 and λ 02 These represent the weight values corresponding to the dynamic range and bandwidth of the spectrum, respectively. The audio quality evaluation coefficient is obtained by combining the audio quality adjustment coefficient with the total harmonic distortion value and the ratio between the byte audio intensity and the background audio intensity. The audio quality evaluation coefficient is obtained using the following formula: Where R represents the audio quality evaluation coefficient; T h B represents the total harmonic distortion value; B represents the ratio between the byte audio intensity and the background audio intensity; U represents the audio quality adjustment factor. The audio quality evaluation coefficient is compared with a preset first coefficient threshold and a second coefficient threshold; When the audio quality evaluation coefficient exceeds a preset first coefficient threshold, the audio segment data with the audio quality evaluation coefficient exceeding the preset first coefficient threshold is determined as a high-quality audio segment. If the audio quality evaluation coefficient does not exceed the preset first coefficient threshold, but exceeds the second coefficient threshold, then the audio segment data whose audio quality evaluation coefficient does not exceed the preset first coefficient threshold but exceeds the second coefficient threshold is determined as a medium quality audio segment. If the audio quality evaluation coefficient does not exceed the preset second coefficient threshold, then the speech segment data whose audio quality evaluation coefficient does not exceed the preset second coefficient threshold is determined as a low-quality speech segment.
6. The intelligent recognition and control method based on multiple voice assistants according to claim 4, characterized in that, Based on the current speech recognition quality of multiple voice assistants and the audio quality levels of multiple speech segments, the target voice assistant corresponding to each audio quality level of the speech segment is obtained, including: Extract the recognition accuracy and recognition processing time for each voice assistant; The recognition evaluation coefficient is obtained based on the recognition accuracy and recognition processing time corresponding to each voice assistant; The identification rating coefficient is obtained using the following formula: Where J represents the recognition rating coefficient; n represents the number of times each voice assistant completes recognition; T i P represents the processing time for the i-th speech recognition attempt by each voice assistant; zi C represents the recognition accuracy for the i-th speech recognition attempt by each voice assistant; i C represents the amount of voice data corresponding to the i-th speech recognition of each voice assistant; d Indicates the preset unit voice data volume; T c This represents the theoretical recognition and processing time for each unit of voice data corresponding to each voice assistant; The voice assistant corresponding to the maximum value of the recognition evaluation coefficient is used as the target voice assistant for the low-quality voice segment. The voice assistant whose recognition rating coefficient is second only to the maximum recognition rating coefficient is selected as the target voice assistant for the medium-quality voice segment. The voice assistant whose recognition rating coefficient is second only to that of the medium-quality speech segment is considered as the target voice assistant for the high-quality speech segment.
7. The intelligent recognition and control method based on multiple voice assistants according to claim 2, characterized in that, The speech recognition process of the second speech processing strategy includes: When the number of users currently sending voice data is multiple users, and the number of multiple users is less than or equal to the number of voice assistants, then the users are assigned voice assistants according to the audio intensity ratio, and the assigned voice assistants are used to perform voice recognition processing on the voice data of their corresponding users. When the number of users currently sending voice data is multiple users, and the number of multiple users is higher than the number of voice assistants, then the multiple voice assistants are controlled to collaboratively perform voice recognition processing through user tags.
8. The intelligent recognition and control method based on multiple voice assistants according to claim 7, characterized in that, Voice assistants are assigned to users based on audio intensity ratios, including: When there are multiple users currently sending voice data, the voice data of multiple users is split to obtain the voice data corresponding to each user; Extract the audio intensity corresponding to the voice data of each user; The audio intensity ratio is obtained by comparing the audio intensity of each user's voice data with the overall audio intensity of multiple users. The first audio recognition difficulty coefficient is obtained by combining the audio intensity ratio with the byte frequency corresponding to each user's voice data. The difficulty coefficient of the first audio recognition is obtained by the following formula: Among them, J 01 P represents the difficulty level of the first audio recognition; d01 This represents the audio intensity ratio for each user; f represents the byte frequency of each user's voice data; f c This indicates the preset byte frequency reference value; The voice assistant corresponding to the highest recognition evaluation coefficient is taken as the target voice assistant of the user corresponding to the highest first audio recognition difficulty coefficient. According to the allocation strategy of matching the recognition evaluation coefficient from low to high based on the first audio recognition difficulty coefficient, users other than those with the highest first audio recognition difficulty coefficient are matched with the remaining voice assistants in turn.
9. The intelligent recognition and control method based on multiple voice assistants according to claim 7, characterized in that, Controlling multiple voice assistants to collaboratively perform speech recognition processing based on user tags includes: Each user's voice data is uniquely identified, so that each user's voice data has a unique identifier that is uniquely associated with the user; The voice data of all users is divided into multiple voice segment data, and each voice segment data is assigned a unique identifier that is consistent with the voice data. By utilizing the audio intensity ratio corresponding to each speech segment data and the byte audio intensity variation amplitude in each speech segment, the second audio recognition difficulty coefficient corresponding to each speech segment data is obtained. The difficulty coefficient for the second audio recognition is obtained using the following formula: Among them, J 02 Indicates the difficulty level of the second audio recognition; P d02 Q represents the audio intensity ratio corresponding to each speech segment data; k represents the number of bytes contained in each speech segment data; Q i Q represents the audio intensity corresponding to the i-th byte; i+1 Q represents the audio intensity corresponding to the (i+1)th byte; b This represents the standard deviation of audio intensity corresponding to k bytes; The second audio recognition difficulty coefficient is compared with a preset recognition difficulty coefficient threshold; The speech segment data whose second audio recognition difficulty coefficient exceeds the preset recognition difficulty coefficient threshold is assigned to the voice assistant corresponding to the maximum recognition evaluation coefficient for speech recognition processing, the speech recognition result corresponding to each speech segment data is obtained, and the speech recognition result is assigned a unique identifier consistent with the speech segment data. The speech segment data whose second audio recognition difficulty coefficient does not exceed the preset recognition difficulty coefficient threshold are allocated to the remaining voice assistants for speech recognition processing according to the principle of near-uniform distribution, the speech recognition result corresponding to each speech segment data is obtained, and the speech recognition result is assigned a unique identifier consistent with the speech segment data. The speech results are filtered according to the identifier corresponding to each speech segment data, and the speech results with the same unique identifier are integrated to obtain the recognition result corresponding to the speech data uniquely associated with the user.
10. The intelligent recognition and control method based on multiple voice assistants according to claim 1, characterized in that, Intelligent control of the target system according to the voice commands includes: When the number of users is one, the target system is intelligently controlled according to the voice result recognized by the voice assistant; When there are multiple users, it is determined whether the voice commands corresponding to the multiple users are for the same controlled target parameter; when there are control commands for the same controlled target parameter among the voice commands corresponding to the multiple users, the voice commands of the user with higher permission are selected according to the user permission priority to intelligently control the target system.