An intention recognition method, device and equipment based on multi-agent, and a medium

By preprocessing and extracting features from the original speech using a multi-agent system, combined with quantum state analysis, the accuracy problem of single-agent intention recognition is solved, achieving more efficient and accurate intention recognition.

CN121034296BActive Publication Date: 2026-04-24湖南工商大学
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
湖南工商大学
Filing Date
2025-10-30
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

In existing technologies, single intelligent agents often fail to accurately or effectively recognize intentions, especially when dealing with complex semantics and ambiguous speech, making it difficult to improve accuracy.

Method used

A multi-agent system is employed to preprocess, enhance, and extract features from the original speech. By combining Mel frequency cepstral coefficients and quantum state analysis, multiple agents process the target information and determine the ground state with the highest probability of quantum state collapse to identify the intent.

Benefits of technology

It improves the accuracy and efficiency of intent recognition, reduces the time required to obtain intent recognition results, and enhances the ability to analyze complex contexts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121034296B_ABST
    Figure CN121034296B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of natural language processing, the technical field of multi-agent and the technical field of intelligent speech recognition, and discloses a multi-agent-based intention recognition method, device, equipment and medium, which comprises the following steps: obtaining a frame-level vector of each frame of speech signal; splicing the frame-level vector of each frame of speech signal to obtain a sequence tensor of original speech; fusing the sequence tensor of original speech, a model instruction, a preset label and external context to obtain target information; processing the target information through multiple agents to obtain multiple real vectors, a confidence of each real vector, an amplitude and a phase angle of each real vector; determining a quantum state based on the multiple real vectors, the confidence of each real vector, the amplitude and the phase angle of each real vector; determining a probability of quantum state collapse to each ground state; and selecting a semantic category corresponding to a ground state with the largest probability as an intention recognition result of the original speech. The application can effectively improve the accuracy of intention recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of natural language processing, multi-agent technology, and intelligent speech recognition, and in particular to a method, apparatus, device, and medium for intent recognition based on multi-agent technology. Background Technology

[0002] With the rapid development of artificial intelligence technology, speech recognition and natural language understanding systems have been widely used in fields such as intelligent interaction, in-vehicle navigation, and intelligent customer service. Among them, intent recognition, as a core link in the speech understanding chain, can convert user language input into executable commands for the system.

[0003] However, existing technologies often rely on a single agent to perform intent recognition on raw speech. This method of intent recognition depends on a single agent, which has limited model capabilities and struggles to handle complex semantics and ambiguous speech. Furthermore, the training data for a single agent is insufficient to fully cover various real-world application scenarios. Therefore, when a single agent performs intent recognition on raw speech, it may result in inaccurate or ineffective recognition, which is detrimental to improving the accuracy of intent recognition on raw speech. Summary of the Invention

[0004] This application provides a method, apparatus, device, and medium for intent recognition based on multiple agents, in order to solve the technical problem that when a single agent performs intent recognition on raw speech, inaccurate or ineffective recognition may occur, which is detrimental to improving the accuracy of intent recognition of raw speech.

[0005] In a first aspect, embodiments of this application provide a multi-agent-based intent recognition method applied to an electronic device, the intent recognition method comprising:

[0006] Obtain the original speech, delete invalid segments from the original speech, and obtain the updated original speech;

[0007] The high-frequency signals in the updated original speech are enhanced to obtain the enhanced original speech.

[0008] The enhanced original speech is divided into multiple frames of speech signals. The Mel frequency cepstral coefficients of each frame of speech signal are obtained. The Mel frequency cepstral coefficients, the first sequence, and the second sequence of each frame of speech signal are concatenated to obtain the frame-level vector of each frame of speech signal.

[0009] The frame-level vectors of each frame of speech signal are concatenated to obtain the sequence tensor of the original speech. The sequence tensor of the original speech, model instructions, preset labels, and external context are fused to obtain the target information. The target information is processed by multiple agents to obtain multiple real vectors, the confidence of each real vector, and the magnitude and phase angle of each real vector.

[0010] Based on multiple real vectors, the confidence level of each real vector, the magnitude and phase angle of each real vector, a quantum state is determined. The probability of the quantum state collapsing to each ground state is determined in a predefined way. The semantic category corresponding to the ground state with the highest probability is selected as the intention recognition result of the original speech.

[0011] In one possible implementation of the first aspect,

[0012] The process of obtaining the original speech and deleting invalid segments from the original speech to obtain the updated original speech includes:

[0013] Obtain the original audio, and delete silent segments, breathing sound segments, and background noise segments from the original audio to obtain the updated original audio.

[0014] In one possible implementation of the first aspect, enhancing the high-frequency signals in the updated original speech to obtain enhanced original speech includes:

[0015] The updated original speech is input into the pre-emphasis filter, which enhances the high-frequency signals in the updated original speech to obtain the enhanced original speech.

[0016] In one possible implementation of the first aspect, the step of dividing the enhanced original speech into multiple frames of speech signals, obtaining the Mel-frequency cepstral coefficients of each frame of speech signal, and concatenating the Mel-frequency cepstral coefficients, the first sequence, and the second sequence of each frame of speech signal to obtain a frame-level vector of each frame of speech signal includes:

[0017] The enhanced original speech is divided into multiple frames of speech signals. The Mel frequency cepstral coefficients of each frame of speech signal are obtained. The first-order difference of the Mel frequency cepstral coefficients of each frame of speech signal is performed to obtain the first sequence. The second-order difference of the Mel frequency cepstral coefficients of each frame of speech signal is performed to obtain the second sequence.

[0018] The Mel frequency cepstral coefficients, the first sequence, and the second sequence of each frame of speech signal are concatenated to obtain the frame-level vector of each frame of speech signal.

[0019] In one possible implementation of the first aspect, the step of concatenating the frame-level vectors of each frame of speech signal to obtain a sequence tensor of the original speech, fusing the sequence tensor of the original speech, model instructions, preset labels, and external context to obtain target information, and processing the target information through multiple agents to obtain multiple real vectors, the confidence level of each real vector, the magnitude and phase angle of each real vector, including:

[0020] The frame-level vectors of each frame of speech signal are concatenated to obtain the sequence tensor of the original speech. The model instructions, preset labels, and external context are obtained. The sequence tensor of the original speech, the model instructions, the preset labels, and the external context are fused to obtain the target information.

[0021] The target information is processed by multiple agents, the output of each agent is recorded, and the output of each agent is merged into a comprehensive result. From the comprehensive result, multiple real vectors, the confidence of each real vector, the magnitude and phase angle of each real vector are obtained.

[0022] In one possible implementation of the first aspect, the step of determining a quantum state based on multiple real vectors, the confidence level of each real vector, the magnitude and phase angle of each real vector, determining the probability of the quantum state collapsing to each ground state in a predefined manner, and selecting the semantic category corresponding to the ground state with the highest probability as the intent recognition result of the original speech includes:

[0023] Based on the vector model, the magnitude and phase angle of each real vector, generate the complex vector corresponding to each real vector;

[0024] A normalization model is used to normalize the complex vector corresponding to each real vector to obtain a normalized vector for each real vector. A weight model is used to process the confidence of each real vector to obtain the weight coefficient for each real vector.

[0025] By constructing a model and processing the weight coefficients and normalized vectors corresponding to each real vector, a quantum state is obtained.

[0026] A set of ground states is constructed. The interference amplitude of the quantum state in each ground state is obtained through an amplitude model. The interference amplitude of the quantum state in each ground state is processed through a probability model to obtain the probability of the quantum state collapsing to each ground state. The semantic category corresponding to the ground state with the highest probability is selected as the intention recognition result of the original speech.

[0027] In one possible implementation of the first aspect, after determining the quantum state based on multiple real vectors, the confidence level of each real vector, the magnitude and phase angle of each real vector, determining the probability of the quantum state collapsing to each ground state in a predefined manner, and selecting the semantic category corresponding to the ground state with the highest probability as the intention recognition result of the original speech, the intention recognition method includes:

[0028] Obtain the identifiers of the agents that support the intent recognition results, package the identifiers of the agents that support the intent recognition results and the confidence levels corresponding to the intent recognition results, and generate explanation information for the intent recognition results.

[0029] Secondly, embodiments of this application provide a multi-agent intent recognition device, applied to an electronic device, comprising:

[0030] The acquisition module is used to acquire the original speech, perform deletion operations on invalid segments in the original speech, and obtain the updated original speech.

[0031] The enhancement module is used to enhance the high-frequency signals in the updated original speech to obtain the enhanced original speech;

[0032] The splicing module is used to divide the enhanced original speech into multiple frames of speech signals, obtain the Mel frequency cepstral coefficients of each frame of speech signal, and splice the Mel frequency cepstral coefficients, the first sequence, and the second sequence of each frame of speech signal to obtain the frame-level vector of each frame of speech signal.

[0033] The fusion module is used to concatenate the frame-level vectors of each frame of speech signal to obtain the sequence tensor of the original speech. The sequence tensor of the original speech, model instructions, preset labels, and external context are fused to obtain target information. The target information is processed by multiple agents to obtain multiple real vectors, the confidence of each real vector, and the magnitude and phase angle of each real vector.

[0034] The recognition module is used to determine the quantum state based on multiple real vectors, the confidence level of each real vector, the magnitude and phase angle of each real vector, and to determine the probability of the quantum state collapsing to each ground state in a predefined way. The semantic category corresponding to the ground state with the highest probability is selected as the intention recognition result of the original speech.

[0035] Thirdly, embodiments of this application provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the intent recognition method of any of the first aspects described above.

[0036] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the intent recognition method of any one of the first aspects described above.

[0037] Fifthly, embodiments of this application provide a computer program product that, when run on an electronic device, causes the electronic device to execute the intent recognition method described in any of the first aspects above.

[0038] The beneficial effects of this application's embodiments are twofold. Firstly, by concatenating the frame-level vectors of each frame of speech signal to obtain the sequence tensor of the original speech, the sequence tensor of the original speech, model instructions, preset labels, and external context are fused to obtain target information. Multiple agents process the target information to obtain multiple real vectors, the confidence level of each real vector, and the amplitude and phase angle of each real vector. Based on these multiple real vectors, their confidence levels, amplitudes, and phase angles, quantum states are determined. The probability of the quantum state collapsing to each ground state is determined in a predefined manner. The semantic category corresponding to the ground state with the highest probability is selected as the intent recognition result of the original speech. Since multiple agents can leverage their respective strengths in algorithms and processing logic to analyze the target information from multiple dimensions, this allows for a comprehensive and in-depth analysis of the target information, effectively improving the accuracy of intent recognition. Secondly, it eliminates the need for multiple intent recognition operations on the original speech, thus reducing the acquisition time of the intent recognition result and improving the efficiency of obtaining the intent recognition result. Attached Figure Description

[0039] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0040] Figure 1 This is an application scenario diagram of the intent recognition method provided in the embodiments of this application;

[0041] Figure 2 This is a flowchart illustrating the intent recognition method provided in an embodiment of this application;

[0042] Figure 3 A flowchart illustrating the implementation of S203 provided in this application embodiment;

[0043] Figure 4 A schematic block diagram of an intent recognition device provided in the embodiments of this application;

[0044] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0045] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application. All other embodiments obtained by those skilled in the art based on the embodiments in this application without inventive effort are within the scope of protection of this application.

[0046] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0047] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0048] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0049] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."

[0050] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0051] References to "one embodiment" or "some embodiments" in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0052] Furthermore, the technical solutions of the various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such combination of technical solutions does not exist and is not within the scope of protection claimed in this application.

[0053] The flowchart shown in the attached diagram is for illustrative purposes only and does not necessarily include all content and operations / steps, nor does it necessarily have to be performed in the order described. For example, some operations / steps can be broken down, combined, or partially merged, so the actual execution order may change depending on the actual situation.

[0054] The intent recognition method provided in this application can be applied to electronic devices such as mobile phones, tablets, wearable devices, in-vehicle devices, augmented reality (AR) / virtual reality (VR) devices, laptops, ultra-mobile personal computers (UMPCs), netbooks, and personal digital assistants (PDAs). This application does not impose any restrictions on the specific type of electronic device.

[0055] Please see Figure 1 , Figure 1 The application scenario diagram of the intent recognition method provided in the embodiments of this application is described in detail below:

[0056] The electronic device acquires the original speech through an audio sensor, and then preprocesses the original speech to obtain an updated original speech.

[0057] When a user speaks, the audio sensor captures the raw speech and transmits it to the voice recognition device of the electronic device.

[0058] In this embodiment of the application, the electronic device acquires the original speech through an audio sensor. The audio sensor has millisecond-level sampling capability, which can capture the original speech in real time, thereby shortening the acquisition time of the original speech and improving the acquisition efficiency of the original speech.

[0059] Please see Figure 2 , Figure 2 This is a flowchart illustrating the intent recognition method provided in an embodiment of this application, which can be applied to electronic devices.

[0060] like Figure 2 As shown, the intent recognition method provided in this application includes the following steps, which are detailed below:

[0061] S201, Obtain the original speech, perform deletion operations on invalid segments in the original speech, and obtain the updated original speech;

[0062] Invalid segments include silent segments, breathing sound segments, and background noise segments.

[0063] The step of obtaining the original speech and deleting invalid segments from the original speech to obtain the updated original speech includes:

[0064] Obtain the original audio, and delete silent segments, breathing sound segments, and background noise segments from the original audio to obtain the updated original audio.

[0065] The original speech is obtained, and the effective speech region is determined using the energy threshold method and zero crossover rate. The silence segments, breathing sound segments, and background noise segments before and after the effective speech region are deleted. The original speech with the silence segments, breathing sound segments, and background noise segments deleted is selected as the updated original speech.

[0066] In this process, silent segments, breathing sound segments, and background noise segments before and after the effective speech region are deleted, making the updated original speech more concise. The agent only needs to focus on the core part when processing, which greatly saves computing resources and improves processing speed, enabling the agent to complete the analysis and processing of the updated original speech in a shorter time.

[0067] S202, enhance the high-frequency signals in the updated original speech to obtain the enhanced original speech;

[0068] The step of enhancing the high-frequency signals in the updated original speech to obtain the enhanced original speech includes:

[0069] The updated original speech is input into the pre-emphasis filter, which enhances the high-frequency signals in the updated original speech to obtain the enhanced original speech.

[0070] S203, the enhanced original speech is divided into multiple frames of speech signals, the Mel frequency cepstral coefficients of each frame of speech signal are obtained, and the Mel frequency cepstral coefficients, the first sequence and the second sequence of each frame of speech signal are concatenated to obtain the frame-level vector of each frame of speech signal.

[0071] Mel-Frequency Cepstral Coefficients (MFCC) is a core technology for extracting features from speech signals by simulating the auditory characteristics of the human ear. Its core logic is to convert speech from the time domain to the Mel frequency domain, and then separate the characteristics of the sound source and vocal tract through cepstral analysis, ultimately obtaining a set of low-dimensional, robust and physiologically meaningful feature parameters.

[0072] S204: The frame-level vectors of each frame of speech signal are concatenated to obtain the sequence tensor of the original speech. The sequence tensor of the original speech, model instructions, preset labels, and external context are fused to obtain target information. The target information is processed by multiple agents to obtain multiple real vectors, the confidence of each real vector, the magnitude and phase angle of each real vector.

[0073] Specifically, the process involves concatenating the frame-level vectors of each frame of speech signal to obtain a sequence tensor of the original speech; fusing the sequence tensor of the original speech, model instructions, preset labels, and external context to obtain target information; and processing the target information through multiple agents to obtain multiple real vectors, the confidence level of each real vector, and the magnitude and phase angle of each real vector, including:

[0074] The frame-level vectors of each frame of speech signal are concatenated to obtain the sequence tensor of the original speech. The model instructions, preset labels, and external context are obtained. The sequence tensor of the original speech, the model instructions, the preset labels, and the external context are fused to obtain the target information.

[0075] The target information is processed by multiple agents, the output of each agent is recorded, and the output of each agent is merged into a comprehensive result. From the comprehensive result, multiple real vectors, the confidence of each real vector, the magnitude and phase angle of each real vector are obtained.

[0076] Among them, multiple intelligent agents include an intent recognition agent, an emotion recognition agent, a tone recognition agent, and a context understanding agent.

[0077] An intent-recognition intelligent agent is an artificial intelligence system that infers user needs based on user input.

[0078] An emotion recognition agent is an artificial intelligence system that detects emotional states based on user input.

[0079] A tone recognition intelligent agent is an artificial intelligence system that detects the tone of a user's voice based on the user's input.

[0080] A context-aware intelligent agent is an artificial intelligence system that infers the current semantics in real time by dynamically analyzing dialogue history, environmental information, and user profiles.

[0081] Among them, model instructions are task descriptions provided by the user to the agent, and model instructions are used to guide the agent to generate outputs that meet expectations.

[0082] For ease of explanation, the following example is provided:

[0083] For example, when the intelligent voice assistant is an in-vehicle voice system, the model instruction is: As an in-vehicle voice system, recognize the intent recognition result of the original speech in real time. Input is single-channel audio, supporting Mandarin recognition under background noise, and output is structured data.

[0084] For example, when the intelligent voice assistant acts as a smart home control system, the model command would be: "As a smart home control system, recognize the intent of the original speech in real time. Input is single-channel audio, supports Mandarin recognition under background noise, and outputs structured data."

[0085] By controlling the output format through preset tags, a unified approach can be achieved, integrating structured information management, dynamic content adaptation, and improved interaction efficiency. The agent outputs results according to the format defined by the preset tags, ensuring consistency and avoiding formatting issues caused by freely generated agents.

[0086] Among them, the model instructions define the task description, the preset labels add structured identifiers to the output information, and the external context provides dynamic scene data. By fusing the sequence tensor of the original speech, the model instructions, the preset labels, and the external context, target information that conforms to both task specifications and actual scenarios can be generated.

[0087] S205 determines the quantum state based on multiple real vectors, the confidence level of each real vector, the magnitude and phase angle of each real vector, and determines the probability of the quantum state collapsing to each ground state in a predefined way. The semantic category corresponding to the ground state with the highest probability is selected as the intention recognition result of the original speech.

[0088] The process of determining a quantum state based on multiple real vectors, the confidence level of each real vector, the magnitude and phase angle of each real vector, and determining the probability of the quantum state collapsing to each ground state in a predefined manner, and selecting the semantic category corresponding to the ground state with the highest probability as the intent recognition result of the original speech, includes:

[0089] Based on the vector model, the magnitude and phase angle of each real vector, generate the complex vector corresponding to each real vector;

[0090] The vector model is as follows:

[0091] ;

[0092] in, Indicates the first The complex vector corresponding to each real vector. Indicates the first The magnitude of a real vector, Indicates the first The phase angle of a real vector. It is the imaginary unit.

[0093] A normalization model is used to normalize the complex vector corresponding to each real vector to obtain a normalized vector for each real vector. A weight model is then used to process the confidence of each real vector to obtain the weight coefficient for each real vector.

[0094] The normalization model is as follows:

[0095] ;

[0096] in, Indicates the first The normalized vector corresponding to each real vector. Indicates the first The complex vector corresponding to each real vector. Indicates the first The norm of the complex vector corresponding to each real vector.

[0097] The weighting model is as follows:

[0098] ;

[0099] Indicates the first The weight coefficients corresponding to each real vector Indicates the first The confidence level corresponding to each real vector. This represents the total number of real vectors.

[0100] By constructing a model and processing the weight coefficients and normalized vectors corresponding to each real vector, a quantum state is obtained.

[0101] The model is constructed as follows:

[0102] ;

[0103] Representing quantum states, Indicates the first The weight coefficients corresponding to each real vector This represents the total number of real vectors.

[0104] A set of ground states is constructed. The interference amplitude of the quantum state in each ground state is obtained through an amplitude model. The interference amplitude of the quantum state in each ground state is processed through a probability model to obtain the probability of the quantum state collapsing to each ground state. The semantic category corresponding to the ground state with the highest probability is selected as the intention recognition result of the original speech.

[0105] The amplitude model is as follows:

[0106] ;

[0107] Representing the quantum state in the th case Interference amplitude at each ground state Indicates the first A ground state, Represents a quantum state;

[0108] The probability model is as follows:

[0109] ;

[0110] This indicates that the quantum state collapses to the th... The probability of the ground state. Representing the quantum state in the th case Interference amplitude on each ground state.

[0111] Each ground state corresponds to a semantic category, and different ground states correspond to different semantic categories.

[0112] Among them, the quantum state collapses to the first The higher the probability of the ground state, the closer the quantum state is to the first ground state. The higher the correlation of the ground states, the more the quantum state collapses to the first state. The lower the probability of the ground state, the closer the quantum state is to the first ground state. The lower the correlation between ground states.

[0113] For ease of explanation, the following example is provided:

[0114] For example, construct a set of ground states, which includes ground state 1, ground state 2, ground state 3, ground state 4, and ground state 5. Ground state 1, ground state 2, ground state 3, ground state 4, and ground state 5 are different ground states.

[0115] The semantic category corresponding to ground state 1 is query; the semantic category corresponding to ground state 2 is command.

[0116] The semantic category corresponding to ground state 3 is confirmation; the semantic category corresponding to ground state 4 is complaint.

[0117] The semantic category corresponding to ground state 5 is request.

[0118] The probability of a quantum state collapsing into ground state 1 is probability 1; the probability of a quantum state collapsing into ground state 2 is probability 2.

[0119] The probability of a quantum state collapsing into ground state 3 is probability 3; the probability of a quantum state collapsing into ground state 4 is probability 4.

[0120] The probability of a quantum state collapsing into ground state 5 is 5.

[0121] Among probabilities 1, 2, 3, 4, and 5, when probability 1 is the highest, the query is selected as the intent recognition result of the original speech.

[0122] Among probabilities 1, 2, 3, 4, and 5, when probability 2 is the highest, the command is selected as the intention recognition result of the original speech.

[0123] Among probabilities 1, 2, 3, 4, and 5, when probability 3 is the highest, confirmation is selected as the intention recognition result of the original speech.

[0124] Among probabilities 1, 2, 3, 4, and 5, when probability 4 is the highest, the complaint is selected as the intention recognition result of the original speech.

[0125] Among probabilities 1, 2, 3, 4, and 5, when probability 5 is the highest, the request is selected as the intent recognition result of the original speech.

[0126] Specifically, by constructing a model and processing the weight coefficients and normalized vectors corresponding to each real vector, a quantum state is obtained, including:

[0127] Record the normalized vector corresponding to each real vector to obtain a vector set. Randomly select the first vector and the second vector from the vector set, calculate the inner product of the first vector and the second vector, and select the square of the magnitude of the inner product as the coherence parameter between the first vector and the second vector.

[0128] When the coherence parameter is less than the preset value, the confidence of each real vector is processed through the weight model to obtain the weight coefficient corresponding to each real vector. Through the preset construction model, the weight coefficient and the normalized vector corresponding to each real vector are processed to obtain the quantum state.

[0129] The coherence parameter between the first vector and the second vector is used to quantify the correlation strength between them. The higher the coherence parameter, the stronger the correlation between them; conversely, the lower the coherence parameter, the weaker the correlation.

[0130] A coherence parameter lower than the preset value typically indicates a weaker-than-expected correlation between the first and second vectors. This is because the first and second vectors are normalized vectors corresponding to different real vectors.

[0131] This demonstrates that there is no significant linear correlation between the normalized vectors corresponding to different real vectors. If we still want to explore potential deep correlations, we need to use quantum states.

[0132] The intent recognition method, after determining the quantum state based on multiple real vectors, the confidence level of each real vector, the magnitude and phase angle of each real vector, determining the probability of the quantum state collapsing to each ground state in a predefined manner, and selecting the semantic category corresponding to the ground state with the highest probability as the intent recognition result of the original speech, includes:

[0133] Obtain the identifiers of the agents that support the intent recognition results, package the identifiers of the agents that support the intent recognition results and the confidence levels corresponding to the intent recognition results, and generate explanation information for the intent recognition results.

[0134] For ease of explanation, the following example is provided:

[0135] Multiple intelligent agents include an intent recognition agent, an emotion recognition agent, a tone recognition agent, and a context understanding agent;

[0136] The agent for intent recognition is labeled A1, the agent for emotion recognition is labeled A2, the agent for tone recognition is labeled A3, and the agent for context understanding is labeled A4.

[0137] When the agent supporting the intent recognition result is an intent recognition agent and a context understanding agent, A1, A4 and the confidence corresponding to the intent recognition result are packaged together to generate the explanation information of the intent recognition result.

[0138] When the agent supporting the intent recognition result is an intent recognition agent, an emotion recognition agent, or a tone recognition agent, A1, A2, A3, and the confidence scores corresponding to the intent recognition results are packaged together to generate explanatory information for the intent recognition results.

[0139] When the agent supporting the intent recognition result is an emotion recognition agent or a tone recognition agent, A2, A3 and the confidence corresponding to the intent recognition result are packaged together to generate the explanation information of the intent recognition result.

[0140] Among them, the explanatory information of intent recognition results can significantly improve the transparency of human-computer interaction and user experience, because the explanatory information of intent recognition results can help users understand why the agent makes a judgment. In addition, when there is a deviation in recognition, a clear explanation can help users quickly locate the problem and actively adjust the input to improve the efficiency of subsequent interactions.

[0141] The beneficial effects of this application's embodiments are twofold. Firstly, by concatenating the frame-level vectors of each frame of speech signal to obtain the sequence tensor of the original speech, the sequence tensor of the original speech, model instructions, preset labels, and external context are fused to obtain target information. Multiple agents process the target information to obtain multiple real vectors, the confidence level of each real vector, and the amplitude and phase angle of each real vector. Based on these multiple real vectors, their confidence levels, amplitudes, and phase angles, quantum states are determined. The probability of the quantum state collapsing to each ground state is determined in a predefined manner. The semantic category corresponding to the ground state with the highest probability is selected as the intent recognition result of the original speech. Since multiple agents can leverage their respective strengths in algorithms and processing logic to analyze the target information from multiple dimensions, this allows for a comprehensive and in-depth analysis of the target information, effectively improving the accuracy of intent recognition. Secondly, it eliminates the need for multiple intent recognition operations on the original speech, thus reducing the acquisition time of the intent recognition result and improving the efficiency of obtaining the intent recognition result.

[0142] Please see Figure 3 , Figure 3 The implementation flowchart of S203 provided in the embodiments of this application is described in detail below:

[0143] S301, the enhanced original speech is divided into multiple frames of speech signals, the Mel frequency cepstral coefficients of each frame of speech signal are obtained, the first-order difference of the Mel frequency cepstral coefficients of each frame of speech signal is performed to obtain the first sequence, and the second-order difference of the Mel frequency cepstral coefficients of each frame of speech signal is performed to obtain the second sequence.

[0144] Specifically, the first-order difference of the Mel-frequency cepstral coefficients of each frame of speech signal is performed to obtain the first sequence, which can make up for the deficiency of the Mel-frequency cepstral coefficients in representing the dynamic changes of speech.

[0145] Specifically, second-order difference is performed on the Mel frequency cepstral coefficients. Based on the Mel frequency cepstral coefficients and the first sequence, the dynamic rate of change of speech features is further extracted, thereby more comprehensively characterizing the temporal evolution pattern of the speech signal.

[0146] S302, the Mel frequency cepstral coefficients, the first sequence and the second sequence of each frame of speech signal are concatenated to obtain the frame-level vector of each frame of speech signal.

[0147] In this embodiment, the Mel frequency cepstral coefficients, the first sequence, and the second sequence of each frame of speech signal are concatenated to obtain the frame-level vector of each frame of speech signal. In this way, the frame-level vector of each frame of speech signal has richer temporal information, enabling the agent to more accurately capture the semantic and emotional details in the user's expression, thereby improving the accuracy of intent recognition in complex contexts.

[0148] For the intent recognition method described in the above embodiments, please refer to [link / reference]. Figure 4 , Figure 4 This is a schematic block diagram of the intent recognition device provided in the embodiments of this application. Figure 4 The intent recognition device 400 shown can be applied to, for example... Figure 1 The application scenario diagram shows electronic devices. The following section uses electronic devices as an example to illustrate this. Figure 4 The intent recognition device 400 shown will be described in detail. The intent recognition device 400 may include an acquisition module 401, an enhancement module 402, a stitching module 403, a fusion module 404, and a recognition module 405.

[0149] The acquisition module 401 is used to acquire the original speech, perform deletion operations on invalid segments in the original speech, and obtain the updated original speech.

[0150] Enhancement module 402 is used to enhance the high-frequency signals in the updated original speech to obtain the enhanced original speech;

[0151] The splicing module 403 is used to divide the enhanced original speech into multiple frames of speech signals, obtain the Mel frequency cepstral coefficients of each frame of speech signal, and splice the Mel frequency cepstral coefficients, the first sequence, and the second sequence of each frame of speech signal to obtain the frame-level vector of each frame of speech signal.

[0152] The fusion module 404 is used to concatenate the frame-level vectors of each frame of speech signal to obtain the sequence tensor of the original speech, fuse the sequence tensor of the original speech, model instructions, preset labels, and external context to obtain target information, process the target information through multiple agents to obtain multiple real vectors, the confidence of each real vector, the magnitude and phase angle of each real vector;

[0153] The recognition module 405 is used to determine the quantum state based on multiple real vectors, the confidence level of each real vector, the magnitude and phase angle of each real vector, and to determine the probability of the quantum state collapsing to each ground state in a predefined manner. The semantic category corresponding to the ground state with the highest probability is selected as the intention recognition result of the original speech.

[0154] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0155] The beneficial effects of this application's embodiments are twofold. Firstly, by concatenating the frame-level vectors of each frame of speech signal to obtain the sequence tensor of the original speech, the sequence tensor of the original speech, model instructions, preset labels, and external context are fused to obtain target information. Multiple agents process the target information to obtain multiple real vectors, the confidence level of each real vector, and the amplitude and phase angle of each real vector. Based on these multiple real vectors, their confidence levels, amplitudes, and phase angles, quantum states are determined. The probability of the quantum state collapsing to each ground state is determined in a predefined manner. The semantic category corresponding to the ground state with the highest probability is selected as the intent recognition result of the original speech. Since multiple agents can leverage their respective strengths in algorithms and processing logic to analyze the target information from multiple dimensions, this allows for a comprehensive and in-depth analysis of the target information, effectively improving the accuracy of intent recognition. Secondly, it eliminates the need for multiple intent recognition operations on the original speech, thus reducing the acquisition time of the intent recognition result and improving the efficiency of obtaining the intent recognition result.

[0156] Please see Figure 5 , Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.

[0157] like Figure 5 As shown, Figure 5 The electronic device 2 includes: at least one processor 20, a memory 21, and a computer program 22 stored in the memory 21 and executable on the at least one processor 20, wherein the processor 20 executes the computer program 22 to implement the steps in any of the above method embodiments.

[0158] The electronic device 2 may include, but is not limited to, a processor 20 and a memory 21. Those skilled in the art will understand that... Figure 5 This is merely an example of electronic device 2 and does not constitute a limitation on electronic device 2. It may include more or fewer components than shown in the figure, or combine certain components, or different components, such as input / output devices, network access devices, etc.

[0159] The processor 20 is used to run a computer program 22 stored in the memory 21, and performs the following steps when executing the computer program 22:

[0160] Obtain the original speech, delete invalid segments from the original speech, and obtain the updated original speech;

[0161] The high-frequency signals in the updated original speech are enhanced to obtain the enhanced original speech.

[0162] The enhanced original speech is divided into multiple frames of speech signals. The Mel frequency cepstral coefficients of each frame of speech signal are obtained. The Mel frequency cepstral coefficients, the first sequence, and the second sequence of each frame of speech signal are concatenated to obtain the frame-level vector of each frame of speech signal.

[0163] The frame-level vectors of each frame of speech signal are concatenated to obtain the sequence tensor of the original speech. The sequence tensor of the original speech, model instructions, preset labels, and external context are fused to obtain the target information. The target information is processed by multiple agents to obtain multiple real vectors, the confidence of each real vector, and the magnitude and phase angle of each real vector.

[0164] Based on multiple real vectors, the confidence level of each real vector, the magnitude and phase angle of each real vector, a quantum state is determined. The probability of the quantum state collapsing to each ground state is determined in a predefined way. The semantic category corresponding to the ground state with the highest probability is selected as the intention recognition result of the original speech.

[0165] In some embodiments, the processor 20 is configured to implement:

[0166] Obtain the original audio, and delete silent segments, breathing sound segments, and background noise segments from the original audio to obtain the updated original audio.

[0167] In some embodiments, the processor 20 is configured to implement:

[0168] The updated original speech is input into the pre-emphasis filter, which enhances the high-frequency signals in the updated original speech to obtain the enhanced original speech.

[0169] In some embodiments, the processor 20 is configured to implement:

[0170] The enhanced original speech is divided into multiple frames of speech signals. The Mel frequency cepstral coefficients of each frame of speech signal are obtained. The first-order difference of the Mel frequency cepstral coefficients of each frame of speech signal is performed to obtain the first sequence. The second-order difference of the Mel frequency cepstral coefficients of each frame of speech signal is performed to obtain the second sequence.

[0171] The Mel-frequency cepstral coefficients, the first sequence, and the second sequence of each frame of the speech signal are concatenated to obtain the frame-level vector of each frame of the speech signal. In some embodiments, the processor 20 is configured to:

[0172] The frame-level vectors of each frame of speech signal are concatenated to obtain the sequence tensor of the original speech. The model instructions, preset labels, and external context are obtained. The sequence tensor of the original speech, the model instructions, the preset labels, and the external context are fused to obtain the target information.

[0173] The target information is processed by multiple agents, the output of each agent is recorded, and the output of each agent is merged into a comprehensive result. From the comprehensive result, multiple real vectors, the confidence of each real vector, the magnitude and phase angle of each real vector are obtained.

[0174] In some embodiments, the processor 20 is configured to implement:

[0175] Based on the vector model, the magnitude and phase angle of each real vector, generate the complex vector corresponding to each real vector;

[0176] A normalization model is used to normalize the complex vector corresponding to each real vector to obtain a normalized vector for each real vector. A weight model is used to process the confidence of each real vector to obtain the weight coefficient for each real vector.

[0177] By constructing a model and processing the weight coefficients and normalized vectors corresponding to each real vector, a quantum state is obtained.

[0178] A set of ground states is constructed. The interference amplitude of the quantum state in each ground state is obtained through an amplitude model. The interference amplitude of the quantum state in each ground state is processed through a probability model to obtain the probability of the quantum state collapsing to each ground state. The semantic category corresponding to the ground state with the highest probability is selected as the intention recognition result of the original speech.

[0179] In some embodiments, the processor 20 is configured to implement:

[0180] Obtain the identifiers of the agents that support the intent recognition results, package the identifiers of the agents that support the intent recognition results and the confidence levels corresponding to the intent recognition results, and generate explanation information for the intent recognition results.

[0181] The processor 20 may be a Central Processing Unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.

[0182] In some embodiments, the memory 21 may be an internal storage unit of the electronic device 2, such as a hard disk or memory of the electronic device 2. In other embodiments, the memory 21 may be an external storage device of the electronic device 2, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the electronic device 2. Furthermore, the memory 21 may include both internal and external storage units of the electronic device 2. The memory 21 is used to store the operating system, applications, boot loader, data, and other programs, such as the program code of the computer program. The memory 21 can also be used to temporarily store data that has been output or will be output.

[0183] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.

[0184] This application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps described in the various method embodiments above.

[0185] The computer-readable storage medium stores program code that can be called by a processor to execute the intent recognition method described in the above method embodiments.

[0186] Computer-readable storage media have storage space for program code.

[0187] The program code includes the code for any step of the intent recognition method described in the above method embodiments.

[0188] For example, when program code is invoked by the processor, it can perform the following steps:

[0189] Obtain the original speech, delete invalid segments from the original speech, and obtain the updated original speech;

[0190] The high-frequency signals in the updated original speech are enhanced to obtain the enhanced original speech.

[0191] The enhanced original speech is divided into multiple frames of speech signals. The Mel frequency cepstral coefficients of each frame of speech signal are obtained. The Mel frequency cepstral coefficients, the first sequence, and the second sequence of each frame of speech signal are concatenated to obtain the frame-level vector of each frame of speech signal.

[0192] The frame-level vectors of each frame of speech signal are concatenated to obtain the sequence tensor of the original speech. The sequence tensor of the original speech, model instructions, preset labels, and external context are fused to obtain the target information. The target information is processed by multiple agents to obtain multiple real vectors, the confidence of each real vector, and the magnitude and phase angle of each real vector.

[0193] Based on multiple real vectors, the confidence level of each real vector, the magnitude and phase angle of each real vector, a quantum state is determined. The probability of the quantum state collapsing to each ground state is determined in a predefined way. The semantic category corresponding to the ground state with the highest probability is selected as the intention recognition result of the original speech.

[0194] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0195] The computer-readable storage medium may also be an external storage device of an intent recognition device or electronic device, such as a plug-in hard drive, a smart media card (SMC), a secure digital (SD) card, a flash card, or a non-transitory computer-readable storage medium equipped on the intent recognition device or electronic device.

[0196] Since the computer program stored in the computer-readable storage medium can execute any of the multi-agent-based intent recognition methods provided in the embodiments of this application, the computer-readable storage medium can achieve the beneficial effects that any of the multi-agent-based intent recognition methods provided in the embodiments of this application can achieve, as detailed in the preceding embodiments, and will not be repeated here.

[0197] This application provides a computer program product that, when run on an electronic device, causes the electronic device to execute the aforementioned intent recognition method.

[0198] When a computer program is loaded into an electronic device, it can perform the following steps:

[0199] Obtain the original speech, delete invalid segments from the original speech, and obtain the updated original speech;

[0200] The high-frequency signals in the updated original speech are enhanced to obtain the enhanced original speech.

[0201] The enhanced original speech is divided into multiple frames of speech signals. The Mel frequency cepstral coefficients of each frame of speech signal are obtained. The Mel frequency cepstral coefficients, the first sequence, and the second sequence of each frame of speech signal are concatenated to obtain the frame-level vector of each frame of speech signal.

[0202] The frame-level vectors of each frame of speech signal are concatenated to obtain the sequence tensor of the original speech. The sequence tensor of the original speech, model instructions, preset labels, and external context are fused to obtain the target information. The target information is processed by multiple agents to obtain multiple real vectors, the confidence of each real vector, and the magnitude and phase angle of each real vector.

[0203] Based on multiple real vectors, the confidence level of each real vector, the magnitude and phase angle of each real vector, a quantum state is determined. The probability of the quantum state collapsing to each ground state is determined in a predefined way. The semantic category corresponding to the ground state with the highest probability is selected as the intention recognition result of the original speech.

[0204] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium.

[0205] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0206] Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to an electronic device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks.

[0207] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0208] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. A multi-agent intent recognition method, characterized in that, The intent recognition method, applied to electronic devices, includes: Obtain the original speech, delete invalid segments from the original speech, and obtain the updated original speech; The high-frequency signals in the updated original speech are enhanced to obtain the enhanced original speech. The enhanced original speech is divided into multiple frames of speech signals. The Mel frequency cepstral coefficients of each frame of speech signal are obtained. The Mel frequency cepstral coefficients, the first sequence, and the second sequence of each frame of speech signal are concatenated to obtain the frame-level vector of each frame of speech signal. The frame-level vectors of each frame of speech signal are concatenated to obtain the sequence tensor of the original speech. The sequence tensor of the original speech, model instructions, preset labels, and external context are fused to obtain the target information. The target information is processed by multiple agents to obtain multiple real vectors, the confidence of each real vector, and the magnitude and phase angle of each real vector. Based on the vector model, the amplitude and phase angle of each real vector, a complex vector corresponding to each real vector is generated. A normalization model is used to normalize the complex vector corresponding to each real vector to obtain a normalized vector corresponding to each real vector. A weight model is used to process the confidence of each real vector to obtain the weight coefficient corresponding to each real vector. By constructing a model, the weight coefficient and the normalized vector corresponding to each real vector are processed to obtain a quantum state. A set of ground states is constructed. An amplitude model is used to obtain the interference amplitude of the quantum state in each ground state. A probability model is used to process the interference amplitude of the quantum state in each ground state to obtain the probability of the quantum state collapsing to each ground state. The semantic category corresponding to the ground state with the highest probability is selected as the intention recognition result of the original speech. The vector model is as follows: ; in, Indicates the first The complex vector corresponding to each real vector. Indicates the first The magnitude of the real vector, Phase angle, where j is the imaginary unit; The normalization model is as follows: ; in, Indicates the first The normalized vector corresponding to each real vector. Indicates the first The complex vector corresponding to each real vector. Indicates the first The norm of the complex vector corresponding to each real vector; The weighting model is as follows: ; Indicates the first The weight coefficients corresponding to each real vector Indicates the first The confidence scores corresponding to each real vector, where n represents the total number of real vectors; The model is constructed as follows: ; Representing quantum states, Indicates the first The weight coefficients corresponding to each real vector, where n represents the total number of real vectors; The amplitude model is as follows: a k =b k × |Ψ> Representing the quantum state in the th case Interference amplitude at each ground state Indicates the first A ground state, Represents a quantum state; The probability model is as follows: ; The quantum state collapses to the first The probability of the ground state. Representing the quantum state in the th case Interference amplitude at each ground state; Each ground state corresponds to a semantic category, and different ground states correspond to different semantic categories.

2. The intent recognition method according to claim 1, characterized in that, The process of obtaining the original speech and deleting invalid segments from the original speech to obtain the updated original speech includes: Obtain the original audio, and delete silent segments, breathing sound segments, and background noise segments from the original audio to obtain the updated original audio.

3. The intent recognition method according to claim 1, characterized in that, The process of enhancing the high-frequency signals in the updated original speech to obtain enhanced original speech includes: The updated original speech is input into the pre-emphasis filter, which enhances the high-frequency signals in the updated original speech to obtain the enhanced original speech.

4. The intent recognition method according to claim 1, characterized in that, The enhanced original speech is divided into multiple frames of speech signals, the Mel-frequency cepstral coefficients of each frame are obtained, and the Mel-frequency cepstral coefficients, the first sequence, and the second sequence of each frame are concatenated to obtain a frame-level vector for each frame of speech signal, including: The enhanced original speech is divided into multiple frames of speech signals. The Mel frequency cepstral coefficients of each frame of speech signal are obtained. The first-order difference of the Mel frequency cepstral coefficients of each frame of speech signal is performed to obtain the first sequence. The second-order difference of the Mel frequency cepstral coefficients of each frame of speech signal is performed to obtain the second sequence. The Mel frequency cepstral coefficients, the first sequence, and the second sequence of each frame of speech signal are concatenated to obtain the frame-level vector of each frame of speech signal.

5. The intent recognition method according to claim 1, characterized in that, The process involves concatenating the frame-level vectors of each frame of speech signal to obtain a sequence tensor of the original speech. This sequence tensor, model instructions, preset labels, and external context are then fused to obtain target information. This target information is processed by multiple agents to obtain multiple real vectors, the confidence level of each real vector, and the magnitude and phase angle of each real vector, including: The frame-level vectors of each frame of speech signal are concatenated to obtain the sequence tensor of the original speech. The model instructions, preset labels, and external context are obtained. The sequence tensor of the original speech, the model instructions, the preset labels, and the external context are fused to obtain the target information. The target information is processed by multiple agents, the output of each agent is recorded, and the output of each agent is merged into a comprehensive result. From the comprehensive result, multiple real vectors, the confidence of each real vector, the magnitude and phase angle of each real vector are obtained.

6. The intent recognition method according to claim 1, characterized in that, The intent recognition method, after generating a complex vector corresponding to each real vector based on the vector model, the amplitude and phase angle of each real vector, and normalizing the complex vector corresponding to each real vector using a normalization model to obtain a normalized vector, and processing the confidence of each real vector using a weight model to obtain the weight coefficients corresponding to each real vector, and then processing the weight coefficients and normalized vectors corresponding to each real vector through model construction to obtain a quantum state, constructing a set of ground states, obtaining the interference amplitude of the quantum state in each ground state through an amplitude model, processing the interference amplitude of the quantum state in each ground state through a probability model to obtain the probability of the quantum state collapsing to each ground state, and selecting the semantic category corresponding to the ground state with the highest probability as the original language, includes: Obtain the identifiers of the agents that support the intent recognition results, package the identifiers of the agents that support the intent recognition results and the confidence levels corresponding to the intent recognition results, and generate explanation information for the intent recognition results.

7. A multi-agent intent recognition device, characterized in that, Applied to electronic devices, including: The acquisition module is used to acquire the original speech, perform deletion operations on invalid segments in the original speech, and obtain the updated original speech. The enhancement module is used to enhance the high-frequency signals in the updated original speech to obtain the enhanced original speech; The splicing module is used to divide the enhanced original speech into multiple frames of speech signals, obtain the Mel frequency cepstral coefficients of each frame of speech signal, and splice the Mel frequency cepstral coefficients, the first sequence, and the second sequence of each frame of speech signal to obtain the frame-level vector of each frame of speech signal. The fusion module is used to concatenate the frame-level vectors of each frame of speech signal to obtain the sequence tensor of the original speech. The sequence tensor of the original speech, model instructions, preset labels, and external context are fused to obtain target information. The target information is processed by multiple agents to obtain multiple real vectors, the confidence of each real vector, and the magnitude and phase angle of each real vector. The recognition module generates a complex vector corresponding to each real vector based on the vector model, the amplitude and phase angle of each real vector, and uses a normalization model to normalize the complex vector corresponding to each real vector to obtain a normalized vector corresponding to each real vector. It then processes the confidence level of each real vector using a weight model to obtain the weight coefficients corresponding to each real vector. Finally, it constructs a quantum state by processing the weight coefficients and normalized vectors corresponding to each real vector through a model construction process. A set of ground states is then constructed. An amplitude model is used to obtain the interference amplitude of the quantum state in each ground state. A probability model is used to process the interference amplitude of the quantum state in each ground state to obtain the probability of the quantum state collapsing to each ground state. The semantic category corresponding to the ground state with the highest probability is selected as the intention recognition result of the original speech. Based on the vector model, the amplitude and phase angle of each real vector, a complex vector corresponding to each real vector is generated. A normalization model is used to normalize the complex vector corresponding to each real vector to obtain a normalized vector corresponding to each real vector. A weight model is used to process the confidence of each real vector to obtain the weight coefficient corresponding to each real vector. By constructing a model, the weight coefficient and the normalized vector corresponding to each real vector are processed to obtain a quantum state. A set of ground states is constructed. An amplitude model is used to obtain the interference amplitude of the quantum state in each ground state. A probability model is used to process the interference amplitude of the quantum state in each ground state to obtain the probability of the quantum state collapsing to each ground state. The semantic category corresponding to the ground state with the highest probability is selected as the intention recognition result of the original speech. The vector model is as follows: ; in, Indicates the first The complex vector corresponding to each real vector. Indicates the first The magnitude of the real vector, Phase angle, where j is the imaginary unit; The normalization model is as follows: ; in, Indicates the first The normalized vector corresponding to each real vector. Indicates the first The complex vector corresponding to each real vector. Indicates the first The norm of the complex vector corresponding to each real vector; The weighting model is as follows: ; Indicates the first The weight coefficients corresponding to each real vector Indicates the first The confidence scores corresponding to each real vector, where n represents the total number of real vectors; The model is constructed as follows: ; Representing quantum states, Indicates the first The weight coefficients corresponding to each real vector, where n represents the total number of real vectors; The amplitude model is as follows: a k =b k × |Ψ> Representing the quantum state in the th case Interference amplitude at each ground state Indicates the first A ground state, Represents a quantum state; The probability model is as follows: ; The quantum state collapses to the first The probability of the ground state. Representing the quantum state in the th case Interference amplitude at each ground state; Each ground state corresponds to a semantic category, and different ground states correspond to different semantic categories.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the intent recognition method as described in any one of claims 1 to 6.

9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the intent recognition method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Data processing method and related equipment

    CN115221846A

  • Government affair item dialogue recommendation method based on multi-dimensional vector fusion

    CN118227894A

  • Speech recognition method and system based on artificial intelligence

    CN120388575A