Voice interaction method and system, electronic equipment and storage medium

By converting the user's voice instructions into instruction vectors, matching them with the pre-constructed instruction matrix, and combining the recognition model to generate target instructions, the problems of speech recognition accuracy and response speed in the prior art are solved, and efficient and flexible voice interaction control is achieved.

CN119993162APending Publication Date: 2025-05-13CHONGQING JINKANG NEW ENERGY VEHICLE CO LTD

Patent Information

Application Number
CN202510375078.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Existing voice control technologies have limited recognition accuracy when diversifying and complexing voice input processing, and relying on cloud resources leads to slow response speed.

Method used

Using voice interaction method, by receiving user voice commands, performing text processing, inputting and identifying a large model, generating instruction vectors, matching with the pre-constructed instruction matrix, obtaining the target instruction index, and finally generating target instructions for controlling the vehicle by identifying the large model.

Benefits of technology

It improves the adaptability to diversified and complex scenarios, significantly reduces the use of storage and computing resources, and improves the generation speed of target instructions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993162A_ABST
    Figure CN119993162A_ABST
Patent Text Reader

Abstract

The invention provides a voice interaction method and system, electronic equipment and a storage medium, and relates to the technical field of man-machine interaction, and the method comprises the steps: receiving a voice instruction of a user, and carrying out the text processing of the voice instruction, and obtaining a first instruction text; inputting the first instruction text into an identification large model to obtain an instruction vector; matching the instruction vector with a pre-constructed instruction matrix to obtain a target instruction index; based on the target instruction index, calling a target instruction text from a pre-constructed instruction list; and performing text processing on the target instruction text and the first instruction text to obtain a query text, and inputting the query text into the identification large model to generate a target instruction for controlling the vehicle. According to the method, the generation speed of the target instruction is increased, and the adaptability to diversified and complicated scenes is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of human-computer interaction technology, and in particular to a voice interaction method, system, electronic device and storage medium. Background Art

[0002] In recent years, vehicle voice control technology has developed rapidly and has been widely used in fields such as intelligent driving and human-computer interaction. Existing voice control methods mainly include command recognition and database query based on traditional text matching. The former converts user voice into text through voice recognition technology, matches the preset command vocabulary, and uses the natural language processing model for feature extraction and slot filling, and finally generates executable instructions for controlling the vehicle. However, this method relies on a static vocabulary, is difficult to adapt to diverse and complex voice inputs, and has limited recognition accuracy in dynamic scenarios. The latter uses cloud databases such as Qdrant and Faiss to perform text retrieval based on the database, but due to its reliance on cloud resources, it has the problem of slow response speed.

[0003] Therefore, in view of the limitations of existing voice control technology, it is urgent to propose a new voice interaction method. Summary of the invention

[0004] In view of the above problems, the embodiments of the present application provide a voice interaction method, system, electronic device and storage medium to overcome or at least partially solve the problems of poor generalization of instruction recognition based on traditional text matching and slow speed of querying instructions through database.

[0005] In a first aspect of an embodiment of the present application, a voice interaction method is provided, the method comprising: Receiving a voice command from a user, and performing text processing on the voice command to obtain a first command text; Input the first instruction text into the recognition model to obtain an instruction vector; Matching the instruction vector with a pre-built instruction matrix to obtain a target instruction index; Based on the target instruction index, retrieve the target instruction text from a pre-built instruction list; The target instruction text and the first instruction text are subjected to text processing to obtain a query text, and the query text is input into the recognition model to generate a target instruction for controlling the vehicle.

[0006] Optionally, matching the instruction vector with a pre-built instruction matrix to obtain a target instruction index includes: Performing cosine similarity calculation on the instruction vector and each row of preset instruction vectors in the instruction matrix to obtain multiple cosine similarity values; Determining a maximum cosine similarity value from the multiple cosine similarity values; The first index of the preset instruction vector corresponding to the maximum cosine similarity value in the instruction matrix is ​​used as the target instruction index.

[0007] Optionally, the performing text processing on the voice instruction to obtain a first instruction text includes: Performing semantic recognition on the voice command to obtain text corresponding to the voice command; The text is generated according to the prompt instruction of the preset first prompt template to obtain the first instruction text, and the prompt instruction of the first prompt template is used to convert the text into vectorized text.

[0008] Optionally, the step of performing text processing on the target instruction text and the first instruction text to obtain a query text, and inputting the query text into the recognition model to generate a target instruction for controlling the vehicle includes: The target instruction text and the first instruction text are subjected to text generation according to a preset prompt instruction of a second prompt template to obtain the query text, wherein the prompt instruction of the second prompt template is used to generate text from the first instruction text according to a text paradigm of the target instruction text; Inputting the query text into the reasoning module of the recognition large model, so as to match the keywords corresponding to the query text through the reasoning module of the recognition large model, and outputting the vehicle control function according to the keywords; The target instruction is generated according to the vehicle control function.

[0009] Optionally, the method for constructing the instruction matrix includes: Acquire multiple preset vehicle control instructions, and respectively perform vectorization processing on the multiple preset vehicle control instructions to obtain preset instruction vectors corresponding to the multiple preset vehicle control instructions; The plurality of preset instruction vectors are spliced ​​together, and a first index is added to each of the spliced ​​preset instruction vectors in a preset order to obtain the instruction matrix, wherein the first index of each preset instruction vector is different.

[0010] Optionally, the method for constructing the instruction list includes: Obtaining the command text of each of the preset vehicle control commands; Based on the first index of the preset instruction vector, adding a second index to the instruction text of the preset vehicle control instruction, wherein the first index and the second index corresponding to the same preset vehicle control instruction are the same; Based on the order of the second indexes corresponding to the instruction texts of the plurality of preset vehicle control instructions, the instruction texts of all the preset vehicle control instructions are concatenated to obtain the instruction list.

[0011] Optionally, the step of inputting the first instruction text into a recognition model to obtain an instruction vector includes: Inputting the first instruction text into the embedding layer of the recognition large model, and outputting a text vectorization result through the embedding layer of the recognition large model; The text vectorization result is extracted through a hook function to obtain a multi-dimensional vector, and the multi-dimensional vector is averaged to obtain the instruction vector, where the multi-dimensional vector represents a vector containing features of multiple dimensions.

[0012] In a second aspect of the present application, a voice interaction system is provided, the system comprising: A first text processing module, used for receiving a user's voice instruction and performing text processing on the voice instruction to obtain a first instruction text; An input module, used for inputting the first instruction text into the recognition model to obtain an instruction vector; A matching module, used for matching the instruction vector with a pre-built instruction matrix to obtain a target instruction index; A calling module, used for calling a target instruction text from a pre-built instruction list based on the target instruction index; The second text processing module is used to perform text processing on the target instruction text and the first instruction text to obtain a query text, and input the query text into the recognition large model to generate a target instruction for controlling the vehicle.

[0013] Optionally, the instruction vector is matched with a pre-built instruction matrix to obtain a target instruction index, and the matching module includes: A calculation submodule, used for performing cosine similarity calculation on the instruction vector and each row of preset instruction vectors in the instruction matrix to obtain multiple cosine similarity values; A first determining submodule, configured to determine a maximum cosine similarity value from among the plurality of cosine similarity values; The second determining submodule is used to use the first index of the preset instruction vector corresponding to the maximum cosine similarity value in the instruction matrix as the target instruction index.

[0014] Optionally, the voice instruction is subjected to text processing to obtain a first instruction text, and the first text processing module includes: A semantic recognition submodule, used to perform semantic recognition on the voice command to obtain a text corresponding to the voice command; The first generating submodule is used to generate the text according to the prompt instruction of the preset first prompt template to obtain the first instruction text, and the prompt instruction of the first prompt template is used to convert the text into vectorized text.

[0015] Optionally, the target instruction text and the first instruction text are subjected to text processing to obtain a query text, and the query text is input into the recognition model to generate a target instruction for controlling the vehicle, and the second text processing module includes: A second generation submodule is used to generate text from the target instruction text and the first instruction text according to a preset prompt instruction of a second prompt template to obtain the query text, wherein the prompt instruction of the second prompt template is used to generate text from the first instruction text according to a text paradigm of the target instruction text; A first input submodule, used for inputting the query text into the inference module of the recognition large model, so as to match the keywords corresponding to the query text through the inference module of the recognition large model, and output the vehicle control function according to the keywords; The second generating submodule is used to generate the target instruction according to the vehicle control function.

[0016] Optionally, it also includes: A first acquisition submodule is used to acquire a plurality of preset vehicle control instructions, and respectively perform vectorization processing on the plurality of preset vehicle control instructions to obtain preset instruction vectors corresponding to the plurality of preset vehicle control instructions; The first splicing submodule is used to splice the multiple preset instruction vectors and add a first index to each of the spliced ​​preset instruction vectors in a preset order to obtain the instruction matrix, wherein the first index of each preset instruction vector is different.

[0017] Optionally, it also includes: A second acquisition submodule is used to acquire the instruction text of each of the preset vehicle control instructions; An adding submodule, used for adding a second index to the instruction text of the preset vehicle control instruction based on the first index of the preset instruction vector, wherein the first index and the second index corresponding to the same preset vehicle control instruction are the same; The second splicing submodule is used to splice the instruction texts of all preset vehicle control instructions based on the order of the second indexes corresponding to the instruction texts of the plurality of preset vehicle control instructions to obtain the instruction list.

[0018] Optionally, the first instruction text is input into a recognition model to obtain an instruction vector, and the input module includes: A second input submodule, used for inputting the first instruction text into the embedding layer of the recognition large model, and outputting a text vectorization result through the embedding layer of the recognition large model; The extraction submodule is used to extract the text vectorization result through a hook function to obtain a multi-dimensional vector, and average the multi-dimensional vector to obtain the instruction vector, wherein the multi-dimensional vector represents a vector containing features of multiple dimensions.

[0019] In a third aspect of the present application, an electronic device is provided, comprising a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction, when executed by the processor, implements the steps of the voice interaction method as described in the first aspect of the present application.

[0020] In a fourth aspect of the present application, a readable storage medium is provided, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the voice interaction method described in the first aspect of the present application are implemented. Beneficial effects of this application: The present application proposes a voice interaction method, which includes: first receiving a user's voice command, and performing text processing on the voice command to obtain a first command text; then inputting the first command text into a recognition large model to obtain a command vector; then matching the command vector with a pre-built command matrix to obtain a target command index; then, based on the target command index, retrieving the target command text from a pre-built command list; finally, performing text processing on the target command text and the first command text to obtain a query text, and inputting the query text into the recognition large model to generate a target command for controlling a vehicle. The present application uses the command matrix for retrieval to achieve efficient retrieval on resource-constrained end-side devices, significantly reduces storage and computing resource usage, and effectively improves the speed of generating target commands. In addition, by inputting the query text into the recognition large model, the semantic generalization ability of the recognition large model itself can be used to generate target commands for controlling vehicles, thereby improving the adaptability to diverse and complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the description of the embodiments of the present application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0022] Figure 1 This is a schematic diagram of the steps of a voice interaction method provided in an embodiment of the present application; Figure 2 It is a schematic diagram of a construction process of an instruction matrix and an instruction list provided in an embodiment of the present application; Figure 3 It is a flowchart of a voice interaction method provided in an embodiment of the present application; Figure 4 is a schematic diagram of a voice interaction system provided in an embodiment of the present application; Figure 5 It is a schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0023] The exemplary embodiments of the present application will be described in more detail below in conjunction with the accompanying drawings in the embodiments of the present application. Although the exemplary embodiments of the present application are shown in the accompanying drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided in order to enable a more thorough understanding of the present application and to enable the scope of the present application to be fully communicated to those skilled in the art.

[0024] In a first aspect of the present application, a voice interaction method is provided. Figure 1 As shown, the method includes: Step S101, receiving a user's voice command, and performing text processing on the voice command to obtain a first command text.

[0025] In this step, the user's voice command is first received, and the voice signal is usually captured by a voice acquisition device such as a microphone integrated in the vehicle. In order to ensure high-quality voice input, in some embodiments, multi-microphone array technology can be used, combined with signal processing algorithms such as noise suppression, echo cancellation and voice enhancement, to improve the accuracy of voice recognition, especially to adapt to the complex noise environment during vehicle driving.

[0026] Subsequently, the collected voice signal is subjected to text processing, and after the voice-to-text conversion is completed, the initially recognized text can also be preprocessed, including removing noise words, correcting errors and other text processing, to ensure that the first instruction text with clear structure and clear semantics is obtained.

[0027] Step S102: input the first instruction text into a recognition model to obtain an instruction vector.

[0028] In step S102, the first instruction text obtained in step S101 is input into the recognition model, and the first instruction text is vectorized by the recognition model to obtain an instruction vector. The recognition model can be a pre-trained language model based on a deep learning architecture, which has powerful semantic understanding and feature extraction capabilities.

[0029] Specifically, the recognition model first performs word segmentation on the first input instruction text and converts it into a word vector. It then maps each word in the text into a multidimensional vector space through an embedding layer to capture the semantic associations between words. Subsequently, the recognition model deeply encodes the word vector through a multi-layer neural network, extracts the contextual information and deep semantic features of the text, and generates an overall instruction vector.

[0030] In this step, by using the currently mature large recognition model, there is no need for additional model training for specific instructions, which reduces development and maintenance costs and shortens the system deployment cycle. In addition, relying on the excellent semantic generalization ability of the large recognition model itself, it can effectively improve the recognition and understanding ability of diversified and complex super control instructions, and enhance the adaptability and stability of the system in changing scenarios.

[0031] Step S103, matching the instruction vector with a pre-built instruction matrix to obtain a target instruction index.

[0032] In step S103, the instruction vector generated by the recognition model is matched with the pre-built instruction matrix to determine the target instruction index that best meets the user's intention.

[0033] In some embodiments, the instruction matrix is ​​a multidimensional vector set generated by vectorizing multiple instruction texts using the same recognition model. These instruction vectors are stored in a specific order to form a structured instruction matrix, wherein each multidimensional vector in the instruction matrix is ​​added with an index. By matching the instruction vector with the instruction matrix, its index in the instruction matrix is ​​determined as the target instruction index.

[0034] In the following embodiments, in order to further improve the matching efficiency, vector normalization and dimensionality reduction techniques may be used to optimize the density of the vector space and reduce the computational complexity.

[0035] Step S104: based on the target instruction index, retrieve the target instruction text from a pre-built instruction list.

[0036] In step S104, based on the target instruction index determined in step S103, the corresponding target instruction text is retrieved from the pre-built instruction list.

[0037] The instruction list is a structured data set pre-built according to vehicle control requirements, containing various standardized control instruction texts, such as "open the window", "turn off the air conditioner", "start navigation", etc. Each instruction text corresponds to the vector in the instruction matrix one by one and is stored in the same order to ensure the accuracy and efficiency of index matching.

[0038] During the retrieval process, the corresponding instruction text is quickly located according to the target instruction index to ensure the efficiency and low latency of the retrieval process.

[0039] In addition, to improve the flexibility of the system, the instruction matrix and instruction list can support dynamic updates. When adding or optimizing instructions, you only need to update the instruction list and instruction matrix synchronously. There is no need to adjust the recognition model or reconstruct the system process, which reduces the update and maintenance costs.

[0040] Step S105, performing text processing on the target instruction text and the first instruction text to obtain a query text, and inputting the query text into the recognition large model to generate a target instruction for controlling the vehicle.

[0041] In step S105, the retrieved target instruction text and the initial first instruction text are processed to generate a more accurate query text, thereby ensuring that the system can accurately identify and execute the user's intention.

[0042] The text processing process may include operations such as text splicing, semantic alignment, keyword extraction, and format standardization. For example, the first instruction text can be combined with the target instruction text to extract the core semantics, remove redundant information, and optimize the existing fuzzy or ambiguous content to ensure that the query text is clear and accurate, which is convenient for subsequent model recognition.

[0043] The generated query text is then input into the recognition model, which further extracts the user's specific control intention through the deep semantic understanding ability of the recognition model, and finally generates a structured and standardized target control instruction. The instruction has clear execution parameters and format, which is convenient for the vehicle control system to directly recognize and execute.

[0044] This step not only improves the accuracy of instruction generation and reduces the risk of misjudgment and misexecution, but also enhances the system's semantic generalization ability and adaptability, ensuring that the system can accurately handle diverse and complex user instructions.

[0045] The present application obtains the target instruction index corresponding to the user's first instruction text through a pre-built instruction matrix, and determines the target instruction text from the pre-built instruction list based on the target instruction index, and then obtains the target instruction for controlling the vehicle after text processing of the target instruction text and the first instruction text. The present application uses the instruction matrix for retrieval to achieve efficient retrieval on resource-constrained terminal devices, significantly reduces storage and computing resource usage, and effectively improves the generation speed of target instructions. In addition, by inputting the query text into the recognition large model, the semantic generalization ability of the recognition large model itself can be used to generate target instructions for controlling the vehicle, thereby improving the adaptability to diversified and complex scenarios.

[0046] In one embodiment, matching the instruction vector with a pre-built instruction matrix to obtain a target instruction index includes: Performing cosine similarity calculation on the instruction vector and each row of preset instruction vectors in the instruction matrix to obtain multiple cosine similarity values; Determining a maximum cosine similarity value from the multiple cosine similarity values; The first index of the preset instruction vector corresponding to the maximum cosine similarity value in the instruction matrix is ​​used as the target instruction index.

[0047] In this embodiment, the cosine similarity calculation is performed one by one between the instruction vector generated by the identification large model and each row of the preset instruction vector in the instruction matrix. Cosine similarity is used to measure the similarity between two vectors in the vector space, and the value range is between -1 and 1. The closer the value is to 1, the higher the similarity. After completing all cosine similarity calculations, the cosine similarity value with the largest value is selected from the multiple cosine similarity values ​​obtained, and it is considered that the preset instruction vector corresponding to the value is closest to the input instruction vector. The first index of the preset instruction vector corresponding to the maximum cosine similarity value in the instruction matrix is ​​determined as the target instruction index.

[0048] This embodiment has the advantages of high computing efficiency and high matching accuracy, and is suitable for the efficient instruction matching requirements of the terminal device with low resource consumption. At the same time, by utilizing the characteristics of cosine similarity, it can effectively process instruction texts with similar semantics but different expressions, and improve the semantic generalization and matching capabilities of the system.

[0049] In one embodiment, a method for calculating instruction similarity is provided. First, for each row vector in a pre-constructed instruction matrix, , and compare it with the vector to be matched Perform dot product calculation, where the dot product formula (1) is as follows: (1) in, Representation vector In the The components in the dimensions, Represents the matrix Line The value of a dimension.

[0050] Next, calculate the vectors And each row vector in the matrix The norm of : For vector The norm of is shown in the following formula (2): (2) For each row vector in the matrix , its norm is shown in the following formula (3): (3) Finally, using the dot product obtained above and their respective norms, we calculate the vector With each row vector The cosine similarity of is shown in the following formula (4): (4) In one embodiment, the text processing of the voice instruction to obtain the first instruction text includes: Performing semantic recognition on the voice command to obtain text corresponding to the voice command; The text is generated according to the prompt instruction of the preset first prompt template to obtain the first instruction text, and the prompt instruction of the first prompt template is used to convert the text into vectorized text.

[0051] In this embodiment, the user's voice command is first received, and the voice data is processed using voice recognition technology to extract the corresponding text information. In some cases, the context perception mechanism can also be combined to optimize the recognition results through context information to improve accuracy. For example, when the user uses colloquial expressions, the grammar can be automatically corrected or missing words can be filled in to make the text more fluent and understandable. Then, the recognized text is input into the preset first prompt template, and the text is normalized according to the prompt instructions in the template to generate the first instruction text. The design purpose of the prompt template is to guide the system to uniformly format the text content and improve the semantic clarity and consistency. The prompt instructions in the first prompt template are specifically used to optimize the vectorization effect of the text, ensuring that the generated first instruction text has a clear and standard expression form, which is conducive to the subsequent recognition large model for efficient and accurate vector conversion and semantic understanding.

[0052] This embodiment can not only accurately recognize and normalize the user's voice input, but also enhance the semantic clarity of the text, ensure the accuracy and stability of vector matching and instruction generation in subsequent steps, and improve the overall voice interaction experience.

[0053] In one embodiment, the step of performing text processing on the target instruction text and the first instruction text to obtain a query text, and inputting the query text into the recognition model to generate a target instruction for controlling the vehicle includes: The target instruction text and the first instruction text are subjected to text generation according to a preset prompt instruction of a second prompt template to obtain the query text, wherein the prompt instruction of the second prompt template is used to generate text from the first instruction text according to a text paradigm of the target instruction text; Inputting the query text into the reasoning module of the recognition large model, so as to match the keywords corresponding to the query text through the reasoning module of the recognition large model, and outputting the vehicle control function according to the keywords; The target instruction is generated according to the vehicle control function.

[0054] In this embodiment, the target instruction text is combined with the first instruction text, and the text is generated through a preset second prompt template. The second prompt template specifies how to generate a query text based on the paradigm of the target instruction text. This ensures that the structure and content of the query text meet the requirements of the vehicle control task. Then, the generated query text is input into the reasoning module of the recognition model. In order to extract feature words from the query text through the reasoning module, these feature words represent the key information of the text. In this application, the reasoning module may include an attention mechanism, a normalization layer, and a regularization layer.

[0055] Based on the feature words extracted from the query text, the vehicle control functions corresponding to the output of the large model are identified. These vehicle control functions are instructions or decision logic used to specifically control the behavior of the vehicle. Finally, the vehicle control functions are used to generate actual vehicle control target instructions, which will perform specific operations based on the current state of the vehicle or environmental conditions.

[0056] Among them, the text paradigm of the target instruction has a fixed text format, which can convert the expression that is relatively general or does not fully match the user's needs into the input standard that meets the vehicle control task. For example: the first instruction text is: "Open the left window", and the query text converted by the second prompt template is: "Execute the window control instruction, the left window state = open", or the first instruction text is: "I am a little cold, turn up the air conditioner", and the query text converted by the second prompt template is: "Temperature adjustment instruction, target: increase the temperature in the car by 2℃", etc.

[0057] For example, assume that the user enters the command: "Set the temperature inside the car to 16°C" The recall function is filled in as follows: adjust_air_conditioning(temperature=int), and the int value is 16-32.5 degrees. The input of the recognition model is: the user input sets the temperature in the car to 16, the Ha number function is adjust_air_conditioning(temperature=int), and the int value is 16-32.5 degrees The final output of the recognition model is: adjust_air_conditioning(temperature=16) This embodiment combines the feature extraction capabilities of text processing and recognition large models, and can generate target instructions related to vehicle control based on the input instruction text.

[0058] In one embodiment, the method for constructing the instruction matrix includes: Acquire multiple preset vehicle control instructions, and respectively perform vectorization processing on the multiple preset vehicle control instructions to obtain preset instruction vectors corresponding to the multiple preset vehicle control instructions; The plurality of preset instruction vectors are spliced ​​together, and a first index is added to each of the spliced ​​preset instruction vectors in a preset order to obtain the instruction matrix, wherein the first index of each preset instruction vector is different.

[0059] In this embodiment, refer to Figure 2 The schematic diagram of the construction process of the instruction matrix and instruction list shown in the figure first obtains multiple preset vehicle control instructions, which are standardized and suitable vehicle control instructions. For each preset vehicle control instruction obtained, a pre-designed vectorization algorithm is used for conversion, and the natural language description or structured control command is converted into an instruction vector of fixed dimension. Then, the preset instruction vectors corresponding to the multiple preset vehicle control instructions obtained after vectorization processing are spliced ​​in a preset order to obtain an N*M matrix, where N is the number of preset vehicle control instructions and M is the dimension of each converted preset instruction vector. This splicing operation not only realizes the collection of multiple preset instruction vectors, but also maintains the order relationship between each vector, which is convenient for subsequent index addition and retrieval. In the spliced ​​instruction matrix, the first index is added to the preset instruction vector of each preset vehicle control instruction according to the preset order. The first index of each preset instruction vector is used as the unique identifier of the preset instruction vector, so that the index values ​​of each preset instruction vector are different in the entire instruction matrix. This operation not only helps to improve the efficiency of subsequent data matching and retrieval, but also provides an accurate basis for instruction positioning and result feedback.

[0060] The instruction matrix constructed by the above method in this embodiment can effectively integrate the vector information of multiple preset vehicle control instructions to achieve efficient and accurate instruction matching and calling. At the same time, the unique first index ensures the distinction of each preset vehicle control instruction in the matrix, providing effective data support for the real-time response of the system in complex application scenarios.

[0061] In one embodiment, the method for constructing the instruction list includes: Obtaining the command text of each of the preset vehicle control commands; Based on the first index of the preset instruction vector, adding a second index to the instruction text of the preset vehicle control instruction, wherein the first index and the second index corresponding to the same preset vehicle control instruction are the same; Based on the order of the second indexes corresponding to the instruction texts of the plurality of preset vehicle control instructions, the instruction texts of all the preset vehicle control instructions are concatenated to obtain the instruction list.

[0062] In this embodiment, continue to refer to Figure 2 , obtain the text description of each preset vehicle control instruction, and after the preset vehicle control instruction is vectorized, each preset instruction vector of each preset vehicle control instruction has a unique first index, and add a second index to the instruction text of each preset vehicle control instruction according to the first index of the preset instruction vector corresponding to it. In this embodiment, the second index is the same as the corresponding first index, which ensures the consistency and correspondence of the same preset vehicle control instruction at the two levels of vectorization and text description, which is convenient for subsequent data processing and tracing.

[0063] According to the order of the second indexes respectively corresponding to the instruction texts of the plurality of preset vehicle condition instructions, all the instruction texts are sequentially spliced ​​to form a complete instruction list.

[0064] The instruction list constructed by the above method in this embodiment not only realizes the seamless connection between the text and vectorized data of the vehicle control instructions, but also ensures the orderliness and consistency of the instruction data. In addition, by using the index between the instruction matrix and the instruction list for retrieval, the security of information transmission is improved.

[0065] In one embodiment, the step of inputting the first instruction text into a recognition model to obtain an instruction vector includes: Inputting the first instruction text into the embedding layer of the recognition large model, and outputting a text vectorization result through the embedding layer of the recognition large model; The text vectorization result is extracted through a hook function to obtain a multi-dimensional vector, and the multi-dimensional vector is averaged to obtain the instruction vector, where the multi-dimensional vector represents a vector containing features of multiple dimensions.

[0066] In this embodiment, the first instruction text is first taken as input and sent to the embedding layer of the recognition model. The embedding layer is responsible for converting the input text into a multidimensional vector, wherein the multidimensional vector represents a vector containing features of multiple dimensions, so that it can represent the semantic information of the text in the vector space. This process can not only capture the semantic relationship of the text, but also retain the contextual relevance and syntactic structure features, providing a semantically rich representation for subsequent vector calculations. In the embedding layer, the recognition model is used for text vectorization. The model extracts text features through a deep neural network and maps the text to a multidimensional vector space, so that texts with similar semantics are closer in the space. The embedding layer can effectively handle synonyms, polysemous words, and context-related expressions to ensure that natural language instructions of different users can be correctly converted into standardized vector representations.

[0067] Then, after the text is vectorized in the embedding layer, the output of the model in this layer is captured in real time using the preset hook function. The hook function is a mechanism that can dynamically obtain the output of a specified layer during the neural network calculation process, avoiding changes to the overall structure of the model and improving calculation efficiency. This method ensures that the embedding vector of the text in the high-dimensional space is directly obtained without affecting the normal operation of the model.

[0068] The vector extracted by the hook function is usually a multi-dimensional vector, in which each dimension contains semantic information at different levels, such as vocabulary-level, phrase-level, and sentence-level features. In addition, this method can maintain a consistent data acquisition method under different language models or different training methods, and can flexibly adapt to different recognition models.

[0069] Since the multidimensional vectors generated by the model may contain redundant information or be too sparse, this embodiment uses mean calculation to normalize the vector values ​​of all dimensions. The specific steps of mean calculation are as follows: Get the embedding vector matrix: For the input text, the embedding layer may output multiple vectors of different dimensions, such as word-level vectors or sentence-level vectors.

[0070] Calculate the mean of each dimension: Calculate the mean of each dimension of all vectors so that the final output instruction vector can maintain a stable feature representation in each dimension.

[0071] Normalization: Perform L2 regularization or other normalization operations on the obtained vector to ensure the stability of the vector in different models and computing environments and improve the matching accuracy.

[0072] In one embodiment, there is provided a Figure 3 The flow chart of the voice interaction method shown in FIG. Figure 3 As shown: ① Obtaining a voice command, and after voice-to-text processing, obtaining a text corresponding to the voice command, performing text processing on the text and the first prompt template to generate a first command text for vectorization processing, wherein the text processing includes text concatenation and text regularization processing, etc., to remove modal particles and invalid punctuation entered by the user; ②In the inference program of the large recognition model, a hook function is implanted. This function extracts the text vectorization result when the embedding layer of the large recognition model outputs the result. After the first instruction text is input into the model, the embedding layer converts the text into a vector. The hook function captures the multi-dimensional vector and obtains the instruction vector to be processed by taking the mean value, preparing for the next similarity calculation; ③ Calculate the cosine similarity between the instruction vector obtained in step ② and the pre-built instruction matrix to obtain the target instruction index; ④ According to the obtained target instruction index, call the corresponding target instruction text in the pre-built instruction list; ⑤ The target instruction text retrieved in step ④ is spliced ​​with the first instruction text in step ① according to the instruction of the second prompt module to form a query text that guides the recognition large model to complete the instruction slot filling function; ⑥ Input the query text generated in step ⑤ into the recognition model, and the recognition model completes keyword matching according to the instruction guidance and outputs the final executable vehicle control function; ⑦ Finally, the vehicle control function is converted into target instructions to realize vehicle command control.

[0073] The present application proposes a voice interaction method, the method comprising: first receiving a user's voice command, and performing text processing on the voice command to obtain a first command text; then inputting the first command text into a recognition large model to obtain a command vector; then matching the command vector with a pre-built command matrix to obtain a target command index; then, based on the target command index, retrieving the target command text from a pre-built command list; finally, performing text processing on the target command text and the first command text to obtain a query text, and inputting the query text into the recognition large model to generate a target command for controlling a vehicle. The present application obtains a target command index through a pre-built command matrix, and determines a target command text from a pre-built command list based on the target command index, and then obtains a target command for controlling a vehicle after text processing of the target command text. The present application uses the command matrix for retrieval, thereby achieving efficient retrieval on a resource-constrained end-side device, significantly reducing the storage and computing resource usage, and effectively improving the speed of generating target commands. In addition, by inputting the query text into the recognition large model, the semantic generalization ability of the recognition large model itself can be utilized to generate target commands for controlling a vehicle, thereby improving the adaptability to diverse and complex scenarios.

[0074] Based on the same inventive concept, the second aspect of the present application provides a voice interaction system, such as Figure 4 As shown, the system comprises: The first text processing module 201 is used to receive a user's voice command and perform text processing on the voice command to obtain a first command text; An input module 202, used for inputting the first instruction text into a recognition model to obtain an instruction vector; A matching module 203 is used to match the instruction vector with a pre-built instruction matrix to obtain a target instruction index; A calling module 204, configured to call a target instruction text from a pre-built instruction list based on the target instruction index; The second text processing module 205 is used to perform text processing on the target instruction text and the first instruction text to obtain a query text, and input the query text into the recognition large model to generate a target instruction for controlling the vehicle.

[0075] Optionally, the instruction vector is matched with a pre-built instruction matrix to obtain a target instruction index, and the matching module 203 includes: A calculation submodule, used for performing cosine similarity calculation on the instruction vector and each row of preset instruction vectors in the instruction matrix to obtain multiple cosine similarity values; A first determining submodule, configured to determine a maximum cosine similarity value from among the plurality of cosine similarity values; The second determining submodule is used to use the first index of the preset instruction vector corresponding to the maximum cosine similarity value in the instruction matrix as the target instruction index.

[0076] Optionally, the voice instruction is subjected to text processing to obtain a first instruction text, and the first text processing module 201 includes: A semantic recognition submodule, used to perform semantic recognition on the voice command to obtain a text corresponding to the voice command; The first generating submodule is used to generate the text according to the prompt instruction of the preset first prompt template to obtain the first instruction text, and the prompt instruction of the first prompt template is used to convert the text into vectorized text.

[0077] Optionally, the target instruction text and the first instruction text are subjected to text processing to obtain a query text, and the query text is input into the recognition model to generate a target instruction for controlling the vehicle. The second text processing module 205 includes: A second generation submodule is used to generate text from the target instruction text and the first instruction text according to a preset prompt instruction of a second prompt template to obtain the query text, wherein the prompt instruction of the second prompt template is used to generate text from the first instruction text according to a text paradigm of the target instruction text; A first input submodule, used for inputting the query text into the inference module of the recognition large model, so as to match the keywords corresponding to the query text through the inference module of the recognition large model, and output the vehicle control function according to the keywords; The second generating submodule is used to generate the target instruction according to the vehicle control function.

[0078] Optionally, it also includes: A first acquisition submodule is used to acquire a plurality of preset vehicle control instructions, and respectively perform vectorization processing on the plurality of preset vehicle control instructions to obtain preset instruction vectors corresponding to the plurality of preset vehicle control instructions; The first splicing submodule is used to splice the multiple preset instruction vectors and add a first index to each of the spliced ​​preset instruction vectors in a preset order to obtain the instruction matrix, wherein the first index of each preset instruction vector is different.

[0079] Optionally, it also includes: A second acquisition submodule is used to acquire the instruction text of each of the preset vehicle control instructions; An adding submodule, used for adding a second index to the instruction text of the preset vehicle control instruction based on the first index of the preset instruction vector, wherein the first index and the second index corresponding to the same preset vehicle control instruction are the same; The second splicing submodule is used to splice the instruction texts of all preset vehicle control instructions based on the order of the second indexes corresponding to the instruction texts of the plurality of preset vehicle control instructions to obtain the instruction list.

[0080] Optionally, the first instruction text is input into a recognition model to obtain an instruction vector, and the input module 202 includes: A second input submodule, used for inputting the first instruction text into the embedding layer of the recognition large model, and outputting a text vectorization result through the embedding layer of the recognition large model; The extraction submodule is used to extract the text vectorization result through a hook function to obtain a multi-dimensional vector, and average the multi-dimensional vector to obtain the instruction vector, wherein the multi-dimensional vector represents a vector containing features of multiple dimensions.

[0081] Based on the same inventive concept, in a third aspect of the present application, an electronic device 100 is provided, such as Figure 5 As shown, it includes a processor 120, a memory 110, and a program or instruction stored in the memory 110 and executable on the processor 120. When the program or instruction is executed by the processor 120, the steps of the voice interaction method described in the first aspect of the present application are implemented.

[0082] Based on the same inventive concept, the fourth aspect of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the voice interaction method described in the first aspect of the present application are implemented. Each embodiment in this specification focuses on the differences from other embodiments, and the same or similar parts between the embodiments can be referenced to each other.

[0083] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, devices, or computer program products. Therefore, the embodiments of the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the embodiments of the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0084] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0085] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0086] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable terminal device. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0087] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present application.

[0088] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or terminal device including the elements.

[0089] The above is a detailed introduction to the provided voice interaction method, system, electronic device and storage medium. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for general technical personnel in this field, according to the idea of ​​the present application, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. A voice interaction method, characterized in that: The method comprises: Receiving a voice command from a user, and performing text processing on the voice command to obtain a first command text; Inputting the first instruction text into a recognition model to obtain an instruction vector; Matching the instruction vector with a pre-built instruction matrix to obtain a target instruction index; Based on the target instruction index, retrieve the target instruction text from a pre-built instruction list; The target instruction text and the first instruction text are subjected to text processing to obtain a query text, and the query text is input into the recognition model to generate a target instruction for controlling the vehicle.

2. The voice interaction method according to claim 1, characterized in that: The step of matching the instruction vector with a pre-built instruction matrix to obtain a target instruction index includes: Performing cosine similarity calculation on the instruction vector and each row of preset instruction vectors in the instruction matrix to obtain multiple cosine similarity values; Determining a maximum cosine similarity value from the multiple cosine similarity values; The first index of the preset instruction vector corresponding to the maximum cosine similarity value in the instruction matrix is ​​used as the target instruction index.

3. The voice interaction method according to claim 1, characterized in that: The text processing of the voice instruction to obtain a first instruction text includes: Performing semantic recognition on the voice command to obtain text corresponding to the voice command; The text is generated according to the prompt instruction of the preset first prompt template to obtain the first instruction text, and the prompt instruction of the first prompt template is used to convert the text into vectorized text.

4. The voice interaction method according to claim 3, characterized in that: The step of performing text processing on the target instruction text and the first instruction text to obtain a query text, and inputting the query text into the recognition model to generate a target instruction for controlling the vehicle includes: The target instruction text and the first instruction text are subjected to text generation according to a preset prompt instruction of a second prompt template to obtain the query text, wherein the prompt instruction of the second prompt template is used to generate text from the first instruction text according to a text paradigm of the target instruction text; Inputting the query text into the reasoning module of the recognition large model, so as to match the keywords corresponding to the query text through the reasoning module of the recognition large model, and outputting the vehicle control function according to the keywords; The target instruction is generated according to the vehicle control function.

5. The voice interaction method according to any one of claims 1 to 4, characterized in that: The method for constructing the instruction matrix comprises: Acquire a plurality of preset vehicle control instructions, and respectively perform vectorization processing on the plurality of preset vehicle control instructions to obtain preset instruction vectors corresponding to the plurality of preset vehicle control instructions; The plurality of preset instruction vectors are spliced ​​together, and a first index is added to each of the spliced ​​preset instruction vectors in a preset order to obtain the instruction matrix, wherein the first index of each preset instruction vector is different.

6. The voice interaction method according to claim 5, characterized in that: The method for constructing the instruction list includes: Obtaining the command text of each of the preset vehicle control commands; Based on the first index of the preset instruction vector, adding a second index to the instruction text of the preset vehicle control instruction, wherein the first index and the second index corresponding to the same preset vehicle control instruction are the same; Based on the order of the second indexes corresponding to the instruction texts of the plurality of preset vehicle control instructions, the instruction texts of all the preset vehicle control instructions are concatenated to obtain the instruction list.

7. The voice interaction method according to claim 1, characterized in that: The step of inputting the first instruction text into a recognition model to obtain an instruction vector includes: Inputting the first instruction text into the embedding layer of the recognition large model, and outputting a text vectorization result through the embedding layer of the recognition large model; The text vectorization result is extracted through a hook function to obtain a multi-dimensional vector, and the multi-dimensional vector is averaged to obtain the instruction vector, where the multi-dimensional vector represents a vector containing features of multiple dimensions.

8. A voice interaction system, characterized in that: The system comprises: A first text processing module, used for receiving a user's voice instruction and performing text processing on the voice instruction to obtain a first instruction text; An input module, used for inputting the first instruction text into the recognition model to obtain an instruction vector; A matching module, used for matching the instruction vector with a pre-built instruction matrix to obtain a target instruction index; A calling module, used for calling a target instruction text from a pre-built instruction list based on the target instruction index; The second text processing module is used to perform text processing on the target instruction text and the first instruction text to obtain a query text, and input the query text into the recognition large model to generate a target instruction for controlling the vehicle.

9. An electronic device, characterized in that: It includes a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction, when executed by the processor, implements the steps of the voice interaction method as described in any one of claims 1 to 7.

10. A readable storage medium storing a program or an instruction, characterized in that: When the program or instruction is executed by the processor, the steps of the voice interaction method as described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Method and device for releasing advertisements on basis of search engines, and search engine system

    CN108280689A

  • Information extraction method and equipment and storage medium

    CN110797012A

  • Audio signal identification method, device and system, equipment and readable medium

    CN111429903A

  • Voice-based equipment operation method and system

    CN113948069A

  • Semantic matching model design method

    CN114298056A

Cited By

  • Voice interaction method and device and intelligent terminal

    CN122083986A