Intelligent voice interconnection method for automobile cabin control
Through intelligent voice interconnection methods, the problem of difficulty in understanding voice commands in the car cockpit is solved in traditional systems in complex acoustic environments, and more efficient and reliable voice interaction is achieved, improving driving safety and user experience.
Patent Information
- Application Number
- CN202510496272.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-04-21
AI Technical Summary
In the complex acoustic environment of the car cockpit, traditional on-board voice interaction systems are difficult to effectively capture and understand voice commands, resulting in drivers being distracted during long-distance driving or complex road conditions, and traditional control methods are difficult to meet consumers' pursuit of convenience and comfort.
The intelligent voice interconnection method is adopted to capture voice signals through microphone array technology, and perform denoising, echo cancellation and voice enhancement processing. The linear predictive cepspectral coefficient algorithm is used to extract speech features, combine speech recognition model and natural language understanding module, decode speech signals, convert them into vehicle control instructions, and feedback results are generated through speech synthesis.
It reduces the interference of ambient noise and echoes, improves the signal-to-noise ratio and clarity of the voice signal, ensures the accuracy of the voice signal, provides a reliable foundation for subsequent processing, enhances the interactive experience between users and cars, improves the usability and reliability of the system, improves driving safety, and improves the overall satisfaction of users.
Smart Images

Figure CN120048262A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent voice technology, and specifically relates to an intelligent voice interconnection method for vehicle cockpit control. Background Art
[0002] With the rapid development of intelligent cockpit technology, in-vehicle voice interaction systems have become the core configuration for enhancing driving safety and optimizing the human-machine interaction experience. In the complex acoustic environment of the vehicle cockpit, accurately capturing and understanding voice commands through microphone array technology is the basic link for building a complete voice ecosystem. With the rapid development of technology, intelligence has become one of the core trends in the modern automotive industry. In this context, intelligent voice interconnection methods, as an innovative interaction method, are gradually changing people's driving experiences.
[0003] In the fast-paced modern life, people are increasingly pursuing efficient and convenient travel methods. Traditional vehicle control methods, such as manual buttons and knobs, although stable and reliable, may seem cumbersome and distract the driver's attention in some cases. Especially during long-distance driving or in complex road conditions, the driver needs to be more focused on the road conditions, and the intelligent voice interconnection method can effectively reduce this burden. In addition, with the continuous increase in consumers' demands for vehicle intelligence and personalization, traditional control methods have been difficult to meet market demands. The intelligent voice interconnection method captures the user's voice commands and converts them into vehicle control commands, achieving a more intuitive and natural interaction method, meeting consumers' pursuit of convenience and comfort. Summary of the Invention
[0004] To solve the above technical problems, an intelligent voice interconnection method for vehicle cockpit control is provided, and this technical solution solves the problems raised in the above background art.
[0005] To achieve the above object, the technical solution adopted by the present invention is as follows: An intelligent voice interconnection method for vehicle cockpit control, including: Capturing the user's voice signal through the vehicle cockpit microphone array technology and converting it into a digital signal, and the intelligent voice interconnection system performs denoising, echo cancellation, and voice enhancement processing on the digital signal; The preprocessed digital signal is input into the feature extraction module, and the linear prediction cepstral coefficient algorithm is used to extract the voice features in the digital signal; The extracted voice features are input into the speech recognition model, and the digital signal is decoded based on the acoustic model, mapping the input voice features to the corresponding phonemes or words, and using the word-to-text mapping relationship provided by the dictionary to output the corresponding text content; Input the text content into the natural language understanding module, perform lexical analysis on the text content, split the text into words and phrases, and determine their parts of speech; Based on the fuzzy instruction parsing of the conversation history, perform syntactic analysis and semantic analysis on the text content, analyze the grammatical relationships between words and phrases, and parse the user's intentions and requirements; Based on the user's intentions and requirements, the intelligent voice interconnection system parses specific instructions and operations, and converts the natural language understanding result into a vehicle control instruction; During and after the execution of the instruction, the intelligent voice interconnection system feeds back relevant information or results to the user through speech synthesis technology.
[0006] Preferably, inputting the extracted speech features into the speech recognition model, decoding the digital signal based on the acoustic model, mapping the input speech features to the corresponding phonemes or words, and using the mapping relationship from words to text provided by the dictionary to output the corresponding text content specifically includes: Input the extracted speech feature sequence into the trained acoustic model, align the feature sequence frame by frame, and retain the time dimension information; Convert the frame-level output to a phoneme sequence through a blank symbol and repetition merging mechanism; Use an encoder-decoder structure to align the acoustic features with the phoneme sequence; Obtain the mapping table from phonemes to words and record it as a dictionary, perform dictionary constraint processing on the dictionary, and only retain the phoneme combinations that can form valid words; Calculate the matching probability between the current phoneme and the speech features and output it as an acoustic score; Perform weighted fusion of the acoustic score and the language model score; Obtain the global optimal path through the A* algorithm, use frames as nodes, phoneme transitions as edges, and the total score as the edge weight to construct a search space; Design a heuristic function to estimate the minimum cost from the current node to the end point, and preferentially expand the node with the lowest sum of the minimum cost and the actual cost; Judge whether the end of the sequence is reached. If so, end the global optimal path search. If not, do not make an output; Restore out-of-vocabulary words in the dictionary through subword units encoded by byte pair encoding; Obtain a punctuation prediction model, input the phoneme sequence and context, and output the punctuation positions; Insert capital letters and line breaks based on the text structure to generate formatted text and complete the conversion from speech to text.
[0007] Preferably, the syntactic analysis and semantic analysis processing of the text content based on the fuzzy instruction parsing of the conversation history, analyzing the grammatical relationships between words and phrases, and parsing the user's intentions and requirements specifically include: Maintain a dialogue state machine to record the user instructions, system responses, and slot filling results of historical turns; Use a neural network model to identify the pronoun references in the user's speech and construct an entity co-reference chain; Copy the unknown parameters not mentioned in the user's speech based on historical actions, where the unknown parameters include time parameters and location parameters; Update the domain dictionary in real time based on the user's speech and cache the context-related entities; Construct a phrase structure tree, identify complex nested structures, and perform pruning guided by the dialogue history, restricting candidate structures according to historical actions; Map the current parsing result to the dialogue state variables; Query the knowledge graph to verify the existence of entities, filter unreasonable requests based on business rules, and ensure that the current request is consistent with the dialogue goal; Use multi-task learning to jointly optimize the domain and sub-intentions; Inherit the dialogue history parameters, perform demand conflict detection, and generate corresponding guiding questions after discovering conflicting parameters; Use the user's clarification result as new training data and optimize the parameter extraction strategy using reinforcement learning; Optimize the slot filling model in combination with the user click log and analyze the user's speech stress features to improve intent recognition.
[0008] Compared with the prior art, the beneficial effects of the present invention are as follows: Reduce the interference of environmental noise and echo, improve the signal-to-noise ratio and clarity of the speech signal, ensure the accuracy and clarity of the speech signal, provide a reliable basis for subsequent processing, extract speech features in the digital signal using the linear prediction cepstral coefficient algorithm, which is crucial for subsequent speech recognition and natural language understanding, enhance the interaction experience between the user and the car, improve the usability and reliability of the system, improve driving safety, and enhance the overall satisfaction of the user. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Figure 1 It is a flowchart of the intelligent voice interconnection method for vehicle cockpit control according to the present invention; Figure 2 It is a flowchart of the method for denoising, echo cancellation, and speech enhancement processing of digital signals according to the present invention; Figure 3 It is a diagram of extracting speech features in digital signals using the linear prediction cepstral coefficient algorithm according to the present invention; Figure 4 It is a flowchart of decoding digital signals and mapping the input speech features to corresponding phonemes or words according to the present invention; Figure 5Flowchart of the method for lexical analysis of text content according to the present invention; Figure 6 Flowchart of the method for syntactic analysis and semantic analysis processing of text content according to the present invention; Figure 7 Flowchart of the method for parsing specific instructions and operations and converting the natural language understanding result into a vehicle control instruction according to the present invention. Specific embodiments
[0010] The following description is used to disclose the present invention so that those skilled in the art can implement the present invention. The preferred embodiments in the following description are only examples, and other obvious variations can be thought of by those skilled in the art.
[0011] Referring to Figure 1 As shown, an intelligent voice interconnection method for automotive cockpit control includes: Capturing the user's voice signal through the automotive cockpit microphone array technology and converting it into a digital signal, and the intelligent voice interconnection system performs denoising, echo cancellation, and voice enhancement processing on the digital signal; The preprocessed digital signal is input into the feature extraction module, and the linear prediction cepstral coefficient algorithm is used to extract the voice features in the digital signal; The extracted voice features are input into the speech recognition model, decoded based on the acoustic model for the digital signal, the input voice features are mapped to the corresponding phonemes or words, and the mapping relationship from words to text provided by the dictionary is used to output the corresponding text content; The text content is input into the natural language understanding module, and lexical analysis is performed on the text content, splitting the text into words and phrases and determining their part of speech; Based on the fuzzy instruction parsing of the dialogue history, syntactic analysis and semantic analysis processing are performed on the text content to analyze the grammatical relationship between words and phrases and parse the user's intention and requirements; The intelligent voice interconnection system parses specific instructions and operations based on the user's intention and requirements, and converts the natural language understanding result into a vehicle control instruction; During and after the execution of the instruction, the intelligent voice interconnection system feeds back relevant information or results to the user through speech synthesis technology.
[0012] Referring to Figure 2 As shown, capturing the user's voice signal through the automotive cockpit microphone array technology and converting it into a digital signal, and the specific process of the intelligent voice interconnection system performing denoising, echo cancellation, and voice enhancement processing on the digital signal includes: Adopting a distributed microphone array and designing the microphone spacing based on the matching of the sound source frequency and the array aperture; Suppress low-frequency road noise and engine harmonic interference through a hardware high-pass filter, and dynamically adjust the microphone gain based on signal saturation; Achieve time-domain alignment of at least one microphone channel through a field-programmable gate array, transmit the data to the main processor, and manage real-time data streams using a circular buffer; Use the navigation audio as a reference signal to construct an echo path model, and perform echo cancellation processing through an adaptive algorithm; Based on noise spectrum estimation, distinguish speech segments from non-speech segments, and use Wiener filtering to attenuate noise components in the frequency domain while preserving the harmonic structure of the speech segments; Calculate the time delay difference between microphones, perform weighted superposition after signal alignment, enhance the signal in the target direction, combine direction-of-arrival estimation to locate the speaker's position, and optimize the beam direction; Detect the fundamental frequency of speech and enhance the formant energy to improve the clarity of the digital signal.
[0013] Calculate the time delay difference between microphones, perform weighted superposition after signal alignment, enhance the signal in the target direction, combine direction-of-arrival estimation technology to locate the speaker's position, optimize the beam direction according to the speaker's position to ensure accurate capture of the speech signal, detect the fundamental frequency of speech, that is, the basic frequency component in the speech signal, and enhance the formant energy, that is, the harmonic component energy in the speech signal, to improve the clarity of the digital signal.
[0014] Refer to Figure 3 As shown, the preprocessed digital signal is input into the feature extraction module, and the linear prediction cepstral coefficient algorithm is used to extract the speech features in the digital signal, specifically including: Cut the preprocessed speech signal into at least one short-time frame, apply a Hamming window function to each frame of the signal to smooth the signal edge and reduce spectral leakage; Based on an approximate estimate of the vocal tract response, use covariance to calculate the linear prediction coefficients of each frame of the signal; Convert the linear prediction coefficients to linear prediction cepstral coefficients, map the prediction coefficients to the cepstral domain through recursive operations to generate a linear prediction cepstral vector; Detect the periodic peak interval in the cepstral domain, and this interval corresponds to the glottal vibration period; By obtaining the peak points in the cepstral domain, calculate the distance between adjacent peak points, and output it as the fundamental frequency of the digital signal; Use the linear prediction coefficients to construct a transfer function, calculate its frequency response curve, obtain the first three local maximum points in the amplitude spectrum, and output the formants corresponding to this set of frequencies for the vocal tract; Combine the linear prediction cepstral coefficients, fundamental frequency, and formant parameters into a feature vector to form a complete speech feature representation. The feature vector includes spectral envelope information, excitation source characteristics, and vocal tract resonance characteristics.
[0015] The preprocessed speech signal is segmented into multiple short-time frames. The length of each frame is usually selected to be 10 - 30 milliseconds, and the frame shift (the overlapping part between adjacent frames) is usually 1 / 2 to 1 / 3 of the frame length. A Hamming window function is applied to each frame signal to reduce spectral leakage and smooth the signal edges. The Hamming window function is: , where is the Hamming window function, is the sample index within the frame, is the frame length.
[0016] Referring to Figure 4 as shown, the extracted speech features are input into a speech recognition model. Based on the acoustic model, the digital signal is decoded, mapping the input speech features to the corresponding phonemes or words, and using the word-to-text mapping relationship provided by the dictionary to output the corresponding text content, which specifically includes: Input the extracted speech feature sequence into the trained acoustic model. The feature sequence is frame-aligned, retaining the time dimension information; Convert the frame-level output into a phoneme sequence through a blank symbol and repetition merging mechanism; Use an encoder-decoder structure to align the acoustic features with the phoneme sequence; Obtain a mapping table from phonemes to words and denote it as a dictionary. Perform dictionary constraint processing on the dictionary, only retaining the phoneme combinations that can form valid words; Calculate the matching probability between the current phoneme and the speech features and output it as an acoustic score; Perform weighted fusion of the acoustic score and the language model score; Obtain the globally optimal path through the A* algorithm. Using frames as nodes, phoneme transitions as edges, and the total score as the edge weight to construct a search space; Design a heuristic function to estimate the minimum cost from the current node to the end point, and preferentially expand the node with the lowest sum of the minimum cost and the actual cost; Judge whether the end of the sequence is reached. If so, end the search for the globally optimal path. If not, do not make an output; Restore out-of-vocabulary words in the dictionary through byte pair encoding subword units; Obtain a punctuation prediction model, input the phoneme sequence and context, and output the punctuation positions; Insert capital letters and line breaks based on the text structure to generate formatted text, completing the conversion from speech to text.
[0017] Using an encoder-decoder structure, the acoustic feature sequence is encoded into hidden states and then decoded into a phoneme sequence. The encoder part is usually used to extract the deep representation of acoustic features, while the decoder part is used to generate the phoneme sequence, obtain the mapping table (dictionary) from phonemes to words, and perform dictionary constraint processing to retain only the phoneme combinations that can form valid words, so as to reduce the search space and improve the decoding efficiency. For the phoneme prediction corresponding to each frame of features, calculate the matching probability between it and the true phoneme and output it as an acoustic score, obtain the language model score, which reflects the probability that the phoneme sequence forms legal words and sentences, and perform weighted fusion on the acoustic score and the language model score to obtain the total score.
[0018] Refer to Figure 5 As shown, input the text content into the natural language understanding module, and perform lexical analysis on the text content, split the text into words and phrases, and determine their parts of speech, specifically including: Split the text into the smallest semantic units, retaining the linguistic structure. The smallest semantic units include words, phrases, and symbols; For each smallest semantic unit, use the rules in the rule base based on the context of the current word for matching; By finding the rule in the rule base that best matches the token, assign a part-of-speech tag to each smallest semantic unit. The part-of-speech tags include nouns, verbs, adjectives, and adverbs.
[0019] In practical applications, the rule base needs to be continuously updated and optimized to adapt to the text content in different fields and contexts. At the same time, machine learning or deep learning techniques can also be combined to improve the accuracy and working efficiency of part-of-speech tagging.
[0020] Refer to Figure 6 As shown, based on the fuzzy instruction parsing of the dialogue history, perform syntactic analysis and semantic analysis processing on the text content, analyze the grammatical relationships between words and phrases, and parse the user's intentions and requirements, specifically including: Maintain a dialogue state machine to record the user instructions, system responses, and slot filling results of historical rounds; Use a neural network model to identify the pronoun references in the user's speech and construct an entity co-reference chain; Based on historical actions, copy the unknown parameters not mentioned in the user's speech. The unknown parameters include time parameters and location parameters; Update the domain dictionary in real time based on the user's speech and cache the context-related entities; Construct a phrase structure tree, identify complex nested structures, and perform pruning based on the dialogue history, restricting the candidate structures according to historical actions; Map the current parsing result to the dialogue state variables; Query the knowledge graph to verify the existence of entities, filter unreasonable requests based on business rules, and ensure that the current request is consistent with the conversation goal; Use multi-task learning to jointly optimize the domain and sub-intentions; Inherit the dialogue history parameters, detect demand conflicts, and generate corresponding guiding questions after discovering conflicting parameters; Use the user's clarification result as new training data and optimize the parameter extraction strategy using reinforcement learning; Optimize the slot filling model in combination with the user's click log and analyze the user's speech stress characteristics to improve intent recognition.
[0021] According to the result of pronoun resolution, construct an entity co-reference chain to connect pronouns and nouns that refer to the same entity. The entity co-reference chain helps to accurately identify the user's intent and slot filling in subsequent processing. Use a multi-task learning framework to jointly optimize the domain recognition and sub-intent recognition tasks, and improve the generalization ability and recognition accuracy of the model by sharing neural network layers or joint loss functions.
[0022] Refer to Figure 7 As shown, the intelligent voice interconnection system analyzes specific instructions and operations based on the user's intent and requirements, and converts the natural language understanding result into vehicle control instructions, which specifically include: Analyze the semantic content of the text based on the intent recognition model, extract information related to the intent from the text, and determine the user's intent; Track and manage the interaction history between the user and the system, and maintain the current dialogue state; Process the clarification requests, correction requests, and multi-round interactions proposed by the user to ensure that the system understands the user's complete requirements; Based on the user's intent and context information, decide the next operation and response; Generate specific vehicle control instructions based on the decision result of the dialogue management; Parse the generated instructions into a format that the vehicle control system can understand and encapsulate them into corresponding control signals.
[0023] Parse the generated vehicle control instructions into a format that the vehicle control system can understand, which involves converting natural language instructions into specific command codes, protocols, or signals, and encapsulating the parsed instructions into corresponding control signals for sending to the vehicle control system for execution.
[0024] Furthermore, this solution also proposes a computer-readable storage medium, on which a computer-readable program is stored. When the computer-readable program is called, it executes the above-mentioned intelligent voice interconnection method for vehicle cockpit control.
[0025] It can be understood that the storage medium can be a magnetic medium, such as a floppy disk, a hard disk, or a magnetic tape; an optical medium such as a DVD; or a semiconductor medium such as a solid state disk (SSD), etc.
[0026] In summary, the advantages of the present invention are as follows: reducing the interference of environmental noise and echo, improving the signal-to-noise ratio and clarity of the voice signal, ensuring the accuracy and clarity of the voice signal, providing a reliable basis for subsequent processing, extracting voice features in the digital signal using the linear prediction cepstral coefficient algorithm, which is crucial for subsequent voice recognition and natural language understanding, enhancing the interaction experience between the user and the vehicle, improving the usability and reliability of the system, improving driving safety, and enhancing the overall satisfaction of the user.
[0027] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification is only the principle of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. An intelligent voice interconnection method for automobile cockpit control, characterized in that: include: The car cabin microphone array technology is used to capture the user's voice signal and convert it into a digital signal. The intelligent voice interconnection system performs denoising, echo cancellation and voice enhancement on the digital signal. The pre-processed digital signal is input into the feature extraction module, and the speech features in the digital signal are extracted using the linear prediction cepstral coefficient algorithm; The extracted speech features are input into the speech recognition model, the digital signal is decoded based on the acoustic model, the input speech features are mapped to the corresponding phonemes or words, and the corresponding text content is output using the mapping relationship from words to text provided by the dictionary; Input the text content into the natural language understanding module, and perform lexical analysis on the text content, split the text into words and phrases, and determine their parts of speech; Based on the fuzzy instruction parsing of the conversation history, the text content is subjected to syntactic and semantic analysis, the grammatical relationship between words and phrases is analyzed, and the user's intention and needs are analyzed; The intelligent voice interconnection system analyzes specific instructions and operations based on the user's intentions and needs, and converts the results of natural language understanding into vehicle control instructions; During and after the execution of instructions, the intelligent voice interconnection system uses speech synthesis technology to feed back relevant information or results to the user.
2. The intelligent voice interconnection method for automobile cockpit control according to claim 1, characterized in that: The method of capturing the user's voice signal by using the car cabin microphone array technology and converting it into a digital signal, and the intelligent voice interconnection system performing denoising, echo cancellation and voice enhancement processing on the digital signal specifically includes: Use a distributed microphone array and design the spacing between microphones based on the matching of the sound source frequency and the array aperture; The hardware high-pass filter suppresses low-frequency road noise and engine harmonic interference, and dynamically adjusts the microphone gain based on signal saturation; Time-domain alignment of at least one microphone channel is achieved through a field programmable gate array, data is transmitted to a main processor, and a ring buffer is used to manage real-time data flow; The navigation audio is used as a reference signal to build an echo path model, and the echo cancellation process is performed through an adaptive algorithm; Based on noise spectrum estimation, speech segments and non-speech segments are distinguished, and Wiener filtering is used to attenuate noise components in the frequency domain to retain the harmonic structure of the speech segment; Calculate the delay difference between microphones, align the signals, perform weighted superposition, enhance the target direction signal, estimate the speaker's position based on the arrival direction, and optimize the beam pointing; Detect the fundamental frequency of speech and enhance the energy of formant to improve the clarity of digital signals.
3. The intelligent voice interconnection method for automobile cockpit control according to claim 2 is characterized in that: The pre-processed digital signal is input into the feature extraction module, and the speech features in the digital signal are extracted by using the linear prediction cepstral coefficient algorithm, which specifically includes: The preprocessed speech signal is cut into at least one short time frame, and a Hamming window function is applied to each frame signal to smooth the signal edge and reduce spectrum leakage; Based on the approximate estimation of the vocal tract response, the linear prediction coefficients of each frame signal are calculated using covariance; The linear prediction coefficients are converted into linear prediction cepstral coefficients, and the prediction coefficients are mapped to the cepstral domain through recursive operation to generate a linear prediction cepstral vector; The periodic peak intervals are detected in the cepstrum domain, which correspond to the glottal vibration period; By obtaining the peak points in the cepstrum domain, the distance between adjacent peak points is calculated and output as the fundamental frequency of the digital signal; The transfer function is constructed using the linear prediction coefficients, its frequency response curve is calculated, the first three local maximum points in the amplitude spectrum are obtained, and the resonance peaks of the corresponding vocal channels of this group of frequencies are output; The linear prediction cepstral coefficients, fundamental frequency and formant parameters are combined into a feature vector to form a complete speech feature representation, wherein the feature vector includes spectrum envelope information, excitation source characteristics and vocal tract resonance characteristics.
4. The intelligent voice interconnection method for automobile cockpit control according to claim 3 is characterized in that: The extracted speech features are input into the speech recognition model, the digital signal is decoded based on the acoustic model, the input speech features are mapped to corresponding phonemes or words, and the mapping relationship from words to texts provided by the dictionary is used to output the corresponding text content, which specifically includes: The extracted speech feature sequence is input into the trained acoustic model, and the feature sequence is aligned by frame to retain the time dimension information; Convert the frame-level output into a phoneme sequence through a whitespace and repetition merging mechanism; Using the encoder-decoder structure, align acoustic features with phoneme sequences; Obtaining a mapping table from phonemes to words and recording it as a dictionary, performing dictionary constraint processing on the dictionary to retain only phoneme combinations that can form valid words; Calculate the matching probability between the current phoneme and the speech feature and output it as an acoustic score; Perform weighted fusion of acoustic score and language model score; The global optimal path is obtained through the A* algorithm, and the search space is constructed with frames as nodes, phoneme transfers as edges, and edge weights as total scores; Design a heuristic function to estimate the minimum cost from the current node to the end point, and give priority to expanding the node with the lowest sum of the minimum cost and the actual cost; Determine whether the sequence end point has been reached. If so, end the global optimal path search. If not, do not output. Recover the missing words in the dictionary through byte-pair encoded subword units; Get the punctuation prediction model, input the phoneme sequence and context, and output the punctuation position; Insert capital letters and line breaks based on the text structure to generate formatted text and complete speech-to-text conversion.
5. The intelligent voice interconnection method for automobile cockpit control according to claim 4 is characterized in that: The step of inputting the text content into a natural language understanding module, performing lexical analysis on the text content, splitting the text into words and phrases, and determining the parts of speech specifically includes: Splitting the text into the smallest semantic units, preserving the linguistic structure, wherein the smallest semantic units include words, phrases and symbols; For each minimum semantic unit, use the rules in the rule base based on the context of the current word to match; By searching the rule in the rule base that best matches the tag, a part-of-speech tag is assigned to each minimum semantic unit, and the part-of-speech tags include nouns, verbs, adjectives, and adverbs.
6. The intelligent voice interconnection method for automobile cockpit control according to claim 5, characterized in that: The fuzzy instruction parsing based on the conversation history performs syntactic and semantic analysis on the text content, analyzes the grammatical relationship between words and phrases, and parses the user's intentions and needs, specifically including: Maintain the dialogue state machine and record the user commands, system responses, and slot filling results of historical rounds; Use a neural network model to identify pronoun references in user speech and build entity coreference chains; Copy unknown parameters not mentioned in the user's voice based on historical actions, wherein the unknown parameters include time parameters and location parameters; Update domain dictionaries in real time based on user voice and cache context-related entities; Build a phrase structure tree to identify complex nested structures, guide pruning based on conversation history, and restrict candidate structures based on historical actions; Map the current parsing result to the dialog state variable; Query the knowledge graph to verify the existence of entities, filter unreasonable requests based on business rules, and ensure that the current request is consistent with the dialogue goal; Jointly optimize domains and sub-intents using multi-task learning; The dialogue history parameters are inherited and demand conflicts are detected. When conflicting parameters are found, corresponding guiding questions are generated. Use user clarification results as new training data and use reinforcement learning to optimize the parameter extraction strategy; The slot filling model is optimized based on user click logs, and the user's voice stress characteristics are analyzed to improve intent recognition.
7. The intelligent voice interconnection method for automobile cockpit control according to claim 6, characterized in that: The intelligent voice interconnection system analyzes specific instructions and operations based on the user's intentions and needs, and converts the natural language understanding results into vehicle control instructions, including: Analyze the semantic content of the text based on the intent recognition model, extract information related to the intent from the text, and determine the user's intent; Track and manage the interaction history between the user and the system and maintain the current dialogue status; Handle clarification requests, correction requests, and multiple rounds of interactions from users to ensure the system understands the user's complete needs; Determine the next action and response based on the user's intent and contextual information; Generate specific vehicle control instructions based on the decision results of dialogue management; The generated instructions are parsed into a format that can be understood by the vehicle control system and encapsulated into corresponding control signals.
Citation Information
Patent Citations
Intelligent human-machine interaction semantic analysis method and interaction system
CN102968409A
Semantic analysis method for secondary matching semantic
CN106970909A
Multi-level dialogue analysis method for intelligent voice dialogue system
CN107679042A
Double-microphone array-based intelligent voice interaction system
CN109817209A
Voice awakening method and device, equipment and medium
CN109949810A
Cited By
Conversational interaction method and device suitable for tobacco machinery and medium thereof
CN121256001A