Voice processing method and device, electronic equipment and storage medium

Through the parallel processing mechanism of online links and offline links, the online and offline processing results of voice data are integrated, and the stability and response speed of the vehicle voice interaction system in complex network environments are solved, and the processing accuracy and driving safety of the vehicle voice interaction system are improved.

CN120260552APending Publication Date: 2025-07-04XIAOMI EV TECH CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510495967.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

It is difficult for vehicle voice interaction systems to ensure stable operation and rapid response in complex network environments, which affects driving safety and operational convenience.

Method used

The parallel processing mechanism of online links and offline links is adopted to process voice data to be processed in parallel through online links and offline links, and the online and offline processing results are integrated to generate the optimal target processing results.

Benefits of technology

Real-time stable processing of voice data is realized in complex network environments, improving the response speed and processing accuracy of the vehicle voice interaction system, and enhancing the operation convenience and safety during driving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260552A_ABST
    Figure CN120260552A_ABST
Patent Text Reader

Abstract

The invention provides a voice processing method and device, electronic equipment and a storage medium, and relates to the field of artificial intelligence. The method comprises the following steps: acquiring to-be-processed voice data; determining a current available link; in response to the current available links including an online link and an offline link, performing parallel processing on the to-be-processed voice data through the online link and the offline link to obtain an online processing result and an offline processing result; and determining a target processing result corresponding to the to-be-processed voice data according to the online processing result and the offline processing result. According to the method, the collected voice data can be stably processed in real time in the complex network environment of the vehicle, the response speed and the processing accuracy of the vehicle-mounted voice interaction system are improved, and the operation convenience and safety in the driving process are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence, and in particular, to a voice processing method, apparatus, electronic device, and storage medium. Background Art

[0002] With the rapid development of intelligent connected vehicles, in-vehicle voice interaction systems have become one of the core functions of intelligent cockpits. Drivers can control in-vehicle functions such as navigation, multimedia entertainment, and air conditioning through voice commands, improving driving safety and operation convenience. However, in complex driving environments, voice interaction systems face multiple technical challenges: the vehicle network environment is complex and variable, and stable operation under various network conditions needs to be ensured; voice processing has extremely high requirements for real-time performance, and quick responses need to be made with limited computing resources. Summary of the Invention

[0003] The purpose of the present disclosure is to provide a voice processing method, apparatus, electronic device, and storage medium.

[0004] According to a first aspect of an embodiment of the present disclosure, a voice processing method is provided. The method includes: obtaining voice data to be processed; determining a currently available link; in response to the currently available link including an online link and an offline link, performing parallel processing on the voice data to be processed through the online link and the offline link to obtain an online processing result and an offline processing result; and determining a target processing result corresponding to the voice data to be processed according to the online processing result and the offline processing result.

[0005] In some embodiments of the present disclosure, the performing parallel processing on the voice data to be processed through the online link and the offline link to obtain an online processing result and an offline processing result includes: sending the voice data to be processed to a server, and receiving the online processing result returned by the server in response to the voice data to be processed; the online processing result includes at least one of an online speech recognition result, an online semantic understanding result, and online voice data; performing speech recognition on the voice data to be processed to obtain an offline speech recognition result; performing fusion processing on the online speech recognition result and the offline speech recognition result to obtain a fused speech recognition result; using a local semantic understanding model to perform semantic understanding on the fused speech recognition result to obtain an offline semantic understanding result; and performing voice conversion on the offline semantic understanding result to obtain offline voice data.

[0006] In some embodiments of the present disclosure, the local semantic understanding model includes a local multi-modal large model; wherein, using the local semantic understanding model to perform semantic understanding on the fused speech recognition result to obtain an offline semantic understanding result includes: performing intent distribution on the fused speech recognition result to determine the task to which the fused speech recognition result belongs; using the local multi-modal large model to perform semantic understanding on the fused speech recognition result according to the basic parameters of the local multi-modal large model and the weight parameters corresponding to the task to which it belongs to obtain the offline semantic understanding result; the weight parameters corresponding to the task to which it belongs are pre-trained.

[0007] In some embodiments of the present disclosure, after obtaining the offline semantic understanding result, the method further includes: determining whether the offline speech recognition result is a user request according to the speech data to be processed, the offline speech recognition result, and the offline semantic understanding result; in response to the offline speech recognition result being a user request, determining to perform speech conversion on the offline semantic understanding result.

[0008] In some embodiments of the present disclosure, after obtaining the offline semantic understanding result, the method further includes: performing fusion processing on the online semantic understanding result and the offline semantic understanding result to obtain a fused semantic understanding result.

[0009] In some embodiments of the present disclosure, determining the target processing result corresponding to the speech data to be processed according to the online processing result and the offline processing result includes: in response to the fused semantic understanding result being the online semantic understanding result, using the online speech data as the target processing result; in response to the fused semantic understanding result being the offline semantic understanding result, using the offline speech data as the target processing result.

[0010] In some embodiments of the present disclosure, the speech data to be processed includes one or more data packets; wherein, the method further includes: for the current data packet, determining whether the online speech recognition result of the current data packet is received within a first preset time; in response to the online speech recognition result of the current data packet being received within the first preset time, displaying the online speech recognition result of the speech data packet; in response to the online speech recognition result of the current data packet not being received within the first preset time, displaying the offline speech recognition result of the speech data packet.

[0011] In some embodiments of the present disclosure, before determining whether the online speech recognition result of the current data packet is received within a first preset time, the method further includes: for the previous data packet of the current data packet, determining whether to display the offline speech recognition result of the previous data packet; in response to displaying the offline speech recognition result of the previous data packet, displaying the offline speech recognition result of the current data packet; in response to displaying the online speech recognition result of the previous data packet, analyzing the display manner of the speech recognition result of the current data packet.

[0012] In some embodiments of the present disclosure, the fusing the online speech recognition result and the offline speech recognition result to obtain a fused speech recognition result includes: in response to not receiving the online speech recognition result within a second preset time, using the offline speech recognition result as the fused speech recognition result; in response to receiving the online speech recognition result within the second preset time, fusing the online speech recognition result and the offline speech recognition result to obtain the fused speech recognition result.

[0013] In some embodiments of the present disclosure, the fusing the online semantic understanding result and the offline semantic understanding result to obtain a fused semantic understanding result includes: in response to not receiving the online semantic understanding result within a third preset time, using the offline semantic understanding result as the fused semantic understanding result; in response to receiving the online semantic understanding result within the third preset time, fusing the online semantic understanding result and the offline semantic understanding result to obtain the fused semantic understanding result.

[0014] In some embodiments of the present disclosure, the method further includes: in response to the current available link including the online link, processing the speech data to be processed through the online link, and using the processing result of the online link as the target processing result.

[0015] In some embodiments of the present disclosure, the method further includes: in response to the current available link including the offline link, processing the speech data to be processed through the offline link, and using the processing result of the offline link as the target processing result.

[0016] According to a second aspect of the embodiments of the present disclosure, there is provided a voice processing device, the device including: a voice data acquisition module configured to acquire voice data to be processed; an available link determination module configured to determine a currently available link; a voice data processing module configured to, in response to the currently available link including an online link and an offline link, perform parallel processing on the voice data to be processed through the online link and the offline link to obtain an online processing result and an offline processing result; and a processing result determination module configured to determine a target processing result corresponding to the voice data to be processed according to the online processing result and the offline processing result.

[0017] In some embodiments of the present disclosure, the voice data processing module is further configured to: send the voice data to be processed to a server and receive the online processing result returned by the server in response to the voice data to be processed; the online processing result includes at least one of an online speech recognition result, an online semantic understanding result, and online voice data; perform speech recognition on the voice data to be processed to obtain an offline speech recognition result; perform fusion processing on the online speech recognition result and the offline speech recognition result to obtain a fused speech recognition result; use a local semantic understanding model to perform semantic understanding on the fused speech recognition result to obtain an offline semantic understanding result; and perform speech conversion on the offline semantic understanding result to obtain offline voice data.

[0018] In some embodiments of the present disclosure, the local semantic understanding model includes a local multimodal large model; wherein, the voice data processing module is further configured to: perform intent distribution on the fused speech recognition result to determine a task to which the fused speech recognition result belongs; according to the basic parameters of the local multimodal large model and the weight parameters corresponding to the task, use the local multimodal large model to perform semantic understanding on the fused speech recognition result to obtain the offline semantic understanding result; and the weight parameters corresponding to the task are pre-trained.

[0019] In some embodiments of the present disclosure, the voice data processing module is further configured to: perform fusion processing on the online semantic understanding result and the offline semantic understanding result to obtain a fused semantic understanding result.

[0020] In some embodiments of the present disclosure, the voice data processing module is further configured to: in response to not receiving the online speech recognition result within a second preset time, use the offline speech recognition result as the fused speech recognition result; and in response to receiving the online speech recognition result within the second preset time, perform fusion processing on the online speech recognition result and the offline speech recognition result to obtain the fused speech recognition result.

[0021] In some embodiments of the present disclosure, the voice data processing module is further configured to: in response to not receiving the online semantic understanding result within a third preset time, use the offline semantic understanding result as the fused semantic understanding result; in response to receiving the online semantic understanding result within the third preset time, perform a fusion process on the online semantic understanding result and the offline semantic understanding result to obtain the fused semantic understanding result.

[0022] In some embodiments of the present disclosure, the voice data processing module is further configured to: in response to the current available link including the online link, process the voice data to be processed through the online link, and use the processing result of the online link as the target processing result.

[0023] In some embodiments of the present disclosure, the voice data processing module is further configured to: in response to the current available link including the offline link, process the voice data to be processed through the offline link, and use the processing result of the offline link as the target processing result.

[0024] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, characterized by comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to implement the voice processing method described in the above embodiments.

[0025] According to a fourth aspect of the embodiments of the present disclosure, there is provided a non-transitory computer-readable storage medium, which when the instructions in the storage medium are executed by a processor of a terminal device, enables the terminal device to execute the voice processing method described in the above embodiments.

[0026] According to a fifth aspect of the embodiments of the present disclosure, there is provided a computer program product, including a computer program, which when executed by a processor, implements the voice processing method described in the above embodiments.

[0027] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects:

[0028] Obtain the voice data to be processed, detect the current network and link status to determine the currently available link. When it is detected that the currently available link includes an online link and an offline link, adopt a dual-link parallel processing mechanism to perform online voice processing and offline voice processing; fuse the online processing results provided by the online link and the offline processing results obtained from the offline link, comprehensively consider factors such as recognition accuracy and response latency, and generate the optimal target processing result. In this way, it is possible to process the collected voice data in real time and stably in the complex network environment of the vehicle, improve the response speed and processing accuracy of the in-vehicle voice interaction system, and enhance the operation convenience and safety during driving.

[0029] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure.

[0031] Figure 1 is a flowchart of a voice processing method shown according to some embodiments of the present disclosure Figure 1 .

[0032] Figure 2 is a flowchart of a voice processing method shown according to some embodiments of the present disclosure Figure 2 .

[0033] Figure 3 is a flowchart of parallel processing of voice data to be processed through an online link and an offline link shown according to some embodiments of the present disclosure.

[0034] Figure 4 is a schematic diagram of performing offline semantic understanding shown according to some embodiments of the present disclosure.

[0035] Figure 5 is a schematic diagram of a flowchart of a method for processing voice data through an in-vehicle voice assistant system and a cloud server shown according to some exemplary embodiments of the present disclosure.

[0036] Figure 6 is a block diagram of a voice processing device shown according to some embodiments of the present disclosure.

[0037] Figure 7 is a functional block diagram of a vehicle shown according to some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0038] Exemplary embodiments of the present disclosure will be described in detail herein, and examples thereof are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. Various changes, modifications, and equivalents of the methods, apparatuses, and / or systems described herein will become apparent after understanding the present disclosure. For example, the order of operations described herein is merely exemplary and is not limited to those set forth herein, but may be changed as will be apparent after understanding the present disclosure, except for operations that must be performed in a specific order. Additionally, descriptions of features known in the art may be omitted for the sake of clarity and conciseness.

[0039] The embodiments described in the following exemplary embodiments of the present disclosure do not represent all embodiments consistent with the present invention of the present disclosure. On the contrary, they are merely examples of apparatuses and methods consistent with some aspects of the present invention of the present disclosure as detailed in the appended claims.

[0040] It should be noted that the acquisition, storage, use, processing, etc. of data in the technical solutions of the present disclosure all comply with the relevant provisions of national laws and regulations. In the embodiments of the present disclosure, various types of data such as personal identity data, operation data, and behavior data related to individuals, customers, and populations have been authorized.

[0041] Figure 1 is a flowchart of a voice processing method shown according to some embodiments of the present disclosure Figure 1 . The method shown in the embodiments of the present disclosure can be applied to an electronic device, which can be a mobile intelligent agent, such as a vehicle or a robot, or an intelligent non-mobile agent, such as an air conditioner or a refrigerator. Referring to Figure 1 , the voice processing method may include the following steps.

[0042] In step S110, voice data to be processed is acquired.

[0043] In the embodiments of the present disclosure, an in-vehicle device can collect voice emitted by a user through a microphone array inside and outside the vehicle. Among them, the microphones can be distributed at different positions on the vehicle body to collect voice signals omnidirectionally and with high fidelity, reducing noise interference.

[0044] In step S120, the currently available link is determined.

[0045] In the embodiments of the present disclosure, the network connection status can be monitored in real time, including parameters such as signal strength, latency, and bandwidth. After acquiring the voice to be processed, the operating conditions of the offline link, such as model integrity and resource availability, can be checked. If the network connection status is normal and the offline link is available, it is determined that the currently available link includes an online link and an offline link.

[0046] Exemplarily, an online link refers to a path through which in-vehicle devices upload voice data to be processed to a cloud server for processing by means of a network connection. The cloud server has powerful computing resources and a vast amount of data, and is capable of performing complex tasks such as speech recognition, semantic understanding, and speech synthesis. For example, when a user says "Find the western restaurant with the highest rating nearby", the online link will transmit the voice data to the cloud server. The cloud server uses rich map data, merchant evaluation data, and complex algorithms to accurately identify the voice content, understand the user's intention, and plan the restaurant information that meets the requirements, and finally returns the result to the in-vehicle device. This process relies on a stable network environment to ensure the rapid transmission of data and the timely feedback of processing results.

[0047] Exemplarily, an offline link is a way for in-vehicle devices to process voice data locally using pre-installed and stored speech processing models, data resources, etc. The local device has a certain computing power and can deploy speech recognition models, semantic understanding models, etc. For example, during offline speech recognition, the local model is used to convert speech into text; during the semantic understanding stage, the rules and models stored locally are used to parse the meaning of the text. When the vehicle is in an area with poor network signals, such as a tunnel or a remote mountainous area, and the user says "Turn on the air conditioner", the offline link relies on local resources to quickly recognize the speech, understand the instruction, and control the air conditioner to turn on, ensuring that the voice interaction function is not restricted by the network and can continuously provide voice services to users.

[0048] In step S130, in response to the current available links including an online link and an offline link, the voice data to be processed is processed in parallel through the online link and the offline link to obtain an online processing result and an offline processing result.

[0049] In the embodiments of the present disclosure, when it is determined that the current available links cover an online link and an offline link, a parallel processing mechanism is started. On the one hand, the online link quickly transmits the voice data to be processed to the cloud server through the network and processes the voice data to be processed to obtain an online processing result. On the other hand, the offline link synchronously processes the voice data to be processed locally to obtain an offline processing result.

[0050] In step S140, based on the online processing result and the offline processing result, the target processing result corresponding to the voice data to be processed is determined.

[0051] In the embodiments of the present disclosure, the online processing result and the offline processing result are fused to finally obtain the target processing result corresponding to the voice data to be processed. Exemplarily, factors such as recognition accuracy and response latency can be comprehensively considered. For example, for recognition accuracy, the accuracy of the online processing result and the offline processing result can be compared. In terms of response latency, the time consumed from data upload to result return in the online link and the local processing time of the offline link can be measured. Based on the comprehensive trade-off of these factors, the advantageous parts in the online processing result and the offline processing result are integrated to generate the most optimized target processing result.

[0052] The voice processing method provided by the embodiments of the present disclosure obtains the voice data to be processed, detects the current network and link status to determine the currently available links, and when it is detected that the currently available links include an online link and an offline link, a dual-link parallel processing mechanism is adopted to perform online voice processing and offline voice processing; the online processing result provided by the online link and the offline processing result obtained by the offline link are fused, and factors such as recognition accuracy and response latency are comprehensively considered to generate the optimal target processing result. In this way, it is possible to process the collected voice data in real time and stably in the complex network environment of the vehicle, improve the response speed and processing accuracy of the in-vehicle voice interaction system, and enhance the operation convenience and safety during driving.

[0053] Figure 2 is a flowchart of a voice processing method shown according to some embodiments of the present disclosure Figure 2 In the embodiments of the present disclosure, Figure 2 In the voice processing method shown, steps S210, S220, S230, and S240 respectively correspond to Figure 1 steps S110, S120, S130, and S140 in the voice processing method shown, and will not be repeated here.

[0054] In the embodiments of the present disclosure, on the basis of Figure 1 the voice processing method shown, Figure 2 the voice processing shown may further include the following steps.

[0055] In step S250, in response to the currently available link including an online link, the voice data to be processed is processed through the online link, and the processing result of the online link is used as the target processing result.

[0056] In the embodiments of the present disclosure, when it is detected that the only available connection currently is the online link, such as in the case of offline link upgrade or offline link failure, it is determined to use the online link for voice processing. Exemplarily, the voice data to be processed is sent to the server, and the server processes the voice data to be processed, and the result returned by the server is used as the target processing result. For example, when the vehicle is driving on an urban road with good network signal and the offline link has a temporary failure, the user says "Please turn on intelligent driving", and the voice data is sent to the server through the online link. After the server processes it and returns the result, this result is played to the user as the target processing result.

[0057] In step S260, in response to the currently available link including an offline link, the voice data to be processed is processed through the offline link, and the offline link processing result is used as the target processing result.

[0058] In the embodiments of the present disclosure, when it is detected that the only available link currently is the offline link, such as when in a network-free environment, it is determined to use the offline link for voice processing, and its processing result is used as the target processing result. For example, when the vehicle drives into a remote mountainous area and the network is completely interrupted, the user says "Play local music", and the voice instruction is processed through the offline link to play local music, and its processing result is the target processing result.

[0059] Through the above steps, flexible and diverse utilization can be carried out according to the actual situation of the currently available link. When both the online link and the offline link are available, the voice data will be processed by means of the online link and the offline link, and the results of both will be integrated to obtain a better target processing result, thereby improving the accuracy and comprehensiveness of voice processing; if only the online link is available, it will switch to the online link, send the voice data to be processed to the server for processing, and use the result returned by the server as the target processing result; if only the offline link is available, it will rely on the offline link to complete the voice processing and use its processing result as the target processing result. In this way, no matter how the network and link status of the in-vehicle environment change, it can ensure that the voice processing continues to work, bringing a stable voice interaction experience to the user and enhancing the practicality of the intelligent in-vehicle voice assistant in various scenarios.

[0060] Figure 3 is a flowchart showing parallel processing of voice data to be processed through an online link and an offline link according to some embodiments of the present disclosure. Refer to Figure 3 , and may include the following steps.

[0061] In step S310, the voice data to be processed is sent to the server, and the online processing result returned by the server in response to the voice data to be processed is received; the online processing result includes at least one of an online speech recognition result, an online semantic understanding result, and online voice data.

[0062] In the embodiments of the present disclosure, the online processing results may include at least one of an online speech recognition result, an online semantic understanding result, and online speech data. After determining that the online link is the currently available link, the speech data to be processed is sent to the server. After receiving the speech data to be processed, the server performs speech recognition, converts the speech data to be processed into text to obtain an online speech recognition result, then performs semantic understanding on the online speech recognition result to obtain an online semantic understanding result, and then performs a speech conversion operation on the online semantic understanding result to generate online speech data.

[0063] Exemplarily, an online Automatic Speech Recognition (ASR) model, an online Natural Language Processing (NLP) model, and an online Text-To-Speech (TTS) module are deployed on the server. The online ASR model is used to perform speech recognition on the speech data to be processed to obtain an online speech recognition result, then the online NLP model is used to perform semantic understanding on the online speech recognition result to obtain an online semantic understanding result, and then the online TTS module is used to convert the online semantic understanding result into speech to obtain online speech data.

[0064] In step S320, speech recognition is performed on the speech data to be processed to obtain an offline speech recognition result.

[0065] In the embodiments of the present disclosure, in the offline link, first, speech recognition is performed on the speech data to be processed to obtain an offline speech recognition result. Exemplarily, a local ASR model is deployed on the in-vehicle device, and the local ASR model is used to perform speech recognition on the speech data to be processed and convert it into text to obtain an offline speech recognition result.

[0066] In some embodiments of the present disclosure, the speech data to be processed includes one or more data packets. Among them, the speech processing method further includes: for the current data packet, determining whether an online speech recognition result of the current data packet is received within a first preset time; in response to receiving the online speech recognition result of the current data packet within the first preset time, displaying the online speech recognition result of the speech data packet; in response to not receiving the online speech recognition result of the current data packet within the first preset time, displaying the offline speech recognition result of the speech data packet.

[0067] In some embodiments of the present disclosure, before determining whether the online speech recognition result of the current data packet is received within a third preset time, the speech processing method further includes: for the previous data packet of the current data packet, determining whether to display the offline speech recognition result of the previous data packet; in response to displaying the offline speech recognition result of the previous data packet, displaying the offline speech recognition result of the current data packet; in response to displaying the online speech recognition result of the previous data packet, analyzing the display manner of the speech recognition result of the current data packet.

[0068] In the embodiments of the present disclosure, the speech data to be processed exists in the form of one or more data packets. For the online link, the online ASR model of the server can adopt a streaming recognition method. Each time a data packet is received, the online speech recognition result corresponding to the data packet is output and sent to the in-vehicle device. Correspondingly, for the offline link, the local ASR model can adopt a streaming recognition method. Each time a data packet is received, the offline speech recognition result corresponding to the data packet is output.

[0069] In the embodiments of the present disclosure, the display manner of each data packet is analyzed. For the current data packet, a time threshold for receiving the online speech recognition result, that is, the first preset time, is set for it. The first preset time can be timed starting from the moment when the data packet is sent. During the timing, it can be monitored whether the online speech recognition result corresponding to the data packet can be received within the first preset time. If it can be received in time, the online speech recognition result is displayed on the in-vehicle display screen to ensure that the user can obtain the most accurate recognition information. If the online speech recognition result is not received after exceeding the first preset time, the offline speech recognition result corresponding to the data packet is displayed. And once such a timeout situation occurs, all subsequent data packets will directly display the corresponding offline speech recognition results, so as to ensure the coherence and timeliness of the display of the recognition results.

[0070] In addition, before judging the display manner of the online speech recognition result of the current data packet, first check the display situation of the previous data packet. If the previous data packet displays the offline speech recognition result, the offline speech recognition result of the current data packet is directly displayed. If the previous data packet displays the online speech recognition result, the display manner of the speech recognition result of the current data packet is further analyzed.

[0071] For example, when a user says to the in-vehicle voice assistant, "Navigate me to the nearest cinema and then play a pop song", the in-vehicle voice assistant splits the voice data into multiple data packets. The first data packet contains the voice information "Navigate me to", and this data packet is sent to the server for online speech recognition, and the first preset time, such as 3 s, starts timing from the sending moment. Within these 3 seconds, the server successfully completes the recognition of this data packet and returns the recognition result "Navigate me to" to the in-vehicle voice assistant, and the in-vehicle display screen then shows this online speech recognition result. The second data packet contains the voice information "the nearest cinema" and is sent out, and the timing starts again. However, since the vehicle enters an area with weak network signal at this time, 3 seconds pass and the online speech recognition result corresponding to this data packet is not received. Therefore, the in-vehicle display screen shows the result "the nearest cinema" obtained by the local ASR model. And starting from this data packet, for all subsequent data packets, such as the data packet containing "and then play a pop song", they no longer wait for the online speech recognition result and directly use the offline speech recognition result for display.

[0072] Through the above display method for each data packet, by using the first preset time and the monitoring of each data packet, the display of the online and offline speech recognition results can be flexibly switched, avoiding a poor experience for the user due to waiting. Further, by referring to the display result of the previous data packet, the display method of the current data packet can be determined more intelligently. When it is found that the online link is unstable (the previous packet is offline display), offline display is quickly and uniformly adopted, reducing unnecessary network waiting time.

[0073] In step S330, the online speech recognition result and the offline speech recognition result are fused to obtain a fused speech recognition result.

[0074] In step S340, the local semantic understanding model is used to perform semantic understanding on the fused speech recognition result to obtain an offline semantic understanding result.

[0075] In the embodiment of the present disclosure, the online speech recognition result and the offline speech recognition result are fused to obtain a fused speech recognition result, and then in the offline link, semantic understanding is performed on the fused speech recognition result to obtain an offline semantic understanding result. Among them, online speech recognition is obtained based on the powerful computing power and massive data of the server. Therefore, the accuracy of the online speech recognition result is higher than that of the offline speech recognition result. Therefore, in the offline link, the online speech recognition result and the offline speech recognition result are fused, and then semantic understanding is performed using the fused speech recognition result of the two, so as to improve the accuracy of semantic understanding.

[0076] In some embodiments of the present disclosure, the online speech recognition result and the offline speech recognition result are fused to obtain a fused speech recognition result, including: in response to not receiving the online speech recognition result within the second preset time, using the offline speech recognition result as the fused speech recognition result; in response to receiving the online speech recognition result within the second preset time, fusing the online speech recognition result and the offline speech recognition result to obtain the fused speech recognition result.

[0077] In the embodiments of the present disclosure, a time limit for receiving the online speech recognition result is set, that is, the second preset time, which can be timed starting from the sending of the speech data to be processed. If the online speech recognition result is not received after exceeding the second preset time, the offline speech recognition result is directly used as the fused result. If the online speech recognition result is received within the second preset time, the online speech recognition result and the offline speech recognition result can be aligned to obtain the fused speech recognition result. Exemplarily, the online speech recognition result can be given priority, or the text contents recognized by both can be compared to retain the more accurate and complete part to obtain the fused speech recognition result.

[0078] For example, the user says "Turn on the radio" and sends a speech data packet for processing. If the online speech recognition result is not received within the second preset time, the offline speech recognition result is directly used as the fused result. If the online speech recognition result is received within the second preset time, and it is found that both the online and offline recognition results are "Turn on the radio", then the fused result is "Turn on the radio"; if the online recognition result is "Turn on the radio" and the offline recognition result is "Turn on the music", it can be determined by comparing historical data, speech features, etc. that "Turn on the radio" is more in line with the user's intention to form the fused speech recognition result.

[0079] Through the above fusion processing of the online speech recognition result and the offline speech recognition result, if the online speech recognition result is received within the second preset time, the contents of the two are compared through alignment processing to determine the fused speech recognition result; if the online speech recognition result is not received after timing out, the offline speech recognition result is directly determined as the fused speech recognition result. In this way, when the network condition is good, the high-precision characteristics of online recognition can be utilized to provide an accurate basis for subsequent semantic understanding using the local semantic understanding model, effectively improving the accuracy of semantic understanding; and when the offline speech recognition result cannot be obtained in time, the offline speech recognition result is directly determined as the fused speech recognition result to ensure that the subsequent offline semantic understanding work is not interfered with.

[0080] In some embodiments of the present disclosure, the local semantic understanding model includes a local multimodal large model. Among them, the local semantic understanding model is used to perform semantic understanding on the fused speech recognition result to obtain an offline semantic understanding result, including: performing intent distribution on the fused speech recognition result to determine the task to which the fused speech recognition result belongs; using the local multimodal large model to perform semantic understanding on the fused speech recognition result according to the basic parameters of the local multimodal large model and the weight parameters corresponding to the task to which it belongs, to obtain an offline semantic understanding result; the weight parameters corresponding to the task to which it belongs are pre-trained.

[0081] Among them, semantic understanding aims to convert the user's natural language expression into structured information, and this structured information can be based on a predefined standard protocol. Exemplarily, in the control semantic understanding scenario, the semantic understanding result can be represented by code. The code mainly includes two parts: "function" (action) and "object" (object). "Function" represents the operation to be performed, such as "open", "close", "turn up", "turn down", etc., and "object" represents the target of the operation, such as a specific device or mode, etc. For ease of understanding, a specific example is provided: x0 = Attribute(name = "Driver Assistance"); Open(object = x0). When the user requests "Open intelligent driving", "Attribute" is used as the "object", referring to the "Driver Assistance" function to be operated, and "Open" is used as the "function", indicating that the open operation is to be performed on this function.

[0082] In the embodiments of the present disclosure, the local semantic understanding model may include a local multimodal large model, that is, a multimodal large model is deployed locally. The application of the multimodal large model in the field of autonomous driving is mainly reflected in its ability to simultaneously process and understand multiple types of data, such as images, text, audio, etc., thereby improving the environmental perception and decision-making ability of the autonomous driving system. It has strong scene reasoning and generalization capabilities. For example, the Vision-Language-Action (VLA) model exhibits higher scene reasoning and generalization capabilities. Deploying a multimodal large model on a vehicle-mounted device can use this multimodal large model to perform semantic understanding on the fused speech recognition result to obtain an offline semantic understanding result.

[0083] Specifically, after obtaining the fused speech recognition result, intent distribution is performed on it. Among them, intent distribution is like an intelligent classifier that analyzes the true intent of the user behind the speech recognition result and then classifies this result into the corresponding task, that is, determines which task the speech recognition result belongs to, such as a control task, a navigation task, a music playback task, an information query task, etc.

[0084] After determining the task to which the fused speech recognition result belongs, the local multi-modal large model will be called for semantic understanding. Among them, the basic parameters of the local multi-modal large model are derived from the pre-training of the model on a large amount of data, and these parameters contain rich general knowledge and feature representations. For each task, dedicated weight parameters are pre-trained, and these parameters can fine-tune the basic parameters specifically, enabling the model to perform more accurately and efficiently when processing specific tasks.

[0085] Taking the "control" task as an example, the weight parameters corresponding to this task will guide the model to focus on the semantic information related to control operations. By combining the basic parameters of the multi-modal large model with the weight parameters corresponding to the task, the local multi-modal large model can deeply analyze the speech recognition result, accurately grasp the semantic connotation therein, and finally successfully obtain the offline semantic understanding result.

[0086] Exemplarily, for different tasks, weight parameters corresponding to different tasks are pre-trained. During the training process, a large amount of task-related sample data will be used to let the model learn the patterns and rules in these data, thereby adjusting the weight parameters to enable the model to achieve better performance when processing this task. For example, when training the weight parameters of the "music playback" task, a large amount of speech data related to music search, playback control, etc. can be used to let the model learn how to accurately understand these instructions, and then optimize the weight parameters.

[0087] It should be noted that in the embodiments of the present disclosure, determining which task the speech recognition result belongs to is essentially equivalent to determining which vertical domain the speech recognition result belongs to, that is, the vertical domain. The vertical domain covers various categories such as the control vertical domain, the navigation vertical domain, the music playback vertical domain, the information query vertical domain, etc. For example, when the user issues the voice "open the window", after the intent distribution determines that the speech recognition result belongs to the control vertical domain, the weight parameters pre-trained for the control vertical domain (i.e., the control task) in the local multi-modal large model will be called, combined with the basic parameters of the local multi-modal large model, for semantic understanding. If the user requests "find the nearest gas station", it will be classified into the navigation vertical domain, and the weight parameters of the navigation vertical domain (i.e., the navigation task) and the basic parameters of the local multi-modal large model will be used for semantic understanding to obtain the offline semantic understanding result.

[0088] Through the above steps of semantic understanding using the local multimodal large model, the fused speech recognition results can be accurately classified into different tasks through intent distribution, and dedicated processing can be performed according to the unique requirements of each task; moreover, the basic parameters of the local multimodal large model carry rich general knowledge, and combined with the weight parameters pre-trained for each task, accurate and efficient semantic understanding can be achieved in specific task processing; in addition, the multimodal large model can process multiple types of data simultaneously, improving the vehicle environment perception and decision-making ability.

[0089] In some embodiments of the present disclosure, after obtaining the offline semantic understanding result, the speech processing method further includes: determining whether the offline speech recognition result is a user request according to the speech data to be processed, the offline speech recognition result, and the offline semantic understanding result; in response to the offline speech recognition result being a user request, determining to perform speech conversion on the offline semantic understanding result.

[0090] In the embodiments of the present disclosure, the original speech data to be processed, the result obtained by offline speech recognition, and the offline semantic understanding result can be comprehensively analyzed to determine whether the recognized content is a valid request from the user. For example, by checking keywords, sentence structures, etc. If it is determined to be a user request, the offline semantic understanding result is converted into speech; otherwise, it can be determined that there is no need to convert the offline semantic understanding result into speech. For example, when the user coughs accidentally and it is also recognized and semantically understood, but through analysis and judgment, it is found that the speech recognition result converted from the cough sound does not conform to the pattern of common user requests, so the semantic understanding result is not converted into speech output. When the user says "raise the air conditioner temperature", which is determined to be a user request after judgment, the semantic understanding result (i.e., the instruction to raise the air conditioner temperature) is converted into speech and fed back to the user.

[0091] Figure 4 is a schematic diagram of performing offline semantic understanding shown according to some embodiments of the present disclosure. Refer to Figure 4 , and can include parts such as intent distribution, vertical domain semantic understanding, intent selection, result output, and post-rejection recognition.

[0092] Among them, intent distribution is carried out according to the content (query) of the user request to judge the category to which the request belongs, and then the judgment result is assigned to the corresponding vertical domain, and then more detailed semantic understanding is carried out, for example, the navigation vertical domain, the music vertical domain, the call vertical domain, the control vertical domain. If the category to which the request belongs cannot be accurately judged, the request will be sent to multiple possibly relevant vertical domains at the same time. In subsequent intent selection processing, according to the analysis results of each vertical domain, the most appropriate vertical domain semantic understanding result is selected for decision-making.

[0093] From Figure 4It can be seen that semantic understanding can be carried out according to different vertical domains, and independent models and professional knowledge in each domain are set. Before the introduction of the multimodal large model, semantic understanding models based on BERT can be adopted in each vertical domain. Considering the high frequency of control requests in the vehicle-mounted scenario, as Figure 4 shown, the model in the control vertical domain can be upgraded from a semantic understanding model based on BERT to a local multimodal large model to achieve more accurate semantic understanding. Among them, the upgraded control vertical domain is named the Control Agent Parser.

[0094] When the central control intention distributes requests to multiple vertical domains, during the intention selection and processing, based on the scoring features of each vertical domain, following the highest score principle, the optimal semantic understanding result is selected from the semantic understanding results of multiple vertical domains.

[0095] The process of semantic understanding is to convert the user's natural language expression into structured information. After obtaining the semantic understanding result of the vertical domain, the semantic understanding result of this vertical domain can be analyzed, and the analyzed semantic understanding result is output as an instruction set. These instruction sets are operation information that can be directly executed on the client according to the protocol defined in advance. For example, for the semantic understanding result of the navigation vertical domain, an instruction set containing operation information such as destination planning and route guidance may be output; for the semantic understanding result of the music vertical domain, an instruction set related to music playback control such as playing a specific song and adjusting the volume will be generated; for the semantic understanding result of the call vertical domain, an instruction set such as dialing a specified number and hanging up the phone can be output; for the semantic understanding result of the control vertical domain, an instruction set for device operations such as window lifting, door opening and closing, and air conditioning temperature adjustment will be generated.

[0096] The specific implementation of post-rejection recognition is to comprehensively consider the content of the user's request (query), the original audio data, and the semantic understanding result of the vertical domain to judge whether the current request is a real human-computer interaction request. If it is a real human-computer interaction request, the result corresponding to the request is output. If it is not a real human-computer interaction request, it is determined not to output the result corresponding to the request. In this way, non-effective requests such as background noise and conversations between people are blocked, reducing unnecessary disturbances.

[0097] Figure 4In the semantic understanding architecture on the vehicle-mounted device side shown, for other vertical domains outside the control vertical domain, it is also possible to upgrade from the semantic understanding model based on BERT to a local multi-modal large model. For example, in the music vertical domain, the semantic understanding model based on BERT may be limited to relatively simple text matching in understanding the user's music instructions. After upgrading to a local multi-modal large model, it can not only understand the text meaning in the user's voice instructions, but also combine multi-modal information such as the melody hummed by the user and historical data on music style preferences to accurately identify the user's needs and recommend more satisfactory music. Another example is the navigation vertical domain. The semantic understanding model based on BERT may only be able to process simple text instructions such as place names, while the local multi-modal large model can integrate multi-modal data such as map image information and temporal and spatial characteristics of the user's travel habits to more accurately understand the user's navigation request and plan a more optimized route.

[0098] In step S350, the offline semantic understanding result is subjected to speech conversion to obtain offline speech data.

[0099] In the embodiments of the present disclosure, in the offline link, the offline semantic understanding result is subjected to speech conversion to obtain offline speech data. Exemplarily, a local TTS module is deployed on the vehicle-mounted device, and through the local TTS module, the offline semantic understanding result is converted into speech to obtain offline speech data.

[0100] In some embodiments of the present disclosure, after obtaining the offline semantic understanding result, the speech processing method further includes: performing a fusion process on the online semantic understanding result and the offline semantic understanding result to obtain a fused semantic understanding result.

[0101] In the embodiments of the present disclosure, the online semantic understanding result and the offline semantic understanding result are subjected to a fusion process to obtain a fused semantic understanding result. Among them, the online semantic understanding is obtained based on the powerful computing power and massive data of the server. Therefore, the accuracy of the online semantic understanding result is higher than that of the offline semantic understanding result. So, in the offline link, fusing the online semantic understanding result and the offline semantic understanding result can obtain a fused semantic understanding result with high accuracy.

[0102] In some embodiments of the present disclosure, performing a fusion process on the online semantic understanding result and the offline semantic understanding result to obtain a fused semantic understanding result includes: in response to not receiving the online semantic understanding result within a third preset time, using the offline semantic understanding result as the fused semantic understanding result; in response to receiving the online semantic understanding result within the third preset time, performing a fusion process on the online semantic understanding result and the offline semantic understanding result to obtain a fused semantic understanding result.

[0103] In the embodiments of the present disclosure, a time limit for receiving the online semantic understanding result is set, that is, the third preset time, which can be timed from the start of sending the voice data to be processed. If the online semantic understanding result is not received after exceeding the third preset time, the offline semantic understanding result is used as the fused offline semantic understanding result. If the online semantic understanding result is received within the third preset time, the online semantic understanding result and the offline semantic understanding result can be aligned to obtain the fused semantic understanding result. Exemplarily, the online semantic understanding result can be given priority, or the in-depth analysis of complex semantics in the online semantic understanding result can be combined with the understanding of the local specific scenario in the offline semantic understanding result to obtain the fused semantic understanding result.

[0104] In some embodiments of the present disclosure, according to the online processing result and the offline processing result, the target processing result corresponding to the voice data to be processed is determined, including: in response to the fused semantic understanding result being the online semantic understanding result, taking the online voice data as the target processing result; in response to the fused semantic understanding result being the offline semantic understanding result, taking the offline voice data as the target processing result.

[0105] In the embodiments of the present disclosure, if the fused semantic understanding result comes from online processing, the online voice data is selected as the final target processing result. Conversely, if the fused semantic understanding result is the offline semantic understanding result, the offline voice data is selected as the target processing result.

[0106] Through the above steps, if the online semantic understanding result is received within the third preset time, the content of the two is compared through alignment processing to determine the fused semantic understanding result, making the fused semantic understanding result more comprehensive and accurate; if the online semantic understanding result is not received after the timeout, the offline semantic understanding result is directly determined as the fused semantic understanding result to ensure the continuous operation of the semantic understanding function and maintain service coherence. In the stage of determining the target processing result, the corresponding online or offline voice data is selected according to the fused semantic understanding result, ensuring the consistency of voice processing and semantic understanding, improving the user interaction experience. Whether the network is smooth or blocked, users can obtain reliable and adaptable processing results, enhancing the stability and applicability of the system.

[0107] The following takes the application of the method provided in the embodiments of the present disclosure to a vehicle as an example, such as it can be applied to a car cockpit. However, the present disclosure is not limited thereto. For example, it can also be applied to the following fields: a smart home system, where users control smart home appliances at home through voice commands, such as smart lights, air conditioners, smart curtains, etc.; a smart education platform, where students can ask questions to the system through voice when using learning devices, such as subject knowledge answers, homework tutoring, etc., and the system can understand the semantics of the questions and accurately push relevant learning materials and explanation videos to achieve personalized learning tutoring.

[0108] Exemplarily, the vehicle may be an intelligent vehicle, which has an intelligent cockpit. An intelligent vehicle is a new generation of vehicle that is equipped with advanced sensors, controllers, actuators and other devices, integrates new technologies such as information communication, Internet, big data, cloud computing, artificial intelligence, etc., has partial or fully autonomous driving functions, and transforms from a traditional means of transportation to an intelligent mobile space. An intelligent cockpit is a new type of in-vehicle space that integrates driving information, in-vehicle applications and interaction technologies based on the background of intelligentization and Internet of Everything to create an efficient and technological driving experience for users. The intelligent voice interaction technology has not only become a standard configuration of intelligent vehicles but is also developing rapidly, continuously improving the user experience by introducing more technologies.

[0109] Next, an embodiment of voice data processing through an in-vehicle voice assistant system and a cloud server will be described with the server being a cloud server. Figure 5 It is a schematic diagram of the process of a voice data processing method through an in-vehicle voice assistant system and a cloud server shown according to some exemplary embodiments of the present disclosure. Refer to Figure 5 , a cloud Speech module, a cloud ASR model, a cloud NLP model and a cloud TTS module are deployed on the cloud server, and a local Speech module, a local ASR model, a local NLP model and a local TTS module are deployed locally.

[0110] As Figure 5 shown, after the in-vehicle voice assistant client receives the voice data input by the user, it determines the currently available link according to the current network status and the link availability status. If it is determined that the offline link is unavailable and the device is in an online state, the user voice data is processed through the online link. If it is determined that the device is currently in an offline state, the user voice data is processed through the offline link. If it is determined that the offline link is available and the device is in an online state, the user voice data is processed through both the online link and the offline link.

[0111] For the online link, the user voice data is distributed to the cloud Speech module for processing. First, the cloud ASR model performs speech recognition processing on the user voice data to convert the audio into text, obtaining an online speech recognition result. For the offline link, the user voice data is distributed to the local Speech module for processing. First, the local ASR model performs speech recognition processing on the user voice data to convert the audio into text, obtaining an offline speech recognition result.

[0112] Among them, the cloud ASR model and the local ASR model adopt a streaming recognition method, and a speech recognition result is output every time 1 packet of voice data is received. From Figure 4It can be seen that each time the cloud-based ASR model outputs a speech recognition result, it is returned to the in-vehicle voice assistant system so that the ASR screen selection can be made between the online speech recognition result obtained by the cloud-based ASR model and the offline speech recognition result obtained by the local ASR model. Exemplarily, if the online speech recognition result obtained by the cloud-based ASR model is received within a preset time, the online speech recognition result is directly displayed on the in-vehicle large screen. If the online speech recognition result obtained by the cloud-based ASR model is not received within a timeout, the offline speech recognition result obtained by the local ASR model is displayed on the in-vehicle large screen, and for subsequent data packets, the offline speech recognition result is displayed.

[0113] Figure 5 In the process, the cloud ASR model and the local ASR module determine whether the sound reception is completed. For example, if the cloud ASR model determines that no sound is received within a period of time after receiving a certain voice data packet, it is considered that the sound reception is completed. For the online link, after the sound reception is determined to be completed, the online speech recognition result recognized by the cloud ASR model is transmitted to the cloud NLP model, and semantic understanding is performed based on the cloud NLP model to obtain the online semantic understanding result, which is then transmitted to the cloud TTS module for text-to-speech conversion to obtain online voice data.

[0114] For offline links, after determining that the sound reception is completed, the online speech recognition result is fused with the offline speech recognition result, and the online speech recognition result recognized by the cloud ASR model is given priority. If the recognition result returned by the cloud ASR model times out, the offline speech recognition result obtained by the offline ASR model is used. The fused speech recognition result is then transmitted to the local NLP model, and semantic understanding is performed based on the local NLP model to obtain the online semantic understanding result, which is then transmitted to the local TTS module for text-to-speech conversion to obtain local speech data. In the disclosed embodiment, the local NLP model can be a local multimodal large model.

[0115] from Figure 5It can be seen that the cloud returns the online semantic understanding result and the online voice data to the in-vehicle voice assistant client for the fusion processing of the semantic understanding result and the fusion processing of the voice data. When fusing the online semantic understanding result and the offline semantic understanding result, the online semantic understanding result recognized by the cloud NLP model is given priority. If the understanding result returned by the cloud NLP model times out, the offline semantic understanding result obtained by using the offline NLP model is used. When fusing the online voice data and the offline voice data, it can be kept consistent with the fusion result of the semantic understanding result. If the fusion result of the semantic understanding result is the online semantic understanding result, the online voice data is selected as the target processing result. If the fusion result of the semantic understanding result is the offline semantic understanding result, the offline voice data is selected as the target processing result. Finally, the target processing result is sent to the in-vehicle voice assistant client.

[0116] For the voice processing method provided by the embodiments of the present disclosure, after obtaining the voice data input by the user, the current network and link status are detected to determine the currently available link. When it is detected that the currently available link includes an online link and an offline link, a dual-link parallel processing mechanism is adopted to perform online voice processing and offline voice processing; the online processing result provided by the online link is fused with the offline processing result obtained by the offline link, and factors such as recognition accuracy and response latency are comprehensively considered to generate an optimal target processing result. In this way, it is possible to process the collected voice data in real time and stably in the complex network environment of the vehicle, improve the response speed and processing accuracy of the in-vehicle voice interaction system, and enhance the operation convenience and safety during driving.

[0117] It should be noted that the above-mentioned drawings are only schematic illustrations of the processes included in the methods according to some embodiments of the present disclosure, rather than for limiting purposes. It is easy to understand that the processes shown in the above-mentioned drawings do not indicate or limit the time sequence of these processes. In addition, it is also easy to understand that these processes can be executed synchronously or asynchronously in, for example, multiple modules.

[0118] The following are the embodiments of the device of the present disclosure, which can be used to execute the embodiments of the method of the present disclosure. For the details not disclosed in the embodiments of the device of the present disclosure, please refer to the embodiments of the method of the present disclosure.

[0119] Figure 6 is a block diagram of a voice processing device shown according to some embodiments of the present disclosure. Referring to Figure 6 , the device 600 is applied to an in-vehicle device and may include: a voice data acquisition module 610, an available link determination module 620, a voice data processing module 630, and a processing result determination module 640.

[0120] Among them, the voice data acquisition module 610 is configured to: acquire voice data to be processed. The available link determination module 620 is configured to: determine the currently available link. The voice data processing module 630 is configured to: in response to the currently available link including an online link and an offline link, perform parallel processing on the voice data to be processed through the online link and the offline link to obtain an online processing result and an offline processing result. The processing result determination module 640 is configured to: determine the target processing result corresponding to the voice data to be processed according to the online processing result and the offline processing result.

[0121] In some embodiments of the present disclosure, the voice data processing module 630 is further configured to: send the voice data to be processed to a server, and receive the online processing result returned by the server in response to the voice data to be processed; the online processing result includes at least one of an online speech recognition result, an online semantic understanding result, and online voice data; perform speech recognition on the voice data to be processed to obtain an offline speech recognition result; perform fusion processing on the online speech recognition result and the offline speech recognition result to obtain a fused speech recognition result; use a local semantic understanding model to perform semantic understanding on the fused speech recognition result to obtain an offline semantic understanding result; perform speech conversion on the offline semantic understanding result to obtain offline voice data.

[0122] In some embodiments of the present disclosure, the local semantic understanding model includes a local multimodal large model; among them, the voice data processing module 630 is further configured to: perform intent distribution on the fused speech recognition result to determine the task to which the fused speech recognition result belongs; according to the basic parameters of the local multimodal large model and the weight parameters corresponding to the task to which it belongs, use the local multimodal large model to perform semantic understanding on the fused speech recognition result to obtain an offline semantic understanding result; the weight parameters corresponding to the task to which it belongs are pre-trained.

[0123] In some embodiments of the present disclosure, the voice data processing module 630 is further configured to: determine whether the offline speech recognition result is a user request according to the voice data to be processed, the offline speech recognition result, and the offline semantic understanding result; in response to the offline speech recognition result being a user request, determine to perform speech conversion on the offline semantic understanding result.

[0124] In some embodiments of the present disclosure, the voice data processing module 630 is further configured to: perform fusion processing on the online semantic understanding result and the offline semantic understanding result to obtain a fused semantic understanding result.

[0125] In some embodiments of the present disclosure, the processing result determination module 640 is further configured to: in response to the fused semantic understanding result being an online semantic understanding result, use the online voice data as the target processing result; in response to the fused semantic understanding result being an offline semantic understanding result, use the offline voice data as the target processing result.

[0126] In some embodiments of the present disclosure, the voice data to be processed includes one or more data packets; wherein, the voice data processing module 630 is further configured to: for the current data packet, determine whether the online speech recognition result of the current data packet is received within a first preset time; in response to receiving the online speech recognition result of the current data packet within the first preset time, display the online speech recognition result of the voice data packet; in response to not receiving the online speech recognition result of the current data packet within the first preset time, display the offline speech recognition result of the voice data packet.

[0127] In some embodiments of the present disclosure, the voice data processing module 630 is further configured to: for the previous data packet of the current data packet, determine whether to display the offline speech recognition result of the previous data packet; in response to displaying the offline speech recognition result of the previous data packet, display the offline speech recognition result of the current data packet; in response to displaying the online speech recognition result of the previous data packet, analyze the display method of the speech recognition result of the current data packet.

[0128] In some embodiments of the present disclosure, the voice data processing module 630 is further configured to: in response to not receiving the online speech recognition result within a second preset time, use the offline speech recognition result as the fused speech recognition result; in response to receiving the online speech recognition result within the second preset time, perform a fusion process on the online speech recognition result and the offline speech recognition result to obtain the fused speech recognition result.

[0129] In some embodiments of the present disclosure, the voice data processing module 630 is further configured to: in response to not receiving the online semantic understanding result within a third preset time, use the offline semantic understanding result as the fused semantic understanding result; in response to receiving the online semantic understanding result within the third preset time, perform a fusion process on the online semantic understanding result and the offline semantic understanding result to obtain the fused semantic understanding result.

[0130] In some embodiments of the present disclosure, the voice data processing module 630 is further configured to: in response to the current available link including an online link, process the voice data to be processed through the online link, and use the processing result of the online link as the target processing result.

[0131] In some embodiments of the present disclosure, the voice data processing module 630 is further configured to: in response to the currently available link including an offline link, process the voice data to be processed through the offline link, and use the offline link processing result as the target processing result.

[0132] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.

[0133] Figure 7 It is a functional block diagram of a vehicle shown according to some embodiments of the present disclosure. For example, the vehicle 700 may be a hybrid vehicle, or a non - hybrid vehicle, an electric vehicle, a fuel cell vehicle, or other types of vehicles. The vehicle 700 may be an autonomous vehicle, a semi - autonomous vehicle, or a non - autonomous vehicle.

[0134] Referring to Figure 7 , the vehicle 700 may include various subsystems. For example, the infotainment system 710, the perception system 720, the decision - making and control system 730, the drive system 740, and the computing platform 750. Among them, the vehicle 700 may further include more or fewer subsystems, and each subsystem may include multiple components. In addition, each subsystem and each component of the vehicle 700 may be interconnected in a wired or wireless manner.

[0135] In some embodiments, the infotainment system 710 may include a communication system, an entertainment system, and a navigation system, etc.

[0136] The perception system 720 may include several sensors for sensing information about the environment around the vehicle 700. For example, the perception system 720 may include a global positioning system (the global positioning system may be a GPS system, or a Beidou system, or other positioning systems), an inertial measurement unit (IMU), a lidar, a millimeter - wave radar, an ultrasonic radar, and a camera device.

[0137] The decision - making and control system 730 may include a computing system, a vehicle controller, a steering system, an accelerator, and a braking system.

[0138] The drive system 740 may include components that provide power movement for the vehicle 700. In one embodiment, the drive system 740 may include an engine, an energy source, a transmission system, and wheels. The engine may be one or a combination of an internal combustion engine, an electric motor, and an air compression engine. The engine can convert the energy provided by the energy source into mechanical energy.

[0139] Some or all of the functions of vehicle 700 are controlled by computing platform 750. Computing platform 750 may include at least one processor 751 and a memory 752, and processor 751 may execute instructions 753 stored in memory 752.

[0140] Processor 751 can be any conventional processor, such as a commercially available CPU. The processor may also include, for example, a Graphic Process Unit (GPU), a Field Programmable Gate Array (FPGA), a System on Chip (SOC), an Application Specific Integrated Circuit (ASIC), or a combination thereof.

[0141] Memory 752 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.

[0142] In addition to instructions 753, memory 752 may also store data, such as road maps, route information, data on the position, direction, speed, etc. of the vehicle. The data stored in memory 752 can be used by computing platform 750.

[0143] In an embodiment of the present disclosure, processor 751 may execute instructions 753 to complete all or part of the steps of the above-described voice processing method.

[0144] The present disclosure also provides a computer-readable storage medium having computer program instructions stored thereon, and when the program instructions are executed by a processor, the steps of the voice processing method provided by the present disclosure are implemented.

[0145] In addition, as used herein, the word "exemplary" is used to mean serving as an example, instance, or illustration. Any aspect or design described herein as "exemplary" is not necessarily to be construed as advantageous over other aspects or designs. Rather, the word exemplary is intended to present concepts in a concrete fashion. As used herein, the term "or" is intended to mean an inclusive "or" rather than an exclusive "or". That is, unless specified otherwise, or clear from the context, "X applies A or B" is intended to mean any of the natural inclusive permutations. That is, if X applies A; X applies B; or X applies both A and B, then "X applies A or B" is satisfied under any of the foregoing instances. Additionally, unless specified otherwise or clear from the context that it is referring to the singular form, the articles "a" and "an" as used in this application and the appended claims are generally understood to mean "one or more".

[0146] Likewise, although the present disclosure has been shown and described with respect to one or more implementations, equivalent variations and modifications will occur to those skilled in the art upon reading and understanding this specification and the drawings. The present disclosure includes all such modifications and variations and is limited only by the scope of the claims. Specifically with respect to the various functions performed by the components (e.g., elements, resources, etc.) described above, unless otherwise indicated, the terms used to describe such components are intended to correspond to any component (functionally equivalent) that performs the specific function of the described component, even if not structurally equivalent to the disclosed structure. Additionally, although certain features of the present disclosure may have been disclosed with respect to only one of several implementations, such features may, as may be desired and advantageous for any given or particular application, be combined with one or more other features of other implementations. Further, with respect to the use of "comprises", "comprising", "has", "having", "includes", or variants thereof in the detailed description or claims, such terms are intended to be inclusive in a manner similar to the term "including".

[0147] Other embodiments of the present disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of the present disclosure are pointed out by the following claims.

[0148] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes may be made without departing from its scope. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A voice processing method, characterized in that, The method includes: Obtaining the voice data to be processed; Determining the currently available link; In response to the currently available link including an online link and an offline link, processing the voice data to be processed in parallel through the online link and the offline link to obtain an online processing result and an offline processing result; Determining the target processing result corresponding to the voice data to be processed according to the online processing result and the offline processing result.

2. The method according to claim 1, characterized in that The processing the voice data to be processed in parallel through the online link and the offline link to obtain an online processing result and an offline processing result includes: Sending the voice data to be processed to a server and receiving the online processing result returned by the server in response to the voice data to be processed; the online processing result includes at least one of an online speech recognition result, an online semantic understanding result, and online voice data; Performing speech recognition on the voice data to be processed to obtain an offline speech recognition result; Performing fusion processing on the online speech recognition result and the offline speech recognition result to obtain a fused speech recognition result; Performing semantic understanding on the fused speech recognition result by using a local semantic understanding model to obtain an offline semantic understanding result; Performing speech conversion on the offline semantic understanding result to obtain offline voice data.

3. The method according to claim 2, wherein The local semantic understanding model includes a local multimodal large model; Wherein, the performing semantic understanding on the fused speech recognition result by using the local semantic understanding model to obtain an offline semantic understanding result includes: Performing intent distribution on the fused speech recognition result to determine the task to which the fused speech recognition result belongs; According to the basic parameters of the local multimodal large model and the weight parameters corresponding to the task to which it belongs, using the local multimodal large model to perform semantic understanding on the fused speech recognition result to obtain the offline semantic understanding result; the weight parameters corresponding to the task to which it belongs are pre-trained.

4. The method according to claim 2, wherein After obtaining the offline semantic understanding result, the method further includes: Determining whether the offline speech recognition result is a user request according to the voice data to be processed, the offline speech recognition result, and the offline semantic understanding result; In response to the offline speech recognition result being a user request, determining to perform speech conversion on the offline semantic understanding result.

5. The method according to claim 2, wherein After obtaining the offline semantic understanding result, the method further includes: Performing fusion processing on the online semantic understanding result and the offline semantic understanding result to obtain a fused semantic understanding result.

6. The method according to claim 5, characterized in that, The determining the target processing result corresponding to the voice data to be processed according to the online processing result and the offline processing result includes: In response to the fused semantic understanding result being the online semantic understanding result, using the online voice data as the target processing result; In response to the fused semantic understanding result being the offline semantic understanding result, using the offline voice data as the target processing result.

7. The method according to claim 2, characterized in that, The voice data to be processed includes one or more data packets; wherein, the method further includes: For the current data packet, determine whether the online speech recognition result of the current data packet is received within a first preset time; In response to receiving the online speech recognition result of the current data packet within the first preset time, display the online speech recognition result of the speech data packet; In response to not receiving the online speech recognition result of the current data packet within the first preset time, display the offline speech recognition result of the speech data packet.

8. The method according to claim 7, wherein Before determining whether the online speech recognition result of the current data packet is received within the first preset time, the method further includes: For the previous data packet of the current data packet, determine whether to display the offline speech recognition result of the previous data packet; In response to displaying the offline speech recognition result of the previous data packet, display the offline speech recognition result of the current data packet; In response to displaying the online speech recognition result of the previous data packet, determine to analyze the display method of the speech recognition result of the current data packet.

9. The method according to claim 2, characterized in that The fusion processing of the online speech recognition result and the offline speech recognition result to obtain the fused speech recognition result includes: In response to not receiving the online speech recognition result within a second preset time, use the offline speech recognition result as the fused speech recognition result; In response to receiving the online speech recognition result within the second preset time, perform fusion processing on the online speech recognition result and the offline speech recognition result to obtain the fused speech recognition result.

10. The method according to claim 5, wherein The fusion processing of the online semantic understanding result and the offline semantic understanding result to obtain the fused semantic understanding result includes: In response to not receiving the online semantic understanding result within a third preset time, use the offline semantic understanding result as the fused semantic understanding result; In response to receiving the online semantic understanding result within the third preset time, perform fusion processing on the online semantic understanding result and the offline semantic understanding result to obtain the fused semantic understanding result.

11. The method according to claim 1, characterized in that The method further includes: In response to the current available link including the online link, process the speech data to be processed through the online link, and use the processing result of the online link as the target processing result.

12. The method according to claim 1, wherein The method further includes: In response to the current available link including the offline link, process the speech data to be processed through the offline link, and use the processing result of the offline link as the target processing result.

13. A voice processing device, characterized in that, The device includes: A speech data acquisition module configured to acquire speech data to be processed; An available link determination module configured to determine the current available link; A speech data processing module configured to, in response to the current available link including an online link and an offline link, perform parallel processing on the speech data to be processed through the online link and the offline link to obtain an online processing result and an offline processing result; A processing result determination module configured to determine the target processing result corresponding to the speech data to be processed according to the online processing result and the offline processing result.

14. The device according to claim 13, wherein The speech data processing module is further configured to: Send the speech data to be processed to the server and receive the online processing result returned by the server in response to the speech data to be processed; the online processing result includes at least one of an online speech recognition result, an online semantic understanding result, and online speech data; Perform speech recognition on the speech data to be processed to obtain an offline speech recognition result; Perform fusion processing on the online speech recognition result and the offline speech recognition result to obtain a fused speech recognition result; Use a local semantic understanding model to perform semantic understanding on the fused speech recognition result to obtain an offline semantic understanding result; Perform speech conversion on the offline semantic understanding result to obtain offline speech data.

15. The device according to claim 14, characterized in that, The local semantic understanding model includes a local multimodal large model; Wherein, the speech data processing module is further configured to: Perform intent distribution on the fused speech recognition result to determine the task to which the fused speech recognition result belongs; According to the basic parameters of the local multimodal large model and the weight parameters corresponding to the task to which it belongs, use the local multimodal large model to perform semantic understanding on the fused speech recognition result to obtain the offline semantic understanding result; the weight parameters corresponding to the task to which it belongs are pre-trained.

16. The device according to claim 14, characterized in that, The speech data processing module is further configured to: Perform fusion processing on the online semantic understanding result and the offline semantic understanding result to obtain a fused semantic understanding result.

17. The device according to claim 14, characterized in that, The speech data processing module is further configured to: In response to not receiving the online speech recognition result within a second preset time, use the offline speech recognition result as the fused speech recognition result; In response to receiving the online speech recognition result within the second preset time, perform fusion processing on the online speech recognition result and the offline speech recognition result to obtain the fused speech recognition result.

18. The device according to claim 17, characterized in that, The speech data processing module is further configured to: In response to not receiving the online semantic understanding result within a third preset time, use the offline semantic understanding result as the fused semantic understanding result; In response to receiving the online semantic understanding result within the third preset time, perform fusion processing on the online semantic understanding result and the offline semantic understanding result to obtain the fused semantic understanding result.

19. The device according to claim 13, characterized in that, The speech data processing module is further configured to: In response to the current available link including the online link, process the speech data to be processed through the online link and use the processing result of the online link as the target processing result.

20. The device according to claim 13, characterized in that The speech data processing module is further configured to: In response to the current available link including the offline link, process the speech data to be processed through the offline link and use the processing result of the offline link as the target processing result.

21. An electronic device, characterized in that, It includes: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to implement the speech processing method according to any one of claims 1-12.

22. A non-transitory computer-readable storage medium, when the instructions in the storage medium are executed by a processor of a terminal device, enable the terminal device to execute the voice processing method according to any one of claims 1-12.

23. A computer program product, characterized in that, Comprising a computer program, which when executed by a processor, implements the voice processing method according to any one of claims 1-12.