Voice interaction system, method, device and medium of cloud AI large model vehicle-mounted TBox
Patent Information
- Application Number
- CN202611097384.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-23
- Publication Date
- 2026-10-09
AI Technical Summary
此类方案虽降低车端硬件要求,但仍存在以下缺陷:(1)硬件冗余——需额外加装麦克风阵列、语音处理器及通信模组,未复用车辆已广泛搭载的TBox基础能力;(2)链路不可靠——语音数据经手机中转导致多跳传输,端到端延迟常超2.5秒,且易受蓝牙/WiFi连接稳定性影响;(3)安全合规风险——原始语音流未经端侧预处理与加密即上传,存在声纹信息泄露、对话内容被截获等数据隐患
Smart Images

Figure CN122888979A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent connected vehicle technology, and in particular to voice interaction systems, methods, devices and media for in-vehicle TBoxes with cloud-based AI large models. Background Technology
[0002] Currently, voice interaction in smart cockpits has become a core indicator of user experience, but its implementation faces a significant cost-performance-adaptability triangle. Mainstream solutions can be divided into two categories: The first is a localized AI deployment solution: relying on a high-performance SoC to deploy a lightweight large model or a voice-specific model in the intelligent cockpit domain controller (IVI) or cockpit-driver integrated domain controller. While this solution offers advantages such as low latency and offline availability, it has high hardware barriers and a sharp increase in BOM costs (over ¥2000 per terminal), making it difficult to cover A-class and lower-end economy vehicles, let alone benefit mass-produced traditional vehicles. The second category is a pure cloud-based AI solution: relying on smartphone apps or aftermarket smart terminals to collect voice data and upload it to a public cloud large model for ASR / TTS / NLP processing. While such solutions reduce the hardware requirements of the vehicle, they still have the following drawbacks: (1) Hardware redundancy – additional microphone arrays, voice processors, and communication modules are required, and the basic capabilities of the TBox already widely used in vehicles are not reused; (2) Unreliable links – voice data is relayed through mobile phones, resulting in multi-hop transmission, end-to-end delays often exceeding 2.5 seconds, and is easily affected by the stability of Bluetooth / WiFi connections; (3) Security and compliance risks – the original voice stream is uploaded without preprocessing and encryption on the device side, posing data risks such as leakage of voiceprint information and interception of dialogue content. Therefore, how to improve the data security and responsiveness of voice interaction in smart cockpits has become a technical issue that cannot be ignored. Summary of the Invention
[0003] In view of this, the purpose of this application is to provide a voice interaction system, method, device and medium for an in-vehicle TBox with a cloud-based AI large model. By utilizing the existing in-vehicle TBox hardware resources, AI voice interaction functions can be realized through software upgrades and cloud integration, making full use of existing hardware resources to improve the data security and responsiveness of voice interaction in the smart cockpit.
[0004] This application provides a voice interaction system for an in-vehicle TBox based on a cloud-based AI large model. The voice interaction system includes a cloud-based AI large model and a communication connection between the cloud-based AI large model and the in-vehicle TBox. The in-vehicle TBox is used to collect the driver's voice commands, and after noise reduction, optimization and encryption processing, the encrypted voice data is obtained and uploaded to the cloud AI large model module; wherein, the in-vehicle TBox includes a voice collection and optimization module, a data encryption and transmission module, a command parsing and execution module and a voiceprint recognition module; The cloud-based AI model is used to decrypt, recognize, analyze, and synthesize the received encrypted voice data to generate voice response data. The voice response data is then encrypted and transmitted to the in-vehicle TBox module, so that the in-vehicle TBox module can decrypt and play the voice response data.
[0005] In one possible implementation, the voice acquisition optimization module is used for: The built-in noise reduction algorithm filters environmental noise and suppresses driving noise on the collected driver's voice, and outputs the noise-reduced voice signal to the data encryption transmission module.
[0006] In one possible implementation, the data encryption transmission module is used for: The uplink voice data is encrypted using an encryption protocol, and the downlink voice response data is decrypted from the cloud. All voice data transmitted between the vehicle-mounted TBox module and the cloud-based AI large model module remains encrypted.
[0007] In one possible implementation, the voiceprint recognition module is used for: The system collects the vehicle owner's voiceprint information and compares it with the voiceprint features pre-stored in the cloud database to achieve identity verification. Voice commands from unauthorized personnel will not be responded to.
[0008] In one possible implementation, the cloud-based AI big model includes: The speech-to-text module is used to perform automatic speech recognition on the decrypted speech data and output text commands by combining the noise-reduced and optimized speech signal. The semantic understanding and instruction generation module is used to perform intent recognition and contextual reasoning on the text instructions based on a large language model, and generate a structured text response that meets the user's needs. The speech synthesis module is used to convert structured text responses into natural and fluent speech waveform data.
[0009] In one possible implementation, the cloud-based AI big data model further includes a data storage and update module, which is used for: Securely store user voiceprint features, interaction history, and personalized preferences, and support algorithm updates for AI models.
[0010] This application embodiment also provides a voice interaction method for an in-vehicle TBox with a cloud-based AI large model, the voice interaction method including: The vehicle-mounted TBox collects the owner's voice commands, which are then processed through noise reduction, optimization, and encryption to obtain encrypted voice data, which is then uploaded to the cloud-based AI large model module. The cloud-based AI model decrypts, performs speech recognition, semantic analysis, and speech synthesis on the received encrypted voice data to generate voice response data. The voice response data is then encrypted and transmitted to the in-vehicle TBox module, which decrypts and plays the voice response data.
[0011] In one possible implementation, the voice interaction method further includes: The vehicle-mounted TBox module is controlled to receive upgrade packages pushed by the cloud-based AI large model via OTA remote online upgrade; wherein, the upgrade package includes upgrade packages for the voice acquisition optimization module, the data encryption transmission module, the command parsing and execution module, and the voiceprint recognition module.
[0012] This application also provides an electronic device, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the memory via the bus. When the machine-readable instructions are executed by the processor, the steps of the voice interaction method of the cloud AI big model vehicle TBox described above are performed.
[0013] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the voice interaction method for the in-vehicle TBox of the cloud-based AI big data model described above.
[0014] This application provides a voice interaction system, method, device, and medium for an in-vehicle TBox based on a cloud-based AI large-scale model. The voice interaction system includes a cloud-based AI large-scale model and a communication connection between the cloud-based AI large-scale model and the in-vehicle TBox. The in-vehicle TBox is used to collect the driver's voice commands, obtain encrypted voice data after noise reduction optimization and encryption processing, and upload it to the cloud-based AI large-scale model module. The in-vehicle TBox includes a voice acquisition and optimization module, a data encryption and transmission module, a command parsing and execution module, and a voiceprint recognition module. The cloud-based AI large-scale model is used to decrypt, recognize, semantically parse, and synthesize the received encrypted voice data to generate voice response data. This voice response data is then encrypted and transmitted to the in-vehicle TBox module, allowing the in-vehicle TBox module to decrypt and play the voice response data. Utilizing existing in-vehicle TBox hardware resources, AI voice interaction functionality can be achieved through software upgrades and cloud integration, fully leveraging existing hardware resources to improve the data security and responsiveness of intelligent cockpit voice interaction. To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0015] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 This is one of the structural schematic diagrams of a voice interaction system for a cloud-based AI large model vehicle TBox provided in an embodiment of this application; Figure 2 The second schematic diagram of the structure of a voice interaction system for a cloud-based AI large model vehicle TBox provided in this application embodiment; Figure 3 A flowchart illustrating a voice interaction method for a cloud-based AI large-scale model in-vehicle TBox provided in an embodiment of this application; Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0017] Icons: 100 - Cloud-based AI large-scale model in-vehicle TBox voice interaction system; 110 - In-vehicle TBox; 111 - Voice acquisition and optimization module; 112 - Data encryption and transmission module; 113 - Voiceprint recognition module; 114 - Command parsing and execution module; 120 - Cloud-based AI large-scale model; 121 - Speech-to-text module; 122 - Semantic understanding and command generation module; 123 - Speech synthesis module; 124 - Data storage and update module; 400 - Electronic device; 410 - Processor; 420 - Memory; 430 - Bus. Detailed Implementation To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of this application. Based on the embodiments of this application, every other embodiment obtained by those skilled in the art without inventive effort falls within the scope of protection of this application.
[0018] First, the applicable application scenarios of this application will be introduced. This application can be applied to the field of intelligent connected vehicle technology.
[0019] Research has revealed the following shortcomings in existing technologies: (1) Hardware redundancy – additional microphone arrays, voice processors, and communication modules are required, failing to reuse the basic capabilities of the widely deployed TBox in vehicles; (2) Unreliable links – voice data is relayed via mobile phones, resulting in multi-hop transmission, with end-to-end delays often exceeding 2.5 seconds, and is susceptible to the stability of Bluetooth / WiFi connections; (3) Security and compliance risks – the raw voice stream is uploaded without preprocessing and encryption on the device side, posing data risks such as voiceprint information leakage and interception of dialogue content. Therefore, improving the data security and responsiveness of voice interaction in smart cockpits has become a significant technical issue.
[0020] Based on this, the embodiments of this application provide a voice interaction system for an in-vehicle TBox with a cloud-based AI large model. By utilizing the existing in-vehicle TBox hardware resources, AI voice interaction functions can be realized through software upgrades and cloud integration. This fully utilizes existing hardware resources to improve the data security and responsiveness of voice interaction in the smart cockpit.
[0021] Please see Figure 1 , Figure 1 This is one of the structural schematic diagrams of a voice interaction system 100 for a cloud-based AI large-scale model vehicle T-box provided in an embodiment of this application. Figure 1 As shown in the figure, the voice interaction system 100 of the cloud AI big model vehicle Tbox provided in this application embodiment includes a cloud AI big model 120 and a communication connection between the cloud AI big model 120 and the vehicle Tbox 110.
[0022] Specifically, the in-vehicle Tbox 110 is used to collect the driver's voice commands, obtain encrypted voice data after noise reduction optimization and encryption processing, and upload it to the cloud-based AI big data model 120; wherein, the in-vehicle Tbox 110 includes a voice acquisition and optimization module 111, a data encryption and transmission module 112, a command parsing and execution module 114, and a voiceprint recognition module 113; the cloud-based AI big data model 120 is used to decrypt the received encrypted voice data, perform voice recognition, semantic parsing, and voice synthesis to generate voice response data, encrypt the voice response data, and transmit it to the in-vehicle Tbox 110 so that the in-vehicle Tbox 110 module can decrypt and play the voice response data.
[0023] In this application, the traditional Tbox on the vehicle side and the AI big model in the cloud establish a two-way communication connection through a 4G / 5G network. No modification to the traditional Tbox110 hardware is required; only a software upgrade is needed to add functions such as optimized voice acquisition, encrypted data transmission, cloud command reception, and voice playback control. This is compatible with most existing traditional Tbox110 models (those with basic network communication and audio / video acquisition and playback functions are sufficient). At the same time, it can be pre-installed as a Tbox110 function in low-configuration vehicles, helping low-configuration vehicles achieve low-cost AI voice interaction and improve the competitiveness of vehicle products.
[0024] For further details, please refer to Figure 2 , Figure 2 This is a second structural schematic diagram of a voice interaction system 100 for a cloud-based AI large-scale model in-vehicle T-box, provided as an embodiment of this application. (See diagram below.) Figure 2 As shown, the voice acquisition and optimization module 111 is used to: use a built-in noise reduction algorithm to filter environmental noise and suppress driving noise in the acquired driver's voice, and output the noise-reduced voice signal to the data encryption and transmission module 112. The data encryption and transmission module 112 is used to: encrypt the uplink voice data using an encryption protocol, and decrypt the downlink voice response data from the cloud. All voice data transmitted between the in-vehicle Tbox 110 module and the cloud AI large model 120 module is kept encrypted. The voiceprint recognition module 113 is used to: collect the driver's voiceprint information and compare it with the voiceprint features pre-stored in the cloud database to achieve identity verification. Voice commands from unauthorized personnel will not be responded to.
[0025] Here, the in-vehicle Tbox110 serves as the core execution unit of the system, retaining all its original basic functions (network communication, sound acquisition, and sound playback). Through software upgrades, three core new functions have been added: ① Voice acquisition optimization: This function reduces noise in the acquired driver's voice, minimizing environmental noise interference and improving voice recognition accuracy, thus resolving the issue of noise-induced command misinterpretation in existing technologies; ② Encrypted data transmission: This function encrypts the acquired voice data using an encryption protocol before uploading it to the cloud, while simultaneously decrypting the downlink voice response data from the cloud, ensuring data transmission security and mitigating the risk of sensitive voice data leakage; ③ Command parsing and execution: This function receives interactive commands returned by the cloud-based AI model 120, parses them, and plays the corresponding voice response, forming a complete interactive loop. In addition, a new voiceprint recognition function has been added. By collecting the driver's voiceprint information and comparing it with the cloud database, identity verification is achieved, preventing unauthorized personnel from operating related functions and further enhancing security. The vehicle-mounted Tbox110 establishes stable two-way communication with the cloud-based AI large model 120 via 4G / 5G networks. No hardware interface modifications are required; functional expansion can be achieved solely through software upgrades. It is compatible with most traditional Tbox110 models on the market that have basic network and audio / video capture and playback capabilities.
[0026] Furthermore, the cloud-based AI large model 120 includes: a speech-to-text module 121, used to perform automatic speech recognition on the decrypted speech data and output text commands by combining the noise-reduced and optimized speech signals; a semantic understanding and command generation module 122, used to perform intent recognition and contextual reasoning on the text commands based on the large language model 120, and generate structured text responses that meet user needs; and a speech synthesis module 123, used to convert the structured text responses into natural and fluent speech waveform data. The data storage and update module 124 is used to: securely store user voiceprint features, interaction history, and personalized preferences, and support algorithm updates for the cloud-based AI large model 120.
[0027] Here, the cloud-based AI large-scale model module serves as the core decision-making and analysis unit of the system. It is responsible for receiving encrypted voice data uploaded from the in-vehicle TBox module and performing core operations such as speech recognition, semantic understanding, command generation, and speech synthesis. Specific functions include: ① Automatic Speech-to-Text (ASR) function: accurately recognizing the decrypted voice data and converting voice commands into text information. Combined with noise-reduced and optimized voice data, it significantly improves recognition accuracy, accurately capturing driver commands even in ambient noise generated by the vehicle; ② Semantic understanding and command generation function: deeply analyzing text commands based on the AI large-scale model and generating tailored voice responses based on the driver's daily consultation needs. For example, if the driver asks "What's the weather like tomorrow?", the model will generate a corresponding response based on local weather data and simultaneously synthesize a voice feedback such as "Tomorrow will be sunny, with temperatures between 18-28℃, a light breeze, suitable for travel"; ③ Text-to-Speech (TTS) function: converting the generated text responses into natural and fluent voice data, encrypting it, and transmitting it downlink to the in-vehicle TBox module; ④ Data storage and update functions store vehicle owner voiceprint information, frequently used command preferences, and other data. At the same time, the AI model algorithm is updated regularly to optimize recognition accuracy and interaction response speed, adapting to the usage scenarios of different vehicle models; ⑤ Multi-terminal adaptation function optimizes communication protocols and supports the access of traditional TBoxes of different brands and models, eliminating the need to develop separate docking solutions for a single model and improving the system's versatility.
[0028] In a specific implementation, the process includes: ① Voice acquisition and optimization: The driver issues a voice command, and the in-vehicle TBox acquires the voice data through its built-in microphone. Simultaneously, a noise reduction algorithm is activated to filter vehicle noise and ambient noise, resulting in a clear voice signal. ② Data encryption and uploading: The TBox module uses an encryption protocol to encrypt the optimized voice data and uploads it to the cloud-based AI model via a 4G / 5G network. ③ Voice recognition and semantic analysis: The cloud-based AI model module decrypts the encrypted data, converts the voice into text commands using ASR (Automatic Speech Recognition) functionality, and then uses a semantic understanding algorithm to analyze the command intent and generate a corresponding text response. ④ Voice synthesis and downlink transmission: The cloud module converts the text response into voice data using TTS (Text-to-Speech) functionality, encrypts it, and transmits it downlink to the in-vehicle TBox. ⑤ Voice feedback: The in-vehicle TBox decrypts the downlink voice data, controls the built-in speaker to play the voice response, informing the driver of the inquiry result, and simultaneously uploads the completed response to the cloud, completing a full interaction. The entire process takes less than one second, meeting the real-time interaction requirements of the vehicle and solving the problem of high latency in existing technologies.
[0029] This application provides a voice interaction system for an in-vehicle TBox based on a cloud-based AI large-scale model. The voice interaction system includes a cloud-based AI large-scale model and a communication connection between the cloud-based TBox and the in-vehicle TBox. The in-vehicle TBox is used to collect the driver's voice commands, obtain encrypted voice data after noise reduction optimization and encryption processing, and upload it to the cloud-based AI large-scale model module. The in-vehicle TBox includes a voice acquisition and optimization module, a data encryption and transmission module, a command parsing and execution module, and a voiceprint recognition module. The cloud-based AI large-scale model is used to decrypt the received encrypted voice data, perform voice recognition, semantic parsing, and voice synthesis to generate voice response data. This voice response data is then encrypted and transmitted to the in-vehicle TBox module, allowing the in-vehicle TBox module to decrypt and play the voice response data. Utilizing the vehicle's existing in-vehicle TBox hardware resources, AI voice interaction functionality can be achieved through software upgrades and cloud integration, fully leveraging existing hardware resources to improve the data security and responsiveness of intelligent cockpit voice interaction.
[0030] Please see Figure 2 , Figure 3 This is a flowchart illustrating a voice interaction method for a cloud-based AI large-scale model in-vehicle TBox, as provided in an embodiment of this application. Figure 3 As shown in the illustration, this application provides a method for voice interaction, including: S301: Controls the in-vehicle TBox to collect the owner's voice commands, and after noise reduction, optimization and encryption processing, obtains encrypted voice data, which is then uploaded to the cloud AI large model module.
[0031] In this step, the in-vehicle Tbox performs noise reduction optimization and encryption processing on the received voice commands from the car owner to obtain encrypted voice data, which is then uploaded to the cloud-based AI big model module.
[0032] S302: Control the cloud-based AI model to decrypt, recognize, analyze semantics, and synthesize the received encrypted voice data to generate voice response data. After encrypting the voice response data, transmit it to the vehicle-mounted TBox module so that the vehicle-mounted TBox module can decrypt and play the voice response data.
[0033] In this step, the cloud-based AI model decrypts the received encrypted voice data, performs speech recognition, semantic analysis, and speech synthesis to generate voice response data. After encrypting the voice response data, it is transmitted to the in-vehicle TBox module.
[0034] In one possible implementation, the voice interaction method further includes: The vehicle-mounted TBox module is controlled to receive upgrade packages pushed by the cloud-based AI large model via OTA remote online upgrade; wherein, the upgrade package includes upgrade packages for the voice acquisition optimization module, the data encryption transmission module, the command parsing and execution module, and the voiceprint recognition module.
[0035] Here, the cloud-based AI model generates a corresponding OTA upgrade package based on the current software version information of the in-vehicle TBox. The upgrade package includes a voice acquisition optimization module, a data encryption transmission module, a command parsing and execution module, and a voiceprint recognition module. The cloud uses an encryption key to encrypt the OTA upgrade package, generating ciphertext. The cloud employs a high-strength hash algorithm to calculate the hash value of the upgrade package and combines it with asymmetric encryption to generate a digital signature, ensuring the integrity and legitimacy of the upgrade package's origin.
[0036] The in-vehicle TBox receives upgrade commands pushed by the cloud-based AI big data model module via a 4G / 5G network. After confirming that the vehicle status meets the upgrade conditions (such as the vehicle being parked and the battery voltage being normal), the in-vehicle TBox sends an upgrade request to the cloud-based AI big data model. The cloud-based AI big data model then downloads the encrypted upgrade package to the in-vehicle TBox via the 4G / 5G network based on the upgrade request. The in-vehicle TBox supports a resume download function; if a network interruption occurs during the download process, the download can resume from the point of interruption after the network is restored, avoiding repeated transmissions.
[0037] In one possible implementation, the voice interaction method further includes: The built-in noise reduction algorithm filters environmental noise and suppresses driving noise on the collected driver's voice, and outputs the noise-reduced voice signal to the data encryption transmission module.
[0038] In one possible implementation, the voice interaction method further includes: The uplink voice data is encrypted using an encryption protocol, and the downlink voice response data is decrypted from the cloud. All voice data transmitted between the vehicle-mounted TBox module and the cloud-based AI large model module remains encrypted.
[0039] In one possible implementation, the voice interaction method further includes: The system collects the vehicle owner's voiceprint information and compares it with the voiceprint features pre-stored in the cloud database to achieve identity verification. Voice commands from unauthorized personnel will not be responded to.
[0040] In one possible implementation, the voice interaction method further includes: The speech-to-text module automatically performs speech recognition on the decrypted speech data and outputs text commands by combining the noise-reduced and optimized speech signal. The control semantic understanding and instruction generation module performs intent recognition and contextual reasoning on the text instructions based on a large language model, and generates a structured text response that meets the user's needs; The speech synthesis module controls the conversion of structured text responses into natural and fluent speech waveform data.
[0041] In one possible implementation, the voice interaction method further includes: Securely store user voiceprint features, interaction history, and personalized preferences, and support algorithm updates for AI models.
[0042] This application provides a voice interaction method for an in-vehicle TBox based on a cloud-based AI large-scale model. The method includes: controlling the in-vehicle TBox to collect the driver's voice commands, obtaining encrypted voice data after noise reduction, optimization, and encryption, and uploading it to a cloud-based AI large-scale model module; controlling the cloud-based AI large-scale model to decrypt, recognize, analyze semantics, and synthesize the received encrypted voice data to generate voice response data; encrypting the voice response data and transmitting it to the in-vehicle TBox module, so that the in-vehicle TBox module can decrypt and play the voice response data. Utilizing the vehicle's existing in-vehicle TBox hardware resources, AI voice interaction functionality can be achieved through software upgrades and cloud integration, fully leveraging existing hardware resources to improve the data security and responsiveness of intelligent cockpit voice interaction.
[0043] Please see Figure 4 , Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 4 As shown, the electronic device 400 includes a processor 410, a memory 420, and a bus 430.
[0044] The memory 420 stores machine-readable instructions executable by the processor 410. When the electronic device 400 is running, the processor 410 communicates with the memory 420 via the bus 430. When the machine-readable instructions are executed by the processor 410, they can perform the operations described above. Figure 3 The steps of the voice interaction method for the cloud-based AI large model in the vehicle TBox in the method embodiment shown are described in detail in the method embodiment, and will not be repeated here.
[0045] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, can perform the above-described actions. Figure 3 The steps of the voice interaction method for the cloud-based AI large model in the vehicle TBox in the method embodiment shown are described in detail in the method embodiment, and will not be repeated here.
[0046] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0047] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the shown or discussed mutual couplings, direct couplings, or communication connections may be through some communication interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.
[0048] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0049] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0050] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0051] Finally, it should be noted that the above-described embodiments are merely specific implementations of this application, used to illustrate the technical solutions of this application, and not to limit them. The scope of protection of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this application. Such modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A voice interaction system for an in-vehicle T-Box based on a cloud-based AI large-scale model, characterized in that, The voice interaction system includes a cloud-based AI large model that is communicatively connected to the in-vehicle TBox, wherein; The in-vehicle TBox is used to collect the driver's voice commands, and after noise reduction, optimization and encryption processing, the encrypted voice data is obtained and uploaded to the cloud AI large model module; wherein, the in-vehicle TBox includes a voice collection and optimization module, a data encryption and transmission module, a command parsing and execution module and a voiceprint recognition module; The cloud-based AI model is used to decrypt, recognize, analyze, and synthesize the received encrypted voice data to generate voice response data. The voice response data is then encrypted and transmitted to the in-vehicle TBox module, so that the in-vehicle TBox module can decrypt and play the voice response data.
2. The voice interaction system according to claim 1, characterized in that, The voice acquisition optimization module is used for: The built-in noise reduction algorithm filters environmental noise and suppresses driving noise on the collected driver's voice, and outputs the noise-reduced voice signal to the data encryption transmission module.
3. The voice interaction system according to claim 1, characterized in that, The data encryption transmission module is used for: The uplink voice data is encrypted using an encryption protocol, and the downlink voice response data is decrypted from the cloud. All voice data transmitted between the vehicle-mounted TBox module and the cloud-based AI large model module remains encrypted.
4. The voice interaction system according to claim 1, characterized in that, The voiceprint recognition module is used for: The system collects the vehicle owner's voiceprint information and compares it with the voiceprint features pre-stored in the cloud database to achieve identity verification. Voice commands from unauthorized personnel will not be responded to.
5. The voice interaction system according to claim 1, characterized in that, The cloud-based AI large model includes: The speech-to-text module is used to perform automatic speech recognition on the decrypted speech data and output text commands by combining the noise-reduced and optimized speech signal. The semantic understanding and instruction generation module is used to perform intent recognition and contextual reasoning on the text instructions based on a large language model, and generate a structured text response that meets the user's needs. The speech synthesis module is used to convert structured text responses into natural and fluent speech waveform data.
6. The voice interaction system according to claim 1, characterized in that, The cloud-based AI large model also includes a data storage and update module, which is used for: Securely store user voiceprint features, interaction history, and personalized preferences, and support algorithm updates for AI models.
7. A voice interaction method for an in-vehicle TBox with a cloud-based AI large model, characterized in that, The voice interaction method is applied to the voice interaction system of the in-vehicle TBox of the cloud-based AI large model as described in any one of claims 1-6, and the voice interaction method includes: The vehicle-mounted TBox collects the owner's voice commands, which are then processed through noise reduction, optimization, and encryption to obtain encrypted voice data, which is then uploaded to the cloud-based AI large model module. The cloud-based AI model decrypts, performs speech recognition, semantic analysis, and speech synthesis on the received encrypted voice data to generate voice response data. The voice response data is then encrypted and transmitted to the in-vehicle TBox module, which decrypts and plays the voice response data.
8. The voice interaction method according to claim 7, characterized in that, The voice interaction method further includes: The vehicle-mounted TBox module is controlled to receive upgrade packages pushed by the cloud-based AI large model via OTA remote online upgrade; wherein, the upgrade package includes upgrade packages for the voice acquisition optimization module, the data encryption transmission module, the command parsing and execution module, and the voiceprint recognition module.
9. An electronic device, characterized in that, include: The device includes a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the memory via the bus. The machine-readable instructions are executed by the processor to perform the steps of the voice interaction method for the in-vehicle TBox of the cloud-based AI large model as described in claim 7 or 8.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the voice interaction method for the in-vehicle TBox of the cloud-based AI large model as described in claim 7 or 8.