Intelligent cabin interaction method and related device

CN122551784APending Publication Date: 2026-08-11IFLYTEK CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-13
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0003]基于上述现有技术的缺陷和不足,本申请提出一种智能座舱交互方法及相关装置,能够解决现有技术中的车辆智能座舱交互系统易发生误响应的问题

Benefits of technology

根据所述第一音频信息和所述第二音频信息的目标特征,确定所述第一音频信息所属的目标类型;其中,所述目标特征包括以下至少一项:语义特征、声学特征、预先为每一音频信息设置的座舱角色标签,所述座舱角色标签用于标识座舱内不同音区的语音输出用户。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122551784A_ABST
    Figure CN122551784A_ABST
Patent Text Reader

Abstract

This application proposes a smart cockpit interaction method and related apparatus, relating to the field of voice interaction technology. The smart cockpit interaction method includes: acquiring first audio information and context information currently collected within the smart cockpit; determining the target type of the first audio information based on the first audio information and context information; wherein the target type is one of an instruction type, a question type, an exclusive private chat type, and an invalid type; determining a processing strategy for the first audio information based on the target type of the first audio information, and executing the processing strategy; wherein the processing strategy is one of executing a vehicle control command corresponding to the first audio information, responding to a question corresponding to the first audio information, remaining silent on the first audio information, and refusing to recognize the first audio information. The technical solution provided by this application can solve the problem of erroneous responses easily occurring in existing vehicle smart cockpit interaction systems.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of voice interaction technology, and more specifically, to an intelligent cockpit interaction method and related devices. Background Technology

[0002] The current intelligent cockpit interaction system of vehicles generally uses voice wake-up technology to wake up a certain voice zone and lock that voice zone. That is, it only responds to the voice information of the user in that voice zone and responds to all the voice content of the user in that voice zone. Therefore, when the voice zone is locked, if the user in that voice zone speaks to other users, it will also be recognized and responded to by the system, thus causing a false response event. Summary of the Invention

[0003] Based on the defects and shortcomings of the prior art, this application proposes an intelligent cockpit interaction method and related device, which can solve the problem of false responses in the vehicle intelligent cockpit interaction system in the prior art.

[0004] According to a first aspect of the embodiments of this application, a smart cockpit interaction method is provided, the method comprising: Acquire first audio information and context information currently collected within the smart cockpit; wherein, the context information includes second audio information collected during the historical process through pickup devices in multiple sound zones within the smart cockpit; Based on the first audio information and the context information, the target type to which the first audio information belongs is determined; wherein, the target type is one of the following: instruction type, question type, exclusive private chat type, and invalid type; Based on the target type to which the first audio information belongs, a processing strategy for the first audio information is determined and the processing strategy is executed; wherein, the processing strategy is one of executing a vehicle control command corresponding to the first audio information, responding to a question corresponding to the first audio information, remaining silent on the first audio information, and refusing to recognize the first audio information.

[0005] In this application, the currently acquired first audio information is classified by combining contextual information including historically acquired second audio information. This classification categorizes the first audio information into instruction type, question type, exclusive private chat type, or invalid type. The processing strategy for the first audio information is then determined based on the classification results. By classifying the first audio information, the system can accurately distinguish between different situations such as human-computer instructions, casual conversations, private dialogues, and invalid sounds. This improves the accuracy of the processing strategy when determining the strategy based on the classification results. This not only reduces the occurrence of false responses but also allows for timely participation in user conversations, enhancing the naturalness and intelligence of voice interaction.

[0006] In some optional embodiments, determining the target type to which the first audio information belongs based on the first audio information and the context information includes: Based on the target features of the first audio information and the second audio information, the target type to which the first audio information belongs is determined; wherein, the target features include at least one of the following: semantic features, acoustic features, and a cockpit role label pre-set for each audio information, wherein the cockpit role label is used to identify the voice output user in different sound zones within the cockpit.

[0007] This application utilizes multi-dimensional information from both current and historical audio to perform a comprehensive analysis of the current audio type. For example, if the source tag of the first audio information is the driver role, and the semantics contain vehicle control command vocabulary, while the context information shows that the audio region was previously silent, the system can determine it as a command type with high confidence. By using multi-dimensional feature judgment, the classification confidence in different interaction scenarios can be effectively improved, and misjudgment events can be reduced.

[0008] In some optional embodiments, determining the target type of the first audio information based on the target features of the first audio information and the second audio information includes: If the first audio information contains words that match vehicle control commands and the audio energy is concentrated in the words, the first audio information is determined to be a command type. or, If the current conversation is determined to be a human-to-human conversation based on the context information, and if the first audio information is detected to include interrogative intonation features and the topic of the current conversation belongs to the target topic, then the first audio information is determined to be a question type. or, If the first audio information is detected to be non-human speech or a sudden human speech shock signal, the first audio information is determined to be invalid. or, If, based on the first audio information and the second audio information, it is determined that the first audio information does not belong to any of the instruction type, question type, or invalid type, then the first audio information is determined to belong to the exclusive private chat type.

[0009] In this application, by classifying the current audio information into four types, the system can accurately classify the current audio information into four different scenarios: human-machine instructions, dialogue participation, private dialogue, and invalid interference. This allows for the determination of targeted response strategies to overcome the shortcomings of related technologies where the wake-up sound zone responds indiscriminately to all audio.

[0010] In some alternative embodiments, after acquiring the target audio information, the method further includes: Based on the identification number of the audio pickup device that collects the target audio information, a corresponding cockpit role tag is added to the target audio information; wherein, the target audio information is either the first audio information or the second audio information.

[0011] In this application, by directly mapping the pickup device number to the cockpit location and automatically adding role tags, it is possible to assign accurate spatial attributes to each audio segment without relying on an additional recognition process.

[0012] In some alternative embodiments, executing the processing strategy includes: When the processing strategy includes outputting voice interaction information, the cockpit role label corresponding to the first audio information is determined; wherein, the cockpit role label is used to identify voice output users in different voice zones within the cockpit; Based on the cockpit role labels, determine the broadcasting equipment for the corresponding area; The voice interaction information is played through the broadcasting device.

[0013] In this application, the corresponding area's broadcasting equipment is selected for directional voice feedback based on the spatial label of the audio source, achieving precise sound delivery and effectively avoiding auditory interference to passengers in other areas of the cabin, thus helping to improve the acoustic comfort and user experience in the cabin.

[0014] In some optional embodiments, the context information also includes vehicle status information acquired during the acquisition of the first audio information; When the first audio information belongs to an instruction type, determining the processing strategy for the first audio information based on the target type to which the first audio information belongs includes: Based on the vehicle control command corresponding to the first audio information and the vehicle status information, a processing strategy for the first audio information is determined.

[0015] In this application, when the first audio information is a vehicle control command, the introduction of vehicle status provides additional constraints and optimization conditions for vehicle control decisions. This enables the system to dynamically adjust the degree and method of control execution by combining real-time vehicle status information such as vehicle speed, window status, and air conditioning status. While responding to commands in a timely manner, it can also improve driving safety and control rationality.

[0016] In some optional embodiments, the steps of determining the target type to which the first audio information belongs and determining the processing strategy for the first audio information include: The first audio information, the context information, and the set of vehicle control commands are input into a pre-trained multimodal large model. The multimodal large model determines the target type to which the first audio information belongs and the processing strategy for the first audio information.

[0017] In this application, by utilizing the end-to-end reasoning capabilities of a multimodal large model, the system can automatically and simultaneously complete the identification of audio types and the generation of processing strategies, deeply integrating acoustic, semantic, spatial, and state information, thereby improving the intelligence level, generalization ability, and overall interactive fluency of decision-making.

[0018] According to a second aspect of the embodiments of this application, a smart cockpit interaction device is provided, the device comprising: The acquisition module is used to acquire first audio information and context information currently collected within the smart cockpit; wherein, the context information includes second audio information collected by pickup devices in multiple sound zones within the smart cockpit during the historical process; The classification module is used to determine the target type to which the first audio information belongs based on the first audio information and the context information; wherein the target type is one of the following: instruction type, question type, exclusive private chat type, and invalid type; The processing module is configured to determine a processing strategy for the first audio information based on the target type to which the first audio information belongs, and execute the processing strategy; wherein the processing strategy is one of executing a vehicle control command corresponding to the first audio information, responding to a question corresponding to the first audio information, keeping the first audio information silent, and refusing to recognize the first audio information.

[0019] According to a third aspect of the embodiments of this application, an electronic device is provided, including: a memory and a processor; The memory is connected to the processor and is used to store programs; The processor is used to implement the smart cockpit interaction method as described in the first aspect by running the program in the memory.

[0020] According to a fourth aspect of the embodiments of this application, a storage medium is provided, on which a computer program is stored, and when the computer program is run by a processor, it implements the intelligent cockpit interaction method as described in the first aspect.

[0021] According to a fifth aspect of the embodiments of this application, a computer program product or computer program is provided, the computer program product including the computer program, wherein when a processor executes the computer program, it implements the steps in the intelligent cockpit interaction method as described in the first aspect. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0023] Figure 1 This is a flowchart illustrating an intelligent cockpit interaction method provided in an embodiment of this application.

[0024] Figure 2 A block diagram of an intelligent cockpit interaction device provided in an embodiment of this application.

[0025] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0026] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0027] With the rapid development of vehicle intelligence and connectivity technologies, the smart cockpit has become the core hub for human-vehicle interaction. Traditional physical buttons and touchscreen operations, due to issues such as visual distraction and inconvenience, can no longer meet users' dual needs for convenience and safety. Therefore, voice interaction technology, with its advantages of being natural, direct, and requiring no eye movement, is gradually becoming the mainstream interaction method in smart cockpits. Furthermore, to further enhance the driving experience, some current models deploy multi-zone independent voice pickup technology. By arranging microphone arrays in different seating areas within the vehicle, such as the driver's seat, front passenger seat, left rear seat, and right rear seat, and combining sound source localization and beamforming algorithms, it achieves accurate differentiation and directional response to voice commands from users in different seats.

[0028] Current intelligent cockpit voice interaction systems mainly adopt directional wake-up and voice zone locking control strategies, and their workflow can be described as follows: The system monitors wake-up words in real time or at set intervals using a microphone array. When a user in a specific vocal range (such as the driver's seat) utters a preset wake-up word (e.g., "Hello, XX"), the system immediately wakes up and locks onto that vocal range, picking up and processing the speech signal from that locked range while filtering out or suppressing speech signals from other unlocked vocal ranges (such as the passenger seat or rear seats). The picked-up speech signal can be converted into text using Automatic Speech Recognition (ASR) technology, and then fed into a language model for processing. This language model will respond indiscriminately to all speech signals within that vocal range.

[0029] When the locked voice zone is in the wake-up state, regardless of whether the voice content in that voice zone is a command to the vehicle system (i.e., human-machine dialogue), a chat with fellow passengers (human-to-human dialogue), or meaningless self-talk, the system will default to treating it as a valid command to be processed, perform voice recognition, and perform intent parsing and action response, which can easily lead to false response events.

[0030] To address this, this application provides a solution that analyzes and categorizes currently collected audio information into command types, question types, exclusive private chat types, or invalid types. Based on the categorization results, a processing strategy is determined for the audio information to reduce false responses and avoid interfering with normal communication between users in the vehicle. Furthermore, this solution allows the voice interaction system to participate in user conversations in a timely manner, enhancing the naturalness and intelligence of voice interaction.

[0031] Exemplary methods This application provides an intelligent cockpit interaction method applied to an electronic device. The electronic device can be a terminal device in a vehicle, such as an in-vehicle terminal, main controller, or cockpit domain controller; it can also be a server, such as a cloud server. When the electronic device is a server, information collected by the vehicle (such as audio information) can be sent to the server, which then processes the uploaded information and returns the result to the vehicle.

[0032] The vehicle's cabin is equipped with multiple audio pickup devices, which can be microphone arrays, and are positioned in different locations within the cabin, such as the driver's seat, front passenger seat, left rear seat, and right rear seat. Each pickup device corresponds to an audio acquisition area, referred to as a sound zone. Each pickup device is used to collect audio signals within its corresponding sound zone in real time.

[0033] The method is described in detail below through some embodiments. The following embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0034] like Figure 1 As shown, the intelligent cockpit interaction method may include steps 101 to 103, as described below.

[0035] Step 101: Obtain the first audio information and context information currently collected in the smart cockpit.

[0036] The first audio information mentioned here is the latest audio information collected, that is, the audio information to be processed.

[0037] The context information mentioned here refers to information used to assist in determining the type of the first audio information, and may include, but is not limited to, the second audio information collected by sound pickup devices in multiple sound zones within the smart cockpit during a historical process, specifically audio information from a preset historical round. One audio message constitutes one round; for example, if user A outputs a sentence and then user B outputs a sentence, these two sentences constitute two rounds. Furthermore, the context information may also include text information corresponding to the second audio information, that is, text information obtained through speech recognition of the second audio information.

[0038] Optionally, the audio information in the embodiments of this application can be the original audio signal or a processed high-dimensional audio feature vector.

[0039] In the case where the audio information in this embodiment is a high-dimensional audio feature vector, the high-dimensional audio feature vector can be obtained through the following process: After acquiring the raw audio stream through a sound pickup device, basic echo cancellation and bypass voice cancellation can be performed. The noise-suppressed raw audio stream is then input into a low-level audio encoder, which transforms it into high-dimensional audio feature vectors (Audio Embeddings) containing features such as volume, semantics, speech rate, tone, emotion, timbre, audio energy distribution, and environmental acoustics. These high-dimensional audio feature vectors are used as the audio information required to implement the intelligent cockpit interaction method provided in this application embodiment.

[0040] Step 102: Determine the target type to which the first audio information belongs based on the first audio information and the context information.

[0041] The target type mentioned here refers to the classification result of the interactive intent or sound source nature carried by the first audio information. The target type can be one of the following: instruction type, question type, exclusive private chat type, and invalid type.

[0042] The instruction type refers to information in the first audio message that matches the vehicle control instruction, and its main purpose is human-machine interaction. For example, if the first audio message is "It's very sunny and hot today, please turn on the air conditioner," and the phrase "It's very hot, please turn on the air conditioner" matches the vehicle control instruction "Turn on the air conditioner's cooling mode," then the type of the first audio message can be an instruction type.

[0043] The "problem type" refers to the inclusion of problem-related information in the first audio message. For example, if the first audio message is "How much longer until we reach XX scenic spot?", which includes questions unrelated to vehicle control, then the type of the first audio message can be "problem type".

[0044] Exclusive private chat refers to voice input from the user that does not require a response from the intelligent cockpit voice interaction system. For example, the first audio information might be the content of a phone conversation being conducted by the user, or a private conversation between users.

[0045] Invalid types refer to audio information that does not require speech recognition, such as non-human speech (e.g., musical rhythms, horn sounds) or sudden human speech shock signals (e.g., coughing sounds).

[0046] In this embodiment, the target type of the current audio information can be determined by comprehensively considering the acoustic features and semantic content of the current audio information itself, as well as the historical dialogue state reflected by the context information.

[0047] Step 103: Determine the processing strategy for the first audio information based on the target type to which the first audio information belongs, and execute the processing strategy.

[0048] The processing strategy described here is one of the following: executing the vehicle control command corresponding to the first audio information, responding to the question corresponding to the first audio information, remaining silent on the first audio information, and refusing to recognize the first audio information.

[0049] For example, when the first audio information is an instruction type, the processing strategy may include executing the vehicle control instruction corresponding to the first audio information, such as adjusting the windows or turning on the air conditioning, and generating a corresponding control message to send to the corresponding domain controller to execute the control instruction.

[0050] For example, when the first audio information belongs to a question type, the processing strategy may include generating a response text and / or voice for the question and broadcasting it through a display device and / or a broadcasting device. The response content may be generated based on a model or retrieved from a knowledge base.

[0051] For example, when the first audio message is an exclusive private chat, the processing strategy can be to remain silent, that is, to continue listening but interrupt text generation, without performing any voice broadcast or signaling. In other words, semantic analysis (such as speech recognition) will be performed on the currently collected audio information, but no response will be given to it, thereby avoiding interference with the user.

[0052] For example, when the first audio information is of an invalid type, the processing strategy may include refusing to recognize it, that is, directly truncating the reasoning process and not occupying the computing resources for subsequent language generation. In other words, the first audio information is not subject to speech recognition, and there is no response to it.

[0053] The intelligent cockpit interaction method provided in this application analyzes the currently collected audio information in conjunction with contextual information to classify it into instruction type, question type, exclusive private chat type, or invalid type. By classifying the first audio information, the system can accurately distinguish different application scenarios such as human-machine instructions, casual conversations, private dialogues, and invalid sounds. This improves the accuracy of the processing strategy when determining the processing strategy based on the classification results. This not only reduces the occurrence of false responses but also allows for timely participation in dialogues between users, enhancing the naturalness and intelligence of voice interaction. Furthermore, this application embodiment directly utilizes audio information rather than text information for the classification task, which preserves more acoustic features and helps improve the accuracy of the classification results.

[0054] In some optional embodiments, step 102: determining the target type to which the first audio information belongs based on the first audio information and context information may include: Based on the target characteristics of the first audio information and the second audio information, determine the target type to which the first audio information belongs.

[0055] The target features mentioned herein may include, but are not limited to, at least one of the following: semantic features, acoustic features, and cockpit role labels pre-set for each audio information.

[0056] Semantic features refer to features that reflect the meaning of the language content of audio information, that is, what the user said. These features may include, but are not limited to, at least one of the following: text content obtained from speech recognition, word vectors, intent, emotion, and other features.

[0057] Acoustic features refer to the characteristics of audio signals at the physical attribute level, that is, how the user speaks, which may include at least one of the following: volume, tone, speech rate, energy distribution, fundamental frequency, spectral envelope, etc.

[0058] The cockpit role label is used to identify voice output users in different audio zones within the cockpit. Cockpit roles can include, but are not limited to: driver role (Role_Driver), co-pilot role (Role_CoPilot), left rear seat role (Role_RearLeft), right rear seat role (Role_RearRight), etc. The specific roles can be set according to the division of audio zones, with each audio zone corresponding to one cockpit role.

[0059] By directly mapping the pickup device number to the cockpit location and automatically adding role tags, accurate spatial attributes can be assigned to each audio segment without relying on an additional recognition process.

[0060] Optionally, a pre-established correspondence between the serial numbers of the pickup devices in each audio zone and the cockpit role labels can be established. For example, a correspondence can be established between the serial numbers of the pickup devices in the driver's seat zone and the driver's role, the serial numbers of the pickup devices in the passenger's seat zone and the passenger's role, the serial numbers of the pickup devices in the left rear seat zone and the left rear seat role, and the serial numbers of the pickup devices in the right rear seat zone and the right rear seat role. In practical applications, based on this pre-established correspondence and the serial numbers of the pickup devices collecting the target audio information, corresponding cockpit role labels can be added to the target audio information, where the target audio information is either the first audio information or the second audio information.

[0061] In this embodiment, multi-dimensional information from both current and historical audio can be comprehensively utilized for analysis of the current audio type. For example, if the source tag of the first audio information is the driver role, and the semantics contain vehicle control command vocabulary, while the context information shows that the audio region was previously silent, the system can determine it as a command type with high confidence. Through multi-dimensional feature judgment, the classification confidence in different interaction scenarios can be effectively improved, and misjudgment events can be reduced.

[0062] Optionally, if the first audio information contains a word that matches a vehicle control command and the audio energy is concentrated on that word, the first audio information is determined to be a command type.

[0063] For example, if, based on contextual information, the driver character speaks with high energy and distinct emphasis after a period of silence, such as, "It's so stuffy in the car, let's open the window to let some fresh air in," and by combining semantic and acoustic judgments to determine that the audio energy is concentrated in the instructional words, then the audio information can be identified as an instruction type.

[0064] Optionally, if the current conversation is determined to be a human-to-human conversation based on contextual information, and if the first audio information is detected to include interrogative intonation features and the topic of the current conversation belongs to the target topic, then the first audio information can be determined to be a question type.

[0065] The target topics mentioned here may include, but are not limited to, topics that are likely to occur in a vehicle environment, such as navigation topics, food topics, and tourist attraction topics. The specific topics can be set according to actual needs.

[0066] For example, when the context information shows that the passengers in the left and right rear seats are speaking alternately over a period of time with a coherent topic that falls within the target topic, and the current audio information exhibits a rising interrogative tone and involves knowledge exploration, it can be determined that the current audio information belongs to the question type. Specifically, for example, two passengers in the back are discussing a weekend outing, and the passenger in the left rear seat asks, "Is it possible to camp in that forest park over there?" Acoustic analysis reveals that this tone is a typical rising interrogative tone, with a slow speaking speed and exploratory characteristics. Furthermore, the context information indicates a high degree of interpersonal interaction and that the current topic pertains to tourist attractions. In this case, it can be determined that the topic has high knowledge-addressing value, and machine intervention will not seem abrupt; therefore, the audio information can be identified as a question type.

[0067] Optionally, if the first audio information is detected to be non-human speech or a sudden human speech shock signal, the first audio information is determined to be invalid.

[0068] For example, if music is playing in the car and the driver coughs continuously due to a cold, or if a loud ambulance siren is heard outside the car window, and the acoustic feature vector of a sudden human impact signal (coughing) or a periodic signal of non-human language (siren) is detected, and the distance between the acoustic feature vector and the feature space of natural language exceeds a preset distance, then the audio information is considered invalid and no speech recognition or response is required.

[0069] Optionally, if it is determined based on the first audio information and context information that the first audio information does not belong to any of the instruction type, question type, and invalid type, then the first audio information can be determined to belong to the exclusive private chat type.

[0070] Since voice interaction systems are often required to remain silent, such as when a driver is taking a private call or talking quietly with a passenger, it is difficult to fully cover such situations by setting rules. Therefore, the type of audio information can be determined by elimination, thereby improving the accuracy of the judgment.

[0071] This application embodiment divides the current audio information into four types, enabling the system to accurately classify the current audio information into four different scenarios: human-machine instructions, dialogue participation, private dialogue, and invalid interference. This allows for the determination of targeted response strategies, overcoming the defect in related technologies where the wake-up sound zone responds indiscriminately to all audio.

[0072] In some alternative embodiments, the context information may also include vehicle status information acquired during the acquisition of the first audio information.

[0073] Accordingly, when the first audio information belongs to an instruction type, the aforementioned step "determine the processing strategy for the first audio information based on the target type to which the first audio information belongs" may include: Based on the vehicle control command and vehicle status information corresponding to the first audio information, a processing strategy for the first audio information is determined.

[0074] In this embodiment of the application, when it is determined that the first audio information belongs to the instruction type, a specific processing strategy can be determined in combination with the current vehicle status.

[0075] For example, if the control intent of the first audio message is to open the car window, and the current vehicle status information shows a current speed of 100 km / h, based on a preset safety strategy, it is determined that it is not suitable to fully open the window at this time. Therefore, the processing strategy based on the current vehicle speed can be determined as: open the window by 10%, and generate voice feedback, such as "The current vehicle speed is relatively high, and the window has been slightly opened for ventilation." The processing strategies corresponding to different vehicle statuses can be pre-set or predicted by a pre-trained model.

[0076] In this embodiment, when the first audio information is a vehicle control command, the introduction of vehicle status provides additional constraints and optimization conditions for vehicle control decisions. This enables the system to dynamically adjust the degree and method of control execution by combining real-time vehicle status information such as vehicle speed, window status, and air conditioning status. While responding to commands in a timely manner, it can also improve driving safety and control rationality.

[0077] In some alternative embodiments, when the processing strategy includes outputting voice interaction information, a precise targeted broadcast method can be adopted to avoid causing unnecessary sound interference to other passengers in the cabin, as described below.

[0078] When the processing strategy includes outputting voice interaction information, the "execute processing strategy" step 103 may include steps A1 to A3, as described below: Step A1: Determine the cockpit role label corresponding to the first audio information.

[0079] Step A2: Determine the broadcasting equipment for the corresponding area based on the cockpit role label.

[0080] Step A3: Play voice interaction information through the broadcasting equipment.

[0081] In this embodiment, each audio zone is equipped with both a pickup device and a broadcasting device. When the processing strategy includes outputting voice interaction information, this embodiment can determine the broadcasting device for the corresponding audio zone based on the cockpit role tag corresponding to the first audio information, and have that broadcasting device play the voice interaction information corresponding to the first audio information.

[0082] For example, if the first audio information comes from the driver, the corresponding broadcasting device can select the directional speaker inside the headrest of that seat. If the first audio information comes from the left rear seat user, the corresponding broadcasting device can select the directional speaker inside the left door.

[0083] If the first audio message is a question and requires the participation of the entire vehicle, then all the vehicle's speakers can be used to participate in casual conversation.

[0084] This embodiment selects the broadcast area based on the location of the audio source, which can accurately deliver the sound to the target user. For example, only the driver hears the vehicle control confirmation information, while other passengers continue to be in a quiet rest or conversation environment, improving the acoustic comfort of the cabin and the overall user experience.

[0085] In some alternative embodiments, the intelligent cockpit interaction method provided in this application embodiment can be implemented through a multimodal large model, as described below.

[0086] The steps of determining the target type of the first audio information and determining the processing strategy for the first audio information may include: The first audio information, context information, and vehicle control command set are input into a pre-trained multimodal large model. The multimodal large model determines the target type of the first audio information and the processing strategy for the first audio information.

[0087] A multimodal large model (also known as a multimodal large language model) refers to an artificial intelligence model capable of simultaneously processing and understanding multiple different types of information (i.e., modalities). This application embodiment utilizes this feature of a multimodal large model, inputting first audio information, contextual information, and a set of vehicle control commands into a pre-trained multimodal large model. The model analyzes this multimodal information, performing end-to-end inference by identifying semantic features, acoustic features, spatial sources, and the current vehicle state in the audio, and outputting specific processing strategies to improve the accuracy of the first audio information classification and the accuracy of the processing strategies. The set of vehicle control commands mentioned here is pre-set and used to determine whether the first audio information belongs to a vehicle control command, and if so, which specific vehicle control command it belongs to.

[0088] During training, this multimodal large model can receive a mixed sequence consisting of a set of vehicle control commands, vehicle status, multi-role audio interaction stack, and current original audio input. Through reinforcement learning and other methods, the model is guided to learn the features of different types of audio information and the corresponding processing strategies, thereby enabling it to respond to command control messages in a timely manner, respond to multi-person topic chat messages in a timely manner, and suppress false responses to private conversations and noise.

[0089] In some optional embodiments, the audio vector (i.e., high-dimensional audio feature vector) extracted from the audio information collected by the sound pickup device and its associated attributes (such as vehicle status) can be encapsulated into a multimodal interactive event object and pushed onto the global sliding window interaction stack maintained by the system to obtain context information. The size of this sliding window can be preset, and its size is the preset number of dialogue interaction rounds.

[0090] Suppose that at a certain moment, the context information stores the chat history of the two passengers in the back row, and its underlying data structure can be as follows: [ { "Frame_ID": "F_001", "Timestamp": "2026-4-22 09:12:05.100", "Location_Role (Location Role, i.e., cockpit role)": "Role_RearLeft (Rear Left Seat Role)", "Audio_Embeddings": "[Vector Matrix 1x1024... Contains timbre A, normal speaking speed, and interrogative intonation]", "Text_Transcript (recognized text)": "What are we going to eat later?" "Vehicle_State": { "Speed": 45, "Nav_Status": "Idle"} }, { "Frame_ID": "F_002", "Timestamp": "2026-4-22 09:12:07.300", "Location_Role": "Role_RearRight (Right rear seat character)", "Audio_Embeddings": "[Vector Matrix 1x1024... Contains a B tone with a slightly excited feel]", "Text_Transcript": "I heard that a new Chongqing hot pot restaurant opened in the mall ahead, and it's very good." "Vehicle_State": { "Speed": 48, "Nav_Status": "Idle"} } ] Subsequently, the information in the stack can rely on the self-attention mechanism of the multimodal large model to automatically align the encoding dimensions during inference and calculate the pointing relationships between historical audio segments, such as who is speaking to whom and the coherence of the topic, so as to combine contextual information to make better inferences and judgments about the first audio information.

[0091] Optionally, the output actions, query results, or audio broadcasts corresponding to the current processing strategy can be stored in the interaction stack as an AI role (Role_AI), and the sliding window can be updated to complete the system loop of a real multi-party dialogue scenario, retaining complete contextual information so that more accurate prediction and reasoning can be performed based on the contextual information.

[0092] In summary, the intelligent cockpit interaction method provided in this application, by combining contextual information including historically collected second audio information, classifies the currently collected first audio information into instruction type, question type, exclusive private chat type, or invalid type, and then determines the processing strategy for the first audio information based on the classification results. By classifying the first audio information, the system can accurately distinguish different application scenarios such as human-machine instructions, casual conversations, private dialogues, and invalid sounds. This improves the accuracy of the processing strategy when determining the processing strategy based on the classification results, reduces the occurrence of false response events, and allows for timely participation in dialogues between users, enhancing the naturalness and intelligence of voice interaction.

[0093] Exemplary device Accordingly, this application also provides an intelligent cockpit interaction device applied to a first electronic device. This electronic device can be a terminal device in a vehicle, such as an in-vehicle terminal, main controller, cockpit domain controller, etc., or it can be a server, such as a cloud server. When the electronic device is a server, information collected by the vehicle (such as audio information) can be sent to the server, which then processes the information uploaded by the vehicle and returns the processing result to the vehicle.

[0094] The vehicle's cabin is equipped with multiple audio pickup devices, which can be microphone arrays, and are positioned in different locations within the cabin, such as the driver's seat, front passenger seat, left rear seat, and right rear seat. Each pickup device corresponds to an audio acquisition area, referred to as a sound zone. Each pickup device is used to collect audio signals within its corresponding sound zone in real time.

[0095] like Figure 2 As shown, the device may include: The acquisition module 201 is used to acquire the first audio information and context information currently collected in the smart cockpit.

[0096] The context information includes second audio information collected during the historical process by pickup devices in multiple sound zones of the smart cockpit.

[0097] The classification module 202 is used to determine the target type to which the first audio information belongs based on the first audio information and the context information.

[0098] The target type is one of the following: instruction type, question type, exclusive private chat type, and invalid type.

[0099] The processing module 203 is used to determine a processing strategy for the first audio information based on the target type to which the first audio information belongs, and to execute the processing strategy.

[0100] The processing strategy includes executing a vehicle control command corresponding to the first audio information, responding to a question corresponding to the first audio information, remaining silent on the first audio information, and refusing to recognize the first audio information.

[0101] In some alternative embodiments, the classification module 202 may include: The classification unit is used to determine the target type to which the first audio information belongs based on the target features of the first audio information and the second audio information.

[0102] The target features include at least one of the following: semantic features, acoustic features, and a cockpit role label pre-set for each audio information, wherein the cockpit role label is used to identify the voice output user in different audio zones within the cockpit.

[0103] In some alternative embodiments, the classification unit may be specifically used for: If the first audio information contains words that match vehicle control commands and the audio energy is concentrated in the words, the first audio information is determined to be a command type. or, If the current conversation is determined to be a human-to-human conversation based on the context information, and if the first audio information is detected to include interrogative intonation features and the topic of the current conversation belongs to the target topic, then the first audio information is determined to be a question type. or, If the first audio information is detected to be non-human speech or a sudden human speech shock signal, the first audio information is determined to be invalid. or, If, based on the first audio information and the second audio information, it is determined that the first audio information does not belong to any of the instruction type, question type, or invalid type, then the first audio information is determined to belong to the exclusive private chat type.

[0104] In some alternative embodiments, after acquiring the target audio information, the device may further include: Based on the identification number of the audio pickup device that collects the target audio information, a corresponding cockpit role tag is added to the target audio information; wherein, the target audio information is either the first audio information or the second audio information.

[0105] In some alternative embodiments, the processing module 203 may include: The first determining unit is configured to determine the cockpit role label corresponding to the first audio information when the processing strategy includes outputting voice interaction information.

[0106] The cockpit role labels are used to identify voice output users in different audio zones within the cockpit.

[0107] The second determining unit is used to determine the broadcasting equipment in the corresponding area based on the cockpit role label.

[0108] A playback control unit is used to play the voice interaction information through the broadcasting device.

[0109] In some alternative embodiments, the context information also includes vehicle status information acquired during the acquisition of the first audio information.

[0110] The processing module 203 may include: The third determining unit is used to determine a processing strategy for the first audio information based on the vehicle control command corresponding to the first audio information and the vehicle status information when the first audio information belongs to the command type.

[0111] In some optional embodiments, the steps of determining the target type to which the first audio information belongs and determining the processing strategy for the first audio information may include: The first audio information, the context information, and the set of vehicle control commands are input into a pre-trained multimodal large model. The multimodal large model is used to determine the target type of the first audio information and the processing strategy for the first audio information.

[0112] The intelligent cockpit interaction device provided in this embodiment belongs to the same application concept as the intelligent cockpit interaction method provided in the above embodiments of this application. It can execute the intelligent cockpit interaction method provided in any of the above embodiments of this application and has the corresponding functional modules and beneficial effects of the execution method. Technical details not described in detail in this embodiment can be found in the specific processing content of the intelligent cockpit interaction method provided in the above embodiments of this application, and will not be repeated here.

[0113] It should be understood that the modules in the above-described intelligent cockpit interaction device can be implemented in the form of a processor calling software. For example, the device includes a processor connected to a memory containing instructions. The processor calls the instructions stored in the memory to implement any of the above methods or to realize the functions of each unit of the device. The processor can be a general-purpose processor, such as a CPU or microprocessor, and the memory can be internal or external to the device. Alternatively, the units in the device can be implemented in the form of hardware circuits. By designing the hardware circuits, some or all of the unit functions can be realized. The hardware circuit can be understood as one or more processors. For example, in one implementation, the hardware circuit is an ASIC, and the functions of some or all of the above units are realized by designing the logical relationships between the components within the circuit. In another implementation, the hardware circuit can be implemented using a PLD, such as an FPGA, which can include a large number of logic gates. The connection relationships between the logic gates are configured through a configuration file, thereby realizing the functions of some or all of the above units. All units of the above device can be implemented entirely through processor calling software, entirely through hardware circuits, or partially through processor calling software with the remaining parts implemented through hardware circuits.

[0114] In this application embodiment, a processor is a circuit with signal processing capabilities. In one implementation, the processor can be a circuit with instruction reading and execution capabilities, such as a CPU, microprocessor, GPU, or DSP. In another implementation, the processor can implement certain functions through the logical relationships of hardware circuits. These logical relationships are fixed or reconfigurable. For example, the processor may be a hardware circuit implemented as an ASIC or PLD, such as an FPGA. In a reconfigurable hardware circuit, the process of the processor loading a configuration document and configuring the hardware circuit can be understood as the processor loading instructions to implement the functions of some or all of the above units. Furthermore, it can also be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC, such as an NPU, TPU, or DPU.

[0115] As can be seen, each unit in the above device can be one or more processors (or processing circuits) configured to implement the above methods, such as: CPU, GPU, NPU, TPU, DPU, microprocessor, DSP, ASIC, FPGA, or a combination of at least two of these processor forms.

[0116] Furthermore, the units in the above devices can be integrated in whole or in part, or they can be implemented independently. In one implementation, these units are integrated together and implemented in the form of a System-on-Chip (SoC). The SoC may include at least one processor for implementing any of the above methods or implementing the functions of the units in the device. The at least one processor may be of different types, such as CPU and FPGA, CPU and artificial intelligence processor, CPU and GPU, etc.

[0117] Exemplary electronic devices This application also provides an electronic device, such as... Figure 3 As shown, the electronic device includes a memory 300 and a processor 310.

[0118] The memory 300 is connected to the processor 310 and is used to store programs.

[0119] The processor 310 is used to implement the intelligent cockpit interaction method in the above embodiments by running the program stored in the memory 300.

[0120] Specifically, the aforementioned electronic device may also include: a communication interface 320, an input device 330, an output device 340, and a bus 350.

[0121] The processor 310, memory 300, communication interface 320, input device 330, and output device 340 are interconnected via a bus. Among them: Bus 350 may include a pathway for transmitting information between various components of a computer system.

[0122] The processor 310 can be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, etc., or an application-specific integrated circuit (ASIC), or one or more integrated circuits used to control the execution of the program of the present invention. It can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0123] Processor 310 may include a main processor, as well as a baseband chip, modem, etc.

[0124] The memory 300 stores a program that executes the technical solution of this invention, and may also store an operating system and other key business functions. Specifically, the program may include program code, which includes computer operation instructions. More specifically, the memory 300 may include read-only memory (ROM), other types of static storage devices capable of storing static information and instructions, random access memory (RAM), other types of dynamic storage devices capable of storing information and instructions, disk storage, flash memory, etc.

[0125] Input device 330 may include a device for receiving user input data and information, such as a keyboard, mouse, camera, scanner, light pen, voice input device, touch screen, pedometer, or gravity sensor.

[0126] Output device 340 may include devices that allow information to be output to a user, such as a display screen, printer, speaker, etc.

[0127] The communication interface 320 may include a device that uses any transceiver to communicate with other devices or communication networks, such as Ethernet, Radio Access Network (RAN), Wireless Local Area Network (WLAN), etc.

[0128] The processor 310 executes the program stored in the memory 300 and calls other devices, which can be used to implement the various steps in the smart cockpit interaction method provided in the above embodiments of this application.

[0129] Exemplary computer program products and storage media In addition to the methods and devices described above, embodiments of this application may also be computer program products, which include computer program instructions that, when executed by a processor, cause the processor to perform the steps in the intelligent cockpit interaction method described in the embodiments of this application.

[0130] The aforementioned computer program product can be implemented through hardware, software, or a combination thereof. In one optional embodiment, the computer program product is specifically embodied in a computer storage medium; in another optional embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.

[0131] The computer program product can be written in any combination of one or more programming languages ​​to perform the operations of the embodiments of this application. The programming languages ​​include object-oriented programming languages ​​such as Java and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0132] Furthermore, embodiments of this application may also be storage media storing a computer program, which is executed by a processor using the steps of the intelligent cockpit interaction method described in the embodiments of this application.

[0133] In addition, embodiments of this application may also be chips, which include processors and data interfaces. The processor reads instructions stored in the memory through the data interface to execute the steps in the smart cockpit interaction method described in the embodiments of this application.

[0134] For the foregoing method embodiments, in order to simplify the description, they are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, because according to this application, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0135] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For apparatus embodiments, since they are basically similar to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.

[0136] The steps in the methods of the various embodiments of this application can be adjusted, merged, or deleted in order according to actual needs, and the technical features described in each embodiment can be replaced or combined.

[0137] The modules and sub-modules in the devices and terminals in the various embodiments of this application can be merged, divided, and deleted according to actual needs.

[0138] It should be understood that the disclosed terminals, devices, and methods can be implemented in other ways, given the several embodiments provided in this application. For example, the terminal embodiments described above are merely illustrative. For instance, the division of modules or sub-modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple sub-modules or modules may be combined or integrated into another module, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or modules, and may be electrical, mechanical, or other forms.

[0139] The modules or submodules described as separate components may or may not be physically separate. The components that constitute a module or submodule may or may not be physical modules or submodules; that is, they may be located in one place or distributed across multiple network modules or submodules. Some or all of the modules or submodules can be selected to achieve the purpose of this embodiment's solution, depending on actual needs.

[0140] Furthermore, the functional modules or sub-modules in the various embodiments of this application can be integrated into one processing module, or each module or sub-module can exist physically separately, or two or more modules or sub-modules can be integrated into one module. The integrated modules or sub-modules described above can be implemented in hardware or in the form of software functional modules or sub-modules.

[0141] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0142] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software unit executed by a processor, or a combination of both. The software unit can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0143] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

Claims

1. A smart cockpit interaction method, characterized in that, The method includes: Acquire first audio information and context information currently collected within the smart cockpit; wherein, the context information includes second audio information collected during the historical process through pickup devices in multiple sound zones within the smart cockpit; Based on the first audio information and the context information, the target type to which the first audio information belongs is determined; wherein, the target type is one of the following: instruction type, question type, exclusive private chat type, and invalid type; Based on the target type to which the first audio information belongs, a processing strategy for the first audio information is determined and the processing strategy is executed; wherein, the processing strategy is one of executing a vehicle control command corresponding to the first audio information, responding to a question corresponding to the first audio information, remaining silent on the first audio information, and refusing to recognize the first audio information.

2. The intelligent cockpit interaction method according to claim 1, characterized in that, The step of determining the target type to which the first audio information belongs based on the first audio information and the context information includes: Based on the target features of the first audio information and the second audio information, the target type to which the first audio information belongs is determined; wherein, the target features include at least one of the following: semantic features, acoustic features, and a cockpit role label pre-set for each audio information, wherein the cockpit role label is used to identify the voice output user in different sound zones within the cockpit.

3. The intelligent cockpit interaction method according to claim 2, characterized in that, The step of determining the target type of the first audio information based on the target features of the first audio information and the second audio information includes: If the first audio information contains words that match vehicle control commands and the audio energy is concentrated in the words, the first audio information is determined to be a command type. or, If the current conversation is determined to be a human-to-human conversation based on the context information, and if the first audio information is detected to include interrogative intonation features and the topic of the current conversation belongs to the target topic, then the first audio information is determined to be a question type. or, If the first audio information is detected to be non-human speech or a sudden human speech shock signal, the first audio information is determined to be invalid. or, If, based on the first audio information and the second audio information, it is determined that the first audio information does not belong to any of the instruction type, question type, or invalid type, then the first audio information is determined to belong to the exclusive private chat type.

4. The intelligent cockpit interaction method according to claim 2, characterized in that, After acquiring the target audio information, the method further includes: Based on the identification number of the audio pickup device that collects the target audio information, a corresponding cockpit role tag is added to the target audio information; wherein, the target audio information is either the first audio information or the second audio information.

5. The intelligent cockpit interaction method according to claim 1 or 4, characterized in that, The execution of the processing strategy includes: When the processing strategy includes outputting voice interaction information, the cockpit role label corresponding to the first audio information is determined; wherein, the cockpit role label is used to identify voice output users in different voice zones within the cockpit; Based on the cockpit role labels, determine the broadcasting equipment for the corresponding area; The voice interaction information is played through the broadcasting device.

6. The intelligent cockpit interaction method according to claim 1, characterized in that, The context information also includes vehicle status information acquired during the acquisition of the first audio information; When the first audio information belongs to an instruction type, determining the processing strategy for the first audio information based on the target type to which the first audio information belongs includes: Based on the vehicle control command corresponding to the first audio information and the vehicle status information, a processing strategy for the first audio information is determined.

7. The intelligent cockpit interaction method according to claim 1, characterized in that, The steps of determining the target type to which the first audio information belongs and determining the processing strategy for the first audio information include: The first audio information, the context information, and the set of vehicle control commands are input into a pre-trained multimodal large model. The multimodal large model is used to determine the target type of the first audio information and the processing strategy for the first audio information.

8. An electronic device, characterized in that, include: Memory and processor; The memory is connected to the processor and is used to store programs; The processor is used to implement the smart cockpit interaction method as described in any one of claims 1 to 7 by running the program in the memory.

9. A computer program product, characterized in that, The computer program product stores a computer program, which, when executed by a processor, implements the intelligent cockpit interaction method as described in any one of claims 1 to 7.

10. A storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the intelligent cockpit interaction method as described in any one of claims 1 to 7.