Vehicle-mounted intelligent man-machine interaction method, device, equipment and medium

By selecting an appropriate level of model for processing based on the complexity and urgency of the instruction signal, the problem of high load of multimodal interactive processors in the vehicle is solved, and the rational utilization of resources and power consumption saving is achieved.

CN120207365APending Publication Date: 2025-06-27VOYAH AUTOMOBILE TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510266092.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the prior art, when multimodal interaction processing is performed in vehicles, the processor is often in unnecessary high load states, resulting in waste of resources and increased power consumption.

Method used

By pre-training models of different levels and determining which level of model is used to process the instruction signal based on the complexity and urgency of the instruction signal, thereby matching the processor's computing power.

Benefits of technology

It realizes that the processor provides more matching computing power when processing instruction signals, and rationally utilizes the processor computing power to save power consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120207365A_ABST
    Figure CN120207365A_ABST
Patent Text Reader

Abstract

The invention discloses a vehicle-mounted intelligent man-machine interaction method, device and equipment and a medium, and belongs to the technical field of vehicles. The method comprises the steps that instruction signals sent by passengers on a vehicle and state information of the vehicle are obtained, the instruction signals comprise voice signals, gesture signals and expression signals, and the state information comprises working condition information, a driver attention value and alarm information; determining the complexity of the instruction signal; determining the emergency degree of the instruction signal according to the state information; according to the complexity and the emergency degree, the target level of a model for processing the instruction signal is determined from models of multiple levels, the models of different levels need different computing power of a processor during operation, and the models are obtained through pre-training of a training set; the model of the target level is controlled to process the instruction signal, an interaction task output by the model is obtained, and the vehicle is controlled to execute the interaction task. According to the method, the processor for processing the instruction signal can provide more matched computing power, the computing power of the processor is reasonably utilized, and power consumption is saved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of vehicles, and in particular, to an in-vehicle intelligent human-computer interaction method, device, equipment and medium. Background Art

[0002] With the rapid development of intelligent vehicle systems, the demand for users to interact with vehicles through voice or gestures is increasing day by day. As a new interaction method, multimodal technology has been gradually applied to various vehicle usage scenarios, such as multimodal driving technology, multimodal autonomous driving technology, etc.; multimodal interaction can achieve more natural interaction with machines or computers.

[0003] In the prior art, the multimodal interaction method will identify commands through gestures and then combine voice to more accurately control the in-vehicle system.

[0004] However, in order to ensure the processing efficiency of commands, a large amount of computing power of the processor is used for command processing in any scenario, resulting in the processor sometimes being in an unnecessary high-load state, causing resource waste and increasing power consumption. Summary of the Invention

[0005] In view of the above problems, the present invention is proposed to provide an in-vehicle intelligent human-computer interaction method, device, equipment and medium that can solve the above problems. Different levels of models can be pre-trained, and then according to the complexity and urgency of the command signal, it is determined which level of model is used to process the command signal, so that the processor can provide more matching computing power when processing the command signal, reasonably utilize the computing power of the processor, and save power consumption.

[0006] In a first aspect, the present invention provides an in-vehicle intelligent human-computer interaction method, and the method includes:

[0007] Obtain an instruction signal issued by an occupant in the vehicle and the status information of the vehicle, where the instruction signal includes a voice signal, a gesture signal and an expression signal, and the status information includes working condition information, driver attention value and alarm information;

[0008] Determine the complexity of the instruction signal;

[0009] According to the status information, determine the urgency of the instruction signal;

[0010] According to the complexity and the urgency, determine the target level of the model for processing the instruction signal from multiple levels of models. Different levels of models require different computing powers of the processor when running, and the model is pre-trained through a training set;

[0011] The model that controls the target level processes the instruction signal to obtain an interaction task output by the model, and controls the vehicle to execute the interaction task.

[0012] Optionally, the multiple levels include a deep level and a shallow level. Determining the target level of the model for processing the instruction signal according to the complexity and the urgency includes:

[0013] If the urgency is greater than a preset urgency threshold, determine that the target level of the model for processing the instruction signal is the deep level, and the computing power of the processor required when the model at the deep level runs is greater than that required when the model at the shallow level runs;

[0014] If the complexity is greater than a preset complexity threshold, obtain the load rate of the processor;

[0015] If the load rate is greater than a preset load rate threshold, determine that the target level of the model for processing the instruction signal is the shallow level or the deep level; if the load rate is less than or equal to the load rate threshold, determine that the target level of the model for processing the instruction signal is the deep level;

[0016] If the urgency is less than or equal to the urgency threshold and the complexity is less than or equal to the complexity threshold, determine that the target level of the model for processing the instruction signal is the shallow level.

[0017] Optionally, there are multiple instruction signals, and the model is used to:

[0018] Obtain the priority and the issuing time of the instruction signal;

[0019] Sort the multiple instruction signals according to the priority, the issuing time, and the urgency, and process the instruction signals in sequence according to the sorting.

[0020] Optionally, the method further includes:

[0021] Determine the instruction type of the instruction signal, and different instruction types correspond to different priorities;

[0022] Determine the priority corresponding to the instruction type to obtain the priority of the instruction signal.

[0023] Optionally, the model is further used to:

[0024] Obtain the privilege level of the occupant who issues the instruction signal;

[0025] Sort multiple instruction signals according to the permission level, the priority, the sending time, and the urgency, and process the instruction signals in sequence according to the sorting.

[0026] Optionally, the method further includes:

[0027] Obtain the identity characteristic information of the occupant, where the identity characteristic information includes voiceprint and face image;

[0028] Determine the permission level of the occupant according to the identity characteristic information.

[0029] Optionally, if the instruction signal includes a first instruction signal with a complexity greater than the complexity threshold and a second instruction signal with a complexity less than or equal to the complexity threshold generated at the same time, the controlling the model of the target level to process the instruction signal to obtain the interactive task output by the model includes:

[0030] If the urgency is less than or equal to the urgency threshold and the load rate is greater than the load rate threshold, first control the model of the shallow level to process the second instruction signal to obtain the first interactive task output by the model, and then control the model of the deep level to process the first instruction signal to obtain the second interactive task output by the model.

[0031] In a second aspect, the present invention provides an in-vehicle intelligent human-machine interaction device, and the device includes:

[0032] An acquisition module, configured to acquire instruction signals sent by an occupant in the vehicle and the state information of the vehicle, where the instruction signals include voice signals, gesture signals, and expression signals, and the state information includes working condition information, driver attention value, and alarm information;

[0033] A first determination module, configured to determine the complexity of the instruction signal;

[0034] A second determination module, configured to determine the urgency of the instruction signal according to the state information;

[0035] A third determination module, configured to determine the target level of the model for processing the instruction signal from multiple levels of models according to the complexity and the urgency, where different levels of models require different computing powers of the processor, and the model is pre-trained through a training set;

[0036] A control module, configured to control the model of the target level to process the instruction signal to obtain the interactive task output by the model, and control the vehicle to execute the interactive task.

[0037] In a third aspect, the present invention provides an electronic device, comprising: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to execute the method as described in the first aspect.

[0038] In a fourth aspect, the present invention provides a computer-readable storage medium storing computer instructions for causing a computer to execute the method as described in the first aspect.

[0039] The technical solutions provided in the embodiments of the present invention have at least the following technical effects or advantages:

[0040] A vehicle-mounted intelligent human-computer interaction method, device, equipment and medium provided in an embodiment of the present invention obtain an instruction signal sent by an occupant in a vehicle and status information of the vehicle, understand the interaction requirements of the occupant and the status of the vehicle. The instruction signal includes a voice signal, a gesture signal and an expression signal, and the status information includes working condition information, a driver attention value and alarm information; determine the complexity of the instruction signal to understand whether the instruction signal is easy to process; determine the urgency of the instruction signal according to the status information to understand whether the status of the vehicle is good; determine the target level of the model for processing the instruction signal from multiple levels of models according to the complexity and the urgency, and use appropriate computing power to process the instruction signal. Different levels of models require different computing power of the processor when running, and the models are pre-trained through a training set; control the model of the target level to process the instruction signal to obtain an interaction task output by the model, and control the vehicle to execute the interaction task to realize the interaction requirements of the occupant. This method enables the processor to provide more matching computing power when processing the instruction signal, rationally utilize the computing power of the processor, and save power consumption.

[0041] The above description is only an overview of the technical solutions of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the description. And in order to make the above and other objects, features and advantages of the present invention more obvious and understandable, the following specific embodiments of the present invention are specifically given. Description of the Drawings

[0042] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0043] Figure 1 is a flowchart of a vehicle-mounted intelligent human-computer interaction method provided in an embodiment of the present invention;

[0044] Figure 2It is a structural block diagram of an in-vehicle intelligent human-computer interaction device provided by an embodiment of the present invention. Detailed implementation manners

[0045] To make the objectives, technical solutions and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings. It should be understood that the embodiments of the present disclosure and the specific features in the embodiments are detailed descriptions of the technical solutions of the present application, rather than limitations on the technical solutions of the present application. Without conflict, the technical features in the embodiments of the present application and the embodiments can be combined with each other.

[0046] Figure 1 It is a flowchart of an in-vehicle intelligent human-computer interaction method provided by an embodiment of the present invention. As Figure 1 shown, the method includes:

[0047] Step S110, obtaining an instruction signal issued by an occupant in the vehicle and status information of the vehicle. Among them, the instruction signal includes a voice signal, a gesture signal and an expression signal, and the status information includes working condition information, a driver attention value and alarm information.

[0048] In an embodiment of the present application, the voice signal of the occupant can be collected through a microphone array, the gesture signal of the occupant can be collected through a camera or a radar, the expression signal can be collected through a camera, the driver attention value can be collected through a camera or an infrared sensor, the working condition information can be collected through a sensor, and the alarm information can be obtained from the vehicle controller. Among them, the working condition information may include information such as vehicle speed and steering wheel angle, and the alarm information may include collision warning, tire pressure abnormal alarm, door lock abnormal alarm and other information.

[0049] In an embodiment of the present application, the interface of the sensor is connected to the data acquisition and preprocessing module, and status information such as vehicle speed, steering wheel angle, and driver attention value collected by the sensor is transmitted to the data acquisition and preprocessing module as context information.

[0050] Step S120, determining the complexity of the instruction signal.

[0051] In an embodiment of the present application, a complexity analyzer trained by a large amount of voice, text, expression and gesture data can be used to quickly score the complexity of the instruction signal. For example, the complexity analyzer identifies keywords and semantic tags of the voice signal, analyzes the gesture signal, obtains the interpretation results of the voice signal and the gesture signal, and performs a complexity score according to the interpretation results to reflect the complexity of the instruction signal through the complexity.

[0052] Among them, the magnitude range of complexity is 0 - 1, and the complexity levels can be divided according to the magnitude of complexity. For example, when the complexity is between 0.0 and 0.4, the complexity level is simple; when the complexity is between 0.4 and 0.7, the complexity level is medium; when the complexity is between 0.7 and 1.0, the complexity level is high. The greater the complexity or the higher the complexity level, the more complex the instruction signal, and the more difficult it is to analyze the interaction task expressed by the instruction signal.

[0053] Step S130: Determine the urgency of the instruction signal according to the status information.

[0054] In the embodiment of the present application, the status information can reflect the current running safety degree of the vehicle. Specifically, the driving scenario of the vehicle can be reflected by the working condition information, the emergency response ability of the driver can be reflected by the driver attention value, and the health condition of the vehicle can be reflected by the alarm information.

[0055] In the embodiment of the present application, the quantitative determination standard of the urgency is as follows:

[0056] The first step: Determine the vehicle speed urgency a according to the vehicle speed. The vehicle speed is positively correlated with the vehicle speed urgency a. For example, when the vehicle speed is 0, a = 0; when the vehicle speed is 220 km / h, a = 1; or calculate the ratio of the vehicle speed to the preset vehicle speed threshold, and determine the vehicle speed urgency a according to the ratio. The ratio is positively correlated with the vehicle speed urgency a. For example, the vehicle speed threshold is 80 km / h. When the vehicle speed is greater than the vehicle speed threshold, it can be regarded as high-speed driving. The greater the ratio, the faster the vehicle speed, and the greater the vehicle speed urgency a; or calculate the acceleration or deceleration according to the vehicle speed, and determine the vehicle speed urgency a according to the acceleration or deceleration. The acceleration or deceleration is positively correlated with the vehicle speed urgency a. For example, when the vehicle suddenly accelerates or brakes suddenly, the acceleration or deceleration will be relatively large, and the vehicle speed urgency a will be greater.

[0057] The second step: Determine the driving urgency b according to the driver attention value. The driver attention value is negatively correlated with the driving urgency b. For example, based on the data monitored by the camera, such as the driver's eye fixation time, blink frequency, head posture, and steering wheel grip force, determine the driver attention value; among them, if abnormal behaviors such as the driver closing eyes for a long time, yawning, or both hands leaving the steering wheel are detected, it means that the driver is fatigued or distracted, and the determined driver attention value is small. The value range of the driver attention value is 0 - 1, and the lower it is, the less concentrated the driver's attention is, and the greater the driving urgency b is.

[0058] The third step: Determine the alarm urgency c according to the alarm information. For example, when the vehicle has collision warning, lane departure warning, tire pressure abnormal warning, obstacle warning, in-vehicle / out-of-vehicle environment alarm, etc., determine the magnitude of the alarm urgency c according to the severity of these situations. The more serious, the greater the alarm urgency.

[0059] Among them, the in-vehicle / outside-vehicle environment alarms include abnormal door locks, airbag failures, high-temperature alarms, etc.

[0060] Step 4: Perform weighted summation on the vehicle speed urgency a, driving urgency b, and alarm urgency c to obtain the urgency f of the command signal. For example, the value ranges of a, b, and c are all 0 - 1, and f = 0.3a + 0.3b + 0.4c.

[0061] Exemplarily, the vehicle speed is 220 km / h, a = 1; the driver attention value is 0.4, b = 0.6; there is a collision warning triggered, c = 0.8; the calculated urgency f of the command signal is 0.8.

[0062] Step S140: Determine the target level of the model for processing the command signal from multiple levels of models according to the complexity and urgency. Different levels of models require different computing powers of the processor when running, and the models are pre-trained through a training set.

[0063] In the embodiments of the present application, the greater the complexity, the more difficult it is to parse the meaning of the command signal, and the higher the urgency, the more urgently it is required to quickly parse the meaning of the command signal. Therefore, different levels of models are selected according to the complexity and urgency for processing and parsing the command signal. Among them, the greater the complexity or the higher the urgency, the model that requires more computing power is used to process the command signal to more quickly and accurately parse the meaning expressed by the command signal.

[0064] Among them, a large number of sample data can be collected in advance, and the sample data is used as a training set to train the initial model to obtain a model for processing the command signal.

[0065] Step S150: Control the model of the target level to process the command signal to obtain the interaction task output by the model, and control the vehicle to execute the interaction task.

[0066] In the embodiments of the present application, after the model processes the command signal, the control intention of the occupant can be known, an interaction task is generated according to the control intention, and then the vehicle is controlled to execute the interaction task to meet the control requirements of the user.

[0067] The method in the embodiments of the present application can pre-train models of different levels, and then determine which level of model to use for processing the command signal according to the complexity and urgency of the command signal, so that the processor can provide more matching computing power when processing the command signal, rationally utilize the computing power of the processor, and save power consumption.

[0068] In the embodiments of the present application, the multiple levels include a deep level and a shallow level. The computing power required by the model of the deep level when running is greater than that required by the model of the shallow level when running.

[0069] The models at the shallow level are applicable to instruction signals with low complexity and low urgency, such as simple music playback, basic question answering, light control, etc.; they provide basic speech recognition and the ability to parse simple intents, without performing multi-round deep semantic reasoning.

[0070] The models at the deep level are applicable to instruction signals with high complexity or high urgency. For example: The instruction signals of multi-modal fusion are high-complexity instruction signals, that is, speech signals, gesture signals, and facial expression signals are generated simultaneously. In emergency situations such as sudden changes in vehicle speed and hard braking, the generated instruction signals are high-urgency instruction signals. It is applicable to in-depth inference of conflicting instruction signals for multiple occupants (such as simultaneously generating instruction signals that the back row wants to watch a video and the driver requires quiet). The models at the deep level usually include Multi-Head Attention (MHA) or a longer Transformer (neural network architecture) depth, and can achieve cross-modal and multi-round reasoning; they will only be activated when system resources are available or the instruction is evaluated as "urgent or complex" to avoid occupying a large amount of computing power all the time.

[0071] Optionally, step S140 includes step S1401, step S1402, step S1403, and step S1404.

[0072] Step S1401: If the urgency is greater than the preset urgency threshold, determine that the target level of the model for processing the instruction signal is the deep level.

[0073] In the embodiments of the present application, if the urgency is greater than the preset urgency threshold, it means that the urgency is relatively large and the meaning of the instruction signal needs to be processed quickly. And the processing efficiency of the models at the deep level is relatively fast. Therefore, the models at the deep level are used to process the instruction signals with relatively high urgency to improve the processing efficiency. For example, the urgency threshold is 0.8.

[0074] Step S1402: If the complexity is greater than the preset complexity threshold, obtain the load rate of the processor.

[0075] In the embodiments of the present application, if the complexity is greater than the preset complexity threshold, it means that the instruction signal is relatively complex and difficult to understand. In this case, according to the load situation of the processor, it is determined what level of model to use for processing the instruction signal. For example, the complexity threshold is 0.6.

[0076] Among them, the processor can be a Central Processing Unit (CPU) or a Graphics Processing Unit (GPU). By detecting the usage rate, available memory, and bandwidth of the processor, the load rate is determined.

[0077] Step S1403: If the load rate is greater than a preset load rate threshold, determine that the target level of the model for processing the instruction signal is the shallow level or the deep level; if the load rate is less than or equal to the load rate threshold, determine that the target level of the model for processing the instruction signal is the deep level.

[0078] In the embodiment of the present application, if the load rate is greater than the preset load rate threshold, it indicates that the load of the processor is high and the remaining available computing power is very small. It is okay to use either the shallow-level or deep-level model to process the instruction signal. Because using the shallow-level model can save the computing power of the processor, while using the deep-level model can provide the accuracy of instruction signal parsing. If the load rate is less than or equal to the load rate threshold, it indicates that the load of the processor is low and the remaining available computing power is sufficient, then determine that the target level of the model for processing the instruction signal is the deep level.

[0079] Step S1404: If the urgency is less than or equal to the urgency threshold and the complexity is less than or equal to the complexity threshold, determine that the target level of the model for processing the instruction signal is the shallow level.

[0080] In the embodiment of the present application, if the urgency is less than or equal to the urgency threshold and the complexity is less than or equal to the complexity threshold, it indicates that the current situation is not urgent and the instruction signal is not complex and is easy to analyze. Therefore, determine that the target level of the model for processing the instruction signal is the shallow level, which can save the computing power of the processor and reduce power consumption.

[0081] Optionally, the method further includes:

[0082] Determine the instruction type of the instruction signal, and different instruction types correspond to different priorities; determine the priority corresponding to the instruction type to obtain the priority of the instruction signal.

[0083] In the embodiment of the present application, the instruction signals can be classified, and the instruction signals of different instruction types correspond to different priorities. For example, the instruction types include safety instructions, vehicle control instructions, navigation instructions, and entertainment instructions with decreasing priority levels. The instruction signals related to "turn on the light" or "adjust the volume" are entertainment instructions, and the instruction signals related to "emergency braking" are safety instructions.

[0084] Specifically, the instruction type of the instruction signal can be analyzed by a complexity analyzer.

[0085] Optionally, there are multiple instruction signals, and the model is used for:

[0086] Obtain the priority and the sending time of the instruction signal; sort the multiple instruction signals according to the priority, the sending time, and the urgency, and process the instruction signals in sequence according to the sorting.

[0087] In an embodiment of the present application, if there are multiple instruction signals to be processed, these instruction signals to be processed are sorted and stored in an instruction queue to be processed. When sorting, the sorting is performed according to the urgency and priority of the instruction signals. The higher the urgency, the higher the sorting order. Under the same urgency, the higher the priority, the higher the sorting order. Under the same priority and the same urgency, the earlier the emission time, the higher the sorting order.

[0088] In an embodiment of the present application, when the model at the depth level is processing an instruction signal with low urgency, suddenly, the occupant issues an instruction signal with high urgency. Then, the depth model can interrupt the processing of the instruction signal with low urgency and use all computing power to process the newly issued instruction signal with high urgency.

[0089] Exemplarily, if there is a previously generated instruction signal with low priority in the instruction queue to be processed of the model at the depth level, and at this time, a newly generated instruction signal with high priority is added. The urgencies of these two instruction signals are the same. Then, the instruction signal with high priority is processed first, and then the instruction signal with low priority is processed.

[0090] Exemplarily, when the vehicle is traveling at a speed of 100 km / h and there are alarm messages such as radar warning and lane departure, the driver issues a voice signal: "Emergency braking assistance" or automatically recognizes a scenario that requires emergency braking. At this time, the urgency of the driver's voice instruction is relatively high. The processing process of the voice signal that is currently playing a movie in the back row is stopped, and all available computing power is occupied to load the model at the depth level for voice signal processing. After processing, an interactive task of generating linkage Electronic Stability Control (ESC for short), Antilock Brake System (ABS), or issuing a warning is generated. The latency of the entire processing process is controlled within hundreds of milliseconds to give priority to ensuring the real-time nature of safety instructions.

[0091] Optionally, the method further includes:

[0092] Obtaining the identity characteristic information of the occupant, where the identity characteristic information includes voiceprint and face image; and determining the authority level of the occupant according to the identity characteristic information.

[0093] In an embodiment of the present application, according to the voice signals of each occupant, the actual voiceprint of each occupant is extracted, and the actual voiceprint is compared with the pre-stored standard voiceprint to obtain the voiceprint matching degree of each occupant; according to the comparison between the actual face image of each occupant and the pre-stored standard face image, the face credibility of each occupant is obtained; it is determined that the occupant with a voiceprint matching degree greater than the matching degree threshold or a face credibility greater than the credibility threshold is a registered user, and the privilege level of the registered user is pre-stored, and the corresponding privilege level of the registered user can be directly obtained; it is determined that the occupant with a voiceprint matching degree less than or equal to the matching degree threshold and a face credibility less than or equal to the credibility threshold is a new user, and the privilege level of the new user is the lowest.

[0094] In an embodiment of the present application, the seat weight can also be collected by a pressure sensor. If the seat weight is greater than the weight threshold, it means that there is someone on the seat. Then, combined with a microphone array for sound source localization and face recognition, the identity of the occupant on the seat is determined to achieve an accurate match of "position and identity". Each occupant is provided with a context cache, and the command signal, historical preference, and privilege level of each occupant are stored in the corresponding context cache to reduce misidentification and role confusion. When the occupant leaves the seat or the door lock state changes, it means that there is an occupant getting on or off the vehicle or changing seats, then the context cache automatic cleaning and transfer mechanism is triggered to avoid occupant privacy leakage or misidentification. For example, when the occupant changes seats, the context cache can be migrated to the new seat / new identity to maintain the semantic continuity across seats.

[0095] Optionally, the model is further used for:

[0096] Obtaining the privilege level of the occupant who issues the command signal; sorting multiple command signals according to the privilege level, priority, issuance time, and urgency, and processing the command signals in sequence according to the sorting.

[0097] In an embodiment of the present application, if there are multiple command signals to be processed, these command signals to be processed are sorted and stored in a queue of command signals to be processed. When sorting, it is sorted according to the issuance time, priority, urgency, and privilege level of the command signal. The higher the urgency, the higher the sorting. Under the same urgency, the higher the priority, the higher the sorting. Under the same urgency and the same priority, the earlier the issuance time, the higher the sorting. Under the same issuance time, the same priority, and the same urgency, the higher the privilege level, the higher the sorting.

[0098] Exemplarily, when the vehicle is in a low-speed driving state, the driver and the rear-seat passenger almost simultaneously send instruction signals; the driver sends a gesture signal of "turn on the air conditioner and play the navigation", and the rear-seat passenger sends a voice signal of "I want to watch the latest movie trailer". There are both gesture signals and voice signals, indicating that the current instruction signal is a multimodal instruction signal, and the complexity is greater than the complexity threshold. If the load rate of the processor is low at this time, it is considered that there is relatively idle computing power, and a deep-level model is called to process the above voice instruction. If the load rate is large at this time, a shallow-level model is called to process the above voice instruction. Among them, the gesture signal and the voice signal are sent at the same time and have the same urgency, but the driver's authority level is higher than that of the rear-seat passenger. "Turn on the air conditioner" belongs to a vehicle control instruction, and "play the navigation" belongs to a navigation instruction. The priority of the navigation instruction is lower than that of the vehicle control instruction. Therefore, the model first satisfies the driver's high-priority request to turn on the air conditioner, then satisfies the navigation request, and finally satisfies the rear-seat passenger's movie entertainment request.

[0099] Optionally, the instruction signal includes a first instruction signal with a complexity greater than the complexity threshold and a second instruction signal with a complexity less than or equal to the complexity threshold generated at the same time. Step S150 includes:

[0100] If the urgency is less than or equal to the urgency threshold and the load rate is greater than the load rate threshold, first control the shallow-level model to process the second instruction signal, and after obtaining the first interaction task output by the model, then control the deep-level model to process the first instruction signal to obtain the second interaction task output by the model.

[0101] In the embodiment of the present application, in the case of low urgency and high load, the shallow-level model can be called first to process the second instruction signal, and then the deep-level model can be called to process the first instruction signal, which can not only meet the processing efficiency of the instruction signal, but also make full use of the computing power of the processor.

[0102] In the embodiment of the present application, the model can also identify whether the intentions of multiple instruction signals conflict. If there is a conflict, the instruction signal with a higher processing priority or a higher authority level is selected. It is also possible to store and analyze the data generated in the whole method, dynamically adjust each threshold, and the computing power allocation strategy of different-level models, so as to improve the accuracy, response delay and reduce energy consumption of the method in this embodiment.

[0103] In the method of the embodiment of the present application, the state information of the vehicle, such as vehicle speed, steering wheel angle, driver attention value, alarm information, etc., is deeply bound, rather than just general multi-modal recognition; when a sudden safety problem occurs, the system can quickly preempt computing power to perform in-depth reasoning and priority response through a deep-level model, realizing double guarantees for vehicle safety and human-machine interaction. Strengthen the quantization determination threshold of multi-modal fusion. For voiceprint, face recognition, seat pressure, gestures, etc., there are clear thresholds / score intervals; effectively distinguish identities through quantitative scoring, reduce algorithm ambiguity, improve the accuracy rate, and highlight the executable depth of the method. The hierarchical reasoning mechanism dynamically switches between "shallow - deep" and combines vehicle state information, reflecting the efficient application of the vehicle-side large model under resource-constrained conditions.

[0104] Based on the same inventive concept, an embodiment of the present invention further provides an in-vehicle intelligent human-machine interaction device. Figure 2 It is a structural block diagram of an in-vehicle intelligent human-machine interaction device provided by an embodiment of the present invention. As Figure 2 shown, the device 200 includes an acquisition module 201, a first determination module 202, a second determination module 203, a third determination module 204, and a control module 205.

[0105] The acquisition module 201 is configured to acquire an instruction signal sent by an occupant in the vehicle and the state information of the vehicle. The instruction signal includes a voice signal, a gesture signal, and an expression signal, and the state information includes operating condition information, driver attention value, and alarm information.

[0106] The first determination module 202 is configured to determine the complexity of the instruction signal.

[0107] The second determination module 203 is configured to determine the urgency of the instruction signal according to the state information.

[0108] The third determination module 204 is configured to determine the target level of the model for processing the instruction signal from multiple levels of models according to the complexity and urgency. Different levels of models require different computing powers of the processor when running, and the models are pre-trained through a training set.

[0109] The control module 205 is configured to control the model of the target level to process the instruction signal, obtain the interaction task output by the model, and control the vehicle to execute the interaction task.

[0110] Optionally, the multiple levels include a deep level and a shallow level. The third determination module 203 is further configured to:

[0111] If the urgency is greater than a preset urgency threshold, determine that the target level of the model for processing the instruction signal is the deep level. The deep-level model requires a greater computing power of the processor when running than the shallow-level model.

[0112] If the complexity is greater than a preset complexity threshold, obtain the load rate of the processor;

[0113] If the load rate is greater than a preset load rate threshold, determine that the target level of the model for processing the instruction signal is the shallow level or the deep level; if the load rate is less than or equal to the load rate threshold, determine that the target level of the model for processing the instruction signal is the deep level;

[0114] If the urgency is less than or equal to the urgency threshold and the complexity is less than or equal to the complexity threshold, determine that the target level of the model for processing the instruction signal is the shallow level.

[0115] Optionally, there are multiple instruction signals, and the model is used to:

[0116] Obtain the priority and issuance time of the instruction signal;

[0117] Sort the multiple instruction signals according to the priority, issuance time, and urgency, and process the instruction signals in sequence according to the sorting.

[0118] Optionally, the apparatus 200 further includes a fourth determination module, which is used to:

[0119] Determine the instruction type of the instruction signal, and different instruction types correspond to different priorities;

[0120] Determine the priority corresponding to the instruction type to obtain the priority of the instruction signal.

[0121] Optionally, the model is further used to:

[0122] Obtain the permission level of the occupant who issues the instruction signal;

[0123] Sort the multiple instruction signals according to the permission level, priority, issuance time, and urgency, and process the instruction signals in sequence according to the sorting.

[0124] Optionally, the apparatus 200 further includes a fifth determination module, which is used to:

[0125] Obtain the identity characteristic information of the occupant, and the identity characteristic information includes voiceprint and face image;

[0126] Determine the permission level of the occupant according to the identity characteristic information.

[0127] Optionally, the instruction signal includes a first instruction signal with a complexity greater than the complexity threshold and a second instruction signal with a complexity less than or equal to the complexity threshold generated at the same time, and the control module 205 is further used to:

[0128] If the urgency level is less than or equal to the urgency threshold and the load rate is greater than the load rate threshold, first control the model at the shallow level to process the second instruction signal. After obtaining the first interaction task output by the model, then control the model at the deep level to process the first instruction signal to obtain the second interaction task output by the model.

[0129] It can be understood that for the device provided in the above embodiment, only the division of the above functional modules is used for illustration. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.

[0130] An embodiment of the present invention further provides an electronic device, which may include a processor and a memory, and the processor and the memory may be communicatively connected to each other through a bus or other means.

[0131] The processor may be a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application, or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. chips, or combinations of the above types of chips.

[0132] The memory may include a mass storage for data or instructions. By way of example and not limitation, the memory may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. In a suitable case, the memory may include a removable or non-removable (or fixed) medium. In a suitable case, the memory may be internal or external to the electronic device. In a particular embodiment, the memory may be a non-volatile solid state memory.

[0133] In one example, the memory may be a Read Only Memory (ROM). In one example, the ROM may be a mask-programmed ROM, a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically Erasable PROM (EEPROM), an Electrically Alterable ROM (EAROM), or a flash memory, or a combination of two or more of these.

[0134] The processor reads and executes the computer program instructions stored in the memory to implement any one of the in-vehicle intelligent human-machine interaction methods in the above embodiments.

[0135] In one example, the electronic device may further include a communication interface and a bus. Among them, the processor, the memory, and the communication interface are connected through the bus and complete communication with each other. The communication interface is mainly used to implement communication between various modules, devices, units, and / or devices in the embodiments of the present application. In a suitable case, the bus may include one or more buses.

[0136] In addition, in combination with the in-vehicle intelligent human-machine interaction method in the above embodiments, the embodiments of the present invention may provide a computer-readable storage medium to implement. Computer program instructions are stored on the computer-readable storage medium; when the computer program instructions are executed by the processor, any one of the in-vehicle intelligent human-machine interaction methods in the above embodiments is implemented.

[0137] Those skilled in the art can understand that to implement all or part of the processes in the methods of the above embodiments, it can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it may include the processes of the embodiments of the above methods. Among them, the storage medium may be a Read-Only Memory (ROM), a Random Access Memory (RAM), a Flash Memory, a Hard Disk Drive (HDD), or a Solid-State Drive (SSD), etc.; the storage medium may also include a combination of the above types of memories.

[0138] The technical solutions in the embodiments of the present application at least have the following technical effects or advantages:

[0139] An in-vehicle intelligent human-computer interaction method, device, equipment and medium provided by an embodiment of the present invention obtain an instruction signal sent by an occupant in a vehicle and the state information of the vehicle, understand the interaction requirements of the occupant and the state of the vehicle, where the instruction signal includes a voice signal, a gesture signal and an expression signal, and the state information includes working condition information, a driver attention value and alarm information; determine the complexity of the instruction signal to understand whether the instruction signal is easy to process; determine the urgency of the instruction signal according to the state information to understand whether the state of the vehicle is good; determine the target level of the model for processing the instruction signal from multiple levels of models according to the complexity and the urgency, and use appropriate computing power to process the instruction signal. Different levels of models require different computing power of the processor when running, and the models are pre-trained through a training set; control the model of the target level to process the instruction signal to obtain an interaction task output by the model, and control the vehicle to execute the interaction task to realize the interaction requirements of the occupant. This method enables the processor for processing the instruction signal to provide more matching computing power, reasonably utilize the computing power of the processor, and save power consumption.

[0140] In the specification provided herein, a large number of specific details are set forth. It is, however, understood that embodiments of the present invention may be practiced without these specific details. In some instances, well-known methods, structures and technologies have not been shown in detail so as not to obscure an understanding of the present specification.

[0141] Similarly, it should be understood that in order to streamline this disclosure and assist in understanding one or more of the various inventive aspects, in the foregoing description of exemplary embodiments of the present invention, the various features of the present invention are sometimes grouped together in a single embodiment, figure, or description thereof. However, the disclosed method should not be construed as reflecting an intention that: the claimed invention requires more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive aspects lie in less than all the features of a single foregoing disclosed embodiment. Thus, the claims following the detailed description are hereby expressly incorporated into this detailed description, with each claim standing on its own as a separate embodiment of the present invention.

[0142] It should be noted that the above embodiments illustrate the present invention rather than limit the present invention, and those skilled in the art can design alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in the claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present invention can be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In the unit claims listing several means, several of these means can be embodied by one and the same item of hardware. The use of the words first, second, and third, etc. does not denote any order. These words can be interpreted as names.

Claims

1. A vehicle-mounted intelligent human-computer interaction method, characterized in that: The method comprises: Acquire a command signal issued by a passenger on the vehicle and status information of the vehicle, wherein the command signal includes a voice signal, a gesture signal and an expression signal, and the status information includes operating condition information, a driver's attention value and alarm information; determining the complexity of the command signal; Determining the urgency of the command signal according to the status information; According to the complexity and the urgency, determine a target level of a model for processing the instruction signal from models of multiple levels, different levels of models requiring different processor computing power when running, and the model is pre-trained through a training set; The model of the target level is controlled to process the command signal to obtain an interactive task output by the model, and the vehicle is controlled to perform the interactive task.

2. The vehicle-mounted intelligent human-computer interaction method according to claim 1, characterized in that: The multiple levels include a depth level and a shallow level, and determining a target level of a model for processing the instruction signal from the multiple levels of models according to the complexity and the urgency, comprises: If the urgency is greater than a preset urgency threshold, determining that the target level of the model for processing the instruction signal is the depth level, and the computing power of the processor required when the model at the depth level is running is greater than the computing power of the processor required when the model at the shallow level is running; If the complexity is greater than a preset complexity threshold, obtaining a load rate of the processor; If the load rate is greater than a preset load rate threshold, the target level of the model for processing the command signal is determined to be the shallow level or the deep level; if the load rate is less than or equal to the load rate threshold, the target level of the model for processing the command signal is determined to be the deep level; If the urgency is less than or equal to the urgency threshold, and the complexity is less than or equal to the complexity threshold, then determining that the target level of the model for processing the instruction signal is the shallow level.

3. The vehicle-mounted intelligent human-computer interaction method according to claim 1, characterized in that: The command signal includes a plurality of signals, and the model is used for: Obtaining the priority and issuing time of the command signal; The plurality of instruction signals are sorted according to the priority, the issuing time and the urgency, and the instruction signals are processed in sequence according to the sorting.

4. The vehicle-mounted intelligent human-computer interaction method according to claim 3, characterized in that: The method further comprises: Determining the instruction type of the instruction signal, different instruction types correspond to different priorities; Determine the priority corresponding to the instruction type, and obtain the priority of the instruction signal.

5. The vehicle-mounted intelligent human-computer interaction method according to claim 3, characterized in that: The model is also used to: obtaining the authority level of the occupant who issues the command signal; The plurality of instruction signals are sorted according to the authority level, the priority, the issuing time and the urgency, and the instruction signals are processed in sequence according to the sorting.

6. The vehicle-mounted intelligent human-computer interaction method according to claim 5, characterized in that: The method further comprises: Acquiring identity feature information of the occupant, wherein the identity feature information includes a voiceprint and a facial image; The authority level of the occupant is determined based on the identity feature information.

7. The vehicle-mounted intelligent human-computer interaction method according to claim 2, characterized in that: The instruction signal includes a first instruction signal with a complexity greater than the complexity threshold and a second instruction signal with a complexity less than or equal to the complexity threshold generated at the same time, and the model controlling the target level processes the instruction signal to obtain an interactive task output by the model, including: If the urgency is less than or equal to the urgency threshold, and the load rate is greater than the load rate threshold, the shallow level model is first controlled to process the second instruction signal to obtain the first interactive task output by the model, and then the deep level model is controlled to process the first instruction signal to obtain the second interactive task output by the model.

8. An in-vehicle intelligent human-computer interaction device, characterized in that: The device comprises: An acquisition module, used to acquire a command signal issued by a passenger on the vehicle and status information of the vehicle, wherein the command signal includes a voice signal, a gesture signal and an expression signal, and the status information includes working condition information, a driver's attention value and alarm information; A first determining module, used to determine the complexity of the instruction signal; A second determination module, configured to determine the urgency of the command signal according to the state information; A third determination module is used to determine a target level of a model for processing the instruction signal from models of multiple levels according to the complexity and the urgency, where different levels of models require different processor computing power when running, and the model is pre-trained through a training set; A control module is used to control the model of the target level to process the command signal, obtain the interactive task output by the model, and control the vehicle to perform the interactive task.

9. An electronic device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the method according to any one of claims 1 to 7 by executing the computer instructions.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Vehicle-mounted braille communication system and braille coding method

    CN120891946A