Vehicle interaction method, server and computer readable storage medium
By obtaining and integrating multimodal perceptual information in the car cockpit, determining the cockpit environment and safety risks, and generating corresponding control instructions, the existing system's insufficient understanding of complex environments is solved, and the cockpit safety and passenger experience are improved.
Patent Information
- Application Number
- CN202510171315.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-05-13
AI Technical Summary
The existing car cockpit safety system has insufficient understanding of the complex environment in the cockpit, resulting in limited safety risk identification capabilities and poor user experience.
By obtaining multimodal perceptual information in the vehicle cockpit, including in-vehicle visual information, in-vehicle audio information and vehicle status information, fusion processing and feature extraction are carried out to determine the cockpit environmental information and safety risks. According to the safety risk level, target measures are generated and vehicle control instructions are generated.
It effectively improves the safety of the cockpit, improves the passenger experience, and reduces accidents. Through real-time monitoring of multimodal perceived information and accurate assessment of security risks, the system can comprehensively and accurately identify potential security risks.
Smart Images

Figure CN119975219A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of vehicle interaction technology, and in particular to a vehicle interaction method, a server and a computer-readable storage medium. Background Art
[0002] In the related technologies, the car cockpit safety system mainly relies on single-modal data analysis and physical sensors to obtain data and identify safety risks. However, such a car cockpit safety system has insufficient understanding of the complex environment inside the car cockpit, limited ability to identify safety risks, and poor user experience. Summary of the invention
[0003] The present application provides a vehicle interaction method, a server and a computer-readable storage medium.
[0004] The present application provides a vehicle interaction method, the method comprising:
[0005] Acquiring multimodal perception information in a vehicle cabin, the multimodal perception information comprising at least two of in-vehicle visual information, in-vehicle audio information, and vehicle status information;
[0006] Determining security risks based on the multimodal perception information;
[0007] According to the security risks, determine the target measures for the security risks;
[0008] Based on the target measures, a vehicle control instruction is generated to complete the interaction.
[0009] In this way, the server obtains multimodal perception information in the vehicle cabin, and the multimodal perception information includes at least two of the in-vehicle visual information, in-vehicle audio information, and vehicle status information. Next, the server determines the safety risk based on the multimodal perception information. Then, the server determines the target measures for the safety risk based on the safety risk. Finally, the server generates vehicle control instructions based on the target measures to complete the interaction. In this way, by real-time monitoring of multimodal perception information, evaluating safety risks, and determining target measures for safety risks, the cabin safety can be effectively improved, the passenger experience can be enhanced, and the occurrence of accidents can be reduced.
[0010] In some embodiments, determining the security risk according to the multimodal perception information includes:
[0011] fusing the multimodal perception information to determine cockpit environment information;
[0012] The safety risk is determined according to the cabin environment information, where the safety risk includes vehicle safety risk and / or vehicle occupant safety risk.
[0013] In this way, the server fuses the multimodal perception information to determine the cabin environment information. Then, the server determines the safety risk based on the cabin environment information, and the safety risk includes the vehicle safety risk and / or the vehicle occupant safety risk. Then, the server constructs the first prompt word based on the first sub-prompt word. In this way, by fusing the multimodal perception information and performing a safety risk assessment based on the fused cabin environment information, potential safety risks can be comprehensively and accurately identified, thereby effectively improving cabin safety and preventing accidents.
[0014] In some embodiments, the fusing the multimodal perception information to determine the cabin environment information includes:
[0015] Performing time synchronization processing on the multimodal perception information to align the timestamps of the in-vehicle visual information and / or the in-vehicle audio information with the vehicle status information;
[0016] Performing feature extraction on the multimodal perception information after the time synchronization processing to determine visual feature information and / or audio feature information;
[0017] The cockpit environment information is determined according to the visual feature information and / or the audio feature information.
[0018] In this way, the server performs time synchronization processing on the multimodal perception information to align the timestamps of the in-vehicle visual information and / or in-vehicle audio information with the vehicle status information. Next, the server performs feature extraction on the multimodal perception information after time synchronization processing to determine the visual feature information and / or audio feature information. Finally, the server determines the cabin environment information based on the visual feature information and / or audio feature information. In this way, through time synchronization processing, the time difference between different modal data can be eliminated, thereby improving the accuracy of data fusion and enabling the system to accurately understand the cabin environment information. Through feature extraction, key features are extracted from the multimodal perception information to determine the cabin environment information, thereby comprehensively and accurately understanding the cabin environment, identifying potential safety risks, effectively improving cabin safety, and preventing accidents.
[0019] In some implementations, determining the safety risk according to the cabin environment information includes:
[0020] Based on the first preset model, determining safety hazards according to the cabin environment information;
[0021] Based on a second preset model, the security risk is determined according to the potential security risk.
[0022] In this way, based on the first preset model, the server determines the safety hazard according to the cabin environment information. Then, based on the second preset model, the server determines the safety risk according to the safety hazard. In this way, through the preset model, the safety hazard can be quantitatively evaluated, so as to accurately determine the safety risk level.
[0023] In some embodiments, determining the security risk based on the second preset model and according to the potential safety hazard includes:
[0024] Performing a first-dimensional assessment process on the potential safety hazard to determine a first-dimensional risk, where the first-dimensional risk represents the severity of the potential safety hazard;
[0025] Performing a second dimension assessment process on the potential safety hazard to determine a second dimension risk, where the second dimension risk represents the urgency of the potential safety hazard;
[0026] The security risk is determined according to the first dimensional risk and the second dimensional risk.
[0027] In this way, the server performs a first-dimensional assessment of the potential safety hazard to determine the first-dimensional risk, which characterizes the severity of the potential safety hazard. Next, the server performs a second-dimensional assessment of the potential safety hazard to determine the second-dimensional risk, which characterizes the urgency of the potential safety hazard. Finally, the server determines the security risk based on the first-dimensional risk and the second-dimensional risk. In this way, by performing a multi-dimensional assessment of the potential safety hazard, it is possible to fully understand the security risk and take effective preventive measures.
[0028] In some embodiments, determining a target measure for the security risk based on the security risk includes:
[0029] Based on a preset safety knowledge base, when the safety risk is at a first level, determining a first target measure according to the cockpit environment information;
[0030] In the case where the safety risk is at a second level, a second target measure is determined according to the cabin environment information, and the risk of the second level is greater than that of the first level.
[0031] In this way, based on the preset safety knowledge base, when the safety risk is at the first level, the server determines the first target measure based on the cabin environment information. Then, when the safety risk is at the second level, the server determines the second target measure based on the cabin environment information, and the risk of the second level is greater than the first level. In this way, different measures are taken according to different risk levels, which is in line with the actual situation, avoiding overreaction or underreaction, and thus providing a safe and comfortable driving experience.
[0032] In certain embodiments, the method further comprises:
[0033] Generate voice announcement information associated with the safety risk and announce it through the vehicle.
[0034] In this way, the server generates voice broadcast information related to safety risks and broadcasts it through the vehicle. In this way, by generating voice broadcast information related to safety risks and communicating and interacting with passengers, accidents can be effectively prevented, thereby providing a safe and comfortable driving experience.
[0035] In some embodiments, generating voice broadcast information associated with the security risk includes:
[0036] Determining a user's emotional state based on the in-car visual information and the in-car audio information;
[0037] The voice broadcast information is generated according to the emotional state of the user.
[0038] In this way, the server describes the visual information and audio information in the car to determine the user's emotional state. Then, the server generates voice broadcast information based on the user's emotional state. In this way, by identifying the emotional state of the passenger and taking corresponding measures, accidents can be effectively prevented, the safety of drivers and passengers can be improved, and the user experience can be improved.
[0039] An embodiment of the present application provides a server, which includes a processor and a memory, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the above-mentioned vehicle interaction method is implemented.
[0040] An embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the vehicle interaction method as described above are implemented.
[0041] Additional aspects and advantages of the embodiments of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the embodiments of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the description of the embodiments in conjunction with the following drawings, in which:
[0043] Figure 1 It is one of the flowcharts of the vehicle interaction method of certain embodiments of the present application;
[0044] Figure 2 This is a second flow chart of a vehicle interaction method according to certain embodiments of the present application;
[0045] Figure 3 This is a third flow chart of a vehicle interaction method according to certain embodiments of the present application;
[0046] Figure 4 This is a fourth flow chart of a vehicle interaction method according to certain embodiments of the present application;
[0047] Figure 5 This is a fifth flow chart of a vehicle interaction method according to certain embodiments of the present application;
[0048] Figure 6 This is the sixth flow chart of the vehicle interaction method of certain embodiments of the present application;
[0049] Figure 7 This is the seventh flow chart of the vehicle interaction method of certain embodiments of the present application;
[0050] Figure 8 This is the eighth flow chart of the vehicle interaction method of certain embodiments of the present application. DETAILED DESCRIPTION
[0051] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals represent the same or similar elements or elements having the same or similar functions from beginning to end. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the embodiments of the present application, and cannot be understood as limiting the embodiments of the present application.
[0052] In the current automotive cockpit safety systems, most systems mainly rely on single-modal data analysis, that is, obtaining data through physical sensors (such as cameras, radars, ultrasonic sensors, etc.) and identifying safety risks based on these data. However, this method that relies on single-modal data analysis and physical sensors has problems such as insufficient environmental understanding, limited safety risk identification capabilities, and poor user experience.
[0053] Insufficient environmental understanding means that due to relying solely on physical sensors, the systems often cannot fully understand the complex environment within the cabin. For example, they may not be able to accurately judge the emotional state, attention level or physiological condition of the passengers, which may affect driving safety.
[0054] Limited safety risk identification capability refers to the fact that single-modal data analysis may lead to insufficient identification of certain safety risks. For example, the system may not be able to effectively identify non-physical safety hazards, such as driver fatigue, drunk driving, or passengers not wearing seat belts.
[0055] Poor user experience means that due to the lack of in-depth understanding of the complex environment within the cockpit, the system may not be able to provide sufficiently accurate and personalized safety reminders and interactions, resulting in reduced user trust in the system, thus affecting the user experience.
[0056] Based on the above questions, please refer to Figure 1 , the embodiment of the present application provides a vehicle interaction method, the method comprising:
[0057] 01: Obtain multimodal perception information in the vehicle cabin;
[0058] 02: Determine security risks based on multimodal perception information;
[0059] 03: Determine the target measures for security risks based on security risks;
[0060] 04: Generate vehicle control instructions based on the target measures to complete the interaction.
[0061] The embodiment of the present application also provides a server, including a memory and a processor. The vehicle interaction method of the embodiment of the present application can be implemented by the server of the embodiment of the present application. Specifically, a computer program is stored in the memory, and the processor is used to obtain multimodal perception information in the vehicle cabin. And determine the safety risk based on the multimodal perception information. The processor is also used to determine the target measures for the safety risk based on the safety risk. And generate vehicle control instructions based on the target measures to complete the interaction.
[0062] The embodiments of the present application also provide a vehicle interaction device. The vehicle interaction method of the embodiments of the present application can be implemented by the vehicle interaction device of the embodiments of the present application. Specifically, the vehicle interaction device includes an acquisition module, a determination module and a generation module. The acquisition module is used to acquire multimodal perception information in the vehicle cabin. The determination module is used to determine safety risks based on the multimodal perception information. The determination module is also used to determine target measures for safety risks based on safety risks. The generation module is used to generate vehicle control instructions based on the target measures to complete the interaction.
[0063] Specifically, multimodal perception information refers to multiple types of data collected from the car cabin environment, including at least two of in-car visual information, in-car audio information, and vehicle status information. Visual data refers to the behavior, expression, and status of the driver and passengers captured by the on-board camera, such as face, gestures, and body posture. Sound data refers to in-car conversations and abnormal sounds captured by the microphone, such as voice, intonation, coughing, screaming, and glass breaking. Sensor data refers to vehicle status information collected by various sensors, such as speed, acceleration, steering angle, seat belt usage status, in-car temperature and humidity, etc.
[0064] It should be noted that the acquisition of multimodal perception information needs to be collected based on the existing equipment or permissions in the car cockpit. If certain equipment (such as infrared sensors and biometric sensors, etc.) does not exist in the vehicle, the corresponding information cannot be collected. If the vehicle system does not have permission to access certain sensor data (such as certain OBD data), only the data with permission can be used for security analysis.
[0065] Safety risk refers to the danger of possible injury to passengers or property loss caused by the cabin environment and user behavior. It is obtained by analyzing the multimodal perception information obtained by the multimodal large model, including driver driving behavior risks, passenger behavior risks, vehicle status risks and emergency risks.
[0066] A multimodal large model refers to a model that can process and analyze multiple types of data. It can process data in multiple dimensions such as vision, sound, and sensor data to obtain more comprehensive cockpit environment information and be used to identify and assess potential safety risks in the cockpit.
[0067] Targeted measures refer to specific actions taken automatically based on the safety risk level and the status of passengers and vehicles, aimed at reducing safety risks or responding to safety incidents, thereby improving cabin safety.
[0068] Vehicle control instructions refer to specific instructions generated based on target measures, which can guide the vehicle to perform specific operations or services to address safety risks.
[0069] The server collects at least two of the in-vehicle visual information, in-vehicle audio information and vehicle status information through on-board cameras, microphones, sensors and other devices. Multi-dimensional data collection ensures the system's comprehensive understanding of the in-vehicle environment.
[0070] Next, the multimodal large model is used to analyze the acquired multimodal perception information, identify potential safety hazards, and evaluate the safety risk level based on the identified safety hazards.
[0071] Then, according to the assessed safety risk level, the system automatically takes corresponding measures, including soothing measures and alarm measures. Soothing measures refer to low-risk situations, such as minor driving distraction, where the system may only provide gentle reminders. Alarm measures refer to high-risk situations, such as knife robbery, where the system automatically calls the police and notifies emergency contacts.
[0072] Finally, the server generates vehicle control instructions that can be understood and executed by the vehicle based on the decided target measures, and sends them to the vehicle to complete the interaction.
[0073] In summary, in the vehicle interaction method and server provided in the embodiments of the present application, the server obtains multimodal perception information in the vehicle cabin, and the multimodal perception information includes at least two of the in-vehicle visual information, in-vehicle audio information, and vehicle status information. Next, the server determines the safety risk based on the multimodal perception information. Then, the server determines the target measures for the safety risk based on the safety risk. Finally, the server generates a vehicle control instruction based on the target measures to complete the interaction. In this way, by real-time monitoring of multimodal perception information, evaluating safety risks, and determining target measures for safety risks, the cabin safety can be effectively improved, the passenger experience can be enhanced, and the occurrence of accidents can be reduced.
[0074] See also Figure 2 In some embodiments, step 02 (determining security risks based on multimodal perception information) includes:
[0075] 021: Fuse multi-modal perception information to determine the cockpit environment information;
[0076] 022: Determine safety risks based on cockpit environment information.
[0077] In some implementations, the determination module is used to perform fusion processing on the multi-modal perception information to determine the cabin environment information and determine the safety risk based on the cabin environment information.
[0078] In some implementations, the processor is further configured to perform fusion processing on the multimodal perception information to determine the cabin environment information and determine the safety risk based on the cabin environment information.
[0079] Specifically, fusion processing refers to fusing the acquired visual, sound and sensor data in a multimodal large model to obtain more comprehensive and accurate information to analyze the environment and user behavior in the cabin. In some embodiments, fusion processing includes feature extraction, data association and the use of deep learning models. Feature extraction refers to extracting key features from data of different modalities, such as facial features, voice intonation, etc. Data association refers to associating data of different modalities, such as associating vehicles in visual data with speeds in sensor data. Using a deep learning model refers to using a deep learning model for feature fusion, such as using a convolutional neural network (CNN) to process visual data and a recurrent neural network (RNN) to process sound data. Fusion processing is one of the core technologies of a multimodal large model. Fusion processing can effectively improve data utilization efficiency and information accuracy, thereby improving the performance and robustness of a multimodal large model.
[0080] The acquired in-car visual information, audio information and vehicle status information are integrated to determine more comprehensive and accurate cabin environment information. That is, key features are extracted from each modal data, such as facial features, behavioral features and expression features from visual data, and intonation, volume and abnormal sounds from audio data.
[0081] Next, based on the fused cockpit environment information, potential safety risks are identified, such as vehicle safety risks such as vehicle failure, collision risk and road hazards, or vehicle occupant safety risks such as drunk driving, fatigue driving and improper use of child safety seats.
[0082] In this way, by fusing multimodal perception information and performing safety risk assessment based on the fused cockpit environment information, potential safety risks can be comprehensively and accurately identified, thereby effectively improving cockpit safety and preventing accidents.
[0083] See also Figure 3 In some implementations, step 021 (fusion processing of multimodal perception information to determine cabin environment information) includes:
[0084] 0211: Perform time synchronization processing on multimodal perception information to align the timestamps of in-vehicle visual information and / or in-vehicle audio information with vehicle status information;
[0085] 0212: Extract features from the multimodal perception information after time synchronization processing to determine visual feature information and / or audio feature information;
[0086] 0213: Determine cabin environment information based on visual feature information and / or audio feature information.
[0087] In some embodiments, the determination module is further used to perform time synchronization processing on the multimodal perception information to align the timestamps of the in-vehicle visual information and / or in-vehicle audio information with the vehicle status information. And to perform feature extraction on the multimodal perception information after time synchronization processing to determine visual feature information and / or audio feature information. And to determine the cabin environment information based on the visual feature information and / or audio feature information.
[0088] In some embodiments, the processor is further configured to perform time synchronization processing on the multimodal perception information to align the timestamps of the in-vehicle visual information and / or in-vehicle audio information with the vehicle status information, and to perform feature extraction on the multimodal perception information after time synchronization processing to determine visual feature information and / or audio feature information, and to determine the cabin environment information based on the visual feature information and / or audio feature information.
[0089] Specifically, time synchronization processing refers to ensuring that visual data, sound data, and sensor data are synchronized in time for effective fusion analysis. In some embodiments, methods that can achieve time synchronization processing include timestamps, clock synchronization technology, and data synchronization algorithms. Timestamps refer to adding timestamps to each data point during data collection for time synchronization. Clock synchronization refers to the use of precise clock synchronization technology, such as the Network Time Protocol (NTP), to ensure that the time of different devices remains consistent. Data synchronization algorithms refer to the use of data synchronization algorithms, such as Kalman filters, to synchronize data in time. Through time synchronization processing, the accuracy of multi-dimensional data fusion and the real-time nature of security analysis are ensured.
[0090] Since different modal data acquisition devices may have time differences, time synchronization processing is required to ensure that the data of each modality is aligned in time. That is, when the in-car visual information and / or in-car audio information is collected, the timestamp information will be recorded and the in-car visual information and / or in-car audio information will be aligned with the timestamp of the sensor device.
[0091] Next, key features are extracted from the multimodal perception information after time synchronization, such as visual feature information such as facial features, behavioral features, expression features, and object features, or audio feature information such as intonation, volume, sound frequency, and keywords. Feature extraction can effectively extract key information from the data, such as facial features can be used to identify the identity of passengers, and behavioral features can be used to determine the status of passengers, thereby improving the accuracy and efficiency of recognition.
[0092] Finally, based on the extracted visual features and / or audio features, combined with sensor features, the cabin environment information is determined, such as passenger status, vehicle status, and road conditions. Passenger status includes whether the passenger is wearing a seat belt, whether the passenger is playing with a mobile phone, whether the passenger is tired, etc. Vehicle status includes whether there is a fault, whether the passenger has deviated from the lane, and whether the passenger is too close to the vehicle in front. Road conditions include road conditions, traffic conditions, weather conditions, etc.
[0093] For example, if the driver is driving fatigued and the vehicle's driving trajectory begins to deviate, the system will analyze the visual features "driver's facial mental state and vehicle's driving trajectory" and vehicle status information "speed and acceleration" to identify potential collision risks and take reminder measures, such as playing a warning sound and reminding the driver to take a rest.
[0094] In this way, through time synchronization processing, the time difference between different modal data can be eliminated, thereby improving the accuracy of data fusion and enabling the system to accurately understand the cockpit environment information. Through feature extraction, key features are extracted from multimodal perception information to determine the cockpit environment information, thereby fully and accurately understanding the cockpit environment and identifying potential safety risks, effectively improving cockpit safety and preventing accidents.
[0095] See also Figure 4 In some implementations, step 022 (determining safety risks based on cabin environment information) includes:
[0096] 0221: Based on the first preset model and according to the cockpit environment information, determine the safety hazard;
[0097] 0222: Based on the second preset model, determine the security risk according to the security hazard.
[0098] In some embodiments, the determination module is further configured to determine safety hazards based on the first preset model and the cabin environment information, and to determine safety risks based on the safety hazards based on the second preset model.
[0099] In some embodiments, the processor is further configured to determine safety hazards based on the first preset model and the cabin environment information, and to determine safety risks based on the safety hazards based on the second preset model.
[0100] Specifically, the first preset model is responsible for identifying potential safety hazards based on the cabin environment information. For example, it can identify signs of drunkenness by analyzing the driver's behavior and facial features, identify whether children are using safety seats correctly by analyzing the images in the car, and identify emergency situations through sound and image analysis.
[0101] The second preset model is responsible for assessing the security risk level based on the identified security risks. For example, according to the risk assessment model, the risk level is assessed as low risk or high risk, taking into account the severity and urgency of the security risks. The output of the first preset model is the input of the second preset model, and the two work together to complete the identification and assessment of security risks.
[0102] Safety hazards refer to factors that may threaten the safety of people in the cabin, including physical safety hazards such as vehicle failure, collision, fire, and non-physical safety hazards such as drunk drivers, passengers not wearing seat belts, improper use of child safety seats, and emergency situations. Through the multimodal large model, these safety hazards can be identified and evaluated in real time, and corresponding response measures can be automatically taken according to the safety level, thereby improving cabin safety and reducing accidents.
[0103] The first preset model is used to identify potential safety hazards based on the cabin environment information, such as driver conditions such as drunkenness, fatigue, and distraction, or passenger conditions such as not wearing seat belts and children not using safety seats correctly, and vehicle equipment conditions such as brake failure. In some embodiments, the first preset model can be a rule-based method or a machine learning method, such as a deep learning model.
[0104] Next, the second preset model is used to evaluate the safety risk level according to the identified safety hazards. In some embodiments, the second preset model may be based on a risk assessment model, such as a logistic regression model, a decision tree model, and the like.
[0105] In this way, through the preset model, it is possible to quantitatively evaluate safety hazards and accurately determine the level of safety risks.
[0106] See also Figure 5 In some embodiments, step 0222 (determining the security risk based on the second preset model and the security hazard) includes:
[0107] 02221: Conduct first-dimensional assessment and processing of safety hazards and determine first-dimensional risks;
[0108] 02222: Conduct second-dimensional assessment and processing of safety hazards and determine second-dimensional risks;
[0109] 02223: Determine security risks based on first dimension risks and second dimension risks.
[0110] In some embodiments, the determination module is further used to perform a first dimension assessment process on the safety hazard to determine the first dimension risk, and to perform a second dimension assessment process on the safety hazard to determine the second dimension risk, and to determine the safety risk based on the first dimension risk and the second dimension risk.
[0111] In some embodiments, the processor is further configured to perform a first dimension assessment process on the safety hazard to determine the first dimension risk, perform a second dimension assessment process on the safety hazard to determine the second dimension risk, and determine the safety risk based on the first dimension risk and the second dimension risk.
[0112] Specifically, the first dimension assessment process is responsible for assessing the severity of safety hazards, for example, the severity of drunk driving may be higher than that of minor driving distraction.
[0113] The second dimension assessment process is responsible for evaluating the urgency of the safety hazard. For example, a vehicle breakdown may require immediate parking, while a distracted driver may be able to be dealt with later under safe circumstances. The first dimension assessment process and the second dimension assessment process are two important dimensions of safety risk assessment, which respectively assess the severity and urgency of safety hazards, and together determine the final safety risk level.
[0114] It should be noted that the dimensional risk and security risk can be expressed in the form of scores or descriptions. When the dimensional risk is expressed as a score, the scores of the first dimensional risk and the second dimensional risk are combined to obtain the final security risk score, and the corresponding response measures are triggered according to the security risk score. High risk may trigger an emergency alarm, and low risk may trigger a gentle reminder. For example, the severity is divided into 5 points, 3 points and 1 point, and the urgency is divided into 5 points, 3 points and 1 point. When the severity is ≥4 and the urgency is ≥3, the security risk level is high risk. In other cases, the security risk level is low risk. In some embodiments, the scores of the first dimensional risk and the second dimensional risk can also be combined to obtain the final security risk score, and the weights can be adjusted according to the specific application scenario.
[0115] When the dimensional risk is expressed in the form of description, the first dimension risk and the second dimension risk are mapped. The first dimension risk is divided into "serious" and "minor", and the second dimension risk is divided into "urgent" and "non-urgent". When the first dimension risk is "serious" and the second dimension risk is "urgent", the security risk level is high risk, and other situations are low risk.
[0116] Conduct a first-dimensional assessment of the identified safety hazards, determine the first-dimensional risk, and characterize the severity of the safety hazards. The assessment can be based on the nature of the safety hazard, possible consequences, and other factors. For example, high severity includes drunk driving and vehicle failure. Low severity includes minor distracted driving and children not wearing seat belts.
[0117] Next, the identified safety hazards are evaluated in the second dimension to determine the second dimension risk and characterize the urgency of the safety hazard. The evaluation can be based on factors such as the timing of the occurrence and the speed of development of the safety hazard. For example, a high urgency level includes an imminent vehicle collision, while a low urgency level includes signs of driver fatigue, etc.
[0118] Finally, combine the first dimension risk and the second dimension risk to determine the safety risk and take corresponding response measures. For example, if the severity of drunk driving is 5 points and its urgency is 5 points, then the safety risk level is high risk.
[0119] In the case of high safety risk, such as drunk driving or vehicle breakdown, the system will take emergency braking, alarm and other measures. In the case of low safety risk, such as slight driving distraction, the system will play a warning sound and remind the driver to pay attention to safety.
[0120] In this way, by conducting a multi-dimensional assessment of safety hazards, we can fully understand the safety risks and take effective preventive measures.
[0121] See also Figure 6 In some implementations, step 03 (determining target measures for security risks based on security risks) includes:
[0122] 031: Based on the preset safety knowledge base, when the safety risk is at the first level, determine the first target measure according to the cockpit environment information;
[0123] 032: When the safety risk is at the second level, determine the second target measure based on the cockpit environment information.
[0124] In some embodiments, the determination module is further configured to determine, based on a preset security knowledge base, a first target measure according to the cabin environment information when the security risk is at a first level, and a second target measure according to the cabin environment information when the security risk is at a second level.
[0125] In some embodiments, the processor is further configured to determine, based on a preset security knowledge base, a first target measure according to the cabin environment information when the security risk is at a first level, and to determine a second target measure according to the cabin environment information when the security risk is at a second level.
[0126] Specifically, the preset safety knowledge base refers to a database that includes various safety knowledge and countermeasures, which is used to guide the system to conduct safety risk assessment and response measures decisions, including safety risk types, corresponding safety risk levels, and countermeasures for different safety risk types and risk levels. Safety risk types include driver fatigue, passengers not wearing seat belts, vehicle failure, emergency situations, etc. Risk levels include driver fatigue as low risk, passengers not wearing seat belts as low risk, and vehicle failure as high risk. Countermeasures refer to providing corresponding countermeasure suggestions for different levels of safety risks, including reminders, alarms, emergency contact notifications, and safe area recommendations.
[0127] Level 1 refers to one level in the security risk level, representing low risk.
[0128] Level 2 refers to another level in the security risk level and represents high risk.
[0129] The first target measure refers to the response to the first level of safety risks. For example, if the driver is slightly distracted, the system may issue a gentle voice reminder to remind the driver to pay attention.
[0130] Secondary target measures refer to the response measures taken for the second level of safety risks. For example, if the driver is drunk driving, the system may issue an emergency alarm, automatically contact emergency contacts, and park the vehicle in a safe area.
[0131] The preset safety knowledge base, the first level, the second level, the first target measure and the second target measure are interrelated to jointly ensure cockpit safety.
[0132] The system first uses a multimodal large model to analyze the cabin environment information, identify potential safety hazards, and conduct risk assessment to determine the safety risk level.
[0133] Next, the target measures are determined. According to the safety risk level, the system selects the corresponding target measures from the preset safety knowledge base. The safety risk level is the first level risk, which corresponds to a lower risk, such as slight driving distraction. The system may take gentle reminder measures, such as voice prompts or slight vibration of the seat. The safety risk level is the second level risk, which corresponds to a higher risk, such as drunk driving. The system may take stronger measures, such as calling the police and notifying emergency contacts, or even automatically controlling the vehicle to slow down or stop.
[0134] For example, the driver is slightly distracted, such as looking down at a mobile phone. The system recognizes distracted driving and determines it as level 1 based on the risk level. The system issues a voice prompt: "Please pay attention to driving." Or, the driver is suspected of drunk driving, such as blurred eyes and abnormal expression. The system recognizes signs of drunkenness and determines it as level 2 based on the risk level. The system issues an emergency alarm and automatically contacts emergency contacts, while controlling the vehicle to slow down and stop safely.
[0135] In this way, different measures are taken according to different risk levels, in line with the actual situation, avoiding overreaction or underreaction, thereby providing a safe and comfortable driving experience.
[0136] See also Figure 7 In some embodiments, the method further comprises:
[0137] 05: Generate voice broadcast information associated with safety risks and broadcast it through the vehicle.
[0138] In some embodiments, the vehicle interaction device further includes a generation module, which is further configured to generate voice broadcast information associated with the safety risk and broadcast the information through the vehicle.
[0139] In some embodiments, the processor is further configured to generate voice announcement information associated with the safety risk and announce the information through the vehicle.
[0140] Specifically, voice broadcast information refers to the use of the vehicle's voice broadcast function to deliver information and prompts related to safety risks to the driver or passengers. For example, when the system detects that the driver is driving fatigued, it can broadcast: "You have been driving for a while, please take a break." Or, when the system detects that the vehicle has a malfunction, it can broadcast: "Your vehicle has a malfunction, please stop and check immediately." Or, when the system detects an emergency, such as a car accident, it can broadcast: "Please stay calm, don't panic, we will help you."
[0141] The system selects appropriate voice broadcast information from the preset safety knowledge base based on the identified safety risks. The system then broadcasts the generated voice information to the passengers through the vehicle's voice broadcast system.
[0142] In this way, by generating voice broadcast information related to safety risks and communicating and interacting with passengers, accidents can be effectively prevented, thereby providing a safe and comfortable driving experience.
[0143] See also Figure 8 In some embodiments, step 05 (generating voice broadcast information associated with security risks) includes:
[0144] 051: Determine the user's emotional state based on the in-car visual information and in-car audio information;
[0145] 052: Generate voice broadcast information based on the user's emotional state.
[0146] In some embodiments, the generating module is further used to determine the user's emotional state based on the in-car visual information and the in-car audio information, and to generate voice broadcast information based on the user's emotional state.
[0147] In some embodiments, the processor is further configured to determine the user's emotional state based on the in-vehicle visual information and the in-vehicle audio information, and to generate voice broadcast information based on the user's emotional state.
[0148] Specifically, the system uses in-car cameras and microphones to collect passengers' visual and audio information, such as facial expressions, voice intonation, etc., and analyzes passengers' emotional states through a large multimodal model, including anxiety, irritability, fatigue, and excitement.
[0149] Then, the system generates appropriate voice broadcast information based on the identified passenger's emotional state and the preset safety knowledge base. For example, if the passenger's emotional state is anxious, the system will use a relaxed tone to say "Relax, park the car on the side of the road, and listen to some music."
[0150] The system then broadcasts the generated voice information to the passengers through the vehicle's voice broadcast system.
[0151] In this way, by identifying the emotional state of passengers and taking corresponding measures, accidents can be effectively prevented, the safety of drivers and passengers can be improved, and the user experience can be enhanced.
[0152] The present application also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of the vehicle interaction method described above are implemented.
[0153] It is understood that a computer program includes computer program code. The computer program code may be in source code form, object code form, executable file or some intermediate form. Computer readable storage media may include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), and software distribution medium.
[0154] In the description of this specification, the descriptions with reference to the terms "specifically", "further", "particularly", "understandably", etc. are intended to mean that the specific features, structures, materials or characteristics described in conjunction with the embodiments or examples are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms are not intended to refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, unless they are contradictory.
[0155] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code that includes one or more executable requests for implementing specific logical functions or steps of a process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present application belong.
[0156] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limitations on the present application. Ordinary technicians in this field can change, modify, replace and modify the above embodiments within the scope of the present application.
Claims
1. A vehicle interaction method, characterized in that: The method comprises: Acquiring multimodal perception information in a vehicle cabin, the multimodal perception information comprising at least two of in-vehicle visual information, in-vehicle audio information, and vehicle status information; Determining security risks based on the multimodal perception information; According to the security risks, determine the target measures for the security risks; Based on the target measures, a vehicle control instruction is generated to complete the interaction.
2. The method according to claim 1, characterized in that The determining of the security risk according to the multimodal perception information includes: fusing the multimodal perception information to determine cockpit environment information; The safety risk is determined according to the cabin environment information, where the safety risk includes vehicle safety risk and / or vehicle occupant safety risk.
3. The method according to claim 2, characterized in that The fusing and processing the multimodal perception information to determine the cabin environment information includes: Performing time synchronization processing on the multimodal perception information to align the timestamps of the in-vehicle visual information and / or the in-vehicle audio information with the vehicle status information; Performing feature extraction on the multimodal perception information after the time synchronization processing to determine visual feature information and / or audio feature information; The cockpit environment information is determined according to the visual feature information and / or the audio feature information.
4. The method according to claim 2, characterized in that: The determining the safety risk according to the cabin environment information includes: Based on the first preset model, determining safety hazards according to the cabin environment information; Based on a second preset model, the security risk is determined according to the potential security risk.
5. The method according to claim 4, characterized in that The determining the security risk based on the second preset model and according to the potential safety hazard includes: Performing a first-dimensional assessment process on the potential safety hazard to determine a first-dimensional risk, where the first-dimensional risk represents the severity of the potential safety hazard; Performing a second-dimensional assessment process on the potential safety hazard to determine a second-dimensional risk, where the second-dimensional risk represents the urgency of the potential safety hazard; The security risk is determined according to the first dimensional risk and the second dimensional risk.
6. The method according to claim 5, characterized in that Determining target measures for the security risk based on the security risk includes: Based on a preset safety knowledge base, when the safety risk is at a first level, determining a first target measure according to the cockpit environment information; In the case where the safety risk is at a second level, a second target measure is determined according to the cabin environment information, and the risk of the second level is greater than that of the first level.
7. The method according to claim 1, characterized in that The method further comprises: Generating voice broadcast information associated with the safety risk and broadcasting it through the vehicle.
8. The method according to claim 7, characterized in that The generating voice broadcast information associated with the safety risk and broadcasting it through the vehicle includes: Determining a user's emotional state based on the in-car visual information and the in-car audio information; The voice broadcast information is generated according to the emotional state of the user.
9. A server, characterized in that: The server includes a processor and a memory, wherein a computer program is stored in the memory. When the computer program is executed by the processor, the method according to any one of claims 1 to 8 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.