Vehicle control method, vehicle and electronic equipment
By using in-vehicle terminals and multimodal large models to identify the responsiveness and accident status of people in the vehicle, rescue suggestions are automatically generated, solving the problem of insufficient information in the eCall system and improving the efficiency and accuracy of emergency rescue.
Patent Information
- Application Number
- CN202511992076.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-26
- Publication Date
- 2026-02-24
AI Technical Summary
The existing eCall system cannot provide detailed real-time information after an accident, especially when the driver and passengers are disabled due to the accident, it cannot effectively convey the accident situation, affecting rescue efficiency and decision-making quality.
After the emergency rescue system is triggered by the vehicle terminal, the system acquires the capability characteristics of the people in the vehicle, uses a pre-trained multimodal large model to identify the people's response capabilities, and automatically acquires vehicle environmental data when people cannot respond, infers the accident status, generates rescue suggestions or response information, converts it into voice data and feeds it back to the rescue party; when people can respond, the emergency call function is activated.
It improves communication efficiency and information accuracy in emergency rescue, shortens rescue response time, reduces the risk of decision-making delays, alleviates the burden on rescuers, and ensures a balance between human-centered and system-assisted approaches.
Smart Images

Figure CN121565137A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of vehicle control technology, and in particular to a vehicle control method, a vehicle, and electronic equipment. Background Technology
[0002] In modern automobiles, with the development of intelligence and connectivity, in-vehicle communication terminals (Telematics Box, or Tbox for short) are gradually becoming standard equipment. As the core communication module of intelligent connected vehicles, the Tbox also integrates the eCall emergency call function, making it a crucial component for ensuring passenger safety. The eCall system is a vehicle emergency assistance system. Traditional eCall systems rely on the Tbox to automatically collect a Minimum Set of Data (MSD) after an accident, including GPS location, Vehicle Identification Number (VIN), and accident time, and transmit this data to the rescue center via 4G / 5G networks. Simultaneously, the system establishes a voice channel to enable two-way voice communication between the driver and the rescue center.
[0003] However, existing eCall systems still have certain technical limitations in practical applications. The types of data they automatically upload are limited, basically confined to fixed fields such as GPS coordinates, VIN codes, and accident time, failing to provide more detailed real-time information such as the vehicle's environment and occupant status. Especially when the driver or occupants are incapacitated due to an accident, the inability to effectively communicate the accident situation via voice makes it difficult for rescuers to comprehensively and accurately assess the scene, impacting rescue efficiency and decision-making quality. Therefore, existing eCall systems still need improvement in terms of information richness and scenario adaptability. Summary of the Invention
[0004] This application addresses, to at least some extent, one of the technical problems in the related art.
[0005] In a first aspect, embodiments of this application provide a vehicle control method applied to an in-vehicle terminal, wherein the in-vehicle terminal is communicatively connected to an in-vehicle communication terminal, and the method includes: When the vehicle-mounted communication terminal detects that the emergency rescue system automatically triggers an emergency call after an accident, it obtains the capability characteristics of the people in the vehicle. The system acquires the rescuer's voice data in real time and calls a pre-trained multimodal large model to respond to the rescuer's voice data. The use of the multimodal large model to respond to the rescuer's voice data includes: Based on the aforementioned capability characteristics, the responsiveness of the people on board the vehicle can be identified; If the occupants are unable to respond, the accident status is inferred and identified after acquiring vehicle environmental data, and rescue suggestions and / or response information are generated based on the accident status. The rescue suggestions and / or the response information are converted into audio data, and the dialogue is conducted with the rescuer through the emergency rescue system based on the audio data.
[0006] Conversely, if the people in the vehicle are able to respond, the emergency call function can be used to allow them to communicate directly with the rescuers.
[0007] Based on the above steps, the control method of this application improves the vehicle's perception and intelligent processing capabilities. When an emergency call is detected, the method acquires the capability characteristics of the occupants, automatically invokes a pre-trained multimodal large model, receives and understands the voice data from the rescuer, and comprehensively judges whether the occupants have normal response capabilities. If it is identified that the occupants cannot respond effectively, environmental data (such as collision, location, etc.) is further collected from the vehicle's sensor network, and the possible accident state is judged through reasoning. Based on this reasoning result, semantic information including rescue suggestions and / or response information is automatically generated and converted into voice format, which is fed back to the rescuer through the emergency rescue system. If the recognition result indicates that the occupants have response capabilities, the vehicle's emergency communication module is directly activated, enabling them to communicate with the rescuer via voice to assist in handling the emergency situation.
[0008] The multimodal large model, fine-tuned using the Low-Rank Adaptation (LoRA) method, is pre-deployed on the vehicle-mounted terminal. By combining the multimodal large model with the vehicle-mounted sensor network, intelligent judgment and interactive response to the state of personnel in emergency situations are achieved, improving the efficiency of rescue communication when occupants are unable to express themselves. This embodiment can automatically determine the response method based on the actual situation, thereby shortening the rescue response time, improving the accuracy and usability of rescue information, and reducing the risk of delayed rescue decisions. Simultaneously, automatic voice feedback reduces the burden on rescuers in obtaining information, helping them make targeted rescue decisions in the first instance.
[0009] Furthermore, if the occupants are capable of responding, their direct communication with the rescue team will not be interfered with, ensuring a balance between human-centered design and system assistance.
[0010] In some embodiments, the capability characteristics include: physiological data, appearance data, and / or voice data, and identifying the responsiveness of on-board personnel based on the capability characteristics further includes: Facial features of people in the vehicle are identified based on the apparent data; Identify the physical condition of passengers based on the aforementioned physiological data; Based on the sound data, it can be identified whether the people in the vehicle are making voice outputs; Determine whether the facial features, body state, and voice output meet the preset response capability conditions. If they do, determine that the person on the vehicle is unable to respond.
[0011] Based on the above configuration, this embodiment of the application refines the capability characteristics of the occupants into three categories: physiological data, appearance data, and voice data, and comprehensively judges the occupants' physical state, external performance, and voice behavior, respectively. After integrating the three types of information, it is compared with preset response capability conditions. If the conditions are met, it is determined that the occupants do not have the ability to respond, and thus the occupants communicate with the rescue team by voice on their behalf. The multi-dimensional comprehensive recognition method helps to achieve stable assessment even in complex accident environments, with high noise interference, or poor lighting conditions, improving the reliability of the judgment, reducing the false judgment rate, and enhancing the robustness in real-world application scenarios.
[0012] In some embodiments, LoRA technology is used to fine-tune the multimodal model. During fine-tuning, a multimodal training set is constructed, containing appearance data, physiological data, sound data, and corresponding labels (such as "pain," "closed eyes," "abnormal heart rate," "no speech output," etc.). The LoRA mechanism is used to perform low-rank adjustment on the key weight matrix in the multimodal model parameters, significantly reducing the number of training parameters and resource consumption, improving training efficiency, and enhancing the model's recognition ability under specific tasks. After fine-tuning, the model can accurately recognize facial features, body states, and the presence or absence of speech output of people in the vehicle.
[0013] In some embodiments, when an emergency call is detected by the emergency rescue system, the capability characteristics of the occupants are acquired, including: when an emergency call is detected by the emergency rescue system, a driver detection system is activated, which calls the vehicle sensor system to collect the physiological data, the appearance data, and the voice data.
[0014] Based on the above configuration, the emergency assistance system and the driver detection system work together to automate the collection and analysis of the occupants' capability characteristics. By automatically activating the driver detection system the moment an emergency call is triggered, the timeliness and comprehensiveness of occupant status perception are effectively improved.
[0015] In some embodiments, the physiological data includes body temperature characteristics, heart rate, etc. The physiological data of this application can monitor the breathing rate and heart rate of the people in the vehicle through electromagnetic wave reflection, which is suitable for detecting abnormal heartbeat and / or breathing caused by accidents. It can also be obtained directly by calling the infrared camera of the vehicle's sensing system, or it can be based on indirect methods such as subtle changes in facial skin color to assist in the judgment.
[0016] At this point, the infrared camera can be deployed above the steering wheel, dashboard, or inside the A-pillar to capture facial infrared images, ensuring clear facial thermal images under different sitting postures and lighting conditions, and operating effectively at night or in low-light environments, avoiding ambient light interference. The infrared camera can work in low light (such as at night) to avoid ambient light interference.
[0017] In some embodiments, the appearance data is used to display facial expressions, eye states, and head posture. The eye states include open and closed eyes. The appearance data in this application can be directly acquired through the optical camera of the vehicle sensing system.
[0018] In some embodiments, the sound data is acquired through the audio acquisition unit of the vehicle sensing system.
[0019] In some embodiments, the preset response capability conditions are configured as follows: if the facial features include pain features, closed eyes features and no voice output, and at the same time, the physical state includes abnormal body temperature features or abnormal heart rate features, then it is determined that the person on the vehicle has lost effective response capability, and the multimodal big data model automatically responds to generate rescue information.
[0020] Based on the above implementation methods, the embodiments of this application achieve highly sensitive and accurate abnormal state recognition by setting abnormal combinations, and perform semantic-level judgment by combining large models, thereby improving the ability to understand and reason about complex interactive information and effectively reducing erroneous judgments caused by perception blind spots or one-dimensional misjudgments.
[0021] This method enables rapid decision-making in time-sensitive accident scenarios, providing timely and reliable data support for subsequent rescue efforts while maintaining the intelligence and controllability of the system response.
[0022] In some embodiments, the step of inferring and identifying the accident state after acquiring vehicle environmental data includes: The vehicle sensing system is invoked to collect images of the road ahead and vehicle collision data. The multimodal large model analyzes the road image ahead of the vehicle and the vehicle collision data to identify whether a collision has occurred. If so, it identifies the collision location and the severity of the collision, and outputs collision information including the collision location and the severity of the collision.
[0023] This solution significantly enhances a vehicle's ability to understand changes in the external environment during sudden accidents. Compared to traditional sensor threshold judgment methods, this method uses a large model to deeply understand the semantic relationship between road images and vehicle body data, effectively reducing false triggering rates and the risk of missed detections, and is particularly accurate in ambiguous scenarios such as minor scrapes and offset collisions.
[0024] LoRA technology is used to fine-tune the process of acquiring collision information for the multimodal large model. A multimodal training set covering a variety of typical collision scenarios is constructed. This set includes images of the road in front of the vehicle taken by the forward-looking camera, vehicle collision data such as acceleration direction and deformation depth, and pre-labeled accident results, including whether there is a collision, the collision location and the severity of the collision, such as "moderate in front" and "minor on the right".
[0025] During fine-tuning, LoRA was used to inject low-rank parameters into the key attention layer and visual embedding layer within the model, optimizing only the subset of parameters related to collision recognition. This retained the original model's general capabilities while endowing it with fine-grained judgment capabilities for collision recognition. After fine-tuning, the model possesses stronger image-sensor data collaborative reasoning capabilities, enabling it to quickly identify collision situations and their detailed states under complex on-site conditions, for use in subsequent rescue response generation and information transmission.
[0026] In some embodiments, the inference and identification of the accident state after acquiring vehicle environmental data includes: The vehicle sensing system is invoked to acquire images of the area surrounding the vehicle body; The multimodal large model extracts road features, terrain features, and obstacle features around the vehicle based on the images around the vehicle. Based on the road features, it identifies the distance to the road edge, the angle and displacement of the vehicle from the road, the terrain features identify the type of terrain the vehicle is in, and the obstacle features identify the types of surrounding obstacles and whether there are safety protection facilities. Based on the distance to the road edge, the angle and displacement of the vehicle from the road, the type of terrain the vehicle is in, the types of surrounding obstacles, and the presence of safety protection facilities, it obtains vehicle scene information, thereby determining whether the vehicle has run off the road or is in a dangerous area, such as near a cliff or river.
[0027] Based on the above steps, multi-dimensional data fusion analysis effectively identifies accident types such as vehicles deviating from their path, running off the road, or entering dangerous areas due to loss of driving control, collision impact, or poor road conditions. The introduction of vehicle scene information not only improves the comprehensiveness of accident assessment but also helps in the rapid adjustment of rescue plans and route planning, such as prioritizing the deployment of hoisting equipment and bypassing dangerous roads.
[0028] To achieve the aforementioned comprehensive scene understanding capability, the LoRA fine-tuning method was used to enhance the training of the multimodal large model. Specific steps included constructing a training dataset containing multiple types of road scenes. Data samples included images of the vehicle's surroundings, corresponding vehicle trajectory and pose data, and labeled road types, terrain attributes, obstacle types, and distribution information. During training, a structured prompt template was used to pair image information with labels and input them into the model. For example: "The vehicle in the image is located on a slope, with a road edge to the left and a metal guardrail to the right, with an 8-degree deviation angle. There are no obvious obstacles in the scene." The multimodal large model uses these samples to jointly model semantic labels and spatial data. The LoRA module performs low-rank matrix insertion on key attention layers and trains only the inserted parameters, thereby endowing it with specific recognition capabilities for features such as roads, terrain, and obstacles without compromising the original model's capabilities. After fine-tuning, the model can automatically infer the comprehensive environmental state of the vehicle in real-world in-vehicle scenarios, such as road edge lines, lane line shapes, unpaved surfaces, terrain contours, slopes, and potential obstacles such as trees, guardrails, rocks, and water bodies, providing crucial information for subsequent rescue strategy generation and accident assessment.
[0029] In some embodiments, the inference and identification of the accident state after acquiring vehicle environmental data includes: The vehicle sensing system is invoked to collect images of the vehicle cabin and acquire the sound data. The multimodal large model extracts the bleeding features of the personnel in the vehicle cabin image and the sound content in the sound data, identifies the bleeding features and the sound content, and outputs the personnel injury information.
[0030] Based on the above steps, the multimodal large model of this application receives and processes the aforementioned vehicle cabin image and sound data. It extracts visual features from key areas such as the face, arms, clothing, and seat area in the images to identify the presence of bloodstains, obvious wounds, blood-stained clothing, and other signs of bleeding. Simultaneously, it analyzes the language content in the sound data, including non-verbal signals such as cries for help, groans, and intermittent speech, fusing and inferring whether the occupants are currently injured, the severity of their injuries, and whether they are conscious, and outputs structured information about their injuries. For example, "The passenger on the left is bleeding significantly, accompanied by continuous groans; the injury may be serious."
[0031] This application significantly improves the accuracy and timeliness of occupant injury identification after an accident by introducing joint analysis of vehicle cabin image and sound data. Through a multimodal large-scale model, it automatically determines whether an occupant is injured and the severity of their injury. Especially when occupants are unable to express themselves, the system can still proactively initiate effective intelligent responses and rescue requests.
[0032] When constructing the training dataset for the multimodal large model, multimodal samples containing real or simulated post-accident images inside the vehicle, occupant injury audio recordings, and their annotation information are collected. The image portion should include samples with bleeding characteristics under different genders, clothing, locations, and postures. The audio portion includes speech content under different tones, emotions, and background noise. Labels include "bleeding," "call for help," and "state of consciousness." During the fine-tuning phase, input data is organized using prompt templates, for example: "The image shows bloodstains on the occupant's face and clothing, and intermittent calls for help are detected in the audio. Is there a serious injury?" The model inserts a low-rank weight matrix using a LoRA structure, adjusting only some parameters of the key visual semantic layer and speech understanding module to quickly adapt to the injury recognition task requirements. After fine-tuning, the model possesses the ability to efficiently and accurately identify injury features and generate rescue decision-making criteria in complex in-vehicle environments.
[0033] This application employs the LoRA fine-tuning method, which not only ensures high model training efficiency and low deployment cost, but also possesses strong task adaptability and scalability, making it suitable for various vehicle models and in-vehicle configurations. Furthermore, the structured output injury information can be directly used for subsequent rescue command generation, resource allocation, and medical response decisions, constructing a complete closed loop from perception to action.
[0034] In some embodiments, the inference and identification of the accident state after acquiring vehicle environmental data includes: The collision information, vehicle scene information, and personnel injury information, or any combination thereof, are input into the multimodal large model to infer the vehicle accident state. The vehicle accident state includes the vehicle location, number of people in the vehicle, personnel status, vehicle collision state, and vehicle scene state. The vehicle location is determined with the assistance of vehicle GPS data.
[0035] To achieve this information fusion and reasoning capability, this application employs LoRA fine-tuning to enhance the multimodal large model's ability to model the semantic combination and logical relationships of accidents. During training, cross-modal integrated training samples are constructed. Each sample contains collision information, vehicle scene information, and the aforementioned personnel injury information, with the corresponding output being a structured accident state label. Data is organized using a prompt template, such as: "Input: Severe collision, road edge close ahead, occupant unconscious, GPS X; Output: Vehicle status: off-road, severe frontal collision, 2 people inside, 1 unconscious." LoRA is used to insert low-rank weight modules only into the model's key fusion layer and task output layer, rapidly transferring and enhancing the ability to reason about accident state combinations. After training, the model can quickly infer a complete accident scene based on partial or complete input information, providing a complete data foundation for emergency communication, rescue resource scheduling, and other applications.
[0036] This application, based on multimodal information fusion, generates a comprehensive vehicle accident status through reasoning, providing crucial support for subsequent emergency response and information exchange. Compared to traditional rule-matching-based analysis methods, this approach offers greater flexibility and scalability, and can handle scenarios with missing, incomplete, or abnormal information.
[0037] In some embodiments, the vehicle sensing system is equipped with a front-view camera, a surround-view camera, and an in-vehicle camera. The front-view camera is used to acquire images of the road ahead of the vehicle to determine road conditions and obstacle information. The surround-view camera is used to acquire images of the area around the vehicle to determine the vehicle's position and surrounding environmental features. The in-vehicle camera is used to acquire images of the vehicle's cabin to identify the number of occupants and their injury status, etc. Vehicle sensing data is acquired through vehicle collision sensors, speed sensors, etc.
[0038] Based on this, the embodiments of this application achieve high-precision perception of the vehicle's surrounding environment and internal state, effectively improving the ability to identify accident scenarios. The standardized installation positions of the cameras ensure the integrity and stability of image acquisition, guaranteeing continuous operation in various driving and collision environments. The collaborative processing of surround-view images and cockpit images not only determines whether the vehicle is in a dangerous area but also infers the number of occupants and potential injuries, enhancing the comprehensiveness and personalization of emergency response. Sensor deployment supplements the shortcomings of image recognition, forming an information loop and improving the accuracy and real-time performance of reasoning.
[0039] In some embodiments, the rescue recommendations include at least: medical rescue recommendations, safety warning recommendations, and traffic control recommendations.
[0040] This application embodiment utilizes a multimodal large model to perform scene understanding on (inside and outside the vehicle) images, voice, and vehicle data, analyze the situation inside and outside the vehicle, provide on-site situation response content, communicate with humans, avoid rescue difficulties due to the incapacity of drivers or passengers, and improve rescue efficiency.
[0041] Secondly, this application also provides a vehicle, including a vehicle-mounted communication terminal and a vehicle-mounted terminal connected in communication. The vehicle-mounted communication terminal is equipped with an emergency rescue system or is connected in communication with an emergency rescue system. The vehicle-mounted terminal is used to execute the control method described in the first aspect above.
[0042] Thirdly, this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the control method described in the first aspect above. As can be seen from the above technical solutions, additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description
[0043] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments of this application and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 This is a flowchart illustrating the control method according to an embodiment of this application; Figure 2 This is a step-by-step flowchart of the control method according to an embodiment of this application; Figure 3 This is a structural block diagram of a vehicle according to an embodiment of this application; Figure 4 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of this application.
[0044] 40. Bus; 41. Processor; 42. Memory; 43. Communication interface. Detailed Implementation
[0045] To make the objectives, technical solutions, and advantages of this application clearer, the application is described and illustrated below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application. All other embodiments obtained by those skilled in the art based on the embodiments provided in this application without inventive effort are within the scope of protection of this application.
[0046] Obviously, the accompanying drawings described below are merely some examples or embodiments of this application. Those skilled in the art can apply this application to other similar scenarios based on these drawings without any inventive effort. Furthermore, it is understood that although the efforts made in this development process may be complex and lengthy, for those skilled in the art related to the content disclosed in this application, any changes to design, manufacturing, or production based on the technical content disclosed in this application are merely conventional technical means and should not be construed as insufficient disclosure of the content of this application.
[0047] In this application, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment that is mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described in this application may be combined with other embodiments without conflict.
[0048] Unless otherwise defined, the technical or scientific terms used in this application shall have the ordinary meaning understood by one of ordinary skill in the art to which this application pertains. The terms “a,” “an,” “an,” “the,” and similar words used in this application do not indicate quantity limitation and may indicate singular or plural. The terms “comprising,” “including,” “having,” and any variations thereof used in this application are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that includes a series of steps or modules (units) is not limited to the listed steps or units, but may also include steps or units not listed, or may include other steps or units inherent to these processes, methods, products, or devices. The terms “connected,” “linked,” “coupled,” and similar words used in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. “Multiple” used in this application refers to two or more. “And / or” describes the relationship between related objects, indicating that three relationships may exist; for example, “A and / or B” can represent: A alone, A and B simultaneously, and B alone. The character " / " generally indicates that the preceding and following objects are in an "or" relationship. The terms "first," "second," and "third" used in this application are merely to distinguish similar objects and do not represent a specific ordering of the objects.
[0049] Figure 1 This is a flowchart of a control method according to an embodiment of this application, such as... Figure 1 As shown, this method is applied to an in-vehicle terminal, which is communicatively connected to an in-vehicle communication terminal. The in-vehicle communication terminal is equipped with an emergency rescue system or is communicatively connected to an emergency rescue system. After an accident occurs, the emergency rescue system automatically triggers an emergency call. The control method of this embodiment includes the following steps: Step S100: When an emergency call is detected by the emergency rescue system, acquire the capability characteristics of the people on board; Step S200: Acquire the rescuer's voice data in real time and call the pre-trained multimodal large model to respond to the rescuer's voice data; Optionally, the multimodal large model is based on Qwen 2.5 VLM 7B, fine-tuned using the LoRA parameter efficiency fine-tuning method, and pre-deployed on the vehicle terminal. It can be locally optimized according to actual deployment needs to adapt to different computing power requirements.
[0050] This application focuses on emergency rescue communication in a Chinese speech environment, involving professional rescue terminology, colloquial expressions, and emotional language input. The Qwen series models have advantages in training with Chinese corpora, enabling them to more accurately understand rescuers' voice commands and generate response information that conforms to rescue procedures and communication habits. Furthermore, Qwen 2.5 VLM 7B achieves a good balance between parameter size, inference performance, and computational resource consumption. Compared to larger-scale models, this model can meet the real-time and stability requirements after an accident while ensuring semantic understanding and generation quality.
[0051] Figure 2 This is a step-by-step flowchart of the control method according to an embodiment of this application, with reference to... Figure 2 As shown, in step S200, responding to the rescuer's voice data using a multimodal large model includes: Step S210: Identify the responsiveness of passengers based on capability features; Step S220: If the occupants of the vehicle are unable to respond, the accident status is identified by inferring after acquiring vehicle environmental data, and rescue suggestions and / or response information are generated based on the accident status. Step S230: Convert the rescue suggestion and / or the response information into audio data, and communicate with the rescuer through the emergency rescue system based on the audio data.
[0052] Conversely, if the people in the vehicle are able to respond, the emergency call function can be used to allow them to communicate directly with the rescuers.
[0053] In other possible implementations, the Qwen 2.5 VLM 7B can be replaced with a multimodal model with a larger parameter scale to enhance inference depth, or a lightweight version can be deployed to adapt to resource-constrained local in-vehicle environments.
[0054] Based on the above steps, the control method of this application leverages the vehicle's perception and intelligent processing capabilities. When an emergency call is detected, it acquires the capability characteristics of the occupants, automatically invokes a pre-trained multimodal large model, receives and understands voice data from the rescue team, and comprehensively judges whether the occupants possess normal response capabilities. If it is determined that the occupants cannot respond effectively, it further collects environmental data (such as collision intensity, location, etc.) from the vehicle's sensor network and infers the possible state. Based on this inference result, it automatically generates semantic information including rescue suggestions and / or response information, converts it into voice format, and feeds it back to the rescue team through the emergency rescue system. If the recognition result indicates that the occupants possess response capabilities, it directly activates the onboard emergency communication module, enabling them to communicate with the rescue team via voice to assist in handling the emergency situation.
[0055] By combining a multimodal large-scale model with an onboard sensor network, intelligent judgment and interactive response to the status of people in emergency situations are achieved, improving the efficiency of rescue communication when people in the vehicle are unable to express themselves. This application embodiment can automatically determine the response method based on the actual situation, thereby shortening the rescue response time, improving the accuracy and usability of rescue information, and reducing the risk of delays in rescue decisions. At the same time, automatic voice feedback reduces the burden on rescuers in obtaining information, helping them to make targeted rescue decisions in the first instance.
[0056] Furthermore, if the occupants are capable of responding, their direct communication with the rescue team will not be interfered with, ensuring a balance between human-centered design and system assistance.
[0057] In some embodiments, the capability characteristics include: physiological data, appearance data, and / or voice data, and identifying the responsiveness of on-board personnel based on the capability characteristics further includes: Facial features of people in the vehicle are identified based on the apparent data; Identify the physical condition of passengers based on the aforementioned physiological data; Based on the sound data, it can be identified whether the people in the vehicle are making voice outputs; Determine whether the facial features, body state, and voice output meet the preset response capability conditions. If they do, determine that the person on the vehicle is unable to respond.
[0058] Based on the above configuration, this embodiment of the application refines the capability characteristics of the occupants into three categories: physiological data, appearance data, and voice data, and comprehensively judges the occupants' physical state, external performance, and voice behavior, respectively. After integrating the three types of information, it is compared with preset response capability conditions. If the conditions are met, it is determined that the occupants do not have the ability to respond, and thus the occupants communicate with the rescue team by voice on their behalf. The multi-dimensional comprehensive recognition method helps to achieve stable assessment even in complex accident environments, with high noise interference, or poor lighting conditions, improving the reliability of the judgment, reducing the false judgment rate, and enhancing the robustness in real-world application scenarios.
[0059] In some embodiments, LoRA technology is used to fine-tune the multimodal model. During fine-tuning, a multimodal training set is constructed, containing appearance data, physiological data, sound data, and corresponding labels (such as "pain," "closed eyes," "abnormal heart rate," "no speech output," etc.). The LoRA mechanism is used to perform low-rank adjustment on the key weight matrix in the multimodal model parameters, significantly reducing the number of training parameters and resource consumption, improving training efficiency, and enhancing the model's recognition ability under specific tasks. After fine-tuning, the model can accurately recognize facial features, body states, and the presence or absence of speech output of people in the vehicle.
[0060] In some embodiments, when an emergency call is detected by the emergency assistance system, the capability characteristics of the occupants are acquired, including: when an emergency call is detected by the emergency assistance system, the driver detection system (DMS) is activated, and the driver detection system calls the vehicle sensor system to collect the physiological data, the appearance data, and the voice data.
[0061] Based on the above configuration, the emergency assistance system and the driver detection system work together to automate the collection and analysis of the occupants' capability characteristics. By automatically activating the driver detection system the moment an emergency call is triggered, the timeliness and comprehensiveness of occupant status perception are effectively improved.
[0062] In the above embodiments, DMS (Driver Monitoring System) is an active safety technology that monitors the driver's status in real time through multimodal sensors and algorithms. Its core objective is to prevent traffic accidents caused by driver fatigue, distraction, operational errors, or sudden health problems (such as coma).
[0063] In some embodiments, the physiological data includes body temperature characteristics, heart rate, etc. The physiological data of this application can monitor the driver's breathing rate and heart rate through electromagnetic wave reflection, which is suitable for detecting abnormal heartbeat and / or breathing caused by accidents. It can also be obtained directly by calling the infrared camera of the vehicle's sensing system, or it can be based on indirect methods such as subtle changes in facial skin color to assist in the judgment.
[0064] At this point, the infrared camera can be deployed above the steering wheel, dashboard, or inside the A-pillar to capture facial infrared images, ensuring clear facial thermal images under different sitting postures and lighting conditions, and operating effectively at night or in low-light environments, avoiding ambient light interference. The infrared camera can work in low light (such as at night) to avoid ambient light interference.
[0065] In some embodiments, the appearance data is used to display facial expressions, eye states, and head posture. The eye states include open and closed eyes. The appearance data in this application can be directly acquired through the optical camera of the vehicle sensing system.
[0066] In some embodiments, the sound data is acquired through an audio acquisition unit of the vehicle sensing system. Optionally, the audio acquisition unit may be a single microphone, an array microphone, or a bone conduction microphone.
[0067] Among them, the array microphone uses multiple microphones to focus on the target sound source (such as the direction of the driver's speech) through beamforming technology, suppresses environmental noise (such as wind noise and air conditioning noise), and improves the accuracy of speech recognition.
[0068] Among them, the bone conduction microphone collects vocal cord vibration signals through a vibration sensor, avoiding air noise interference.
[0069] Based on the above configuration, the embodiments of this application effectively improve the accuracy and real-time performance of physiological data acquisition by introducing multiple non-invasive monitoring methods such as electromagnetic wave reflection, physiological thermal imaging, and facial skin color difference detection. This makes it particularly suitable for detecting the unconscious or immobile state of occupants in accident scenarios. The application of an infrared camera enables the system to maintain stable identification even at night or in low-light environments, improving overall robustness.
[0070] By comprehensively analyzing facial expressions, eye states, and head posture, it is possible to more accurately determine whether a person is conscious or experiencing impaired consciousness. Audio recording directly assesses a person's language abilities, providing the system with more complete multimodal support when determining responsiveness. The integrated application of these feature analysis methods enables a highly accurate and adaptable intelligent assisted rescue response system.
[0071] In another embodiment, the infrared camera can be implemented using different technologies such as a thermopile infrared sensor or an uncooled focal plane array infrared imaging module, and a temperature calibration mechanism can be added to improve measurement accuracy.
[0072] In another embodiment, electromagnetic wave reflection monitoring can employ millimeter-wave radar, UWB radar technology, or a biosensing module based on 5.8 GHz microwaves to achieve physiological monitoring under different installation conditions.
[0073] In another embodiment, in terms of facial state recognition, eye state can be further quantified into parameters such as eye opening and closing time and blinking frequency to improve the accuracy of consciousness state recognition.
[0074] In some embodiments, the preset response capability conditions are configured as follows: if the facial features include pain features, closed eyes features and no voice output, and at the same time, the physical state includes abnormal body temperature features or abnormal heart rate features, then it is determined that the driver has lost effective response capability. The multimodal big data model automatically responds to generate rescue information. For example, but not limited to, pain features may include furrowed brows and drooping corners of the mouth; abnormal body temperature features are determined by being higher than a preset safe temperature range, such as 37.5 degrees Celsius; and abnormal heart rate features are determined by being higher or lower than a safe heart rate threshold.
[0075] Based on the above implementation method, if the driver's facial expression is painful, his eyes are closed and he makes no sound, and the infrared camera detects an abnormally high body temperature and an extremely unstable heart rate, the large model will determine that the driver is unable to respond; if the driver has a normal facial reaction, his eyes are open (such as continuously looking ahead or blinking naturally) and he can make clear sounds (the sound acquisition unit detects continuous and clear speech output), then he will determine that he is able to respond.
[0076] This application's embodiments achieve highly sensitive and accurate abnormal state recognition by setting typical abnormal combinations (such as facial expression of pain + closed eyes + no speech + high body temperature + irregular heart rate). Combined with a large model for semantic-level judgment, it enhances the understanding and reasoning ability of complex interactive information, effectively reducing erroneous judgments caused by perceptual blind spots or one-dimensional misjudgments. This method enables rapid decision-making in time-sensitive accident scenarios, providing timely and reliable data support for subsequent rescue efforts, while maintaining the intelligence and controllability of the system response.
[0077] In another embodiment, facial expression recognition can incorporate a Facial Action Coding System (FACS) for standardized judgment, while eye state recognition can be further refined to indicators such as eye movement trajectory and pupil response. Body temperature and heart rate thresholds can be personalized with upper and lower limits set based on individual historical data to improve recognition accuracy.
[0078] In some embodiments, the step of inferring and identifying the accident state after acquiring vehicle environmental data includes: The vehicle sensing system is invoked to collect images of the road ahead and vehicle collision data. The multimodal large model analyzes the road image ahead of the vehicle and the vehicle collision data to identify whether a collision has occurred. If so, it identifies the collision location and the severity of the collision, and outputs collision information including the collision location and the severity of the collision.
[0079] This solution significantly enhances a vehicle's ability to understand changes in the external environment during sudden accidents. Compared to traditional sensor threshold judgment methods, this method uses a large model to deeply understand the semantic relationship between road images and vehicle body data, effectively reducing false triggering rates and the risk of missed detections, and is particularly accurate in ambiguous scenarios such as minor scrapes and offset collisions.
[0080] LoRA technology is used to fine-tune the process of acquiring collision information for the multimodal large model. A multimodal training set covering a variety of typical collision scenarios is constructed. This set includes images of the road in front of the vehicle taken by the forward-looking camera, vehicle collision data such as acceleration direction and deformation depth, and pre-labeled accident results, including whether there is a collision, the collision location and the severity of the collision, such as "moderate in front" and "minor on the right".
[0081] During fine-tuning, LoRA was used to inject low-rank parameters into the key attention layer and visual embedding layer within the model, optimizing only the subset of parameters related to collision recognition. This retained the original model's general capabilities while endowing it with fine-grained judgment capabilities for collision recognition. After fine-tuning, the model possesses stronger image-sensor data collaborative reasoning capabilities, enabling it to quickly identify collision situations and their detailed states under complex on-site conditions, for use in subsequent rescue response generation and information transmission.
[0082] In some embodiments, the inference and identification of the accident state after acquiring vehicle environmental data includes: The vehicle sensing system is invoked to acquire images of the area surrounding the vehicle body; The multimodal large model extracts road features, terrain features, and obstacle features around the vehicle based on the images around the vehicle. Based on the road features, it identifies the distance to the road edge, the angle and displacement of the vehicle from the road, the terrain features identify the type of terrain the vehicle is in, and the obstacle features identify the types of surrounding obstacles and whether there are safety protection facilities. Based on the distance to the road edge, the angle and displacement of the vehicle from the road, the type of terrain the vehicle is in, the types of surrounding obstacles, and the presence of safety protection facilities, it obtains vehicle scene information, thereby determining whether the vehicle has run off the road or is in a dangerous area, such as near a cliff or river.
[0083] Based on the above steps, multi-dimensional data fusion analysis effectively identifies accident types such as vehicles deviating from their path, running off the road, or entering dangerous areas due to loss of driving control, collision impact, or poor road conditions. The introduction of vehicle scene information not only improves the comprehensiveness of accident assessment but also helps in the rapid adjustment of rescue plans and route planning, such as prioritizing the deployment of hoisting equipment and bypassing dangerous roads.
[0084] To achieve the aforementioned comprehensive scene understanding capability, the LoRA fine-tuning method was used to enhance the training of the multimodal large model. Specific steps included constructing a training dataset containing multiple types of road scenes. Data samples included images of the vehicle's surroundings, corresponding vehicle trajectory and pose data, and labeled road types, terrain attributes, obstacle types, and distribution information. During training, a structured prompt template was used to pair image information with labels and input them into the model. For example: "The vehicle in the image is located on a slope, with a road edge to the left and a metal guardrail to the right, with an 8-degree deviation angle. There are no obvious obstacles in the scene." The multimodal large model uses these samples to jointly model semantic labels and spatial data. The LoRA module performs low-rank matrix insertion on key attention layers and trains only the inserted parameters, thereby endowing it with specific recognition capabilities for features such as roads, terrain, and obstacles without compromising the original model's capabilities. After fine-tuning, the model can automatically infer the comprehensive environmental state of the vehicle in real-world in-vehicle scenarios, such as road edge lines, lane line shapes, unpaved surfaces, terrain contours, slopes, and potential obstacles such as trees, guardrails, rocks, and water bodies, providing crucial information for subsequent rescue strategy generation and accident assessment.
[0085] In some embodiments, the inference and identification of the accident state after acquiring vehicle environmental data includes: The vehicle sensing system is invoked to collect images of the vehicle cabin and acquire the sound data. The multimodal large model extracts the bleeding features of the personnel in the vehicle cabin image and the sound content in the sound data, identifies the bleeding features and the sound content, and outputs the personnel injury information.
[0086] Based on the above steps, the multimodal large model of this application receives and processes the aforementioned vehicle cabin image and sound data. It extracts visual features from key areas such as the face, arms, clothing, and seat area in the images to identify the presence of bloodstains, obvious wounds, blood-stained clothing, and other signs of bleeding. Simultaneously, it analyzes the language content in the sound data, including non-verbal signals such as cries for help, groans, and intermittent speech, fusing and inferring whether the occupants are currently injured, the severity of their injuries, and whether they are conscious, and outputs structured information about their injuries. For example, "The passenger on the left is bleeding significantly, accompanied by continuous groans; the injury may be serious."
[0087] This application significantly improves the accuracy and timeliness of occupant injury identification after an accident by introducing joint analysis of vehicle cabin image and sound data. Through a multimodal large-scale model, it automatically determines whether an occupant is injured and the severity of their injury. Especially when occupants are unable to express themselves, the system can still proactively initiate effective intelligent responses and rescue requests.
[0088] When constructing the training dataset for the multimodal large model, multimodal samples containing real or simulated post-accident images inside the vehicle, occupant injury audio recordings, and their annotation information are collected. The image portion should include samples with bleeding characteristics under different genders, clothing, locations, and postures. The audio portion includes speech content under different tones, emotions, and background noise. Labels include "bleeding," "call for help," and "state of consciousness." During the fine-tuning phase, input data is organized using prompt templates, for example: "The image shows bloodstains on the occupant's face and clothing, and intermittent calls for help are detected in the audio. Is there a serious injury?" The model inserts a low-rank weight matrix using a LoRA structure, adjusting only some parameters of the key visual semantic layer and speech understanding module to quickly adapt to the injury recognition task requirements. After fine-tuning, the model possesses the ability to efficiently and accurately identify injury features and generate rescue decision-making criteria in complex in-vehicle environments.
[0089] This application employs the LoRA fine-tuning method, which not only ensures high model training efficiency and low deployment cost, but also possesses strong task adaptability and scalability, making it suitable for various vehicle models and in-vehicle configurations. Furthermore, the structured output injury information can be directly used for subsequent rescue command generation, resource allocation, and medical response decisions, constructing a complete closed loop from perception to action.
[0090] In some embodiments, the inference and identification of the accident state after acquiring vehicle environmental data includes: The collision information, vehicle scene information, and personnel injury information, or any combination thereof, are input into the multimodal large model to infer the vehicle accident state. The vehicle accident state includes the vehicle location, number of people in the vehicle, personnel status, vehicle collision state, and vehicle scene state. The vehicle location is determined with the assistance of vehicle GPS data.
[0091] To achieve this information fusion and reasoning capability, this application employs LoRA fine-tuning to enhance the multimodal large model's ability to model the semantic combination and logical relationships of accidents. During training, cross-modal integrated training samples are constructed. Each sample contains collision information, vehicle scene information, and the aforementioned personnel injury information, with the corresponding output being a structured accident state label. Data is organized using a prompt template, such as: "Input: Severe collision, road edge close ahead, occupant unconscious, GPS X; Output: Vehicle status: off-road, severe frontal collision, 2 people inside, 1 unconscious." LoRA is used to insert low-rank weight modules only into the model's key fusion layer and task output layer, rapidly transferring and enhancing the ability to reason about accident state combinations. After training, the model can quickly infer a complete accident scene based on partial or complete input information, providing a complete data foundation for emergency communication, rescue resource scheduling, and other applications.
[0092] This application, based on multimodal information fusion, generates a comprehensive vehicle accident status through reasoning, providing crucial support for subsequent emergency response and information exchange. Compared to traditional rule-matching-based analysis methods, this approach offers greater flexibility and scalability, and can handle scenarios with missing, incomplete, or abnormal information.
[0093] To acquire the aforementioned vehicle environmental data, the vehicle sensing system is equipped with a front-view camera, a surround-view camera, and an in-vehicle camera. The front-view camera is used to acquire images of the road in front of the vehicle to determine road conditions and obstacle information. The surround-view camera is used to acquire images of the area around the vehicle to determine the vehicle's location and surrounding environmental features. The in-vehicle camera is used to acquire images of the vehicle's cabin to identify the number of people and their injury status. Vehicle sensing data is acquired through vehicle collision sensors, speed sensors, and other sensors.
[0094] In some embodiments, the forward-facing camera is typically mounted in the center of the upper edge of the windshield (near the rearview mirror) or in the center of the front bumper, providing a wide-angle view directly in front of the vehicle. The forward-facing camera can capture real-time images of the road ahead to identify road types (such as highways and urban roads), lane markings, vehicles ahead, and obstacles, enabling assessment of road conditions and potential hazards.
[0095] In some embodiments, surround-view cameras are installed around the vehicle: on both sides of the front bumper, in the center or on the left and right sides of the rear bumper, and below the left and right exterior rearview mirrors, etc., to form a 360-degree image stitching without blind spots, for real-time acquisition of images around the vehicle.
[0096] Based on the above embodiments, image analysis technology is used to identify the vehicle's location features, such as whether it is deviating from the lane, close to the guardrail, or in the emergency lane; and the surrounding environment features, such as whether there are obstacles or whether it is surrounded by other vehicles, to assist the large model in generating highly reliable accident description information.
[0097] In some embodiments, the in-vehicle camera is typically located inside the A-pillar, on the center console of the roof, or below the rearview mirror to capture cabin images for occupant number recognition and injury status recognition.
[0098] In some embodiments, vehicle collision sensors are typically deployed inside the front and rear bumpers, in the inner side beams of the doors, and at key structural points in the vehicle chassis to detect impact data of different directions and intensities. Speed sensors are usually integrated with the wheel ABS module or installed separately on the drive shaft or transmission output port to measure the vehicle's instantaneous speed and acceleration changes. Upon receiving this sensor data, the multimodal large model can identify the direction of the collision (frontal, rear, side impact, etc.), its intensity (minor or severe), and whether sudden deceleration, rollover, or other accident scenarios occurred.
[0099] Based on this, the embodiments of this application achieve high-precision perception of the vehicle's surrounding environment and internal state, effectively improving the ability to identify accident scenarios. The standardized installation position of the camera ensures the integrity and stability of image acquisition, guaranteeing continuous operation in various driving and collision environments. Sensor deployment supplements the shortcomings of image recognition, forming an information loop and improving the accuracy and real-time performance of reasoning.
[0100] Furthermore, this application embodiment automatically identifies accident states by collecting and processing multi-source vehicle environmental data. It can not only determine the existence of an accident when occupants cannot directly express their situation, but also analyze the accident type and severity, greatly enhancing intelligent assisted rescue capabilities. Through structured extraction of accident features, a multimodal large model can output differentiated rescue strategies based on scenario differences, improving the practicality and relevance of rescue information.
[0101] In another embodiment, the front-view camera of this application may support night vision or HDR mode to enhance the image acquisition quality under extreme lighting conditions.
[0102] In another embodiment, the quantification of the degree of collision can be achieved by combining factors such as airbag triggering state, collision duration, and synchronization signals from multiple sensors to construct a multi-dimensional collision severity scoring system.
[0103] Furthermore, embodiments of this application may also incorporate a sound recognition module to monitor the audio characteristics of impact sounds on the vehicle body as a source of collision assistance signals. The collision location determination can also be combined with an inertial navigation system to identify changes in vehicle attitude (such as tilt angle or rollover) to assist in determining whether a serious accident such as a vehicle rollover has occurred.
[0104] In another embodiment, positioning data provided by the vehicle's GPS or high-precision GNSS module can be invoked to obtain the vehicle's current geographic coordinates and altitude information. This information is then matched with a preset map database or Geographic Information System (GIS) to identify whether the vehicle is located in a dangerous area (such as mountain roads, waterfront sections, bridge edges, or the outside of curves). After integrating image and geographic location information, high-risk conditions such as "the vehicle has deviated 30 degrees from its lane, and there is an unguarded cliff 2 meters ahead" or "the vehicle has left the road and entered a riverside grassland area" can be identified and used as one of the accident characteristics in subsequent judgment and response.
[0105] Based on this, by using multi-dimensional data fusion to determine vehicle location status, it is possible to effectively identify accident types such as vehicles deviating from their path, running off the road, or entering dangerous areas due to loss of control, collision impact, or poor road conditions. The introduction of vehicle positioning data not only improves the comprehensiveness of accident assessment but also helps in the rapid adjustment of rescue plans and route planning, such as prioritizing the deployment of hoisting equipment and bypassing dangerous roads.
[0106] By combining with static map data, vehicles can still possess a certain degree of geographical awareness even without a network connection, improving the reliability of the model's independent operation. Comparing road structure with the actual vehicle location can also help determine potential causes of accidents (such as oversteering, brake failure, etc.), providing foundational data support for subsequent liability analysis.
[0107] In addition, this application can also realize automatic injury classification, assigning different hazard level labels to different occupants (such as serious, minor injury, conscious), which helps to rationally allocate rescue resources and improve emergency response efficiency.
[0108] In other alternatives, image recognition algorithms can use color space conversion and texture analysis techniques specifically designed for blood testing (such as HSV space separation and edge enhancement processing) to improve the blood flow recognition rate under low-light conditions.
[0109] In another embodiment, if serious injuries are detected, a predefined high-priority emergency channel can be automatically triggered, for example, by directly reporting to the rescue team that "multiple occupants in the vehicle are seriously injured and a rapid response is required."
[0110] In some embodiments, the rescue recommendations include at least: medical rescue recommendations, safety warning recommendations, and traffic control recommendations.
[0111] In some embodiments, during the process of generating rescue suggestions and / or response information based on the accident status, the multimodal large model acquires structured collision information, vehicle scene information, and personal injury information, and then encodes and converts them to adapt the model input in the form of vectorization or semantic tags. The model configuration integrates a rule base to achieve dynamic matching between structured data and semantic generation templates, thereby improving the consistency of the generation logic. The large model is also configured to support multi-layer semantic analysis capabilities to understand the accident risk level, occupant urgency, etc., under different data combinations.
[0112] To improve response speed and content standardization, multiple semantic generation templates can be preset, such as: Medical advice template: "The {N}th occupant in the vehicle has {injury characteristics}, and priority treatment is recommended." Safety advice template: "The vehicle is located in a {high-risk type} area. It is recommended to avoid approaching." Traffic suggestion template: "The accident is affecting the passage of {road name}. It is recommended to contact {traffic police / platform} for traffic management." After generating suggested content, the natural language processing module is used for language clarification to avoid repetition, ambiguity, and semantic jumps. If necessary, a model fine-tuning layer is configured to focus on language style alignment training for three types of rescue discourse: medical, safety, and transportation.
[0113] To enable voice broadcast of rescue suggestions, a TTS (Text-To-Speech) module can be configured to interface with a multimodal large model, so that the content generated by the model is automatically converted into multiple audio formats (supporting multiple languages and multiple tones), which is suitable for various rescue scenarios.
[0114] This application embodiment can combine vehicle cabin images, sound data, physiological state and appearance features. If it is identified that the occupant has obvious injury (such as bleeding, groaning, or coma), it will trigger the generation of medical assistance suggestions and output structured text based on the injury level and location, such as "The rear passenger is suspected of having facial trauma, accompanied by coma".
[0115] Secondly, the control method of this application embodiment, if it is identified through vehicle positioning data, map information and environmental images that the vehicle is approaching a dangerous area, such as a gas station, chemical plant, or a section of road prone to landslides, generates safety warning suggestions to remind rescue personnel to pay attention to self-protection and on-site safety control measures.
[0116] Furthermore, this application can also utilize traffic sensing modules or real-time cloud traffic data to determine whether an accident has caused traffic disruption. If an increase in queuing vehicles or traffic interruption is detected at key nodes such as intersections and bridge passages, traffic guidance suggestions will be automatically generated to prompt the rescue team to coordinate with the traffic police to guide traffic and alleviate regional congestion.
[0117] The aforementioned suggestions are semantically optimized and converted into voice data, which is then transmitted to the rescue team in real time via the eCall system.
[0118] This method enables the automatic generation and classification of accident rescue information, allowing accident response to extend beyond the minimum dataset MSD to provide comprehensive guidance across three dimensions: medical, safety, and transportation. This significantly improves the efficiency of rescue organization and the scientific nature of on-site handling.
[0119] Among them, the medical advice section can help rescuers allocate resources rationally and prepare targeted equipment or medicines in advance; the safety advice provides advance judgment for personnel protection and reduces the risk of secondary injuries; and the traffic advice can avoid large-scale traffic congestion caused by accidents and improve the robustness of the overall transportation system.
[0120] In addition, providing suggestions via voice prompts makes it easier for rescuers to quickly obtain key information in noisy or limited-visibility environments.
[0121] In other implementations, medical recommendations can be combined with historical injury treatment cases and medical databases to generate personalized treatment suggestions, such as "recommend carrying a hemostatic clip and neck brace"; safety recommendations can include a regional risk level quantification mechanism to dynamically adjust warning intensity based on the safety level of the area where the vehicle is located; traffic recommendations can access local V2X infrastructure or vehicle-to-everything (V2X) platform information to obtain more granular congestion routes, feasible detours, and other information. Recommendation content can be generated based on preset rule sets, semantic template matching, or automatic generation mechanisms using large language models, and supports multilingual output to adapt to different rescue systems or usage scenarios.
[0122] In another embodiment, utilizing the multimodal large model to respond to the rescuer's voice data includes: Based on the aforementioned capability characteristics, the responsiveness of the people on board the vehicle can be identified; If the occupants are unable to respond, the system will acquire vehicle environmental data, infer and identify the accident status, and generate rescue suggestions based on the accident status. The rescue suggestions are converted into audio data and played through the emergency rescue system to provide feedback to the rescue team; The system receives voice data from the rescuer's call, inquiring about current information. It then repeats the steps of obtaining vehicle environmental data and inferring the accident status, extracting key information corresponding to the inquiring voice data and generating corresponding response information. It interacts with the rescuer until the call ends.
[0123] For example, if the rescue team asks for the "specific age and gender of the injured person inside the vehicle," the AI model will again use information such as the vehicle's in-vehicle cameras to make inferences and provide a corresponding answer, such as "the injured person inside the vehicle is a woman around 30 years old." Through this closed-loop interaction, the rescue team can obtain comprehensive and accurate information, improving rescue efficiency.
[0124] In one possible implementation, the control method is encapsulated as an intelligent agent module, operating in the vehicle terminal with autonomous perception, reasoning, interaction, and response capabilities. The intelligent agent is used to invoke a multimodal large model and complete the entire process of accident status identification and rescue suggestion generation when an emergency rescue system is detected. This intelligent agent includes a perception input module, a status evaluation module, a multimodal large model interface module, a semantic response generation module, and a general scheduling control module.
[0125] The perception input module collects and standardizes multi-source environmental data from the vehicle, including images of the road ahead, the surrounding area, the passenger compartment, collision sensor data, speed data, occupant voice data, and physiological parameters. This data is then encapsulated into structured data packets for subsequent inference and generation. The state assessment module determines whether an emergency has been entered. Once the emergency call system is activated, it immediately initiates the multimodal processing flow.
[0126] The multimodal large model interface module constitutes the core inference engine of the intelligent agent. It is used to call pre-trained multimodal large models (such as Qwen 2.5 VLM 7B) deployed on local or remote servers. The model receives multimodal fusion data provided by the perception input module and outputs the following structured inference results: the responsiveness status of the occupants, the type of accident, the location and severity of the collision, whether the vehicle's location is dangerous, the severity of occupant injuries, and traffic environmental risks. This module also has task distribution capabilities, and can select different sub-paths of the large model for processing according to different tasks (state judgment or language generation).
[0127] The semantic response generation module automatically constructs various types of emergency response content based on reasoning results, including medical assistance suggestions, safety tips, and traffic management suggestions. This module is configured with language generation strategies or semantic template systems and supports natural language optimization and speech synthesis. The generated response information will be sent to the rescuer through the in-vehicle audio system or an emergency communication channel connected to an external rescue platform, realizing the "system automatic response" function.
[0128] The general scheduling and control module serves as the internal control center of the agent, responsible for monitoring emergency event processes, state judgment logic, data flow paths, and access permissions for large model resources. It also manages the overall operation of the agent in real time, such as handling abnormal situations like large model response timeouts, missing modal data, and duplicate calls, ensuring that the agent can operate stably under different accident conditions.
[0129] The agent's packaging method supports modular deployment, enabling it to run on in-vehicle platforms with heterogeneous computing resources (such as CPU+GPU / NPU), and supports OTA updates and cloud-based model collaborative inference expansion. Furthermore, the agent supports multi-language and multi-regional map environment adaptation configurations to meet the regulatory and language requirements of in-vehicle systems in different countries or regions.
[0130] In one possible implementation, to achieve efficient collaboration between the intelligent agent and the vehicle system, communication system, and cloud platform, the intelligent agent is further configured with a system interaction adaptation module and a cloud-edge collaboration interface module. As an independent operating unit within the vehicle system, the intelligent agent achieves a complete closed-loop chain of accident recognition, model invocation, information generation, and external communication through deep integration with the local perception system, information processing system, and remote communication platform.
[0131] The system interaction adaptation module manages the data interface between the intelligent agent and the vehicle's local system, including protocol adaptation and data standardization processing with the CAN bus system, central gateway, camera module, sensor integration system, physiological monitoring equipment, and vehicle operating system (such as QNX, Linux, Android Automotive). This module supports multiple sensor interface protocols (such as CAN, LIN, FlexRay, and Ethernet), abstracting and encapsulating various data source signals into a unified data format for the intelligent agent's core module to call and process, while ensuring compatible operation with existing vehicle controller systems (such as ECU and ADAS domain controllers).
[0132] The cloud-edge collaboration interface module is used for communication coordination between the agent and the cloud-based large model platform or remote service platform. This module supports data transmission via vehicular cellular communication (such as 4G / 5G modules), V2X (Vehicle-to-Everything) technology, and Low-Power Wide Area Network (LPWAN), and possesses network status awareness and switching mechanisms. When onboard computing resources are limited or the model's local inference pressure is high, the agent can upload data to the cloud-based model processing platform through the cloud-edge collaboration interface, call upon a high-performance large model (such as the Qwen 14B VLM) to complete inference or language generation tasks, and then send the results back to the local machine.
[0133] To ensure stable operation under weak or no network conditions, the agent is equipped with a local inference capability caching mechanism. When the network is unavailable, a local lightweight model (such as a distilled version of Qwen VLM) can be used to complete core emergency functions, and the latest data and model updates are automatically synchronized after the network is restored. This mechanism ensures that the system can complete critical functions without relying on the network in accident scenarios.
[0134] Furthermore, when collaborating with external rescue platforms, the intelligent agent can connect with traffic police systems, emergency platforms, medical emergency centers, or OEM cloud platforms via predefined communication protocols (such as HTTP / MQTT / WebSocket). All outputs of accident status descriptions, injury assessments, and rescue suggestions are sent via encrypted communication, complying with vehicle-to-everything (V2X) communication security standards (such as TLS, ISO / SAE 21434, GB / T 38693, etc.), ensuring the integrity and confidentiality of data transmission.
[0135] The system architecture of this intelligent agent supports multi-task parallelism and hot-swappable modules, has good scalability and platform portability, is suitable for various vehicle platforms (passenger vehicles, commercial vehicles, new energy vehicles) and system integration scenarios of different OEMs, and has the feasibility and scalability for engineering implementation.
[0136] This application embodiment utilizes an AI multimodal large model to perform scene understanding on (inside and outside the vehicle) images, voice, and vehicle data, analyze the situation inside and outside the vehicle, provide on-site situation response content, communicate with humans, avoid rescue difficulties due to the incapacity of drivers or passengers, and improve rescue efficiency.
[0137] It should be noted that the steps shown in the above process or in the flowchart of the accompanying figures can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.
[0138] This application also provides a vehicle, as shown in the embodiments. Figure 3 As shown, the vehicle includes a vehicle-mounted communication terminal and a vehicle-mounted terminal with communication connection. The vehicle-mounted communication terminal is equipped with an emergency rescue system or is connected to the emergency rescue system. The vehicle-mounted terminal is used to execute the control method described above. In order to obtain the vehicle environment data required for the method calculation, the vehicle is equipped with image acquisition devices such as a front-view camera, a surround-view camera and an in-vehicle camera, as well as a vehicle collision sensor (not shown in the figure) and a speed sensor (not shown in the figure).
[0139] The vehicle provided in this embodiment is used to execute the control method described above, and therefore can achieve the same effect as the implementation method described above.
[0140] The beneficial effects of the above embodiments can be referred to the beneficial effects of the corresponding methods provided above, and will not be repeated here.
[0141] Combination Figures 1-2 The vehicle control method described in this application embodiment can be implemented by an electronic device. Figure 4 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of this application.
[0142] refer to Figure 4 As shown, the electronic device may include a processor 41 and a memory 42 storing computer program instructions.
[0143] Specifically, the processor 41 may include a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.
[0144] The memory 42 may include a large-capacity memory for data or instructions. For example, and not limitingly, the memory 42 may include a hard disk drive (HDD), a floppy disk drive, a solid-state drive (SSD), flash memory, an optical disk drive, a magneto-optical disk drive, magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory 42 may include removable or non-removable (or fixed) media. Where appropriate, the memory 42 may be internal or external to a data processing device. In a particular embodiment, the memory 42 is non-volatile memory. In a particular embodiment, the memory 42 includes read-only memory (ROM) and random access memory (RAM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), an electrically alterable read-only memory (EAROM), or flash memory, or a combination of two or more of these. Where appropriate, the RAM can be Static Random-Access Memory (SRAM) or Dynamic Random-Access Memory (DRAM). DRAM can be Fast Page Mode Dynamic Random-Access Memory (FPMDRAM), Extended Data Out Dynamic Random-Access Memory (EDODRAM), Synchronous Dynamic Random-Access Memory (SDRAM), etc.
[0145] The memory 42 can be used to store or cache various data files that need to be processed and / or used for communication, as well as possible computer program instructions executed by the processor 41.
[0146] The processor 41 implements any of the vehicle control methods described in the above embodiments by reading and executing computer program instructions stored in the memory 42.
[0147] In some embodiments, the electronic device may further include a communication interface 43 and a bus 40. (Referring to...) Figure 4 As shown, the processor 41, memory 42, and communication interface 43 are connected through bus 40 and complete communication with each other.
[0148] The electronic device can execute the vehicle control method in the embodiments of this application based on the acquired computer program instructions.
[0149] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0150] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. A method for controlling a vehicle, characterized in that, Applied to vehicle-mounted terminals, the method includes: When an emergency call is triggered by the emergency assistance system, the capability characteristics of the occupants in the vehicle are obtained; The system acquires the rescuer's voice data in real time and calls a pre-trained multimodal large model to respond to the rescuer's voice data. The use of the multimodal large model to respond to the rescuer's voice data includes: Based on the aforementioned capability characteristics, the responsiveness of the people on board the vehicle can be identified; If the occupants are unable to respond, the accident status is inferred and identified after acquiring vehicle environmental data, and rescue suggestions and / or response information are generated based on the accident status. The rescue suggestions and / or the response information are converted into audio data, and the dialogue is conducted with the rescuer through the emergency rescue system based on the audio data.
2. The control method according to claim 1, characterized in that, The capability characteristics include: physiological data, appearance data, and / or voice data. Identifying the responsiveness of on-board personnel based on these capability characteristics further includes: Facial features of people in the vehicle are identified based on the apparent data; Identify the physical condition of passengers based on the aforementioned physiological data; Based on the sound data, it can be identified whether the people in the vehicle are making voice outputs; Determine whether the facial features, body state, and voice output meet the preset response capability conditions. If they do, determine that the person on the vehicle is unable to respond.
3. The control method according to claim 2, characterized in that, When an emergency call is triggered by the emergency assistance system, the capability characteristics of the occupants in the vehicle are obtained, including: When an emergency call is detected by the emergency assistance system, the driver detection system is activated, which calls the vehicle sensor system to collect the physiological data, the appearance data, and the sound data.
4. The control method according to claim 3, characterized in that, The process of inferring and identifying the accident state after acquiring vehicle environmental data includes: The vehicle sensing system is invoked to collect images of the road ahead and vehicle collision data. The multimodal large model analyzes the road image ahead of the vehicle and the vehicle collision data to identify whether a collision has occurred. If so, it identifies the collision location and the severity of the collision, and outputs collision information including the collision location and the severity of the collision.
5. The control method according to claim 4, characterized in that, The process of inferring and identifying the accident state after acquiring vehicle environmental data includes: The vehicle sensing system is invoked to acquire images of the area surrounding the vehicle body; The multimodal large model extracts road features, terrain features, and obstacle features around the vehicle based on the images surrounding the vehicle. It then identifies the road edge distance, the vehicle's deviation angle and displacement from the road based on the road features, the terrain type of the terrain the vehicle is in based on the terrain features, and the types of surrounding obstacles and the presence of safety protection facilities based on the obstacle features. Vehicle scene information is obtained based on the distance to the road edge, the angle and displacement of the vehicle from the road, the terrain type where the vehicle is located, the type of surrounding obstacles, and whether there are safety protection facilities.
6. The control method according to claim 5, characterized in that, The process of inferring and identifying the accident state after acquiring vehicle environmental data includes: The vehicle sensing system is invoked to collect images of the vehicle cabin and acquire the sound data. The multimodal large model extracts the bleeding features of the personnel in the vehicle cabin image and the sound content in the sound data, identifies the bleeding features and the sound content, and outputs the personnel injury information.
7. The control method according to claim 6, characterized in that, The process of inferring and identifying the accident state after acquiring vehicle environmental data includes: The collision information, vehicle scene information, and personal injury information, or any combination thereof, are input into the multimodal large model to infer the vehicle accident state.
8. The control method according to claim 7, characterized in that, The multimodal large model is fine-tuned based on an efficient parameter fine-tuning method and then pre-deployed on the vehicle terminal.
9. A vehicle, characterized in that, The system includes a vehicle-mounted communication terminal and a vehicle-mounted terminal with communication connection. The vehicle-mounted communication terminal is equipped with an emergency rescue system or is connected to an emergency rescue system. The vehicle-mounted terminal is used to execute the control method as described in any one of claims 1 to 8.
10. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the control method as described in any one of claims 1 to 8.