Vehicle unlocking method, device and equipment based on multi-modal large model and medium
By fusing multimodal features from vehicle environment image data using a multimodal large model, the system proactively predicts user intent, solving the accuracy and robustness issues of existing intelligent vehicle unlocking solutions and achieving higher unlocking accuracy and privacy protection.
Patent Information
- Application Number
- CN202511098538.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2025-11-18
AI Technical Summary
Existing intelligent vehicle unlocking solutions are not very accurate in recognizing user intent and are prone to mis-locking. They are particularly lacking in robustness in dynamic movement scenarios and under obstructed conditions, and also pose a risk of privacy leakage.
The system employs a multimodal large model to fuse motion trajectories, body postures, and facial dynamic features from vehicle environmental image data. It uses the multimodal large model for intent recognition to proactively predict users' vehicle unlocking needs. The system is then fine-tuned by combining users' historical data and preferences, and the confidence threshold is dynamically adjusted to control vehicle unlocking.
It improves the accuracy of intelligent vehicle unlocking, avoids accidental locking, enhances robustness in dynamic and obstructed scenarios, reduces the risk of privacy leaks, and improves the accuracy of recognizing user unlocking intentions.
Smart Images

Figure CN120977026A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of vehicles, and particularly relates to a vehicle unlocking method and device based on a multi-modal large model, equipment and a medium. BACKGROUND
[0002] With the development of intelligent networked vehicle technology, the vehicle unlocking mode gradually evolves from the traditional mechanical key, remote key to the Bluetooth unlocking, near field communication (NFC) unlocking, mobile phone application (App) remote control, static face recognition and fingerprint recognition and other intelligent unlocking schemes. However, most of the intelligent unlocking schemes in the related art only trigger the vehicle unlocking through a single action such as "entering a fixed area", which may trigger the unlocking in the case that a passerby mistakenly enters the fixed area but actually does not need to unlock the vehicle, resulting in poor accuracy of the intelligent unlocking vehicle. SUMMARY
[0003] The main purpose of the embodiments of the present application is to propose a vehicle unlocking method, device, equipment and medium based on a multi-modal large model, which aims to actively predict user intent based on a multi-modal large model to determine the user's vehicle unlocking demand, thereby improving the accuracy of the intelligent unlocking vehicle.
[0004] To achieve the above purpose, a first aspect of the embodiments of the present application proposes a vehicle unlocking method based on a multi-modal large model, the method comprising:
[0005] obtaining image data of an environment where a vehicle is located;
[0006] performing multi-modal feature extraction based on the image data to obtain a multi-modal feature vector of a person in the image data;
[0007] performing intent recognition based on a multi-modal large model to fuse the multi-modal feature vector to obtain an intent recognition result of the person in the image data;
[0008] controlling the vehicle unlocking based on the intent recognition result.
[0009] In some embodiments, the multi-modal feature vector includes a motion trajectory feature vector, a limb posture feature vector and a face dynamic feature vector.
[0010] The multi-modal feature extraction based on the image data comprises:
[0011] In the case where it is monitored that the person in the image data enters a vehicle perception area, the motion trajectory feature vector, the limb posture feature vector and the face dynamic feature vector of the person in the image data are extracted based on the image data.
[0012] In some embodiments, the extracting, based on the image data, a motion trajectory feature vector, a limb posture feature vector and a face dynamic feature vector of a person in the image data comprises:
[0013] performing optical flow calculation on the continuous multiple frames of the image data to obtain an optical flow vector representing motion information of the person in the image data, and performing spatio-temporal feature fusion on the optical flow vector to extract the motion trajectory feature vector of the person in the image data;
[0014] performing key point detection on the person in the image data to obtain key point coordinates of the person in the image data, and performing posture semantic analysis based on the key point coordinates to extract the limb posture feature vector of the person in the image data;
[0015] performing face recognition on the person in the image data to obtain aligned face data of the person in the image data, and performing dynamic feature encoding on the face data to extract the face dynamic feature vector of the person in the image data.
[0016] In some embodiments, the performing, based on the multi-modal large model, fusion on the multi-modal feature vectors to perform intent recognition to obtain an intent recognition result of the person in the image data comprises:
[0017] performing self-attention fusion encoding on the motion trajectory feature vector, the limb posture feature vector and the face dynamic feature vector based on a multi-modal large model to obtain a fused context feature; the context feature represents the correlation between the motion trajectory, the limb posture and the face dynamic of the person in the image data;
[0018] performing intent recognition on the context feature based on the multi-modal large model to obtain an intent recognition result of the person in the image data.
[0019] In some embodiments, the multi-modal large model is deployed on one side of the vehicle, and the method further comprises:
[0020] obtaining historical unlocking record data and user preference data of the vehicle;
[0021] performing fine-tuning training on the multi-modal large model based on the historical unlocking record data and the user preference data to obtain a fine-tuned multi-modal large model;
[0022] performing fusion encoding on the multi-modal feature vectors based on the fine-tuned multi-modal large model.
[0023] In some embodiments, the intent recognition result includes an unlocking intent confidence, and the controlling the vehicle to unlock based on the intent recognition result comprises:
[0024] The confidence level of the unlocking intention is compared with a preset confidence threshold to obtain the comparison result;
[0025] If the comparison result indicates that the confidence level of the unlocking intent is greater than or equal to the confidence threshold, facial authentication is performed on the person in the image data to obtain the facial authentication result;
[0026] If the facial recognition result indicates that the facial recognition is successful, a vehicle unlocking command is generated to control the vehicle to unlock.
[0027] In some embodiments, the method further includes:
[0028] Obtain vehicle user history behavior data;
[0029] The confidence threshold is dynamically adjusted based on the user's historical behavior data to obtain the adjusted confidence threshold.
[0030] The step of comparing the confidence level of the unlocking intent with the preset confidence threshold is performed based on the adjusted confidence threshold.
[0031] In some embodiments, performing facial authentication on persons in the image data includes:
[0032] Obtain facial template features from the vehicle's local facial feature database;
[0033] The facial dynamic feature vector in the multimodal feature vector is matched with the face template feature to perform face authentication on the person in the image data.
[0034] To achieve the above objectives, a second aspect of this application provides a vehicle unlocking device based on a multimodal large model, the device comprising:
[0035] The data acquisition module is used to acquire image data of the vehicle's environment;
[0036] A multimodal feature extraction module is used to extract multimodal features based on the image data to obtain multimodal feature vectors of people in the image data;
[0037] The multimodal large model inference module is used to perform intent recognition based on the fusion of the multimodal feature vectors of the multimodal large model, and to obtain the intent recognition result of the person in the image data;
[0038] The control execution module is used to control the vehicle to unlock based on the intent recognition result.
[0039] To achieve the above object, a third aspect of the embodiments of the present application provides a vehicle unlocking device based on a multi-modal large model, which comprises a memory and a processor, the memory stores a computer program, and the processor implements the multi-modal large model based vehicle unlocking method of the first aspect when executing the computer program.
[0040] To achieve the above object, a fourth aspect of the embodiments of the present application provides a vehicle, which is configured with a vehicle unlocking device based on a multi-modal large model, the vehicle unlocking device based on a multi-modal large model comprises a memory and a processor, the memory stores a computer program, and the processor implements the multi-modal large model based vehicle unlocking method of the first aspect when executing the computer program.
[0041] To achieve the above object, a fifth aspect of the embodiments of the present application provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the multi-modal large model based vehicle unlocking method of the first aspect.
[0042] To achieve the above object, a sixth aspect of the embodiments of the present application provides a computer program product, which comprises a computer program, and the computer program is executed by a processor to implement the multi-modal large model based vehicle unlocking method of the first aspect.
[0043] The multi-modal large model based vehicle unlocking method, device, equipment, vehicle, computer readable storage medium and computer program product provided by the present application, by acquiring image data of an environment where a vehicle is located, performing multi-modal feature extraction based on the image data to obtain a multi-modal feature vector of a person in the image data, performing intention recognition based on a multi-modal large model to fuse the multi-modal feature vector to obtain an intention recognition result of the person in the image data, and controlling the vehicle to be unlocked based on the intention recognition result.
[0044] Compared to methods that trigger vehicle unlocking solely through the single action of "entering a fixed area," this embodiment acquires image data of the vehicle's environment, extracts multimodal feature vectors of individuals within that image data, and then uses a multimodal large-scale model to fuse these feature vectors for intent recognition. This yields the intent recognition result of the individuals in the image data, and finally, the vehicle unlocking is controlled based on this intent recognition result. Thus, this embodiment can proactively predict user intent. Even when the vehicle user (the individual in the image data) has not actively operated (e.g., pulled the door handle) to unlock the vehicle, the multimodal large-scale model fuses their multimodal feature vectors for behavioral analysis to determine in advance whether they have a need to unlock the vehicle. In other words, this embodiment proactively predicts user intent based on a multimodal large-scale model to determine the user's need to unlock the vehicle, thereby improving the accuracy of intelligent vehicle unlocking.
[0045] Furthermore, the embodiments of this application actively predict user intent based on the fusion of multimodal feature vectors of users using a multimodal large model. This can also avoid the problem of insufficient robustness of vehicle unlocking based on static face recognition in related technologies in scenarios involving side faces, occlusion (such as wearing masks), and dynamic movement. It achieves robust intent determination in multiple scenarios with interference factors such as occlusion, changes in lighting, and dynamic behavior, thereby improving the accuracy of recognizing user intent to unlock vehicles and further enhancing the accuracy of intelligent vehicle unlocking. Attached Figure Description
[0046] Figure 1 A flowchart illustrating the steps of the vehicle unlocking method based on a multimodal large model provided in this application in some embodiments;
[0047] Figure 2 The vehicle unlocking method based on a multimodal large model provided in the embodiments of this application is illustrated in the detailed steps of spatiotemporal feature analysis of image data in some embodiments.
[0048] Figure 3 for Figure 1 A detailed flowchart of step S103;
[0049] Figure 4 A flowchart illustrating the steps of the vehicle unlocking method based on a multimodal large model provided in this application in some other embodiments;
[0050] Figure 5 for Figure 1 A detailed flowchart of step S104;
[0051] Figure 6 A flowchart illustrating the steps of the vehicle unlocking method based on a multimodal large model provided in this application in some other embodiments;
[0052] Figure 7 A module cooperation flowchart involved in a complete embodiment of the vehicle unlocking method based on a multi-modal large model provided by the embodiments of the present application is shown in the following figure:
[0053] Figure 8 A multi-modal large model inference flowchart involved in a complete embodiment of the vehicle unlocking method based on a multi-modal large model provided by the embodiments of the present application is shown in the following figure:
[0054] Figure 9 A timing logic diagram involved in a complete embodiment of the vehicle unlocking method based on a multi-modal large model provided by the embodiments of the present application is shown in the following figure:
[0055] Figure 10 A structure diagram of the vehicle unlocking device based on a multi-modal large model provided by the embodiments of the present application is shown in the following figure:
[0056] Figure 11 A hardware structure diagram of the vehicle unlocking device based on a multi-modal large model provided by the embodiments of the present application is shown in the following figure. DETAILED DESCRIPTION
[0057] In order to make the purpose, technical scheme and advantages of the present application clearer, the present application will be further described in detail below in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.
[0058] It should be noted that although the functional modules are divided in the device schematic diagram, and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the device or the order in the flowchart. The terms "first", "second", etc. in the specification and claims and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence.
[0059] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application, and are not intended to limit the present application.
[0060] First, the overall concept of the vehicle unlocking method based on a multi-modal large model provided by the embodiments of the present application is described.
[0061] With the development of intelligent networked vehicle technology, the vehicle unlocking mode gradually evolves from traditional mechanical key, remote key to Bluetooth unlocking, NFC near field unlocking, mobile phone App remote control, static face recognition, fingerprint recognition and other intelligent solutions. However, the existing solutions still have many pain points. For example: passive triggering, that is, the user needs to actively operate (such as pulling the door handle, taking out the mobile phone) to complete the vehicle unlocking, lacking the ability to "predict in advance" the user's intention. For another example: single mode dependence, that is, the traditional visual unlocking solution is only based on static face features, which is easily affected by occlusion, light changes, and imitation attacks (such as photos / videos), with high misrecognition rate. For another example: poor scene adaptability, that is, it cannot distinguish between "user approaching to get something", "passing by the vehicle", "stranger probing" and other different behavior intentions, resulting in mislocking the vehicle or unnecessary security response. For another example: privacy and delay problems, that is, some solutions rely on cloud computing, which requires uploading user biometric features to the cloud, thereby there is a risk of privacy leakage, and the network delay of transmitting data may also affect the timeliness of unlocking the vehicle.
[0062] In related technologies, the vehicle active unlocking solution based on vision deploys a monocular camera outside the vehicle to capture images of personnel approaching the vehicle, and detects whether the personnel enters a preset "wake-up area" (such as 2-3 meters away from the vehicle) in the image through a traditional computer vision algorithm (such as HOG features + SVM classifier), if so, triggers static face recognition, so that after the face recognition matches the pre-stored user feature library, the vehicle unlocking operation is performed.
[0063] In this way, only through the single action of "entering the fixed area" to trigger the vehicle unlocking, it is actually impossible to distinguish the user's real intention, for example, it may trigger unlocking in the case of a passerby mistakenly entering the fixed area but actually not needing to unlock the vehicle, thereby leading to poor accuracy of intelligent unlocking of the vehicle. In addition, static face recognition relies on clear frontal face images, and has insufficient robustness to profile, occlusion (such as wearing a mask), and dynamic moving scenes.
[0064] To this end, the embodiments of the present application provide a vehicle unlocking method and device based on a multi-modal large model, equipment, vehicle, computer readable storage medium and computer program product, aiming to overcome the shortcomings of the above-mentioned related technologies, and actively predict the user's intention based on a multi-modal large model to determine the user's vehicle unlocking demand, thereby improving the accuracy of intelligent unlocking of the vehicle.
[0065] Considering that the logic of determining whether the user unlocks the vehicle by processing the image data of a single modality in the traditional scheme is simple, resulting in a high triggering rate, embodiments of the present application perform multi-dimensional feature fusion reasoning by fusing user behavior context (such as walking trajectory, limb movement), obtain image data of the environment where the vehicle is located, extract a multi-modal feature vector of the person in the image data based on the image data, perform intent recognition based on the multi-modal large model by fusing the multi-modal feature vector, obtain the intent recognition result of the person in the image data, and finally control the vehicle to unlock based on the intent recognition result. Thus, embodiments of the present application can actively predict the user's intent, and when the user does not actively operate (such as does not pull the door handle) to unlock the vehicle, the multi-modal large model is used to fuse the multi-modal feature vector of the user to perform behavior analysis, so as to determine whether the user has a vehicle unlocking demand in advance. That is, when the vehicle is unlocked, embodiments of the present application actively predict the user's intent based on the multi-modal large model to determine the user's vehicle unlocking demand, thereby improving the accuracy of intelligent unlocking of the vehicle.
[0066] In addition, based on the multi-modal large model fusing the multi-modal feature vector of the user to actively predict the user's intent, embodiments of the present application can also avoid the problem of insufficient robustness in related technologies based on static face recognition to unlock the vehicle in side face, shielding (such as wearing a mask), dynamic moving scene, and achieve robust intent determination in multiple scenes with shielding, light changes, dynamic behaviors and other interference factors, thereby improving the recognition accuracy of the user's unlocking vehicle intent and further improving the accuracy of intelligent unlocking of the vehicle.
[0067] It should be noted that in embodiments of the present application, the multi-modal large model can be a pre-trained multi-modal large model, or a multi-modal large model specially fine-tuned for the pre-trained multi-modal large model. The multi-modal large model can be, for example, a multi-modal encoder based on the deep learning model architecture Transformer, such as ViT (Vision Transformer) + long short-term memory network LSTM architecture. In some embodiments, a lightweight multi-modal model (such as MobileNet + LSTM) can be used instead of the multi-modal large model, thereby further reducing the computing resource requirement of the vehicle end, but the accuracy of intent reasoning for complex scenes (such as multiple people approaching at the same time) may decrease.
[0068] Next, the vehicle unlocking method, device, equipment, vehicle, computer readable storage medium and computer program product based on the multi-modal large model provided by embodiments of the present application are specifically described as follows, and first, the vehicle unlocking method based on the multi-modal large model provided by embodiments of the present application is described in detail.
[0069] It should be noted that in various specific embodiments of the present application, when relevant processing needs to be performed on data related to the identity or characteristics of the user, such as user information, user behavior data, user history data, and user location information, the user's permission or consent will be obtained first, and the collection, use, and processing of such data will comply with relevant laws, regulations, and standards. In addition, when the embodiments of the present application need to obtain sensitive personal information of the user, the separate permission or separate consent of the user will be obtained through a pop-up window or a jump to a confirmation page, and after obtaining the separate permission or separate consent of the user, the necessary user-related data for the normal operation of the embodiments of the present application will be obtained.
[0070] It should be noted that the vehicle unlocking method based on a multi-modal large model provided by the embodiments of the present application can be applied to a terminal, can also be applied to a server end, and can also be software running in a terminal or a server end. In some embodiments, the terminal can be a vehicle-mounted terminal (such as a vehicle-mounted computing platform) on a vehicle, or a computer device such as a smartphone, a tablet computer, a notebook computer, a desktop computer, etc. associated with the vehicle, and the terminal is associated with the vehicle means that the terminal can communicate and interact with the vehicle based on a network. The server end can be a background server terminal device of the vehicle, which can be configured as an independent physical server, or as a server cluster or distributed system composed of multiple physical servers, or as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content distribution networks (CDN), and big data and artificial intelligence platforms. The software can be an application, a computer program, and a storage medium carrying the computer program, etc. that implements the vehicle unlocking method based on a multi-modal large model. It should be understood that the terminal, server end, and software, etc. applying the vehicle unlocking method based on a multi-modal large model provided by the embodiments of the present application can of course also be other forms not listed here, and the vehicle unlocking method based on a multi-modal large model provided by the embodiments of the present application is not specifically limited in this regard.
[0071] Moreover, embodiments can be practiced in a distributed computing environment where tasks are performed by a remote processing device that is located at a remote computer system, over a communications network. In a distributed computing environment, program modules can be located in both local and remote computer storage media including memory storage devices.
[0072] For the sake of understanding and elaboration, the terminal device (directly configured on the vehicle or associated with the vehicle) applies the vehicle unlocking method based on the multi-modal large model provided by the embodiments of the present application in the following. Each specific embodiment of the present application is described in detail. Any form of subject matter described above applies the vehicle unlocking method based on the multi-modal large model provided by the embodiments of the present application. The process of the terminal device applying the vehicle unlocking method based on the multi-modal large model can be referred to the process of the terminal device applying the vehicle unlocking method based on the multi-modal large model described below.
[0073] Please refer to Figure 1 , Figure 1 The step flow diagram of the vehicle unlocking method based on the multi-modal large model provided by the embodiments of the present application in some embodiments. It should be understood that, although Figure 1 and subsequent other step flow diagrams show the execution order of some method steps, the vehicle unlocking method based on the multi-modal large model provided by the embodiments of the present application can of course adopt an execution order of method steps different from that shown in the figure. That is, Figure 1 The order of the method steps shown does not constitute a limitation on the execution logic order of the vehicle unlocking method based on the multi-modal large model provided by the embodiments of the present application. Any reasonable change in the order of the method steps shown Figure 1 should be included in the protection scope of the vehicle unlocking method based on the multi-modal large model provided by the embodiments of the present application.
[0074] As Figure 1 shown, in some embodiments, the vehicle unlocking method based on the multi-modal large model provided by the embodiments of the present application can include the steps S101 to S104 as shown below.
[0075] Step S101: Obtain image data of the environment in which the vehicle is located.
[0076] The terminal device continuously collects image data of the environment in which the vehicle is located based on the vehicle-side camera during parking or running of the vehicle.
[0077] It should be noted that the vehicle-side camera can be 3-5 wide-angle cameras deployed outside the vehicle, which cover the entire vehicle. It should be understood that the 3-5 wide-angle cameras listed here are only examples, and the specific number of cameras can be flexibly adjusted based on actual needs.
[0078] In some embodiments, the terminal device can continuously collect image data of the environment in which the vehicle is located by the vehicle-side camera frame by frame.
[0079] In other embodiments, the terminal device can also collect continuous video stream data of the environment in which the vehicle is located by the vehicle-side camera, for example, the terminal device collects continuous video stream by 3-5 wide-angle cameras covering the entire vehicle at 30fps (frame rate). Then, the terminal device obtains continuous multi-frame image data in the video stream data.
[0080] In some embodiments, the terminal device can collect continuous video stream by the data collection module in the system architecture (also referred to as the device) that actively predicts the user's intention to unlock the vehicle through 3-5 wide-angle cameras deployed outside the vehicle at 30fps.
[0081] It should be noted that the system architecture can be a vehicle-side localized processing architecture, and the modules in the system architecture communicate through a high-speed vehicle bus (such as CAN / LIN), thereby ensuring real-time data transmission. Among them, the hardware deployment of the data collection module can be: 3-5 wide-angle cameras with a resolution of 1920x1080 (field of view angle 120°-180°) are deployed outside the vehicle in a ring shape, and the specific layout of these cameras can be: 1 in front bumper (covering the front of the vehicle and the area 5 meters in front), 1 in each side mirror (covering an area of 3 meters on both sides), and 1 in rear bumper (covering the rear of the vehicle and the area 5 meters behind). In addition, these cameras are connected to the terminal device through GMSL (Gigabit Multimedia Serial Link), supporting continuous video stream collection at 30fps (video code rate of each road is about 8Mbps). In this way, by selecting wide-angle cameras to cover the entire vehicle without dead angles, and using 30fps frame rate for video stream data collection, real-time performance (single frame processing time ≤ 33ms) and motion trajectory continuity (avoiding frame-to-frame blur when moving at high speed) are considered. Moreover, the collected video stream can be directly input to the vehicle-mounted graphics processing unit (GPU), such as NVIDIA Orin, without compression storage, reducing latency.
[0082] Step S102: performing multi-modal feature extraction based on the image data to obtain a multi-modal feature vector of the person in the image data.
[0083] After obtaining the image data of the environment in which the vehicle is located, the terminal device can recognize the walking trajectory, body movement, and face dynamics of the person in the image data, perform multi-modal feature extraction on the person in the image data, and thus obtain a multi-modal feature vector of the person in the image data.
[0084] In some embodiments, the terminal device can extract multi-dimensional features such as the action trajectory (such as walking direction and speed), body posture (such as hand raising and approaching the vehicle head), and facial features (dynamic expression and gaze direction) of the person in the image data based on the image data in the continuous video stream captured by the vehicle camera, and fuse time sequence information (such as the behavior change of the person in the image data within 5 seconds) to form a multi-modal feature vector of the person in the image data.
[0085] In some embodiments, the terminal device can perform multi-modal feature extraction on the person in the image data through a multi-modal feature extraction module in the system architecture. The multi-modal feature extraction module can extract the action trajectory (optical flow method to calculate the motion direction), body posture (OpenPose key point detection), and face dynamic features (MTCNN+FaceNet to extract continuous frame face features) of the person in the image data (such as a video stream) based on ResNet+3D-CNN.
[0086] In some embodiments, the terminal device can also fuse data such as the motion speed and distance of the person based on the millimeter wave radar while capturing image data based on the camera in a multi-sensor (such as camera+millimeter wave radar) fusion manner, fuse these data with the image data, and then perform multi-modal feature extraction based on the fused data to obtain a multi-modal feature vector of the person in the image data. In this way, the problem of poor image data clarity in low-light scenarios such as night, rain, and fog, which makes it difficult to accurately extract features such as the action trajectory and body posture of the person, can be avoided, and the robustness of the user intent determination in low-light scenarios such as night, rain, and fog can be further improved.
[0087] Step S103: performing intent recognition based on the multi-modal large model and the multi-modal feature vector to obtain an intent recognition result of the person in the image data.
[0088] After the terminal device extracts the multi-modal feature vector of the person in the image data, the terminal device inputs the multi-modal feature vector into the multi-modal large model, fuses the multi-modal feature vector based on the multi-modal large model, jointly encodes the action, behavior, and face data of the person in the image data, captures the correlation between different modal feature vectors (for example, "accelerating to the driver's seat + gaze at the door handle" is more likely to be an unlocking intention than "passing by the co-driver at a constant speed"), and performs intention "context understanding" recognition on the person in the image data, to obtain an intention recognition result of the person in the image data.
[0089] In some embodiments, the terminal device can perform intention recognition based on the multi-modal large model fusing the multi-modal feature vector of the person in the image data through a multi-modal large model inference module in the system architecture. The multi-modal large model inference module deploys a lightweight multi-modal large model (such as a ViT+LSTM architecture) on the vehicle side, fuses the multi-modal feature vector based on the multi-modal large model inputting the multi-modal feature vector, and performs intention "context understanding" on the person in the image data based on the multi-modal large model.
[0090] Step S104: Controlling the vehicle to be unlocked based on the intention recognition result.
[0091] After the terminal device obtains the intention recognition result of the person in the image data based on the multi-modal large model actively predicting the user's intention, if the intention recognition result indicates that the person in the image data has a vehicle unlocking demand, the terminal device immediately sends an unlocking instruction to a body controller (BCM), to control the vehicle to be unlocked.
[0092] In some embodiments, if the intention recognition result of the person in the image data does not indicate that the person in the image data has a vehicle unlocking demand, the terminal device can control the vehicle to maintain a locked state.
[0093] In some embodiments, if the intention recognition result of the person in the image data does not indicate that the person in the image data has a vehicle unlocking demand, the terminal device can also trigger a preset abnormal handling mechanism, that is, send a 0x00 instruction and trigger "secondary confirmation". The terminal device displays prompt information "Do you want to unlock? Nod to confirm" on the vehicle screen, and triggers an unlocking instruction to control the vehicle to be unlocked after detecting a user's nodding action through the camera. In this way, the secondary confirmation mechanism can reduce the risk of misjudgment and further improve the accuracy of intelligent unlocking of the vehicle.
[0094] In some embodiments, the terminal device can send an unlocking instruction to the vehicle body controller through the control execution module in the system architecture to control the vehicle to unlock. Among them, the control execution module can convert the intention recognition result indicating that the personnel in the image data have a vehicle unlocking demand into an unlocking instruction, so as to control the vehicle body controller to unlock the vehicle. The control execution module can communicate with the vehicle body controller BCM through the CAN bus (500 kbps rate). In addition, the unlocking instruction is a CAN frame with ID = 0x123, and the data field is 1 byte (0x01 = unlock, 0x00 = maintain locking).
[0095] In some embodiments, the terminal device can set the total delay from the intention recognition result output to the unlocking instruction sending to be ≤500ms through the control execution module, wherein the multi-modal feature extraction is 150ms, the multi-modal large model inference is 100ms, the face authentication after the intention recognition result output is 50ms, and the communication is 200ms, thereby meeting the experience demand of the vehicle user "close to unlock". Among them, the 500ms delay conforms to the ergonomics (upper limit of user imperceptible waiting time).
[0096] In the embodiments of the present application, the terminal device continuously collects image data of the environment where the vehicle is located based on the vehicle-end camera during the parking or running of the vehicle. Then, the terminal device recognizes the walking track, body movement and face dynamics of the personnel in the image data to extract multi-modal features of the personnel in the image data, thereby obtaining a multi-modal feature vector of the personnel in the image data. Then, the terminal device inputs the multi-modal feature vector into a multi-modal large model, thereby fusing the multi-modal feature vector based on the multi-modal large model, capturing the association between different modal feature vectors (for example, "accelerating towards the driver's seat + gaze at the door handle" is more likely to be an unlocking intention than "uniformly passing the co-driver"), and identifying the intention "context understanding" of the personnel in the image data, thereby obtaining an intention recognition result of the personnel in the image data. Finally, the terminal device sends an unlocking instruction to the vehicle body controller BCM immediately when the intention recognition result indicates that the personnel in the image data have a vehicle unlocking demand, thereby controlling the vehicle to unlock.
[0097] Therefore, compared with a manner of triggering the active unlocking of the vehicle only through a single action of "entering the fixed area", the embodiment of the present application can actively predict the user's intention, and when the user of the vehicle (the person in the image data) does not actively operate (for example, pulls the door handle) to unlock the vehicle, the multi-modal large model is used to fuse the multi-modal feature vectors of the user to perform behavior analysis, so as to determine whether the user has a vehicle unlocking demand in advance. That is, when the vehicle is unlocked, the embodiment of the present application actively predicts the user's intention based on the multi-modal large model to determine the user's vehicle unlocking demand, so that the accuracy of the intelligent unlocking of the vehicle can be improved.
[0098] In addition, based on the multi-modal large model and the fusion of the multi-modal feature vectors of the user, the embodiment of the present application can actively predict the user's intention, and can also avoid the problem of insufficient robustness in a side face, shielding (such as wearing a mask), and dynamic moving scene in the related art based on static face recognition to unlock the vehicle. Robust intention determination can be achieved in multiple scenes with interference factors such as shielding, light changes, and dynamic behaviors, so that the recognition accuracy of the user's intention to unlock the vehicle is improved, and the accuracy of the intelligent unlocking of the vehicle is further improved.
[0099] In some embodiments, the multi-modal feature vector can include a motion trajectory feature vector, a body posture feature vector, and a face dynamic feature vector. In this case, the step of "performing multi-modal feature extraction based on the image data" in step S102 described above can include but is not limited to the following steps:
[0100] In the case where it is monitored that the person in the image data enters the vehicle perception area, the motion trajectory feature vector, the body posture feature vector, and the face dynamic feature vector of the person in the image data are extracted based on the image data.
[0101] It should be noted that the vehicle perception area can be a 5-meter radius annular area (covering the effective field of view of the camera) centered on the vehicle and pre-set by the terminal device. Among them, the terminal device sets the radius of the area to 5 meters, which can be mainly based on two considerations. First, at a normal walking speed (1.2 m / s) of a person, 5 meters takes about 4 seconds, which can reserve enough time for subsequent feature extraction and inference. Second, the camera can clearly capture a face within 5 meters (0.1 m x 0.1 m face corresponds to an image resolution of about 20 x 20 pixels, which meets the lower limit of MTCNN detection). It should be understood that the terminal device can of course set other sizes of the vehicle perception area for the various functional modules of the perception wake-up system architecture based on different design needs of actual applications.
[0102] When collecting image data, the terminal device can continuously monitor whether people in the image data enter the vehicle perception area (e.g., within 5 meters of the vehicle). When it detects that people in the image data have entered the vehicle perception area, the terminal device triggers multimodal feature extraction to extract the motion trajectory feature vector, limb posture feature vector, and facial dynamic feature vector of the people in the image data by performing spatiotemporal feature analysis on the people in the image data.
[0103] For example, the terminal device can capture video at a low frame rate of 1fps by default using a vehicle-mounted camera (to reduce computational power consumption), and detect the "pedestrian" category in real time using a lightweight object detection model (such as YOLOv5s, with 2.7M parameters). When the vehicle-mounted camera detects a pedestrian's bounding box entering a 5-meter area (measured by monocular vision): Where f is the camera focal length, H is the average height of people (1.7m), and h is the number of pixels representing the height of pedestrians in the image, the camera frame rate is immediately increased to 30fps to acquire image data, and a "wake-up signal" is sent to the terminal device. After the terminal device responds to the wake-up signal and wakes up, the multimodal feature extraction module, inference module, etc., switch from sleep mode to full-speed operation (GPU / TPU computing power usage increases from 10% to 70%).
[0104] In this embodiment, monitoring pedestrians at a low frame rate using the vehicle-mounted camera reduces the vehicle's static power consumption (≤5W). Then, dynamically waking up the terminal device to perform feature extraction and subsequent inference based on the image data avoids high computing power consumption around the clock, meeting the energy efficiency constraints of the vehicle system. Furthermore, by maintaining low power consumption when no user is nearby, and waking up the core computing module to extract multimodal features upon detecting a potential user entering the sensing area, energy efficiency and response speed can be effectively balanced.
[0105] Please refer to Figure 2 , Figure 2 The vehicle unlocking method based on a multimodal large model provided in this application is illustrated in some embodiments with a detailed flowchart of spatiotemporal feature analysis of image data.
[0106] In some embodiments, such as Figure 2 As shown, the above-mentioned step of "extracting the motion trajectory feature vector, limb posture feature vector and facial dynamic feature vector of the person in the image data based on the image data" may include steps S201 to S203 as shown below.
[0107] Step S201: optical flow calculation is performed on the continuous multiple frames of image data to obtain optical flow vectors representing the motion information of the person in the image data, and spatio-temporal feature fusion is performed on the optical flow vectors to extract the motion trajectory feature vector of the person in the image data.
[0108] When the terminal device performs spatio-temporal feature analysis on the person in the image data to extract the multi-modal feature vector of the person in the image data, the terminal device can perform optical flow calculation on the acquired continuous multiple frames of image data by using an optical flow algorithm to obtain optical flow vectors representing the motion information (motion direction and speed, etc.) of the person in the image data. Then, the terminal device further performs spatio-temporal feature fusion on the optical flow vectors, and captures the time dimension motion pattern between the multiple frames of image data to extract the motion trajectory feature vector of the person in the image data.
[0109] For example, the terminal device can use the dense optical flow algorithm Farneback in the multi-modal feature extraction module in the system framework to calculate the pixel-level optical flow field of the continuous 3 frames of image data (t-1, t, t+1) in the video stream, thereby extracting the center of gravity of the person's contour in the image data (by using the background difference method to segment the foreground), and counting the displacement vector (Δx, Δy) and speed (Δx / Δt, Δy / Δt) of the center of gravity in the 3 frames of image data, thereby obtaining the optical flow vectors representing the motion information (motion direction and speed, etc.) of the person in the image data, such as approaching / away from the vehicle, and 0.5 m / s→1.2 m / s indicating acceleration and approach.
[0110] After that, the terminal device inputs the optical flow vectors into the 3D-CNN (such as I3D network, 3x3x3 convolution kernel) through the multi-modal feature extraction module to perform spatio-temporal feature fusion on the optical flow vectors, so as to capture the time dimension motion pattern (such as the abnormal trajectory of "straight approaching→sudden turning") between the 3 frames of image data, thereby obtaining the 128-dimensional action trajectory feature vector output by the 3D-CNN.
[0111] Step S202: key point detection is performed on the person in the image data to obtain the key point coordinates of the person in the image data, and pose semantic analysis is performed based on the key point coordinates to extract the limb pose feature vector of the person in the image data.
[0112] When the terminal device performs spatio-temporal feature analysis on the person in the image data to extract the multi-modal feature vector of the person in the image data, the terminal device can also perform key point detection on the person in the image data by using a key point detection algorithm, so as to obtain key point coordinates (for example, skeletal key point coordinates) of the person in the image data. Then, the terminal device performs pose semantic analysis based on the key point coordinates, that is, normalizes the key point coordinates, and calculates joint angles and relative positions, so as to extract a limb pose feature vector of the person in the image data.
[0113] Exemplarily, the terminal device can use the OpenPose network (VGG19 backbone) based on deep learning human pose estimation technology through the multi-modal feature extraction module in the system framework, predict 18 limb key point heat maps (such as shoulders, elbows, wrists, hips, knees, and ankles) and part of the affinity field (PAF, a vector field connecting key points) of the person in the image data through a multi-stage CNN, and output skeletal key point coordinates (x, y, confidence) of the person. Then, the terminal device normalizes the key point coordinates (with the hip as the origin) through the multi-modal feature extraction module, calculates joint angles (such as the elbow angle = the included angle of wrist-elbow-shoulder) and relative positions (such as the vertical distance between the wrist and the door handle), and extracts pose features (256 dimensions) through ResNet-34, so as to identify behavior features such as “arm lifting” (wrist higher than shoulder by 20 cm) and “slowing down” (ankle key point displacement ≤5 cm / frame) performed by the person in the image data. These behavior features are limb pose feature vectors of the person in the image data.
[0114] Step S203: performing face recognition on the person in the image data to obtain aligned face data of the person in the image data, and performing dynamic feature coding on the face data to extract a facial dynamic feature vector of the person in the image data.
[0115] When the terminal device performs spatio-temporal feature analysis on the person in the image data to extract the multi-modal feature vector of the person in the image data, the terminal device can perform face detection and alignment on the continuous multiple frames of image data to obtain aligned face data of the person in the image data. Then, the terminal device further performs dynamic feature coding on the aligned face data, so as to extract a facial dynamic feature vector of the person in the image data.
[0116] Exemplarily, the terminal device can use a three-stage cascade network (P-Net candidate frame generation→R-Net screening→O-Net refinement) of a face detection network MTCNN through a multi-modal feature extraction module in the system architecture to detect a face region in 5 continuous image data frames, and perform affine transformation alignment (calibrate rotation and scaling) based on the eye corner and nose tip key points to output an aligned face of 160*160 pixels (aligned face data). Then, the terminal device inputs the aligned face into FaceNet (Inception-ResNet-v1 backbone) through the multi-modal feature extraction module to perform dynamic feature coding, extract a 512-dimensional static feature vector, and use a long short-term memory network LSTM to code the time dimension of the static feature vectors extracted from the 5 image data frames (capture the dynamic process of “turning head→fixing gaze on the door handle”), thereby outputting a 256-dimensional face dynamic feature vector (fusion of static features and time changes).
[0117] Please refer to Figure 3 , Figure 3 for Figure 1 a detailed step flowchart of step S103.
[0118] In some embodiments, as shown in Figure 3 , the above step S103: performing intent recognition on the multi-modal feature vectors based on a multi-modal large model to obtain an intent recognition result of a person in the image data can include steps S301 and S302 as shown below.
[0119] Step S301: performing self-attention fusion coding on the motion trajectory feature vector, the limb posture feature vector, and the face dynamic feature vector based on a multi-modal large model to obtain a fused context feature; the context feature represents the correlation between the motion trajectory, the limb posture, and the face dynamic of the person in the image data.
[0120] When the terminal device performs intent recognition on the multi-modal feature vectors of a person in the image data based on a multi-modal large model, it can input the motion trajectory feature vector, the limb posture feature vector, and the face dynamic feature vector of the person in the image data extracted into the multi-modal large model, so that the multi-modal large model performs self-attention fusion coding on the motion trajectory feature vector, the limb posture feature vector, and the face dynamic feature vector based on a self-attention mechanism to obtain a fused context feature. Among them, the fused context feature can represent the correlation between the motion trajectory, the limb posture, and the face dynamic of the person in the image data, for example, the attention weight of “accelerating and approaching (action) + arm lifting (posture)” is higher than that of “passing at a constant speed (action) + arm drooping (posture)”.
[0121] Step S302: performing intent recognition on the context feature based on the multi-modal large model, to obtain an intent recognition result of the person in the image data.
[0122] After the terminal device obtains the fused context feature, the terminal device further performs intent recognition on the context feature based on the multi-modal large model, to perform context understanding on the intent of the person in the image data, and to obtain an intent recognition result of the person in the image data.
[0123] Exemplarily, the intent recognition result of the person in the image data can be an unlock intent confidence output by the multi-modal large model (a confidence score of 0 to 1, for example, a value greater than or equal to 0.8 can represent a "high probability of unlock intent"). The terminal device can receive a 640-dimensional multi-modal feature vector based on an input layer of the multi-modal large model in the system architecture, and map the multi-modal feature vector to a 3x256-dimensional modal-specific embedding (256-dimensional for each of action, posture, and face) based on the input layer through linear projection, and add a positional encoding (Positional Encoding) to identify the time sequence (for example, a feature sequence of the first second to the fifth second). Then, the ViT encoder of the multi-modal large model adopts an 8-head self-attention mechanism (Multi-Head Attention) to calculate the correlation between the modal embeddings (for example, the attention weight of "accelerating and approaching (action) + arm lifting (posture)" is higher than that of "uniformly passing by (action) + arm lowering (posture)"), and outputs a fused 3x256-dimensional context feature. Then, the 2-layer LSTM of the multi-modal large model performs time modeling, that is, the sequence feature output by the ViT encoder is input into the 2-layer LSTM (hidden layer of 256 dimensions) to capture the behavior evolution of the person in the image data within 5 seconds (for example, the coherent intent of "approaching → pausing → raising hand"), and output a 256-dimensional temporal context feature. Finally, the output layer of the multi-modal large model outputs a "unlock intent confidence" of 0 to 1 (for example, 0.92, indicating a high probability of unlocking) through a fully connected layer (256→1) + a Sigmoid activation function.
[0124] In this embodiment, the model architecture of the multi-modal large model is designed based on ViT and LSTM, so that the self-attention mechanism of ViT explicitly models the multi-modal correlation, and LSTM captures the time dependence, and the combination of the two improves the reasoning ability of the multi-modal large model for complex intents.
[0125] Please refer to Figure 4 , Figure 4 The step flowchart of the vehicle unlocking method based on the multi-modal large model provided in the embodiments of the present application in another embodiment.
[0126] In some embodiments, the multi-modal large model is deployed on the side of the vehicle. In this case, as shown in FIG. 8, the multi-modal large model-based vehicle unlocking method provided by the embodiments of the present application can further include steps S401 to S403 as shown below. Figure 4
[0127] Step S401: Obtain historical unlocking record data and user preference data of the vehicle.
[0128] The terminal device can deploy the multi-modal large model on the side of the vehicle, for example, on a vehicle-side computing platform (such as a vehicle-mounted GPU / TPU), so as to support local feature extraction and intent reasoning. Among them, the terminal device can optimize the large model calculation amount through model lightening technology (such as knowledge distillation, quantization compression), and at the same time, the terminal device can obtain vehicle-side real-time data such as historical unlocking record data and user preference data of the vehicle locally, for fine-tuning the multi-modal large model.
[0129] It should be noted that the user preference data can be user self-set preference data obtained by the terminal device through vehicle user-oriented human-computer interaction results in advance. Alternatively, the terminal device can also automatically generate user preference data for storage by performing deep learning processing on the vehicle use behavior of the vehicle user through the large model.
[0130] Step S402: Fine-tune the multi-modal large model based on the historical unlocking record data and the user preference data, and obtain a fine-tuned multi-modal large model.
[0131] After obtaining the historical unlocking record data and the user preference data of the vehicle, the terminal device continuously fine-tunes the multi-modal large model using the historical unlocking record data and the user preference data, and obtains a fine-tuned multi-modal large model. Among them, the fine-tuned multi-modal large model will be more in line with the personalized behavior of the vehicle user (for example, a user habit of "stopping for 1 second before unlocking", and the model can learn this pattern) when determining the intent.
[0132] Step S403: Fuse and encode the multi-modal feature vector based on the fine-tuned multi-modal large model.
[0133] After obtaining the fine-tuned multi-modal large model, the terminal device can use the fine-tuned multi-modal large model to perform the above-mentioned operation step of fusing and encoding the multi-modal feature vector of the person in the image data in the current round of intent reasoning based on the multi-modal large model, so as to obtain an intent recognition result that is more in line with the personalized behavior of the person in the image data.
[0134] After obtaining the fine-tuned multi-modal large model, the terminal device can also use the fine-tuned multi-modal large model to perform the operation step of fusing and encoding the multi-modal feature vectors of the person in the image data in the next round of intent reasoning based on the multi-modal large model, so as to obtain an intent recognition result that is more in line with the personalized behavior of the person in the image data.
[0135] In some embodiments, when the terminal device optimizes the large model calculation amount through model lightening technology such as knowledge distillation and quantization compression, the terminal device can use a cloud large model (12-layer Transformer+4-layer LSTM) as a teacher model, train a vehicle terminal small model (4-layer Transformer+2-layer LSTM), and learn the soft label of the teacher model (such as decomposing the confidence of 0.92 into the contribution weight of each modality) through KL divergence loss. In addition, when the terminal device quantizes and compresses the multi-modal large model, the terminal device can quantize the floating point parameter (FP32) to INT8, reduce the memory occupation (the model volume is reduced from 200MB to 50MB), and optimize the quantization error through a calibration data set (1000 groups of real scene data). In addition, the terminal device can also cut the redundant heads (from 8 heads to 4 heads) of the self-attention mechanism of the multi-modal large model based on the importance score of the attention head (such as the attention weight of a head to the "gaze direction" <0.1), and only keep the core reasoning logic.
[0136] In this embodiment, by using lightening technology to optimize and process the multi-modal large model, the terminal device can ensure the real-time performance (inference delay ≤100ms) of the multi-modal large model on the vehicle-mounted GPU (such as Jetson Orin Nano).
[0137] Please refer to Figure 5 , Figure 5 for Figure 1 the detailed step flowchart of step S104.
[0138] In some embodiments, as shown in Figure 5 , when the intent recognition result is an unlock intent confidence, the step S104 of controlling the vehicle to unlock based on the intent recognition result can include steps S501 to S503 as shown below.
[0139] Step S501: comparing the unlock intent confidence with a preset confidence threshold to obtain a comparison result.
[0140] It should be noted that the preset confidence threshold can be a threshold preset for the terminal device, which is used to determine whether the confidence of the unlocking intention output by the multi-modal large model based on the fused multi-modal features is triggered to control the vehicle to be unlocked, that is, in the case that the confidence of the unlocking intention is greater than or equal to the confidence threshold, the terminal device confirms to trigger the process of controlling the vehicle to be unlocked.
[0141] In the case that the terminal device performs intention recognition on the motion trajectory feature vector, the body posture feature vector and the face dynamic feature vector through the multi-modal large model to obtain the confidence of the unlocking intention of the person in the image data (for example, 0.92) as the intention recognition result, when controlling the vehicle to be unlocked based on the intention recognition result, first, the confidence of the unlocking intention is compared with the confidence threshold (for example, any value between 0.85 and 0.90) to obtain the comparison result about the size between the confidence of the unlocking intention and the confidence threshold.
[0142] Step S502: In the case that the comparison result indicates that the confidence of the unlocking intention is greater than or equal to the confidence threshold, performing face authentication on the person in the image data to obtain a face authentication result.
[0143] After the terminal device obtains the comparison result about the size between the confidence of the unlocking intention and the confidence threshold, if the comparison result indicates that the confidence of the unlocking intention is greater than or equal to the confidence threshold, in this case, the terminal device triggers the face authentication to enter the process of controlling the vehicle to be unlocked, that is, the terminal device immediately performs face authentication on the person in the image data to obtain a face authentication result.
[0144] Step S503: In the case that the face authentication result indicates that the face authentication is passed, generating a vehicle unlocking instruction to control the vehicle to be unlocked.
[0145] After the terminal device obtains the face authentication result of performing face authentication on the person in the image data, if the face authentication result indicates that the face authentication is passed, the terminal device immediately generates a vehicle unlocking instruction and issues it to the corresponding vehicle body controller in this case, so as to control the vehicle to be unlocked, for example, by sending the unlocking instruction to the vehicle body controller through the control execution module in the system architecture to control the vehicle to be unlocked.
[0146] In some embodiments, before comparing the confidence of the unlocking intention output by the multi-modal large model with the confidence threshold, the terminal device can also dynamically adjust the confidence threshold, and then use the adjusted confidence threshold to compare with the confidence of the unlocking intention.
[0147] Please refer to Figure 6 , Figure 6The method for unlocking a vehicle based on a multi-modal large model provided in the embodiments of the present application can further include steps S601-S603 as shown in the flowchart of the step process in some embodiments.
[0148] In some embodiments, as shown in Figure 6 The method for unlocking a vehicle based on a multi-modal large model provided in the embodiments of the present application can further include steps S601-S603 as shown in the flowchart of the step process in some embodiments.
[0149] Step S601: Obtain user historical behavior data of the vehicle.
[0150] The terminal device can obtain the user historical behavior data of the vehicle from a vehicle local record database, for example, the user of the vehicle often quickly approaches the vehicle, etc.
[0151] Step S602: Dynamically adjust the confidence threshold based on the user historical behavior data to obtain an adjusted confidence threshold.
[0152] After obtaining the user historical behavior data, the terminal device can dynamically adjust the confidence threshold based on the user historical behavior data to obtain the adjusted confidence threshold. For example, based on the user historical behavior data that “the user often quickly approaches”, the confidence threshold is reduced by 0.05 to obtain the adjusted confidence threshold.
[0153] In some embodiments, the terminal device can also dynamically adjust the confidence threshold in combination with the current time. For example, the terminal device can increase the confidence threshold from 0.85 to 0.90 at night, etc.
[0154] In some embodiments, the terminal device can also dynamically adjust the confidence threshold based on the current time and the user historical behavior data to obtain the adjusted confidence threshold. For example, the terminal device can first increase the confidence threshold from 0.85 to 0.90 at night, and then further reduce the threshold 0.90 by 0.05 based on the historical behavior that “the user often quickly approaches” to obtain the adjusted confidence threshold.
[0155] Step S603: Perform the step of comparing the unlocking intention confidence with the preset confidence threshold based on the adjusted confidence threshold.
[0156] After obtaining the adjusted confidence threshold, the terminal device can use the adjusted confidence threshold to perform the step of “comparing the unlocking intention confidence with the preset confidence threshold” in step S501.
[0157] In this embodiment, by dynamically adjusting the confidence threshold by the terminal device in combination with the user historical behavior and time factors, the risk of false unlocking can be reduced, thereby further improving the accuracy of intelligent unlocking of the vehicle.
[0158] In some embodiments, when the terminal device performs face authentication on the person in the image data, the terminal device can perform fast face authentication through the local feature library of the vehicle, thereby achieving privacy protection on the person in the image data.
[0159] It should be noted that the local feature library can be constructed by the terminal device pre-collecting face data of the user of the vehicle through human-computer interaction, for example, the terminal device can collect 100-200 pieces of multi-scene face data of the user when the user registers the vehicle, including face data of front face / side face ± 30°, face data of day / night, face data of wearing a mask / wearing glasses, etc., then extract 512-dimensional feature vectors through the face recognition technology FaceNet, and take the average of the features of the same type of scene (such as “front face in the day” template = average of 10 pieces of front face in the day features), and finally generate 8 types of face template features (4 angles × 2 scenes) to construct the face feature library, which can be stored in the vehicle safety chip (such as eSE, supporting AES-256 encryption).
[0160] In some embodiments, the step S502 of “performing face authentication on the person in the image data” can include the following steps:
[0161] obtaining face template features from the local face feature library of the vehicle;
[0162] performing feature matching between the face dynamic feature vector in the multi-modal feature vector and the face template features to perform face authentication on the person in the image data.
[0163] When the terminal device performs face authentication on the person in the image data, the terminal device first obtains the face template features of the user of the vehicle from the local face feature library of the vehicle, and then performs feature matching between the face dynamic feature vector in the multi-modal feature vector of the person in the image data extracted previously and the face template features, thereby performing fast privacy-protected face authentication on the person in the image data locally in the vehicle.
[0164] In some embodiments, the terminal device can perform fast face authentication on the person in the image data based on the local face feature library of the vehicle through the face authentication module in the above-mentioned system architecture.
[0165] For example, the terminal device inputs the face dynamic features (256 dimensions) of the person in the image data to the face authentication module, and the face authentication module calculates the cosine similarity between the input face dynamic features (256 dimensions) and the face template features (512 dimensions) in the face feature library to obtain the face authentication result.
[0166] In some embodiments, the face authentication module can perform multi-template fusion face authentication operations, for example, simultaneously matching the facial dynamic features (256 dimensions) with the "side face at night" and "wearing a mask" face template features, and then taking the highest similarity obtained by matching. In addition, the face authentication module can also perform occlusion adaptation face authentication operations. For example, when the terminal device pre-stores the "wearing a mask" face template feature in the face feature library, the feature invariance of the unoccluded face region (such as the eye, eyebrow, and nose) is learned through the model, that is, the high correlation of the features before and after the occlusion is constrained through the feature pairs of "full face → wearing a mask" in the training data. In this way, when the face authentication module performs face authentication on the personnel in the image data, it automatically matches the facial dynamic features with the closest occlusion template, for example, detects that the facial dynamic features are occluded by a mask, and the face authentication module preferentially matches the "wearing a mask" face template feature.
[0167] In this embodiment, by constructing a face feature library with multiple face template features to cover the real use scenarios of vehicle users, the matching failure of a single scene template can be avoided, thereby further improving the accuracy of intelligent unlocking of the vehicle.
[0168] Next, the complete embodiment of the vehicle unlocking method based on the multi-modal large model provided by the embodiments of the present application is proposed.
[0169] First, the complete embodiment of the vehicle unlocking method based on the multi-modal large model using the above-mentioned system architecture and various functional templates on the vehicle terminal hardware platform is proposed.
[0170] Please refer to Figure 7 and Figure 8 , Figure 7 The module coordination process schematic diagram involved in the vehicle unlocking method based on the multi-modal large model provided by the embodiments of the present application in a complete embodiment; Figure 8 The multi-modal large model inference process schematic diagram involved in the vehicle unlocking method based on the multi-modal large model provided by the embodiments of the present application in a complete embodiment.
[0171] As shown in Figure 7 , when the terminal device performs intelligent unlocking of the vehicle, the data acquisition module on the vehicle terminal hardware platform provides the original video, that is, 3-5 wide-angle cameras are deployed outside the vehicle (covering the four directions around the vehicle) to collect continuous video streams at 30fps. Then, the multi-modal feature extraction module analyzes the spatio-temporal features: based on ResNet+3D-CNN to extract the motion trajectory (optical flow method to calculate the motion direction), the body posture (OpenPose key point detection), and the face dynamic features (MTCNN+FaceNet to extract the face features of continuous frames) in the video. Then, the large model inference module outputs the intention confidence. At this time, as shown inFigure 8 As shown, the multi-modal large model deployed at the vehicle end (such as the ViT+LSTM architecture) inputs the multi-modal feature vectors (action trajectory features, limb posture features, and face dynamic features), encodes them through the self-attention mechanism for cross-modal management, then fuses the feature vectors for spatio-temporal feature analysis, and finally outputs the “unlock intention confidence” (a confidence score between 0 and 1, such as 0.92).
[0172] Then, if the unlock intention confidence is greater than or equal to a threshold, the face authentication module calls the local face feature library (pre-stored multi-angle and multi-scene face features of the user) for 1:1 matching to verify the identity through face authentication. After the face authentication is passed, the control execution module sends an unlock instruction to the vehicle body controller to complete the vehicle unlocking. In addition, if the unlock intention confidence is insufficient (less than the threshold) or the face authentication fails, the control execution module controls the vehicle to maintain the locked state.
[0173] In this embodiment, the entire process of vehicle intelligent unlocking performed by the vehicle end hardware platform using various functional modules is processed locally at the vehicle end, thereby balancing privacy protection, real-time response, and the robustness of complex scene intention recognition.
[0174] The core process of the vehicle unlocking method based on a multi-modal large model provided by the embodiments of the present application will be described in detail below in combination with the timing and technical implementation.
[0175] Please refer to Figure 9 , Figure 9 The timing logic diagram involved in the vehicle unlocking method based on a multi-modal large model provided by the embodiments of the present application in a complete embodiment.
[0176] As Figure 9 shown, when the terminal device applies the vehicle unlocking method based on a multi-modal large model provided by the embodiments of the present application for vehicle intelligent unlocking, in the perception wake-up stage, the camera collects vehicle external video in real time, and when detecting that a person enters the “perception area” (5 meters away from the vehicle), multi-modal feature extraction is triggered.
[0177] Then, when the terminal device performs multi-modal feature extraction, for the action trajectory, the terminal device calculates the motion vectors of the continuous 3 frames through the optical flow method to determine whether the person is continuously approaching the vehicle; for the limb posture, the terminal device detects whether the arm is raised (such as preparing to pull the door handle) and whether the pace is slowed down (such as staying to observe); finally, for the face dynamic features, the terminal device extracts the face features (including the contour, relative positions of eyes, nose, and mouth) of the continuous 5 frames to generate a dynamic feature vector.
[0178] It should be noted that the terminal device can start the multi-modal feature extraction within 100 ms after self-waking up when performing multi-modal feature extraction, and can process the action trajectory, body posture and face dynamic three types of features in parallel to provide key input for subsequent intent reasoning.
[0179] Among them, the action trajectory feature extraction is the quantification of the user's motion intention, and the goal is to determine whether the personnel is "continuously approaching the vehicle" or to exclude "passing by" and "temporary stay" and other non-unlocking behaviors. Specifically, the terminal device can use the optical flow field calculation, foreground segmentation and motion trend determination technical process to extract the action trajectory feature. Among them, the optical flow field calculation is: using the Farneback dense optical flow algorithm on the continuous 3 frames of video (t-1, t, t+1), to generate a 2D optical flow vector diagram (the displacement direction and rate of each pixel); the foreground segmentation is to extract the moving foreground (person outline) by the background difference method (cumulative 30 frames to generate a background model), calculate its barycenter coordinates ((xc, yc) = N1∑(xi, yi)), N is the number of foreground pixels); the motion trend determination is to calculate the barycenter displacement between 3 frames (Δx = xt+1-xt-1, Δy = yt+1-yt-1), combined with the camera perspective projection matrix, to convert into actual physical displacement (such as Δx = 10 pixels → ΔX = 0.5 m); if the displacement of the continuous 2 groups of 3 frames (i.e. 6 frames, 0.2 seconds) are all directed to the center of the vehicle (the included angle between the displacement direction and the center of the vehicle is ≤ threshold 30°), it is marked as "continuous approach". Among them, the included angle threshold 30° is a key parameter to ensure that only the motion directly towards the vehicle is captured, avoiding the misjudgment of side passing; the continuous determination window of 0.2 seconds is a key parameter to balance the real-time and accuracy.
[0180] Furthermore, body posture feature extraction can be used to analyze the semantics of user behavior, with the goal of identifying postures strongly related to unlocking, such as "raising an arm" or "slowing down." Specifically, terminal devices can use keypoint detection and posture parameter calculation techniques to extract body posture features. During keypoint detection, the terminal device can use the OpenPose network to detect 18 limb keypoints (points with a confidence level ≥ 0.5 are retained) and output coordinates (x, y). Then, during pose parameter calculation, for the parsing of arm raising action, the terminal device can calculate the vertical distance between the wrist (keypoint 4 / 7) and the shoulder (keypoint 2 / 5) (\\(y_{wrist}-y_{shoulder}\\)). If it is > the raising threshold of 20cm (based on image scale transformation) and lasts for 2 frames (0.067 seconds), it is judged as "arm raising". For the parsing of foot slowing action, the terminal device can calculate the displacement of the left and right ankles (keypoint 15 / 16) in 3 consecutive frames (\\(\\sqrt{(x_{t+1}-x_t)^2+(y_{t+1}-y_t)^2}\\)). If it is ≤ 5cm / frame (corresponding to actual speed ≤ 0.15m / s), it is judged as "foot slowing" (a typical behavior when approaching a doorknob). The lifting threshold is designed to be 20cm, which is ergonomic (the common lifting height for pulling door handles), and the continuous judgment of 0.067 seconds can avoid misjudgment due to hand tremors.
[0181] Furthermore, facial dynamic feature extraction can correlate user identity with their intent over time. Its goal is to capture dynamic processes such as "turning head → looking at a doorknob" and "from side profile to frontal view," distinguishing users from strangers. Specifically, terminal devices can employ techniques such as face sequence alignment and dynamic feature encoding to extract facial dynamic features. The face sequence alignment can be performed as follows: The terminal device detects face regions in 5 consecutive frames (one face per frame, confidence ≥ 0.8) based on MTCNN, performs affine transformation based on the corners of the eyes (key points 1-2) and the tip of the nose (key point 3), and outputs a 160×160 pixel aligned face. For dynamic feature encoding, for single-frame features, the terminal device can extract 512-dimensional static features based on FaceNet (L2 normalization, modulus = 1; normalization ensures the stability of subsequent similarity calculations). For temporal features, the terminal device can input the 5 frames of static features into an LSTM (256-dimensional hidden layer) to capture the changing trend of the feature sequence (e.g., the cosine similarity of the feature vector changes from 0.6 to 0.9, indicating that the face changes from a side profile to a frontal view), and output a 256-dimensional dynamic feature vector. The time step of the LSTM can be set to 5 (corresponding to 0.17 seconds), thus balancing the integrity and real-time performance of the dynamic process.
[0182] After the above multi-modal features are extracted, the terminal device further performs intention reasoning based on the multi-modal features, that is, the terminal device performs fusion coding on the above multi-modal features based on the multi-modal large model, and outputs an intention confidence (such as 0.92). In addition, the terminal device first performs dynamic threshold determination, that is, the threshold is adjusted according to time (such as a night threshold of 0.85→0.90) and user history (such as a certain user often approaches quickly, and the threshold is reduced by 0.05), and then the terminal device compares the intention confidence output by the multi-modal large model with the adjusted threshold, and if the intention confidence is greater than or equal to the threshold, the terminal device enters face authentication.
[0183] It should be noted that the terminal device can adopt a fusion decision of multi-modal features in the stage of intention reasoning, that is, a 640-dimensional (128+256+256) multi-modal feature vector is received, and a "unlock intention confidence" (0-1) is output by a lightweight large model, and the time consumption is less than or equal to 100 ms. Specifically, the input layer of the multi-modal large model splits the 640-dimensional feature into three categories of action (128), posture (256), and face (256), and projects them into 256-dimensional embeddings (retaining modality specificity) through linear layers; then, the multi-modal large model performs self-attention fusion on the input, that is, the correlation between modalities is calculated using 8 heads of self-attention mechanism (such as the attention weight of "continuous approach (action) + arm lifting (posture)" = 0.7, which is much higher than the 0.2 of "uniform speed passing (action) + arm drooping (posture)"), and outputs a fused context feature; then, the multi-modal large model performs time modeling: the LSTM processes the feature sequence within 5 seconds (1 group of features is input every 0.1 second), and captures the coherent intention of "approaching → pausing → lifting hand" (such as the LSTM output of the sequence feature from 0.3→0.6→0.9); finally, the multi-modal large model outputs the confidence, that is, based on a fully connected layer (256→1) + Sigmoid activation, the final confidence (such as 0.92 indicating a high probability of unlocking intention) is output.
[0184] In this embodiment, it is found through comparative experiments that the model of fusion self-attention and LSTM has a higher recognition accuracy (92%) for complex intentions than the LSTM (85%) or the self-attention (88%) alone.
[0185] It should be noted that the terminal device can adopt a scene-adaptive decision calibration when performing dynamic threshold determination, that is, the determination threshold is adjusted according to the environment and user habits to balance the "unlock success rate" and the "misunlock risk". Specifically, the terminal device can adopt the threshold adjustment strategy as shown below:
[0186] Time dimension: poor light conditions at night (18:00-6:00), face detection and posture recognition errors increase, threshold value increases from the default 0.85 to 0.90 (reduces the probability of unlocking);
[0187] User history: record the user's last 30 unlocking behaviors (such as "quick approach (average approach time 2 seconds)" and "slow approach (average 5 seconds)") through the vehicle-mounted system, reduce the threshold value of "quick approach" users by 0.05 (0.85→0.80), and improve the unlocking speed;
[0188] Environmental interference: in rainy and snowy weather (detected by vehicle-mounted sensors), the error of the optical flow method increases, and the threshold value increases by 0.03 (0.85→0.88).
[0189] Exemplarily, the terminal device can pre-store the threshold value in the vehicle-mounted EEPROM, and dynamically adjust it through the rule engine (IF-THEN) when making a dynamic threshold judgment (such as "time = night → threshold + 0.05"). The adjustment logic can be updated through OTA.
[0190] After that, when performing face authentication, the terminal device compares the dynamic face features and the pre-stored features (supporting side face and wearing mask scenarios, learning feature invariance under occlusion through large models) locally in the vehicle. If the face feature matching is successful, the vehicle is controlled to be unlocked. Among them, the terminal device can perform face authentication and unlocking operation based on privacy-protected identity verification, that is, if the confidence ≥ dynamic threshold (such as 0.92≥0.90), the terminal device starts face authentication, and the overall startup time is ≤50ms. In addition, in the process of face authentication, the terminal device can calculate the cosine similarity (sim=(a.b), because the features have been normalized, the length=1) between the face dynamic features (256 dimensions) and the pre-stored templates (8 categories, such as "front face in daytime", "side face at night", "wearing mask") to perform feature matching. Then, the terminal device performs multi-template fusion, that is, the similarity of 8 templates is calculated at the same time during matching, and the maximum value is taken as the final score (such as "wearing mask" template similarity=0.85, higher than other templates, final score=0.85). Finally, if the similarity ≥0.7 (industry general threshold), the terminal device sends an unlock instruction (ID=0x123, data=0x01) to the BCM through the CAN bus, and the door is unlocked within 1 second. If it fails (for example, similarity<0.7), the terminal device maintains the lock and triggers a secondary confirmation (the vehicle screen prompts "nod to confirm unlocking", and the instruction is sent after detecting the user's nodding action through posture recognition).
[0191] In addition, the terminal device further determines the locking intention after controlling the vehicle to be unlocked, that is, if the user appears behaviors such as “back away + frequent look back” and “stay for 2 seconds after closing the door” after unlocking the vehicle, the terminal device determines the user's intention as “locking intention” based on the model, and thus actively triggers the vehicle to be locked.
[0192] It should be noted that the terminal device determines the locking intention by using the same operation process as the aforementioned active prediction of the user's unlocking intention, that is, the dynamic reasoning of the user's intention by the multi-modal large model when the user locks the vehicle. Alternatively, the terminal device can use other models to determine the user's intention after the user unlocks, so as to determine the user's intention as “locking intention” and thus actively trigger the vehicle to be locked. The multi-modal large model or other models can be trained by 1000 sets of labeled data (positive samples: user's return behavior after forgetting to lock; negative samples: normal leaving without return), and the verification set accuracy reaches 90%.
[0193] The determination of the locking intention by the terminal device can realize active and safe closed-loop optimization, that is, after the user unlocks, the scenario of “need to actively lock” can be identified by behavior analysis, and the risk of forgetting to lock can be avoided. Among them, the terminal device can use different determination logics to determine the locking intention of the user for different scenarios. For example, scenario 1: abnormal behavior when leaving: the user unlocks and leaves the vehicle, the terminal device detects “back away” (displacement direction away from the vehicle, speed > 0.5 m / s) by the motion trajectory module, and at the same time, the body posture module detects “frequent look back” (head rotation angle > 30°, frequency ≥ 2 times / s), the model determines “possible forgetting to lock”, and outputs the locking confidence 0.85; scenario 2: stay after closing the door: the user closes the door (triggered by the door sensor signal), and if the terminal device detects “stay ≥ 2 seconds” (ankle key point displacement ≤ 2 cm / frame) by the body posture module, the model determines “confirm locking”, and outputs the confidence 0.90. Regardless of the determination logic used by the terminal device to determine the locking intention of the user, the same execution strategy can be used to control the vehicle to be locked. For example: if the locking confidence ≥ 0.8, send the locking instruction (ID = 0x123, data = 0x00) to the BCM, and prompt “automatically locked” through the vehicle machine screen.
[0194] In this embodiment, by adopting the above process cooperation: whole process timing such as: perception wake-up (≤200 ms) → feature extraction (150 ms) → intention reasoning (100 ms) → threshold determination (10 ms) → authentication (50 ms) → execution (50 ms), the total delay ≤560 ms, which can meet the experience demand of the user "unlocking as close". Moreover, each module avoids repeated transmission by sharing feature data through the vehicle-mounted memory, and the CAN bus can adopt a high-priority queue (ID=0x123 priority higher than ordinary signals), to ensure the real-time performance of the instruction, so as to guarantee the overall performance of the active pre-judgment of the user's unlocking / locking intention for vehicle unlocking / locking control. That is, through the closed-loop design of "dynamic wake-up-multi-modal feature extraction-adaptive reasoning-privacy authentication-active locking", the embodiment can realize the whole cycle intelligent service from "user close intention perception" to "locking safety guarantee", and balance the real-time performance, robustness and user experience.
[0195] Please refer to Figure 10 The embodiment of the present application also provides a vehicle unlocking device based on a multi-modal large model, which can realize the vehicle unlocking method based on the multi-modal large model.
[0196] As Figure 10 shown, the vehicle unlocking device based on the multi-modal large model provided by the embodiment of the present application includes a data acquisition module 1001, a multi-modal feature extraction module 1002, a multi-modal large model reasoning module 1003 and a control execution module 1004. Among them,
[0197] The data acquisition module 1001 is configured to acquire image data of an environment in which a vehicle is located.
[0198] The multi-modal feature extraction module 1002 is configured to perform multi-modal feature extraction based on the image data to obtain a multi-modal feature vector of a person in the image data.
[0199] The multi-modal large model reasoning module 1003 is configured to perform intention recognition based on a multi-modal large model to fuse the multi-modal feature vector to obtain an intention recognition result of the person in the image data.
[0200] The control execution module 1004 is configured to control vehicle unlocking based on the intention recognition result.
[0201] In some embodiments, image data of an environment in which a vehicle is located is acquired.
[0202] Multi-modal feature extraction is performed based on the image data to obtain a multi-modal feature vector of a person in the image data.
[0203] Intention recognition is performed based on a multi-modal large model to fuse the multi-modal feature vector to obtain an intention recognition result of the person in the image data.
[0204] controlling the vehicle to be unlocked based on the intention recognition result.
[0205] In some embodiments, the multi-modal feature vector includes a motion trajectory feature vector, a body posture feature vector, and a facial dynamic feature vector; the multi-modal feature extraction module 1002 is further configured to, in a case where it is monitored that a person in the image data enters a vehicle perception area, extract a motion trajectory feature vector, a body posture feature vector, and a facial dynamic feature vector of the person in the image data based on the image data.
[0206] In some embodiments, the multi-modal feature extraction module 1002 is further configured to perform optical flow calculation on continuous multiple frames of the image data to obtain an optical flow vector representing motion information of the person in the image data, and perform spatio-temporal feature fusion on the optical flow vector to extract a motion trajectory feature vector of the person in the image data; perform key point detection on the person in the image data to obtain key point coordinates of the person in the image data, and perform posture semantic analysis based on the key point coordinates to extract a body posture feature vector of the person in the image data; and perform face recognition on the person in the image data to obtain aligned face data of the person in the image data, and perform dynamic feature coding on the face data to extract a facial dynamic feature vector of the person in the image data.
[0207] In some embodiments, the multi-modal large model inference module 1003 is further configured to perform self-attention fusion coding on the motion trajectory feature vector, the body posture feature vector, and the facial dynamic feature vector based on a multi-modal large model to obtain a fused context feature; the context feature represents a correlation relationship between the motion trajectory, the body posture, and the facial dynamic of the person in the image data; and perform intention recognition on the context feature based on the multi-modal large model to obtain an intention recognition result of the person in the image data.
[0208] In some embodiments, the multi-modal large model is deployed on a side of a vehicle, and the multi-modal large model inference module 1003 is further configured to obtain historical unlocking record data and user preference data of the vehicle; perform fine-tuning training on the multi-modal large model based on the historical unlocking record data and the user preference data to obtain a fine-tuned multi-modal large model; and perform fusion coding on the multi-modal feature vector based on the fine-tuned multi-modal large model.
[0209] In some embodiments, the intention recognition result includes an unlock intention confidence, and the control execution module 1004 is further configured to compare the unlock intention confidence with a preset confidence threshold to obtain a comparison result; in a case where the comparison result indicates that the unlock intention confidence is greater than or equal to the confidence threshold, perform face authentication on the person in the image data to obtain a face authentication result; and in a case where the face authentication result indicates that the face authentication is passed, generate a vehicle unlock instruction to control the vehicle to be unlocked.
[0210] In some embodiments, the control execution module 1004 is further configured to obtain user historical behavior data of the vehicle; dynamically adjust the confidence threshold based on the user historical behavior data to obtain an adjusted confidence threshold; and perform the step of comparing the unlock intention confidence with the preset confidence threshold based on the adjusted confidence threshold.
[0211] In some embodiments, the control execution module 1004 is further configured to obtain a face template feature from a face feature library locally of the vehicle; and perform feature matching between the face dynamic feature vector in the multi-modal feature vector and the face template feature to perform face authentication on the person in the image data.
[0212] It should be noted that the specific embodiments of the vehicle unlocking device based on the multi-modal large model provided by the embodiments of the present application are basically the same as the specific embodiments of the vehicle unlocking method based on the multi-modal large model described above, and will not be repeated here.
[0213] Please refer to Figure 11 The embodiments of the present application also provide a vehicle unlocking device based on a multi-modal large model. The vehicle unlocking device based on a multi-modal large model comprises a memory and a processor. The memory stores a computer program. The processor implements the evaluation method of the visual language model when executing the computer program.
[0214] In some embodiments, the vehicle unlocking device based on a multi-modal large model can be any intelligent terminal such as a vehicle-mounted hardware platform (e.g., a vehicle-mounted computer), a tablet computer, a smart phone, a wearable device, etc.
[0215] As Figure 11 shown, the vehicle unlocking device based on a multi-modal large model provided by the embodiments of the present application can comprise:
[0216] The processor 1101 can be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits, and is configured to execute related programs to implement the technical solutions provided by the embodiments of the present application.
[0217] The memory 1102 can be implemented by a ROM (ReadOnly Memory), a static storage device, a dynamic storage device, or a RAM (Random Access Memory), and the like. The memory 1102 can store an operating system and other application programs. When the technical solutions provided by the embodiments of the present application are implemented by software or firmware, the related program codes are stored in the memory 1102 and are called and executed by the processor 1101 to implement the vehicle unlocking method based on the multi-modal large model.
[0218] The input / output interface 1103 is configured to realize information input and output.
[0219] The communication interface 1104 is configured to realize the communication interaction between the device and other devices. The communication can be realized by a wired manner (for example, a USB, a network cable, and the like) or a wireless manner (for example, a mobile network, WIFI, Bluetooth, and the like).
[0220] The bus 1105 is configured to transmit information between various components (for example, the processor 1101, the memory 1102, the input / output interface 1103, and the communication interface 1104) of the device.
[0221] The processor 1101, the memory 1102, the input / output interface 1103, and the communication interface 1104 are connected to each other through the bus 1105 to realize the communication connection between the device.
[0222] The embodiments of the present application also provide a vehicle. The vehicle is configured with a vehicle unlocking device based on a multi-modal large model. The vehicle unlocking device based on the multi-modal large model includes a memory and a processor. The memory stores a computer program. The processor implements the vehicle unlocking method based on the multi-modal large model when executing the computer program.
[0223] The embodiments of the present application also provide a computer readable storage medium. The computer readable storage medium stores a computer program. The computer program is executed by a processor to implement the vehicle unlocking method based on the multi-modal large model.
[0224] The memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state memory device. In some embodiments, the memory can optionally include memory that is remotely located with respect to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0225] The embodiment of the present application further provides a computer program product comprising a computer program which, when executed by a processor, implements steps substantially the same as the specific embodiments of the above vehicle unlocking method based on a multi-modal large model, and thus will not be described herein.
[0226] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art can know that, with the evolution of technology and the appearance of new application scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.
[0227] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and can include more or fewer steps than those shown in the figures, or combine certain steps or different steps.
[0228] The device embodiments described above are merely illustrative, and the units described as separate components can or can not be physically separated, i.e., can be located in one place or distributed on multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the embodiments of the present application.
[0229] Those skilled in the art can understand that all or some of the steps in the above disclosed method, the functional modules / units in the system and the device can be implemented as software, firmware, hardware and their appropriate combinations.
[0230] The terms "first", "second", "third", "fourth", and the like in the description and in the claims of this application, if any, are used for distinguishing between similar elements and not necessarily for describing a particular sequential or chronological order. It is to be understood that the use of the terms so termed is interchangeable under appropriate circumstances such that the embodiments of the application described herein are, for example, capable of orderly or chronological mundane operation, reverse order operation, based on circuitry availability, based on stated preference or the like, and that "default" or other orderings are thus permissible. Further, the terms "comprise", "comprising", "include", "including", and the like, are specifically intended to be open-ended. That is, references to individual steps and the like do not suhstantially exclude the presence of two or more of a given step or its integral presence in the process, method, system, article, or apparatus having been made with a wider scope. The use of notation such as "first", "second", "third", etc. does not generally limit the areas, but is used to connect like elements or to distinguish one claim from another. These terms can be used interchangeably when appropriate. Terms concerning the relative position of elements can be interpreted such that their use adheres to their normal meaning, but they can also be interpreted to mean the opposite according to specific claims.
[0231] It should be understood that, in the application, "at least one" refers to one or more, and "multiple" refers to two or more. "And / or" is used to describe the relationship between associated objects, which means that there can be three relationships, for example, "A and / or B" can represent three cases: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the front and rear associated objects. "At least one of the following" or similar expressions means any combination of these items, including any combination of single or multiple items. For example, at least one of a, b or c can mean a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0232] In several embodiments provided in the application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the above-described device embodiments are only illustrative, for example, the division of the above-mentioned units is only a logical functional division, and actual implementation can have another division manner, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed each other can be indirect coupling or communication connection through some interface, device or unit, which can be electrical, mechanical or other forms.
[0233] The units described above as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, that is, they can be located in one place, or they can be distributed on multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment of the application.
[0234] In addition, each function unit in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software function unit.
[0235] If the integrated unit is realized in the form of a software function unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application, essentially or in part, or all or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes multiple instructions used to cause a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods in the embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various other media that can store programs.
[0236] The preferred embodiments of the embodiments of the present application are described above with reference to the accompanying drawings, and are not limited to the scope of the embodiments of the present application. Any modifications, equivalent replacements and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of the embodiments of the present application.
Claims
1. A vehicle unlocking method based on a multimodal large model, characterized in that, The method includes: Acquire image data of the vehicle's surroundings; Multimodal feature extraction is performed on the image data to obtain the multimodal feature vector of the people in the image data; Intent recognition is performed by fusing the multimodal feature vectors based on a multimodal large model to obtain the intent recognition result of the person in the image data; The vehicle is unlocked based on the intent recognition result.
2. The method according to claim 1, characterized in that, The multimodal feature vectors include motion trajectory feature vectors, limb posture feature vectors, and facial dynamic feature vectors; The multimodal feature extraction based on the image data includes: When a person is detected entering the vehicle's sensing area in the image data, the motion trajectory feature vector, limb posture feature vector, and facial dynamic feature vector of the person in the image data are extracted based on the image data.
3. The method according to claim 2, characterized in that, The step of extracting motion trajectory feature vectors, limb posture feature vectors, and facial dynamic feature vectors of a person from the image data includes: Optical flow calculation is performed on multiple consecutive frames of image data to obtain optical flow vectors that characterize the motion information of people in the image data. Spatiotemporal feature fusion is then performed on the optical flow vectors to extract the motion trajectory feature vectors of people in the image data. Key point detection is performed on the people in the image data to obtain the coordinates of the key points of the people in the image data, and pose semantic parsing is performed based on the key point coordinates to extract the limb pose feature vector of the people in the image data. Face recognition is performed on the people in the image data to obtain the aligned face data of the people in the image data, and dynamic feature encoding is performed on the face data to extract the dynamic facial feature vector of the people in the image data.
4. The method according to claim 2, characterized in that, The intention recognition based on the fusion of the multimodal feature vectors using a multimodal large model to obtain the intention recognition result of the person in the image data includes: The motion trajectory feature vector, the limb posture feature vector, and the facial dynamic feature vector are self-attentionally fused and encoded based on a multimodal large model to obtain fused contextual features; the contextual features characterize the correlation between the motion trajectory, limb posture, and facial dynamics of a person in the image data. Intent recognition is performed on the contextual features based on the multimodal large model to obtain the intent recognition results of the people in the image data.
5. The method according to claim 4, characterized in that, The multimodal large model is deployed on one side of the vehicle, and the method further includes: Obtain vehicle historical unlock records and user preference data; The multimodal large model is fine-tuned and trained based on the historical unlock record data and the user preference data to obtain the fine-tuned multimodal large model; The multimodal feature vectors are fused and encoded based on the fine-tuned multimodal large model.
6. The method according to claim 1, characterized in that, The intent recognition result includes an unlock intent confidence level, and the step of controlling vehicle unlocking based on the intent recognition result includes: The confidence level of the unlocking intention is compared with a preset confidence threshold to obtain the comparison result; If the comparison result indicates that the confidence level of the unlocking intent is greater than or equal to the confidence threshold, facial authentication is performed on the person in the image data to obtain the facial authentication result; If the facial recognition result indicates that the facial recognition is successful, a vehicle unlocking command is generated to control the vehicle to unlock.
7. The method according to claim 6, characterized in that, The method further includes: Obtain historical user behavior data for vehicles; The confidence threshold is dynamically adjusted based on the user's historical behavior data to obtain the adjusted confidence threshold. The step of comparing the confidence level of the unlocking intent with the preset confidence threshold is performed based on the adjusted confidence threshold.
8. The method according to claim 6, characterized in that, The step of performing facial recognition on individuals in the image data includes: Obtain facial template features from the vehicle's local facial feature database; The facial dynamic feature vector in the multimodal feature vector is matched with the face template feature to perform face authentication on the person in the image data.
9. A vehicle unlocking device based on a multimodal large model, characterized in that, The device includes: The data acquisition module is used to acquire image data of the vehicle's environment; A multimodal feature extraction module is used to extract multimodal features based on the image data to obtain multimodal feature vectors of people in the image data; The multimodal large model inference module is used to perform intent recognition based on the fusion of the multimodal feature vectors of the multimodal large model, and to obtain the intent recognition result of the person in the image data; The control execution module is used to control the vehicle to unlock based on the intent recognition result.
10. A vehicle unlocking device based on a multimodal large model, characterized in that, The vehicle unlocking device based on a multimodal large model includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the vehicle unlocking method based on a multimodal large model as described in any one of claims 1 to 8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the vehicle unlocking method based on a multimodal large model as described in any one of claims 1 to 8.
12. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the vehicle unlocking method based on a multimodal large model as described in any one of claims 1 to 8.
Citation Information
Cited By
Method for determining risk in identity authentication process and computing equipment
CN121706074A