Vehicle cabin passenger body type classification method and device, electronic equipment and storage medium

By combining target multimodal large model with cockpit image data and indicator annotation information, the problem of high labor cost and low accuracy of traditional convolutional neural networks in vehicle cockpit occupant body shape classification is solved, and accurate occupant body shape classification and multi-task adaptation are achieved.

CN121236737APending Publication Date: 2025-12-30CHONGQING JINKANG NEW ENERGY VEHICLE CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511425088.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-30
Publication Date
2025-12-30

AI Technical Summary

Technical Problem

Traditional convolutional neural networks require extensive data collection and labeling for vehicle cabin occupant body shape classification, resulting in high labor costs and low classification accuracy.

Method used

A target multimodal large model is generated by acquiring vehicle cabin image data and indicator labeling information, and fine-tuning the initial multimodal large model. This model is used to calculate the occupant's sitting height data and determine the occupant's body type by combining posture and facial features parameters.

Benefits of technology

It enables accurate prediction of occupant sitting height data and body type classification, reduces data collection and labeling costs, improves classification accuracy, and expands the functions of seat belt wearing detection and cabin occupant number detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121236737A_ABST
    Figure CN121236737A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a vehicle cabin passenger body type classification method and device, electronic equipment and a storage medium. The method comprises the following steps: acquiring cabin image data of a target passenger in a vehicle; determining indication annotation information corresponding to the cabin image data; inputting the cockpit image data and the indication labeling information into a trained target multi-mode large model, and outputting target sitting height data of the target passenger; and determining a target body type category of the target passenger according to the target sitting height data. Through the embodiment of the invention, the target multi-mode large model can be assisted to focus the cabin reference object by introducing the indication labeling information, so that the intention of the user is understood through the description of the user, the sitting height data of the passenger is accurately predicted, and the accurate body type classification is further obtained.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of vehicles, in particular to a vehicle cabin occupant body type classification method and device, an electronic device and a storage medium. BACKGROUND

[0002] Traditional OMS occupant body type classification methods in the cabin are mostly based on traditional convolutional neural networks for training. This method can use target detection, key point extraction and other methods. However, the traditional convolutional neural network needs to collect and label a large amount of data, involving a large amount of human cost. Moreover, the accuracy of body type classification based on traditional neural networks is also low. SUMMARY

[0003] In view of the above problems, a vehicle cabin occupant body type classification method and device, an electronic device and a storage medium are provided to overcome the above problems or at least partially solve the above problems, comprising: A vehicle cabin occupant body type classification method, the method comprising: obtaining cabin image data of a target occupant in a vehicle; determining indication label information corresponding to the cabin image data; the indication label information is label information used to distinguish different cabin reference objects in the vehicle cabin; inputting the cabin image data and the indication label information into a target multi-modal large model generated by training, and outputting target seat height data of the target occupant, the target multi-modal large model being a model for calculating target seat height based on cabin image data and indication label information; determining a target body type category of the target occupant according to the target seat height data, the target body type category being a category for dividing the body type of an occupant on a seat in the cabin.

[0004] Optionally, the determination of the indication label information corresponding to the cabin image data comprises: determining a plurality of cabin reference objects based on the cabin image data; obtaining actual size information of each cabin reference object; generating corresponding indication label information for each cabin reference object according to the actual size information.

[0005] Optionally, further comprising: inputting the cabin image data and the indication label information into a target multi-modal large model generated by training, and outputting the body state parameters and / or facial feature parameters of the target occupant; The determination of the target body type category of the target occupant according to the target seat height data comprises: determine a target body type category of the target occupant based on the target sitting height data and the body shape parameter and / or the facial feature parameter.

[0006] Optionally, a corresponding relationship between sitting height data and body type categories is acquired. A target body type category corresponding to the target sitting height data is determined based on the corresponding relationship.

[0007] Optionally, before inputting the cabin image data and the indication annotation information into the target multi-modal large model generated by training, the method further comprises: processing the cabin image data and the indication annotation information according to a preset format corresponding to the target multi-modal large model.

[0008] Optionally, the processing the cabin image data and the indication annotation information according to the preset format corresponding to the target multi-modal large model comprises: performing data processing on the cabin image data according to Alpaca format data or Share GPT format data corresponding to the target multi-modal large model; performing data processing on the indication annotation information according to json format data corresponding to the target multi-modal large model.

[0009] Optionally, the method further comprises: determining a model size type of the target multi-modal large model; determining a target deployment mode corresponding to the target multi-modal large model according to the model size type; deploying the target multi-modal large model based on the target deployment mode.

[0010] Optionally, the determining the target deployment mode corresponding to the target multi-modal large model according to the model size type comprises: when the model size type is greater than or equal to 7B, determining that the target deployment mode corresponding to the target multi-modal large model is cloud deployment; when the model size type is 3B, determining that the target deployment mode corresponding to the target multi-modal large model is vehicle machine end deployment.

[0011] Optionally, the method further comprises: inputting the cabin image data and the indication annotation information into the target multi-modal large model to output a safety belt wearing state of the target occupant and / or a number of the target occupants.

[0012] A vehicle cabin occupant body type classification device, the device comprising: a cabin image data acquisition module configured to acquire cabin image data of a target occupant in a vehicle; An indication standard information determination module is configured to determine indication label information corresponding to the cabin image data, the indication label information being label information used to distinguish different cabin references in the vehicle cabin; A target sitting height data determination module is configured to input the cabin image data and the indication label information into a target multi-modal large model trained and generated, and output target sitting height data of the target occupant, the target multi-modal large model being a model for calculating target sitting height based on cabin image data and indication label information; A target body type determination module is configured to determine a target body type category of the target occupant according to the target sitting height data, the target body type category being a category classified according to the body type of the occupant on the seat in the cabin.

[0013] An electronic device includes a processor, a memory, and a computer program stored on the memory and capable of running on the processor, and the computer program, when executed by the processor, implements the vehicle cabin occupant body type classification method described above.

[0014] A computer readable storage medium stores a computer program, and the computer program, when executed by a processor, implements the vehicle cabin occupant body type classification method described above.

[0015] The embodiments of the present application have the following advantages: In the embodiments of the present application, cabin image data of a target occupant in a vehicle can be obtained, and then indication label information corresponding to the cabin image data can be determined, and then the cabin image data and the indication label information can be input into a target multi-modal large model, so as to output target sitting height data of the target occupant, and then a target body type category of the target occupant can be obtained, so that the indication label information can be introduced to assist the target multi-modal large model to focus on the cabin references, and then the user's intention can be understood through the user's description, the sitting height data of the occupant can be accurately predicted, and an accurate body type classification can be obtained.

[0016] In addition, the target multi-modal large model in the embodiments of the present application also realizes the expansion of functions such as seat belt wearing detection and the number of people in the cabin detection, and compared with the traditional model, one multi-modal large model can adapt to multiple tasks at the same time. BRIEF DESCRIPTION OF DRAWINGS

[0017] In order to more clearly illustrate the technical solutions of the present application, the following will briefly introduce the drawings needed to be used in the description of the present application. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0018] Figure 1aThis is a flowchart of the steps of a method for classifying the body types of vehicle cabin occupants according to an embodiment of the present invention; Figure 1b This is a schematic diagram illustrating various data annotations for occupants according to an embodiment of the present invention; Figure 2 This is a flowchart of another method for classifying the body types of vehicle cabin occupants according to an embodiment of the present invention; Figure 3 This is a flowchart of another method for classifying the body types of vehicle cabin occupants according to an embodiment of the present invention; Figure 4 This is a flowchart of another method for classifying the body types of vehicle cabin occupants according to an embodiment of the present invention; Figure 5 This is a schematic diagram of the structure of a vehicle cabin occupant body type classification device provided in an embodiment of the present invention. Detailed Implementation

[0019] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0020] Reference Figure 1a The diagram illustrates a flowchart of a method for classifying the body types of vehicle cabin occupants according to an embodiment of the present invention, which may specifically include the following steps: Step S101: Acquire cockpit image data of the target occupant inside the vehicle; In practical applications, occupants of different body types in a vehicle may have different sitting heights. The core concept of this invention is to introduce a target multimodal large model trained based on indicator label information and cabin image data to accurately predict the occupant's sitting height, and then predict the occupant's actual body type based on the sitting height.

[0021] In this embodiment of the invention, the vehicle may be equipped with an image recording device for recording cabin images. During vehicle operation, the image recording device can record cabin image data, including the target occupant, to facilitate the determination of the occupant's sitting height through analysis of the cabin image data. The occupant's sitting height data refers to the actual sitting height of the target occupant, that is, the height of the occupant after sitting in the seat, specifically the height from the seat to the head. Figure 1b As shown, the number 23 indicates the seating height of the occupant.

[0022] Step S102: Determine the indication annotation information corresponding to the cockpit image data; After obtaining the cockpit image data, corresponding indicator labeling information can be determined based on the cockpit image data. The indicator labeling information is used to distinguish different cockpit reference objects in the vehicle cockpit. The cockpit reference objects can be objects in the cockpit with measurable values, such as the armrest box, the vehicle roof, etc.

[0023] After determining the cockpit image data, the cockpit reference objects in the cockpit image data can be marked so that the target multimodal large model can accurately predict the sitting height of the target occupant.

[0024] In one embodiment of the present invention, the process of generating indication label information is as follows: determining multiple cockpit reference objects based on the cockpit image data; obtaining the actual size information of each cockpit reference object; and generating corresponding indication label information for each cockpit reference object according to the actual size information.

[0025] After acquiring cockpit image data, cockpit reference objects in the cockpit image data can be identified, and the actual size information of each identified cockpit reference object can be obtained. The actual size information of each cockpit reference object can be determined according to the vehicle model. Since the position and size of the basic vehicle components inside the cockpit of a uniform vehicle model are relatively fixed, some vehicle components can be selected as cockpit reference objects. The actual size of the cockpit reference object itself can be measured in advance or determined according to the design drawings of the vehicle model.

[0026] After determining the cockpit reference objects and their actual dimensions, corresponding indicator labels can be generated for each cockpit reference object. The indicator labels can include the name of the cockpit reference object, its actual dimensions, and its label frame.

[0027] In one example, different colors can be used to indicate cockpit references, and the color of each cockpit reference can be preset. For example, the armrest box can be set to a green frame.

[0028] Step S103: Input the cockpit image data and the indicator annotation information into the trained target multimodal large model, and output the target occupant's target sitting height data; The target multimodal large model is a model that calculates the target sitting height based on cockpit image data and indicator label information, obtained by fine-tuning the initial multimodal large model using a preset fine-tuning tool.

[0029] After determining the cockpit image data and indicator label information, both can be input into a pre-trained target multimodal large model. The target multimodal large model can analyze the cockpit image data and indicator label information based on the trained model parameters, and then output the target occupant's target sitting height data.

[0030] The target multimodal large model can be obtained by fine-tuning the initial multimodal large model using a pre-defined fine-tuning tool. The initial multimodal large model (e.g., the Qwen2.5-vl model) is trained on a large dataset and possesses general knowledge capabilities. Therefore, a small amount of training data can be used to train a target multimodal large model that can output the target occupant's seat height data in the vehicle based on the input cockpit image data and indicator annotation information.

[0031] In this embodiment of the invention, the initial multimodal large model can be trained according to the data of different vehicle models to obtain the target multimodal large model corresponding to different vehicle models. Then, the corresponding target multimodal large model can be called according to the vehicle model to realize the prediction of passenger seating height.

[0032] The initial training data for the multimodal large model may include: real occupant image data collected by OMS, real height and weight information of occupants, seat position information, real height of seat back, and reference information from other locations, such as the real height information of the armrest box and roof.

[0033] After collecting training data, the aforementioned multi-dimensional information can be processed and organized into a data format suitable for supervised fine-tuning, such as Alpaca or ShareGPT format. Then, based on the training data, fine-tuning training of the general initial multimodal large model can be initiated to obtain the target multimodal large model used in this embodiment of the invention for measuring occupant seating height.

[0034] In one embodiment of the present invention, the fine-tuning tool can employ the llamafactory fine-tuning tool to perform fine-tuning training on the initial multimodal large model. The llamafactory fine-tuning tool can perform both full fine-tuning and partial fine-tuning.

[0035] Specifically, the training data and the already trained initial multimodal large model can be imported into the llamafactory fine-tuning tool. In the llamafactory fine-tuning tool, a LORA (Low-Rank Adaptation of Large Language Models) matrix containing variable parameters can be constructed. During the training process, the parameters of the initial multimodal large model itself are kept unchanged. By iteratively training the variable parameters in the LORA matrix, a target multimodal large model that can calculate the target sitting height based on cockpit image data and indicator annotation information can be obtained.

[0036] In one embodiment of the present invention, before using the target multimodal large model for seat height prediction, the multimodal large model can be deployed. The process of determining the deployment method of the multimodal large model is as follows: first, the model size type of the target multimodal large model needs to be determined; the target deployment method corresponding to the target multimodal large model is determined according to the model size type; and the multimodal large model is deployed based on the target deployment method.

[0037] In this embodiment of the invention, the multimodal large model can be deployed in the cloud or on the vehicle side, and the specific deployment method can be determined according to the model size.

[0038] In one embodiment of the present invention, determining the target deployment method corresponding to the multimodal large model according to the model size type includes: when the model size type is greater than or equal to 7B, determining the target deployment method corresponding to the multimodal large model as cloud deployment.

[0039] In one embodiment of the present invention, determining the target deployment method corresponding to the multimodal large model according to the model size type includes: when the model size type is 3B, determining the target deployment method corresponding to the multimodal large model as vehicle-mounted deployment.

[0040] In practical applications, OMS data is first collected in the cockpit, simulating a standard sitting posture. The collected data is then manually labeled using drawing tools. The labeled images are then processed using manual methods or image recognition algorithms to obtain annotation information and create JSON-formatted label data. Fine-tuning tools (such as Llamafactory) are used to fine-tune the initial multimodal model. Depending on computing power, some or all parameters can be fine-tuned. The input is images and JSON-formatted label data. The final training objective is for the multimodal large model to have the following capabilities: given image data and annotation information, the multimodal large model can output approximate sitting height information of the human body in the image, and determine the body's plethysmometric shape based on this information. If a multimodal large model of 7 bytes or more is selected for fine-tuning, it can be deployed to the cloud using VLLM. If a 3-byte multimodal large model is used for fine-tuning, after quantization, it can be deployed on the vehicle's infotainment system using MNN.

[0041] In this embodiment of the invention, the process of predicting the sitting height of the target multimodal large model is as follows: "instruction": You are a vehicle infotainment system that answers user questions based on the user's description; "input": The red pull box represents the roof, and the green pull box represents the armrest box. The height of the roof from the armrest box is 110cm, and the height of the armrest box is 40cm. Please estimate the approximate seating height of the driver on the right based on these descriptions. "output": According to the user's description, the height of the roof from the armrest box is 110cm, the height of the armrest box is 40cm, and the height of the roof from the ground is 110+40=150cm. The driver on the right is about 2 / 3 of the distance between the roof and the ground, so the driver's seat height is 100cm. "images": Path to the image size storage.

[0042] The JSON format data in the target multimodal large model requires engineers to describe it based on the annotation information in the images. This can be done manually or by batch processing the annotation information in the images using a recognition algorithm. The JSON format data can be in Alpaca format, which meets the data format requirements of the llamafactory fine-tuning large model framework. Furthermore, the image content can be described in the input field of the data, and the content the model should respond to can be specified in the output field. `images` specifies the absolute path to the images, which can be the storage path for the target's elevation and height images.

[0043] Step S104: Determine the target body type category of the target occupant based on the target sitting height data.

[0044] Among them, the target body type category is a category that classifies the body types of occupants seated in the cabin.

[0045] In practical applications, there is a correlation between occupant body types and sitting height. Based on this correlation, it is possible to determine the target body type category of the corresponding occupant based on the target sitting height data.

[0046] In one embodiment of the present invention, the cockpit image data and the indication label information can also be input into the target multimodal large model generated by training, and the body posture parameters and / or facial features parameters of the target occupant can be output; thereby, the target body type category of the target occupant can be determined based on the target sitting height data and the body posture parameters and / or facial features parameters.

[0047] In practical applications, the target multimodal large model can also predict the body posture parameters and / or facial features parameters of the target occupant based on cockpit image data and indicator annotation information. Among them, body posture parameters may include user shoulder width data, and facial features parameters may include any one of eye height data, eyebrow height data, etc.

[0048] Specifically, based on the known dimensions of the cockpit reference object, relevant data of the occupants can be predicted from both lateral and longitudinal angles, thereby comprehensively determining the occupants' sitting height. In particular, the length or height of the cockpit reference object in the image can be compared with a certain lateral or longitudinal data of the occupants. If the actual length or height of the cockpit reference object has been determined, the relevant lateral or longitudinal values ​​of the occupants can be predicted.

[0049] After obtaining the occupant's sitting height data, posture parameters, and facial features parameters, since different body types correspond to different sitting height, posture parameters, and facial features parameters, the occupant's body type can be inferred based on the sitting height data, posture parameters, and facial features parameters.

[0050] In this embodiment of the invention, determining body type category through the combined effect of multiple data can improve the accuracy of prediction.

[0051] In one embodiment of the present invention, the target body type category is used to adjust the airbag configuration parameters of the vehicle. That is, when the occupants are of different body types, the airbag configuration parameters can be adjusted according to the body type of the occupants in the current cabin, so that when the vehicle is involved in an accident, the airbags suitable for the occupants in the current cabin can be deployed to ensure the safety of the occupants.

[0052] In this embodiment of the invention, cabin image data of the target occupant inside the vehicle can be acquired, and then the corresponding indicator label information can be determined. The cabin image data and indicator label information can then be input into the target multimodal large model to output the target occupant's target sitting height data, thereby obtaining the target occupant's target body type category. By introducing indicator label information, the target multimodal large model can be assisted in focusing on cabin reference objects, thereby understanding the user's intention through the user's description, achieving accurate prediction of the occupant's sitting height data, and thus obtaining an accurate body type classification.

[0053] In this embodiment of the invention, the generalization ability of the initial multimodal large model is used to solve the classification of occupant body shape in the cabin. The initial multimodal large model is trained on a large dataset and has general knowledge capabilities. Therefore, only a small amount of data is needed to train the target multimodal large model that can predict sitting height in this case.

[0054] Reference Figure 2 The diagram illustrates a flowchart of another method for classifying the body types of vehicle cabin occupants according to an embodiment of the present invention, which may specifically include the following steps: Step S201: Acquire cockpit image data of the target occupant inside the vehicle; Step S202: Determine the indication and annotation information corresponding to the cockpit image data; The indicated labeling information is used to distinguish different cabin reference objects within the vehicle cabin.

[0055] Step S203: Input the cockpit image data and the indicator annotation information into the trained target multimodal large model, and output the target occupant's target sitting height data; The target multimodal large model is a model that calculates the target's sitting height based on cockpit image data and indicator label information.

[0056] Step S204: Obtain the correspondence between sitting height data and body type category; In this embodiment of the invention, sitting height data and body type data can be collected in advance to obtain the correspondence between sitting height data and body type.

[0057] Table 1 shows a correspondence between sitting height data and body type categories in an embodiment of the present invention:

[0058] Specifically, when the occupant's sitting height is greater than or equal to 78 and less than 85, the occupant corresponds to the P5 body type; when the occupant's sitting height is greater than or equal to 85 and less than 92, the occupant corresponds to the P50 body type; and when the occupant's sitting height is greater than or equal to 92, the occupant corresponds to the P95 body type.

[0059] Among them, the human body type (i.e., body shape category) can be classified according to the body shape classification in GB / T1000-2023.

[0060] Step S205: Determine the target body type category corresponding to the target sitting height data based on the correspondence.

[0061] Among them, the target body type category is a category that classifies the body types of occupants seated in the cabin.

[0062] In this embodiment of the invention, the introduction of indicator labeling information can assist the target multimodal large model in focusing on the cockpit reference object, thereby understanding the user's intention through the user's description, achieving accurate prediction of the occupant's sitting height data, and thus obtaining accurate body type classification.

[0063] Reference Figure 3 The diagram illustrates a flowchart of another method for classifying the body types of vehicle cabin occupants according to an embodiment of the present invention, which may specifically include the following steps: Step S301: Acquire cockpit image data of the target occupant inside the vehicle; Step S302: Determine the indication and annotation information corresponding to the cockpit image data; The indicated labeling information is used to distinguish different cabin reference objects within the vehicle cabin.

[0064] Step S303: Process the cockpit image data and the indication annotation information according to the preset format corresponding to the target multimodal large model generated during training.

[0065] The target multimodal large model is a model that calculates the target sitting height based on cockpit image data and indicator label information, obtained by fine-tuning the initial multimodal large model using a preset fine-tuning tool.

[0066] The preset format can be determined based on the target multimodal large model, that is, the preset format is the format that can be applied in the target multimodal large model.

[0067] In one embodiment of the present invention, the step of processing the cockpit image data and the indication annotation information according to the preset format corresponding to the target multimodal large model includes: processing the cockpit image data according to the Alpaca format data or Share GPT format data corresponding to the target multimodal large model; and processing the indication annotation information according to the JSON format data corresponding to the target multimodal large model.

[0068] In this embodiment of the invention, the data format that can be used in the target multimodal large model can be determined in advance. For example, the cockpit image data can be in Alpaca format or Share GPT format, and the indicator standard information can be in JSON format. After obtaining the cockpit image data and indicator label data, the data in the target multimodal large model can be processed quickly to obtain the sitting height of the target occupant.

[0069] Step S304: Input the cockpit image data and the indication label information into the target multimodal large model, and output the target occupant's target sitting height data; Step S305: Determine the target body type category of the target occupant based on the target sitting height data.

[0070] In practical applications, there is a correlation between occupant body types and seating height. Based on this correlation, it is possible to determine the target body type of the corresponding occupant according to the target seating height data. The target body type category is a classification based on the body type of the occupant seated in the cabin.

[0071] In this embodiment of the invention, processing the input data of the target multimodal large model according to a preset format allows the target multimodal large model to better understand the input data and obtain accurate seat height prediction data. Furthermore, this embodiment of the invention introduces indicator labeling information to assist the target multimodal large model in focusing on cabin reference objects, thereby understanding the user's intentions through the user's description, achieving accurate prediction of occupant seat height data, and ultimately obtaining accurate body type classification.

[0072] Meanwhile, before inputting data into the target multimodal large model, the cockpit image data and indicator annotation information can be preprocessed to obtain data applicable to the target multimodal large model, thereby facilitating the multimodal large model to perform occupant seating height prediction.

[0073] Reference Figure 4 The diagram illustrates a flowchart of another method for classifying the body types of vehicle cabin occupants according to an embodiment of the present invention, which may specifically include the following steps: Step S401: Acquire cockpit image data of the target occupant inside the vehicle; Step S402: Determine the indication annotation information corresponding to the cockpit image data; The indicated labeling information is used to distinguish different cabin reference objects within the vehicle cabin.

[0074] Step S403: Input the cockpit image data and the indicator annotation information into the trained target multimodal large model, and output the target occupant's target sitting height data; The target multimodal large model is a model that calculates the target's sitting height based on cockpit image data and indicator label information.

[0075] Step S404: Determine the target body type category of the target occupant based on the target sitting height data.

[0076] Among them, the target body type category is a category that classifies the body types of occupants seated in the cabin.

[0077] Step S405: Input the cockpit image data and the indication label information into the target multimodal large model, and output the seat belt wearing status of the target occupant and / or the number of the target occupants; In this embodiment of the invention, the target multimodal large model may also have the functions of occupant seat belt wearing status and / or the number of target occupants.

[0078] Specifically, the target multimodal large model can be obtained by fine-tuning the initial multimodal large model using a preset fine-tuning tool. The initial multimodal large model can be an scalable model; with a small amount of data, a target multimodal large model capable of performing different functions can be trained. In this embodiment of the invention, the initial multimodal large model data can be trained to a model for detecting seatbelt wearing status, or it can be trained to a model for detecting the number of occupants in a vehicle, or it can be trained to a model that simultaneously has the functions of detecting seatbelt wearing status and detecting the number of occupants in a vehicle.

[0079] In this embodiment of the invention, the trained target multimodal large model realizes the expansion of functions such as seat belt wearing detection and number of people in the cabin. Compared with traditional models, a target multimodal large model can adapt to multiple tasks at the same time.

[0080] It should be noted that, for the sake of simplicity, the method embodiments are described as a series of actions. However, those skilled in the art should understand that the embodiments of the present invention are not limited to the described order of actions, because according to the embodiments of the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present invention.

[0081] Reference Figure 5 The diagram shows a structural schematic of a vehicle cabin occupant body shape classification device according to an embodiment of the present invention, which may specifically include the following modules: The cockpit image data acquisition module 501 is used to acquire cockpit image data of the target occupant inside the vehicle; The indicator standard information determination module 502 is used to determine the indicator annotation information corresponding to the cockpit image data, wherein the indicator annotation information is annotation information used to distinguish different cockpit reference objects in the vehicle cockpit; The target sitting height data determination module 503 is used to input the cockpit image data and the indicator label information into the trained target multimodal large model and output the target sitting height data of the target occupant. The target multimodal large model is a model that calculates the target sitting height based on the cockpit image data and indicator label information. The target body type determination module 504 is used to determine the target body type category of the target occupant based on the target sitting height data. The target body type category is a category that classifies the body type of the occupant sitting in the cabin.

[0082] In one embodiment of the present invention, the indication standard information determination module 502 may include: A cockpit reference object determination submodule is used to determine multiple cockpit reference objects based on the cockpit image data; The actual size information determination submodule is used to obtain the actual size information of each cockpit reference object; The indicator labeling information determination submodule is used to generate corresponding indicator labeling information for each cockpit reference object according to the actual size information.

[0083] In one embodiment of the present invention, the device further includes: The body posture or facial feature parameter determination module is used to input the cockpit image data and the indicator annotation information into the target multimodal large model generated by training, and output the body posture parameters and / or facial feature parameters of the target occupant; Target body type determination module 504 may include: The second target body type determination submodule is used to determine the target body type of the target occupant based on the target sitting height data, as well as the posture parameters and / or facial feature parameters.

[0084] In one embodiment of the present invention, the device further includes: The configuration parameter determination module is used to adjust the configuration parameters of the airbags in the vehicle according to the body type.

[0085] In one embodiment of the present invention, the target body type determination module 504 may include: The mapping relationship acquisition submodule is used to obtain the mapping relationship between sitting height data and body type category; The target body shape acquisition submodule is used to determine the target body shape category corresponding to the target sitting height data based on the correspondence.

[0086] In one embodiment of the present invention, the device may include: The format processing submodule is used to process the cockpit image data and the indication and annotation information according to the preset format corresponding to the multimodal large model.

[0087] In one embodiment of the present invention, the format processing submodule may include: The image data processing unit is used to process the cockpit image data according to the Alpaca format data or Share GPT format data corresponding to the target multimodal large model. The indicator label information processing submodule is used to process the indicator label information according to the JSON format data corresponding to the multimodal large model.

[0088] In one embodiment of the present invention, the device may further include: The size type determination module is used to determine the model size type of the target multimodal large model; The deployment method determination module is used to determine the target deployment method corresponding to the target multimodal large model according to the model size type. The deployment module is used to deploy the target multimodal large model based on the target deployment method.

[0089] In one embodiment of the present invention, the deployment method determination module may include: The cloud deployment submodule is used to determine the target deployment method of the multimodal large model as cloud deployment when the model size type is greater than or equal to 7B. The terminal deployment submodule is used to determine the target deployment method of the multimodal large model as terminal deployment when the model size type is 3B.

[0090] In one embodiment of the present invention, the device may further include: The functional expansion module is used to input the cockpit image data and the indication label information into the target multimodal large model, and output the seat belt wearing status of the target occupants and / or the number of target occupants; In this embodiment of the invention, cabin image data of the target occupant inside the vehicle can be acquired, and then the corresponding indicator label information can be determined. The cabin image data and indicator label information can then be input into the target multimodal large model to output the target occupant's target sitting height data, thereby obtaining the target occupant's target body type category. By introducing indicator label information, the target multimodal large model can be assisted in focusing on cabin reference objects, thereby understanding the user's intention through the user's description, achieving accurate prediction of the occupant's sitting height data, and thus obtaining an accurate body type classification.

[0091] An embodiment of the present invention also provides an electronic device, which may include a processor, a memory, and a computer program stored in the memory and capable of running on the processor. When the computer program is executed by the processor, it implements the above-mentioned vehicle cabin occupant body type classification method.

[0092] An embodiment of the present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the above-mentioned vehicle cabin occupant body type classification method.

[0093] As the device embodiment is basically similar to the method embodiment, the description is relatively simple, and relevant parts can be found in the description of the method embodiment.

[0094] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0095] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, apparatus, or computer program products. Therefore, embodiments of the present invention can take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of the present invention can take the form of computer program products implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0096] Embodiments of the present invention are described with reference to flowchart illustrations and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, create means for implementing the functions specified in one or more blocks of the flowchart illustrations and / or one or more blocks of the block diagrams.

[0097] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means that implement the functions specified in one or more flowcharts and / or one or more block diagrams.

[0098] These computer program instructions may also be loaded onto a computer or other programmable data processing terminal equipment to cause a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable terminal equipment, provide steps for implementing the functions specified in one or more flowcharts and / or one or more block diagrams.

[0099] Although preferred embodiments of the present invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present invention.

[0100] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.

[0101] The above provides a detailed description of the vehicle cabin occupant body type classification method and device, electronic equipment, and storage medium. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A vehicle cabin occupant size classification method characterized by, The method comprises: acquiring cabin image data of a target occupant in a vehicle; determining indication label information corresponding to the cabin image data, the indication label information being label information used to distinguish different cabin references in the vehicle cabin; inputting the cabin image data and the indication label information into a target multi-modal large model trained and generated, and outputting target seat height data of the target occupant, the target multi-modal large model being a model for calculating target seat height based on cabin image data and indication label information; determining a target body type category of the target occupant according to the target seat height data, the target body type category being a category classified according to the body type of an occupant on a seat in the cabin.

2. The method of claim 1, wherein, The determination of the indication label information corresponding to the cabin image data comprises: determining a plurality of cabin references based on the cabin image data; acquiring actual size information of each cabin reference; generating corresponding indication label information for each cabin reference according to the actual size information.

3. The method of claim 1, wherein, Further comprising: inputting the cabin image data and the indication label information into a target multi-modal large model trained and generated, and outputting body state parameters and / or facial feature parameters of the target occupant; The determination of the target body type category of the target occupant according to the target seat height data comprises: determining the target body type category of the target occupant based on the target seat height data, the body state parameters and / or the facial feature parameters.

4. The method of claim 1, wherein, Further comprising: adjusting configuration parameters of an airbag in the vehicle according to the body type category.

5. The method of claim 1, wherein, The determination of the target body type category of the target occupant according to the target seat height data comprises: acquiring a corresponding relationship between seat height data and body type categories; determining the target body type category corresponding to the target seat height data based on the corresponding relationship.

6. The method of claim 1, wherein, Before inputting the cabin image data and the indication label information into the target multi-modal large model trained and generated, comprising: processing the cabin image data and the indication label information according to a preset format corresponding to the target multi-modal large model.

7. The method of claim 6, wherein, The processing of the cabin image data and the indication label information according to the preset format corresponding to the target multi-modal large model comprises: performing data processing on the cabin image data according to Alpaca format data or ShareGPT format data corresponding to the target multi-modal large model; performing data processing on the indication label information according to json format data corresponding to the target multi-modal large model.

8. The method of claim 1, wherein, Further comprising: determining a model size type of the target multi-modal large model; determining a target deployment mode corresponding to the target multi-modal large model according to the model size type; deploying the target multi-modal large model based on the target deployment mode.

9. The method of claim 8, wherein, The determination of the target deployment mode corresponding to the target multi-modal large model according to the model size type comprises: when the model size type is greater than or equal to 7B, determining that the target deployment mode corresponding to the target multi-modal large model is cloud deployment; when the model size type is 3B, determining that the target deployment mode corresponding to the target multi-modal large model is vehicle machine end deployment.

10. The method of claim 1, wherein, Further comprising: The cabin image data and the indication annotation information are input into the target multi-modal large model, and a safety belt wearing state of the target occupant and / or a number of the target occupant are output.

11. A vehicle cabin occupant size classification device characterized by comprising: The device comprises: a cabin image data acquisition module configured to acquire cabin image data of a target occupant in a vehicle; an indication standard information determination module configured to determine indication annotation information corresponding to the cabin image data, the indication annotation information being annotation information used to distinguish different cabin reference objects in the vehicle cabin; a target seat height data determination module configured to input the cabin image data and the indication annotation information into a target multi-modal large model trained and generated, and output target seat height data of the target occupant, the target multi-modal large model being a model for calculating target seat height based on cabin image data and indication annotation information; a target body type type determination module configured to determine a target body type category of the target occupant according to the target seat height data, the target body type category being a category classified according to the body type of the occupant on the seat in the cabin.

12. An electronic device, comprising: A computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the vehicle cabin occupant body type classification method according to any one of claims 1 to 10.

13. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the vehicle cabin occupant body type classification method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Method for calculating passenger sitting height and classifying seat cabin members based on in-vehicle OMS camera

    CN114581893A

  • Seated Passenger Height Estimation

    US20240013419A1

  • KR20210012491A