Vehicle-mounted personal image optimization method based on multi-modal perception

By collecting user image features through multimodal sensing devices and making composite judgments, personalized intervention strategies are triggered, which solves the problem of insufficient image management in in-vehicle systems and improves user experience and service quality.

CN120747137BActive Publication Date: 2026-04-14RIVOTEK TECH (JIANGSU) CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing in-vehicle human-machine interaction systems lack multimodal perception capabilities, cannot accurately identify user appearance features, and lack personalized response and dynamic feedback mechanisms, resulting in inadequate image management.

Method used

By synchronously collecting user image features through multimodal sensing devices, preprocessing and anomaly detection are performed. Combined with destination scene tags, a composite judgment is made to trigger personalized intervention strategies, including seat heating and wet wipe release.

Benefits of technology

It enables precise quantitative management and personalized response of user profiles, thereby enhancing the proactive service capabilities and user satisfaction of the in-vehicle intelligent cockpit.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120747137B_ABST
    Figure CN120747137B_ABST
Patent Text Reader

Abstract

The application discloses a vehicle-mounted personal image optimization method based on multi-modal perception and relates to the technical field of vehicle-mounted intelligent cabin and personalized interaction, comprising the following steps: synchronously collecting user image features by a multi-modal perception device and preprocessing to obtain collar profile data, clothing wrinkle level and sebum secretion index; performing abnormality judgment based on the collar profile data, the clothing wrinkle level and the sebum secretion index, and performing composite judgment in combination with a destination scene label to output a composite judgment flag; and triggering an intervention execution strategy according to the composite judgment flag to perform vehicle-mounted personal image optimization. The application overcomes the defects of low single-dimensional recognition accuracy, insufficient scene adaptability and lagging intervention measures in the prior art, enhances passenger image management, improves the active service capability of the vehicle-mounted intelligent cabin and the user satisfaction, and optimizes the travel experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of in-vehicle intelligent cockpit and personalized interaction technology, and in particular to an in-vehicle personal image optimization method based on multimodal perception. Background Technology

[0002] With the development of in-vehicle intelligent technology, vehicles are gradually evolving from traditional transportation tools into intelligent mobile spaces that integrate perception, computing, and interaction. In particular, continuous breakthroughs in areas such as autonomous driving, intelligent cockpits, and personalized services have enabled in-vehicle systems to increasingly focus on deeper needs such as user health, comfort, and social image, in addition to meeting driving safety and navigation assistance requirements. The maturity of in-vehicle multimodal perception technologies, such as the integrated application of sensing components like wide-angle cameras, piezoelectric fabric sensors, and near-infrared spectrometers, provides the hardware foundation for real-time status recognition and behavioral responses to drivers and passengers. Simultaneously, as the capabilities of artificial intelligence at the edge computing level increase, intelligent data fusion and personalized recommendation services are gradually becoming more real-time and contextualized, driving the transformation of in-vehicle human-machine interaction systems from "information-providing" to "proactive service-oriented." Against this backdrop, how to utilize in-vehicle systems to achieve intelligent perception and dynamic management of users' appearance and other external aspects has become a new direction for improving the quality of in-vehicle intelligent services and user experience.

[0003] However, existing in-vehicle human-machine interaction or smart cockpit systems still have several shortcomings in user image management. First, most systems lack the ability to accurately identify and structurally quantify user appearance features, failing to effectively perceive and analyze dimensions such as clothing wrinkles, sebum levels, or hairstyle integrity, resulting in the system's inability to effectively "understand" the user's image status. Second, existing technologies mostly remain at the level of single-point intervention (such as ventilation and deodorization, mirror lighting), lacking a systematic judgment mechanism based on multi-source perception data fusion, making it difficult to provide personalized responses and suggestions based on specific travel scenarios (such as business, leisure, and emergencies). Furthermore, the lack of a continuous learning and dynamic feedback mechanism for changes in user status prevents the system from adjusting parameters and optimizing interventions over long-term use, thus limiting the accuracy and adaptability of the service. Summary of the Invention

[0004] In view of the problems existing in existing in-vehicle personal image optimization methods based on multimodal perception, this invention is proposed. Therefore, the problem to be solved by this invention is how to provide an in-vehicle personal image optimization method based on multimodal perception.

[0005] To solve the above-mentioned technical problems, the present invention provides the following technical solution:

[0006] In a first aspect, the present invention provides a vehicle-mounted personal image optimization method based on multimodal perception, which includes synchronously collecting user image features through a multimodal perception device and performing preprocessing to obtain collar outline data, clothing wrinkle level and sebum secretion index.

[0007] Anomaly detection is performed based on collar outline data, clothing wrinkle level, and sebum secretion index, and combined with destination scene tags for composite detection, outputting composite detection flags;

[0008] Based on the composite judgment flag, an intervention execution strategy is triggered to optimize the in-vehicle personal image.

[0009] As a preferred embodiment of the in-vehicle personal image optimization method based on multimodal perception described in this invention, the multimodal perception device includes a roof camera, a seat piezoelectric fabric sensor, and an armrest contact spectrometer; the user image features include an upper body image of the user acquired by the roof camera; pressure data acquired by the seat piezoelectric fabric sensor; and skin reflectance spectrum acquired by the armrest contact spectrometer.

[0010] As a preferred embodiment of the in-vehicle personal image optimization method based on multimodal perception described in this invention, the acquisition of collar contour data, clothing wrinkle level, and sebum secretion index includes:

[0011] The user's upper body image is acquired and converted to grayscale. Canny edge detection is used to extract the image edges. Image template matching and morphological operations are used for region segmentation to obtain the collar outline mask and hairstyle boundary mask.

[0012] The collected pressure data is normalized. Information entropy is calculated based on the normalized pressure data. The clothing wrinkle level is then calculated based on the information entropy. The information entropy is expressed as:

[0013] ;

[0014] in: For information entropy, The stress data is after normalization. It is a constant. and For index variables;

[0015] Obtain the user's skin reflectance spectrum and extract reflectance data in the near-infrared band to calculate the sebum secretion index.

[0016] As a preferred embodiment of the in-vehicle personal image optimization method based on multimodal perception described in this invention, the expression for the clothing wrinkle level is:

[0017] ;

[0018] in: The level of wrinkles in clothing. This represents the upper limit of historical information entropy. This is the lower bound of historical information entropy;

[0019] The expression for the sebum secretion index is:

[0020] ;

[0021] in: This refers to the sebum secretion index. The current user's skin reflectivity, Reflectance of the standard sample for low sebum secretion; This refers to the number of bands.

[0022] As a preferred embodiment of the in-vehicle personal image optimization method based on multimodal perception described in this invention, the anomaly judgment based on collar contour data, clothing wrinkle level, and sebum secretion index includes:

[0023] The edge continuity and left-right symmetry of the collar outline coordinates are obtained. Edge continuity is the ratio of the actual collar outline length to the ideal collar outline length; left-right symmetry is the average difference in the horizontal distance between corresponding points on the left and right sides of the collar outline.

[0024] The edge continuity and left-right symmetry are fed into a pre-trained support vector machine classifier to obtain the collar abnormality probability. When the abnormality probability exceeds the predetermined probability, the collar is identified as abnormal.

[0025] The system assesses wrinkles and sebum thresholds. If the wrinkle level of clothing exceeds a predetermined level, the wrinkles are considered abnormal; if the sebum secretion index exceeds a predetermined index, the sebum level is considered excessive.

[0026] As a preferred embodiment of the in-vehicle personal image optimization method based on multimodal perception described in this invention, the step of combining destination scene tags for composite judgment includes:

[0027] Obtain the current destination scenario tag from the vehicle navigation system or related calendar application, categorizing it into three types: business, leisure, or emergency.

[0028] In the local knowledge base, each scenario has three preset thresholds: the maximum acceptable wrinkle level; the maximum acceptable sebum index; and whether the collar is guaranteed to be free of obvious defects.

[0029] The detection results are compared with the scene thresholds. If any threshold is exceeded, a dress non-compliance flag is triggered.

[0030] The three abnormal flags are assigned initial weights and a weighted sum is calculated. When the weighted sum exceeds a predetermined limit, it is determined to be a composite judgment abnormality and the composite judgment flag is triggered.

[0031] As a preferred embodiment of the in-vehicle personal image optimization method based on multimodal perception described in this invention, the step of triggering the intervention execution strategy based on composite judgment flags includes:

[0032] The vehicle receives composite judgment flags and scene tags through the vehicle's intranet and issues control commands.

[0033] When the wrinkle level of the clothing is greater than the preset level, and the current scene label is "business", the composite judgment flag has been triggered. The target temperature command is sent to the seat heating module to start the heating module. After heating continues for a preset time, a stop heating command is issued and the system enters the waiting cooling protection mode.

[0034] When the sebum secretion index exceeds the predetermined index and the composite judgment flag is triggered, power is supplied to the storage box ejection mechanism. The spring pushes the box lid and wet wipes to slide out. After the limit switch confirms the ejection position, the power is cut off, and a prompt pops up in the corner of the central control screen and the instrument panel.

[0035] In a second aspect, the present invention provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of a vehicle-mounted personal image optimization method based on multimodal perception.

[0036] Thirdly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements the steps of a vehicle-mounted personal image optimization method based on multimodal perception.

[0037] The beneficial effects of this invention are as follows: it not only overcomes the shortcomings of low single-dimensional recognition accuracy, insufficient scene adaptability and lagging intervention measures in the existing technology, but also realizes the objective quantification, efficient integration and real-time correction of appearance parameters; while enhancing passenger image management, it also improves the proactive service capabilities and user satisfaction of the in-vehicle intelligent cockpit, enhances image consistency, improves social etiquette compatibility and optimizes the travel experience. Attached Figure Description

[0038] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0039] Figure 1This is a flowchart of a method for optimizing in-vehicle personal image based on multimodal perception. Detailed Implementation

[0040] To make the above-mentioned objects, features, and advantages of the present invention more readily understood, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.

[0041] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0042] Secondly, the term "one embodiment" or "example" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the invention. An embodiment appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single embodiment or an embodiment that selectively excludes other embodiments.

[0043] Reference Figure 1 This is the first embodiment of the present invention, which provides a method for optimizing in-vehicle personal image based on multimodal perception, including:

[0044] S1: Simultaneously collect user image features through multimodal sensing devices and preprocess them to obtain collar outline data, clothing wrinkle level and sebum secretion index;

[0045] Specifically, after the vehicle is started, user image data is collected simultaneously through the ceiling camera integrated in the cabin, the seat piezoelectric fabric sensor, and the armrest contact spectrometer.

[0046] The ceiling camera acquires images of the user's upper body and uses an edge detection algorithm to extract the collar outline and hairstyle boundary;

[0047] The piezoelectric fabric sensor of the seat generates a pressure distribution matrix and outputs a quantitative value of the clothing wrinkle level (range 0-5) based on the entropy calculation module.

[0048] The spectrometer captures the skin reflectance spectrum, and the sebum secretion excess index is obtained by near-infrared characteristic band analysis.

[0049] After the vehicle is started, the system automatically activates the multimodal sensing devices integrated into the cabin, including: a roof camera (FOV 120°, resolution 1920×1080); a seat piezoelectric fabric sensor (based on Pyralux® AP material, array size 64×64); and an armrest contact spectrometer (operating wavelength 400–1000nm, sampling interval 2nm).

[0050] The three devices simultaneously collect images of the user's upper body, seat pressure distribution maps, and skin spectral data, and transmit them to the edge computing module for preprocessing.

[0051] The user's upper body image is acquired and converted to grayscale. Canny edge detection is used to extract the image edges. Based on positional and structural features (the hairline is usually located in the center of the top of the head, and the collar edge is U-shaped at the shoulder and neck position), image template matching and morphological operations are used to segment the region to obtain the collar outline mask and the hairstyle boundary mask.

[0052] The collected pressure data is normalized, and the information entropy is calculated based on the normalized pressure data, expressed as:

[0053] ;

[0054] in: For information entropy, The stress data is after normalization. It is a constant. and For index variables;

[0055] The calculation of garment wrinkle levels is expressed as follows:

[0056] ;

[0057] in: This refers to the level of wrinkles in clothing, ranging from 0 to 5. This represents the upper limit of historical information entropy. This is the lower bound of historical information entropy;

[0058] A contact spectrometer acquires the reflectance spectrum of the user's skin, extracts reflectance data in the near-infrared band, and calculates the sebum secretion index, expressed as:

[0059] ;

[0060] in: This refers to the sebum secretion index. The current user's skin reflectivity, Reflectance of the standard sample for low sebum secretion; This refers to the number of bands.

[0061] S2: Anomaly detection is performed based on collar outline data, clothing wrinkle level, and sebum secretion index, combined with destination scene tags for composite detection, and a composite detection flag is output.

[0062] Specifically, collar contour anomaly detection is performed to obtain the edge continuity and left-right symmetry of the collar contour coordinates. Edge continuity reflects the ratio of the detected actual collar contour length to the ideal collar contour length; left-right symmetry measures the average difference in horizontal distance between corresponding points on the left and right sides of the collar contour.

[0063] The edge continuity and left-right symmetry are fed into a pre-trained support vector machine (SVM) classifier to obtain the collar anomalous probability. When the probability exceeds the predetermined probability (95%), the collar is identified as anomalous.

[0064] The wrinkle and sebum threshold are judged. If the wrinkle level of the clothing is greater than the predetermined level (3), the wrinkle is considered abnormal; if the sebum secretion index exceeds the predetermined index (20%), the sebum is considered excessive.

[0065] Obtain the current destination scenario tag from the vehicle navigation system or related calendar application, which can be categorized into three types: business, leisure, or emergency.

[0066] In the local knowledge base, each scenario has three preset thresholds: the maximum acceptable wrinkle level; the maximum acceptable sebum index; and whether it is necessary to ensure that the collar has no obvious defects.

[0067] For example, business settings typically require a wrinkle level of no more than 2, a sebum index of no more than 20%, and an intact collar. The test results are compared with the corresponding thresholds for each scenario, listing all items that exceed the limits. If any item exceeds the limit, a dress code violation is triggered.

[0068] Assign initial weights to the three abnormal flags (collar, wrinkle, sebum) (e.g., collar 0.4, wrinkle 0.3, sebum 0.3), and calculate the weighted sum. When the weighted sum reaches or exceeds a predetermined limit (0.8), it is judged as a composite judgment anomaly, and the composite judgment flag is set. Output the judgment flag.

[0069] S3: Trigger the intervention execution strategy based on the composite judgment flag to optimize the in-vehicle personal image.

[0070] Specifically, the system receives composite judgment flags and scene tags through the vehicle's intranet and issues control commands. The first-level strategy (seat heating) is based on the following conditions: clothing wrinkle level ≥ 3, current scene tag is business, and composite judgment flag has been triggered.

[0071] Action: The ECU sends a command to the seat heating module: target temperature 45℃, allowable error ±1℃. The heating module starts, with the PWM initial duty cycle at 80%, gradually adjusting to approach the target temperature. The temperature sensor samples at 1Hz and feeds back to the ECU, forming a closed-loop control. After heating continues for 180 seconds, the ECU issues a stop heating command and enters a cooling protection mode (protection is released when the temperature cools to ≤35℃).

[0072] Secondary strategy (release of oil-control wipes), condition: sebum secretion index > 25%, compound judgment indicator has been triggered;

[0073] Action: The ECU powers on the storage box ejection mechanism, the spring pushes the box cover and wet wipes out, the limit switch confirms the ejection position and then cuts off the power, a prompt pops up in the corner of the central control screen and the instrument panel: We have prepared minty oil-control wet wipes for you, please take them and wipe your face gently.

[0074] If multiple actions are triggered and executed in parallel, the release of wet wipes has the highest priority, followed by heating, in order to avoid abruptly interfering with the user's experience.

[0075] Each time it is triggered, the dashboard and central control screen will display brief text and icons to prompt the user with the physical intervention project being performed and the operation instructions, ensuring that the user is aware of and cooperates with the operation.

[0076] This embodiment also provides a computer device applicable to the in-vehicle personal image optimization method based on multimodal perception, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement all or part of the steps of the method described in the above embodiments of the present invention.

[0077] This embodiment also provides a storage medium storing a computer program thereon. When the computer program is executed by a processor, it performs the method in any optional implementation of the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0078] The storage medium proposed in this embodiment and the data storage method proposed in the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.

[0079] In summary, this method constructs a complete closed loop from multimodal perception and scenario-based diagnosis to personalized intervention. It not only overcomes the shortcomings of existing technologies, such as low accuracy of single-dimensional recognition, insufficient scenario adaptability, and lagging intervention measures, but also achieves objective quantification, efficient fusion, and real-time correction of appearance parameters. While enhancing passenger image management, it also improves the proactive service capabilities and user satisfaction of the in-vehicle intelligent cockpit, enhances image consistency, improves social etiquette compatibility, and optimizes the travel experience.

[0080] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A method for optimizing in-vehicle personal image based on multimodal perception, characterized in that: include, By synchronously collecting user image features through multimodal sensing devices and preprocessing them, data such as collar outline, clothing wrinkle level, and sebum secretion index can be obtained. Anomaly detection is performed based on collar outline data, clothing wrinkle level, and sebum secretion index, and combined with destination scene tags for composite detection, outputting composite detection flags; The anomaly detection based on collar contour data, clothing wrinkle level, and sebum secretion index includes: The edge continuity and left-right symmetry of the collar outline coordinates are obtained. Edge continuity is the ratio of the actual collar outline length to the ideal collar outline length; left-right symmetry is the average difference of the horizontal distances between corresponding points on the left and right sides of the collar outline. The edge continuity and left-right symmetry are fed into a pre-trained support vector machine classifier to obtain the collar abnormality probability. When the abnormality probability exceeds the predetermined probability, the collar is identified as abnormal. The system assesses wrinkles and sebum thresholds. If the wrinkle level of clothing exceeds a predetermined level, the wrinkles are considered abnormal; if the sebum secretion index exceeds a predetermined index, the sebum level is considered excessive. The combined judgment based on destination scene tags includes: Obtain the current destination scenario tag from the vehicle navigation system or related calendar application, categorizing it into three types: business, leisure, or emergency. In the local knowledge base, each scenario has three preset thresholds: the maximum acceptable wrinkle level; the maximum acceptable sebum index; and whether the collar is guaranteed to be free of obvious defects. The detection results are compared with the scene thresholds. If any threshold is exceeded, a dress non-compliance flag is triggered. The three abnormal flags are assigned initial weights and a weighted sum is calculated. When the weighted sum exceeds a predetermined limit, it is determined to be a composite judgment abnormality and the composite judgment flag is triggered. Based on the composite judgment flag, an intervention execution strategy is triggered to optimize the in-vehicle personal image.

2. The in-vehicle personal image optimization method based on multimodal perception as described in claim 1, characterized in that: The multimodal sensing device includes a ceiling camera, a seat piezoelectric fabric sensor, and an armrest contact spectrometer; the user image features include an image of the user's upper body acquired by the ceiling camera; Pressure data acquired by a piezoelectric fabric sensor in the seat; skin reflectance spectrum acquired by a contact spectrometer in the armrest.

3. The in-vehicle personal image optimization method based on multimodal perception as described in claim 2, characterized in that: The acquisition of collar contour data, garment wrinkle level, and sebum secretion index includes: The user's upper body image is acquired and converted to grayscale. Canny edge detection is used to extract the image edges. Image template matching and morphological operations are used to segment the region to obtain the collar outline mask and hairstyle boundary mask. The collected pressure data is normalized. Information entropy is calculated based on the normalized pressure data. The clothing wrinkle level is then calculated based on the information entropy. The information entropy is expressed as: in: For information entropy, The stress data is after normalization. It is a constant. and For index variables; Obtain the user's skin reflectance spectrum and extract reflectance data in the near-infrared band to calculate the sebum secretion index.

4. The in-vehicle personal image optimization method based on multimodal perception as described in claim 3, characterized in that: The expression for the garment wrinkle level is: in: The level of wrinkles in clothing. This represents the upper limit of historical information entropy. This is the lower bound of historical information entropy; The expression for the sebum secretion index is: in: This refers to the sebum secretion index. The current user's skin reflectivity, The reflectance of the standard sample with low sebum secretion; This refers to the number of bands.

5. The in-vehicle personal image optimization method based on multimodal perception as described in claim 1, characterized in that: The intervention execution strategy triggered based on the composite judgment flag includes: The vehicle receives composite judgment flags and scene tags through the vehicle's intranet and issues control commands. When the wrinkle level of the clothing is greater than the preset level, and the current scene label is "business", the composite judgment flag has been triggered. The target temperature command is sent to the seat heating module to start the heating module. After heating continues for a preset time, a stop heating command is issued and the system enters the waiting cooling protection mode. When the sebum secretion index exceeds the predetermined index and the composite judgment flag is triggered, power is supplied to the storage box ejection mechanism. The spring pushes the box lid and wet wipes to slide out. After the limit switch confirms the ejection position, the power is cut off, and a prompt pops up in the corner of the central control screen and the instrument panel.

6. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the in-vehicle personal image optimization method based on multimodal perception as described in any one of claims 1 to 5.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the in-vehicle personal image optimization method based on multimodal perception as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Personal image management method and device of vehicle cabin, vehicle and storage medium

    CN117251609A

  • Digital printing fabric defect detection method

    CN119784740A

  • Multifunctional probe for measuring skin property

    JP2009153728A

  • Methods and systems for improving human facial skin conditions by leveraging vehicle cameras and skin data ai analytics

    US20210338146A1