Vehicle lamp functionality detection method and system based on multi-modal fusion model

CN122072703BActive Publication Date: 2026-08-21FITOW (TIANJIN) DETECTION TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610525637.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-04-21
Publication Date
2026-08-21
Estimated Expiration
2046-04-21

AI Technical Summary

Technical Problem

[0011]因此,本发明的目的在于提供一种基于多模态融合模型的车灯功能性检测方法及系统,以解决现有技术中因依赖单张图像检测而无法完整捕捉动态灯光状态、缺乏操作信息与车辆信号融合机制以及在复杂灯光功能场景下检测精度低和鲁棒性差的技术问题

Benefits of technology

本发明通过在操作时间段内连续采集多张图像,并对多帧图像的检测结果进行时间段聚合分析,能够完整捕捉转向灯闪烁、AUTO模式切换、组合灯联动等动态灯光功能的全过程状态,避免了单张图像因拍摄时机或角度限制导致的漏检或误判。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122072703B_ABST
    Figure CN122072703B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on multi-modal fusion model's car light functionality detection method and system, the method includes: through continuous acquisition multiple vehicle images in operating time period;Each image is input into the visual detection submodel of pre-training, and the state matrix of each lamp in each image is output;State matrix is aggregated and analyzed in time period, and the comprehensive state matrix of each car light in operating time period is obtained;Comprehensive state matrix, control modal data, signal modal data are input into multi-modal fusion model, and the final state of each car light function and abnormal identification are output.The application solves the problem that single image cannot capture dynamic light state completely by multi-period image acquisition, and the image detection result, operation instruction and vehicle signal are integrated by multi-modal information fusion, which significantly improves the accuracy and robustness of complex light function detection, and can be widely applied to vehicle research and development test, quality detection, regulation compliance detection and intelligent traffic monitoring scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of vehicle intelligent detection technology, and in particular to a method and system for detecting the functionality of vehicle lights based on a multimodal fusion model. Background Technology

[0002] With the development of automotive intelligence, vehicle lighting systems are becoming increasingly complex and diverse in function. For example, in AUTO mode, the headlights and daytime running lights automatically switch illumination states based on ambient light; turn signals involve multiple lights at the front, rear, and sides of the vehicle and need to flash at specific frequencies; brake light status is closely related to driving operations; and combination lights may illuminate multiple lights simultaneously under different operations, forming complex linkage logic. These complex functional states place higher demands on detection methods.

[0003] In existing technologies, vehicle lighting function detection mainly relies on the following methods: Firstly, there's the single-image-based light status recognition method. This method captures a single image of the vehicle and uses image processing or deep learning models to identify the on / off state of each light in the image. However, this method is essentially static detection. For dynamic functions such as turn signal flashing, AUTO mode switching, and combined light linkage, a single image can only capture the light status at a specific moment and cannot fully reflect the changes in light throughout the entire operation. For example, a turn signal might be exactly during an off interval in a single image, leading to missed detection; the process of light changing with ambient light in AUTO mode cannot be verified using a single image. This lack of information makes it difficult for the detection results to accurately reflect the actual performance of the lighting functions.

[0004] Secondly, there is the detection method based on multi-angle image acquisition. This method deploys multiple cameras to capture images of the vehicle from different angles, aiming to comprehensively cover the status of the turn signals at the front, rear, and sides. However, this method still relies on the independent analysis of multiple images, lacking the ability to model the temporal dimension of the image sequence and failing to capture the dynamic characteristics of the turn signal status changing over time. Furthermore, multi-camera systems are costly to deploy and complex to calibrate, and still cannot solve the problem of simultaneously capturing the front and rear status of the turn signals when the shooting angle is limited.

[0005] Thirdly, there is the detection method based on vehicle signal diagnostics. This method reads the vehicle's CAN bus data to obtain lighting control commands and feedback signals, and then determines whether there is a fault in the lighting system. However, this method relies on the vehicle's own diagnostic system and cannot verify whether the lights actually illuminate as instructed. For example, it cannot distinguish between a damaged bulb and a wiring fault, nor can it detect whether the light brightness meets the standard or whether the flashing frequency complies with regulatory requirements, or other physical functional states.

[0006] The core flaw of existing methods lies in: First, there is a lack of dynamic modeling capabilities over time. Existing technologies are essentially still static detection or independent multi-image analysis, unable to correlate the detection results of multiple frames to form a complete description of the changes in lighting status over time. For dynamic functions such as turn signal flashing frequency, AUTO mode switching response time, and combination light linkage sequence, existing methods lack fundamental information, resulting in limited detection accuracy.

[0007] Second, there is a lack of a multimodal information fusion mechanism. Existing methods typically process visual detection, operational commands, and vehicle signals in isolation: the visual detection system doesn't know what operation should be performed, the operational command system cannot verify whether the lights are actually responding, and the vehicle signal system cannot perceive the physical state of the lights. This information silo phenomenon makes it difficult to comprehensively judge whether the lighting function is working normally according to the expected logic based on the detection results. For example, it cannot be combined with the brake pedal signal to verify whether the brake lights are illuminated, nor can it be combined with the turn signal lever signal to verify whether the turn signal flashing frequency meets the standard.

[0008] Third, there is a lack of operational guidance and a closed-loop linkage in the detection process. Existing methods are mostly passive detection methods, meaning they analyze images offline after capture, lacking a mechanism to actively guide operators to perform specified operations and simultaneously collect data during the detection process. This makes it difficult to standardize the detection process, leading to significant operational arbitrariness and poor repeatability and comparability of the detection results.

[0009] Fourth, the ability to diagnose anomalies is insufficient. Existing methods typically only output whether the light is on or off, failing to provide fine-grained anomaly type identification, such as lights that should be on but are not, lights that remain constantly on, insufficient brightness, response delays, or abnormal flickering frequencies. This makes the test results difficult to use for fault location and quality traceability, limiting their practical value in R&D testing and quality control.

[0010] In summary, existing technologies are insufficient to meet the requirements of high-precision, multi-functional, multi-angle, and dynamic lighting function detection. There is an urgent need for an intelligent detection method that can combine multi-time period image acquisition, operation guidance, temporal modeling, and multimodal information fusion to achieve comprehensive, accurate, and automated judgment of complex lighting functions. Summary of the Invention

[0011] Therefore, the purpose of this invention is to provide a vehicle lighting functionality detection method and system based on a multimodal fusion model, in order to solve the technical problems in the prior art that rely on single image detection and cannot fully capture dynamic lighting states, lack operation information and vehicle signal fusion mechanism, and have low detection accuracy and poor robustness in complex lighting functional scenarios.

[0012] To achieve the above objectives, the present invention provides a vehicle headlight functionality detection method based on a multimodal fusion model, comprising: Acquire multimodal data; the multimodal data includes control modal data, visual modal data, and signal modal data; the control modal data includes operation prompt information, a specified operation to be performed according to the operation prompt information, and the corresponding operation time period; the signal modal data includes vehicle signal data fed back by the vehicle when the specified operation is performed; the visual modal data includes multiple images of vehicle headlight changes continuously acquired during the operation time period; Each image of the vehicle headlight change is input into a pre-trained visual detection sub-model, which outputs a single-image headlight state matrix. The single-image headlight state matrix is ​​then aggregated and analyzed according to the operation time period to obtain the comprehensive state matrix of each headlight throughout the entire operation time period. The integrated state matrix, control modal data, and signal modal data are input into the multimodal fusion model. The multimodal fusion model is used to extract structured visual features, control command features, and vehicle sensing features. After fusing the structured visual features, control command features, and vehicle sensing features, the normal probability of each vehicle light function is calculated, and the final state and abnormality indicator of each vehicle light function are output.

[0013] Preferably, the visual detection sub-model is trained separately for different vehicle light functions, including one or more of the following: headlights, turn signals, brake lights, and combination lights.

[0014] Preferably, each image showing a change in vehicle headlights is input into a pre-trained visual detection sub-model, which outputs a single-image headlight state matrix, including: For each type of vehicle light function to be detected, an independent deep learning target visual detection sub-model is pre-trained. Each visual detection sub-model is used to identify the on / off state of the corresponding vehicle light in the image. For each input image showing changes in vehicle lights, inference is performed sequentially through all visual detection sub-models to obtain the on / off probability p of each vehicle light. i,j i represents the image number, and j represents the headlight number; The probability p of each car light being on or off i,j Compared with the preset confidence threshold θ, if p i,j If the value is greater than θ, the light is determined to be on, with a state value of 1; otherwise, it is off, with a state value of 0. Construct a single-image light state matrix; the rows of the single-image light state matrix correspond to each image, the columns correspond to each vehicle light, and each element in the single-image light state matrix is ​​the binarized on / off state of the corresponding vehicle light in that image.

[0015] Preferably, the single-image light state matrix is ​​aggregated and analyzed according to the operation time period to obtain the comprehensive state matrix of each light during the entire operation time period, including the following steps: Using the aggregated analysis results of the single-image light state matrix, the number of frames in which each vehicle light is determined to be lit in all image frames within the operation time period is counted, and the lighting ratio is calculated. The lighting ratio of each vehicle light is compared with a preset threshold to obtain the static lighting determination result of that vehicle light; The sequence of the on / off states of each vehicle light during the operation time period is input into the timing model to extract timing features, which include one or more of the following: flashing period, on-time, on-time duration, and off-time. Based on the preset lamp group association rules and extracted timing features, the time synchronization of the on / off states of the lamps in the same group is calculated to obtain the linkage features; The obtained lighting ratio, timing characteristics, and linkage characteristics are integrated into a comprehensive state matrix.

[0016] Preferably, the lighting ratio of each vehicle light during the operating time period is calculated according to the following formula: Where, p j represents the illumination ratio of any vehicle headlight; n represents the total number of image frames acquired during the operation time period. This represents the number of frames the vehicle light illuminates during the current operating time period.

[0017] Preferably, the multimodal fusion model formula is as follows: in, This represents a multimodal fusion model; This represents the comprehensive state matrix obtained by aggregation analysis based on the operation time period. O represents operation prompt information, and V represents vehicle signal data.

[0018] Preferably, when the turn signals or combination lights cannot be fully captured in a single image, a time weight w is introduced into the aggregation analysis. i The probability of lights turning on in each frame is weighted: in, s i , j This represents the detection status of the j-th car light in the i-th image; This represents the probability of lights being on in each frame of the image after weighting.

[0019] Preferably, the integrated state matrix, control mode data, and signal mode data are input into a multimodal fusion model, and the multimodal fusion model is used to extract structured visual features, control command features, and vehicle sensing features, including: For the comprehensive state matrix Numerical normalization is performed to map the lighting ratio, temporal features, and linkage features in the comprehensive state matrix to a unified numerical range, thereby obtaining structured visual features. The operation prompt information O is encoded, wherein the operation type is converted into a binary feature vector using one-hot encoding or embedding encoding, the duration of the operation time period is extracted, and minimum-maximum normalization is performed to map the operation duration to the [0,1] interval to obtain the control command features; The vehicle signal data V is normalized or statistically feature-extracted to obtain the vehicle sensing features; the normalization process includes one of min-max normalization or Z-score normalization. After the above processing, the integrated state matrix, the control modal data, and the signal modal data are converted into structured visual feature vectors, control command feature vectors, and vehicle sensor feature vectors of a unified scale, which are used as inputs to the multimodal fusion model.

[0020] Preferably, the structured visual features, control command features, and vehicle sensor features are input into a multimodal fusion module, and multimodal feature fusion is performed using feature stitching, attention mechanisms, or decision-level fusion strategies to obtain fused features; the fused features are used to calculate the normal probability of each vehicle light function, and the final state and anomaly indicator of each vehicle light function are output, including: The fused features are input into a fully connected classification network, and after layer-by-layer nonlinear transformation by the fully connected classification network, a logical value vector is output. The normal probability P of each vehicle light function is output through a Sigmoid or Softmax activation function. The normal probability P is compared with a preset judgment threshold. If P ≥ the preset judgment threshold, the light is judged to be functioning normally; otherwise, it is judged to be abnormal, and the corresponding abnormality flag is output. The anomaly identifier includes one or more fine-grained anomaly types, such as not being lit when it should be lit, always lit, insufficient brightness, or response delay.

[0021] This invention also provides a vehicle headlight functionality testing system based on a multimodal fusion model, used to implement the steps of the above-mentioned vehicle headlight functionality testing method based on a multimodal fusion model, including: A multimodal data acquisition layer is used to acquire multimodal data, including control modal data, visual modal data, and signal modal data. The control modal data includes operation prompts, specified operations to be performed according to the operation prompts, and corresponding operation time periods. The signal modal data includes vehicle signal data fed back by the vehicle when the specified operation is performed. The visual modal data includes multiple images of vehicle headlight changes continuously acquired during the operation time period. Visual detection and aggregation layer: This layer is used to input each image of the vehicle headlight change into a pre-trained visual detection sub-model and output a single-image headlight state matrix; the single-image headlight state matrix is ​​aggregated and analyzed according to the operation time period to obtain the comprehensive state matrix of each headlight throughout the entire operation time period. Multimodal fusion layer: Utilizes a multimodal fusion model to extract structured visual features, control command features, and vehicle sensor features. After fusing the structured visual features, control command features, and vehicle sensor features, calculates the normal probability of each vehicle light function and outputs the final state and anomaly indicator of each vehicle light function.

[0022] The vehicle headlight functionality detection method and system disclosed in this application, based on a multimodal fusion model, has at least the following advantages compared to existing technologies: This invention captures multiple images continuously over an operating period and performs time-segment aggregation analysis on the detection results of multiple frames. This allows for the complete capture of the entire process of dynamic lighting functions such as turn signal flashing, AUTO mode switching, and combined light linkage, avoiding missed detections or misjudgments caused by limitations in the timing or angle of a single image capture.

[0023] This invention constructs independent visual detection sub-models for different light functions. Each sub-model focuses on the on / off state recognition of a specific light type, improving the recognition accuracy in complex light combination scenarios. At the same time, by using a time period aggregation strategy to perform statistical analysis on the results of multiple frames, it effectively filters out false detections in a single frame and improves the stability of the detection results.

[0024] This invention performs multimodal fusion of image aggregation results, front-end operation prompts, and vehicle signal data, enabling a comprehensive assessment of whether the lighting functions respond as expected. For example, by combining brake pedal signals and brake light detection results, it can accurately determine whether the brake lights are functioning correctly; by combining turn signal lever signals and turn signal detection results, it can verify whether the turn signals illuminate as instructed. This fusion mechanism significantly improves the intelligence level and reliability of the detection system. Attached Figure Description

[0025] Figure 1 This is a schematic diagram of the main process of the vehicle headlight functionality detection method based on a multimodal fusion model according to the present invention.

[0026] Figure 2 This is a structural block diagram of the vehicle lighting functionality detection system based on a multimodal fusion model according to the present invention.

[0027] Figure 3 This is a complete flowchart of the vehicle headlight functionality detection method based on a multimodal fusion model according to the present invention. Detailed Implementation

[0028] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0029] like Figure 1 and Figure 3 As shown, one embodiment of the present invention provides a method for detecting the functionality of vehicle lights based on a multimodal fusion model, comprising the following steps: S1. Acquire multimodal data; the multimodal data includes control modal data, visual modal data, and signal modal data; The control modal data includes operation prompts, specified operations to be performed according to the operation prompts, and corresponding operation time periods; representing the functional logic that the tester expects to trigger.

[0030] The signal modal data includes vehicle signal data fed back by the vehicle when a specified operation is performed; and data from the vehicle bus, such as physical sensing information like brake pedal signals and steering lever positions, reflecting the true response state of the vehicle's underlying system.

[0031] The visual modal data originates from the image acquisition module, specifically manifested as a single-image light state matrix and a comprehensive state matrix aggregated over time periods, representing the actual on / off state of the lights in the physical world; specifically, it includes multiple images of vehicle light changes continuously acquired during the operation time period; each image is input into a pre-trained visual detection sub-model, which outputs a single-image light state matrix; the single-image light state matrix is ​​aggregated and analyzed according to the operation time period to obtain the comprehensive state matrix of each vehicle light throughout the entire operation time period; The three modalities correspond to the three levels of visual representation, operational intent, and physical mechanism, respectively, and together they form a complete information loop from data collection to fusion and judgment.

[0032] The visual detection sub-model is trained separately for different light functions, including one or more of the following: headlights, turn signals, brake lights, and combination lights.

[0033] During implementation, the system provides operational prompts to testers through the front-end interface, such as prompting them to turn on the left turn signal, brake, or switch to AUTO mode. It also records the start and end times of these operations, marking them as the detection time period [ts,te][t_s, t_e][ts,te]. Within this time period, the system continuously acquires images I1, I2,...,InI_1, I_2, ..., I_nI1,I2,...,In. Each image is input into the visual detection sub-model for processing, resulting in a single-image light state matrix S.

[0034] S2. Input each image of the vehicle headlight change into a pre-trained visual detection sub-model and output a single-image headlight state matrix; perform aggregate analysis on the single-image headlight state matrix according to the operation time period to obtain the comprehensive state matrix of each headlight in the entire operation time period. Each image is input into a pre-trained visual detection sub-model, which outputs a single-image light state matrix, including: For each light function to be detected, an independent deep learning visual detection sub-model for object detection is pre-trained. Each visual detection sub-model is used to identify the on / off state of the corresponding light in the image. For each input image, inference is performed sequentially through all visual detection sub-models to obtain the on / off probability p of each car light. i,j i represents the image number, and j represents the headlight number; The probability p of each car light being on or off i,j Compared with the preset confidence threshold θ, if p i,j If the value is greater than θ, the light is determined to be on, with a state value of 1; otherwise, it is off, with a state value of 0. According to the preset threshold Determine if the light is on: Arrange the binarized states of all lights according to the image rows and light columns to construct a single-image light state matrix. j Let p represent the lighting ratio of any vehicle headlight. After the visual detection sub-model determines the probability of a vehicle headlight being on or off, the ratio of the number of image frames with the headlight on to the total number of captured frames is considered the lighting ratio of any vehicle headlight. j, Meets preset threshold Only then will the headlights be considered lit; if the number of frames with the lights on is too low, they will be considered not lit.

[0035] The single-image lamp state matrix S is represented as follows: Among them, s i,j This represents the detection status of the j-th car light in the i-th image, where 1 indicates the light is on and 0 indicates the light is off.

[0036] Further preferably, the single-image light state matrix is ​​aggregated and analyzed according to the operation time period to obtain a comprehensive state matrix for each vehicle light throughout the entire operation time period: the comprehensive state matrix is ​​a condensed record of the state of all lights during the entire operation time period. It not only records how many times each vehicle light is lit throughout the process (lighting ratio), but also reflects the linkage relationship between lights (whether the turn signals on the same side are synchronized), and the change pattern of the state over time (such as flashing frequency and response speed). Simply put, it is a complete behavioral data table describing "how the lights are lit and for how long" during the entire operation process.

[0037] Specifically, the following steps are included: Calculate the lighting percentage of each vehicle light during the operating time period; further, the lighting percentage of each vehicle light during the operating time period is calculated according to the following formula: Where, p j represents the illumination ratio of any vehicle headlight; n represents the total number of image frames acquired during the operation time period. This represents the number of frames the vehicle light illuminates during the current operating time period.

[0038] The lighting ratio of each vehicle light is compared with a preset threshold to obtain the static lighting determination result of that vehicle light.

[0039] It should be noted that the time-period aggregation analysis in this application uses statistical methods, such as calculating the lighting ratio and then determining it through a threshold. This is static aggregation. This application also provides another embodiment that uses temporal modeling, taking a multi-frame detection state sequence as input, inputting the on / off state sequence of each vehicle light within the operation time period into the temporal model, and extracting temporal features. The temporal features include one or more of the following: flashing period, lighting time, lighting duration, and extinguishing time. Based on the preset lamp group association rules and extracted timing features, the time synchronization of the on / off states of the lamps in the same group is calculated to obtain the linkage features; The obtained lighting ratio, timing characteristics, and linkage characteristics are integrated into a comprehensive state matrix. In addition, it may be necessary to combine operation time period information (such as operation start and end times) to align the timing.

[0040] The state sequence of each headlight in multiple frames of images is input into the time series model to extract time series features, which are used to determine whether the dynamic characteristics of the headlight function are normal.

[0041] The temporal model can be LSTM, GRU, or Transformer. The input is a sequence of states (possibly with confidence levels or probabilities), and the output is the normal probability or anomaly type (such as abnormal flicker frequency, response delay, etc.) to obtain the comprehensive state matrix of each vehicle light throughout the entire operation period. The output of the temporal model can be further combined with a multimodal fusion model, or directly used as one of the final decision criteria. Temporal modeling can also be combined with multimodal fusion, for example, by inputting temporal features as another modality into the multimodal fusion model.

[0042] S3. Input the comprehensive state matrix, control modal data and signal modal data into the multimodal fusion model, use the multimodal fusion model to extract structured visual features, control command features and vehicle sensing features, and after fusing the structured visual features, control command features and vehicle sensing features, calculate the normal probability of each vehicle light function, and output the final state and abnormality indicator of each vehicle light function.

[0043] A further preferred embodiment of the multimodal fusion model is as follows: in, Represents a multimodal fusion model This represents a comprehensive matrix based on aggregated analysis of the operation time period. O represents operation prompt information, and V represents vehicle signal data.

[0044] The integrated state matrix, control mode data, and signal mode data are further input into the multimodal fusion model. The multimodal fusion model is then used to extract structured visual features, control command features, and vehicle-mounted sensor features, including: For the comprehensive state matrix Numerical normalization is performed to map the lighting ratio, temporal features, and linkage features in the comprehensive state matrix to a unified numerical range, thereby obtaining structured visual features. The operation prompt information O is encoded, wherein the operation type is converted into a binary feature vector using one-hot encoding or embedding encoding, the duration of the operation time period is extracted and subjected to minimum-maximum normalization, so that the operation duration is mapped to the [0,1] interval to obtain the control command feature; The vehicle signal data V is normalized or statistically feature-extracted to obtain the vehicle sensing features; the normalization process includes one of min-max normalization or Z-score normalization. After the above processing, the integrated state matrix, the control modal data, and the signal modal data are converted into structured visual feature vectors, control command feature vectors, and vehicle sensing feature vectors of a unified scale, which are used as inputs to the multimodal fusion model.

[0045] The structured visual features, control command features, and vehicle sensor features are input into the multimodal fusion module. Multimodal feature fusion is performed using feature concatenation, attention mechanisms, or decision-level fusion strategies to obtain fused features. The structured visual features originate from image aggregation results and are already light-up ratios within the 0-1 range, thus possessing a uniform scale. The operation type in the control command features is converted to 0 / 1 binary features using one-hot encoding, while the operation duration is separately mapped to the 0-1 range using minimum-maximum normalization. Although the vehicle signal data in the vehicle sensor features has different dimensions such as voltage and speed in its original form, the influence of these dimensions has been eliminated during feature extraction through Z-score standardization or normalization. All three feature vectors are dimensionless or at a uniform scale when entering the fusion module, preventing model training imbalance due to dimensional differences.

[0046] The fused features are input into a fully connected classification network, and the normal probability P of each headlight function is output through a Sigmoid or Softmax activation function. It should be noted that the process from feature fusion to final probability output is a gradual mapping process of feature transformation and dimensionality compression, specifically including three distinct sub-steps: The first step is feature encoding and alignment. Before being input into the fully connected network, the three feature vectors are first processed through a uniform embedding layer or normalization layer to ensure that they are on the same numerical scale (e.g., mean 0, variance 1), in preparation for subsequent fusion.

[0047] The second step is fusion and dimensionality reduction. The three feature vectors are concatenated or weighted and then fed into a classification head composed of fully connected layers. This process typically involves 2-3 layers of fully connected networks, each followed by a ReLU activation function and a Dropout layer. The dimensionality of the hidden layers decreases layer by layer (e.g., 512 → 256 → 128) to extract a high-dimensional abstract representation from the fused features, while gradually compressing the dimensionality and preventing overfitting.

[0048] The third step is probability mapping and judgment. The output dimension of the last fully connected network layer is equal to the number of light functions to be detected. By using a Sigmoid activation function (for multi-label classification) or a Softmax activation function (for multi-class classification), the output values ​​are mapped to the [0,1] interval to obtain the normal probability P of each light function. Finally, P is compared with a preset threshold to obtain the final judgment result.

[0049] The normal probability P is compared with a preset judgment threshold. If P ≥ the threshold, the light is judged to be functioning normally; otherwise, it is judged to be abnormal, and the corresponding abnormality identifier is output. The anomaly identifier includes one or more fine-grained anomaly types, such as not being lit when it should be lit, always lit, insufficient brightness, or response delay.

[0050] Specifically, the fully connected classification network acts as a trained scorer. It receives the concatenated fused features as input and gradually extracts key information and compresses dimensionality through nonlinear transformations of multiple neural networks. The output of the final layer then passes through a Sigmoid or Softmax activation function, forcibly transforming the original values ​​into a probability range between 0 and 1, ultimately obtaining the normal probability P for each headlight function. This probability value represents the model's confidence in the normality of the headlight function. Subsequently, by comparing it with a preset threshold, the model can automatically determine whether the function is normal, thus completing a complete inference loop from multimodal data input to the final detection result.

[0051] In practical applications, the system displays the status of each vehicle light function on the front-end interface to determine if it is functioning correctly. For any abnormal light functions detected, the system can generate a detection report or save a log record, achieving comprehensive monitoring and verification of the lighting functions. By combining front-end operation prompts, time-segment image acquisition, a deep learning-based visual detection sub-model, and multimodal information fusion, this invention effectively solves the problem that a single image cannot fully capture the lighting status and significantly improves the accuracy and reliability of lighting function detection.

[0052] In optional implementations, the visual detection sub-model can be trained independently for each light function module, adapting to different light patterns and combinations. The multimodal fusion model can employ feature stitching, attention mechanisms, or other deep learning fusion methods to comprehensively analyze image features, operation commands, and vehicle signals, outputting the final light function status. Time-segment aggregation strategies can use majority voting, lighting ratio thresholds, or weighted statistical methods to ensure the accuracy of the aggregation results in judging the light status. The system can also be expanded to detect combined functions such as high beams, fog lights, and logo lights to adapt to different vehicle models and functional requirements, and can be widely applied in R&D testing, quality control, and vehicle function verification scenarios to achieve intelligent detection of vehicle lighting systems.

[0053] A further preferred embodiment includes introducing time weights into the aggregation analysis when turn signals or combination lights cannot be fully captured in a single image. The probability of lights turning on in each frame of the image is weighted: in, s i , j This indicates the on / off status of the vehicle lights during the current operation period. This represents the probability of a light being on in each frame of the image after weighting. The numerator is the on / off state of each light multiplied by its weight (which can also be considered as confidence level), representing the number of reliable lights. The denominator is the sum of all confidence levels, representing the total number of valid observations. The final result is the probability value between [0,1].

[0054] This invention can be widely applied to R&D testing, quality control, vehicle function verification, regulatory compliance testing, and intelligent transportation systems. In R&D testing, it can guide operators to complete specified operations through front-end prompts, and automatically determine the normality of lighting functions through multi-time period data collection and multimodal fusion. In quality control, the system can automatically perform testing after the vehicle rolls off the production line, quickly generating functional status reports and improving testing efficiency and coverage. In regulatory compliance testing, this method ensures that lighting functions meet standard requirements. In intelligent transportation scenarios, this method can also provide data support for vehicle function monitoring and remote diagnostics.

[0055] like Figure 2 As shown, the present invention also provides a vehicle headlight functionality detection system based on a multimodal fusion model, comprising: A multimodal data acquisition layer is used to acquire multimodal data, including control modal data, visual modal data, and signal modal data. The control modal data includes operation prompts, specified operations to be performed according to the operation prompts, and corresponding operation time periods. The signal modal data includes vehicle signal data fed back by the vehicle when the specified operation is performed. The visual modal data includes multiple images of vehicle headlight changes continuously acquired during the operation time period. Visual detection and aggregation layer: This layer is used to input each image of the vehicle headlight change into a pre-trained visual detection sub-model and output a single-image headlight state matrix; the single-image headlight state matrix is ​​aggregated and analyzed according to the operation time period to obtain the comprehensive state matrix of each headlight throughout the entire operation time period. Multimodal fusion layer: Utilizes a multimodal fusion model to extract structured visual features, control command features, and vehicle sensor features. After fusing the structured visual features, control command features, and vehicle sensor features, calculates the normal probability of each vehicle light function and outputs the final state and anomaly indicator of each vehicle light function.

[0056] The following example, using the detection of a vehicle's left turn signal function, provides a detailed explanation of the technical solution of this invention in conjunction with specific data. This example is only used to illustrate the working principle of this invention and does not constitute a limitation on the scope of protection.

[0057] Suppose we need to test the left turn signal function of a test vehicle. The testing targets include: Check if the left turn signal can be turned on normally according to the operating instructions; Does the flashing frequency of the left turn signal meet the standard (e.g., 60-120 times per minute)? Does the left turn signal turn off normally after the turn signal lever returns to its original position?

[0058] During the operation period, the system continuously acquired images of the left front area of ​​the vehicle at a frame rate of 10 frames per second, for a total of 30 images.

[0059] Each image is input into a pre-trained left turn signal detection sub-model. The model outputs the state probability of the left turn signal in each image.

[0060] In actual testing, the left turn signal illuminated approximately 0.3 seconds after the operation began, remained off for about 0.8 seconds, and then remained off until the data collection was completed. This reflects one normal flashing cycle of the left turn signal.

[0061] The percentage of the left turn signal that is illuminated during the entire operation period is calculated, for example, if the number of frames is 8. If the percentage of illuminated signals is greater than a preset threshold, the aggregated result determines that the left turn signal was illuminated during the operation. If only a single image is used (e.g., the first image), it will be mistakenly judged as the light not being on, while the aggregated analysis correctly reflects the actual illumination status of the light.

[0062] To further verify whether the turn signal flashing frequency is normal, a timing analysis was performed on the state sequence. The length of the continuous lighting interval was extracted: the number of lighting frames was 8, corresponding to a time length of 8 × 0.1 = 0.88 × 0.1 = 0.8 seconds; the number of off frames was 22, but this needs to be considered in conjunction with the flashing period. Assuming the standard requires a flashing frequency of 1.5Hz (i.e., a period of approximately 0.67 seconds, with each on / off period approximately 0.33 seconds), the actual measured single-light duration of 0.8 seconds is too long. However, a single flash is insufficient for judgment; multiple flashes need to be observed. In this example, only one activation operation was collected, which may only include one flash. Therefore, the timing analysis can output "single-light duration 0.8 seconds" as a timing feature.

[0063] Use the following data as multimodal input: Aggregation status (indicating whether the light is on); Operation prompt: Operation type is "left turn", operation duration is 3 seconds; Vehicle signal data: The steering lever signal was read from the CAN bus, showing that the lever was in the "left turn" position during the test and returned to its original position 0.1 seconds later.

[0064] The above information is encoded and then input into the multimodal fusion model. The fusion model uses an attention mechanism to assign weights to different modalities. For example, for turn signal detection, the operation type and the lever signal have higher weights.

[0065] Fusion model output: The probability P that the left turn signal function is normal; The judgment threshold is set to 0.7. If P > 0.7, the final state is judged as "normal". Anomaly indicator: No serious anomaly, but the timing characteristics indicate "light-up duration is too long", therefore the fine-grained output prompts "flickering frequency is slightly low, check recommended".

[0066] For brake light detection, the system will prompt the tester to press the brake pedal, acquire images, aggregate and analyze them, and perform multimodal fusion in conjunction with the brake pedal signal. For AUTO mode detection, the system will prompt the tester to switch to AUTO mode, simulate changes in ambient light (such as obstructing the light sensor), acquire images, aggregate and analyze them, and verify whether the headlights automatically illuminate in conjunction with the ambient light signal.

[0067] As can be seen from the above examples, this invention solves the problem of insufficient information in a single frame image by acquiring images over multiple time periods, improves detection robustness by aggregating and analyzing images over time periods, and integrates operation commands, vehicle signals, and visual results through multimodal fusion, thereby achieving high-precision and intelligent judgment of complex lighting functions.

[0068] Obviously, the above embodiments are merely illustrative examples for clear explanation and are not intended to limit the implementation. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is neither necessary nor possible to exhaustively list all possible implementations here. However, obvious variations or modifications derived therefrom are still within the scope of protection of this invention.

Claims

1. A method for detecting the functionality of vehicle lights based on a multimodal fusion model, characterized in that, Including the following steps: Acquire multimodal data; the multimodal data includes control modal data, visual modal data, and signal modal data; the control modal data includes operation prompt information, as well as the specified operation to be performed according to the operation prompt information and the corresponding operation time period; the signal modal data includes the vehicle's feedback signal when the specified operation is performed; the visual modal data includes multiple images of vehicle headlight changes continuously acquired during the operation time period; Each image showing a change in vehicle headlights is input into a pre-trained visual detection sub-model, which outputs a single-image headlight state matrix. The single-image light state matrix is ​​aggregated and analyzed according to the operation time period to obtain the comprehensive state matrix of each light during the entire operation time period; The single-image light state matrix is ​​aggregated and analyzed according to the operation time period to obtain the comprehensive state matrix of each light during the entire operation time period, including the following steps: Using the aggregated analysis results of the single-image light state matrix, the number of frames in which each vehicle light is determined to be lit in all image frames within the operation time period is counted, and the lighting ratio is calculated. The lighting ratio of each vehicle light is compared with a preset threshold to obtain the static lighting determination result of that vehicle light; The sequence of the on / off states of each vehicle light during the operation time period is input into the timing model to extract timing features, which include one or more of the following: flashing period, on-time, on-time duration, and off-time. Based on the preset combination light association rules and extracted timing features, the time synchronization of the on / off states of the lights in the same group is calculated to obtain the linkage features; The obtained lighting ratio, timing characteristics, and linkage characteristics are integrated into a comprehensive state matrix. The integrated state matrix, control modal data, and signal modal data are input into the multimodal fusion model. The multimodal fusion model is used to extract structured visual features, control command features, and vehicle sensing features. After fusing the structured visual features, control command features, and vehicle sensing features, they are mapped to the normal probability of each vehicle light function, and the final state and abnormality identifier of each vehicle light function are output.

2. The method for detecting vehicle headlight functionality based on a multimodal fusion model according to claim 1, characterized in that, The visual detection sub-model is trained separately for different vehicle light functions, including one or more of the following: headlights, turn signals, brake lights, and combination lights.

3. The method for detecting vehicle headlight functionality based on a multimodal fusion model according to claim 2, characterized in that, Each image showing a change in vehicle headlights is input into a pre-trained visual detection sub-model, which outputs a single-image headlight state matrix, including: For each type of vehicle light function to be detected, an independent deep learning visual detection sub-model for object detection is pre-trained. Each visual detection sub-model is used to identify the on / off state of the corresponding vehicle light in the image. For each input image showing changes in vehicle lights, inference is performed sequentially through all visual detection sub-models to obtain the on / off probability of each vehicle light. pi,j i represents the image number, and j represents the headlight number; The probability of each car light turning on or off pi,j Compared with the preset confidence threshold θ If a comparison is made, pi,j >θ If the light is on, the status value is 1; otherwise, it is off, and the status value is 0. Construct a single-image light state matrix; the rows of the single-image light state matrix correspond to each image, the columns correspond to each vehicle light, and each element in the single-image light state matrix is ​​the binarized on / off state of the corresponding vehicle light in that image.

4. The method for detecting vehicle headlight functionality based on a multimodal fusion model according to claim 3, characterized in that, The lighting ratio of each vehicle light during the operating time period is calculated according to the following formula: in, pj represents the illumination ratio of any vehicle headlight; n represents the total number of image frames acquired during the operation time period. This represents the number of frames the vehicle light illuminates during the current operating time period.

5. The method for detecting vehicle headlight functionality based on a multimodal fusion model according to claim 1, characterized in that, The formula for the multimodal fusion model is as follows: in, This represents a multimodal fusion model; This represents the comprehensive state matrix obtained by aggregation analysis based on the operation time period. O represents operation prompt information, and V represents optional vehicle signal data.

6. The method for detecting vehicle headlight functionality based on a multimodal fusion model according to claim 3, characterized in that, This also includes introducing time weights w into the aggregation analysis when turn signals or combination lights cannot be fully captured in a single image. i The probability of lights turning on in each frame of the image is weighted: in, si,j To represent the detection state of the j-th car light in the i-th image; is the weighted probability of the light turning on for each frame, and n is the total number of image frames acquired during the operation time period.

7. The method for detecting vehicle headlight functionality based on a multimodal fusion model according to claim 6, characterized in that, The integrated state matrix, control mode data, and signal mode data are input into a multimodal fusion model. The multimodal fusion model is then used to extract structured visual features, control command features, and vehicle-mounted sensor features, including: For the comprehensive state matrix Numerical normalization is performed to map the lighting ratio, temporal features, and linkage features in the comprehensive state matrix to a unified numerical range, resulting in structured visual features. The operation prompt information O is encoded, wherein the operation type is converted into a binary feature vector using one-hot encoding or embedding encoding, the duration of the operation time period is extracted and subjected to minimum-maximum normalization, so that the operation duration is mapped to the [0,1] interval, and the control command features are obtained. The vehicle signal data V is normalized or statistically feature-extracted to obtain the vehicle sensing features; the normalization process includes one of min-max normalization or Z-score normalization. The integrated state matrix, control modal data, and signal modal data processed as described above are converted into structured visual feature vectors, control command feature vectors, and vehicle sensor feature vectors of a unified scale, which serve as inputs to the multimodal fusion model.

8. The method for detecting vehicle headlight functionality based on a multimodal fusion model according to claim 7, characterized in that, The structured visual features, control command features, and vehicle sensing features are input into the multimodal fusion module, and multimodal feature fusion is performed using feature stitching, attention mechanism, or decision-level fusion strategy to obtain fused features; The fused features are mapped to the normal probability of each vehicle light function, and the final state and anomaly identifier of each vehicle light function are output, including: The fused features are input into a fully connected classification network, and after layer-by-layer nonlinear transformation by the fully connected classification network, a logical value vector is output. The normal probability P of each vehicle light function is output through a Sigmoid or Softmax activation function. The normal probability P is compared with a preset judgment threshold. If P ≥ the preset judgment threshold, the light is judged to be functioning normally; otherwise, it is judged to be abnormal, and the corresponding abnormality flag is output. The anomaly identifier includes one or more fine-grained anomaly types, such as the one that is not lit, is constantly lit, has insufficient brightness, or has a delayed response.

9. A vehicle headlight functionality testing system based on a multimodal fusion model, used to implement the steps of the vehicle headlight functionality testing method based on a multimodal fusion model as described in any one of claims 1-8, characterized in that, include: A multimodal data acquisition layer is used to acquire multimodal data, including control modal data, visual modal data, and signal modal data. The control modal data includes operation prompts, specified operations to be performed according to the operation prompts, and corresponding operation time periods. The signal modal data includes vehicle feedback signals when the specified operations are performed. The visual modal data includes multiple continuously acquired images of vehicle headlight changes during the operation time period. Visual detection and aggregation layer: This layer is used to input each image of the vehicle headlight change into a pre-trained visual detection sub-model and output a single-image headlight state matrix. The single-image light state matrix is ​​aggregated and analyzed according to the operation time period to obtain the comprehensive state matrix of each light during the entire operation time period; Multimodal fusion layer: Utilizes a multimodal fusion model to extract structured visual features, control command features, and vehicle sensing features. After fusing the structured visual features, control command features, and vehicle sensing features, the results are mapped to the normal probability of each vehicle light function, and the final state and anomaly identifier of each vehicle light function are output.

Citation Information

Patent Citations

  • Vehicle lamp auditing method and device

    CN113869106A

  • Automobile light state automatic detection system and method

    CN120056865A