Intelligent vehicle-mounted interaction system and method based on multi-sensor fusion

By using a light anomaly detection model and a time synchronization model, the spatiotemporal synchronization problem in multi-sensor data fusion was solved, achieving deep collaboration between vision and radar, improving the robustness and response speed of gesture recognition, and ensuring the stable operation of the in-vehicle interactive system in complex environments.

CN122018671APending Publication Date: 2026-05-12JIANGSU HAIDA DYEING & PRINTING MACHINERY
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JIANGSU HAIDA DYEING & PRINTING MACHINERY
Filing Date
2025-12-08
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing gesture recognition systems suffer from spatiotemporal synchronization issues in multi-sensor data fusion, resulting in insufficient recognition accuracy and real-time performance in complex environments, failing to meet the requirements for efficient and stable gesture recognition.

Method used

By establishing a light anomaly detection model, a feature action encoding module, and a time synchronization model, and utilizing multi-sensor fusion technology, light anomalies can be accurately identified, a three-dimensional spatial distribution can be constructed, deep collaboration between vision and radar can be achieved, time deviation can be eliminated, and the stability and response speed of action recognition can be improved.

Benefits of technology

The robustness and response speed of gesture recognition have been improved in complex environments, ensuring stable operation and high-precision recognition under different lighting conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122018671A_ABST
    Figure CN122018671A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent vehicle-mounted interaction system and method based on multi-sensor fusion, and relates to the technical field of multi-source data fusion. According to the intelligent vehicle-mounted interaction system and method based on the multi-sensor fusion, interaction records with characteristic action codes are obtained, the difference value of timestamps between continuous image frames and recognition actions is calculated, and based on characteristic light scores, movement scores and the difference value of the timestamps, the interaction records are obtained; establishing a time synchronization model; according to real-time data, whether light is abnormal or not is judged, if light is abnormal, a real-time timestamp difference value is calculated, and continuous image frames and recognition actions are synchronized, so that the robustness and response speed of gesture recognition in a complex environment are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of multi-source data fusion technology, specifically to an intelligent in-vehicle interaction system and method based on multi-sensor fusion. Background Technology

[0002] Existing gesture recognition technology is widely used in the field of multi-sensor data fusion, especially for the combination of visual cameras and radar sensors, which has become an important part of intelligent vehicle gesture interaction systems. Existing gesture recognition systems are mostly based on traditional single sensors or simple sensor data fusion strategies, which often cannot operate stably in complex environments. Existing gesture recognition systems typically rely on simple feature-level or decision-level fusion methods, failing to adequately consider the temporal synchronization issues of data from different sensors. In multi-sensor systems, clock drift and differences in sampling frequencies lead to time deviations between data, causing error accumulation and response delays during the fusion process, which affects recognition accuracy and real-time performance. Existing spatiotemporal synchronization methods suffer from high algorithm complexity and poor real-time performance when dealing with accurate synchronization of heterogeneous data, making it impossible to meet the requirements for efficient and stable gesture recognition. Therefore, how to solve the problem of spatiotemporal synchronization among multiple sensors, reduce errors in the data fusion process, and achieve deep collaboration between vision and millimeter-wave radar to improve the robustness and response speed of gesture recognition in complex environments has become a technical challenge that urgently needs to be overcome in the field of intelligent interaction. Summary of the Invention

[0003] The purpose of this invention is to provide an intelligent in-vehicle interaction system and method based on multi-sensor fusion to solve the problems raised in the prior art.

[0004] To address the aforementioned technical problems, this invention provides the following technical solution: an intelligent in-vehicle interaction method based on multi-sensor fusion, the method comprising: Step S100: Calculate the feature light score based on the light parameters of the historical recognition records, count the abnormalities of the historical recognition records, calculate the label value of the historical recognition records, identify the historical abnormal recognition records, assign values ​​to the historical recognition records, and establish a light anomaly judgment model based on the feature light score and the assigned values ​​of the historical recognition records. Step S200: Set the training period, determine the light anomaly recognition record based on the light anomaly judgment model, obtain the action information collected by the radar sensing device in the light anomaly recognition record, identify the user's action according to the action database, calculate the motion score according to the motion data of the action, combine the action with the action score level and encode it, and determine the feature action code according to the frequency of occurrence of the action code. Step S300: Obtain interaction records with feature action encoding, calculate the difference between the timestamps of consecutive image frames and the recognized actions, and establish a time synchronization model based on the difference between feature light score, motion score, and timestamp. Step S400: Based on real-time data, determine whether there is an abnormality in lighting. If there is an abnormality in lighting, calculate the real-time timestamp difference and synchronize the continuous image frames with the recognition action.

[0005] Furthermore, step S100 includes: Step S101: Deploy several devices in the in-vehicle interactive device, preset the action recognition cycle, when the sound sensor recognizes the user's voice call to the in-vehicle interactive device, collect the current time point as the initial time point, start the radar sensing device and visual acquisition device in the in-vehicle interactive device, collect the user's action video according to the action recognition cycle, and analyze the collected video to obtain the user's action command. At the same time, through the light sensing device, collect the light parameters in the vehicle, combine the collected action video, the analyzed action command and the light parameters to generate the interaction record corresponding to the action recognition cycle. When the in-vehicle interactive device determines that the voice call has ended, collect the current time point as the end time point, set the time period between the initial time point and the end time point as the interaction time period, number the interaction records sequentially according to the timestamp of the interaction records within the interaction time period, construct a complete recognition record, and upload the recognition record to the in-vehicle interactive platform; Step S102: Obtain the historical recognition record set, collect the light parameters of each interaction record in the historical recognition record, normalize each light parameter, preset the weight of each light parameter, and perform a weighted sum of the normalized light parameters and the corresponding weights to calculate the light score of each interaction record. Step S103: Collect the recognition result of each interaction record in the historical recognition records, mark the interaction records with normal recognition results as normal interaction records, and mark the interaction records with abnormal recognition results as abnormal interaction records. Calculate the feature ray score of the historical recognition records according to the following formula: ; Where A represents the characteristic ray score, B a Let B represent the ray score of the a-th normal interaction record. d Let C1 represent the light score of the dth abnormal interaction record, C2 represent the weight of the normal interaction record, b represent the total number of normal interaction records, and e represent the total number of abnormal interaction records. Step S104: Count the number of normal interaction records and abnormal interaction records in each historical identification record, and calculate the tag value of the historical identification record according to the following formula: ; Where D represents the tag value, F1 represents the weight of the number of normal interaction records, and F2 represents the weight of the number of abnormal interaction records. The average value of the tag values ​​is calculated as f1 and the standard deviation as f2, and k is preset as the threshold coefficient. The tag value threshold D1 is calculated as D1 = f1 + k × f2. Step S105: If the label value of a historical identification record does not exceed the label value threshold, the historical identification record is marked as a historical normal identification record. If the label value of a historical identification record exceeds the label value threshold, the historical identification record is marked as a historical abnormal identification record. The historical identification records are assigned values ​​according to historical normal identification records and historical abnormal identification records. The feature ray score and the assigned value of each historical identification record are combined to obtain a feature ray score training set. The feature ray score is used as input and the assigned value is used as output. The model is trained through a random forest model. The training process involves setting several candidate probability thresholds and evaluating the performance index of the corresponding model for each candidate probability threshold. The performance index score corresponding to each candidate probability threshold is calculated. The candidate probability threshold corresponding to the best performance index score is selected as the probability threshold to generate a ray anomaly judgment model. By combining motion video, motion commands, and lighting parameters, it is possible to accurately identify and judge lighting anomalies in in-vehicle interactive devices, thereby improving the user experience. Automatically acquire and process historical interaction records, using light scoring, feature light scoring, and label values ​​to effectively distinguish between normal and abnormal records, and ensure efficient data labeling and classification; By calculating and setting adaptive label value thresholds, the system can dynamically adjust according to different usage environments, thereby improving the accuracy of judgment. The light anomaly detection system was trained using a random forest model, and the best performance metrics were evaluated based on candidate probability thresholds to ensure the efficiency and stability of the detection model.

[0006] Furthermore, step S200 includes: Step S201: Select several consecutive days as the training period, obtain the recognition records within the training period, collect the light parameters of each interaction record in each recognition record, and calculate the feature light score of the recognition record. Input the feature light score into the light anomaly judgment model to obtain the light anomaly probability. If the light anomaly probability exceeds the probability threshold, mark the recognition record as a light anomaly recognition record and summarize it. Step S202: In the light anomaly identification record set, obtain the action information collected by the radar sensing device in the light anomaly identification record. The action information includes the distance, azimuth angle and elevation angle of the target point. Collect the action information in each action identification cycle. Through the conversion from polar coordinates to Cartesian coordinates, convert the distance, azimuth angle and elevation angle of each target point into three-dimensional spatial coordinates. Obtain the continuous three-dimensional spatial coordinates in the action identification cycle and construct the three-dimensional spatial distribution of the action identification cycle. Step S203: Based on the three-dimensional spatial distribution of the continuous action recognition cycle, extract the feature data of each target point, and match the feature data with the action database according to the preset action database to identify the actions made by the user, and calculate the motion data of the actions. The motion data includes the velocity and acceleration of the action spatial trajectory. Normalize the motion data, preset the weights of the motion data, and perform weighted summation on the normalized motion data to calculate the motion score. Preset several action score levels and determine the action score range corresponding to each action score level. Step S204: Obtain the abnormal recognition records in the light abnormal recognition record set, extract the actions and corresponding motion rating levels corresponding to the abnormal interaction records in the abnormal recognition records, combine the actions and motion rating levels and encode them, count the number of abnormal interaction records corresponding to each action code, calculate the occurrence frequency of the action code, preset the occurrence frequency threshold, and mark the action codes that exceed the occurrence frequency threshold as feature action codes. The light anomaly detection model filters out the recognition records during the training period one by one, and only includes them in the analysis when the probability of light anomaly exceeds the threshold, which effectively avoids interference from irrelevant data and improves the accuracy of subsequent action recognition and analysis. By using radar sensing equipment to collect distance, azimuth and elevation angles, and by converting polar coordinates to Cartesian coordinates to construct a continuous three-dimensional spatial distribution, the system can still accurately capture human movements in abnormal lighting conditions, thus improving the stability and robustness of action recognition. By matching the three-dimensional spatial trajectory with the preset motion database, it can not only identify the motion type, but also perform a refined evaluation based on motion data such as speed and acceleration, so as to achieve a comprehensive judgment on the motion quality, amplitude and dynamic characteristics. By normalizing, weighting, and calculating motion scores, a standardized motion scoring system is constructed, making the analysis of motion features more objective and comparable, and providing a unified quantitative basis for subsequent model training and anomaly analysis. By statistically analyzing the frequency of each action code in the light anomaly recognition records and filtering out the feature action codes using frequency thresholds, it is possible to effectively extract "key action patterns that occur with high probability under light anomaly conditions," thereby providing a basis for tracing the causes of errors in light anomaly recognition and optimizing the model. By identifying the characteristic action codes corresponding to abnormal lighting, the in-vehicle interactive system can adjust the recognition logic or trigger compensation strategies accordingly, thereby improving the accuracy of action recognition in scenarios such as strong light, low light, and shadow.

[0007] Furthermore, step S300 includes: Step S301: In the light anomaly recognition record set, obtain the interaction record with characteristic action code, extract the action video collected by the visual acquisition device, identify and extract the continuous image frame corresponding to each action of the user through visual human motion capture technology, obtain the timestamp of each continuous image frame, extract the timestamp corresponding to the action identified by the radar sensing device, match the continuous image frame with the identified action, and calculate the difference between the timestamps of the continuous image frame and the identified action. Step S302: Classify the interaction records according to different actions. In the interaction record set of a certain action, use the feature light score and motion score as input and the difference of timestamps as output, train through a linear regression model to generate a time synchronization model. By extracting the timestamps of continuous image frames from the vision device and the action recognition timestamps from the radar sensing device, and calculating the time difference between the two, the system can accurately calibrate the time offset between different sensors, thereby improving the reliability of action recognition synchronization. For feature action encoding in abnormal lighting scenarios, deep fusion of visual and radar data is achieved through timestamp matching, which significantly reduces action recognition errors caused by insufficient lighting, occlusion or noise, and improves the consistency of multimodal data. The method takes the feature light score and motion score as input, and establishes the mapping relationship between "motion features - light state - time offset" by training a linear regression model to generate a time synchronization model that can predict time deviation, so that the model has adaptive compensation capability. Abnormal lighting conditions can lead to visual recognition delays or unstable frame capture. This method uses feature lighting scoring for correlation modeling, which can automatically compensate for the time shift caused by lighting differences, thereby improving the stability of visual recognition under different lighting conditions. By using a time synchronization model to calibrate the timing of action recognition, the differences in response speed, frame rate, and sampling delay between visual and radar devices can be eliminated, making the final action recognition results more accurate and consistent.

[0008] Furthermore, step S400 includes: Step S401: When the user makes a real-time voice call to the in-vehicle interactive device, the real-time light parameters inside the vehicle during the action recognition cycle are collected, the real-time light score is calculated, and the real-time light anomaly probability of the current action recognition cycle is determined by the light anomaly judgment model. If the real-time light anomaly probability exceeds the probability threshold, then step S402 is executed. If the real-time light anomaly probability does not exceed the probability threshold, then step S401 is repeated. Step S402: Collect real-time motion information through radar sensing device, calculate real-time motion score, input real-time light score and real-time motion score into time synchronization model, calculate real-time timestamp difference, and synchronize continuous image frames with recognized actions according to the real-time timestamp difference; By combining real-time light scoring with a light anomaly judgment model, the system can dynamically monitor and judge the impact of light changes on action recognition. When the probability of light anomaly exceeds the threshold, the system enters the compensation mode in time, thereby effectively dealing with the impact of ambient light fluctuations on recognition accuracy. By using real-time light score and motion score as input, the time synchronization model accurately calculates the timestamp difference, ensuring a high degree of synchronization between visual data and radar data, and significantly reducing motion recognition errors caused by sensor delay or time deviation. Under different lighting conditions, the system can adjust the recognition strategy in a timely manner through the collaboration of real-time light anomaly detection and motion scoring, thereby maintaining efficient interactive recognition performance, whether in strong light during the day or low light at night. The system can adaptively adjust its strategy in each motion recognition cycle based on changes in real-time light and motion scores, avoiding misidentification or delays caused by changes in light and motion, and ensuring that the in-vehicle interactive system can respond quickly to user needs. In cases of abnormal lighting conditions, real-time monitoring and synchronous model compensation can effectively eliminate the impact of lighting changes on image frame capture, maintain high accuracy and consistency in motion recognition, and avoid performance degradation of the recognition system due to lighting fluctuations.

[0009] To better implement the above methods, an intelligent vehicle interaction system based on multi-sensor fusion is also proposed. The system includes a light anomaly determination model module, a feature action encoding module, a time synchronization model module, and a real-time synchronization module. Light Anomaly Detection Model Module: Based on the light parameters of historical recognition records, the module calculates the feature light score, counts the anomalies in historical recognition records, calculates the label values ​​of historical recognition records, identifies historical anomaly recognition records, assigns values ​​to historical recognition records, and establishes a light anomaly detection model based on the feature light score and the assigned values ​​of historical recognition records. Feature Action Encoding Module: Set training period, determine light anomaly recognition records based on light anomaly judgment model, obtain action information collected by radar sensing device in light anomaly recognition records, identify user actions according to action database, calculate motion score based on motion data, combine actions with action score level and encode them, and determine feature action codes based on the frequency of occurrence of action codes. Time synchronization model module: acquires interaction records with coded features, calculates the difference between timestamps of consecutive image frames and recognized actions, and establishes a time synchronization model based on the difference between feature light score, motion score, and timestamp. Real-time synchronization module: Based on real-time data, it determines whether there is an abnormality in lighting. If there is an abnormality in lighting, it calculates the difference in real-time timestamps and synchronizes the continuous image frames with the recognition action.

[0010] Furthermore, the lighting anomaly detection model module includes a feature lighting score calculation unit and a lighting anomaly detection model building unit: The feature light scoring unit calculates the following: Several devices are deployed in the in-vehicle interactive device. A preset action recognition cycle is established. When the sound sensor recognizes a user's voice call to the in-vehicle interactive device, the current time point is collected as the initial time point. The radar sensing device and visual acquisition device in the in-vehicle interactive device are activated to collect user action videos according to the action recognition cycle. The collected videos are analyzed to obtain the user's action commands. Simultaneously, the light sensor collects the light parameters inside the vehicle. The collected action videos, analyzed action commands, and light parameters are combined to generate an interaction record corresponding to the action recognition cycle. When the in-vehicle interactive device determines that the voice call has ended, the current time point is collected as the end time point. The initial time point and the end time point are then compared... The time period between each interaction is defined as the interaction time period. Based on the timestamps of the interaction records within the interaction time period, the interaction records are sequentially numbered to construct a complete recognition record. The recognition record is then uploaded to the vehicle interaction platform. A historical recognition record set is obtained, and the light parameters of each interaction record in the historical recognition record are collected. Each light parameter is normalized and calculated. The weight of each light parameter is preset, and the normalized light parameters and their corresponding weights are weighted and summed to calculate the light score of each interaction record. The recognition results of each interaction record in the historical recognition record are collected. Interaction records with normal recognition results are marked as normal interaction records, and interaction records with abnormal recognition results are marked as abnormal interaction records. The feature light score of the historical recognition record is calculated. A lighting anomaly detection model unit is established: The number of normal and abnormal interaction records in each historical recognition record is counted, the label value of each historical recognition record is calculated, historical recognition records with abnormal interaction records are summarized, and the average and standard deviation of the label values ​​are calculated. A preset threshold coefficient is used to calculate the label value threshold. If the label value of a historical recognition record does not exceed the label value threshold, the historical recognition record is marked as a historical normal recognition record; if the label value of a historical recognition record exceeds the label value threshold, the historical recognition record is marked as a historical abnormal recognition record. Historical recognition records are assigned values ​​according to whether they are historical normal or historical abnormal recognition records. The feature lighting score and the assigned value of each historical recognition record are combined to obtain a feature lighting score training set. Using the feature lighting score as input and the assigned value as output, the model is trained using a random forest model. The training process involves presetting several candidate probability thresholds, evaluating the performance index of the corresponding model for each candidate probability threshold, calculating the performance index score corresponding to each candidate probability threshold, selecting the candidate probability threshold corresponding to the best performance index score as the probability threshold, and generating the lighting anomaly detection model.

[0011] Furthermore, the feature action coding module includes a light anomaly identification recording unit and a feature action coding unit: Determine the light anomaly identification record unit: Select several consecutive days as the training period, obtain the identification records within the training period, collect the light parameters of each interaction record in each identification record, and calculate the feature light score of the identification record. Input the feature light score into the light anomaly judgment model to obtain the light anomaly probability. If the light anomaly probability exceeds the probability threshold, mark the identification record as a light anomaly identification record and summarize it. The feature action encoding unit is determined as follows: In the light anomaly identification record set, action information collected by the radar sensing device in the light anomaly identification record is acquired. This action information includes the distance, azimuth, and elevation angle of the target point. Action information within each action recognition cycle is collected. Through polar coordinate to Cartesian coordinate conversion, the distance, azimuth, and elevation angle of each target point are converted into three-dimensional spatial coordinates. Continuous three-dimensional spatial coordinates within the action recognition cycle are obtained, and the three-dimensional spatial distribution of the action recognition cycle is constructed. Based on the three-dimensional spatial distribution of the continuous action recognition cycle, feature data of each target point is extracted. According to a preset action database, the feature data is matched with the action database to identify the user's actions and calculate the motion data of the actions. The motion data includes the velocity and acceleration of the action space trajectory. The motion data is normalized and calculated. The weights of the motion data are preset. The normalized motion data is then weighted and summed to calculate the motion score. Several motion score levels are preset, and the motion score range corresponding to each motion score level is determined. Anomaly recognition records are obtained from the light anomaly recognition record set. The actions corresponding to the abnormal interaction records in the anomaly recognition records and the corresponding motion score levels are extracted. The actions and motion score levels are combined and encoded. The number of abnormal interaction records corresponding to each action code is counted, and the occurrence frequency of the action code is calculated. An occurrence frequency threshold is preset, and action codes that exceed the occurrence frequency threshold are marked as feature action codes.

[0012] Furthermore, the time synchronization model module includes a unit for calculating the difference and a unit for establishing the time synchronization model: The difference calculation unit: In the light anomaly recognition record set, the interaction record with characteristic action code is obtained, the action video captured by the visual acquisition device is extracted, and the continuous image frame corresponding to each action of the user is identified and extracted by the visual human motion capture technology. The timestamp of each continuous image frame is obtained, the timestamp corresponding to the action identified by the radar sensing device is extracted, the continuous image frame is matched with the identified action, and the difference between the timestamps of the continuous image frame and the identified action is calculated. Establish a time synchronization model unit: classify the interaction records according to different actions. In the interaction record set of a certain action, use the feature light score and motion score as input and the difference of timestamps as output. Train the model through a linear regression model to generate a time synchronization model.

[0013] Furthermore, the real-time synchronization module includes a unit for determining light anomalies and a unit for implementing synchronization: Lighting anomaly determination unit: When the user makes a real-time voice call to the in-vehicle interactive device, the real-time lighting parameters inside the vehicle during the action recognition cycle are collected, the real-time lighting score is calculated, and the real-time lighting anomaly probability of the current action recognition cycle is determined through the lighting anomaly determination model. If the real-time lighting anomaly probability exceeds the probability threshold, step S402 is executed. If the real-time lighting anomaly probability does not exceed the probability threshold, step S401 is executed repeatedly. Synchronization Unit: Real-time motion information is collected through radar sensing equipment, real-time motion score is calculated, real-time light score and real-time motion score are input into time synchronization model, real-time timestamp difference is calculated, and synchronization between continuous image frames and recognized actions is performed based on the real-time timestamp difference.

[0014] Compared with the prior art, the beneficial effects of the present invention are: by integrating light parameters, historical recognition records and light scores, a light anomaly judgment model is established, which can more accurately identify light anomalies in the in-vehicle environment, thereby optimizing the response of in-vehicle interactive devices under different lighting conditions; By combining radar sensors, visual acquisition devices, motion scoring, and light scoring, and using a multi-sensor fusion method, the user's action commands can be identified and understood more accurately, avoiding the limitations of a single sensor and improving the reliability of action recognition. This invention, by constructing a time synchronization model, can achieve precise synchronization between continuous image frames and action recognition. This is of great importance for improving the real-time performance and response speed of in-vehicle interactive systems, especially the synchronization between abnormal lighting and action recognition, which is crucial for the real-time response of in-vehicle systems. By integrating multiple layers of models, including light anomaly detection, motion recognition, and time synchronization, this invention effectively improves the robustness of in-vehicle interactive systems in complex and changing environments. Even under complex lighting or motion conditions, the system can still operate stably. Attached Figure Description

[0015] Figure 1 This is a flowchart illustrating the intelligent in-vehicle interaction method based on multi-sensor fusion of the present invention. Figure 2 This is a schematic diagram of the structure of the intelligent in-vehicle interaction system based on multi-sensor fusion according to the present invention. Detailed Implementation

[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0017] Please see Figure 1 and Figure 2 This invention provides a technical solution: an intelligent in-vehicle interaction method based on multi-sensor fusion, the method comprising: Step S100: Calculate the feature light score based on the light parameters of the historical recognition records, count the abnormalities of the historical recognition records, calculate the label value of the historical recognition records, identify the historical abnormal recognition records, assign values ​​to the historical recognition records, and establish a light anomaly judgment model based on the feature light score and the assigned values ​​of the historical recognition records. Step S100 includes: Step S101: Deploy several devices in the in-vehicle interactive device, preset the action recognition cycle, when the sound sensor recognizes the user's voice call to the in-vehicle interactive device, collect the current time point as the initial time point, start the radar sensing device and visual acquisition device in the in-vehicle interactive device, collect the user's action video according to the action recognition cycle, and analyze the collected video to obtain the user's action command. At the same time, through the light sensing device, collect the light parameters in the vehicle, combine the collected action video, the analyzed action command and the light parameters to generate the interaction record corresponding to the action recognition cycle. When the in-vehicle interactive device determines that the voice call has ended, collect the current time point as the end time point, set the time period between the initial time point and the end time point as the interaction time period, number the interaction records sequentially according to the timestamp of the interaction records within the interaction time period, construct a complete recognition record, and upload the recognition record to the in-vehicle interactive platform; Step S102: Obtain the historical recognition record set, collect the light parameters of each interaction record in the historical recognition record, normalize each light parameter, preset the weight of each light parameter, and perform a weighted sum of the normalized light parameters and the corresponding weights to calculate the light score of each interaction record. Step S103: Collect the recognition result of each interaction record in the historical recognition records, mark the interaction records with normal recognition results as normal interaction records, and mark the interaction records with abnormal recognition results as abnormal interaction records. Calculate the feature ray score of the historical recognition records according to the following formula: ; Where A represents the characteristic ray score, B a Let B represent the ray score of the a-th normal interaction record. d Let C1 represent the light score of the dth abnormal interaction record, C2 represent the weight of the normal interaction record, b represent the total number of normal interaction records, and e represent the total number of abnormal interaction records. Step S104: Count the number of normal interaction records and abnormal interaction records in each historical identification record, and calculate the tag value of the historical identification record according to the following formula: ; Where D represents the tag value, F1 represents the weight of the number of normal interaction records, and F2 represents the weight of the number of abnormal interaction records. The average value of the tag values ​​is calculated as f1 and the standard deviation as f2, and k is preset as the threshold coefficient. The tag value threshold D1 is calculated as D1 = f1 + k × f2. Step S105: If the label value of a historical identification record does not exceed the label value threshold, the historical identification record is marked as a historical normal identification record. If the label value of a historical identification record exceeds the label value threshold, the historical identification record is marked as a historical abnormal identification record. The historical identification records are assigned values ​​according to historical normal identification records and historical abnormal identification records. The feature ray score and the assigned value of each historical identification record are combined to obtain a feature ray score training set. The feature ray score is used as input and the assigned value is used as output. The model is trained through a random forest model. The training process involves setting several candidate probability thresholds and evaluating the performance index of the corresponding model for each candidate probability threshold. The performance index score corresponding to each candidate probability threshold is calculated. The candidate probability threshold corresponding to the best performance index score is selected as the probability threshold to generate a ray anomaly judgment model. For example, suppose the historical identification records contain the following data: Interaction Log 1: Illumination parameter 200 lux, Action command: Gesture A; Interaction Log 2: Illumination parameter 180 lux, Action command: Gesture B; Interaction Log 3: Illumination parameter 220 lux, Action command: Gesture A; The normalized lighting parameters are calculated as follows: Interaction Log 1: Normalized value of illumination parameters: 0.75; Interaction Log 2: Normalized value of illumination parameters: 0.65; Interaction Log 3: Normalized value of lighting parameters: 0.80; Weighted summation calculation of ray score: Assume the weight of each ray parameter is 0.4: The lighting score for Interactive Record 1 is 0.30; The lighting score for Interactive Record 2 is 0.26; The lighting score for Interactive Record 3 is 0.32; Assuming there are 2 normal interaction records and 1 abnormal interaction record, the light scores for the normal records are interaction record 1 and interaction record 3, and the abnormal record is interaction record 2, then the feature light score A is calculated to be 0.1587.

[0018] Step S200: Set the training period, determine the light anomaly recognition record based on the light anomaly judgment model, obtain the action information collected by the radar sensing device in the light anomaly recognition record, identify the user's action according to the action database, calculate the motion score according to the motion data of the action, combine the action with the action score level and encode it, and determine the feature action code according to the comments of the occurrence of the action code. Step S200 includes: Step S201: Select several consecutive days as the training period, obtain the recognition records within the training period, collect the light parameters of each interaction record in each recognition record, and calculate the feature light score of the recognition record. Input the feature light score into the light anomaly judgment model to obtain the light anomaly probability. If the light anomaly probability exceeds the probability threshold, mark the recognition record as a light anomaly recognition record and summarize it. Step S202: In the light anomaly identification record set, obtain the action information collected by the radar sensing device in the light anomaly identification record. The action information includes the distance, azimuth angle and elevation angle of the target point. Collect the action information in each action identification cycle. Through the conversion from polar coordinates to Cartesian coordinates, convert the distance, azimuth angle and elevation angle of each target point into three-dimensional spatial coordinates. Obtain the continuous three-dimensional spatial coordinates in the action identification cycle and construct the three-dimensional spatial distribution of the action identification cycle. Step S203: Based on the three-dimensional spatial distribution of the continuous action recognition cycle, extract the feature data of each target point, and match the feature data with the action database according to the preset action database to identify the actions made by the user, and calculate the motion data of the actions. The motion data includes the velocity and acceleration of the action spatial trajectory. Normalize the motion data, preset the weights of the motion data, and perform weighted summation on the normalized motion data to calculate the motion score. Preset several action score levels and determine the action score range corresponding to each action score level. Step S204: Obtain the abnormal recognition records in the light abnormal recognition record set, extract the actions and corresponding motion rating levels corresponding to the abnormal interaction records in the abnormal recognition records, combine the actions and motion rating levels and encode them, count the number of abnormal interaction records corresponding to each action code, calculate the occurrence frequency of the action code, preset the occurrence frequency threshold, and mark the action codes that exceed the occurrence frequency threshold as feature action codes. Assuming a three-day consecutive period is selected as the training period, the system collects recognition records during these three days, including the interaction data for each record. Assume the record for the first day is as follows: Interaction Log 1: Illumination parameter 220 lux, Action command: Gesture A; Interaction Log 2: Illumination parameter 180 lux, Action command: Gesture B; The normalized value of the illumination parameters for interactive record 1 is 0.80, and the characteristic ray score is 0.32; The normalized value of the illumination parameters in Interactive Record 2 is 0.72, and the characteristic ray score is 0.288; Input the feature light score into the light anomaly detection model to obtain the light anomaly probability. Assuming the model threshold is 0.3, the anomaly probability of interaction record 1 is 0.35 and the anomaly probability of interaction record 2 is 0.25. Since the probability of abnormal lighting in interaction record 1 exceeds the threshold of 0.3, it is marked as an abnormal lighting identification record. Interaction record 2 does not exceed the threshold and is considered a normal record. Assuming that in the light anomaly recognition record set, the action information corresponding to interaction record 1 is: target point distance: 5 meters, the angle with the front of the vehicle is the azimuth angle: 30°, and the elevation angle relative to the ground is the pitch angle: 10°; Assuming the device records motion information every 5 seconds, and this is repeated 3 times, the following data is obtained: First motion recognition cycle: Target point (5m, 30°, 10°); Second action recognition cycle: target point (4.8m, 32°, 11°); Third action recognition cycle: target point (5.2m, 28°, 9°); To convert polar coordinates to Cartesian coordinates, use the following formula: X = r × cos(θ) × cos(θ); Y = r × sin(θ) × cos(θ); Z = r × sin(θ); The calculated action recognition cycle times are: X = 4.27, Y = 2.46, Z = 0.87 for the first action recognition cycle; X = 3.83, Y = 2.5, Z = 0.91 for the second action recognition cycle; and X = 4.54, Y = 2.44, Z = 0.81 for the third action recognition cycle. Suppose that a pre-set action database identifies the action corresponding to this set of data as "gesture A" attempting to adjust the air conditioner. Calculation of motion data: Spatial trajectory: Calculate the trajectory based on the changes in continuous three-dimensional spatial coordinates; Speed: Assuming the distance between trajectory points is 0.4 meters in 1 second, then the speed is 0.4 m / s; Acceleration: Assuming the velocity increases by 0.1 m / s per second, then the acceleration is 0.1 m / s². Assuming the weights for velocity and acceleration are 0.5, the motion score is 0.30; Assume the preset action rating levels are: 0.0-0.2 is low, 0.2-0.5 is medium, and 0.5-1.0 is high; The action received a score of 0.30, therefore the action rating is "Medium".

[0019] Step S300: Obtain interaction records with feature action encoding, calculate the difference between the timestamps of consecutive image frames and the recognized actions, and establish a time synchronization model based on the difference between feature light score, motion score, and timestamp. Step S300 includes: Step S301: In the light anomaly recognition record set, obtain the interaction record with characteristic action code, extract the action video collected by the visual acquisition device, identify and extract the continuous image frame corresponding to each action of the user through visual human motion capture technology, obtain the timestamp of each continuous image frame, extract the timestamp corresponding to the action identified by the radar sensing device, match the continuous image frame with the identified action, and calculate the difference between the timestamps of the continuous image frame and the identified action. Step S302: Classify the interaction records according to different actions. In the interaction record set of a certain action, use the feature light score and motion score as input and the difference of timestamps as output, train through a linear regression model to generate a time synchronization model.

[0020] Step S400: Based on real-time data, determine whether there is an abnormality in lighting. If there is an abnormality in lighting, calculate the real-time timestamp difference and synchronize the continuous image frames with the recognition action. Step S400 includes: Step S401: When the user makes a real-time voice call to the in-vehicle interactive device, the real-time light parameters inside the vehicle during the action recognition cycle are collected, the real-time light score is calculated, and the real-time light anomaly probability of the current action recognition cycle is determined by the light anomaly judgment model. If the real-time light anomaly probability exceeds the probability threshold, then step S402 is executed. If the real-time light anomaly probability does not exceed the probability threshold, then step S401 is repeated. Step S402: Collect real-time motion information through radar sensing equipment, calculate real-time motion score, input real-time light score and real-time motion score into time synchronization model, calculate real-time timestamp difference, and synchronize continuous image frames with recognized actions based on the real-time timestamp difference.

[0021] To better implement the above methods, an intelligent vehicle interaction system based on multi-sensor fusion is also proposed. The system includes a light anomaly determination model module, a feature action encoding module, a time synchronization model module, and a real-time synchronization module. Light Anomaly Detection Model Module: Based on the light parameters of historical recognition records, the module calculates the feature light score, counts the anomalies in historical recognition records, calculates the label values ​​of historical recognition records, identifies historical anomaly recognition records, assigns values ​​to historical recognition records, and establishes a light anomaly detection model based on the feature light score and the assigned values ​​of historical recognition records. The lighting anomaly detection model module includes a feature lighting score calculation unit and a lighting anomaly detection model establishment unit: The feature light scoring unit calculates the following: Several devices are deployed in the in-vehicle interactive device. A preset action recognition cycle is established. When the sound sensor recognizes a user's voice call to the in-vehicle interactive device, the current time point is collected as the initial time point. The radar sensing device and visual acquisition device in the in-vehicle interactive device are activated to collect user action videos according to the action recognition cycle. The collected videos are analyzed to obtain the user's action commands. Simultaneously, the light sensor collects the light parameters inside the vehicle. The collected action videos, analyzed action commands, and light parameters are combined to generate an interaction record corresponding to the action recognition cycle. When the in-vehicle interactive device determines that the voice call has ended, the current time point is collected as the end time point. The initial time point and the end time point are then compared... The time period between each interaction is defined as the interaction time period. Based on the timestamps of the interaction records within the interaction time period, the interaction records are sequentially numbered to construct a complete recognition record. The recognition record is then uploaded to the vehicle interaction platform. A historical recognition record set is obtained, and the light parameters of each interaction record in the historical recognition record are collected. Each light parameter is normalized and calculated. The weight of each light parameter is preset, and the normalized light parameters and their corresponding weights are weighted and summed to calculate the light score of each interaction record. The recognition results of each interaction record in the historical recognition record are collected. Interaction records with normal recognition results are marked as normal interaction records, and interaction records with abnormal recognition results are marked as abnormal interaction records. The feature light score of the historical recognition record is calculated. A lighting anomaly detection model unit is established: The number of normal and abnormal interaction records in each historical recognition record is counted, the label value of each historical recognition record is calculated, historical recognition records with abnormal interaction records are summarized, and the average and standard deviation of the label values ​​are calculated. A preset threshold coefficient is used to calculate the label value threshold. If the label value of a historical recognition record does not exceed the label value threshold, the historical recognition record is marked as a historical normal recognition record; if the label value of a historical recognition record exceeds the label value threshold, the historical recognition record is marked as a historical abnormal recognition record. Historical recognition records are assigned values ​​according to whether they are historical normal or historical abnormal recognition records. The feature lighting score and the assigned value of each historical recognition record are combined to obtain a feature lighting score training set. Using the feature lighting score as input and the assigned value as output, the model is trained using a random forest model. The training process involves presetting several candidate probability thresholds, evaluating the performance index of the corresponding model for each candidate probability threshold, calculating the performance index score corresponding to each candidate probability threshold, selecting the candidate probability threshold corresponding to the best performance index score as the probability threshold, and generating the lighting anomaly detection model.

[0022] Feature Action Encoding Module: Set training period, determine light anomaly recognition records based on light anomaly judgment model, obtain action information collected by radar sensing device in light anomaly recognition records, identify user actions according to action database, calculate motion score based on motion data, combine actions with action score level and encode them, and determine feature action codes based on comments on the occurrence of action codes. The feature action coding module includes a light anomaly identification recording unit and a feature action coding unit: Determine the light anomaly identification record unit: Select several consecutive days as the training period, obtain the identification records within the training period, collect the light parameters of each interaction record in each identification record, and calculate the feature light score of the identification record. Input the feature light score into the light anomaly judgment model to obtain the light anomaly probability. If the light anomaly probability exceeds the probability threshold, mark the identification record as a light anomaly identification record and summarize it. The feature action encoding unit is determined as follows: In the light anomaly identification record set, action information collected by the radar sensing device in the light anomaly identification record is acquired. This action information includes the distance, azimuth, and elevation angle of the target point. Action information within each action recognition cycle is collected. Through polar coordinate to Cartesian coordinate conversion, the distance, azimuth, and elevation angle of each target point are converted into three-dimensional spatial coordinates. Continuous three-dimensional spatial coordinates within the action recognition cycle are obtained, and the three-dimensional spatial distribution of the action recognition cycle is constructed. Based on the three-dimensional spatial distribution of the continuous action recognition cycle, feature data of each target point is extracted. According to a preset action database, the feature data is matched with the action database to identify the user's actions and calculate the motion data of the actions. The motion data includes the velocity and acceleration of the action space trajectory. The motion data is normalized and calculated. The weights of the motion data are preset. The normalized motion data is then weighted and summed to calculate the motion score. Several motion score levels are preset, and the motion score range corresponding to each motion score level is determined. Anomaly recognition records are obtained from the light anomaly recognition record set. The actions corresponding to the abnormal interaction records in the anomaly recognition records and the corresponding motion score levels are extracted. The actions and motion score levels are combined and encoded. The number of abnormal interaction records corresponding to each action code is counted, and the occurrence frequency of the action code is calculated. An occurrence frequency threshold is preset, and action codes that exceed the occurrence frequency threshold are marked as feature action codes.

[0023] Time synchronization model module: acquires interaction records with coded features, calculates the difference between timestamps of consecutive image frames and recognized actions, and establishes a time synchronization model based on the difference between feature light score, motion score, and timestamp. The time synchronization model module includes a difference calculation unit and a time synchronization model establishment unit: The difference calculation unit: In the light anomaly recognition record set, the interaction record with characteristic action code is obtained, the action video captured by the visual acquisition device is extracted, and the continuous image frame corresponding to each action of the user is identified and extracted by the visual human motion capture technology. The timestamp of each continuous image frame is obtained, the timestamp corresponding to the action identified by the radar sensing device is extracted, the continuous image frame is matched with the identified action, and the difference between the timestamps of the continuous image frame and the identified action is calculated. Establish a time synchronization model unit: classify the interaction records according to different actions. In the interaction record set of a certain action, use the feature light score and motion score as input and the difference of timestamps as output. Train the model through a linear regression model to generate a time synchronization model.

[0024] Real-time synchronization module: Based on real-time data, it determines whether there is an abnormality in lighting. If there is an abnormality in lighting, it calculates the real-time timestamp difference and synchronizes the continuous image frames with the recognition action. The real-time synchronization module includes a light anomaly detection unit and a synchronization implementation unit: Lighting anomaly determination unit: When the user makes a real-time voice call to the in-vehicle interactive device, the real-time lighting parameters inside the vehicle during the action recognition cycle are collected, the real-time lighting score is calculated, and the real-time lighting anomaly probability of the current action recognition cycle is determined through the lighting anomaly determination model. If the real-time lighting anomaly probability exceeds the probability threshold, step S402 is executed. If the real-time lighting anomaly probability does not exceed the probability threshold, step S401 is executed repeatedly. Synchronization Unit: Real-time motion information is collected through radar sensing equipment, real-time motion score is calculated, real-time light score and real-time motion score are input into time synchronization model, real-time timestamp difference is calculated, and synchronization between continuous image frames and recognized actions is performed based on the real-time timestamp difference.

[0025] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims.

Claims

1. A smart vehicle interaction method based on multi-sensor fusion, characterized in that, The methods include: Step S100: Calculate the feature light score based on the light parameters of the historical recognition records, count the abnormalities of the historical recognition records, calculate the label value of the historical recognition records, identify the historical abnormal recognition records, assign values ​​to the historical recognition records, and establish a light anomaly judgment model based on the feature light score and the assigned values ​​of the historical recognition records. Step S200: Set the training period, determine the light anomaly recognition record based on the light anomaly judgment model, obtain the action information collected by the radar sensing device in the light anomaly recognition record, identify the user's action according to the action database, calculate the motion score according to the motion data of the action, combine the action with the action score level and encode it, and determine the feature action code according to the frequency of occurrence of the action code. Step S300: Obtain interaction records with feature action encoding, calculate the difference between the timestamps of consecutive image frames and the recognized actions, and establish a time synchronization model based on the difference between feature light score, motion score, and timestamp. Step S400: Based on real-time data, determine whether there is an abnormality in lighting. If there is an abnormality in lighting, calculate the real-time timestamp difference and synchronize the continuous image frames with the recognition action.

2. The intelligent vehicle interaction method based on multi-sensor fusion according to claim 1, characterized in that, Step S100 includes the following steps: Step S101: Deploy several devices in the in-vehicle interactive device, preset the action recognition cycle, when the sound sensor recognizes the user's voice call to the in-vehicle interactive device, collect the current time point as the initial time point, start the radar sensing device and visual acquisition device in the in-vehicle interactive device, collect the user's action video according to the action recognition cycle, and analyze the collected video to obtain the user's action command. At the same time, collect the light parameters in the vehicle through the light sensor, combine the collected action video, the analyzed action command and the light parameters to generate the interaction record corresponding to the action recognition cycle. When the in-vehicle interactive device determines that the voice call has ended, collect the current time point as the end time point, set the time period between the initial time point and the end time point as the interaction time period, number the interaction records sequentially according to the timestamp of the interaction records within the interaction time period, construct a complete recognition record, and upload the recognition record to the in-vehicle interactive platform; Step S102: Obtain the historical recognition record set, collect the light parameters of each interaction record in the historical recognition record, normalize each light parameter, preset the weight of each light parameter, and perform a weighted sum of the normalized light parameters and the corresponding weights to calculate the light score of each interaction record. Step S103: Collect the recognition result of each interaction record in the historical recognition records, mark the interaction records with normal recognition results as normal interaction records, and mark the interaction records with abnormal recognition results as abnormal interaction records. Calculate the feature ray score of the historical recognition records according to the following formula: ; Where A represents the characteristic ray score, B a Let B represent the ray score of the a-th normal interaction record. d Let C1 represent the light score of the dth abnormal interaction record, C2 represent the weight of the normal interaction record, b represent the total number of normal interaction records, and e represent the total number of abnormal interaction records. Step S104: Count the number of normal interaction records and abnormal interaction records in each historical identification record, and calculate the tag value of the historical identification record according to the following formula: ; Where D represents the tag value, F1 represents the weight of the number of normal interaction records, and F2 represents the weight of the number of abnormal interaction records. The average value of the tag values ​​is calculated as f1 and the standard deviation as f2, and k is preset as the threshold coefficient. The tag value threshold D1 is calculated as D1 = f1 + k × f2. Step S105: If the label value of a historical identification record does not exceed the label value threshold, the historical identification record is marked as a historical normal identification record. If the label value of a historical identification record exceeds the label value threshold, the historical identification record is marked as a historical abnormal identification record. The historical identification records are assigned values ​​according to historical normal identification records and historical abnormal identification records. The feature ray score and the assigned value of each historical identification record are combined to obtain a feature ray score training set. The feature ray score is used as input and the assigned value is used as output. The model is trained through a random forest model. The training process involves setting several candidate probability thresholds and evaluating the performance index of the corresponding model for each candidate probability threshold. The performance index score corresponding to each candidate probability threshold is calculated. The candidate probability threshold corresponding to the best performance index score is selected as the probability threshold to generate the ray anomaly judgment model.

3. The intelligent vehicle interaction method based on multi-sensor fusion according to claim 2, characterized in that, Step S200 includes the following steps: Step S201: Select several consecutive days as the training period, obtain the recognition records within the training period, collect the light parameters of each interaction record in each recognition record, and calculate the feature light score of the recognition record. Input the feature light score into the light anomaly judgment model to obtain the light anomaly probability. If the light anomaly probability exceeds the probability threshold, mark the recognition record as a light anomaly recognition record and summarize it. Step S202: In the light anomaly identification record set, obtain the action information collected by the radar sensing device in the light anomaly identification record. The action information includes the distance, azimuth angle and elevation angle of the target point. Collect the action information in each action identification cycle. Through the conversion from polar coordinates to Cartesian coordinates, convert the distance, azimuth angle and elevation angle of each target point into three-dimensional spatial coordinates. Obtain the continuous three-dimensional spatial coordinates in the action identification cycle and construct the three-dimensional spatial distribution of the action identification cycle. Step S203: Based on the three-dimensional spatial distribution of the continuous action recognition cycle, extract the feature data of each target point, and match the feature data with the action database according to the preset action database to identify the actions made by the user, and calculate the motion data of the actions. The motion data includes the velocity and acceleration of the action spatial trajectory. Normalize the motion data, preset the weights of the motion data, and perform weighted summation on the normalized motion data to calculate the motion score. Preset several action score levels and determine the action score range corresponding to each action score level. Step S204: Obtain the abnormal recognition records in the light abnormality recognition record set, extract the actions and corresponding motion rating levels corresponding to the abnormal interaction records in the abnormal recognition records, combine the actions and motion rating levels and encode them, count the number of abnormal interaction records corresponding to each action code, calculate the occurrence frequency of the action code, preset the occurrence frequency threshold, and mark the action codes that exceed the occurrence frequency threshold as feature action codes.

4. The intelligent vehicle interaction method based on multi-sensor fusion according to claim 3, characterized in that, Step S300 includes the following steps: Step S301: In the light anomaly recognition record set, obtain the interaction record with characteristic action code, extract the action video collected by the visual acquisition device, identify and extract the continuous image frame corresponding to each action of the user through visual human motion capture technology, obtain the timestamp of each continuous image frame, extract the timestamp corresponding to the action identified by the radar sensing device, match the continuous image frame with the identified action, and calculate the difference between the timestamps of the continuous image frame and the identified action. Step S302: Classify the interaction records according to different actions. In the interaction record set of a certain action, use the feature light score and motion score as input and the difference of timestamps as output, train through a linear regression model to generate a time synchronization model.

5. The intelligent vehicle interaction method based on multi-sensor fusion according to claim 4, characterized in that, Step S400 includes the following steps: Step S401: When the user makes a real-time voice call to the in-vehicle interactive device, the real-time light parameters inside the vehicle during the action recognition cycle are collected, the real-time light score is calculated, and the real-time light anomaly probability of the current action recognition cycle is determined by the light anomaly judgment model. If the real-time light anomaly probability exceeds the probability threshold, then step S402 is executed. If the real-time light anomaly probability does not exceed the probability threshold, then step S401 is repeated. Step S402: Collect real-time motion information through radar sensing equipment, calculate real-time motion score, input real-time light score and real-time motion score into time synchronization model, calculate real-time timestamp difference, and synchronize continuous image frames with recognized actions based on the real-time timestamp difference.

6. An intelligent in-vehicle interaction system based on multi-sensor fusion, used to implement the intelligent in-vehicle interaction method based on multi-sensor fusion as described in any one of claims 1-5, characterized in that, The system includes a light anomaly detection model module, a feature action encoding module, a time synchronization model module, and a real-time synchronization module; The light anomaly determination model module calculates the feature light score based on the light parameters of historical recognition records, counts the anomalies in historical recognition records, calculates the label value of historical recognition records, identifies historical anomaly recognition records, assigns values ​​to historical recognition records, and establishes a light anomaly determination model based on the feature light score and the assigned values ​​of historical recognition records. The feature action encoding module: sets a training period, determines light anomaly recognition records based on the light anomaly judgment model, obtains action information collected by radar sensing devices in the light anomaly recognition records, identifies user actions according to the action database, calculates motion scores based on the motion data of the actions, combines and encodes the actions and action score levels, and determines feature action codes based on the frequency of occurrence of the action codes. The time synchronization model module: acquires interaction records with coded features, calculates the difference between timestamps of consecutive image frames and recognized actions, and establishes a time synchronization model based on the difference between feature light score, motion score, and timestamp. The real-time synchronization module determines whether there is an abnormality in lighting based on real-time data. If there is an abnormality in lighting, it calculates the real-time timestamp difference and synchronizes the continuous image frames with the recognition action.

7. The intelligent in-vehicle interaction system based on multi-sensor fusion according to claim 6, characterized in that, The light anomaly detection model module includes a feature light scoring unit and a light anomaly detection model building unit: The computational feature light scoring unit: Deploys several devices within the in-vehicle interactive device, presets an action recognition cycle, and when the sound sensor recognizes a user's voice call to the in-vehicle interactive device, collects the current time point as the initial time point, activates the radar sensing device and visual acquisition device in the in-vehicle interactive device, collects user action videos according to the action recognition cycle, analyzes the collected videos to obtain user action commands, and simultaneously collects light parameters inside the vehicle through the light sensing device. The collected action videos, analyzed action commands, and light parameters are combined to generate an interaction record corresponding to the action recognition cycle. When the in-vehicle interactive device determines that the voice call has ended, it collects the current time point as the end time point, sets the time period between the initial time point and the end time point as the interaction time period, and sequentially numbers the interaction records according to the timestamps of the interaction records within the interaction time period to construct a complete recognition record. The recognition record is then uploaded to the in-vehicle interactive platform. A set of historical recognition records is obtained, and the light parameters of each interaction record in the historical recognition records are collected. Each light parameter is normalized and calculated, and a weight is preset for each light parameter. The normalized light parameters and their corresponding weights are then weighted and summed to calculate the light score for each interaction record. Collect the recognition results of each interaction record in the historical recognition records, mark the interaction records with normal recognition results as normal interaction records, mark the interaction records with abnormal recognition results as abnormal interaction records, and calculate the feature light score of the historical recognition records; The unit for establishing the light anomaly determination model involves: counting the number of normal and abnormal interaction records in each historical recognition record; calculating the label value of each historical recognition record; summarizing historical recognition records with abnormal interaction records; calculating the average and standard deviation of the label values; setting a threshold coefficient; and calculating the label value threshold. If the label value of a historical recognition record does not exceed the label value threshold, the historical recognition record is marked as a historical normal recognition record; if the label value of a historical recognition record exceeds the label value threshold, the historical recognition record is marked as a historical abnormal recognition record. The historical recognition records are assigned values ​​according to whether they are historical normal or historical abnormal recognition records. The feature light score and assigned value of each historical recognition record are combined to obtain a feature light score training set. Using the feature light score as input and the assigned value as output, the model is trained using a random forest model. The training process involves setting several candidate probability thresholds, evaluating the performance index of the corresponding model for each candidate probability threshold, calculating the performance index score corresponding to each candidate probability threshold, selecting the candidate probability threshold corresponding to the best performance index score as the probability threshold, and generating the light anomaly determination model.

8. The intelligent in-vehicle interaction system based on multi-sensor fusion according to claim 6, characterized in that, The feature action coding module includes a light anomaly identification and recording unit and a feature action coding unit: The unit for determining light anomaly identification records: selects several consecutive days as the training period, acquires identification records within the training period, collects light parameters of each interactive record in each identification record, calculates the feature light score of the identification record, inputs the feature light score into the light anomaly determination model to obtain the light anomaly probability, and if the light anomaly probability exceeds the probability threshold, the identification record is marked as a light anomaly identification record and summarized. The defined action encoding unit: In the light anomaly identification record set, it acquires action information collected by the radar sensing device from the light anomaly identification records. This action information includes the distance, azimuth, and elevation angle of the target point. It collects action information within each action recognition cycle, and through polar to Cartesian coordinate conversion, transforms the distance, azimuth, and elevation angle of each target point into three-dimensional spatial coordinates, obtaining continuous three-dimensional spatial coordinates within the action recognition cycle, and constructing the three-dimensional spatial distribution of the action recognition cycle. Based on the three-dimensional spatial distribution of the continuous action recognition cycle, it extracts feature data for each target point, and matches the feature data with a preset action database to identify the user's actions and calculate the motion count of the actions. According to the data, the motion data includes the velocity and acceleration of the motion spatial trajectory. The motion data is normalized and calculated. The weights of the motion data are preset. The normalized motion data is then weighted and summed to calculate the motion score. Several motion score levels are preset, and the motion score range corresponding to each motion score level is determined. Anomaly recognition records are obtained from the light anomaly recognition record set. The actions and corresponding motion score levels corresponding to the abnormal interaction records in the anomaly recognition records are extracted. The actions and motion score levels are combined and encoded. The number of abnormal interaction records corresponding to each action code is counted, and the occurrence frequency of the action code is calculated. An occurrence frequency threshold is preset, and action codes that exceed the occurrence frequency threshold are marked as feature action codes.

9. The intelligent in-vehicle interaction system based on multi-sensor fusion according to claim 6, characterized in that, The time synchronization model module includes a difference calculation unit and a time synchronization model establishment unit: The difference calculation unit: In the light anomaly recognition record set, it obtains the interaction record with characteristic action code, extracts the action video collected by the visual acquisition device, identifies and extracts the continuous image frame corresponding to each action of the user through visual human motion capture technology, obtains the timestamp of each continuous image frame, extracts the timestamp corresponding to the action identified by the radar sensing device, matches the continuous image frame with the identified action, and calculates the difference between the timestamps of the continuous image frame and the identified action. The time synchronization model unit is established by classifying interaction records according to different actions. In the interaction record set of a certain action, the feature light score and motion score are used as inputs, and the difference of timestamps is used as output. The model is trained through a linear regression model to generate a time synchronization model.

10. The intelligent in-vehicle interaction system based on multi-sensor fusion according to claim 6, characterized in that, The real-time synchronization module includes a light anomaly detection unit and a synchronization implementation unit: The light anomaly determination unit: when the user makes a real-time voice call to the in-vehicle interactive device, it collects the real-time light parameters inside the vehicle during the action recognition cycle, calculates the real-time light score, and determines the real-time light anomaly probability of the current action recognition cycle through the light anomaly determination model. If the real-time light anomaly probability exceeds the probability threshold, step S402 is executed. If the real-time light anomaly probability does not exceed the probability threshold, step S401 is executed repeatedly. The synchronization unit implements the following: real-time motion information is collected through radar sensing equipment, real-time motion score is calculated, real-time light score and real-time motion score are input into time synchronization model, real-time timestamp difference is calculated, and synchronization between continuous image frames and recognized actions is performed based on the real-time timestamp difference.