Vehicle data processing method, vehicle, storage medium and program product
Patent Information
- Application Number
- CN202611112366.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-24
- Publication Date
- 2026-09-25
AI Technical Summary
[0005]本申请提供了一种车载数据的处理方法,该车载数据的处理方法能够通过融合多模态数据特征并映射至统一语义空间,解决了传统单一传感器在复杂环境下感知不准、误报率高的问题,再通过计算第一风险熵值和风险突变增益生成第二风险熵值作为综合风险评判指标,据此动态调整车载数据的传输等级,由此,实现了在车辆当前时刻的风险熵值较高时,将数据快速传输给相关系统,在车辆当前时刻的风险熵值较低或一般时存入本地存储或顺序上传,从而在保证高风险场景下数据的完整性与实时性的同时,有效节省低风险场景下车端的通信带宽与计算资源,以解决相关技术中数据分级依赖静态规则、缺乏上下文感知能力及实时性差的问题
[0033]基于上述技术方案,本申请实施例可以通过比对不同模态语义标记中的时间戳信息与空间坐标信息,计算两两之间的时间差与空间距离,量化不同模态语义标记在时空维度上的匹配程度,从而在某一模态的语义标记在时空维度上无法得到其他模态数据的印证时,削弱该模态数据的可信度,有效避免单一模态数据在等不同环境下的感知错误,保障数据的全面性与可靠性。
Smart Images

Figure CN122818249A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, specifically to a method for processing vehicle-mounted data, a vehicle, a storage medium, and a program product. Background Technology
[0002] With the rapid development of intelligent connected vehicle technology, the amount of data generated by a single vehicle as a mobile data collection terminal is growing exponentially.
[0003] In related technologies, in addition to traditional data processing methods, some methods have attempted to introduce large AI models to intelligently classify data in order to improve the accuracy of classification. For example, OCR (Optical Character Recognition) technology is used to recognize the text content in in-vehicle UI screenshots, and combined with large models to generate tags for data deduplication and compression; neural networks are used to identify data service types (such as ADAS (Advanced Driver Assistance System) types, media types, and log types), and select the upload channel accordingly.
[0004] However, in related technologies, traditional data processing solutions and data processing solutions that incorporate large AI models mostly focus on uploading data to the cloud for offline annotation or simulation. The processing cycle is long, making it difficult to meet the real-time requirements of the vehicle and limited by the computing power of the vehicle. The context awareness capability is also weak, making it difficult to guarantee the accuracy of data classification results. Moreover, most AI models are deployed in the cloud or trained offline, making it difficult to quickly deploy them on different vehicle models and different hardware platforms, which urgently needs to be addressed. Summary of the Invention
[0005] This application provides a method for processing vehicle-mounted data. This method solves the problems of inaccurate perception and high false alarm rate of traditional single sensors in complex environments by fusing multimodal data features and mapping them to a unified semantic space. Then, it generates a second risk entropy value as a comprehensive risk assessment index by calculating a first risk entropy value and a risk mutation gain. Based on this, the transmission level of vehicle-mounted data is dynamically adjusted. Thus, when the risk entropy value of the vehicle at the current moment is high, the data is quickly transmitted to the relevant system. When the risk entropy value of the vehicle at the current moment is low or moderate, the data is stored locally or uploaded sequentially. This ensures the integrity and real-time performance of the data in high-risk scenarios while effectively saving the communication bandwidth and computing resources of the vehicle in low-risk scenarios. This solves the problems of data classification relying on static rules, lack of context awareness, and poor real-time performance in related technologies.
[0006] In a first aspect, a method for processing vehicle-mounted data is provided, comprising the following steps: extracting multiple types of data features based on the multimodal data of the vehicle at the current moment, projecting the multiple types of data features onto a target semantic space to obtain spatiotemporally aligned fusion features; calculating a first risk entropy value of the vehicle at the current moment based on the fusion features, and calculating a risk mutation gain of the vehicle at the current moment based on the first risk entropy value; calculating a second risk entropy value at the current moment based on the first risk entropy value and the risk mutation gain, determining the transmission level of the multimodal data based on the second risk entropy value, and transmitting the multimodal data to the vehicle according to the transmission level.
[0007] Based on the above technical solution, the embodiments of this application can solve the problems of inaccurate perception and high false alarm rate of traditional single sensors in complex environments by fusing multimodal data features and mapping them to a unified semantic space. Then, by calculating the first risk entropy value to quantify the static risk of the vehicle at the current moment, a risk mutation gain is introduced to capture the dynamic change trend of the risk entropy value of the vehicle at the current moment compared with the risk entropy value at the previous moment. Finally, a second risk entropy value is generated as a comprehensive evaluation index, and the data transmission level is dynamically adjusted accordingly. Thus, when the risk entropy value of the vehicle at the current moment is high, the data is quickly transmitted to the relevant system, and when the risk entropy value of the vehicle at the current moment is low or average, it is stored locally or uploaded sequentially. This ensures the integrity and real-time performance of the data in high-risk scenarios while effectively saving the communication bandwidth and computing resources of the vehicle in low-risk scenarios, thereby solving the problems of data classification relying on static rules, lack of context awareness, and poor real-time performance in related technologies.
[0008] In combination with the first aspect and the above implementation methods, in some possible implementation methods, before extracting multiple types of data features based on the multimodal data of the vehicle at the current moment, the method further includes: acquiring the vehicle's visual data, radar data, and state data at the current moment; and determining the multimodal data based on the visual data, the radar data, and the state data.
[0009] Based on the above technical solution, the embodiments of this application can acquire multimodal data such as vehicle visual data, radar data, and state data. Visual data can provide environmental information and classification semantics, radar data can provide accurate spatial distance and speed information, and state data can reflect the vehicle's own motion state. The three complement each other, comprehensively covering the vehicle's data and providing a complete and reliable data foundation for subsequent feature extraction and fusion.
[0010] In combination with the first aspect and the above implementation methods, in some possible implementation methods, the step of extracting multiple types of data features based on the multimodal data includes: extracting image features of the vehicle at the current moment based on the visual data; extracting spatial geometric features of the vehicle at the current moment based on the radar data; extracting dynamic features of the vehicle at the current moment based on the state data; and determining multiple types of data features based on the image features, the spatial geometric features, and the dynamic features.
[0011] Based on the above technical solution, the embodiments of this application can extract corresponding features from multimodal data: image features, spatial geometric features and dynamic features, so as to preserve the integrity of the original information of each modality data to the greatest extent. Through this structured processing of feature extraction, the three can provide rich and physically meaningful information input for subsequent semantic alignment and fusion on the basis of carrying the original information of the original multimodal data.
[0012] In combination with the first aspect and the above implementation methods, in some possible implementation methods, the step of projecting multiple types of data features onto a target semantic space to obtain spatiotemporally aligned fusion features includes: projecting the image features, spatial geometric features, and dynamic features onto the target semantic space based on a pre-constructed shared semantic projection layer to generate visual semantic tags corresponding to the image features, spatial geometric semantic tags corresponding to the spatial geometric features, and motion state semantic tags corresponding to the dynamic features; calculating a spatiotemporal consistency score among the visual semantic tags, spatial geometric semantic tags, and motion state semantic tags to determine the fusion weights corresponding to the image features, spatial geometric features, and dynamic features based on the spatiotemporal consistency score; and fusing the visual semantic tags, spatial geometric semantic tags, and motion state semantic tags using the fusion weights to generate the spatiotemporally aligned fusion features.
[0013] Based on the above technical solution, the embodiments of this application can project multi-dimensional heterogeneous features into a unified high-dimensional semantic labeling space through a pre-constructed shared semantic projection layer, and introduce a spatiotemporal consistency score to quantify the degree of mutual corroboration between cross-modal data. Accordingly, the fusion weights between each modality data are adaptively adjusted. Thus, when a certain modality data is falsely reported, its weight will be dynamically reduced, thereby achieving noise suppression of data at the data feature layer and effectively improving the reliability of the final output fusion features.
[0014] In conjunction with the first aspect and the above implementation methods, in some possible implementation methods, the calculation of the spatiotemporal consistency score among the visual semantic marker, the spatial geometric semantic marker, and the motion state semantic marker includes: obtaining timestamp information and spatial coordinate information from the visual semantic marker, the spatial geometric semantic marker, and the motion state semantic marker; calculating the time difference and spatial distance between any two of the visual semantic marker, the spatial geometric semantic marker, and the motion state semantic marker based on the timestamp information and the spatial coordinate information; and calculating the spatiotemporal consistency score based on the spatial distance and the time difference.
[0015] Based on the above technical solution, the embodiments of this application can compare the timestamp information and spatial coordinate information in different modal semantic tags, calculate the time difference and spatial distance between each pair, quantify the matching degree of different modal semantic tags in the spatiotemporal dimension, thereby weakening the credibility of the modal data when the semantic tag of a certain modality cannot be verified by other modal data in the spatiotemporal dimension, effectively avoiding the perception error of single modal data in different environments, and ensuring the comprehensiveness and reliability of the data.
[0016] In combination with the first aspect and the above implementation methods, in some possible implementation methods, the step of calculating the first risk entropy value of the vehicle at the current moment based on the fusion features includes: acquiring the motion state and environmental information of the vehicle at the current moment; identifying the actual driving scenario of the vehicle at the current moment based on the motion state and the environmental information; generating the calculation weight of the scenario classification score of the vehicle at the current moment based on the actual driving scenario, so as to calculate the first risk entropy value based on the scenario classification score and its calculation weight.
[0017] Based on the above technical solution, embodiments of this application can generate a scenario classification score corresponding to the current driving scenario based on the vehicle's current motion state and environmental information. This allows the risk entropy value to effectively reflect the differentiated risk sensitivity requirements of the vehicle in different driving environments. Furthermore, embodiments of this application can incorporate the scenario classification score into the calculation of the first risk entropy value, effectively avoiding false alarms or missed alarms caused by using a uniform threshold to evaluate all driving scenarios.
[0018] In combination with the first aspect and the above implementation methods, in some possible implementation methods, the step of calculating the first risk entropy value based on the scenario classification score includes: obtaining the relative distance and relative speed between the vehicle and the target vehicle located in front of the vehicle; calculating the objective collision risk of the vehicle at the current moment based on the relative distance and relative speed; and calculating the first risk entropy value based on the scenario classification score and its calculation weight and the objective collision risk and its calculation weight.
[0019] Based on the above technical solution, the embodiments of this application can determine the objective collision risk of a vehicle at the current moment based on the relative distance and relative speed between the vehicle and the target vehicle in front of it, and then combine the subjective scene classification score with the objective collision risk to form a risk quantification mechanism that integrates subjective and objective factors, thereby improving the reliability and accuracy of the first risk entropy value.
[0020] In combination with the first aspect and the above implementation methods, in some possible implementation methods, the step of calculating the risk mutation gain of the vehicle at the current moment based on the first risk entropy value includes: obtaining the final risk entropy value of the vehicle at a historical moment adjacent to the current moment; calculating the incremental change rate between the first risk entropy value and the final risk entropy value at the historical moment; and determining the risk mutation gain based on the incremental change rate.
[0021] Based on the above technical solution, this application embodiment can measure the suddenness of risk by comparing the incremental change rate between the current first risk entropy value and the final risk entropy value at the target historical time. This effectively captures the sudden changes that may occur in the risk entropy value over time. That is, although the current first risk entropy value is not high, if it jumps sharply in a very short time, it indicates that the vehicle may be in sudden danger (such as the vehicle in front braking suddenly), and high-speed data transmission is also required. By introducing this risk mutation gain term, the ability of this application to predict dynamic risks in advance can be increased, and the data delay between the occurrence of risk and system response can be significantly shortened.
[0022] In combination with the first aspect and the above implementation methods, in some possible implementation methods, the step of calculating the second risk entropy value at the current moment based on the first risk entropy value and the risk mutation gain includes: determining the calculation weight corresponding to the risk mutation gain and the calculation weight corresponding to the scene classification score and the objective collision risk in the first risk entropy value based on the actual driving scenario of the vehicle at the current moment; and calculating the second risk entropy value based on the risk mutation gain, the scene classification score, the objective collision risk and its corresponding calculation weight.
[0023] Based on the above technical solution, the embodiments of this application can adaptively adjust the calculation weights of risk mutation gain, scene classification score, and objective collision risk according to the actual driving scenario, so as to achieve adaptive weighted fusion of the three in different scenarios. This allows the vehicle to automatically adjust the degree of importance given to each risk factor such as risk mutation gain, scene classification score, and objective collision risk in different driving environments, ensuring that the final generated second risk entropy value can include the consideration of risk trend mutation and scene tolerance quantification while taking into account the physical danger level of the current scenario, thus providing an accurate and reliable decision basis for subsequent data transmission level determination.
[0024] Secondly, a vehicle data processing device is provided, comprising: an extraction module, configured to extract multiple data features based on multimodal data of a vehicle at the current moment, and project the multiple data features onto a target semantic space to obtain spatiotemporally aligned fusion features; a calculation module, configured to calculate a first risk entropy value of the vehicle at the current moment based on the fusion features, and calculate a risk mutation gain of the vehicle at the current moment based on the first risk entropy value; and a processing module, configured to calculate a second risk entropy value at the current moment based on the first risk entropy value and the risk mutation gain, to determine the transmission level of the multimodal data based on the second risk entropy value, and transmit the multimodal data to the vehicle according to the transmission level.
[0025] Based on the above technical solution, the embodiments of this application can solve the problems of inaccurate perception and high false alarm rate of traditional single sensors in complex environments by fusing multimodal data features and mapping them to a unified semantic space. Then, by calculating the first risk entropy value to quantify the static risk of the vehicle at the current moment, a risk mutation gain is introduced to capture the dynamic change trend of the risk entropy value of the vehicle at the current moment compared with the risk entropy value at the previous moment. Finally, a second risk entropy value is generated as a comprehensive evaluation index, and the data transmission level is dynamically adjusted accordingly. Thus, when the risk entropy value of the vehicle at the current moment is high, the data is quickly transmitted to the relevant system, and when the risk entropy value of the vehicle at the current moment is low or average, it is stored locally or uploaded sequentially. This ensures the integrity and real-time performance of the data in high-risk scenarios while effectively saving the communication bandwidth and computing resources of the vehicle in low-risk scenarios, thereby solving the problems of data classification relying on static rules, lack of context awareness, and poor real-time performance in related technologies.
[0026] In combination with the first aspect and the above implementation methods, some possible implementation methods further include: an acquisition module, used to acquire visual data, radar data, and state data of the vehicle at the current moment before extracting multiple types of data features based on the multimodal data of the vehicle at the current moment; and a determination module, used to determine the multimodal data based on the visual data, the radar data, and the state data.
[0027] Based on the above technical solution, the embodiments of this application can acquire multimodal data such as vehicle visual data, radar data, and state data. Visual data can provide environmental information and classification semantics, radar data can provide accurate spatial distance and speed information, and state data can reflect the vehicle's own motion state. The three complement each other, comprehensively covering the vehicle's data and providing a complete and reliable data foundation for subsequent feature extraction and fusion.
[0028] In combination with the first aspect and the above implementation methods, in some possible implementation methods, the extraction module includes: a first extraction unit, used to extract image features of the vehicle at the current moment based on the visual data; a second extraction unit, used to extract spatial geometric features of the vehicle at the current moment based on the radar data; a third extraction unit, used to extract dynamic features of the vehicle at the current moment based on the state data; and a first determination unit, used to determine multiple types of data features based on the image features, the spatial geometric features, and the dynamic features.
[0029] Based on the above technical solution, the embodiments of this application can extract corresponding features from multimodal data: image features, spatial geometric features and dynamic features, so as to preserve the integrity of the original information of each modality data to the greatest extent. Through this structured processing of feature extraction, the three can provide rich and physically meaningful information input for subsequent efficient semantic alignment and fusion on the basis of carrying the original information of the original multimodal data.
[0030] In combination with the first aspect and the above implementation methods, in some possible implementation methods, the extraction module includes: a projection unit, used to project the image features, the spatial geometric features, and the dynamic features onto a target semantic space based on a pre-constructed shared semantic projection layer, to generate visual semantic tags corresponding to the image features, spatial geometric semantic tags corresponding to the spatial geometric features, and motion state semantic tags corresponding to the dynamic features; a first calculation unit, used to calculate the spatiotemporal consistency score between the visual semantic tags, the spatial geometric semantic tags, and the motion state semantic tags, so as to determine the fusion weights corresponding to the image features, the spatial geometric features, and the dynamic features based on the spatiotemporal consistency score; and a fusion unit, used to fuse the visual semantic tags, the spatial geometric semantic tags, and the motion state semantic tags using the fusion weights to generate the spatiotemporally aligned fusion features.
[0031] Based on the above technical solution, the embodiments of this application can project multi-dimensional heterogeneous features into a unified high-dimensional semantic labeling space through a pre-constructed shared semantic projection layer, and introduce a spatiotemporal consistency score to quantify the degree of mutual corroboration between cross-modal data. Accordingly, the fusion weights between each modality data are adaptively adjusted. Thus, when a certain modality data is falsely reported, its weight will be dynamically reduced, thereby achieving noise suppression of data at the data feature layer and effectively improving the reliability of the final output fusion features.
[0032] In combination with the first aspect and the above implementation methods, in some possible implementation methods, the calculation unit includes: a first acquisition subunit, used to acquire timestamp information and spatial coordinate information from the visual semantic tag, the spatial geometric semantic tag, and the motion state semantic tag; a first calculation subunit, used to calculate the time difference and spatial distance between any two of the visual semantic tag, the spatial geometric semantic tag, and the motion state semantic tag based on the timestamp information and the spatial coordinate information; and a second calculation subunit, used to calculate the spatiotemporal consistency score based on the spatial distance and the time difference.
[0033] Based on the above technical solution, the embodiments of this application can compare the timestamp information and spatial coordinate information in different modal semantic tags, calculate the time difference and spatial distance between each pair, quantify the matching degree of different modal semantic tags in the spatiotemporal dimension, thereby weakening the credibility of the modal data when the semantic tag of a certain modality cannot be verified by other modal data in the spatiotemporal dimension, effectively avoiding the perception error of single modal data in different environments, and ensuring the comprehensiveness and reliability of the data.
[0034] In combination with the first aspect and the above implementation methods, in some possible implementation methods, the calculation module includes: a first acquisition unit, used to acquire the vehicle's motion state and environmental information at the current moment; an identification unit, used to identify the actual driving scenario of the vehicle at the current moment based on the motion state and the environmental information; and a generation unit, used to generate a calculation weight for the scenario classification score of the vehicle at the current moment based on the actual driving scenario, so as to calculate the first risk entropy value based on the scenario classification score and its calculation weight.
[0035] Based on the above technical solution, this embodiment can generate a scenario classification score corresponding to the current driving scenario based on the vehicle's current motion state and environmental information. This allows the risk entropy value to effectively reflect the differentiated risk sensitivity requirements of the vehicle in different driving environments. Furthermore, by incorporating the scenario classification score into the calculation of the first risk entropy value, this embodiment can effectively avoid false alarms or missed alarms caused by using a uniform threshold to evaluate all driving scenarios.
[0036] In combination with the first aspect and the above implementation methods, in some possible implementation methods, the generation unit includes: a second acquisition subunit, used to acquire the relative distance and relative speed between the vehicle and a target vehicle located in front of the vehicle; a third calculation subunit, used to calculate the objective collision risk of the vehicle at the current moment based on the relative distance and relative speed; and a fourth calculation subunit, used to calculate the first risk entropy value based on the scene classification score and its calculation weight and the objective collision risk and its calculation weight.
[0037] Based on the above technical solution, the embodiments of this application can determine the objective collision risk of a vehicle at the current moment based on the relative distance and relative speed between the vehicle and the target vehicle in front of it, and then combine the subjective scene classification score with the objective collision risk to form a risk quantification mechanism that integrates subjective and objective factors, thereby improving the reliability and accuracy of the first risk entropy value.
[0038] In combination with the first aspect and the above implementation methods, in some possible implementation methods, the calculation module includes: a second acquisition unit, configured to acquire the final risk entropy value of the vehicle at a historical time adjacent to the current time; a second calculation unit, configured to calculate the incremental change rate between the first risk entropy value and the final risk entropy value at the historical time; and a first determination unit, configured to determine the risk mutation gain based on the incremental change rate.
[0039] Based on the above technical solution, this application embodiment can measure the suddenness of risk by comparing the incremental change rate between the current first risk entropy value and the final risk entropy value at a target historical time. This effectively captures potential abrupt changes in the risk entropy value over time. That is, although the current first risk entropy value is not high, if it experiences a sharp jump in a very short time, it indicates that the vehicle may face a sudden danger (such as a vehicle in front braking suddenly), requiring high-speed data transmission. By introducing this risk mutation gain term, the application's ability to predict dynamic risks in advance can be increased, significantly shortening the data delay between the occurrence of a risk and the system response.
[0040] In combination with the first aspect and the above implementation methods, in some possible implementation methods, the processing module includes: a second determining unit, used to determine the calculation weight corresponding to the risk mutation gain and the calculation weight corresponding to the scene classification score and the objective collision risk in the first risk entropy value based on the actual driving scenario of the vehicle at the current moment; and a third calculating unit, used to calculate the second risk entropy value based on the risk mutation gain, the scene classification score, the objective collision risk and its corresponding calculation weight.
[0041] Based on the above technical solution, the embodiments of this application can adaptively adjust the calculation weights of risk mutation gain, scene classification score, and objective collision risk according to the actual driving scenario, so as to achieve adaptive weighted fusion of the three in different scenarios. This allows the vehicle to automatically adjust the degree of importance given to each risk factor such as risk mutation gain, scene classification score, and objective collision risk in different driving environments, ensuring that the final generated second risk entropy value can include the consideration of risk trend mutation and scene tolerance quantification while taking into account the physical danger level of the current scenario, thus providing an accurate and reliable decision basis for subsequent data transmission level determination.
[0042] Thirdly, a vehicle is provided, the vehicle comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the processor executing the program to implement the method in any possible implementation of the first or second aspect described above.
[0043] Fourthly, a computer-readable storage medium is provided having a computer program stored thereon, characterized in that the program is executed by a processor to implement the vehicle data processing method described above.
[0044] Fifthly, a computer program product is provided, comprising a computer program, characterized in that, when the computer program is executed, it is used to implement the vehicle data processing method described above. Attached Figure Description
[0045] Figure 1 A flowchart illustrating the method for processing vehicle data provided in this application embodiment; Figure 2 This is a schematic diagram of the structure of a vehicle data dynamic hierarchical processing system based on multimodal semantics according to an embodiment of this application; Figure 3 This is a schematic diagram illustrating data hierarchy according to one embodiment of this application; Figure 4 This is a schematic diagram of the vehicle structure provided in an embodiment of this application. Detailed Implementation
[0046] The technical solutions in this application will be clearly and thoroughly described below with reference to the accompanying drawings. In the description of the embodiments of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B. "And / or" in the text is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Furthermore, in the description of the embodiments of this application, "multiple" refers to two or more than two.
[0047] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as implying or suggesting relative importance or implicitly indicating the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature.
[0048] With the rapid development of intelligent connected vehicle technology, the amount of data generated by a single vehicle as a mobile data collection terminal is growing exponentially, including but not limited to: sensor data (such as millimeter-wave radar, cameras, accelerometers, GNSS positioning); control bus data (such as CAN / LIN protocol messages); user behavior logs (such as voice commands, touch operations, navigation routes); system event logs (such as ADAS warnings, AEB triggers, battery abnormalities, communication disconnections), etc.
[0049] Limited by the computing power and network bandwidth of the vehicle communication module (T-Box), how to classify massive amounts of data in real time and ensure that critical data (such as accident evidence and emergency faults) are uploaded first has become the core pain point of vehicle network data transmission.
[0050] With the development of large-scale models, some technical solutions are attempting to introduce AI models to intelligently classify data, thereby improving the accuracy of classification. For example, OCR technology is used to recognize text content in screenshots of in-vehicle UIs, and tags are generated in conjunction with large-scale models for data deduplication and compression; neural networks are used to identify data service types (such as ADAS, media, and logs), and upload channels are selected accordingly.
[0051] However, these AI model-based solutions suffer from the following technical bottlenecks: Poor real-time performance: Most related technical solutions focus on uploading data to the cloud for offline annotation or simulation. This approach cannot meet the real-time requirements of the vehicle and does not solve the problem of efficient inference under the limited computing power of the vehicle. Lack of context awareness: The AI model does not fully integrate contextual features such as vehicle operating status (e.g., vehicle speed, acceleration, steering angle, braking status) and environmental information (e.g., weather, road type, traffic density), resulting in classification results that do not match the actual risk level; Poor platform adaptability: Most of the relevant AI models are deployed in the cloud or trained offline, without considering the needs of lightweight, low latency and multi-platform compatibility of the vehicle end, making it difficult to deploy quickly on different vehicle models and different hardware platforms.
[0052] In summary, current data classification technologies suffer from the following common defects: they are dominated by static rules and lack dynamic adjustment capabilities; AI classification results are isolated and not linked to core processes such as data caching and upload scheduling; they have weak context awareness capabilities, resulting in classification results that are out of sync with actual risk levels; and they have poor platform adaptability, making it difficult to deploy them uniformly in multi-vehicle and multi-platform environments.
[0053] Therefore, there is an urgent need for a new data classification method that can dynamically sense vehicle operating status and environmental changes, intelligently generate data level labels, and directly drive data caching and uploading strategies to solve the problems of slow response, disconnected strategies, and poor adaptability in existing technologies.
[0054] Based on this, this application provides a method for processing vehicle data, which can intelligently and dynamically classify vehicle data based on AI: by introducing a lightweight AI semantic reasoning engine into the vehicle, a closed-loop mechanism of "perception-alignment-evaluation-classification" is constructed, realizing the leap from "preset rule matching" to "real-time intent understanding" in data classification, thereby improving the timeliness of uploading high-priority data, system resource utilization, and multi-platform adaptability. This solves the problems in current vehicle communication caused by data classification relying on static rules, lack of context awareness, and the disconnect between AI classification results and caching / upload strategies, resulting in key data loss, upload delays, and resource waste.
[0055] The methods for processing the vehicle data include, but are not limited to, the following: Multimodal semantic alignment: Mapping heterogeneous visual, radar, and CAN bus data to a unified semantic space to achieve cross-validation and complementarity of features; Dynamic Risk Entropy Calculation: Innovatively introduces the quantitative indicator of "Risk Entropy" and combines semantic severity and spatiotemporal urgency to calculate the value density of data in real time; Fine-grained dynamic transitions: Instantaneous transitions in data level are triggered based on the rate of change of entropy (such as jumping directly from L1 to L3 lifeline), ensuring that core data is not missed.
[0056] The following description, in conjunction with the accompanying drawings, illustrates an embodiment of the vehicle data processing method of this application.
[0057] Figure 1 This is a flowchart illustrating a method for processing vehicle data according to an embodiment of this application.
[0058] like Figure 1 As shown, the vehicle data processing method of this application embodiment includes the following steps: Step S101: Extract multiple data features based on the vehicle's multimodal data at the current moment, and project the multiple data features onto the target semantic space to obtain spatiotemporally aligned fusion features.
[0059] In some real-time examples, this application can collect multimodal data of the vehicle in real time and extract multiple types of data features based on the multimodal data. Here, multimodal data refers to a set of data that can describe the vehicle's state, collected from various types of sensors or information sources on the vehicle (e.g., different physical principles, different data structures).
[0060] Furthermore, in order to unify the dimensions of multiple data features, embodiments of this application can project multiple types of data features onto a target semantic space, thereby obtaining a fused feature obtained by fusing multiple types of data features after cross-modal spatiotemporal alignment. Here, the target semantic space can be understood as a single high-dimensional semantic space.
[0061] The embodiments of this application can collect multimodal data of vehicles and extract multiple types of data features, and then project the multiple types of data features onto a unified high-dimensional semantic space to ensure the dimensionality of the multiple types of data features.
[0062] In some embodiments of this application, before extracting multiple data features based on the vehicle's multimodal data at the current moment, the method further includes: acquiring the vehicle's visual data, radar data, and state data at the current moment; and determining the multimodal data based on the visual data, radar data, and state data.
[0063] In some embodiments, this application needs to collect the vehicle's multimodal data before extracting multiple data features based on the vehicle's current multimodal data.
[0064] The multimodal data of the vehicle obtained in this embodiment includes, but is not limited to, the vehicle's visual data, radar data, and status data at the current moment.
[0065] For example, Figure 2 This is a schematic diagram of the structure of a vehicle data dynamic hierarchical processing system based on multimodal semantics, according to an embodiment of this application. Figure 2 As shown, the system can be deployed on an in-vehicle computing platform (such as EDC) and includes the following modules: a multimodal data acquisition module for receiving data such as camera images, radar point clouds, and CAN bus signals (vehicle speed, acceleration, steering angle); a cross-modal semantic alignment and fusion engine, integrated on the vehicle side, including a visual encoder, a point cloud processing network, and a temporal convolutional network (TCN); and a dynamic hierarchical decision module for calculating risk entropy values and mapping them to data level labels.
[0066] The multimodal data acquisition module can receive vehicle data such as camera frame images, LiDAR point cloud data, CAN bus signals (vehicle speed, acceleration, steering angle), and changes in tone of voice alarm words and geographic information through multiple types of input channels to complete the collection and processing of multimodal data.
[0067] For example, embodiments of this application can acquire three types of core data streams of a vehicle in real time: Visual data ( ): High-definition RGB image frames from the surround-view camera.
[0068] Radar data ( ): Point cloud / target list from lidar or millimeter-wave radar.
[0069] Vehicle status data ( ): Timing signals from the CAN bus, such as vehicle speed, acceleration, steering angle, and brake pedal opening.
[0070] The embodiments of this application can collect data such as vehicle camera images, radar point clouds, and CAN bus signals (vehicle speed, acceleration, steering angle) in real time, ensuring that the collected data can fully cover the data generation range of the vehicle and provide a reliable data foundation for subsequent feature extraction.
[0071] In some embodiments of this application, multiple types of data features are extracted based on multimodal data, including: extracting image features of the vehicle at the current moment based on visual data; extracting spatial geometric features of the vehicle at the current moment based on radar data; extracting dynamic features of the vehicle at the current moment based on state data; and determining multiple types of data features based on image features, spatial geometric features, and dynamic features.
[0072] In actual implementation, corresponding to the collected multimodal data, the various data features extracted by this application based on the multimodal data include, but are not limited to, the image features, spatial geometric features, and dynamic features of the vehicle at the current moment.
[0073] For example, such as Figure 2 As shown, this application can, but is not limited to, extract the following data features using a lightweight neural network autonomously deployed on the vehicle side: Using visual data as input, a visual encoder is used to extract image features. These image features exist in the form of vectors in practical applications, hereinafter referred to as image feature vectors. .
[0074] Using radar data as input, a point cloud processing network is used to extract spatial geometric features. These features exist in the form of vectors in practical applications, and are referred to below as spatial geometric feature vectors. .
[0075] Using state data as input, a temporal convolutional network (TCN) is used to extract vehicle dynamic features. These dynamic features exist in the form of vectors in practical applications, hereinafter referred to as dynamic feature vectors. .
[0076] It should be noted that in practical applications, lightweight neural networks such as visual encoders, point cloud processing networks, and temporal convolutional networks can be replaced with other lightweight neural networks. These can be set by those skilled in the art according to actual needs or determined according to the neural networks actually deployed in the vehicle, as long as they can be deployed in the vehicle and perform the relevant feature extraction functions. The embodiments in this application are only illustrative examples and do not impose specific limitations.
[0077] In some embodiments of this application, multiple types of data features are projected onto a target semantic space to obtain spatiotemporally aligned fusion features, including: projecting image features, spatial geometric features, and dynamic features onto the target semantic space based on a pre-built shared semantic projection layer to generate visual semantic tags corresponding to image features, spatial geometric semantic tags corresponding to spatial geometric features, and motion state semantic tags corresponding to dynamic features; calculating the spatiotemporal consistency score among the visual semantic tags, spatial geometric semantic tags, and motion state semantic tags to determine the fusion weights corresponding to the image features, spatial geometric features, and dynamic features based on the spatiotemporal consistency score; and fusing the visual semantic tags, spatial geometric semantic tags, and motion state semantic tags using the fusion weights to generate spatiotemporally aligned fusion features.
[0078] It is understandable that the dimensions of data collected from different types of sensors or information sources are not uniform, and the corresponding extracted data features are also not uniform in dimension. Therefore, the embodiments of this application can project multiple types of data features into a unified high-dimensional semantic space and further solve the false alarm problem that may be generated by a single modality.
[0079] For example, this application can pre-construct a shared semantic projection layer, and then, based on this pre-constructed shared semantic projection layer, project the image feature vector... Spatial geometric eigenvectors , dynamic eigenvectors These are mapped to the same high-dimensional semantic tokens, which can be referred to here as visual semantic tokens (corresponding to image features), spatial geometric semantic tokens (corresponding to spatial geometric features), and motion state semantic tokens (corresponding to dynamic features), respectively.
[0080] To avoid false alarms from a single modality (such as misjudging obstacles due to visual interference from backlight), this application introduces a cross-attention mechanism to calculate the spatiotemporal consistency score between different modalities. Based on the spatiotemporal consistency score, the fusion weight of each modality is determined, and the semantic tokens of each modality are weighted and summed with the corresponding dynamic fusion weights to obtain the spatiotemporally aligned fusion features.
[0081] Specifically, embodiments of this application can introduce a cross-attention mechanism to calculate the spatiotemporal consistency score among visual semantic tags, spatial geometric semantic tags, and motion state semantic tags. Based on this spatiotemporal consistency score, embodiments of this application can determine the fusion weights that each type of data feature should have when fusing image features, spatial geometric features, and dynamic features. That is, the fusion weights that visual semantic tags, spatial geometric semantic tags, and motion state semantic tags should each have, so as to fuse multiple types of features according to the fusion weights to obtain spatiotemporally aligned fused features.
[0082] In some embodiments of this application, calculating the spatiotemporal consistency score among visual semantic tags, spatial geometric semantic tags, and motion state semantic tags includes: obtaining timestamp information and spatial coordinate information from the visual semantic tags, spatial geometric semantic tags, and motion state semantic tags; calculating the time difference and spatial distance between any two of the visual semantic tags, spatial geometric semantic tags, and motion state semantic tags based on the timestamp information and spatial coordinate information; and calculating the spatiotemporal consistency score based on the spatial distance and time difference.
[0083] In actual implementation, this application can, but is not limited to, use the timestamps and spatial coordinate information of each semantic tag to score the spatiotemporal consistency among computational visual semantic tags, spatial geometric semantic tags, and motion state semantic tags.
[0084] For example, by using the timestamps and spatial coordinates of each semantic tag, this application can calculate the time difference and spatial distance between each semantic tag, thereby identifying whether any semantic tag has deviated in time or space, and then calculating the spatiotemporal consistency score.
[0085] For example, if vision identifies an "obstacle ahead," radar also detects an echo in the same spatiotemporal coordinates, and the CAN signal indicates "rapid deceleration," then the three are semantically highly aligned, significantly increasing confidence and improving the spatiotemporal consistency score. For example, calculating the spatiotemporal consistency score... If visual recognition indicates a "frontal collision," radar detects a "sudden decrease in distance," and the CAN signal indicates "rapid deceleration," then... ≈1 indicates an extremely high confidence level.
[0086] Conversely, if only the visual modality responds while other modalities do not (such as radar and CAN), it can be identified as visual noise. The weight of the visual modality in the spatiotemporal consistency score calculation will be reduced, and the spatiotemporal consistency score will also decrease accordingly.
[0087] Among them, the spatiotemporal consistency score (Confidence level) can be intrinsically calculated by the lightweight multimodal large model on the vehicle and the perception algorithms of each sensor during real-time inference.
[0088] For example, this application can first binarize or normalize the input information from the three sensors, namely visual semantic tags, spatial geometric semantic tags, and motion state semantic tags, and denote them as follows: , and To facilitate calculation. Where: This represents the confidence level of the visual algorithm regarding "obstacles ahead," with a value range of [0, 1], where 0 indicates no obstacle and 1 indicates certainty of an obstacle. This represents the radar echo intensity or probability of existence in the corresponding spatiotemporal coordinates, and its value ranges from [0, 1], where 0 indicates no echo and 1 indicates an echo. This indicates the degree to which the vehicle's dynamic state is confirmed by the obstacle (such as rapid deceleration), and its value ranges from [0, 1], where 0 represents constant speed / acceleration and 1 represents rapid deceleration / braking.
[0089] Dynamic weights The calculation formula can be expressed, but is not limited to, as follows:
[0090] .
[0091] in, The values represent the basic visual weights (e.g., 0.3), 'a' represents the gain coefficient for radar verification (e.g., 0.4), and 'b' represents the gain coefficient for CAN bus verification (e.g., 0.3). Specifically... The coefficients of a, b can be set and adjusted by those skilled in the art according to the actual situation. The embodiments in this application are only illustrative and do not impose specific limitations.
[0092] For example: If neither radar nor CAN responds (i.e.) =0, =0), then Maintaining a low baseline value means that the system considers the visual signal to be noise at this point, and even if the visual system sees something, the impact on the total score is limited.
[0093] If both radar and CAN respond ( =1, =1), then Reaching the maximum value means that the multimodal is highly aligned and the visual signal is identified as a "high-confidence target".
[0094] Final rating The alignment is determined by both the weighted visual and radar signals. Since the CAN signal is mainly used to adjust the weights and is not directly used as the input for spatial position, the alignment between the visual and radar signals can be primarily calculated here.
[0095] The final spatiotemporal consistency scoring formula can be, but is not limited to, expressed as:
[0096] To illustrate the logic more intuitively, the embodiments of this application may, but are not limited to, a scene determination table. Table 1 is a scene determination table of one embodiment of this application, which can be represented as follows:
[0097] It should be noted that the specific calculation model and algorithm used in calculating the spatiotemporal consistency score can be determined by those skilled in the art or by the actual vehicle model. The embodiments in this application are only illustrative and do not impose any specific limitations.
[0098] The following example illustrates the cross-modal semantic alignment and fusion process in this application: When a vehicle is traveling on a highway in a twilight environment with backlighting, the backlight can cause the visual images captured by the onboard camera to be overexposed or have localized brightness saturation. In this situation, the visual perception model may misidentify dark shadow areas or road seams ahead as "crossing obstacles," and this will be reflected in the visual feature vector. A corresponding visual semantic token is generated, which carries the semantic information that "there is an obstacle ahead". At this time, if the judgment is made solely based on the visual modality, it is very easy to trigger unnecessary emergency braking, which seriously affects driving safety and driving experience.
[0099] To avoid the aforementioned single-modal false alarms, embodiments of this application can use visual feature vectors millimeter-wave radar point cloud feature vectors and CAN bus dynamic eigenvectors The inputs are fed into a pre-built shared semantic projection layer. This shared semantic projection layer consists of multiple fully connected networks that can map three feature vectors from different sources and with different original dimensions to a unified 256-dimensional high-dimensional semantic space, generating visual semantic tokens, spatial geometric semantic tokens, and dynamic semantic tokens, respectively.
[0100] Subsequently, embodiments of this application introduce a cross-attention mechanism to perform spatiotemporal consistency calculations on visual semantic tokens, spatial geometric semantic tokens, and dynamic semantic tokens. For example, the cross-attention mechanism can use the visual semantic token as the query vector and the spatial geometric semantic token and dynamic semantic token as the key vector and value vector, respectively, to calculate the feature relevance of any two modality tokens within the same spatial mapping region at the same time, thereby outputting a spatiotemporal consistency score. This score is used to characterize the alignment and mutual corroboration strength between multimodal information at the current moment, reflecting the confidence level of multimodal data.
[0101] In the current scenario, when the visual perception model misjudges an "obstacle ahead" due to overexposure, the cross-attention mechanism simultaneously retrieves radar echo information and CAN signals located in the same spatiotemporal coordinates as the misjudged target. If the millimeter-wave radar does not detect any valid reflection point cloud in that spatiotemporal coordinate, and the CAN bus data shows that the current accelerator pedal opening is stable, the vehicle speed has not changed abruptly, and the braking system has not been triggered, it indicates that neither the radar mode nor the dynamics mode has generated any positive response to the visually recognized "obstacle ahead." At this time, the correlation between the cross-modal token pairs calculated by the cross-attention mechanism is extremely low, which can be used to score the spatiotemporal consistency. Set to a low value close to 0.
[0102] Based on spatiotemporal consistency score After determining that the output of the current visual modality is perceptual noise caused by ambient light interference, this embodiment of the application can dynamically adjust the fusion weights of each modality according to the determination result, and fuse visual semantic tokens, spatial geometric semantic tokens, and dynamic semantic tokens based on the adjusted weights to generate a comprehensive semantic fusion feature. This fusion feature still exists in the form of a vector. Since this fusion feature vector suppresses the interference components caused by visual false alarms, unnecessary emergency braking caused by single visual backlight interference can be avoided.
[0103] Step S102: Calculate the first risk entropy value of the vehicle at the current moment based on the fusion features, and calculate the risk mutation gain of the vehicle at the current moment based on the first risk entropy value.
[0104] In some embodiments, based on the spatiotemporally aligned fusion features, this application can determine the risk entropy value E of the current data segment, which is used to quantify the importance of the data.
[0105] Specifically, in this embodiment, the vehicle's first risk entropy value and risk mutation gain at the current moment can be calculated based on the fused features. Here, the first risk entropy value can be understood as the base risk entropy value reflecting the vehicle's static risk level at the current instant. The risk mutation gain can be understood as the rate of change of risk increment between the vehicle's base risk entropy value at the current moment and the final risk entropy value at the previous moment.
[0106] For example, this application can first calculate the basic risk entropy value at the current time within each sampling period t. .
[0107] Specifically, in this application embodiment, the basic risk entropy value can be calculated by weighted fusion algorithm based on the fusion features obtained from the multimodal data collected at the current time t. This value reflects the static risk level of the vehicle at the current instant, without considering historical change trends.
[0108] Among them, the basic risk entropy value The calculation formula can be expressed, but is not limited to, as follows:
[0109] in, This represents the base risk entropy value at time t, reflecting the weighted sum of the static risk levels of vehicles at the current time. Its value typically ranges from [value missing]. The larger the value, the more important the data and the higher the risk. It is a semantic severity factor; It is a physical risk probability factor. and The weights for the semantic severity factor and the physical risk probability factor are respectively calculated. The specific weights can be set and adjusted by those skilled in the art according to the actual situation. The embodiments in this application are only illustrative and do not impose specific limitations.
[0110] In some embodiments of this application, calculating the first risk entropy value of the vehicle at the current moment based on the fusion features includes: obtaining the vehicle's motion state and environmental information at the current moment; identifying the actual driving scenario of the vehicle at the current moment based on the motion state and environmental information; generating the calculation weight of the scenario classification score of the vehicle at the current moment based on the actual driving scenario, so as to calculate the first risk entropy value based on the scenario classification score and its calculation weight.
[0111] Based on the descriptions of other embodiments, it will be understood that a semantic severity factor is introduced when calculating the basic risk entropy value (first risk entropy value). .
[0112] In actual implementation, this application determines the semantic severity factor. In this case, the scene classification score output by the multimodal large model may be used as the semantic severity factor, but is not limited to. .
[0113] When calculating the scene classification score, this application can determine the vehicle's motion state and environmental information at the current moment based on the fusion features, so as to identify the actual driving scene in which the vehicle is located at the current moment based on the motion state and environmental information.
[0114] Based on the actual driving scenario, this application can generate a calculation weight for the scenario classification score of the overall situation of the actual driving scenario, and then calculate the first risk entropy value based on the scenario classification score and its calculation weight.
[0115] For example, this application can first obtain a scene classification score based on the fusion features. For instance, the fusion features (i.e., the multimodal feature alignment result) can be input into a lightweight neural network on the vehicle side to directly output a scene classification score used to quantify the scene situation of the vehicle. Table 2 is a description of the relationship between the semantic severity factor (scene classification score) and the scene in one embodiment of this application. Table 2 can be represented as follows:
[0116] For example, when the actual driving scenario is "normal driving", In actual driving scenarios where pedestrians are crossing, The actual driving scenario is a collision accident. And so on. The specific actual driving scenarios and their corresponding scenario classification scores can be set and adjusted by professionals in this field according to actual conditions and needs. The embodiments in this application are only illustrative and do not impose specific limitations.
[0117] It should be noted that, The value is not an absolutely discrete value, but is calculated by weighting the probability confidence scores output by the multimodal large model. For example, when the model identifies 'pedestrian crossing' with 80% confidence (baseline score 0.5) and 'road shadow' with 20% confidence (baseline score 0.1), the final output is... This is a weighted sum of the two. This continuous numerical output can effectively avoid data oscillations caused by abrupt changes in scene classification.
[0118] Then, this embodiment of the application can identify the actual driving scenario of the vehicle at the current moment based on the vehicle's current motion state and environmental information, and output the weight of the scenario classification score corresponding to the actual driving scenario. To calculate the first risk entropy value.
[0119] In some embodiments of this application, calculating a first risk entropy value based on a scene classification score includes: obtaining the relative distance and relative speed between the vehicle and a target vehicle located in front of the vehicle; calculating the objective collision risk of the vehicle at the current moment based on the relative distance and relative speed; and calculating the first risk entropy value based on the scene classification score and its calculation weight, and the objective collision risk and its calculation weight.
[0120] Based on the descriptions of other embodiments, it will be understood that this application also introduces a physical risk probability factor when calculating the basic risk entropy value (first risk entropy value). .
[0121] In actual implementation, this application determines the physical risk probability factor. At times, but not limited to, the relative distance between the vehicle and targets ahead can be obtained in real time using vehicle-mounted radar and visual perception modules. and relative velocity The calculated objective collision risk quantification value serves as a physical risk probability factor. .
[0122] Specifically, this application can obtain the relative distance and relative speed between a vehicle and a target vehicle located in front of it, in order to calculate the objective collision risk between the vehicle and the target vehicle in front of it, and then combine the scene classification score and the quantitative value corresponding to the objective collision risk to calculate the first risk entropy value.
[0123] In this context, the target vehicle located in front of the vehicle can be understood as the vehicle ahead of it that is on the same road as the vehicle, and is generally the closest vehicle.
[0124] For example, the physical risk probability factor in this application It can be expressed as follows, but is not limited to: , in, Time to Collision (TTO) is the quantified value corresponding to the objective collision risk. When approaching 0, Approaching 1.
[0125] in, The calculation follows the following physical correlation formula:
[0126] Where d represents the real-time relative distance between the vehicle and the target ahead (unit: meters, m), and v_r represents the real-time relative speed between the vehicle and the target ahead (unit: meters per second, m / s), which can usually be taken as the difference between the two speeds.
[0127] Finally, by combining the scene classification score and its calculation weight, as well as the objective collision risk and its calculation weight, the first risk entropy value can be calculated in this embodiment of the application.
[0128] In some embodiments of this application, calculating the risk mutation gain of the vehicle at the current moment based on the first risk entropy value includes: obtaining the final risk entropy value of the vehicle at a historical moment adjacent to the current moment; calculating the incremental change rate between the first risk entropy value and the final risk entropy value at the historical moment; and determining the risk mutation gain based on the incremental change rate.
[0129] Based on the descriptions of other embodiments, it will be understood that the risk mutation gain can be understood here as the rate of change of the risk increment of the vehicle at the current moment compared with the final risk entropy value at the previous moment.
[0130] Based on this, in some embodiments, when calculating the risk mutation gain of the vehicle at the current moment based on the first risk entropy value (basic risk entropy value), this application may first obtain the final risk entropy value of the vehicle at a historical moment adjacent to the current moment, and then use the incremental change rate between the calculated basic risk entropy value at the current moment and the final risk entropy value at that historical moment as the risk mutation gain.
[0131] Here, the historical moment adjacent to the current moment can be understood as the moment before the current moment, with the current moment as the reference.
[0132] Specifically, to capture the instantaneous nature of dangerous events, this application introduces a one-way mutation compensation mechanism. That is, this application can read the final risk entropy value from the previous moment. Calculate the current base value The rate of change of the incremental risk entropy relative to historical values.
[0133] Among them, risk mutation gain The calculation formula can be expressed, but is not limited to, as follows:
[0134] in, It is the system's data sampling time interval. (Risk Sudden Change Gain): Characterizes the instantaneous suddenness and intensity of risk, defined as the rate of change of the current entropy value relative to the historical average entropy value. Used to capture sudden events (such as a sudden tire blowout), even if the semantics are not fully resolved, the huge rate of change can still increase the entropy value.
[0135] The principle behind this formula is as follows: When danger approaches, the rising edge strengthens: when This indicates that the level of danger is rapidly increasing (such as a sudden tire blowout or a sharp turn of the steering wheel). The system will generate a positive "mutation gain", which will be added to the final result to achieve early warning.
[0136] When the danger is mitigated, the falling edge remains: when This indicates that the level of danger is decreasing or remaining stable. At this time, a mandatory order is issued. .
[0137] Limiting non-negative numbers: This is to prevent "negative compensation" from causing the system to prematurely lower the alert level. For example, although the acceleration drops instantly after a vehicle collision, the accident scene still needs to maintain high-level data uploading. By truncating the negative gain to 0, the system implements level hysteresis, effectively avoiding frequent oscillations of data between certain adjacent levels (such as L2 / L3).
[0138] Step S103: Based on the first risk entropy value and the risk mutation gain, calculate the second risk entropy value at the current moment, determine the transmission level of the multimodal data according to the second risk entropy value, and transmit the multimodal data to the vehicle according to the transmission level.
[0139] As one possible approach, after obtaining the basic risk entropy value and the risk mutation gain, this application can combine the two to calculate the vehicle's second risk entropy value at the current moment. Here, the second risk entropy value can be understood as the vehicle's final risk entropy value at the current moment, reflecting the vehicle's overall risk at that moment and ultimately determining the entropy value for the multimodal data level.
[0140] Specifically, such as Figure 2 As shown in the dynamic risk entropy calculation module, this application can sequentially calculate the basic risk entropy value at the current moment within each sampling period t. Risk mutation gain and final risk entropy value .
[0141] Among them, the final risk entropy value A limiting mechanism can be introduced to superimpose the mutation gain onto the base risk entropy value and output the final risk entropy value for final decision-making.
[0142] The formula for calculating the final risk entropy value can be, but is not limited to, expressed as follows:
[0143] in, The term refers to the basic risk entropy value. ; express The final risk entropy value at time 1, plus the mutation gain, is used to trigger data classification. Its value range is typically [value range missing]. The larger the value, the more important the data and the higher the risk. , , All of these are configurable weighting coefficients, dynamically adjusted variables based on the vehicle's operating scenario. This determines the system's sensitivity to changes in the visual / semantic environment. This determines the system's sensitivity to sudden changes in the vehicle's own motion state. The adjustment system amplifies the impact of "sudden dangers" by a factor; the higher the value, the more aggressive the system. Furthermore... .
[0144] Based on the vehicle's final risk entropy value at the current moment The embodiments of this application can quantify the importance of multimodal data at the current moment, and thus determine the multimodal data transmission level at the current moment.
[0145] For example, Figure 3 This is a schematic diagram illustrating data hierarchy according to one embodiment of this application. Figure 3 As shown, this application can calculate Based on a pre-defined dynamic grading strategy, the final data grade label is output, which represents the transmission grade of the multimodal data at the current moment. L1 Standard Level: When At this time. Corresponding scenarios: highway cruising, stationary parking. At this time, the multimodal data of the current moment can be flowed to the vehicle's local storage or uploaded to the vehicle sequentially.
[0146] L2 core level: When In situations where lane lines are blurred, following too closely, or inclement weather, the multimodal data for the current moment needs to be compressed and periodically uploaded to the vehicle; the upload frequency could be appropriately increased.
[0147] L3 Lifeline Level: When At this time, the corresponding scenarios are: AEB triggering, airbag deployment, and severe collision. In this case, the current multimodal data must be preserved intact and immediately transmitted to the vehicle with the highest priority.
[0148] in, and It is a pre-set entropy value classification threshold, for example, Set it to 0.3. The threshold is set to 0.7, etc. The specific grading threshold can be set and adjusted by professionals in this field according to the actual situation and configuration information. The embodiments in this application are only illustrative and do not impose specific limitations.
[0149] In some embodiments of this application, the calculation of the second risk entropy value at the current moment based on the first risk entropy value and the risk mutation gain includes: determining the calculation weight corresponding to the risk mutation gain and the calculation weight corresponding to the scene classification score and the objective collision risk in the first risk entropy value based on the actual driving scenario of the vehicle at the current moment; and calculating the second risk entropy value based on the risk mutation gain, the scene classification score, the objective collision risk and their corresponding calculation weights.
[0150] Based on the descriptions of other embodiments, it is understood that this application can calculate a first risk entropy value based on the scene classification score and its calculation weight and the objective collision risk and its calculation weight. The final risk entropy value is the weighted sum of the first risk entropy value and the risk mutation gain.
[0151] As one possible approach, this application calculates the final risk entropy value using a formula that includes risk mutation gain and scenario classification score (corresponding to the semantic severity factor). ) and objective collision risk (corresponding to the physical risk probability factor) The corresponding calculation weights can be dynamically adjusted according to the actual driving scenario of the vehicle, but are not limited to.
[0152] Specifically, embodiments of this application may introduce scene context (…). The perception logic adjusts various factors (risk mutation gain) in real time based on the vehicle's current motion state (such as speed and acceleration) and geographical environment (corresponding to the preceding environmental information, such as highway, urban area, etc.). semantic severity factor Physical risk probability factor The weighting of risk mutation gain, scene classification score, and objective collision risk, along with their corresponding calculation weights, is determined by assigning weights to these factors. The weighting adjustment strategy can, but is not limited to, the following: In high-speed driving scenarios, the focus is on physical risks: Scene determination: When the vehicle speed Furthermore, when the lane markings are clearly visible, the system identifies it as a "high-speed driving scenario".
[0153] Weighting strategy: In this scenario, the vehicle has high kinetic energy, and a collision would have severe consequences. Therefore, the system automatically increases the weight of the physical risk probability factor. (For example, adjust to 0.6-0.7), while relatively reducing the semantic severity weight. .
[0154] Therefore, the embodiments of this application can make the system more sensitive to physical indicators such as "relative velocity" and "time of collision (TTC)". For example, even if the visual semantic recognition has not fully confirmed the type of obstacle ahead (low semantic confidence), as long as the radar detects an extremely short TTC, the high-weighted This item can quickly increase the overall entropy value. This triggered an L3 level alarm.
[0155] In urban congestion / complex interaction scenarios, the focus is on semantic understanding: Scene determination: When the vehicle speed Furthermore, when traffic lights / pedestrian detection occur frequently, the system identifies it as an "urban congestion / complex interaction scenario".
[0156] Weighting Adjustment Strategy: In this scenario, vehicles face complex semantic events such as "ghost peeks out" and pedestrians crossing the road. The physical distance may be short but not pose a substantial danger (e.g., following another vehicle). Therefore, the system automatically increases the weight of the semantic severity factor. (For example, adjust to 0.5-0.6), and combine with the risk mutation gain weight. .
[0157] Therefore, the embodiments of this application allow the system to focus more on "scene understanding". When the visual model recognizes "a pedestrian suddenly rushing out" or "a vehicle running a red light", the higher weighting of the visual model will be used to... It will respond quickly, and even if the physical distance is still far (TTC is large), it can anticipate risks and improve the data level in advance.
[0158] Low-speed parking / stationary scenarios, focusing on sudden changes: Scene determination: When the vehicle speed When the vehicle is in park (P gear), the system identifies it as a "low-speed parking / stationary scenario".
[0159] Weighting strategy: In this scenario, physical and semantic risks are typically low (unless a collision occurs). The system automatically increases the weight of the risk mutation gain factor. And set a low entropy value to trigger the threshold. This is the mutation compensation coefficient, used to adjust the system's sensitivity to sudden risks.
[0160] Therefore, embodiments of this application can detect sudden scrapes or collisions. Because... The (rate of change) has a high weight. Once a sudden and violent shock (acceleration mutation) occurs, the entropy value will soar in a very short time, triggering a level jump.
[0161] Example of a dynamically adjusted formula:
[0162] in, It is determined by the vehicle's motion state (vehicle speed, acceleration) and the results of environmental perception (map type, traffic density).
[0163] Figure 4 A schematic diagram of the structure of a vehicle provided in an embodiment of this application. The vehicle may include: The memory 401, the processor 402, and the computer program stored on the memory 401 and capable of running on the processor 402.
[0164] When the processor 402 executes the program, it implements the vehicle data processing method provided in the above embodiments.
[0165] Furthermore, the vehicle also includes: Communication interface 403 is used for communication between memory 401 and processor 402.
[0166] The memory 401 is used to store computer programs that can run on the processor 402.
[0167] The memory 401 may include high-speed RAM (Random Access Memory) memory, and may also include non-volatile memory, such as at least one disk storage.
[0168] If the memory 401, processor 402, and communication interface 403 are implemented independently, then the communication interface 403, memory 401, and processor 402 can be interconnected via a bus to complete communication between them. The bus can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 4 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0169] Optionally, in a specific implementation, if the memory 401, processor 402, and communication interface 403 are integrated on a single chip, then the memory 401, processor 402, and communication interface 403 can communicate with each other through an internal interface.
[0170] Processor 402 may be a CPU (Central Processing Unit), an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of this application.
[0171] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0172] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0173] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of this application.
Claims
1. A method for processing vehicle-mounted data, characterized in that, Includes the following steps: Based on the vehicle's multimodal data at the current moment, multiple data features are extracted, and these multiple data features are projected onto the target semantic space to obtain spatiotemporally aligned fusion features. The first risk entropy value of the vehicle at the current moment is calculated based on the fusion features, and the risk mutation gain of the vehicle at the current moment is calculated based on the first risk entropy value. Based on the first risk entropy value and the risk mutation gain, a second risk entropy value at the current moment is calculated to determine the transmission level of the multimodal data according to the second risk entropy value, and the multimodal data is transmitted to the vehicle according to the transmission level.
2. The method according to claim 1, characterized in that, Before extracting multiple types of data features based on the vehicle's multimodal data at the current moment, the method further includes: Acquire the vehicle's visual data, radar data, and status data at the current moment; The multimodal data is determined based on the visual data, the radar data, and the state data.
3. The method according to claim 2, characterized in that, The extraction of multiple data features based on the vehicle's multimodal data at the current moment includes: Based on the visual data, extract the image features of the vehicle at the current moment; Based on the radar data, the spatial geometric features of the vehicle at the current moment are extracted; Based on the state data, the dynamic characteristics of the vehicle at the current moment are extracted; Multiple types of data features are determined based on the image features, the spatial geometric features, and the dynamic features.
4. The method according to claim 3, characterized in that, The step of projecting multiple types of data features onto the target semantic space to obtain spatiotemporally aligned fusion features includes: Based on a pre-built shared semantic projection layer, the image features, spatial geometric features, and dynamic features are projected onto the target semantic space to generate visual semantic tags corresponding to the image features, spatial geometric semantic tags corresponding to the spatial geometric features, and motion state semantic tags corresponding to the dynamic features. Calculate the spatiotemporal consistency score among the visual semantic markers, the spatial geometric semantic markers, and the motion state semantic markers, and determine the fusion weights corresponding to the image features, the spatial geometric features, and the dynamic features based on the spatiotemporal consistency score; The spatiotemporal aligned fusion feature is generated by fusing the visual semantic markers, the spatial geometric semantic markers, and the motion state semantic markers using the fusion weights.
5. The method according to claim 4, characterized in that, The calculation of the spatiotemporal consistency score among the visual semantic markers, the spatial geometric semantic markers, and the motion state semantic markers includes: Obtain the timestamp information and spatial coordinate information from the visual semantic markers, the spatial geometric semantic markers, and the motion state semantic markers; Based on the timestamp information and the spatial coordinate information, calculate the time difference and spatial distance between any two of the visual semantic marker, the spatial geometric semantic marker and the motion state semantic marker; The spatiotemporal consistency score is calculated based on the spatial distance and the time difference.
6. The method according to claim 1, characterized in that, The step of calculating the first risk entropy value of the vehicle at the current moment based on the fusion features includes: Obtain the vehicle's current motion state and environmental information; Based on the motion state and the environmental information, identify the actual driving scenario of the vehicle at the current moment; The calculation weights for the scenario classification score of the vehicle at the current moment are generated based on the actual driving scenario, and the first risk entropy value is calculated based on the scenario classification score and its calculation weights.
7. The method according to claim 6, characterized in that, The calculation of the first risk entropy value based on the scenario classification score includes: Obtain the relative distance and relative speed between the vehicle and a target vehicle located in front of the vehicle; Calculate the objective collision risk of the vehicle at the current moment based on the relative distance and relative speed; The first risk entropy value is calculated based on the scenario classification score and its calculation weight, and the objective collision risk and its calculation weight.
8. The method according to claim 1, characterized in that, The step of calculating the risk mutation gain of the vehicle at the current moment based on the first risk entropy value includes: Obtain the final risk entropy value of the vehicle at a historical moment adjacent to the current moment; Calculate the rate of change of the increment between the first risk entropy value and the final risk entropy value at the historical moment; The risk mutation gain is determined based on the incremental rate of change.
9. The method according to claim 1, characterized in that, The step of calculating the second risk entropy value at the current moment based on the first risk entropy value and the risk mutation gain includes: Based on the actual driving scenario of the vehicle at the current moment, determine the calculation weight corresponding to the risk mutation gain and the calculation weight corresponding to the scenario classification score and objective collision risk in the first risk entropy value; The second risk entropy value is calculated based on the risk mutation gain, the scene classification score, the objective collision risk and its corresponding calculation weight.
10. A vehicle, characterized in that, The vehicle includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the processor executing the program to implement the vehicle data processing method as described in any one of claims 1-9.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by the processor to implement the method for processing vehicle data as described in any one of claims 1-9.
12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed, it is used to implement the vehicle data processing method as described in any one of claims 1-9.