Vehicle-based CAN data processing method and device, equipment and storage medium

By constructing a CAN data processing method based on a vector database, combined with a large language model and a virtual environment simulator, the problem of low efficiency in CAN data query in autonomous driving systems was solved, achieving fast and accurate data positioning and scene feature extraction, thereby improving processing efficiency and user experience.

CN121456014APending Publication Date: 2026-02-03CENNAVI TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511604242.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-04
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

In autonomous driving systems, the query and processing efficiency of massive CAN data is low, there is a lack of scene description, it is difficult to rely on manual screening, the feature extraction of complex driving scenarios relies on human experience and is costly, the data is disconnected from the context, and it is difficult to achieve cross-modal query.

Method used

The CAN data processing method based on a pre-set vector database is adopted. It receives user query information through human-computer interaction, uses a large language model and retrieval enhancement generation technology, and combines virtual environment simulator to generate simulated data and real data to construct a vector database, thereby realizing natural language interaction and accurate data query.

Benefits of technology

It improves the processing efficiency and accuracy of CAN data, reduces manual query time, enhances user experience, enables rapid and accurate data positioning and scene feature extraction, and reduces labor costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121456014A_ABST
    Figure CN121456014A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a CAN data processing method and device based on a vehicle, equipment and a storage medium. The method comprises the following steps: receiving data query information input by a user; wherein the data query information represents a query demand of a user for CAN data; determining target data based on a preset vector database according to the data query information; wherein the preset vector database comprises CAN data of a vehicle in a preset driving scene and a text vector corresponding to the CAN data, the text vector represents natural language description corresponding to the CAN data, and the target data represents CAN data to be queried by a user; generating and outputting reply information according to the target data; wherein the reply information is used for representing target data. Through man-machine interaction, the CAN data can be quickly queried, and the query efficiency and precision of the CAN data are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, and particularly relates to a CAN data processing method and device based on a vehicle, equipment and a storage medium. BACKGROUND

[0002] A vehicle can collect a large number of control signals in real time through a CAN bus, such as speed, acceleration, steering angle, brake signal, etc. In the development process of an automatic driving system, an engineer needs to quickly locate CAN data in a specific driving scene to optimize an automatic driving algorithm, diagnose a fault, or perform safety testing, etc.

[0003] Therefore, how to efficiently and accurately query a large amount of CAN data is a technical problem to be solved. SUMMARY

[0004] The embodiments of the present application provide a CAN data processing method and device based on a vehicle, equipment and a storage medium to improve the processing efficiency and accuracy of CAN data.

[0005] In a first aspect, the embodiments of the present application provide a CAN data processing method based on a vehicle, comprising:

[0006] receiving data query information input by a user; wherein the data query information represents a query requirement of the user for CAN data;

[0007] determining target data based on a preset vector database according to the data query information; wherein the preset vector database comprises CAN data of a vehicle in a preset driving scene and a text vector corresponding to the CAN data, the text vector represents a natural language description corresponding to the CAN data, and the target data represents CAN data to be queried by the user;

[0008] generating and outputting reply information according to the target data; wherein the reply information is used to represent the target data.

[0009] In a second aspect, the embodiments of the present application provide a CAN data processing device based on a vehicle, comprising:

[0010] an information receiving unit configured to receive data query information input by a user; wherein the data query information represents a query requirement of the user for CAN data;

[0011] The target determining unit is configured to determine target data according to the data query information and based on a preset vector database, wherein the preset vector database comprises CAN data of a vehicle in a preset driving scene and a text vector corresponding to the CAN data, the text vector represents a natural language description corresponding to the CAN data, and the target data represents the CAN data to be queried by the user.

[0012] The information output unit is configured to generate and output reply information according to the target data, wherein the reply information is used to represent the target data.

[0013] In a third aspect, an embodiment of the present application provides an electronic device, comprising a memory and a processor.

[0014] The memory stores computer execution instructions.

[0015] The processor executes the computer execution instructions stored in the memory, so that the processor executes the embodiments of the first aspect.

[0016] In a fourth aspect, an embodiment of the present application provides a computer readable storage medium, wherein the computer readable storage medium stores computer execution instructions, and the computer execution instructions are executed by a processor to implement the embodiments of the first aspect.

[0017] In a fifth aspect, an embodiment of the present application provides a computer program product, comprising a computer program, and the computer program is executed by a processor to implement the embodiments of the first aspect.

[0018] The embodiments of the present application provide a CAN data processing method and device based on a vehicle, an equipment and a storage medium. A user can quickly and accurately query CAN data of a vehicle through human-computer interaction. The user inputs data query information according to actual needs, and the data query information represents the query requirement of the user for the CAN data. The data query information is analyzed, and target data is determined according to a preset vector database. The preset vector database can comprise CAN data of a vehicle in different preset driving scenes and a text vector corresponding to the CAN data. The text vector is vector data of a natural language description corresponding to the CAN data, and the target data represents the CAN data to be queried by the user. That is, the CAN data to be queried by the user can be accurately obtained from the preset vector database. Reply information is generated according to the target data, and is fed back to the user based on human-computer interaction. The embodiments of the present application can directly locate a driving scene required by the user through the preset vector database, effectively reduce manual query time, improve the processing efficiency and accuracy of the CAN data through natural language human-computer interaction, and improve user experience. BRIEF DESCRIPTION OF DRAWINGS

[0019] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0020] Figure 1 A flowchart illustrating a vehicle-based CAN data processing method provided in an embodiment of this application;

[0021] Figure 2 A flowchart illustrating a vehicle-based CAN data processing method provided in an embodiment of this application;

[0022] Figure 3 A flowchart illustrating a vehicle-based CAN data processing method provided in an embodiment of this application;

[0023] Figure 4 A schematic diagram illustrating the CAN data query process provided in an embodiment of this application;

[0024] Figure 5 A schematic diagram of the structure of a vehicle-based CAN data processing device provided in an embodiment of this application;

[0025] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.

[0026] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation

[0027] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0028] First, let me explain the terms used in this application:

[0029] CAN data: CAN stands for Controller Area Network, a vehicle bus standard used for communication between electronic control units within a vehicle. The real-time data it transmits, such as speed and braking signals, is commonly used in autonomous driving functions.

[0030] LLM: LLM stands for Large Language Model. It is a deep learning model that is pre-trained on massive amounts of text data. It can understand and generate natural language and is often used in scenarios such as data description generation and feature summarization.

[0031] RAG: RAG stands for Retrieval-Augmented Generation. It is an AI framework that enhances the accuracy and reliability of large language models by retrieving relevant information from external knowledge bases. It is suitable for processing domain-specific data, such as vehicle scenario queries.

[0032] Bag: A file format used to store sensor and vehicle data, such as CAN signals and camera data. In autonomous driving, it facilitates the recording and playback of scene data to improve query efficiency.

[0033] This application relates to the technical fields of human-computer interaction, autonomous vehicle data management, and intelligent retrieval in artificial intelligence, specifically applied to the intelligent querying and scenario analysis of vehicle CAN bus data. During the development of autonomous driving systems, vehicles collect a large number of control signals in real time via the CAN bus, such as speed, acceleration, steering angle, and braking signals. This data is stored in unstructured or semi-structured form. Engineers need to quickly locate specific driving scenarios, such as high-speed ramp entry, driver fatigue, and abnormal braking, for algorithm optimization, fault diagnosis, or safety testing.

[0034] However, traditional data management systems have the following pain points:

[0035] 1. Massive data volume and lack of scene description: It is difficult to quickly locate the target scene through manual filtering or predefined labels in massive CAN data;

[0036] 2. Low query efficiency: The attribute query function that relies on SQL (Structured Query Language) or key-value databases is limited and cannot support natural language interaction or complex semantic retrieval.

[0037] 3. Scene feature extraction relies on human experience: The feature combination of complex driving scenarios, such as abnormal braking on highway ramps, needs to be manually defined by experts, which is costly and prone to missing key features.

[0038] 4. Data disconnected from context: CAN data lacks semantic association with associated sensor data such as camera and radar data, making it difficult to achieve cross-modal queries.

[0039] This application provides a vehicle-based CAN data processing method, apparatus, device, and storage medium, which aims to solve the above-mentioned technical problems in the prior art.

[0040] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.

[0041] Figure 1 This is a flowchart illustrating a vehicle-based CAN data processing method provided in an embodiment of this application. This method can be executed by a vehicle-based CAN data processing device. Figure 1 As shown, the method includes:

[0042] S101, Receive data query information input by the user; wherein, the data query information represents the user's query requirement for CAN data.

[0043] For example, a vehicle can generate CAN data in real time while in motion, and this CAN data can characterize the vehicle's driving status. By analyzing the CAN data, operations such as vehicle inspection or optimization can be performed. Vehicles generate different CAN data under different driving scenarios, and when CAN data analysis is needed, CAN data from different driving scenarios may be required. Therefore, users need to query the CAN data for a specific scenario. In other words, the amount of CAN data is large, and the CAN data required by the user may vary each time.

[0044] Users can query the required CAN data through human-machine interaction. For example, a human-machine interface can be preset, and users can input data query information into the interface according to their actual needs. The data query information represents the user's query requirements for CAN data. For example, if a user wants to query CAN data under a certain driving scenario, the data query information can include a description of that driving scenario. Users can input data query information in the form of text or voice. For example, a user can input "find CAN data for all abnormal braking events on highway ramps" as the data query information.

[0045] Through human-computer interaction, the system can receive user data query information in real time, thereby determining the driving scenario for the CAN data that the user wants to query, and enabling rapid location of the target scenario through the user's natural language.

[0046] S102. Based on the data query information and a preset vector database, determine the target data; wherein, the preset vector database includes CAN data of the vehicle under a preset driving scenario and the text vector corresponding to the CAN data, the text vector represents the natural language description corresponding to the CAN data, and the target data represents the CAN data that the user wants to query.

[0047] For example, a vector database is pre-set, which stores CAN data under different preset driving scenarios. Each CAN data also corresponds to a text vector, which can represent the natural language description of the CAN data. That is, each CAN data can have its own meaning, that is, a natural language description. The vector form of the natural language description is used as a text vector and associated with the CAN data and stored in the preset vector database.

[0048] After obtaining the data query information, the system can retrieve the required CAN data from a pre-defined vector database. This required CAN data is identified as the target data, meaning it can be determined from the pre-defined vector database. In other words, based on the user's data query information, the target scenario can be determined, and the corresponding target data can be obtained. The target scenario is the pre-defined driving scenario in the user's query request.

[0049] Text vectors in the vector database can represent the natural language description of CAN data. Data query information is also in the form of natural language description. Therefore, when querying target data, the corresponding text vector can be determined based on the data query information, and then the CAN data associated with that text vector can be found as the target data.

[0050] The vector database can also store natural language descriptions corresponding to CAN data. After obtaining data query information, the semantic similarity between the data query information and the natural language descriptions in the vector database can be calculated. The natural language description with the highest semantic similarity is determined, and the CAN data corresponding to that natural language description is identified as the target data.

[0051] S103. Generate and output response information based on the target data; wherein the response information is used to characterize the target data.

[0052] For example, the target data queried from the vector database is the raw CAN data generated by the vehicle. To provide the user with richer and more easily understood information, the target data can be expanded or adjusted to obtain response information. Through human-computer interaction, the response information is fed back to the user, completing the user's query for CAN data.

[0053] For example, based on a vehicle-specific DBC (Database CAN, dictionary of the CAN bus network) file, the content of the target data can be converted into physical values, which represent the specific meaning of the CAN data content. For instance, target data with ID 0x100 and content 0x0055 can be decoded into a speed of 85 km / h based on the DBC file. The decoded data can be sent to the user along with the target data as a response message. Alternatively, after obtaining the physical values, the obtained physical values ​​can be combined into a complete text description, and this combined text description, along with the target data, can be sent to the user as a response message. Furthermore, the storage path of the target data can be obtained, and the original CAN data, text description, and storage path can all be sent to the user as a response message.

[0054] In other words, the query results can be used as external knowledge to expand the information and generate a structured response. This can avoid the illusion of a large model, increase the richness of the response information, and improve the user's human-computer interaction experience.

[0055] This application provides a vehicle-based CAN data processing method, allowing users to quickly and accurately query vehicle CAN data through human-computer interaction. Users input query information based on their actual needs, representing their CAN data query requirements. The query information is parsed, and the target data is determined according to a preset vector database. The preset vector database may include CAN data from different preset driving scenarios and the corresponding text vectors. The text vectors are vector data describing the CAN data in natural language, and the target data represents the CAN data the user wants to query. In other words, the CAN data the user wants to query can be accurately obtained from the preset vector database. A response is generated based on the target data and fed back to the user through human-computer interaction. Natural language interaction allows direct location of the user's desired driving scenario, reducing manual query time and improving the efficiency and accuracy of CAN data processing.

[0056] Figure 2 A flowchart illustrating a vehicle-based CAN data processing method provided in this application embodiment is shown below. Figure 2 As shown, in this embodiment... Figure 1 Based on the embodiments, a vehicle-based CAN data processing method is described in detail, which includes:

[0057] S201, Receive data query information input by the user; wherein, the data query information represents the user's query request for CAN data.

[0058] For example, this step can refer to step S101 above, and will not be repeated here.

[0059] S202. Perform vector transformation on the data query information to obtain the feature vector corresponding to the data query information.

[0060] For example, if the data query information is in text format, after obtaining the data query information, it can be converted from text to vector format to obtain the corresponding feature vector. That is, feature extraction processing can be performed on the data query information to obtain the feature vector, which represents the data query information.

[0061] In this embodiment, a feature extraction network can be pre-configured to perform feature extraction processing on the data query information. The network architecture of the feature extraction network is not specifically limited in this embodiment. For example, when converting the data query information into vectors, a Sentence-Transformer model can be used to generate feature vectors.

[0062] S203. Based on the feature vector corresponding to the data query information, determine the target data from the preset vector database.

[0063] For example, the vector database includes multiple text vectors, and different text vectors can correspond to different CAN data. Based on the feature vector corresponding to the data query information, the text vector corresponding to that feature vector can be determined from the preset vector database, thereby identifying the CAN data corresponding to that text vector as the target data. For instance, a text vector matching the feature vector can be searched to obtain the target data for the driving scenario required by the user.

[0064] In this embodiment, the preset vector database includes multiple text vectors; the target data is determined from the preset vector database based on the feature vector corresponding to the data query information, including: for each text vector in the preset vector database, determining the similarity between the feature vector corresponding to the data query information and the text vector; determining the target vector from each text vector based on the similarity corresponding to each text vector; and determining the CAN data corresponding to the target vector in the preset vector database as the target data.

[0065] Specifically, the preset vector database stores multiple text vectors. For each text vector, the similarity between the feature vector corresponding to the data query information and the text vector can be calculated, thus determining the degree of similarity between the two vector data. Each text vector corresponds to a similarity score. In this embodiment, the method for calculating the similarity score is not specifically limited. For example, K-NN (K-Nearest Neighbors) can be used to calculate the similarity score.

[0066] After obtaining the similarity scores for each text vector, a target text vector is selected from these vectors based on each similarity score. For example, the text vector corresponding to the highest similarity score can be determined as the target vector. The CAN data corresponding to the target vector in the vector database is then selected as the target data.

[0067] The advantage of this setup is that by calculating similarity, searching in the vector database, and returning matching items for the data query information, it enables data retrieval in different scenarios through natural language interaction, thereby improving the accuracy and efficiency of CAN data determination.

[0068] S204. Generate and output response information based on the target data; wherein the response information is used to characterize the target data.

[0069] For example, the target data is the raw CAN data. After obtaining the target data, it can be transformed or modified to obtain the response information, which is then output to the user for easy viewing and analysis.

[0070] In this embodiment, generating and outputting response information based on target data includes: obtaining a preset driving scenario corresponding to the target data; inputting the target data and the preset driving scenario corresponding to the target data into a preset large model to obtain response information; and feeding back the response information to the user.

[0071] Specifically, in the vector database, each CAN data point can correspond to its own preset driving scenario, and different CAN data points can correspond to different preset driving scenarios. The vector database can store the association between CAN data and preset driving scenarios. The preset driving scenario represents the vehicle's driving environment, such as the vehicle's speed, acceleration, braking pressure, steering angle, and surrounding environment.

[0072] After determining the target data, the corresponding preset driving scenario is retrieved from the vector database. This preset driving scenario is the driving scenario for the CAN data that the user needs to query. Based on the target data and the corresponding preset driving scenario, a response message can be generated and delivered to the user via text or voice.

[0073] A large language model, or LLM, is pre-built and trained. In this embodiment, the model architecture of the LLM is not specifically limited. The target data and its corresponding preset driving scenario are input into the preset large model, which then outputs response information. For example, the large model can use the target data and the corresponding preset driving scenario as external knowledge augmentation, and generate response information based on RAG technology.

[0074] The benefits of this setup are that the vector database serves as an external knowledge base for the large model. Combining this external knowledge base helps avoid the illusion created by LLM, ensures that the response information is generated based on real data, improves the accuracy of data processing, and enhances the user's human-computer interaction experience.

[0075] This application provides a vehicle-based CAN data processing method, allowing users to quickly and accurately query vehicle CAN data through human-computer interaction. Users input query information based on their actual needs, representing their CAN data query requirements. The query information is parsed, and the target data is determined according to a preset vector database. The preset vector database may include CAN data from different preset driving scenarios and the corresponding text vectors. The text vectors are vector data describing the CAN data in natural language, and the target data represents the CAN data the user wants to query. In other words, the CAN data the user wants to query can be accurately obtained from the preset vector database. A response is generated based on the target data and fed back to the user through human-computer interaction. Natural language interaction allows direct location of the user's desired driving scenario, reducing manual query time and improving the efficiency and accuracy of CAN data processing.

[0076] Figure 3 A flowchart illustrating a vehicle-based CAN data processing method provided in this application embodiment is shown below. Figure 3 As shown, in this embodiment... Figure 1 and Figure 2 Based on the embodiments, a vehicle-based CAN data processing method is described in detail, which includes:

[0077] S301. Based on a preset virtual environment simulator, CAN data simulated under a preset driving scenario is generated, which is the simulated data.

[0078] For example, before the user engages in human-machine interaction, a vector database needs to be generated in advance to facilitate the determination of the CAN data the user wants to query. The CAN data in the vector database should be the vehicle's actual CAN data and the corresponding preset driving scenarios. While actual CAN data can be collected in real time, accurate preset driving scenarios cannot be obtained. Therefore, simulated CAN data and corresponding preset driving scenarios can be generated first, and then the preset driving scenarios for the actual CAN data can be determined based on the preset driving scenarios of the simulated CAN data.

[0079] Simulated CAN data is CAN data generated in a virtual environment; this simulated CAN data can be called simulated data. Preset virtual environment simulators can be used to generate simulated data. For example, the preset virtual environment simulator is CARLA (Car Learning to Act, an autonomous driving simulator). CARLA supports highly realistic traffic environment simulation, including elements such as vehicles, pedestrians, weather, and roads. When generating simulated data, virtual vehicles and virtual environments can be set through CARLA, i.e., inputting scenario parameters. These input scenario parameters constitute a preset driving scenario. For example, vehicle blueprints can be modified to simulate different vehicle models, and driving scenario parameters, including speed and steering change rate, can be input. Maps can also be loaded, weather settings can be set, and NPC (Non-Player Character) vehicles can be added. Specifically, a preset vehicle type can be set, the speed can be set to greater than 80 km / h, the steering change rate can be set to greater than 5° / s, a highway map can be loaded, and the weather can be set to sunny.

[0080] The virtual environment simulator allows you to set up a variety of preset driving scenarios, and in each preset driving scenario, corresponding simulation data can be generated.

[0081] In this embodiment, based on a preset virtual environment simulator, simulated CAN data under a preset driving scenario is generated. The simulated data includes: generating sensor data under a preset driving scenario based on the preset virtual environment simulator; determining vehicle control information based on the sensor data; wherein the control information represents the control commands issued by the vehicle in response to the sensor data; and converting the control information into CAN data to obtain the simulated data.

[0082] Specifically, the data directly output by the virtual environment simulator is not CAN data, but simulated sensor data. For example, a virtual camera image can be generated using a virtual environment simulator. Based on the simulated sensor data, simulated CAN data can be generated.

[0083] Simulated sensor data can be input into the vehicle's Open Pilot (an open-source autonomous driving system). Open Pilot generates control information based on the sensor data, representing the control commands the vehicle will issue in response to the sensor data, such as acceleration or steering. This Open Pilot control information is then fed back to CARLA's simulated vehicle, generating a simulated CAN signal sequence—i.e., simulated data—based on a preset driving scenario. A scenario label can be generated from the simulated data to represent the preset driving scenario, thus generating both simulated data and the corresponding scenario label. The simulated data and the corresponding preset driving scenario can be used as templates for the content in a vector database, facilitating the subsequent generation of the actual vector database.

[0084] The beneficial effect of this setup is that it obtains virtual sensor data from CARLA, acquires vehicle control information based on this data, converts the control information into CAN frames, simulates CAN data, and obtains corresponding preset driving scenarios. This solves the problem of relying on human experience for the definition of features in complex driving scenarios and improves the accuracy and efficiency of vector database generation.

[0085] S302. Obtain the CAN data generated by the vehicle, which is the actual data.

[0086] For example, a vehicle can generate real-time CAN data during actual driving. This real-time CAN data generated by the vehicle is referred to as real data. The real data from the vehicle is acquired for subsequent construction of the vector database. Real data is typically collected from the vehicle's preset interface or data logger and can be stored in formats such as Bag, containing the original CAN frames.

[0087] S303. Generate a preset vector database based on the simulated data and the real data.

[0088] For example, after obtaining simulated data and real data, the simulated data and real data can be combined to determine the preset driving scenario corresponding to the real data. Then, a vector database can be constructed based on the real data and the preset driving scenario corresponding to the real data. For instance, simulated data and real data can be compared to determine the simulated data that matches the real data. The preset driving scenario corresponding to this simulated data can then be identified as the preset driving scenario for the real data, and the real data and the preset driving scenario for the real data can be stored in the vector database.

[0089] In this embodiment, generating a preset vector database based on simulated data and real data includes: parsing the simulated data to obtain the physical values ​​corresponding to the simulated data; wherein the physical values ​​represent the driving conditions of the vehicle; generating an initial vector database based on the physical values ​​corresponding to the simulated data and the preset driving scenario corresponding to the simulated data; and generating a preset vector database based on the real data and the initial vector database.

[0090] Specifically, both simulated and real data are in CAN data format. For simulated data, it can be parsed to determine the corresponding physical values. These physical values ​​represent the vehicle's driving conditions; that is, they are the load content within the simulated data. For example, simulated data can be read, and the ID and load of each CAN frame can be decoded. Based on a pre-defined DBC file, the load is converted into physical values. Physical values ​​can be represented using structured Data Frames, which may include timestamps, speed, acceleration, brake pressure, steering angle, etc. For example, for simulated data with ID 0×100, the load 0×0055 can be decoded as a speed of 85 km / h.

[0091] For each simulation data point, determine its corresponding physical value and the preset driving scenario. Based on the physical values ​​and preset driving scenarios of each simulation data point, generate an initial vector database. This initial vector database serves as the knowledge base template for the final generated preset vector database.

[0092] Based on the initial vector database, the driving scenarios corresponding to the real data can be obtained. Then, based on the real data and the corresponding driving scenarios, the final preset vector database is obtained. That is, the driving scenarios in the final preset vector database can be the same as those in the initial vector database, but the simulated data is replaced with real data, allowing the real CAN data to be retrieved from the preset vector database.

[0093] The advantage of this setup is that, based on simulated data and corresponding scene labels, it summarizes scene characteristics and generates scene feature summaries. For example, the scene characteristics for a high-speed ramp entry scenario are: speed maintained at 80-100 km / h, steering angle gradually changing by 10-20°, no emergency braking signal, and acceleration remaining stable at 2-4 m / s². Based on these characteristics and simulated data, a knowledge base template is formed, which is the initial vector database. Then, based on this initial vector database, the final vector database is generated. This achieves automatic generation of realistic scene descriptions, solves the problem of defining complex driving scenarios relying on human experience, and improves the efficiency of vector database generation.

[0094] In this embodiment, generating a preset vector database based on real data and an initial vector database includes: parsing the real data to obtain the physical values ​​corresponding to the real data; determining the preset driving scenario corresponding to the real data from the initial vector database based on the physical values ​​corresponding to the real data; and determining the preset vector database based on the physical values ​​corresponding to the real data and the preset driving scenario corresponding to the real data.

[0095] Specifically, for each piece of real data, the real data can be parsed to determine its corresponding physical value. For example, real data can be read, the ID and payload of each CAN frame in the real data can be decoded, and the payload can be converted into a physical value according to a preset DBC file.

[0096] Based on the physical values ​​corresponding to the real data, the preset driving scenario corresponding to the real data is determined from the initial vector database. For example, based on the physical values ​​corresponding to the real data and the physical values ​​corresponding to the simulated data, simulated data consistent with the real data can be determined, and thus the preset driving scenario corresponding to the simulated data can be determined as the preset driving scenario corresponding to the real data.

[0097] Based on the physical values ​​corresponding to the real data and the preset driving scenarios corresponding to the real data, a preset vector database is determined. That is, the physical values ​​corresponding to the real data and the preset driving scenarios corresponding to the real data are associated and stored in the database as the preset vector database.

[0098] The pre-defined vector database includes text vectors corresponding to CAN data, i.e., text vectors corresponding to real data. In other words, real data can be converted into vector form to obtain text vectors, thereby enabling the association and storage of text vectors, real data, and pre-defined driving scenarios.

[0099] The advantage of this setup is that it decodes the payload of the real CAN data, that is, the physical value, and searches for the driving scenario corresponding to the real data in the initial database based on the physical value. Combining the driving scenario and the real data, it automatically builds a vector database without the need for experts to manually define the scenario, thereby reducing labor costs and improving the determination accuracy and efficiency of the vector database.

[0100] In this embodiment, the preset driving scenario corresponding to the real data is determined from the initial vector database based on the physical value corresponding to the real data. This includes: for each physical value corresponding to the simulated data in the initial vector database, determining the difference between the physical value corresponding to the simulated data and the physical value corresponding to the real data; if the difference is within a preset difference range, then the preset driving scenario corresponding to the simulated data is determined as the preset driving scenario corresponding to the real data.

[0101] Specifically, for each simulated data point in the initial vector database, the corresponding physical value is determined. For each real data point, the physical value is determined. For each simulated data point and each real data point, the difference between the physical value corresponding to the simulated data and the physical value corresponding to the real data is determined. CAN data can correspond to multiple types of physical values, and the difference between physical values ​​of the same type between real and simulated data can be determined. For example, the difference between the velocity in real data and the velocity in simulated data can be determined.

[0102] For each type of physical value, a corresponding difference range can be preset. The system determines whether the difference falls within this range. If it does, the real data and simulated data are considered consistent, and the preset driving scenario corresponding to the simulated data can be used as the preset driving scenario corresponding to the real data. If not, the real data and simulated data are considered inconsistent, and the preset driving scenario corresponding to the simulated data cannot be used as the preset driving scenario corresponding to the real data.

[0103] In this embodiment, a similarity search can be performed on simulated data and real data to determine the simulated data with the highest similarity to the real data, thereby obtaining the corresponding preset driving scenario. In this embodiment, the method of similarity search is not specifically limited; for example, Euclidean distance calculation can be used.

[0104] The advantage of this setup is that it allows for the search for simulated data within the same range as real data to obtain corresponding driving scenarios, reducing the need for manual determination of driving scenarios and improving the determination efficiency of the vector database.

[0105] In this embodiment, a preset vector database is determined based on the physical values ​​corresponding to the real data and the preset driving scenarios corresponding to the real data. This includes: determining the natural language description corresponding to the real data based on the physical values ​​corresponding to the real data and the preset driving scenarios corresponding to the real data; performing vector conversion processing on the natural language description corresponding to the real data to obtain the text vector corresponding to the real data; and associating and storing the real data and the text vector corresponding to the real data to obtain the preset vector database.

[0106] Specifically, the pre-defined vector database includes text vectors corresponding to the target data. The corresponding text vectors can be determined based on the physical values ​​of the target data and the corresponding pre-defined driving scenarios. For example, the physical values ​​of real data and the pre-defined driving scenarios can be input into a pre-defined large model. The large model outputs a natural language description representing the real data, and then feature extraction processing is performed on the natural language description to obtain the text vectors corresponding to the real data. During feature extraction of the natural language description, it can be divided into blocks to obtain multiple sentence blocks. Each sentence block is then input into the Sentence-Transformer model to obtain multi-dimensional text vectors. For example, it could be a 768-dimensional text vector.

[0107] Each piece of real data in a vector database can correspond to a unique index, which facilitates subsequent searching and locating.

[0108] The advantage of this setup is that it combines physical values ​​and driving scenarios as context to generate natural language descriptions of real-world data. These natural language descriptions are then converted into text vectors, stored in a vector database, and associated with original CAN data fragments and other data, improving the accuracy and efficiency of vector database generation, as well as enhancing the richness of the data within the vector database.

[0109] S304. Receive data query information input by the user; wherein, the data query information represents the user's query requirement for CAN data.

[0110] For example, this step can refer to step S101 above, and will not be repeated here.

[0111] S305. Based on the data query information and a preset vector database, determine the target data; wherein, the preset vector database includes the CAN data of the vehicle under a preset driving scenario and the text vectors corresponding to the CAN data, the text vectors represent the natural language descriptions corresponding to the CAN data, and the target data represents the CAN data that the user wants to query.

[0112] For example, this step can refer to step S102 above, and will not be repeated here.

[0113] S306. Generate and output response information based on the target data; wherein the response information is used to characterize the target data.

[0114] For example, this step can refer to step S103 above, and will not be repeated here.

[0115] In this embodiment, Figure 4 This is a schematic diagram of the CAN data query process. Figure 4In this system, multiple preset driving scenarios can be pre-set. The scenario parameters of these preset driving scenarios are input into the virtual environment simulator to generate simulated CAN data, i.e., simulated data. Using a large language model, the simulated data and the corresponding preset driving scenarios are analyzed to obtain a summary of scenario features. This summary of scene features characterizes the simulated data and the corresponding preset driving scenarios. Based on the summary of scene features, that is, based on the simulated data and the preset driving scenarios, a knowledge base template, i.e., the initial vector database, is generated.

[0116] Obtain real CAN data, i.e., real data. Parse the CAN frames in the real data to obtain the physical values ​​of the real data. Based on the physical values ​​of the real data and the knowledge base template, find the corresponding preset driving scenario for the real data. Using a large language model, generate a semantic description of the real data, i.e., a natural language description. Combine the semantic description and the real data to generate a vector database. In this embodiment, the vector database can also be in the form of a graph database.

[0117] When a user interacts with a computer, they input natural language into a large language model to query information. The model then performs vector transformations using an embedded vector model to obtain feature vectors. For example, the embedded vector model could be a Sentence-Transformer model. Based on the extracted feature vectors, the system queries a vector database for target data, and then uses the large language model to generate a response, which is then sent back to the user, thus completing the human-computer interaction process.

[0118] This application provides a vehicle-based CAN data processing method, allowing users to quickly and accurately query vehicle CAN data through human-computer interaction. Users input query information based on their actual needs, representing their CAN data query requirements. The query information is parsed, and the target data is determined according to a preset vector database. The preset vector database may include CAN data from different preset driving scenarios and the corresponding text vectors. The text vectors are vector data describing the CAN data in natural language, and the target data represents the CAN data the user wants to query. In other words, the CAN data the user wants to query can be accurately obtained from the preset vector database. A response is generated based on the target data and fed back to the user through human-computer interaction. Natural language interaction allows direct location of the user's desired driving scenario, reducing manual query time and improving the efficiency and accuracy of CAN data processing.

[0119] Figure 5 A schematic diagram of a vehicle-based CAN data processing device provided in this application embodiment is shown below. Figure 5 As shown, the vehicle-based CAN data processing device 50 provided in this embodiment includes:

[0120] The information receiving unit 501 is used to receive data query information input by the user; wherein, the data query information represents the user's query request for CAN data;

[0121] The target determination unit 502 is used to determine the target data based on the data query information and a preset vector database. The preset vector database includes the CAN data of the vehicle in a preset driving scenario and the text vectors corresponding to the CAN data. The text vectors represent the natural language descriptions corresponding to the CAN data, and the target data represents the CAN data that the user wants to query.

[0122] The information output unit 503 is used to generate and output response information based on the target data; wherein the response information is used to characterize the target data.

[0123] In one possible implementation, the target determination unit 502 includes:

[0124] The vector transformation module is used to perform vector transformation processing on the data query information to obtain the feature vector corresponding to the data query information;

[0125] The target determination module is used to determine target data from a preset vector database based on the feature vectors corresponding to the data query information.

[0126] In one possible implementation, the pre-defined vector database includes multiple text vectors; the target determination module is specifically used for:

[0127] For each text vector in the pre-defined vector database, determine the similarity between the feature vector corresponding to the data query information and the text vector;

[0128] The target vector is determined from each text vector based on the similarity between the text vectors.

[0129] The CAN data corresponding to the target vector in the preset vector database is identified as the target data.

[0130] In one possible implementation, the information output unit 503 includes:

[0131] The scene acquisition module is used to acquire the preset driving scene corresponding to the target data;

[0132] The information acquisition module is used to input target data and the corresponding preset driving scenario into a preset large model and obtain response information.

[0133] The information feedback module is used to send reply information back to the user.

[0134] One possible implementation also includes:

[0135] The data simulation unit is used to generate simulated CAN data under a preset driving scenario based on a preset virtual environment simulator; this data is called simulated data.

[0136] The data acquisition unit is used to acquire the CAN data generated by the vehicle, which is real data.

[0137] The database generation unit is used to generate a preset vector database based on simulated data and real data.

[0138] In one possible implementation, the data simulation unit is specifically used for:

[0139] Based on a preset virtual environment simulator, sensor data simulated under preset driving scenarios is generated;

[0140] Based on sensor data, the vehicle's control information is determined; whereby the control information represents the control commands issued by the vehicle in response to the sensor data.

[0141] The control information is converted into CAN data to obtain analog data.

[0142] In one possible implementation, the database generation unit includes:

[0143] The physical value acquisition module is used to parse the simulation data and obtain the physical values ​​corresponding to the simulation data; the physical values ​​represent the vehicle's driving conditions.

[0144] The initial generation module is used to generate an initial vector database based on the physical values ​​corresponding to the simulation data and the preset driving scenarios corresponding to the simulation data.

[0145] The database generation module is used to generate a preset vector database based on real data and an initial vector database.

[0146] In one possible implementation, the database generation module is specifically used for:

[0147] Analyze the real data to obtain the physical values ​​corresponding to the real data;

[0148] Based on the physical values ​​corresponding to the real data, the preset driving scenarios corresponding to the real data are determined from the initial vector database;

[0149] Based on the physical values ​​corresponding to the real data and the preset driving scenarios corresponding to the real data, a preset vector database is determined.

[0150] In one possible implementation, the database generation module is specifically used for:

[0151] For each physical value corresponding to a simulated data in the initial vector database, determine the difference between the physical value corresponding to the simulated data and the physical value corresponding to the real data;

[0152] If the difference is within the preset difference range, the preset driving scenario corresponding to the simulated data will be determined as the preset driving scenario corresponding to the real data.

[0153] In one possible implementation, the database generation module is specifically used for:

[0154] Based on the physical values ​​corresponding to the real data and the preset driving scenarios corresponding to the real data, determine the natural language description corresponding to the real data;

[0155] The natural language description corresponding to the real data is processed by vector conversion to obtain the text vector corresponding to the real data.

[0156] The real data and its corresponding text vectors are associated and stored to obtain a pre-defined vector database.

[0157] This embodiment provides a vehicle-based CAN data processing device that can execute the methods provided in the above-described method embodiments. Its implementation principle and technical effects are similar, and will not be described in detail here.

[0158] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 6 As shown, the electronic device 60 provided in this embodiment includes at least one processor 601 and a memory 602. Optionally, the device 60 further includes a communication component 603. The processor 601, memory 602, and communication component 603 are connected via a bus 604.

[0159] In a specific implementation, at least one processor 601 executes computer execution instructions stored in memory 602, causing at least one processor 601 to perform the above-described method.

[0160] The specific implementation process of processor 601 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.

[0161] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor.

[0162] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.

[0163] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.

[0164] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.

[0165] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the above-described method.

[0166] The aforementioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.

[0167] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an Application Specific Integrated Circuit (ASIC). Alternatively, the processor and the readable storage medium can exist as discrete components in the device.

[0168] The division of units is merely a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.

[0169] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0170] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0171] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0172] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.

[0173] Finally, it should be noted that other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This invention is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein, and is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.

Claims

1. A vehicle-based CAN data processing method, characterized by, The method comprises: receiving data query information input by a user; wherein the data query information represents a user's query requirement for CAN data; determining target data based on a preset vector database according to the data query information; wherein the preset vector database comprises CAN data of a vehicle in a preset driving scenario and a text vector corresponding to the CAN data, the text vector representing a natural language description corresponding to the CAN data, and the target data representing the CAN data to be queried by the user; generating and outputting reply information according to the target data; wherein the reply information represents the target data.

2. The method of claim 1, wherein, The method comprises: performing vector conversion processing on the data query information to obtain a feature vector corresponding to the data query information; determining the target data from the preset vector database according to the feature vector corresponding to the data query information.

3. The method of claim 2, wherein, The preset vector database comprises a plurality of text vectors; and the method comprises: for each text vector in the preset vector database, determining a similarity between the feature vector corresponding to the data query information and the text vector; determining a target vector from each text vector according to the similarity corresponding to each text vector; determining, as the target data, CAN data corresponding to the target vector in the preset vector database.

4. The method of claim 1, wherein, The method comprises: obtaining a preset driving scenario corresponding to the target data; inputting the target data and the preset driving scenario corresponding to the target data into a preset large model to obtain the reply information; feeding back the reply information to the user.

5. The method according to any one of claims 1-4, characterized in that, The method further comprises: generating, based on a preset virtual environment simulator, simulated CAN data in a preset driving scenario as simulation data; obtaining CAN data generated by a vehicle as real data; generating the preset vector database according to the simulation data and the real data.

6. The method of claim 5, wherein, The method comprises: generating, based on a preset virtual environment simulator, simulated sensor data in a preset driving scenario; determining control information of a vehicle according to the sensor data; wherein the control information represents a control instruction issued by the vehicle for the sensor data; converting the control information into CAN data to obtain the simulation data.

7. The method of claim 5, wherein, The method comprises: performing analysis on the simulation data to obtain a physical value corresponding to the simulation data; wherein the physical value represents a driving condition of the vehicle; generating an initial vector database according to the physical value corresponding to the simulation data and the preset driving scenario corresponding to the simulation data; generating the preset vector database according to the real data and the initial vector database.

8. The method of claim 7, wherein, According to the real data and the initial vector database, the preset vector database is generated, including: The real data is parsed to obtain a physical value corresponding to the real data; According to the physical value corresponding to the real data, a preset driving scene corresponding to the real data is determined from the initial vector database; According to the physical value corresponding to the real data and the preset driving scene corresponding to the real data, the preset vector database is determined.

9. A vehicle-based CAN data processing apparatus characterized by comprising: Comprising: An information receiving unit configured to receive data query information input by a user; wherein the data query information represents a query requirement of the user on CAN data; A target determining unit configured to determine target data based on a preset vector database according to the data query information; wherein the preset vector database comprises CAN data of a vehicle in a preset driving scene and a text vector corresponding to the CAN data, the text vector represents a natural language description corresponding to the CAN data, and the target data represents the CAN data to be queried by the user; An information output unit configured to generate and output reply information according to the target data; wherein the reply information is used to represent the target data.

10. An electronic device / computer readable storage medium / computer program product, characterized in that: The electronic device comprises a memory and a processor; the memory stores computer execution instructions; the processor executes the computer execution instructions stored in the memory, so that the processor executes the method of any one of claims 1-8; and / or The computer readable storage medium stores computer execution instructions, and the computer execution instructions are executed by the processor to implement the method of any one of claims 1-8; and / or The computer program product comprises a computer program, and the computer program is executed by the processor to implement the method of any one of claims 1-8.