Data processing method, device, equipment, medium, product, vehicle and system
By acquiring and updating the identification model of abnormal data sets, the problem of low recognition efficiency of vehicle abnormal log data is solved, the accurate identification of similar data and the removal of duplicate data is achieved, and the accuracy of abnormal data positioning is improved.
Patent Information
- Application Number
- CN202510115177.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-06-10
AI Technical Summary
In the prior art, when identifying abnormal log data generated by vehicles, it is difficult to accurately identify similar abnormal log data, resulting in low recognition efficiency and repeated data interfering with the positioning of true abnormal data.
By obtaining the exception data set, including the unmatched first exception data and the second exception data in the preset data set, the target annotation information is obtained and the identification model is updated to improve the identification efficiency of similar data.
By updating the identification model in real time, the efficiency of identifying similar data for abnormal log data is improved, the interference of duplicate data is reduced, and the accuracy of positioning abnormal data is improved.
Smart Images

Figure CN120123925A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and in particular, to a data processing method, apparatus, device, medium, product, vehicle, and system. Background Art
[0002] With the popularization of vehicle electrification, a large amount of log data is generated by vehicles. Among them, the abnormal log data in the log data is particularly important for vehicle research.
[0003] Currently, in the abnormal log data, different vehicles or the same vehicle may report a large number of similar abnormal log data at different times, and the meanings of these abnormal log data are the same.
[0004] However, due to the differences in information such as the abnormal stack location and abnormal service numbers in the abnormal log data, the string matching method used in the prior art is not sufficient to accurately identify such similar abnormal log data. As a result, when the vehicle generates abnormal log data, the recognition efficiency of similar data for the abnormal log data is low, and a large number of duplicate abnormal log data will interfere with relevant personnel in locating truly valuable abnormal log data. Summary of the Invention
[0005] An embodiment of this application provides a data processing method, which can improve the recognition efficiency of similar data for abnormal log data when the vehicle generates abnormal log data, so as to at least partially solve the above technical problems.
[0006] To achieve the above object, according to the first aspect of this application, a data processing method is provided, including:
[0007] Obtain at least one abnormal data group, where the abnormal data group includes first abnormal data and at least one second abnormal data. The first abnormal data is data for which no similar abnormal data is matched in the preset data set by the recognition model, and the second abnormal data is abnormal data in the preset data set;
[0008] Obtain target annotation information corresponding to the abnormal data group, where the target annotation information is used to indicate the similarity between the first abnormal data and the second abnormal data;
[0009] Based on the abnormal data group and the target annotation information, update the model parameters of the recognition model to obtain an updated recognition model.
[0010] Optionally, the obtaining at least one abnormal data group includes:
[0011] Obtain target abnormal data to be recognized generated when the vehicle is abnormal;
[0012] Using the above recognition model, match similar abnormal data for the above target abnormal data in the above preset data set;
[0013] If no similar abnormal data is matched for the above target abnormal data, then determine the above target abnormal data as the first abnormal data;
[0014] Select the second abnormal data combined with the above first abnormal data from the preset data set to obtain an abnormal data group.
[0015] Optionally, the above-mentioned matching of similar abnormal data for the above target abnormal data in the above preset data set includes:
[0016] Input the above target abnormal data into the above recognition model, and obtain the first feature vector of the above target abnormal data through the above recognition model;
[0017] Based on the above first feature vector, match similar abnormal data for the above target abnormal data in the above preset data set.
[0018] Optionally, the above preset data set is stored in a vector database, and the above preset data set includes second feature vectors corresponding to each abnormal data. The above-mentioned matching of similar abnormal data for the above target abnormal data in the above preset data set includes:
[0019] Input the above first feature vector into the above vector database;
[0020] Through the above vector database, perform similarity matching between the above first feature vector and the second feature vectors in the above preset data set to obtain a matching result.
[0021] Optionally, the above-mentioned performing similarity matching between the above first feature vector and the second feature vectors in the above preset data set through the above vector database to obtain a matching result includes:
[0022] Through the above vector database, calculate the similarity between the above first feature vector and the second feature vectors in the above preset data set;
[0023] If there is no second feature vector in the above preset data set whose similarity meets the preset similarity threshold, then determine that no similar abnormal data is matched for the above target abnormal data;
[0024] If there is a second feature vector in the above preset data set whose similarity meets the preset similarity threshold, then determine that similar abnormal data is matched for the above target abnormal data.
[0025] Optionally, selecting the second abnormal data that combines with the first abnormal data from the preset data set to obtain an abnormal data group, including:
[0026] Sending the first abnormal data to the annotation platform to receive the data feedback information of the annotation platform for the first abnormal data;
[0027] Based on the data feedback information, selecting the second abnormal data that combines with the first abnormal data from the preset data set to obtain an abnormal data group.
[0028] Optionally, obtaining the target annotation information corresponding to the abnormal data group, including:
[0029] Based on the data feedback information, obtaining the target annotation information for the first abnormal data.
[0030] Optionally, the method further includes:
[0031] If the target abnormal data matches similar abnormal data, then eliminating the target abnormal data.
[0032] Optionally, the target annotation information includes a similar label and / or a dissimilar label. The similar label is used to indicate that the first abnormal data and the corresponding second abnormal data are similar, and the dissimilar label is used to indicate that the first abnormal data and the corresponding second abnormal data are not similar.
[0033] Optionally, based on the abnormal data group and the target annotation information, updating the model parameters of the recognition model to obtain an updated recognition model, including:
[0034] Inputting the abnormal data group into the recognition model, and obtaining the target feature vectors corresponding to each abnormal data in the abnormal data group through the recognition model;
[0035] Calculating a loss value through a preset loss function based on each target feature vector in the abnormal data group and the target annotation information;
[0036] Updating the model parameters of the recognition model based on the calculated loss value to obtain an updated recognition model.
[0037] Optionally, the recognition model is configured with two identical network structures, the second abnormal data includes one, and inputting the abnormal data group into the recognition model, and obtaining the target feature vectors corresponding to each abnormal data in the abnormal data group through the recognition model, including:
[0038] Input the above first abnormal data and the above second abnormal data into a network structure respectively, and obtain the target feature vectors of the above first abnormal data and the target feature vectors of the above second abnormal data through each network structure respectively;
[0039] The above calculation of the loss value based on each target feature vector in the above abnormal data group and the above target annotation information through a preset loss function includes:
[0040] Through a preset loss function, based on the similarity between the above first abnormal data and the above second abnormal data indicated by the above target annotation information, calculate the cosine similarity between the target feature vector of the above first abnormal data and the target feature vector of the above second abnormal data to obtain the loss value.
[0041] Optionally, the above recognition model is configured with three identical network structures, the above second abnormal data includes two, the above target annotation information indicates that the above first abnormal data is similar to one second abnormal data, and the above first abnormal data is not similar to another second abnormal data;
[0042] The above input of the above abnormal data group into the recognition model to obtain the target feature vectors corresponding to each abnormal data in the above abnormal data group through the above recognition model includes:
[0043] Input the above first abnormal data and each of the above second abnormal data into a network structure respectively, and obtain the target feature vectors of the above first abnormal data and the target feature vectors of each of the above second abnormal data through each network structure respectively;
[0044] The above calculation of the loss value based on each target feature vector in the above abnormal data group and the above target annotation information through a preset loss function includes:
[0045] Through a preset loss function, based on the target feature vector of the above first abnormal data and the target feature vectors of each of the above second abnormal data, calculate the first vector distance between the above first abnormal data and the similar second abnormal data, and the second vector distance between the above first abnormal data and the dissimilar second abnormal data respectively;
[0046] Calculate the loss value based on the above first vector distance, the above second vector distance and a preset distance difference parameter.
[0047] According to the second aspect of the present application, a communication device is provided, including:
[0048] A data acquisition module for acquiring at least one abnormal data group, where the abnormal data group includes first abnormal data and at least one second abnormal data. The first abnormal data is data for which no similar abnormal data is matched in a preset data set by an identification model, and the second abnormal data is abnormal data in the preset data set;
[0049] An information acquisition module for acquiring target annotation information corresponding to the abnormal data group, where the target annotation information is used to indicate the similarity between the first abnormal data and the second abnormal data;
[0050] A model update module for updating model parameters of the identification model based on the abnormal data group and the target annotation information to obtain an updated identification model.
[0051] Optionally, the data acquisition module includes:
[0052] Acquire target abnormal data to be identified generated when the vehicle is abnormal;
[0053] Match similar abnormal data for the target abnormal data in the preset data set through the identification model;
[0054] If no similar abnormal data is matched for the target abnormal data, determine the target abnormal data as the first abnormal data;
[0055] Select second abnormal data combined with the first abnormal data from the preset data set to obtain an abnormal data group.
[0056] Optionally, the data acquisition module includes:
[0057] Input the target abnormal data into the identification model, and obtain a first feature vector of the target abnormal data through the identification model;
[0058] Match similar abnormal data for the target abnormal data in the preset data set based on the first feature vector.
[0059] Optionally, the preset data set is stored in a vector database, and the preset data set includes second feature vectors corresponding to each abnormal data. The data acquisition module includes:
[0060] Input the first feature vector into the vector database;
[0061] Through the vector database, perform similarity matching between the first feature vector and the second feature vectors in the preset data set to obtain a matching result.
[0062] Optionally, the data acquisition module includes:
[0063] Calculate the similarity between the first feature vector and the second feature vector in the preset data set through the above vector database;
[0064] If there is no second feature vector in the preset data set whose similarity meets the preset similarity threshold, it is determined that the target abnormal data does not match similar abnormal data;
[0065] If there is a second feature vector in the preset data set whose similarity meets the preset similarity threshold, it is determined that the target abnormal data matches similar abnormal data.
[0066] Optionally, the above data acquisition module includes:
[0067] Send the above first abnormal data to the annotation platform to receive the data feedback information of the annotation platform for the above first abnormal data;
[0068] Based on the above data feedback information, select the second abnormal data combined with the above first abnormal data from the preset data set to obtain an abnormal data group.
[0069] Optionally, the above information acquisition module includes:
[0070] Based on the above data feedback information, obtain the target annotation information for the above first abnormal data.
[0071] Optionally, the above communication device further includes a data elimination module, and the data elimination module includes:
[0072] If the above target abnormal data matches similar abnormal data, the above target abnormal data is eliminated.
[0073] Optionally, the above target annotation information includes a similar label and / or a dissimilar label. The similar label is used to indicate that the above first abnormal data and the corresponding second abnormal data are similar, and the dissimilar label is used to indicate that the above first abnormal data and the corresponding second abnormal data are not similar.
[0074] Optionally, the above model update module includes:
[0075] Input the above abnormal data group into the recognition model, and obtain the target feature vectors corresponding to each abnormal data in the above abnormal data group through the above recognition model;
[0076] Calculate the loss value based on each target feature vector in the above abnormal data group and the above target annotation information through a preset loss function;
[0077] Update the model parameters of the recognition model based on the calculated loss value to obtain an updated recognition model.
[0078] Optionally, the above recognition model is configured with two identical network structures. The above second abnormal data includes one. The above model update module includes:
[0079] Input the above first abnormal data and the above second abnormal data into a network structure respectively, and obtain the target feature vectors of the above first abnormal data and the target feature vectors of the above second abnormal data through each network structure respectively;
[0080] Through a preset loss function, based on the similarity between the above first abnormal data and the above second abnormal data indicated by the above target annotation information, calculate the cosine similarity between the target feature vector of the above first abnormal data and the target feature vector of the above second abnormal data to obtain a loss value.
[0081] Optionally, the above recognition model is configured with three identical network structures. The above second abnormal data includes two. The above target annotation information indicates that the above first abnormal data is similar to one second abnormal data, and the above first abnormal data is not similar to another second abnormal data;
[0082] The above model update module includes:
[0083] Input the above first abnormal data and each of the above second abnormal data into a network structure respectively, and obtain the target feature vectors of the above first abnormal data and the target feature vectors of each of the above second abnormal data through each network structure respectively;
[0084] Through a preset loss function, based on the target feature vector of the above first abnormal data and the target feature vectors of each of the above second abnormal data, calculate the first vector distance between the above first abnormal data and the similar second abnormal data respectively, and the second vector distance between the above first abnormal data and the dissimilar second abnormal data;
[0085] Based on the above first vector distance, the above second vector distance and a preset distance difference parameter, calculate a loss value.
[0086] According to the third aspect of the present application, there is provided an electronic device, including one or more processors and a memory. The above memory stores a computer program. When the above computer program is executed by the above processor, the above processor is enabled to execute any one of the data processing methods provided by the embodiments of the present application.
[0087] According to the fourth aspect of the present application, there is provided a computer-readable storage medium, including a computer program. When the above computer program runs on a controller, the above computer program is used to enable the above controller to execute any one of the data processing methods provided by the embodiments of the present application.
[0088] According to a fifth aspect of the present application, there is provided a computer program product, including a computer program or instructions, which, when executed by a processor, implement any of the data processing methods provided in the embodiments of the present application.
[0089] According to a sixth aspect of the present application, there is provided a vehicle, which includes an electronic device.
[0090] According to a seventh aspect of the present application, there is provided a data processing system, including a vehicle and an electronic device;
[0091] The above vehicle is used to generate first abnormal data and send the first abnormal data to the electronic device;
[0092] The above electronic device is used to execute the steps of any of the data processing methods provided in the embodiments of the present application on the first abnormal data.
[0093] In the data processing method of the embodiments of the present application, by obtaining at least one abnormal data group, the abnormal data group includes first abnormal data and at least one second abnormal data, the first abnormal data is data that no similar abnormal data is matched in a preset data set by an identification model, the second abnormal data is abnormal data in the preset data set, then, obtaining target annotation information corresponding to the abnormal data group, the target annotation information is used to indicate the similarity between the first abnormal data and the second abnormal data, and finally, based on the abnormal data group and the target annotation information, updating the model parameters of the identification model to obtain an updated identification model, so as to realize real-time updating of the identification model, so that the updated model learns the characteristics corresponding to the target abnormal data, so as to match similar abnormal data for abnormal data similar to the target abnormal data, resulting in when the vehicle generates abnormal log data, continuously updating the identification model to improve the identification efficiency of the identification model for similar data of the abnormal log data.
[0094] Other features and advantages of the present application will be described in detail in the subsequent specific implementation part. Description of the Drawings
[0095] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0096] In order to more completely understand the present application and its beneficial effects, the following will be described in conjunction with the drawings, where the same reference numerals in the following description represent the same parts.
[0097] Figure 1 It is the first flowchart of the data processing method provided in the exemplary embodiment of the present disclosure;
[0098] Figure 2 It is the second flowchart of the data processing method provided in the exemplary embodiment of the present disclosure;
[0099] Figure 3 It is the second flowchart of the data processing method provided in the exemplary embodiment of the present disclosure;
[0100] Figure 4 It is the third flowchart of the data processing method provided in the exemplary embodiment of the present disclosure;
[0101] Figure 5 It is the fourth flowchart of the data processing method provided in the exemplary embodiment of the present disclosure;
[0102] Figure 6 It is a schematic diagram of a network structure provided in the exemplary embodiment of the present disclosure;
[0103] Figure 7 It is another schematic diagram of a network structure provided in the exemplary embodiment of the present disclosure;
[0104] Figure 8 It is a schematic diagram of the structure of the data processing device provided in the exemplary embodiment of the present disclosure;
[0105] Figure 9 It is a schematic diagram of the structure of the electronic device provided in the embodiment of the present application. Detailed implementation manners
[0106] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the protection scope of the present application.
[0107] The embodiments of the present application provide a data processing method, device, equipment, medium, product, vehicle and system. The data processing device can be integrated in electronic devices such as terminal devices and / or cloud servers. For example, the terminal device can be a vehicle-mounted control terminal, an in-vehicle data processing system, an in-vehicle communication system, a drone controller (such as a handle), a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, etc., which are devices installed on a vehicle.
[0108] In addition, "a plurality of" in the embodiments of the present application means two or more. "First", "second", etc. in the embodiments of the present application are used for distinguishing descriptions and should not be construed as implying relative importance.
[0109] The following will be described in detail respectively. It should be noted that the description order of the following embodiments does not limit the preferred order of the embodiments.
[0110] Please refer to Figure 1 , Figure 1 , which is a schematic flowchart of a data processing method provided by an embodiment of the present application. For the convenience of description, in the embodiments of the present application, a terminal device is taken as an example for illustration, that is, the above data processing method may include:
[0111] S101. Obtain at least one abnormal data group, where the abnormal data group includes first abnormal data and at least one second abnormal data. The first abnormal data is data that fails to match similar abnormal data in a preset data set through an identification model, and the second abnormal data is abnormal data in the preset data set.
[0112] Among them, the abnormal data group may be a data group generated in real time based on the abnormalities that occur during vehicle operation, or at least one data group generated for the abnormalities that occur in the vehicle's historical time, that is, an abnormal data group will be generated corresponding to each time the vehicle moves abnormally.
[0113] Among them, the abnormal data may be data in an abnormal log generated when the vehicle has an abnormality. The data in the abnormal log includes but is not limited to the abnormal stack position, abnormal service number, etc., and can be specifically set according to requirements and will not be limited here.
[0114] It can be understood that due to different methods of generating abnormal logs, and there may be more than one invoker of different methods, there are differences in the abnormal stack positions or the abnormalities of the stack in the method in the generated abnormal logs; or, for a certain type of abnormality in terms of business meaning, similar abnormal logs are generated on different codes, but the abnormalities on different codes are actually the same, only the stack positions in the generated logs are different, and the occurrence of this situation may be because the relevant staff did not abstract or it is not easy to abstract a common method to design the log content; or, the output of the abnormal stack in the abnormal log accounts for a large proportion of the total output characters, and the abnormal stack hardly contains natural language information, and only a small amount of natural language information that can be printed by the relevant staff is contained in the abnormal log; or, the amount of data corresponding to the abnormal log is large, but the number of similar abnormal logs with labels is small, and most are abnormal log data without labels.
[0115] In this regard, in order to match similar data for the exception data generated in real time when the vehicle is running, an identification model can be introduced to match the similar data for the exception data generated by the vehicle through the identification model. Among them, the identification model can be used to find similar data of the exception data based on the exception data, and the identification model can also be used to extract features from the exception data to obtain the feature vector of the exception data, so as to prompt the next node to find similar data of the exception data based on the feature vector of the exception data.
[0116] Furthermore, since the neural network model requires a large amount of labeled data to train and learn the text characteristics of exception logs to adapt to the current similarity calculation task, but the amount of labeled data in the exception logs generated when the vehicle is abnormal is extremely small and is not enough to support the neural network model to learn the text characteristics of the exception logs. Or, due to the small amount of natural language information in the exception logs and the large number of exception stack texts without natural language information, the model cannot adapt to the characteristics of the current exception data.
[0117] In this regard, the terminal device can obtain at least one exception data group in real time or periodically to update the identification model in real time based on at least one exception data group, so that the updated model can learn the characteristics corresponding to the target exception data, so as to match similar exception data for the exception data similar to the target exception data.
[0118] Among them, the natural language information in the above exception stack can be understood as the descriptive text contained in the stack trace when the vehicle runs abnormally. This information can help relevant personnel understand the nature and source of the exception.
[0119] Among them, the labeled similar exception logs can be added with similar labels to the similar exception logs manually. For example, when the current exception log is reported to the platform for manual review, and after manual review, it is considered that the newly generated exception log is similar to a certain historical exception log, then the manual can associate these two logs, that is, the newly generated exception log can be labeled to indicate that it is similar to the corresponding historical exception log.
[0120] Among them, the above recognition model can be the SBERT (Sentence-BERT) model. This SBert model does not require additional preprocessing operations on the input data. When it accepts natural language sentences as a whole and tokenizes the input natural language sentences, the data cleaning process has been implemented, so there is no need for manual intervention. In contrast, other methods (such as the combination of MinHash and the LSH algorithm) need to perform data cleaning and standardization processing on the input data in order to achieve better results or accuracy, which consumes more preprocessing time. In addition, the above data processing method can also have the combination of MinHash and the LSH algorithm at the same time.
[0121] In some embodiments, the obtaining of at least one abnormal data group includes: detecting the vehicle data during vehicle operation to obtain the target abnormal data to be recognized generated when the vehicle is abnormal when the vehicle is detected to be abnormal, and then inputting the target abnormal data into the recognition model to match similar abnormal data for the target abnormal data in the above preset data set through the above recognition model to obtain a corresponding matching result, and this matching result is used to indicate whether the target abnormal data matches similar abnormal data.
[0122] Among them, if the above target abnormal data does not match similar abnormal data, the above target abnormal data is determined as the first abnormal data, that is, it means that the target abnormal data is not currently recognized by the recognition model. To match similar abnormal data based on the recognition result of the recognition model, the abnormal data that needs to be combined with other abnormal data to adjust the model parameters of the recognition model at present, so as to prompt the recognition model to learn the characteristics of the target abnormal data, that is, the second abnormal data combined with the above first abnormal data needs to be selected from the preset data set to obtain the abnormal data group.
[0123] Specifically, every time the vehicle has an abnormality, corresponding abnormal data can be generated for processing by the corresponding similar abnormality detection service, that is, the similar abnormality detection service can call the recognition model to receive the corresponding abnormal data to match similar abnormal data for the abnormal data.
[0124] Specifically, a data processing system can be set up. The data processing system includes a vehicle data acquisition module to stipulate the reporting protocol and service processing method of the abnormal data generated when the vehicle is abnormal in the vehicle data acquisition module, so as to obtain the target abnormal data to be recognized generated when the vehicle is abnormal and send the target abnormal data to the recognition model for processing.
[0125] In some embodiments, the matching of similar abnormal data for the target abnormal data in the preset data set through the recognition model may include: inputting the target abnormal data into the recognition model, obtaining a first feature vector of the target abnormal data through the recognition model, and based on the first feature vector, matching similar abnormal data for the target abnormal data in the preset data set.
[0126] Exemplarily, the representation form of the first feature vector may be: [0.15334, 0.23234,..., 0.23533].
[0127] Specifically, each time an abnormality occurs in the vehicle, the similarity detection service can call the recognition model to receive the corresponding abnormal data, perform text feature engineering on the abnormal data, that is, perform feature extraction, to calculate the embedding vector of the abnormal data, that is, the first feature vector, which is used to represent the understanding of the abnormal data by the recognition model.
[0128] Specifically, a data processing system can be set up. The data processing system includes a data representation module, and the data representation module includes a recognition model. The recognition model is used to represent the target abnormal data to generate a corresponding first feature vector. The first feature vector may include semantic word order information of the abnormal data, abnormal stack text information learned by the recognition model, etc., which can be specifically set according to requirements and will not be limited here.
[0129] In some embodiments, the preset data set is stored in a vector database. The preset data set includes second feature vectors corresponding to each abnormal data. The matching of similar abnormal data for the target abnormal data in the preset data set based on the first feature vector includes: inputting the first feature vector into the vector database; through the vector database, performing a similarity match between the first feature vector and the second feature vectors in the preset data set to obtain a matching result.
[0130] Among them, the vector database is a database specifically used for storing and querying vector data. The typical structure of the vector data in the vector database can be presented in the form of a one-dimensional array. Among them, the elements in the one-dimensional array are numerical values (usually floating-point numbers), and these numerical values can represent the positions, features, or attributes of objects or data points in a multi-dimensional space.
[0131] It should be noted that the vector database can support fast similarity search based on the vector distance or similarity of data, that is, it can find the most similar or relevant data according to semantic or context meanings, thus eliminating the need to rely on the query methods of traditional databases based on exact matching or predefined criteria.
[0132] In an embodiment, the vector database can quickly and accurately find similar abnormal data based on the first feature vector generated by the recognition model.
[0133] Among them, the above vector database includes but is not limited to Chroma database, Milvus database, Faiss database, Weaviate database, etc. The above Chroma database has 12.1K on Github Stars, supports horizontal scaling in terms of scalability, is extremely easy to use in terms of key features, can be developed on Jupyter Notebook, and is good at multimedia content, and is mostly applied to the audio field; the above Milvus database has 26.7K on Github Stars, supports both horizontal and vertical scaling in terms of scalability, uses in-memory storage and persistent storage in terms of key features to provide high-speed query and insertion performance, and provides automatic data partitioning and fault tolerance. The Milvus vector database has great advantages in terms of community activity, scalability, security and reliability, etc.; the above Faiss database has 28.1K on Github Stars, does not support scaling in terms of scalability, but supports GPU acceleration in terms of key features, and supports Flat index and high-quality search; the above Weaviate database has 9.5K on Github Stars, supports modularity in terms of scalability, provides GraphQL API in terms of key features, is suitable for graph-structured data interaction, and supports real-time data update.
[0134] Specifically, after using the vector database to perform similarity matching between the first feature vector and the second feature vector in the preset data set to obtain a matching result, the first feature vector can be stored in the preset data set in the vector database for subsequent abnormal data to perform similarity matching of abnormal data.
[0135] It should be noted that since the number of parameters of the recognition model itself is fixed, the first feature vector generated by the recognition model can be stored in the vector database. For example, the above vector database can be the Milvus database, which can support large-scale data processing, provide distributed storage for data, and have fast approximate vector retrieval capabilities, so as to enable this embodiment to provide better performance in terms of efficiency and stability during similar abnormal data matching, so as to avoid that as the input data volume increases, the data storage requirements, data insertion time, and matching time of similar data of the model may increase, and the storage cost will also increase accordingly.
[0136] It can be understood that after the processing of abnormal data in the historical period is completed, the above recognition model can learn the characteristics of the corresponding abnormal data, and the embedded vectors extracted from the abnormal data by the recognition model are also stored in the vector database at the same time, so that they can be directly applied at the vehicle terminal system.
[0137] Exemplarily, as Figure 2 shown, the model parameters of the recognition model can be updated in the cloud in advance using labeled historical data. The recognition model is constructed based on a pre-trained model. After obtaining the recognition model, the recognition model is pushed to the vehicle terminal so that unlabeled historical data can be input into the recognition model to prompt the recognition model to perform text feature expression on the unlabeled historical data to obtain the corresponding embedded vectors, and then the embedded vectors are input into the Milvus database for storage.
[0138] Optionally, the above-mentioned matching the first feature vector with the second feature vector in the above-mentioned preset data set through the above-mentioned vector database to obtain a matching result includes: calculating, through the above-mentioned vector database, the similarity between the first feature vector and the second feature vector in the above-mentioned preset data set; if there is no second feature vector in the above-mentioned preset data set whose similarity meets the preset similarity threshold, it is determined that the above-mentioned target abnormal data does not match similar abnormal data; if there is a second feature vector in the above-mentioned preset data set whose similarity meets the preset similarity threshold, it is determined that the above-mentioned target abnormal data matches similar abnormal data.
[0139] Among them, the method of calculating the similarity between the first feature vector and the second feature vector in the above-mentioned preset data set can be to calculate the distance between the first feature vector and the second feature vector through trigonometric functions to indicate the similarity. For example, if the first feature vector is vector a and the second feature vector is vector b), the similarity between the first feature vector and the second feature vector is cos(vector a, vector b).
[0140] Optionally, the above-mentioned selecting the second abnormal data combined with the above-mentioned first abnormal data from the preset data set to obtain an abnormal data group includes: sending the above-mentioned first abnormal data to a labeling platform to receive data feedback information of the labeling platform for the above-mentioned first abnormal data; based on the above-mentioned data feedback information, selecting the second abnormal data combined with the above-mentioned first abnormal data from the preset data set to obtain an abnormal data group.
[0141] Optionally, the above-mentioned selecting the second abnormal data combined with the above-mentioned first abnormal data from the preset data set to obtain an abnormal data group includes: based on the similarity between the above-mentioned first feature vector and the second feature vector in the above-mentioned preset data set, selecting the preset number of second abnormal data with the highest similarity from the preset data set to obtain an abnormal data group.
[0142] Optionally, selecting the second abnormal data that combines with the first abnormal data from the preset data set to obtain an abnormal data group includes: based on the similarity between the first feature vector and the second feature vector in the preset data set, selecting a preset number of second abnormal data with the highest similarity from the preset data set, and sending the first abnormal data and the selected second abnormal data to the annotation platform to receive the data feedback information of the annotation platform for the first abnormal data and the second abnormal data; based on the data feedback information, selecting the second abnormal data that combines with the first abnormal data from the preset number of second abnormal data to obtain an abnormal data group, and obtaining an abnormal data group.
[0143] Specifically, a data processing system can be set up. The data processing system also includes a similar abnormal data retrieval and update module. The similar abnormal data retrieval and update module includes a vector database. By using the vector database, the characterized embedding vector is used as input, and the distance is calculated between it and the embedding vector corresponding to the historical abnormal data stored in the vector database to obtain the similarity between the new abnormal data and the historical abnormal data. If the similarity does not meet the preset similarity threshold, it is considered that the new abnormal data does not match similar data. Then, a preset number of historical abnormal data with the highest similarity can be selected based on the similarity, and the preset number of historical abnormal data is reported to the corresponding platform via http for relevant staff to annotate. At the same time, the new abnormal data and the embedding vector are inserted into the vector database to update the Milvus database.
[0144] In some embodiments, if the target abnormal data matches similar abnormal data, the target abnormal data is excluded to filter out duplicate similar abnormal data, so as to retain valuable and high-quality non-duplicate new abnormal data, thereby helping relevant staff (such as R & D operation and maintenance personnel) quickly respond to and locate vehicle abnormal problems, improving the repair speed of vehicle abnormalities, and realizing the improvement of the user experience of vehicle users.
[0145] Exemplarily, as Figure 3 shown, when an abnormality occurs in the vehicle and abnormal data is generated at the vehicle end, the recognition model can be called through the similar abnormal detection service to perform text feature expression on the abnormal data to obtain an embedding vector. Then, the embedding vector can be inserted into the Milvus database, and similar vectors can be retrieved from the Milvus database for the embedding vector, such as obtaining several samples with the closest distance, to determine whether there is abnormal data with a similarity that meets the similarity threshold. If the judgment result is yes, it means that similar abnormal data is matched, and the abnormal data is determined as duplicate abnormal and excluded to end the processing flow.
[0146] S102. Obtain the target annotation information corresponding to the above abnormal data group, where the target annotation information is used to indicate the similarity between the above first abnormal data and the above second abnormal data.
[0147] In this embodiment, by obtaining the target annotation information of the obtained abnormal data group, the recognition model can be updated and adjusted based on the target annotation information to optimize the recognition ability of the recognition model for abnormal data.
[0148] Among them, the above similarity includes but is not limited to two situations: the first abnormal data and the second abnormal data are similar, and the first abnormal data and the second abnormal data are not similar.
[0149] Optionally, the target annotation information includes a similarity label and / or a dissimilarity label. The similarity label is used to indicate that the first abnormal data and the corresponding second abnormal data are similar, and the dissimilarity label is used to indicate that the first abnormal data and the corresponding second abnormal data are not similar.
[0150] In some embodiments, the obtaining of the target annotation information corresponding to the above abnormal data group includes: based on the above data feedback information, obtaining the target annotation information for the above first abnormal data, that is, manually annotating the abnormal data group through a annotation platform to obtain the corresponding target annotation information.
[0151] Specifically, a data processing system can be set up. The data processing system includes a similar text pair / dissimilar text pair storage module. According to the manual annotation results of the abnormal data group in the similar text pair / dissimilar text pair storage module, several pairs of similar and / or dissimilar data pairs can be obtained. These data pairs can be stored in a relational database so that corresponding data can be obtained from the relational database when needed.
[0152] Optionally, after obtaining the target annotation information corresponding to the above abnormal data group, it further includes: if it is determined based on the above target annotation information that there is no similar abnormal data for the above first abnormal data in the above preset data set, then report the above first abnormal data as an exception, such as Figure 4 shown, by Figure 4 reporting the first abnormal data, high-value abnormal data that is not repeated can be retained, so as to quickly locate and repair vehicle anomalies from it, and push and update the repaired data to the vehicle end to repair vehicle anomalies at the vehicle end and improve the owner experience.
[0153] S103. Based on the above abnormal data group and the above target annotation information, update the model parameters of the above recognition model to obtain an updated recognition model.
[0154] In this embodiment, based on the above abnormal data group and the above target annotation information, the parameters of the recognition model are adjusted to update the model parameters of the recognition model, so that the recognition model can learn the data characteristics of the new abnormal data, that is, the data characteristics of the first abnormal data.
[0155] It can be understood that since the abnormal data corresponding to the in-vehicle computer abnormal log is continuously generated as the vehicle runs, the characteristics of the data with existing manual annotations in the model still cannot directly judge similar abnormal data for some abnormal data. Therefore, for the abnormal data that cannot directly judge similar abnormal data, it can be handed over to manual judgment for duplicate removal, so as to form a high-quality labeled data set through manual annotation and store it in the relational database, so as to adjust the parameters of the recognition model based on this labeled data set, so that the recognition model can learn the data characteristics of the new abnormal data, so as to enable the recognition model to have the ability of continuous learning. Thus, through continuous learning, the recognition accuracy of similar abnormal data is improved, and on the contrary, the cost of manual annotation is reduced, prompting the recognition model and the manual annotation result to form a positive feedback closed loop.
[0156] It should be noted that in some scenarios, the characteristic of abnormal data is that the stack occupancy ratio is large, and the order and semantic information that can be learned from the stack itself are limited. This makes the effect of directly applying the recognition model for similarity calculation limited. By updating the parameters of the recognition model through the abnormal data group, the recognition model can learn the data characteristics of the abnormal data in the current scenario, so as to achieve a better effect when calculating the similarity for abnormal data similar to the abnormal data in the subsequent process.
[0157] It should be noted that since the above recognition model is updated based on the abnormal data generated during vehicle operation, the parameters of the above updated recognition model can be trained by two parts of data. Among them, one part comes from pre-training the recognition model on a large-scale general data corpus, and the other part comes from updating the recognition model through the abnormal data group in the current scenario, so as to enable the recognition model to not only understand and express a small amount of natural language information in the abnormal log, but also learn the characteristics of the abnormal stack text with a large proportion, so as to enable the recognition model to be effectively applicable to the similar data matching task in the current scenario and learn the text characteristics of the current scenario.
[0158] Among them, the general data corpus contains many types of data, such as people's dialogue language, posts, books, and people's annotation information on abnormal data, etc. For example, the general data corpus can be used to prompt the recognition model to learn the natural language information in the abnormal data, and an abnormal data group obtained from real-time applications can be used to prompt the recognition model to learn the characteristics of the abnormal stack text that has not been learned before.
[0159] Specifically, the method for adjusting the parameters of the recognition model based on the above abnormal data group and the above target annotation information includes, but is not limited to, global fine-tuning, local parameter fine-tuning, or fine-tuning using different specific frameworks, etc., which can be specifically set according to requirements and will not be limited here.
[0160] Specifically, the abnormal data set can be stored in a relational database so that it can be used as the input of the recognition model after a certain period to update the model parameters of the recognition model, so as to strengthen the feature expression of the recognition model for abnormal data and help the recognition model learn the data characteristics of new abnormal data. Each time the parameters of the recognition model are updated, only the newly generated abnormal data group after the previous model parameter update is taken as the input, realizing the incremental learning and updating ability of the model, that is, there is no need for a large number of training samples for training and supporting instant similar text retrieval, and it can learn the characteristics of the text in the current scenario, realizing the incremental learning and updating of the model.
[0161] Specifically, a data processing system can be set up, and the data processing system also includes a model update module to prompt the model update module to consume the abnormal data group regularly for update, improving the learning ability and accuracy of the model update module.
[0162] Specifically, the regular update strategy of the above model update module can include: a periodic update strategy, that is, periodically updating the recognition model; a loss-based update strategy, that is, deciding whether to update the recognition model according to the loss value predicted by the recognition model during the fine-tuning process. Only when the loss value is large, that is, due to the large similarity calculation error of the recognition model for the current new data, the model is updated.
[0163] Exemplarily, based on Figure 3 the example shown, when judging whether there is abnormal data whose similarity meets the similarity threshold, as Figure 5 shown, if the judgment result is negative, the abnormal data is reported to the cloud, that is, the above annotation platform, for manual annotation of similar / dissimilar texts. Then, the processing flow can be ended, or the manually annotated similar / dissimilar texts, that is, the abnormal data group, can be stored in a relational database to prompt the cloud to send data to the vehicle end to fine-tune the recognition model, that is, periodically update the fine-tuning model or update the fine-tuning model based on loss, to obtain the updated recognition model.
[0164] In some embodiments, updating the model parameters of the recognition model based on the above abnormal data group and the above target annotation information to obtain an updated recognition model includes: inputting the above abnormal data group into the recognition model, and obtaining, through the recognition model, target feature vectors corresponding to each abnormal data in the above abnormal data group; calculating a loss value based on each target feature vector in the above abnormal data group and the above target annotation information through a preset loss function; and updating the model parameters of the recognition model based on the calculated loss value to obtain an updated recognition model.
[0165] Specifically, a corresponding network structure can be configured for the recognition model, that is, the recognition model is constructed by applying the corresponding network structure, such as a siamese network structure, that is, the above recognition model is configured with two identical network structures, or a triplet network structure, that is, the above recognition model is configured with three identical network structures, so as to enable the corresponding network structure to be used in the process of updating the model parameters of the recognition model.
[0166] Among them, the weights of the model in the siamese network and the triplet network structure can be shared, and the benefits of sharing the weights of the model are as follows: 1. It can reduce the number of parameters and the amount of calculation; 2. It enhances the consistency of feature extraction and improves the performance of the model in similarity comparison tasks; 3. It focuses on learning similarity metrics.
[0167] In some embodiments, the above recognition model is configured with two identical network structures, the above second abnormal data includes one, and the step of inputting the above abnormal data group into the recognition model and obtaining, through the recognition model, target feature vectors corresponding to each abnormal data in the above abnormal data group includes: inputting the above first abnormal data and the above second abnormal data into a network structure respectively, and obtaining, through each network structure, the target feature vector of the above first abnormal data and the target feature vector of the above second abnormal data.
[0168] Correspondingly, the step of calculating a loss value based on each target feature vector in the above abnormal data group and the above target annotation information through a preset loss function includes: calculating a cosine similarity between the target feature vector of the above first abnormal data and the target feature vector of the above second abnormal data based on the similarity between the above first abnormal data and the above second abnormal data indicated by the above target annotation information through a preset loss function to obtain a loss value.
[0169] Exemplarily, as Figure 6 shown Figure 6As shown above, the recognition model is configured with two identical network structures. The recognition model is set as the SBert model. SentenceA is the first abnormal data, and SentenceB is the second abnormal data. After SentenceA and SentenceB are respectively calculated by SBert, their corresponding embedding vectors are generated. Then, the embedding vectors are pooled to obtain the semantic expressions u and v corresponding to SentenceA and SentenceB respectively. Here, W is the parameter weight matrix of the model. Then, the semantic expressions u and v are calculated through the cosine function, and the result is used as the value of the loss function expression in the current forward propagation of the recognition model. The loss function is as follows:
[0170] loss=cosine-sim(u,v)
[0171] Among them, the expression of the above loss function is used to indicate the calculation of the cosine similarity. After obtaining the loss value, the model parameters can be updated through backpropagation. It should be particularly noted that the weights of the two Berts in the siamese network are shared here.
[0172] In some embodiments, the recognition model is configured with three identical network structures. The second abnormal data includes two. The target annotation information indicates that the first abnormal data is similar to one second abnormal data, and the first abnormal data is not similar to the other second abnormal data.
[0173] Correspondingly, when the abnormal data group is input into the recognition model, the target feature vectors corresponding to each abnormal data in the abnormal data group are obtained through the recognition model, including: inputting the first abnormal data and each second abnormal data into a network structure respectively, and obtaining the target feature vector of the first abnormal data and the target feature vectors of each second abnormal data through each network structure respectively.
[0174] Correspondingly, the loss value is calculated based on each target feature vector in the abnormal data group and the target annotation information through a preset loss function, including: calculating the first vector distance between the first abnormal data and the similar second abnormal data, and the second vector distance between the first abnormal data and the dissimilar second abnormal data respectively through a preset loss function based on the target feature vector of the first abnormal data and the target feature vectors of each second abnormal data; calculating the loss value based on the first vector distance, the second vector distance and a preset distance difference parameter.
[0175] Exemplarily, as Figure 7 shown, when the input is a statement in the form of a triple, Figure 7The shown SBert model is configured with three identical network structures. SentenceA is the first abnormal data, SentenceB and SentenceC are the second abnormal data. The target annotation information corresponding to SentenceA and SentenceB is similar (positive), and the target annotation information corresponding to SentenceA and SentenceC is not similar (negative). After SentenceA, SentenceB, and SentenceC are respectively calculated by SBert, their corresponding embedding vectors are generated, and then the embeddings are passed through a pooling operation to obtain the semantic expressions u, v, and w corresponding to SentenceA, SentenceB, and SentenceC respectively. Then, the semantic expressions u, v, and w are calculated through a loss function to obtain the corresponding loss values. The loss function is shown as follows:
[0176] loss = max(d(u, v) - d(u, w) + margin, 0)
[0177] Among them, d(u, v) represents the distance between vectors u and v. Similarly, d(u, w) represents the distance between vectors u and w; margin is a hyperparameter, indicating how much difference there should be between d(u, v) and d(u, w). Margin is used to keep a minimum difference between d(u, v) and d(u, w).
[0178] It can be understood that since it is clear from the target annotation information that u and v are similar samples, and u and w are dissimilar samples, theoretically, d(u, v) should be made as small as possible, and d(u, w) should be made as large as possible. And it should be particularly noted that the weights of SBert in the triplet network here are shared.
[0179] It should be noted that this method has the ability of MinHash + LSH.
[0180] As can be seen from the above, by obtaining at least one abnormal data group, which includes first abnormal data and at least one second abnormal data, the first abnormal data is data for which no similar abnormal data is matched in the preset data set by the recognition model, and the second abnormal data is abnormal data in the preset data set. Then, obtain the target annotation information corresponding to the abnormal data group, and the target annotation information is used to indicate the similarity between the first abnormal data and the second abnormal data. Finally, based on the abnormal data group and the target annotation information, update the model parameters of the recognition model to obtain an updated recognition model, so as to realize real-time update of the recognition model, so that the updated model learns the characteristics corresponding to the target abnormal data, so as to realize matching similar abnormal data for abnormal data similar to the target abnormal data, so that when the vehicle generates abnormal log data, the recognition model is continuously updated to improve the recognition efficiency of the recognition model for similar data of the abnormal log data.
[0181] To facilitate better implementation of the data processing method provided by the embodiments of the present application, the embodiments of the present application further provide a device based on the above data processing method. The meanings of the nouns are the same as those in the above data processing method, and the specific implementation details can be referred to the description in the method embodiments.
[0182] For example, as Figure 8 shown, the data processing device may include:
[0183] A data acquisition module 801, configured to acquire at least one abnormal data group, the abnormal data group includes first abnormal data and at least one second abnormal data, the first abnormal data is data for which no similar abnormal data is matched in the preset data set by the recognition model, and the second abnormal data is abnormal data in the preset data set;
[0184] An information acquisition module 802, configured to acquire the target annotation information corresponding to the abnormal data group, and the target annotation information is used to indicate the similarity between the first abnormal data and the second abnormal data;
[0185] A model update module 803, configured to update the model parameters of the recognition model based on the abnormal data group and the target annotation information to obtain an updated recognition model.
[0186] In an embodiment of the present application, the data acquisition module 801 includes:
[0187] Acquire the target abnormal data to be recognized generated when the vehicle is abnormal;
[0188] Through the recognition model, match similar abnormal data for the target abnormal data in the preset data set;
[0189] If no similar abnormal data is matched for the above target abnormal data, then determine the above target abnormal data as the first abnormal data;
[0190] Select second abnormal data combined with the above first abnormal data from the preset data set to obtain an abnormal data group.
[0191] Optionally, the above data acquisition module 801 includes:
[0192] Input the above target abnormal data into the above recognition model, and obtain the first feature vector of the above target abnormal data through the above recognition model;
[0193] Based on the above first feature vector, match similar abnormal data for the above target abnormal data in the above preset data set.
[0194] Optionally, the above preset data set is stored in a vector database, the above preset data set includes second feature vectors corresponding to each abnormal data, and the above data acquisition module 801 includes:
[0195] Input the above first feature vector into the above vector database;
[0196] Through the above vector database, perform similarity matching between the above first feature vector and the second feature vectors in the above preset data set to obtain a matching result.
[0197] Optionally, the above data acquisition module 801 includes:
[0198] Through the above vector database, calculate the similarity between the above first feature vector and the second feature vectors in the above preset data set;
[0199] If there is no second feature vector in the above preset data set whose similarity meets the preset similarity threshold, then determine that no similar abnormal data is matched for the above target abnormal data;
[0200] If there is a second feature vector in the above preset data set whose similarity meets the preset similarity threshold, then determine that similar abnormal data is matched for the above target abnormal data.
[0201] Optionally, the above data acquisition module 801 includes:
[0202] Send the above first abnormal data to the annotation platform to receive data feedback information of the annotation platform for the above first abnormal data;
[0203] Based on the above data feedback information, select second abnormal data combined with the above first abnormal data from the preset data set to obtain an abnormal data group.
[0204] Optionally, the above information acquisition module 802 includes:
[0205] Based on the above data feedback information, obtain the target annotation information for the above first abnormal data.
[0206] Optionally, the above communication device further includes a data elimination module, and the above data elimination module includes:
[0207] If the above target abnormal data matches similar abnormal data, then eliminate the above target abnormal data.
[0208] Optionally, the above target annotation information includes a similar label and / or a dissimilar label. The similar label is used to indicate that the above first abnormal data and the corresponding second abnormal data are similar, and the dissimilar label is used to indicate that the above first abnormal data and the corresponding second abnormal data are not similar.
[0209] Optionally, the above model update module 803 includes:
[0210] Input the above abnormal data group into the recognition model, and obtain the target feature vectors corresponding to each abnormal data in the above abnormal data group through the above recognition model;
[0211] Calculate the loss value based on each target feature vector in the above abnormal data group and the above target annotation information through a preset loss function;
[0212] Update the model parameters of the recognition model based on the calculated loss value to obtain an updated recognition model.
[0213] Optionally, the above recognition model is configured with two identical network structures, the above second abnormal data includes one, and the above model update module 803 includes:
[0214] Input the above first abnormal data and the above second abnormal data into a network structure respectively, and obtain the target feature vector of the above first abnormal data and the target feature vector of the above second abnormal data through each network structure respectively;
[0215] Calculate the cosine similarity of the target feature vector of the above first abnormal data and the target feature vector of the above second abnormal data based on the similarity between the above first abnormal data and the above second abnormal data indicated by the above target annotation information through a preset loss function to obtain a loss value.
[0216] Optionally, the above recognition model is configured with three identical network structures, the above second abnormal data includes two, and the above target annotation information indicates that the above first abnormal data is similar to one second abnormal data, and the above first abnormal data is not similar to another second abnormal data;
[0217] The above model update module 803 includes:
[0218] Input the above first abnormal data and each of the above second abnormal data into a network structure respectively, and obtain the target feature vectors of the above first abnormal data and the target feature vectors of each of the above second abnormal data through each network structure;
[0219] Through a preset loss function, based on the target feature vectors of the above first abnormal data and the target feature vectors of each of the above second abnormal data, calculate the first vector distance between the above first abnormal data and the similar second abnormal data respectively, and the second vector distance between the above first abnormal data and the dissimilar second abnormal data;
[0220] Calculate the loss value based on the above first vector distance, the above second vector distance and a preset distance difference parameter.
[0221] The data processing device proposed in this application obtains at least one abnormal data group, where the abnormal data group includes first abnormal data and at least one second abnormal data. The first abnormal data is the data that no similar abnormal data is matched in the preset data set by the recognition model, and the second abnormal data is the abnormal data in the above preset data set. Then, obtain the target annotation information corresponding to the above abnormal data group, and the above target annotation information is used to indicate the similarity between the above first abnormal data and the above second abnormal data. Finally, based on the above abnormal data group and the above target annotation information, update the model parameters of the above recognition model to obtain an updated recognition model, so as to realize real-time update of the recognition model, so that the updated model can learn the features corresponding to the target abnormal data, so as to realize matching similar abnormal data for abnormal data similar to the target abnormal data, so that when the vehicle generates abnormal log data, the recognition model is continuously updated to improve the recognition efficiency of the recognition model for similar data of the abnormal log data.
[0222] In specific implementation, each of the above modules can be implemented as an independent entity, or can be combined arbitrarily to be implemented as the same or several entities. The specific implementation manners of each of the above modules and the corresponding beneficial effects can be referred to the above method embodiments and will not be elaborated here.
[0223] An embodiment of this application provides a data processing system, including a vehicle and an electronic device;
[0224] The above vehicle is used to generate first abnormal data and send the above first abnormal data to the electronic device;
[0225] The above electronic device is used to execute the steps of any data processing method provided in the embodiments of this application for the above first abnormal data.
[0226] The embodiment of the present application further provides an electronic device, such as Figure 9 shown, which shows a schematic structural diagram of the electronic device involved in the embodiment of the present application. Specifically:
[0227] The electronic device may include a processor 901 with one or more processing cores, a memory 902 with one or more computer-readable storage media, a power supply 903, an input unit 904 and other components. Those skilled in the art can understand that Figure 9 the structure of the electronic device shown in
[0228] does not limit the electronic device, and it may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements. Among them:
[0229] The processor 901 is the control center of the electronic device, connecting various parts of the entire electronic device through various interfaces and lines. By running or executing computer programs and / or modules stored in the memory 902, and by calling data stored in the memory 902, it executes various functions of the electronic device and processes data. Optionally, the processor 901 may include one or more processing cores; preferably, the processor 901 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 901 either.
[0230] The electronic device further includes a power supply 903 for supplying power to each component. Preferably, the power supply 903 can be logically connected to the processor 901 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 903 may also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.
[0231] The electronic device may further include an input unit 904, which may be configured to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.
[0232] Although not shown, the electronic device may further include a display unit and the like, which will not be elaborated here. Specifically, in this embodiment, the processor 901 in the electronic device will load the executable files corresponding to the processes of one or more computer programs into the memory 902 according to the following instructions, and the processor 901 will run the computer programs stored in the memory 902 to implement various functions, such as:
[0233] Obtain at least one abnormal data group, where the abnormal data group includes first abnormal data and at least one second abnormal data. The first abnormal data is data for which no similar abnormal data is matched in the preset data set through the recognition model, and the second abnormal data is abnormal data in the preset data set;
[0234] Obtain the target annotation information corresponding to the abnormal data group, where the target annotation information is used to indicate the similarity between the first abnormal data and the second abnormal data;
[0235] Based on the abnormal data group and the target annotation information, update the model parameters of the recognition model to obtain an updated recognition model.
[0236] For the specific implementation manners of the above operations and the corresponding beneficial effects, reference may be made to the detailed description of the data processing method above, which will not be elaborated here.
[0237] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by a computer program, or by a computer program controlling related hardware. The computer program can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0238] Therefore, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored, and the computer program can be loaded by a processor to execute the steps in any of the data processing methods provided by the embodiments of the present application. For example, the computer program can execute the following steps:
[0239] Obtain at least one abnormal data group, where the abnormal data group includes first abnormal data and at least one second abnormal data. The first abnormal data is data for which no similar abnormal data is matched in the preset data set through the recognition model, and the second abnormal data is abnormal data in the preset data set;
[0240] Obtain the target annotation information corresponding to the above abnormal data group, where the above target annotation information is used to indicate the similarity between the above first abnormal data and the above second abnormal data;
[0241] Based on the above abnormal data group and the above target annotation information, update the model parameters of the above recognition model to obtain an updated recognition model.
[0242] For the specific implementation manners of the above various operations and the corresponding beneficial effects, reference may be made to the previous embodiments, which will not be elaborated herein.
[0243] Among them, the computer-readable storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, etc.
[0244] Since the computer program stored in the computer-readable storage medium can execute the steps in any of the data processing methods provided in the embodiments of the present application, the beneficial effects that can be achieved by any of the data processing methods provided in the embodiments of the present application can be realized. For details, reference may be made to the previous embodiments, which will not be elaborated herein.
[0245] Among them, according to one aspect of the present application, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to enable the computer device to execute the above data processing method.
[0246] The above has introduced in detail a data processing method, device, equipment, medium, product, vehicle and system provided by the embodiments of the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A data processing method, characterized in that: The method comprises: Acquire at least one abnormal data group, the abnormal data group including first abnormal data and at least one second abnormal data, the first abnormal data is data that is not matched to similar abnormal data in a preset data set by a recognition model, and the second abnormal data is abnormal data in the preset data set; Acquire target labeling information corresponding to the abnormal data group, where the target labeling information is used to indicate similarities between the first abnormal data and the second abnormal data; Based on the abnormal data group and the target labeling information, the model parameters of the recognition model are updated to obtain an updated recognition model.
2. The data processing method according to claim 1, characterized in that: The obtaining of at least one abnormal data group comprises: Acquire target abnormal data to be identified when a vehicle is abnormal; Matching similar abnormal data for the target abnormal data in the preset data set by using the recognition model; If the target abnormal data does not match similar abnormal data, the target abnormal data is determined as the first abnormal data; Second abnormal data combined with the first abnormal data is selected from a preset data set to obtain an abnormal data group.
3. The data processing method according to claim 2, characterized in that: The matching of similar abnormal data for the target abnormal data in the preset data set by using the recognition model includes: Inputting the target abnormal data into the recognition model, and obtaining a first feature vector of the target abnormal data through the recognition model; Based on the first feature vector, similar abnormal data is matched for the target abnormal data in the preset data set.
4. The data processing method according to claim 3, characterized in that: The preset data set is stored in a vector database, the preset data set includes a second feature vector corresponding to each abnormal data, and matching similar abnormal data for the target abnormal data in the preset data set based on the first feature vector includes: inputting the first feature vector into the vector database; The first feature vector is similarly matched with the second feature vector in the preset data set through the vector database to obtain a matching result.
5. The data processing method according to claim 4, characterized in that: The similarity matching of the first feature vector with the second feature vector in the preset data set through the vector database to obtain a matching result includes: Calculating the similarity between the first feature vector and a second feature vector in the preset data set through the vector database; If there is no second feature vector in the preset data set whose similarity meets the preset similarity threshold, it is determined that the target abnormal data is not matched with similar abnormal data; If there is a second feature vector in the preset data set whose similarity meets a preset similarity threshold, it is determined that the target abnormal data matches the similar abnormal data.
6. The data processing method according to claim 2, characterized in that: The step of selecting second abnormal data combined with the first abnormal data from a preset data set to obtain an abnormal data group includes: Sending the first abnormal data to a labeling platform to receive data feedback information of the labeling platform for the first abnormal data; Based on the data feedback information, second abnormal data combined with the first abnormal data is selected from a preset data set to obtain an abnormal data group.
7. The data processing method according to claim 6, characterized in that: The obtaining target labeling information corresponding to the abnormal data group includes: Based on the data feedback information, target labeling information for the first abnormal data is obtained.
8. The data processing method according to claim 2, characterized in that: The method further comprises: If the target abnormal data matches similar abnormal data, the target abnormal data is removed.
9. The data processing method according to claim 1, characterized in that: The target annotation information includes a similar label and / or a dissimilar label, wherein the similar label is used to indicate that the first abnormal data and the corresponding second abnormal data are similar, and the dissimilar label is used to indicate that the first abnormal data and the corresponding second abnormal data are not similar.
10. The data processing method according to any one of claims 1 to 9, characterized in that: The updating of the model parameters of the recognition model based on the abnormal data group and the target annotation information to obtain an updated recognition model includes: Inputting the abnormal data group into a recognition model, and obtaining target feature vectors corresponding to each abnormal data in the abnormal data group through the recognition model; Calculating the loss value based on each target feature vector in the abnormal data group and the target annotation information through a preset loss function; The model parameters of the recognition model are updated based on the calculated loss value to obtain an updated recognition model.
11. The data processing method according to claim 10, characterized in that: The recognition model is configured with two identical network structures, the second abnormal data includes one, the abnormal data group is input into the recognition model, and the target feature vector corresponding to each abnormal data in the abnormal data group is obtained by the recognition model, including: Inputting the first abnormal data and the second abnormal data into a network structure respectively, and obtaining a target feature vector of the first abnormal data and a target feature vector of the second abnormal data respectively through each network structure; The calculation of the loss value based on each target feature vector in the abnormal data group and the target annotation information by using a preset loss function includes: By using a preset loss function, based on the similarity between the first abnormal data and the second abnormal data indicated by the target annotation information, a cosine similarity calculation is performed on the target feature vector of the first abnormal data and the target feature vector of the second abnormal data to obtain a loss value.
12. The data processing method according to claim 10, characterized in that: The recognition model is configured with three identical network structures, the second abnormal data includes two, the target annotation information indicates that the first abnormal data is similar to one second abnormal data, and the first abnormal data is not similar to another second abnormal data; The step of inputting the abnormal data group into a recognition model and obtaining target feature vectors corresponding to each abnormal data in the abnormal data group through the recognition model includes: Inputting the first abnormal data and each of the second abnormal data into a network structure respectively, and obtaining a target feature vector of the first abnormal data and a target feature vector of each of the second abnormal data respectively through each network structure; The calculation of the loss value based on each target feature vector in the abnormal data group and the target annotation information by using a preset loss function includes: By using a preset loss function, based on the target feature vector of the first abnormal data and the target feature vector of each of the second abnormal data, respectively calculate a first vector distance between the first abnormal data and similar second abnormal data, and a second vector distance between the first abnormal data and dissimilar second abnormal data; The loss value is calculated based on the first vector distance, the second vector distance and a preset distance difference parameter.
13. A data processing device, characterized in that: The device comprises: A data acquisition module, used to acquire at least one abnormal data group, wherein the abnormal data group includes first abnormal data and at least one second abnormal data, wherein the first abnormal data is data that is not matched to similar abnormal data in a preset data set by a recognition model, and the second abnormal data is abnormal data in the preset data set; An information acquisition module, used to acquire target annotation information corresponding to the abnormal data group, wherein the target annotation information is used to indicate the similarity between the first abnormal data and the second abnormal data; The model updating module is used to update the model parameters of the recognition model based on the abnormal data group and the target annotation information to obtain an updated recognition model.
14. An electronic device, characterized in that: The invention comprises one or more processors and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the data processing method according to any one of claims 1 to 12.
15. A storage medium, characterized in that: The invention comprises a computer program, and when the computer program is run on a controller, the computer program is used to make the controller execute the steps of the data processing method according to any one of claims 1 to 12.
16. A computer program product, characterized in that The method comprises a computer program or an instruction, which implements the steps of the data processing method according to any one of claims 1 to 12 when the computer program or the instruction is executed by a processor.
17. A vehicle, characterized in that: The vehicle includes the electronic device as claimed in claim 14.
18. A data processing system, characterized in that: This includes vehicles and electronic equipment; The vehicle is used to generate first abnormal data and send the first abnormal data to the electronic device; The electronic device is used to perform the steps of the data processing method according to any one of claims 1 to 12 on the first abnormal data.