Sample generation method and device, relation discrimination method and device and storage medium

By generating simulated negative samples to enrich the training set, the problem of low efficiency in identifying pairwise relationships of identifiers in vehicle diagnostic communication is solved, and high accuracy and stability in identification under multiple vehicle models and complex network topologies are achieved.

CN120929914APending Publication Date: 2025-11-11LAUNCH TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511039188.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

In existing vehicle diagnostic communication technologies, the determination of pairwise relationships between identifiers relies on limited training data, resulting in low efficiency and difficulty in adapting to the application requirements of multiple vehicle models, multiple platforms, and complex network topologies.

Method used

By obtaining real identifier pairs as the first samples, multiple simulated negative samples are generated based on the vehicle communication protocol to enrich the diversity of the training set. The generative model is then used to select negative sample pairs with high similarity for training, thereby improving the generalization ability of the relationship discrimination model.

Benefits of technology

It improves the accuracy and robustness of the relationship discrimination model in complex communication environments, and can identify pairwise relationships between different vehicle models and ECUs, thereby improving the stability and success rate of vehicle diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120929914A_ABST
    Figure CN120929914A_ABST
Patent Text Reader

Abstract

The invention relates to a sample generation method, a relation discrimination method and device and a storage medium, and the method comprises the steps: obtaining at least one first sample corresponding to a to-be-trained relation discrimination model, the first sample comprises a vehicle identifier and a first identifier pair, and the first identifier pair comprises two identifiers, the relation discrimination model is used for identifying whether the two identifiers in the first sample conform to a pairwise relation corresponding to the vehicle communication protocol; based on at least one negative sample generation rule corresponding to the vehicle communication protocol, the first identifier pair is processed to generate at least one second identifier pair, and the second identifier pair is different from the first identifier pair; and generating a second sample for training the relationship discrimination model based on the second identifier pair. According to the method, at least one second sample is generated through the first sample, and good generalization of the model to different identifier pairs is realized under the condition that the negative samples of the training data are few, so that the discrimination capability of the relation discrimination model to the negative samples is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of vehicle technology, and in particular to a sample generation method, a relationship discrimination method, an apparatus, and a storage medium. Background Technology

[0002] In vehicle diagnostic communication, there is a fixed pairwise relationship between the identifiers carried in the diagnostic request and the identifiers carried in the response message. For example, the diagnostic request sent by the diagnostic tool has a specific request ID, and the response message from the device in the vehicle also has a corresponding response ID. Generally, machine learning models can be used to determine the pairwise relationship of identifiers. Related techniques typically involve extracting paired identifiers from existing configuration files or diagnostic files and manually labeling them to obtain training data. However, the amount of training data obtained is limited, and the efficiency is low. Summary of the Invention

[0003] This application provides a sample generation method, a relationship discrimination method, an apparatus, and a storage medium to solve the above-mentioned problems.

[0004] To achieve the above objectives, according to a first aspect of this application, a sample generation method is provided, the method comprising:

[0005] At least one first sample corresponding to the relationship discrimination model to be trained is obtained. The first sample includes a vehicle identifier and a first identifier pair. The first identifier pair includes two identifiers. The relationship discrimination model is used to identify whether the two identifiers in the first sample conform to the pair relationship corresponding to the vehicle communication protocol. The pair relationship indicates that one of the two identifiers is an identifier carried in a diagnostic request sent to the vehicle indicated by the vehicle identifier, and the other identifier is an identifier carried in the response message of the vehicle to the diagnostic request.

[0006] Based on at least one negative sample generation rule corresponding to the vehicle communication protocol, the first identifier pair is processed to generate at least one second identifier pair, the second identifier pair being different from the first identifier pair;

[0007] Based on the second identifier pair, a second sample is generated for training the relation discrimination model.

[0008] Optionally, the negative sample generation rule includes a replacement rule.

[0009] The first identifier pair is processed based on at least one negative sample generation rule corresponding to the vehicle communication protocol to generate at least one second identifier pair, including:

[0010] Based on at least one identifier in the first identifier pair, generate a replacement identifier that is different from the identifier;

[0011] Based on the obtained at least one replacement identifier, the at least one identifier is replaced to obtain the second identifier pair.

[0012] Optionally, generating a replacement identifier different from the identifier based on at least one identifier in the first identifier pair includes:

[0013] Obtain a functionally addressed identifier with the same length as the at least one identifier, and use it as the replacement identifier.

[0014] Optionally, generating a replacement identifier different from the identifier based on at least one identifier in the first identifier pair includes:

[0015] A sample identifier is randomly selected from the sample set corresponding to the vehicle identifier as the replacement identifier. The sample identifier is an identifier of the same type as either the identifier carried in the diagnostic request or the identifier carried in the response message.

[0016] The step of replacing the at least one obtained replacement identifier to obtain the second identifier pair includes:

[0017] The second identifier pair is obtained by replacing the corresponding type of identifier in the first identifier pair with the sample identifier.

[0018] Optionally, the two identifiers include a first identifier and a second identifier, wherein one of the first identifier and the second identifier is an identifier carried in the diagnostic request, and the other is an identifier carried in the response message;

[0019] The step of generating a replacement identifier different from the first identifier pair based on at least one identifier in the first identifier pair includes:

[0020] Based on the length of the identifier, generate at least one replacement identifier with the same length but different content from the identifier; and / or

[0021] Randomly select at least one replacement identifier with the same length as the identifier from the sample set corresponding to other vehicle identifiers;

[0022] The step of replacing the at least one obtained replacement identifier to obtain the second identifier pair includes:

[0023] Based on the two replacement identifiers obtained, the first identifier and the second identifier are replaced respectively to obtain the second identifier pair.

[0024] Optionally, generating a replacement identifier different from the identifier based on at least one identifier in the first identifier pair includes:

[0025] Determine the target identifier from the two identifiers;

[0026] The replacement identifier is obtained by flipping the contents of the preset mask bits in the target identifier.

[0027] Optionally, the first sample includes the lengths corresponding to the two identifiers respectively, and the negative sample generation rule includes a length adjustment rule and / or a position swapping rule, wherein,

[0028] The length adjustment rules include:

[0029] Determine the target identifier and the length of the target identifier from the two identifiers in the first sample;

[0030] Adjust the length of the target identifier to obtain the identifier content after adjustment;

[0031] Based on the adjusted length of the identifier content, the target identifier in the first sample is replaced to generate the second identifier pair;

[0032] The location exchange rules include:

[0033] The positions of the identifier carried in the diagnostic request and the identifier carried in the response message in the first sample are changed to generate the second identifier pair.

[0034] According to a second aspect of this application, embodiments of this application also provide a relationship determination method, the method comprising:

[0035] Obtain identifier pair information to be judged, the identifier pair information including the identifier carried in the diagnostic request to be judged, the identifier carried in the response message to be judged, and the length of each of the identifiers;

[0036] The identifier to be judged is input into the trained relation discrimination model to obtain the discrimination result, which is used to indicate the degree of matching between the diagnostic request to be judged and the response message to be judged.

[0037] According to a third aspect of this application, embodiments of this application also provide a sample generation apparatus, the apparatus comprising:

[0038] The acquisition module is used to acquire at least one first sample corresponding to the relationship discrimination model to be trained. The first sample includes a vehicle identifier and a first identifier pair. The first identifier pair includes two identifiers. The relationship discrimination model is used to identify whether the two identifiers in the first sample conform to the pairing relationship corresponding to the vehicle communication protocol. The pairing relationship indicates that one of the two identifiers is an identifier carried in a diagnostic request sent to the vehicle indicated by the vehicle identifier, and the other identifier is an identifier carried in the response message of the vehicle to the diagnostic request.

[0039] The first generation module is used to process the first identifier based on at least one negative sample generation rule corresponding to the vehicle communication protocol to generate at least one second identifier pair, wherein the second identifier pair is different from the first identifier pair.

[0040] The second generation module is used to generate a second sample for training the relation discrimination model based on the second identifier pair.

[0041] According to a fourth aspect of this application, embodiments of this application also provide a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of any of the methods provided in embodiments of this application.

[0042] Some embodiments of this specification include at least the following beneficial effects: at least one second sample is generated from the first sample; the relationship discrimination model can be trained based on the first and second samples; the model achieves good generalization of different identifier pairs when there are few negative samples in the training data, thereby improving the relationship discrimination model's ability to discriminate negative samples and can be applied to various vehicle diagnostic scenarios; the trained relationship discrimination model can automatically identify the pairwise relationship of identifier pairs, thereby improving the robustness and stability of vehicle diagnosis.

[0043] Other features and advantages of this application will be described in detail in the following detailed description section. Attached Figure Description

[0044] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0045] To gain a more complete understanding of this application and its beneficial effects, the following description will be provided in conjunction with the accompanying drawings, wherein the same reference numerals in the following description denote the same parts.

[0046] Figure 1 These are application scenario diagrams illustrating the sample generation method according to some embodiments of this specification;

[0047] Figure 2 This is an exemplary flowchart of a sample generation method according to some embodiments of this specification;

[0048] Figure 3 This is an exemplary flowchart of a relationship determination method according to some embodiments of this specification;

[0049] Figure 4 This is a schematic diagram of the sample generation apparatus shown in some embodiments of this specification;

[0050] Figure 5 This is a schematic diagram of the structure of an electronic device according to some embodiments of this specification. Detailed Implementation

[0051] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the protection scope of this application.

[0052] To facilitate understanding of the implementation schemes provided in this application, the relevant application background of the sample generation method provided in this application will be explained first.

[0053] Currently, in automotive electronic systems, the Controller Area Network (CAN) bus is typically used as the primary communication protocol to enable data exchange between various Electronic Control Units (ECUs). In vehicle diagnostic communication, diagnostic request messages are sent by vehicle diagnostic equipment (such as a diagnostic tool with an OBD interface) to a specific ECU; while diagnostic response messages are returned by the addressed ECU, containing the execution result or a negative response. Related diagnostic equipment and gateway devices rely on predefined configuration files or manual input to determine CANID mapping relationships, lacking the ability for automatic learning or dynamic recognition, making it difficult to adapt to the application requirements of multiple vehicle models, multiple platforms, and complex network topologies.

[0054] In view of this, some embodiments of this specification provide a sample generation method that extracts a first sample (i.e., a real-world identifier pair) and generates multiple second samples (i.e., simulated negative samples) based on the first sample, thereby enriching the diversity of the training set. By introducing diverse negative samples, the relationship discrimination model can better distinguish between real and abnormal pairwise relationships, improve its discrimination accuracy in complex communication environments, and make the relationship discrimination model no longer limited to the data distribution of a specific vehicle model. It can learn the association between more general identifier pairs and has high recognition performance for different vehicle models and different ECUs.

[0055] Figure 1 This is an application scenario diagram of the sample generation method shown in some embodiments of this specification.

[0056] The sample generation method of this application can be applied to application scenarios such as terminal communication relationship identification, vehicle diagnostic systems, configuration file anomaly detection, and vehicle log anomaly detection. For example, in a vehicle diagnostic system, the relationship discrimination model can identify the CAN ID pairing relationship between different ECUs and vehicle diagnostic equipment, thereby enabling functions such as replacing flow control frames and replacing negative responses, improving the success rate of remote diagnostics. As another example, in configuration file anomaly detection, the relationship discrimination model can identify the request-response pairing relationship between devices recorded in the configuration file, improving the correctness of the configuration file.

[0057] The execution subject of this application embodiment can be an electronic device, which can be an electronic device, a server, a terminal, or other devices.

[0058] The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, and cloud computing. Terminal devices can include, but are not limited to, mobile phones, computers, intelligent voice interaction devices, and industrial control computers. Terminal devices and servers can be directly or indirectly connected via wired or wireless communication methods; this embodiment does not impose any limitations on this.

[0059] The terminal device (which may be referred to as the first device) can also communicate with another terminal device (which may be referred to as the second device) to obtain a relation discrimination model trained by the second device. The second device can be used to: obtain at least one first sample corresponding to the relation discrimination model to be trained, the first sample including a vehicle identifier and a first identifier pair, the first identifier pair including two identifiers; process the first identifier pair based on at least one negative sample generation rule corresponding to the vehicle communication protocol to generate at least one second identifier pair, the second identifier pair being different from the first identifier pair; and generate a second sample for training the relation discrimination model based on the second identifier pair. Further details can be found in the relevant description below.

[0060] It is understandable that the terminal device used to train the model and the terminal device used to apply the model can be the same or different. For example, a model trained via a second device can be applied to different systems or devices (e.g., the first device). For instance, the first device can be integrated into various mobile devices (wheeled construction equipment, autonomous vehicles, driver-assisted vehicles, etc.), and autonomous vehicles can also be cars, trucks, motorcycles, buses, ships, airplanes, helicopters, lawnmowers, recreational vehicles, amusement park vehicles, construction equipment, trams, golf carts, trains, and handcarts, etc.

[0061] In some embodiments, the first device includes a processing module, which can be used to execute the trained relation discrimination model provided in the embodiments of this application.

[0062] In some embodiments, the application scenario may also include, for example, networks, storage devices, etc. Networks may include any suitable wired or wireless networks that facilitate the exchange of information and / or data. Storage devices are used to store data, instructions, and / or any other information.

[0063] It is important to note that the application scenarios of the sample generation method are provided for illustrative purposes only and are not intended to limit the scope of this specification. Those skilled in the art can make various changes and modifications based on the description in this specification. For example, application scenarios may also include databases, information sources, etc. Furthermore, application scenarios may be implemented on other devices to achieve matching or different functions. However, these changes and modifications will not depart from the scope of this specification.

[0064] Figure 2 This is an exemplary flowchart of a sample generation method according to some embodiments of this specification. In some embodiments, process 200 may be performed on an electronic device. Figure 2 As shown, process 200 includes the following steps.

[0065] Step 210: Obtain at least one first sample corresponding to the relation discrimination model to be trained. The first sample includes a vehicle identifier and a first identifier pair.

[0066] The first identifier pair includes two identifiers. The relationship discrimination model is used to identify whether the two identifiers in the first sample conform to the pairing relationship corresponding to the vehicle communication protocol. The pairing relationship indicates that one of the two identifiers is the identifier carried in the diagnostic request sent to the vehicle indicated by the vehicle identifier, and the other identifier is the identifier carried in the response message of the vehicle to the diagnostic request.

[0067] Vehicle communication protocols refer to a set of rules used to standardize communication formats, data structures, and other aspects of communication within a vehicle. Examples of vehicle communication protocols include the CAN communication protocol and DOIP (Diagnostic Over Internet Protocol, an Ethernet-based diagnostic communication protocol).

[0068] Vehicle identification refers to various codes, symbols, or markings used to identify and distinguish vehicles. For example, vehicle identification may contain information such as the vehicle's manufacturer, model, configuration, and year of manufacture. Exemplarily, vehicle identification includes the Vehicle Identification Number (VIN), which is a code composed of numbers and letters used to identify and distinguish each vehicle.

[0069] A pairwise relationship can be a mapping between the identifier carried in the diagnostic request sent by the requesting terminal and the identifier carried in the response message sent by the responding terminal during communication interaction.

[0070] An identifier is a character or code used in a communication protocol to uniquely identify the source or destination of a message. For example, in vehicle diagnostic communication, an identifier could be the CAN ID of a request message sent by a requesting terminal or the CAN ID of a response message sent by a responding terminal. Identifiers can also be IP addresses, logical addresses, etc.

[0071] A CAN ID is an identifier used in the CAN bus protocol to uniquely identify each message. A CAN ID can represent the message type and the source or destination terminal of a message, among other things.

[0072] For example, in an on-board diagnostic system, the diagnostic tool sends a diagnostic request using CAN ID 0x7E0, and the electronic control unit (ECU) returns a response message using CAN ID 0x7E8. Thus, (0x7E0, 0x7E8) constitutes a first identifier pair.

[0073] A requesting terminal refers to the terminal that initiates the communication request. For example, a requesting terminal may include vehicle diagnostic equipment (such as a diagnostic tool connected to an OBD interface), a remote server (such as a remote cloud diagnostic platform), etc.

[0074] A response terminal is a terminal that receives requests, performs operations, and returns results. For example, a response terminal may include various ECUs in a vehicle (such as engine ECU, body ECU, battery management ECU, etc.), or sensors or actuators in an industrial control system.

[0075] The first sample is at least a portion of the training data used to train the relation discrimination model. In some embodiments, the first sample is a positive sample. A positive sample is a training sample that contains pairs of identifiers that correctly identify the relation. By training the model with positive samples, the model can learn the identifiers carried in diagnostic requests and response messages during normal communication.

[0076] In some embodiments, at least one first sample can be obtained in various ways. For example, the first sample can be obtained through actual communication data, communication protocol specifications, or manually labeled data. For instance, actual communication data in the vehicle communication network can be captured using diagnostic tools or network analysis devices. For example, when communicating with the vehicle ECU using an OBD-II scanning tool, the CAN ID of the diagnostic request and the CAN ID of the response message can be recorded. Another example is that, according to the vehicle diagnostic protocol, the CAN ID of the diagnostic request and the CAN ID of the response message can be directly generated, conforming to the protocol. For instance, the protocol specifies that the CAN ID of the diagnostic request is 0x7E0 and the CAN ID of the response message is 0x7E8.

[0077] Step 220: Based on at least one negative sample generation rule corresponding to the vehicle communication protocol, process the first identifier pair to generate at least one second identifier pair, the second identifier pair being different from the first identifier pair.

[0078] The second identifier pair is an identifier pair based on the first identifier pair, representing a simulated communication relationship between the identifier carried in a non-real diagnostic request and the identifier carried in the response message.

[0079] The first identifier pair is (0x7E0, 0x7E8), and the second identifier pair can be (0x7E0, 0x7EA), which represents data that does not actually involve communication interaction; or the second identifier pair can be (0x7E0, 0x7F1), where 0x7F1 is an invalid ID that does not exist.

[0080] In some embodiments, the second identifier pair can be generated in various ways. For example, the second identifier pair can be obtained based on an existing first identifier pair through random generation, statistical models, or other methods. The second identifier pair may differ from the first identifier pair in length and / or content.

[0081] Negative sample generation rules refer to a series of specific rules for generating second identifier pairs. For example, a negative sample generation rule might be a random substitution rule, replacing the CAN ID of the response message in a first identifier pair with the CAN ID of another known ECU, constructing a false pairing. Another example is a cross-combination rule, cross-combining multiple first identifier pairs to generate CAN ID pairs that do not conform to actual communication logic. Yet another example is a Generative Adversarial Network (GAN) rule, using machine learning methods to generate structurally plausible but unrealistic identifier pairs. Still another example is a random perturbation rule, performing bit operations or offsets on a field in the CAN ID to generate a new CAN ID.

[0082] In some embodiments, the negative sample generation rule to be processed can be determined from a variety of negative sample generation rules by using indicator parameters. The variety of negative sample generation rules and their corresponding indicator parameters include:

[0083] Rule 1: Length mismatch (modify the field length of the identifier);

[0084] Rule 2: Function addressing obfuscation (replace with a default function addressing identifier);

[0085] Rule 3: Reverse the direction of identifiers (swap the identifiers carried in the diagnostic request with the identifiers carried in the response message);

[0086] Rule 4: Random identifier pollution (randomly generating identifiers or using identifiers across vehicle models);

[0087] Rule 5: Cross-combination attack (mixing identifier pairs across logical groups for the same vehicle model);

[0088] Rule 6: Random physical timing perturbations (simulating physical layer interference causing bit errors in certain identifiers), etc. For more information on various negative sample generation rules, please see the relevant descriptions below.

[0089] In some embodiments, for multiple second identifier pairs generated based on the first identifier pair, the generation model can select at least one second identifier pair that satisfies a first preset condition based on the similarity between the first identifier pair and the second identifier pair, and use it as the final second identifier.

[0090] Generative models are models used to generate identifier pairs that meet certain requirements, where there is similarity between the identifiers in the identifier pairs, for example, the similarity is greater than a preset threshold.

[0091] In some embodiments, the generative model includes an encoding model and a similarity calculation model. The encoding model is used to represent the input identifier pairs as vectors. If the input is a character, the input identifier pairs can be represented as vectors to obtain vectors for different characters in the input identifier pairs. For example, for the input identifier pair a1 and b1, the vectors corresponding to a1 and b1 can be obtained.

[0092] The similarity calculation model calculates the similarity between different identifier pairs based on their vectors. That is, the similarity calculation model can calculate the similarity between two identifiers based on their vectors. In some embodiments, similarity can be measured by distance. Specifically, the similarity between the vectors of the first identifier pair and the vectors of the second identifier pair can be obtained. Distance is negatively correlated with similarity; that is, the greater the distance, the smaller the similarity. In some embodiments, the distance may include, but is not limited to, cosine distance, Euclidean distance, Manhattan distance, Mahalanobis distance, or Minkowski distance.

[0093] An identifier pair consists of two identifiers. In some embodiments, an identifier pair may consist of a first identifier and a second identifier. In some embodiments, an identifier pair may be represented by (Sp, Sn), where Sp represents the first identifier in the identifier pair and Sn represents the second identifier in the identifier pair.

[0094] In some embodiments, the generative model can select at least one second identifier pair that satisfies a first preset condition based on the similarity between the first identifier pair and the second identifier pair. Specifically, the first identifier pair and the second identifier pair are input into the generative model, which can calculate the similarity between the first identifier pair and the second identifier pair. When the similarity satisfies the first preset condition, the second identifier pair is taken as the final identifier pair. It can be understood that after the first identifier pair and the second identifier pair are input into the generative model, the encoding model of the generative model first represents the first identifier pair and the second identifier pair as vectors. Then, the similarity calculation model of the generative model can calculate the similarity between the vectors of the first identifier pair and the vectors of the second identifier pair to obtain the similarity between the first identifier pair and the second identifier pair.

[0095] In some embodiments, the first preset condition can be specifically set according to actual needs. For example, the first preset condition can be that the similarity between the second identifier pair and the first identifier pair is greater than or equal to a preset threshold. The preset threshold can be specifically set according to actual needs, such as 0.9 or 0.95, etc. By setting different preset thresholds, second identifier pairs with different similarities can be selected. Another example is that the preset condition is the Top N in similarity ranking, where N can be 1, 2, etc.

[0096] For example, the first preset condition is the maximum similarity. The first identifier pair is a1, and the second identifier pairs are b1, b2, b3, b4, and b5. The generative model calculates that the similarity between a1 and b1, b2, b3, b4, and b5 are 0.65, 0.87, 0.96, 0.34, and 0.67, respectively. b3 is then selected as the final identifier pair.

[0097] It should be noted that by selecting a second identifier pair that is highly similar to the first identifier pair, the model can learn the subtle differences between normal and abnormal communication, thereby reducing false positives (misclassifying normal communication as abnormal) and false negatives (misclassifying abnormal communication as normal). The second identifier pair with high similarity can better simulate actual abnormal communication, enabling the model to detect abnormal situations more accurately and improve the anomaly detection capability.

[0098] Step 230: Based on the second identifier pair, generate a second sample for training the relation discrimination model.

[0099] The second sample is at least a portion of the training data used to train the relation discrimination model. In some embodiments, the second sample is a negative sample. A negative sample is a training sample that does not contain pairs of identifiers with correct pairwise relations. By training the model with positive samples, the model can learn the identifiers carried in diagnostic requests and response messages during normal communication.

[0100] The process of training a machine learning model using positive and negative samples allows the model to learn the differences between positive and negative samples, thereby distinguishing the pairwise relationships between different identifiers in normal and abnormal communication in practical applications.

[0101] In some embodiments, multiple second identifier pairs can be generated based on a first identifier pair and a preset ratio. The preset ratio can be determined by a system preset or by system default; for example, the preset ratio is 1:1, meaning that one second identifier pair is generated based on one first identifier pair. The preset ratio can be determined according to actual circumstances.

[0102] In some embodiments, the relation discriminant model is generated based on the initial model, and the initial model and the relation discriminant model have the same model structure.

[0103] In some embodiments, the initial model can be trained iteratively multiple times based on a training sample set to obtain a relation discrimination model. The training sample set includes multiple pairs of identifiers carrying labels. The label indicates whether each identifier pair matches; for example, if the identifier pairs do not match, the label is 0, and if they match, the label is 1. In some embodiments, the identifier pairs in the training sample set can be any two identifiers. In some embodiments, positive sample pairs are identifier pairs with matching labels, and negative sample pairs are identifier pairs with non-matching labels.

[0104] In some embodiments, the initial model may include any supervised learning model used to train the relation discrimination model, such as a linear regression model, a neural network model, a decision tree model, a support vector machine model, a Naive Bayes model, or a KNN model. In some embodiments, the training sample set may refer to a collection of multiple training samples selected from the training data. The training data may be determined based on different application scenarios, which may include vehicle fault diagnosis, vehicle communication protocol detection, etc. In some embodiments, an iteration round may refer to performing a complete training of the initial model using all training samples in the training set, such as inputting all training samples in the training set into the initial model once and updating the parameters of the initial model. The iterative model may refer to the corresponding model generated by the initial model based on different iteration rounds.

[0105] Specifically, after inputting the training sample set into the initial model, the model parameters can be adjusted based on the model's output predictions and labels. For example, parameters can be adjusted using gradient descent or backpropagation. The model can be continuously trained until the training results converge, at which point training ends, and the initial model is obtained.

[0106] In some embodiments, training ends when the initial model meets preset conditions. These preset conditions may include the loss function result converging or falling below a preset threshold.

[0107] In some embodiments of this specification, by generating a second identifier pair that is different from the first identifier pair, the relationship discrimination model can learn the differences between identifier pairs during normal and abnormal communication, which helps to improve the robustness and reliability of the model.

[0108] In some embodiments, the negative sample generation rule includes a replacement rule.

[0109] Based on at least one negative sample generation rule corresponding to the vehicle communication protocol, the first identifier pair is processed to generate at least one second identifier pair, including:

[0110] Based on at least one identifier in the first identifier pair, generate a replacement identifier that is different from the identifier;

[0111] Based on the obtained at least one replacement identifier, replace at least one identifier to obtain a second identifier pair.

[0112] At least one identifier can be any one or both of the first identifier pair. A replacement identifier is a new identifier used to replace one of the identifiers in the first identifier pair. The replacement identifier can be randomly generated or selected from other datasets.

[0113] In some embodiments, the number of replacement identifiers can be one or two. When the number of replacement identifiers is one, the replacement identifier is used to replace any one identifier in the first identifier pair. When the number of replacement identifiers is two, the two replacement identifiers are used to replace the two identifiers in the first identifier pair respectively.

[0114] In some embodiments, for any one of the first identifier pairs, the length of the identifier remains unchanged, and one or more digits of the value can be randomly changed based on the identifier. If this is performed n times, n new identifiers are generated to replace the original identifier, generating at least one second identifier pair, such that the newly generated second identifier pair no longer satisfies the pairing relationship defined in the communication protocol.

[0115] In some embodiments, two corresponding new identifiers can be generated in a similar manner as described above, and the two new identifiers can be used to replace the two identifiers in the first identifier pair to generate a second identifier pair.

[0116] In some embodiments, the first identifier includes either the CAN ID of the diagnostic request or the CAN ID of the response message. The CAN ID of the diagnostic request appears in the request frame and is used to identify the source or type of the diagnostic request. In an on-board diagnostic system, the ID used by the diagnostic tool or other master device to send a command to an ECU (Electronic Control Unit) is the CAN ID of the diagnostic request. The CAN ID of the response message appears in the response frame and is used to identify the source or type of the response message. In an on-board diagnostic system, when an ECU receives a request, it sends a response message using its specific CAN ID.

[0117] In practical applications, there is a clear pairwise relationship between the CAN ID of the diagnostic request and the CAN ID of the response message. For example, in some vehicle diagnostic scenarios, a CAN ID of 0x7E0 for a diagnostic request corresponds to a CAN ID of 0x7E8 for the response message, indicating that a specific ECU will respond with 0x7E8 after receiving a 0x7E0 request.

[0118] In some embodiments of this specification, by generating replacement identifiers that have the same length as at least one identifier in the first identifier pair but different content, more diverse negative samples can be generated, which helps the model learn a wider range of abnormal patterns, thereby improving the model's generalization ability.

[0119] In some embodiments, generating a replacement identifier that is different from the identifier based on at least one identifier in the first identifier pair includes:

[0120] Obtain a function-addressable identifier with the same length as at least one identifier, and use it as a replacement identifier.

[0121] A function-addressable identifier (FID) is an identifier used in vehicle communication protocols to identify a specific function or service. For example, a FID can be used to broadcast messages to all electronic control units (ECUs) that support that function or service. For instance, in a vehicle's CAN communication, the FID can be a specific CAN ID.

[0122] For example, for rule 2, obtain the length of the CAN ID that meets the requirements, randomly select a functional addressing CAN ID with the required length, replace the CAN ID of the diagnostic request or the CAN ID of the response message with the functional addressing CAN ID, and generate a second identifier pair.

[0123] In some embodiments, generating a replacement identifier that is different from the identifier based on at least one identifier in the first identifier pair includes:

[0124] A sample identifier is randomly selected from the sample set corresponding to the vehicle identifier as a replacement identifier. The sample identifier is an identifier of the same type as either the identifier carried in the diagnostic request or the identifier carried in the response message.

[0125] Based on the obtained at least one replacement identifier, replace at least one identifier to obtain a second identifier pair, including:

[0126] The second identifier pair is obtained by replacing the corresponding type of identifier in the first identifier pair with the sample identifier.

[0127] The sample set corresponding to a vehicle identifier refers to the set of multiple identifier pairs associated with the vehicle to which the first identifier pair belongs. The sample set corresponding to a vehicle identifier can be captured from actual communication data or generated according to the vehicle communication protocol.

[0128] A sample identifier is an identifier randomly selected from the sample set. A sample identifier is an identifier pair that belongs to the same vehicle type as the first identifier pair.

[0129] In some embodiments, a sample identifier may be randomly selected from the sample set corresponding to the vehicle identifier. When the sample identifier is of the type of identifier carried in the diagnostic request, the identifier carried in the diagnostic request of the first identifier pair is replaced based on the sample identifier. When the sample identifier is of the type of identifier carried in the response message, the identifier carried in the response message of the first identifier pair is replaced based on the sample identifier.

[0130] In some embodiments, two sample identifiers can be randomly selected from the sample set corresponding to the vehicle identifier, including a first sample identifier and a second sample identifier. The first sample identifier is the type of identifier carried in the diagnostic request, and the identifier carried in the diagnostic request of the first identifier pair is replaced based on the first sample identifier. The second sample identifier is the type of identifier carried in the response message, and the identifier carried in the response message of the first identifier pair is replaced based on the second sample identifier.

[0131] For example, in rule 5, a CAN ID of a diagnostic request or a CAN ID of a response message is randomly selected from the positive sample dataset of the same VIN code, and the diagnostic request CAN ID or the CAN ID of the response message of the first identifier pair is replaced to generate the second identifier pair. The same VIN code refers to the VIN code of the vehicle to which the first identifier belongs.

[0132] In some embodiments, the first identifier pair includes a first identifier and a second identifier, wherein one of the first identifier and the second identifier is an identifier carried in the diagnostic request and the other is an identifier carried in the response message;

[0133] Based on at least one identifier in the first identifier pair, generate a replacement identifier that is different from the identifier, including:

[0134] Based on the length of the identifier, generate at least one replacement identifier with the same length but different content; and / or

[0135] Randomly select at least one replacement identifier with the same length as the identifier from the sample set corresponding to other vehicle identifiers;

[0136] Based on the obtained at least one replacement identifier, replace at least one identifier to obtain a second identifier pair, including:

[0137] Based on the two replacement identifiers obtained, the first identifier and the second identifier are replaced respectively to obtain the second identifier pair.

[0138] The sample set corresponding to other vehicle identifiers refers to the collection of multiple identifier pairs associated with other vehicles. The sample set corresponding to other vehicle identifiers can be captured from actual communication data or generated according to vehicle communication protocols. The vehicles to which other vehicles belong are different vehicle models from those to which the first identifier pair belongs.

[0139] Multiple replacement identifiers include a first replacement identifier and a second replacement identifier.

[0140] The first substitution identifier and the second substitution identifier are new identifiers generated based on the first identifier and the second identifier, respectively. The first substitution identifier and the second substitution identifier have the same type and length as the first identifier and the second identifier, but their numerical values ​​are different.

[0141] In some embodiments, one or more new identifiers with the same length as the first identifier can be generated randomly; alternatively, certain bits of the first identifier can be flipped or modified according to specific rules to generate one or more new identifiers. The first replacement identifier is then randomly selected from these new identifiers. A second replacement identifier can also be obtained using a similar method. The first replacement identifier has the same length and type as the first identifier, but different content; the second replacement identifier has the same length and type as the first identifier, but different content.

[0142] In addition to the existing identifiers for the first identifier of the vehicle to which it belongs, identifiers of other vehicles with the same length can be introduced to increase sample diversity and improve the model's generalization ability to recognize cross-vehicle types.

[0143] For example:

[0144] Model A uses CAN ID 0x7E0 and CAN ID 0x7E8 to indicate diagnostic requests / responses;

[0145] Model B uses CAN ID 0x7D0 and CAN ID 0x7D8 to implement diagnostic requests / responses;

[0146] The CAN ID 0x7D0 or CAN ID 0x7D8 of model B can be introduced into model A as a replacement identifier.

[0147] In some embodiments, one or more new identifiers with the same length as the first identifier can be generated based on a sample set corresponding to other vehicle identifiers. A first replacement identifier is then randomly selected from these new identifiers. A second replacement identifier can also be obtained in a similar manner, where the first replacement identifier has the same length and type as the first identifier but different content, and the second replacement identifier has the same length and type as the first identifier but different content.

[0148] For example, for rule 4, obtain the CAN ID length of the first identifier in the first identifier pair, and randomly generate a contamination type based on the length. When the contamination type is randomly generated, a CAN ID with the same length is randomly generated as the first replacement identifier. When the contamination type is cross-vehicle type, a CAN ID with a different VIN code but the same length is randomly selected from the positive sample dataset as the second replacement identifier. Replace the CAN ID of the diagnostic request and the CAN ID of the response message to generate the second identifier pair.

[0149] In some embodiments of this specification, diverse negative samples can be generated without extensive manual annotation by automatically selecting replacement identifiers from local and external vehicle databases; by introducing second replacement identifiers of the same length from other vehicle platforms, the model can adapt to the discrimination of communication relationships under different vehicle models and different ECU configurations.

[0150] In some embodiments, generating a replacement identifier that is different from the identifier based on at least one identifier in the first identifier pair includes:

[0151] Determine the target identifier from the first identifier pair;

[0152] Flip the contents of the preset mask bits in the target identifier to obtain the replacement identifier.

[0153] The target identifier is the specific identifier in the first identifier pair that needs to be replaced. The target identifier can be one of the identifiers carried in the request or the response message.

[0154] For example, for rule 6, generate random mask bits (the mask bits can be one bit or multiple bits of data), flip the numbers of the corresponding mask bits in the CAN ID of the diagnostic request or the CAN ID of the response message, generate a replacement identifier, and replace the corresponding CAN ID of the diagnostic request or the CAN ID of the response message to obtain a second identifier pair.

[0155] In some embodiments, the first sample includes the lengths corresponding to the two identifiers, and the negative sample generation rule includes a length adjustment rule.

[0156] Length adjustment rules include:

[0157] Determine the target identifier and its length from the two identifiers in the first sample;

[0158] Adjust the length of the target identifier to obtain the identifier content after adjustment;

[0159] Based on the adjusted length of the identifier content, the target identifier in the first sample is replaced to generate a second identifier pair.

[0160] In some embodiments, a data of a different length than the target identifier can be randomly generated as the adjusted identifier content. Based on the adjusted identifier content, the target identifier in the first sample is replaced to generate a second identifier pair to obtain the second sample.

[0161] For example, for rule 1, a data of a length not equal to the length of the CAN ID of the diagnostic request or the CAN ID of the response message can be randomly generated based on the length of the CAN ID of the compliant diagnostic request or the CAN ID of the response message to obtain the adjusted length identifier content. The adjusted length identifier content is then used to replace the CAN ID of the diagnostic request or the CAN ID of the response message. Based on the replaced identifier pair, a second identifier pair is generated to obtain the second sample.

[0162] In some embodiments of this specification, generating second identifier pairs by replacing identifiers can increase the diversity of training samples, enabling the model to learn a wider range of features and helping to improve the model's generalization ability.

[0163] In some embodiments, the negative sample generation rule includes a position swapping rule.

[0164] Location exchange rules include:

[0165] The positions of the identifier carried in the diagnostic request and the identifier carried in the response message in the first sample are changed to generate a second identifier pair.

[0166] For example, for rule 3, the diagnostic request CAN ID and the response message CAN ID are exchanged to generate a second identifier pair.

[0167] In some embodiments of this specification, a second identifier pair can be quickly generated to obtain a negative sample by exchanging the request ID and response ID.

[0168] Figure 3 This is an exemplary flowchart of a relationship discrimination method according to some embodiments of this specification.

[0169] In some embodiments, process 300 may be executed based on an electronic device. For example... Figure 3 As shown, process 300 includes the following steps.

[0170] Step 310: Obtain the identifier pair information to be judged. The identifier pair information includes the identifier carried in the diagnostic request to be judged, the identifier carried in the response message to be judged, and the length of each identifier.

[0171] The identifier pair information to be judged refers to a set of identifier pairs whose authenticity has not yet been determined. These pairs are used as inputs into the relationship discrimination model to determine whether they constitute a valid request-response relationship.

[0172] In some embodiments, the identifier pair information to be determined may include the identifier and its corresponding length carried in the diagnostic request to be determined, and the identifier and its corresponding length carried in the response message to be determined.

[0173] The diagnostic request to be determined can be, for example, a request message sent by the vehicle diagnostic equipment or control unit.

[0174] The response message to be judged can be a response message sent by the vehicle's electronic control unit.

[0175] Step 320: Input the identifier information to be judged into the trained relation discrimination model to obtain the discrimination result. The discrimination result is used to indicate the degree of matching between the diagnostic request to be judged and the response message to be judged.

[0176] The matching score can be a value between 0 and 1, representing the degree of confidence the relation discriminant model has in the pairwise relationship between the diagnostic request to be discriminated and the response message to be discriminated. A high score indicates a high probability that the pairwise relationship detection result is correct, while a low score indicates a low probability that the pairwise relationship detection result is correct.

[0177] A relation discriminant model is a model or algorithm used to determine the degree of matching between a diagnostic request to be discriminated and a response message to be discriminated.

[0178] In some embodiments, the relation discrimination model is a machine learning model. For example, the relation discrimination model may include any one or a combination of convolutional neural networks (CNN), recurrent neural networks (RNN), deep neural networks (DNN), or other custom model structures.

[0179] In some embodiments, the input to the relation discrimination model includes identifier pairs to be discriminated, and the output may include the degree of matching between the diagnostic request to be discriminated and the response message to be discriminated.

[0180] In some embodiments, the relation discrimination model can be trained using a large number of labeled training samples through various feasible methods. For example, parameters can be updated using gradient descent. An exemplary training process includes: inputting multiple labeled training samples into an initial relation discrimination model; constructing a loss function using the labels and the results of the initial relation discrimination model; and iteratively updating the parameters of the initial relation discrimination model based on the loss function using gradient descent or other methods. Sample generation is complete when preset conditions are met, resulting in a trained relation discrimination model. These preset conditions may include loss function convergence, the number of iterations reaching a threshold, etc.

[0181] In some embodiments, the training samples include at least positive and negative samples. The training samples may be obtained based on historical data.

[0182] In some embodiments, the label may include the actual matching degree corresponding to the training sample. The label can be obtained automatically or manually.

[0183] In some embodiments of this specification, the relation discrimination model helps to efficiently and accurately obtain the matching degree between the diagnostic request to be judged and the response message to be judged.

[0184] It should be noted that the above description of the process is for illustrative purposes only and does not limit the scope of this specification. Those skilled in the art can make various modifications and changes to the process under the guidance of this specification. However, these modifications and changes remain within the scope of this specification.

[0185] Figure 4 This is a schematic diagram of the sample generation apparatus shown in some embodiments of this specification.

[0186] like Figure 4 As shown, one or more embodiments of this specification also provide a schematic diagram of a sample generation device. This sample generation device may include:

[0187] The acquisition module 401 is used to acquire at least one first sample corresponding to the relationship discrimination model to be trained. The first sample includes a vehicle identifier and a first identifier pair. The first identifier pair includes two identifiers. The relationship discrimination model is used to identify whether the two identifiers in the first sample conform to the pair relationship corresponding to the vehicle communication protocol. The pair relationship indicates that one of the two identifiers is the identifier carried in the diagnostic request sent to the vehicle indicated by the vehicle identifier, and the other identifier is the identifier carried in the response message of the vehicle to the diagnostic request.

[0188] The first generation module 402 is used to process the first identifier based on at least one negative sample generation rule corresponding to the vehicle communication protocol to generate at least one second identifier pair, wherein the second identifier pair is different from the first identifier pair.

[0189] The second generation module 403 is used to generate a second sample for training a relation discrimination model based on the second identifier pair.

[0190] The acquisition module 401, the first generation module 402, and the second generation module 403 can be used to execute the corresponding embodiments of the above sample generation method. For the specific implementation methods of these modules and more details, please refer to the corresponding method section, which will not be elaborated here.

[0191] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0192] Figure 5 This is a schematic diagram of the structure of an electronic device according to some embodiments of this specification.

[0193] This application embodiment also provides an electronic device 500, which may include components such as a processor 501 with one or more processing cores, a memory 502 with one or more computer-readable storage media, a power supply 503, and an input unit 504. Those skilled in the art will understand that... Figure 5 The electronic device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein:

[0194] The processor 501 is the sample generation center, connecting various parts of the electronic device via various interfaces and lines. It performs various functions and processes data by running or executing software programs and / or modules stored in the memory 502, and by calling data stored in the memory 502, thereby providing overall monitoring of the electronic device. It is understood that the processor 501 communicates with the controller via signal transmission. Optionally, the processor 501 may include one or more processing cores; preferably, the processor 501 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, while the modem processor mainly handles wireless communication. It is understood that the aforementioned modem processor may also not be integrated into the processor 501.

[0195] The memory 502 can be used to store software programs and modules. The processor 501 executes various functional applications and data processing by running the software programs and modules stored in the memory 502. The memory 502 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 502 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 502 may also include a memory controller to provide the processor 501 with access to the memory 502.

[0196] In some embodiments of this application, the sample generation apparatus can be implemented as a computer program, which can be implemented as follows: Figure 5 The device operates on the electronic device shown. The memory of the electronic device can store various program modules that make up the sample generation apparatus. The computer program, composed of the various program modules, causes the processor to execute the steps of the sample generation methods in the various embodiments of this application described in this specification.

[0197] The electronic device includes a processor, memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface is used to communicate with external electronic devices via a network connection. When the computer program is executed by the processor, it implements a sample generation method.

[0198] The electronic device also includes a power supply 503 that supplies power to various components. Preferably, the power supply 503 can be logically connected to the processor 501 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 503 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0199] The electronic device may also include an input unit 504, which can be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.

[0200] Although not shown, the electronic device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 501 in the electronic device loads the executable files corresponding to the processes of one or more applications into the memory 502 according to computer instructions, and the processor 501 runs the applications stored in the memory 502 to realize various functions, such as the sample generation methods of various embodiments of this application described in this specification.

[0201] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0202] In practice, each of the above units or structures can be implemented as an independent entity or can be arbitrarily combined to be implemented as the same or several entities. For the specific implementation of each of the above units or structures, please refer to the previous method embodiments, which will not be repeated here.

[0203] It should be noted that, Figure 5 This is merely one implementation of the electronic device 500 provided in this application embodiment. In actual applications, the electronic device 500 may include more or fewer components, which is not limited here.

[0204] It should be understood that the various solutions in the embodiments of this application can be used in a reasonable combination, and the explanations or descriptions of the various terms appearing in the embodiments can be referenced or explained to each other in the various embodiments, without limitation.

[0205] It should also be understood that, in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0206] Based on the above embodiments and the same concept, this application also provides a computer-readable storage medium storing a computer program that, when run on a computer, causes the computer to perform the method provided in the above embodiments.

[0207] Based on the above embodiments and the same concept, this application also provides a computer program product, which includes a computer program that, when run on a computer, causes the computer to execute the method provided in the above embodiments.

[0208] This application also provides a vehicle, which includes the sample generation device described in any embodiment; or includes the electronic device described in any embodiment. The vehicle may be a gasoline-powered vehicle, a plug-in hybrid electric vehicle, or a new energy vehicle, etc., and this specification does not specifically limit it.

[0209] In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more features. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.

[0210] The embodiments, implementation methods, and related technical features of this application can be combined and substituted for each other without conflict.

[0211] The above are merely preferred embodiments of this application and are not intended to limit this application in any way. Although the descriptions of each embodiment in this application have different focuses, and the parts not described in detail in a certain embodiment can be referred to the relevant embodiments of other embodiments, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of this application without departing from the content of the technical solution of this application shall still fall within the scope of the technical solution of this application.

Claims

1. A sample generation method, characterized in that, The method includes: At least one first sample corresponding to the relationship discrimination model to be trained is obtained. The first sample includes a vehicle identifier and a first identifier pair. The first identifier pair includes two identifiers. The relationship discrimination model is used to identify whether the two identifiers in the first sample conform to the pair relationship corresponding to the vehicle communication protocol. The pair relationship indicates that one of the two identifiers is an identifier carried in a diagnostic request sent to the vehicle indicated by the vehicle identifier, and the other identifier is an identifier carried in the response message of the vehicle to the diagnostic request. Based on at least one negative sample generation rule corresponding to the vehicle communication protocol, the first identifier pair is processed to generate at least one second identifier pair, the second identifier pair being different from the first identifier pair; Based on the second identifier pair, a second sample is generated for training the relation discrimination model.

2. The method according to claim 1, characterized in that, The negative sample generation rules include replacement rules. The first identifier pair is processed based on at least one negative sample generation rule corresponding to the vehicle communication protocol to generate at least one second identifier pair, including: Based on at least one identifier in the first identifier pair, generate a replacement identifier that is different from the identifier; Based on the obtained at least one replacement identifier, the at least one identifier is replaced to obtain the second identifier pair.

3. The method according to claim 2, characterized in that, The step of generating a replacement identifier different from the first identifier pair based on at least one identifier in the first identifier pair includes: Obtain a functionally addressed identifier with the same length as the at least one identifier, and use it as the replacement identifier.

4. The method according to claim 2, characterized in that, The step of generating a replacement identifier different from the first identifier pair based on at least one identifier in the first identifier pair includes: A sample identifier is randomly selected from the sample set corresponding to the vehicle identifier as the replacement identifier. The sample identifier is an identifier of the same type as either the identifier carried in the diagnostic request or the identifier carried in the response message. The step of replacing the at least one obtained replacement identifier to obtain the second identifier pair includes: The second identifier pair is obtained by replacing the corresponding type of identifier in the first identifier pair with the sample identifier.

5. The method according to claim 2, characterized in that, The first identifier pair includes a first identifier and a second identifier, wherein one of the first identifier and the second identifier is an identifier carried in the diagnostic request, and the other is an identifier carried in the response message; The step of generating a replacement identifier different from the first identifier pair based on at least one identifier in the first identifier pair includes: Based on the length of the identifier, generate at least one replacement identifier with the same length but different content from the identifier; and / or Randomly select at least one replacement identifier with the same length as the identifier from the sample set corresponding to other vehicle identifiers; The step of replacing the at least one obtained replacement identifier to obtain the second identifier pair includes: Based on the two replacement identifiers obtained, the first identifier and the second identifier are replaced respectively to obtain the second identifier pair.

6. The method according to claim 2, characterized in that, The step of generating a replacement identifier different from the first identifier pair based on at least one identifier in the first identifier pair includes: Determine the target identifier from the first identifier pair; The replacement identifier is obtained by flipping the contents of the preset mask bits in the target identifier.

7. The method according to claim 1, characterized in that, The first sample includes the lengths corresponding to the two identifiers respectively, and the negative sample generation rule includes a length adjustment rule and / or a position swapping rule, wherein, The length adjustment rules include: Determine the target identifier and the length of the target identifier from the two identifiers in the first sample; Adjust the length of the target identifier to obtain the identifier content after adjustment; Based on the adjusted length of the identifier content, the target identifier in the first sample is replaced to generate the second identifier pair; The location exchange rules include: The positions of the identifier carried in the diagnostic request and the identifier carried in the response message in the first sample are changed to generate the second identifier pair.

8. A method for determining relationships, characterized in that, The method includes: Obtain identifier pair information to be judged, the identifier pair information including the identifier carried in the diagnostic request to be judged, the identifier carried in the response message to be judged, and the length of each of the identifiers; The identifier to be judged is input into the trained relation discrimination model to obtain the discrimination result, which is used to indicate the degree of matching between the diagnostic request to be judged and the response message to be judged.

9. A sample generation device, characterized in that, The device includes: The acquisition module is used to acquire at least one first sample corresponding to the relationship discrimination model to be trained. The first sample includes a vehicle identifier and a first identifier pair. The first identifier pair includes two identifiers. The relationship discrimination model is used to identify whether the two identifiers in the first sample conform to the pairing relationship corresponding to the vehicle communication protocol. The pairing relationship indicates that one of the two identifiers is an identifier carried in a diagnostic request sent to the vehicle indicated by the vehicle identifier, and the other identifier is an identifier carried in the response message of the vehicle to the diagnostic request. The first generation module is used to process the first identifier based on at least one negative sample generation rule corresponding to the vehicle communication protocol to generate at least one second identifier pair, wherein the second identifier pair is different from the first identifier pair. The second generation module is used to generate a second sample for training the relation discrimination model based on the second identifier pair.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8.