Verification library generation method, format verification method and equipment
By generating and using the target verification library, intelligent format verification of YAML format configuration files is solved, and the problem of inability to effectively deal with complex field relationships and constraints in the existing technology is improved, and the security and consistency of configuration files are improved.
Patent Information
- Application Number
- CN202510067117.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-30
AI Technical Summary
The existing YAML format verification method is single and rough, and it is impossible to effectively identify and handle field relationships and constraints in complex configuration files, resulting in configuration errors and inconsistencies.
By obtaining the feature vector set of configuration files, the initial model is trained to generate a verification library, and the target verification library is determined through preset verification conditions to achieve intelligent format verification of the features of the configuration file field.
Improves the format verification efficiency of configuration files, enhances the security and consistency of configuration files, and can effectively identify and handle complex field relationships and constraints.
Smart Images

Figure CN120067675A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application belong to the technical field of data verification, and particularly relate to a method for generating a verification library, a format verification method, and a device. Background Art
[0002] With the increasing complexity of modern software systems, many applications and services use the YAML (YAML Ain't Markup Language) format as the standard format for configuration files. Due to its strong readability and clear structure, it is widely used in fields such as system configuration, data transmission, and automation tools. However, although the YAML format has obvious advantages in terms of ease of use, the current verification methods are still relatively single and rough in verifying the correctness and consistency of its data content.
[0003] Currently, the mainstream YAML verification methods usually only rely on basic format checks, such as the legality of line breaks, indents, or spaces. These methods mainly focus on checking the syntax structure of the file to ensure that the YAML file conforms to the basic format specifications. In complex configuration files, there may be certain relationships or constraints between fields, and the current verification methods cannot effectively identify and process these logical relationships, resulting in configuration errors and inconsistencies that are likely to occur during use.
[0004] Therefore, there is an urgent need for an innovative method that can intelligently and comprehensively verify configuration files to improve the reliability and security of configuration files. Summary of the Invention
[0005] In view of this, the embodiments of the present application provide a method for generating a verification library, a format verification method, and a device, which are used to improve the verification efficiency of configuration files, as well as improve the security and reliability of configuration files.
[0006] The first aspect of the embodiments of the present application provides a method for generating a verification library, including:
[0007] Obtain a feature vector set; the feature vector set includes a plurality of feature vectors determined from a configuration file; the feature vectors are used to characterize the features of the fields of the configuration file;
[0008] Train an initial model according to the feature vector set;
[0009] Receive the trained verification library output by the initial model for the feature vector set;
[0010] When the trained verification library passes a preset verification condition, determine the trained verification library as the target verification library.
[0011] In some implementation manners of the first aspect, the obtaining the feature vector set includes:
[0012] Obtain a first set of configuration files;
[0013] Extract the feature information of the fields from the first set of configuration files;
[0014] Vectorize the feature information to generate a feature vector corresponding to the field;
[0015] Generate a set of feature vectors based on the feature vectors corresponding to each of the fields;
[0016] The feature information includes at least one of: field type feature, required field feature, field dependency feature, field mutual exclusion feature, field constraint condition feature.
[0017] In some implementation manners of the first aspect, the initial model is used to generate a set of mapping vectors for each of the feature vectors; the set of mapping vectors includes a query vector, a key vector, and a value vector; calculate the attention weights between the target feature vector and other feature vectors; generate context information based on each of the attention weights of the target feature vector and the value vectors corresponding to the other feature vectors; generate a training and validation library based on the context information of each of the target feature vectors.
[0018] Wherein, the target feature vector is any one of the feature vectors.
[0019] In some implementation manners of the first aspect, calculating the attention weights between the target feature vector and other feature vectors includes:
[0020] Determine the target feature vector in the set of feature vectors, and the feature vectors in the set of feature vectors other than the target feature vector are other feature vectors;
[0021] Calculate the similarity between the target feature vector and the other feature vectors;
[0022] Calculate the attention weights between the target feature vector and the other feature vectors based on each of the similarities.
[0023] In some implementation manners of the first aspect, calculating the similarity between the target feature vector and the other feature vectors includes:
[0024] Determine the query vector of the target feature vector as the target query vector;
[0025] Determine the device of the key vector of the other feature vector as the target key vector;
[0026] Calculate the similarity between the target feature vector and the other feature vectors according to the similarity calculation formula, the target query vector, and the target key vector;
[0027] The similarity calculation formula is Q i is the target query vector, is the target key vector.
[0028] In some implementation manners of the first aspect, calculating the attention weights between the target feature vector and the other feature vectors according to the similarities includes:
[0029] Use the similarity transformation formula to perform normalization calculation on each of the similarities to determine the attention weights between the target feature vector and the other feature vectors;
[0030] The similarity transformation formula is
[0031] where n is used to represent the number of feature vectors input to the initial model, k is the index information of the feature variable, i is used to represent the field corresponding to the target feature vector, and j is used to represent the field corresponding to one of the other feature vectors.
[0032] In some implementation manners of the first aspect, generating context information according to the attention weights of the target feature vector and the value vectors corresponding to the other feature vectors includes:
[0033] For each field corresponding to the target feature vector, use the context calculation formula to calculate each of the attention weights and the value vectors corresponding to the other feature vectors to generate context information;
[0034] The context calculation formula is
[0035] C i is used to represent the context information corresponding to the field i corresponding to the target feature vector, j is used to represent the field corresponding to one of the other feature vectors, a ij is used to represent the attention weight corresponding to the target feature vector and one of the other feature vectors, V j is used to represent the value vector of the feature vector corresponding to the field j.
[0036] The second aspect of the embodiments of the present application provides a format verification method, including:
[0037] Obtain a second configuration file set;
[0038] Convert the second configuration file set into JSON data;
[0039] Use a target verification library to perform format verification on the JSON data and generate a verification result;
[0040] Among them, the target verification library is generated by the method described in the first aspect.
[0041] The third aspect of the embodiments of this application provides a device for generating a verification library, including:
[0042] A feature vector set acquisition module, configured to acquire a feature vector set; the feature vector set includes multiple feature vectors determined from a configuration file; the feature vectors are used to characterize the features of the fields of the configuration file;
[0043] A training module, configured to train an initial model based on the feature vector set;
[0044] A trained verification library receiving module, configured to receive the trained verification library output by the initial model for the feature vector set;
[0045] A target verification library determination module, configured to determine the trained verification library as the target verification library when the trained verification library passes a preset verification condition.
[0046] The fourth aspect of the embodiments of this application provides a format verification device, including
[0047] A second configuration file set acquisition module, configured to acquire a second configuration file set;
[0048] A format conversion module, configured to convert the second configuration file set into JSON data;
[0049] A verification module, configured to perform format verification on the JSON data using the target verification library and generate a verification result;
[0050] Among them, the target verification library is generated by the device described in the third aspect.
[0051] The fifth aspect of the embodiments of this application provides an electronic device, including a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the electronic device implements the method for generating a verification library as described in the first aspect above, or the format verification method as described in the second aspect above.
[0052] The sixth aspect of the embodiments of this application provides a computer program product, including a computer program. When the computer program is run, the method for generating a verification library as described in the first aspect above, or the format verification method as described in the second aspect above is executed.
[0053] The seventh aspect of the embodiments of the present application provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the method for generating a verification library as described in the first aspect above, or the format verification method as described in the second aspect above.
[0054] The embodiments of the present application have the following beneficial effects:
[0055] In the embodiments of the present application, by obtaining a feature vector set, which includes a plurality of feature vectors determined from a configuration file and is used to characterize the features of the fields of the configuration file; training an initial model according to the feature vector set; receiving the output training verification library of the trained initial model for the feature vector set; and determining the training verification library as the target verification library when the training verification library passes a preset verification condition, a target verification library for format verification of the configuration file is generated based on the features of the fields of the configuration file. When using the target verification library to perform format verification on the features of the fields of the configuration file, the format verification efficiency of the configuration file is improved, and thus the security and consistency of the configuration files in the same configuration file set are improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0057] Figure 1 It is a schematic diagram of a method for generating a verification library provided by an embodiment of the present application;
[0058] Figure 2 It is a configuration file verification flowchart provided by an embodiment of the present application;
[0059] Figure 3 It is a schematic diagram of a format verification method provided by an embodiment of the present application;
[0060] Figure 4 It is a schematic diagram of a device for generating a verification library provided by an embodiment of the present application;
[0061] Figure 5 It is a schematic diagram of a device for a format verification method provided by an embodiment of the present application;
[0062] Figure 6 It is a schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0063] In the following description, specific details such as specific system architectures, technologies, etc. are presented for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.
[0064] It should be understood that when used in the specification and appended claims of the present application, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0065] It should also be understood that the term "and / or" used in the specification and appended claims of the present application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0066] As used in the specification and appended claims of the present application, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" depending on the context. Similarly, the phrases "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]" depending on the context.
[0067] In addition, in the description of the specification and appended claims of the present application, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0068] The reference to "one embodiment" or "some embodiments" etc. described in the specification of the present application means that a specific feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of the present application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in another way. The terms "comprising", "including", "having", and their variants all mean "including but not limited to", unless otherwise specifically emphasized in another way.
[0069] Configuration files can be set in different formats, such as: INI, XML, JSON, YAML, etc. Among them, the configuration files in YAML format have the characteristics of being concise and easy to read, supporting multiple data structures, and supporting multiple data types. Therefore, the configuration files in YAML format are widely used in multiple scenarios, such as: client programs, servers, cloud services, databases, machine learning, etc.
[0070] In a configuration file in YAML format, a field is the basic unit that makes up a data structure. A field can be the key or value in a key-value pair, or an element in a list or sequence. There may be many fields in a configuration file. In the usage objects of the same configuration file (such as: the same server, the same application), the fields have certain characteristics. By identifying the characteristics of the fields and validating the configuration file, the reliability and security of the configuration file can be verified. In the related art, only the method of manual screening can be used to identify the characteristics of the fields in the configuration file, and the characteristics of the corresponding fields in different configuration files are compared manually. This verification method is extremely inefficient. The embodiment of the present application provides a method for generating a verification library. The generated verification library can quickly perform format verification on the configuration file based on the characteristics of the fields, greatly improving the verification efficiency of the configuration file, and further improving the security and reliability of the configuration file.
[0071] The technical solution of the present application will be described below through specific embodiments.
[0072] Refer to Figure 1 , which shows a schematic diagram of a method for generating a verification library provided by an embodiment of the present application. Specifically, it may include the following steps:
[0073] Step 101, obtain a feature vector set;
[0074] The feature vector set includes a plurality of feature vectors determined from the configuration file; the feature vectors are used to characterize the characteristics of the fields of the configuration file;
[0075] The configuration file set used by the same configuration file usage object can be divided into a first configuration file set and a second configuration file set. The first configuration file set is used to generate a target verification library, and the target verification library is used to perform format verification on the second configuration file set. In practical applications, the first configuration file set accounts for a small part of the configuration file set, so that a target verification library is generated through a small part of the configuration files, and format verification is performed on the remaining configuration files.
[0076] Specifically, a feature vector set can be obtained through a first set of configuration files. The feature vector set includes features of multiple fields used to characterize the configuration files, that is, to extract the features of the fields in the configuration files to generate a feature vector set.
[0077] Step 102: Train the initial model based on the feature vector set.
[0078] The initial model is a pre - constructed deep - learning model, including but not limited to a convolutional neural network model, a recurrent neural network model, a generative adversarial network model, a Transformer (a model that processes data through self - attention mechanism), a deep reinforcement learning model, etc.
[0079] The embodiments of the present application do not limit the deep - learning model used, as long as the initial model can process the feature vector set and output a verification library.
[0080] Before step 102, an initial model can be constructed and the model parameters of the initial model can be initialized. As an example, taking the initial model as a Transformer, the initialization settings can include: (1) Set the number of attention heads: the number of heads in the multi - head attention mechanism, with a default of 8. Increasing it can enhance the ability to capture complex relationships between fields. (2) Set the dimension of the feature vector: the dimension of the feature vector of the input field, with a default of 128. Increasing it can capture more field details. Set the Dropout ratio: Dropout is a regularization method used to reduce overfitting, with a default ratio of 0.1, and this ratio can be set according to the capacity of the training data.
[0081] It can be understood that the above - mentioned initialization process is only for illustrative purposes, and in actual applications, the corresponding model parameters can be set according to actual needs.
[0082] Step 103: Receive the training verification library output by the trained initial model for the feature vector set.
[0083] The initial model can process the feature vector set it receives, generate and output a training verification library.
[0084] The verification library is a JSON Schema, and the JSON Schema is used to describe and verify the JSON data structure.
[0085] Step 104: When the training verification library passes the preset verification conditions, determine the training verification library as the target verification library.
[0086] The feature vector set can be divided into two parts, namely a training set and a test set. The training set is used to train the initial model and receive the training verification library output by the initial model for the training set, and the test set is used to verify the training verification library.
[0087] The preset verification condition is that the training verification library passes the verification of the test set. As an example, passing the verification of the test set means that when the training verification library performs format verification on the test set, no format errors are found.
[0088] When the training verification library passes the verification condition, determine the current training verification library as the target verification library, and the target verification library can be used to perform format verification on the second configuration file set in the configuration file set; when the training verification library does not pass the verification condition, update the model parameters of the initial model, and return to step 102 until the training verification library passes the verification condition.
[0089] In the embodiment of the present application, by obtaining a feature vector set; the feature vector set includes a plurality of feature vectors determined from the configuration file; the feature vectors are used to characterize the features of the fields of the configuration file; training the initial model according to the feature vector set; receiving the training verification library output by the trained initial model for the feature vector set; when the training verification library passes the preset verification condition, determining the training verification library as the target verification library, so as to generate a target verification library for performing format verification on the configuration file based on the features of the fields of the configuration file. When using the target verification library to perform format verification on the features of the fields of the configuration file, the format verification efficiency of the configuration file is improved, and further the security and consistency of the configuration files in the same configuration file set are improved.
[0090] In some implementation manners of the embodiment of the present application, step 101 includes: obtaining a first configuration file set; extracting feature information of fields in the first configuration file set; vectorizing the feature information to generate feature vectors corresponding to the fields; generating a feature vector set according to the feature vectors corresponding to each field; the feature information includes at least one of field type feature, necessary field feature, field dependency relationship feature, field mutual exclusion relationship feature, and field constraint condition feature.
[0091] It is possible to obtain at least some of the multiple configuration files used by a specified configuration file user as the configuration file set, and determine the first configuration file set from the configuration file set according to a certain quantity or proportion.
[0092] By parsing multiple configuration files in the first configuration file set, a data structure tree of the configuration file can be obtained. The data structure tree contains information such as the fields, values, and nested structures of the configuration file, and based on the data structure tree, the feature information of the fields is extracted. The field feature information includes at least one of field type feature, necessary field feature, field dependency relationship feature, field mutual exclusion relationship feature, and field constraint condition feature.
[0093] Field type feature extraction: Analyze these data structure trees. If the field is of integer type, extract it as type: "integer"; if the field is of floating-point type, extract it as type: "number"; if the field is of boolean type, extract it as type: "boolean"; if the field value is null, extract it as type: "null"; if the field is a nested key-value pair, extract it as type: "object" and define the properties of the object in the properties field; if the field is an array (list), extract it as type: "array" and use items to describe the type of the array elements; if each element in the array is an object, the object structure can be defined in items. If a field does not belong to the above types, determine that the field is of the standard type and extract it as type: "string".
[0094] If different types appear for the same field during the analysis process, it can be extracted as including the types that appear, for example, extracted as: anyOf: [{type: "XXX"}, {type: "YYY"}] (where "XXX" and "YYY" can be one of the above "integer", "number", "boolean", "null", "object", "array" respectively). For example: a certain field is extracted as anyOf: [{type: "integer"}, {type: "string"}].
[0095] Field mandatory feature extraction: If a certain field appears in all data structure trees, determine that the field is mandatory and has the necessary field feature.
[0096] Field dependency feature extraction: According to the co-occurrence patterns of fields in these data structure trees, determine whether there is a dependency relationship between fields. For example, if name and age always appear together, then there may be a dependency relationship between them, which can be extracted as {name: name, dependencies: ["age"]}, {name: age, dependencies: ["name"]}.
[0097] Feature extraction of field mutual exclusion relationship: According to the mutual exclusivity of field values in the configuration file, judge the possible mutual exclusion relationships between fields. For example, if field content and field desc never appear in the same configuration file, then it can be inferred that there is a mutual exclusion relationship between field content and field desc. It means that if one field appears, the other field cannot appear, and it is extracted as: [{name: content, mutually_exclusive: ["desc"]}, {name: desc, mutually_exclusive: ["content"]}].
[0098] Feature extraction of constraint conditions: If the value of a certain field is selected from a fixed list, it can be inferred that the field has an enumeration constraint. For example, if the value of a certain field in all data structure trees does not exceed 3 types, then we can infer that it is a field with an enumeration constraint, and the feature of the field constraint condition can be extracted.
[0099] For example: {name: status, constraints: {enum: [success, error, warning]}}; the 3 types (success, error, warning) here can be variable parameters. In the specific implementation, the number of enumeration constraints corresponding to the feature of the field constraint condition is not limited to 3 types, and the feature of the field constraint condition can be extracted according to the actual situation.
[0100] After extracting the feature information of each field, vectorize the feature information of each field (that is, for each field, integrate its corresponding feature information), and the corresponding feature vector of the field, so that the initial model can recognize and process the feature information of the field. For example: the feature information extracted for the field "name" includes type: string, constraints: {enum: [success, error, warning]}, then the feature vector generated for this field is {name: status, type: string, constraints: {enum: [success, error, warning]}}.
[0101] In some implementation manners of the embodiments of the present application, the initial model is used to generate a mapping vector set for each of the feature vectors; the mapping vector set includes a query vector, a key vector, and a value vector; calculate the attention weights between the target feature vector and other feature vectors; generate context information according to the attention weights of the target feature vector and the value vectors corresponding to the other feature vectors; generate a training verification library according to the context information of each of the target feature vectors;
[0102] Among them, the target feature vector is any one of the feature vectors.
[0103] The initial model can map each received feature vector into a set of mapped vectors, including a query vector Q i , a key vector K i and a value vector V i .
[0104] For the set of mapped vectors of each feature vector, the query vector Q i represents the information of other feature vectors (which can also be regarded as other fields) that the feature vector (which can also be regarded as the current field) needs to obtain; the key vector K i represents that the information of this feature vector can be provided to other feature vectors; the value vector V i represents the information of this feature vector to be passed to other feature vectors.
[0105] The initial model can generate vector pairs for the received combinations of feature vectors. Each vector pair includes two different feature vectors. It can be understood that the vector pairs also correspond to two different fields.
[0106] The initial model can sequentially determine the target feature vectors from the received feature vectors, and for each target feature vector, calculate the attention weights between the target feature vector and each vector pair, and generate the context information of the target feature vector based on each attention weight and the value vectors of other feature vectors. Then, based on the context information of each target feature vector, perform deduplication and aggregation to generate a training and verification library.
[0107] In some implementation manners of the embodiments of the present application, calculating the attention weights between the target feature vector and other feature vectors includes: determining the target feature vector in the set of feature vectors, and the feature vectors in the set of feature vectors other than the target feature vector as other feature vectors; calculating the similarity between the target feature vector and the other feature vectors; calculating the attention weights between the target feature vector and the other feature vectors based on each of the similarities.
[0108] Sequentially determine the target feature vectors from the feature vectors received by the initial model, and determine the feature vectors in the received feature vectors other than the target feature vector as other feature vectors. Calculate the similarity between the target feature vector and the other feature vector through the query vector of the target feature vector and the value vector of the other feature vector. Calculate the attention weights between the target feature vector and each of the other feature vectors based on the similarities between the target feature vector and each of the other feature vectors, so that the attention weights between the feature vectors with high similarity are greater.
[0109] In some implementation manners of the embodiments of the present application, for each of the other feature vectors with respect to the target feature vector, calculating the similarity between the target feature vector and the other feature vector includes: determining the query vector of the target feature vector as the target query vector; determining the device for the key vector of the other feature vector as the target key vector; calculating the similarity between the target feature vector and the other feature vector according to the similarity calculation formula, the target query vector, and the target key vector; the similarity calculation formula is Q i is the target query vector, is the target key vector.
[0110] The similarity between the target feature vector and the other feature vector can be calculated using the similarity formula. i is used to represent the field corresponding to the target feature vector, j is used to represent a field corresponding to one of the other feature vectors, and the feature vector corresponding to field j is one of the other feature vectors corresponding to the target feature vector.
[0111] The similarity S ij represents the similarity between the target feature vector and the feature vector corresponding to field j. Taking the dot product result of the target query vector and the target key vector as the value of the similarity, it is used to measure the correlation between the target feature vector and the other feature vector, that is, to measure the correlation between two fields, namely the field corresponding to the target feature vector and the field corresponding to the other feature vector.
[0112] In some implementation manners of the embodiments of the present application, calculating the attention weight between the target feature vector and the other feature vector according to each of the similarities includes:
[0113] Using the similarity transformation formula to perform normalization calculation on each of the similarities to determine the attention weight between the target feature vector and the other feature vector;
[0114] The similarity transformation formula is
[0115] where n is used to represent the number of feature vectors input to the initial model, k is the index information of the feature variable, i is used to represent the field corresponding to the target feature vector, and j is used to represent a field corresponding to one of the other feature vectors.
[0116] In practical applications, the initial model can use the similarity transformation formula to transform each similarity into an attention weight. k represents all the feature vectors involved in calculating the attention weight of a certain target feature vector.
[0117] For each similarity corresponding to the same target feature vector, a similarity formula is used for calculation to determine the attention weights between the target feature vector and each other feature vector. In the above similarity transformation formula, the denominator is used to sum the exponential transformation of the similarity scores calculated by traversing all feature vectors. The numerator is an exponential function that non-linearly amplifies larger values, making the attention weights corresponding to high similarities (large dot product values) larger.
[0118] In some implementation manners of the embodiments of the present application, generating context information based on each of the attention weights of the target feature vector and the value vectors corresponding to the other feature vectors includes:
[0119] For each field corresponding to the target feature vector, a context calculation formula is used to calculate each of the attention weights and the value vectors corresponding to the other feature vectors to generate context information;
[0120] The context calculation formula is
[0121] C i is used to represent the context information corresponding to field i of the target feature vector, j is used to represent a field corresponding to one of the other feature vectors, a ij is used to represent the attention weight corresponding to the target feature vector (the feature vector corresponding to field i) and an other feature vector (the feature vector corresponding to field j), V j is used to represent the value vector of the feature vector corresponding to field j.
[0122] After obtaining each attention weight of the target feature vector, the value vectors are weighted and summed using the attention weights to obtain the context representation of the field corresponding to the target feature vector (i.e., the above C i ).
[0123] After the initial model obtains the context representations corresponding to each field, the context information of each received feature vector is de-duplicated and aggregated to obtain a training verification library, and when it is determined that the training verification library passes the verification condition, the current training verification library is determined to be a target verification library that can be used to verify the format of the configuration file.
[0124] Referring to Figure 2 , a configuration file verification flow chart provided by the embodiments of the present application is shown, including the following steps:
[0125] YAML file collection: Obtain the YAML file that needs to be format-verified.
[0126] YAML file data preprocessing: Parse the YAML file to obtain a data structure tree.
[0127] Dataset division: Based on the above data structure tree, feature extraction of fields is performed to obtain the feature information of the fields, and the feature information is vectorized to obtain the feature vectors corresponding to each field. Multiple feature vectors are divided into a training set and a test set.
[0128] Build a Transformer model: Build a Transformer model as the initial model and initialize the model parameters of the Transformer model.
[0129] Train the model: Input the test set into the Transformer model and receive the training verification library output by the Transformer model.
[0130] Verify the model: Use the test set to verify the training verification library. If it passes, the result is output. If it does not pass, the model parameters are modified;
[0131] Modify the model parameters: Adjust the model parameters of the Transformer model according to the verification result until the training verification library passes the verification condition.
[0132] Output the result (JSON Schema): Determine the training verification library that passes the verification condition as the target verification library.
[0133] File verification: Through the above steps, we obtain a general JSON Schema (target verification library). Since the JSON Schema serves the JSON format, it is necessary to convert the content of the YAML file in the format to be verified into JSON data, and then verify the JSON data through the JSON Schema, so as to obtain whether it conforms to the JSON Schema rule verification and give the error reason defined in the JSON Schema.
[0134] For the use object of the configuration file, by using a part of the associated configuration file for training and testing the model, the model can obtain the format rules of the associated configuration file, and write the format rules into the JSONSchema to generate a final template (target verification library) to verify the formats of all associated configuration files in the project, so as to solve the problem that the existing YAML file verification method cannot effectively handle complex field types, field dependency relationships, and field constraint conditions; and reduce the workload of manually writing JSON Schema rules to verify the TAML file. When the YAML file format changes or new fields are added, the embodiments of the present application can quickly adjust the model and regenerate the rules through a small number of samples, and the adaptability is significantly better than the traditional method of manually writing fixed rules.
[0135] Refer to Figure 3, showing a schematic diagram of a format verification method provided by an embodiment of the present application, which may specifically include the following steps:
[0136] Step 301, obtain a second configuration file set;
[0137] Step 302, convert the second configuration file set into JSON data;
[0138] Step 303, use a target verification library to perform format verification on the JSON data to generate a verification result;
[0139] Among them, the target verification library is generated by the embodiment of the above verification library generation method.
[0140] The verification result can be verification passed or verification failed. In the case of verification failed, the verification result may further include information related to the fields in the configuration files in the second configuration file set that do not conform to the target verification library.
[0141] It should be noted that the magnitudes of the sequence numbers of the above steps do not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0142] Referring to Figure 4 , showing a schematic diagram of a verification library generation device provided by an embodiment of the present application, which may specifically include:
[0143] A feature vector set acquisition module 401, configured to acquire a feature vector set; the feature vector set includes a plurality of feature vectors determined from configuration files; the feature vectors are used to characterize the features of the fields of the configuration files;
[0144] A training module 402, configured to train an initial model according to the feature vector set;
[0145] A trained verification library receiving module 403, configured to receive a trained verification library output by the initial model for the feature vector set;
[0146] A target verification library determination module 404, configured to determine the trained verification library as the target verification library when the trained verification library passes a preset verification condition.
[0147] In some implementation manners of the embodiments of the present application, the feature vector set acquisition module 401 includes:
[0148] A first configuration file set acquisition sub-module, configured to acquire a first configuration file set;
[0149] A feature information extraction sub-module, configured to extract feature information of fields in the first configuration file set;
[0150] A vectorization sub-module, configured to vectorize the feature information to generate a feature vector corresponding to the field;
[0151] A feature vector set generation sub-module, configured to generate a feature vector set according to the feature vectors corresponding to each of the fields;
[0152] The feature information includes at least one of a field type feature, a necessary field feature, a field dependency relationship feature, a field mutual exclusion relationship feature, and a field constraint condition feature.
[0153] In some implementation manners of the embodiments of the present application, the initial model is used to generate a mapping vector set for each of the feature vectors; the mapping vector set includes a query vector, a key vector, and a value vector; calculate an attention weight between a target feature vector and other feature vectors; generate context information according to each of the attention weights of the target feature vector and the value vectors corresponding to the other feature vectors; generate a training and verification library according to the context information of each of the target feature vectors.
[0154] Wherein, the target feature vector is any one of the feature vectors.
[0155] In some implementation manners of the embodiments of the present application, calculating the attention weight between the target feature vector and other feature vectors includes:
[0156] Determine a target feature vector in the feature vector set, and the feature vectors in the feature vector set other than the target feature vector are other feature vectors;
[0157] Calculate the similarity between the target feature vector and the other feature vectors;
[0158] Calculate the attention weight between the target feature vector and the other feature vectors according to each of the similarities.
[0159] In some implementation manners of the embodiments of the present application, calculating the similarity between the target feature vector and the other feature vectors includes:
[0160] Determine the query vector of the target feature vector as the target query vector;
[0161] Determine the device of the key vector of the other feature vector as the target key vector;
[0162] Calculate the similarity between the target feature vector and the other feature vectors according to a similarity calculation formula, the target query vector, and the target key vector;
[0163] The similarity calculation formula is Qi is the target query vector, is the target key vector.
[0164] In some implementation manners of the embodiments of the present application, calculating the attention weights between the target feature vector and the other feature vectors according to each of the similarities includes:
[0165] Using a similarity conversion formula to perform a normalization calculation on each of the similarities to determine the attention weights between the target feature vector and the other feature vectors;
[0166] The similarity conversion formula is
[0167] where n is used to represent the number of feature vectors input to the initial model, k is the index information of the feature variable, i is used to represent the field corresponding to the target feature vector, and j is used to represent the field corresponding to one of the other feature vectors.
[0168] In some implementation manners of the embodiments of the present application, generating context information according to each of the attention weights of the target feature vector and the value vectors corresponding to the other feature vectors includes:
[0169] For each field corresponding to the target feature vector, using a context calculation formula to calculate each of the attention weights and the value vectors corresponding to the other feature vectors to generate context information;
[0170] The context calculation formula is
[0171] C i is used to represent the context information corresponding to the field i corresponding to the target feature vector, j is used to represent the field corresponding to one of the other feature vectors, a ij is used to represent the attention weight corresponding to the target feature vector and one of the other feature vectors, V j is used to represent the value vector of the feature vector corresponding to the field j.
[0172] A verification library generation device provided by an embodiment of the present application. By applying this device, each step in the foregoing embodiments of the verification library generation method can be implemented. Referring to Figure 5 , a schematic diagram of a format verification device provided by an embodiment of the present application is shown, which may specifically include:
[0173] A second configuration file set acquisition module 501, configured to acquire a second configuration file set;
[0174] A format conversion module 502, configured to convert the second configuration file set into JSON data;
[0175] A verification module 503, configured to perform format verification on the JSON data by using a target verification library, and generate a verification result;
[0176] Wherein, the target verification library is generated by the generating device embodiment of the above verification library.
[0177] A format verification device provided by an embodiment of the present application. By applying this device, each step in the foregoing format verification method embodiment can be implemented.
[0178] For the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For related parts, refer to the description in the method embodiment part.
[0179] Referring to Figure 6 , a schematic diagram of an electronic device provided by an embodiment of the present application is shown. As Figure 6 shown, the electronic device 600 in the embodiment of the present application includes: a processor 610, a memory 620, and a computer program 621 stored in the memory 620 and executable on the processor 610. When the processor 610 executes the computer program 621, the steps in each embodiment of the above verification library generation method and / or format verification method are implemented.
[0180] Exemplarily, the computer program 621 can be divided into one or more modules / units. The one or more modules / units are stored in the memory 620 and executed by the processor 610 to complete the present application. The one or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and the instruction segments can be used to describe the execution process of the computer program 621 in the electronic device 600.
[0181] The electronic device 600 can be a computing device such as a desktop computer or a cloud server. The electronic device 600 may include, but is not limited to, a processor 610 and a memory 620. Those skilled in the art can understand that Figure 6 this is only an example of the electronic device 600, and does not constitute a limitation on the electronic device 600. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the electronic device 600 may further include input / output devices, network access devices, a bus, etc.
[0182] The processor 610 may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0183] The memory 620 may be an internal storage unit of the electronic device 600, such as the hard disk or memory of the electronic device 600. The memory 620 may also be an external storage device of the electronic device 600, such as a plug-in hard disk equipped on the electronic device 600, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 620 may also include both the internal storage unit and the external storage device of the electronic device 600. The memory 620 is used to store the computer program 621 and other programs and data required by the electronic device 600. The memory 620 may also be used to temporarily store data that has been output or is to be output.
[0184] The embodiments of the present application also disclose a computer-readable storage medium storing a computer program, which when executed by a processor implements the method for generating a verification library or the format verification method as described in the foregoing various embodiments.
[0185] The embodiments of the present application also disclose a computer program product including a computer program, which when run causes the method for generating a verification library or the format verification method as described in the foregoing various embodiments to be executed.
[0186] The above-described embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit the same. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the various embodiments of the present application and should all be included within the protection scope of the present application.
Claims
1. A method for generating a verification library, characterized in that: include: Acquire a feature vector set; the feature vector set includes a plurality of feature vectors determined from a configuration file; The feature vector is used to characterize the features of the field of the configuration file; Training the initial model according to the feature vector set; Receiving a training validation library of outputs of the trained initial model for the feature vector set; In the case where the training verification library passes the preset verification condition, the training verification library is determined to be the target verification library.
2. The method according to claim 1, characterized in that The step of obtaining a feature vector set comprises: Get a first configuration file set; Extracting feature information of fields in the first configuration file set; Vectorizing the feature information to generate a feature vector corresponding to the field; Generate a feature vector set according to the feature vectors corresponding to each of the fields; The characteristic information includes: at least one of: field type characteristics, necessary field characteristics, field dependency characteristics, field mutual exclusion characteristics, and field constraint condition characteristics.
3. The method according to claim 1, characterized in that The initial model is used to generate a mapping vector set for each of the feature vectors; the mapping vector set includes a query vector, a key vector and a value vector; calculate the attention weight between the target feature vector and other feature vectors; generate context information based on each of the attention weights of the target feature vector and the value vectors corresponding to the other feature vectors; generate a training verification library based on the context information of each of the target feature vectors; The target feature vector is any one of the feature vectors.
4. The method according to claim 3, characterized in that The calculation of the attention weight between the target feature vector and other feature vectors includes: Determine a target feature vector in the feature vector set, and feature vectors other than the target feature vector in the feature vector set are other feature vectors; Calculating the similarity between the target feature vector and the other feature vectors; According to each of the similarities, an attention weight between the target feature vector and the other feature vectors is calculated.
5. The method according to claim 4, characterized in that The calculating the similarity between the target feature vector and the other feature vectors includes: Determining a query vector of the target feature vector as a target query vector; means for determining a key vector of said other feature vectors as a target key vector; Calculate the similarity between the target feature vector and the other feature vectors according to a similarity calculation formula, the target query vector, and the target key vector; The similarity calculation formula is: Q i is the target query vector, is the target key vector.
6. The method according to claim 4, characterized in that The calculating the attention weight between the target feature vector and the other feature vectors according to each of the similarities includes: Using a similarity conversion formula, normalizing and calculating each of the similarities to determine the attention weight between the target feature vector and the other feature vectors; The similarity conversion formula is: Among them, n is used to represent the number of feature vectors input into the initial model, k is the index information of the feature variable, i is used to represent the field corresponding to the target feature vector, and j is used to represent a corresponding field in other feature vectors.
7. The method according to any one of claims 3 to 6, characterized in that: The generating context information according to each of the attention weights of the target feature vector and the value vectors corresponding to the other feature vectors includes: For each field corresponding to the target feature vector, a context calculation formula is used to calculate each attention weight and the value vector corresponding to the other feature vectors to generate context information; The context calculation formula is: C i It is used to represent the context information corresponding to field i of the target feature vector, and j is used to represent a corresponding field in other feature vectors. ij The attention weight V is used to represent the corresponding target feature vector and another feature vector. j A value vector representing the feature vector corresponding to field j.
8. A format verification method, characterized in that: include: Get a second configuration file set; Convert the second configuration file set into JSON data; Use the target verification library to perform format verification on the JSON data and generate a verification result; Wherein, the target verification library is generated by the method described in any one of claims 1-7.
9. An electronic device, characterized in that: The electronic device comprises a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the electronic device implements the method according to any one of claims 1 to 8.
10. A computer program product, characterized in that The invention comprises a computer program, which, when being executed, enables the method according to any one of claims 1 to 8 to be performed.