Generation method and device for sample data and neural network
By converting the high-level language code into intermediate language code and encoding the two, the generated target sample data contains richer feature information, solving the problem of insufficient feature information in the prior art, meeting the needs of complex task scenarios, and improving the generalization ability of neural networks.
Patent Information
- Application Number
- CN202210190136.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-28
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-02-28
AI Technical Summary
When the prior art converts code into feature representation, it contains less feature information, making it difficult to meet the needs of complex tasks.
By obtaining sample code written in high-level languages, converting it into code written in preset intermediate languages, and encoding the two forms of code to generate target sample data containing deeper feature information.
The generated target sample data covers more in-depth and comprehensive feature information, which can meet the usage needs in complex tasks, and improve the generalization ability of subsequent neural networks.
Smart Images

Figure CN114546361B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and in particular, to a method for generating sample data, a method for generating a neural network, a device, a computer device, and a storage medium. Background Art
[0002] With the rapid development of computer and software technologies, people's lives have enjoyed the convenience brought by the Internet era, and the software technologies are embodied in the codes written by thousands of programmers. The development of artificial intelligence today also enables us to analyze and process languages in the form of codes to perform intelligent programming of codes more efficiently, such as code Application Programming Interface (API) mining, naming convention learning, error localization, code summarization, comment generation, code search, code repair, compiler optimization, etc.; to analyze and process languages in the form of codes, it is necessary to pre-convert the codes into feature representations, and use the codes converted into feature representations as sample data for supervised data of various tasks; however, the current method for converting codes into feature representations has the problem of less included feature information. Summary of the Invention
[0003] Embodiments of the present disclosure at least provide a method for generating sample data, a method for generating a neural network, a device, a computer device, and a storage medium.
[0004] In a first aspect, an embodiment of the present disclosure provides a method for generating sample data, including: obtaining a first sample code written in a high-level language; converting the first sample code into a second sample code written in a preset intermediate language; performing a first encoding process on the first sample code to obtain first sample data; and performing a second encoding process on the second sample code to obtain second sample data; generating target sample data corresponding to the first sample code based on the first sample data and the second sample data.
[0005] In this way, the generated target sample data not only includes the first sample data obtained by performing a first encoding process on the first sample code written in a high-level language, but also converts the first sample code into a second sample code written in a preset intermediate language, and the intermediate language is more concise, and the function and role of the code can be inferred from the second sample code in this form. Therefore, the second sample data obtained by performing a second encoding process on the second sample code written in the preset intermediate language is included in the target sample data, so that the feature information covered by the sample data is deeper and more comprehensive. Furthermore, the target sample data generated by using the method for generating sample data provided by the embodiments of the present disclosure can meet the usage requirements in complex task scenarios.
[0006] In an alternative implementation, the conversion of the first sample code into the second sample code written in a preset intermediate language includes: converting the first sample code into the second sample code based on the conversion relationship between the high-level language and the preset intermediate language.
[0007] In this way, the first sample code written in a relatively complex and high-level high-level language is converted into the second sample code written in a relatively simple and low-level preset intermediate language. Since the preset intermediate language is more concise, the function and role of the code can be inferred from the second sample code in this form.
[0008] In an alternative implementation, the first encoding process on the first sample code to obtain the first sample data includes: determining a target string from the first sample code, and converting the target string into a standard string corresponding to the target string to obtain the converted first sample code; the target string includes: variable names and / or function names; performing word segmentation on the converted first sample code to generate a first string sequence corresponding to the converted first sample code; and performing a first encoding process on the first string sequence to obtain the first sample data.
[0009] In an alternative implementation, the first string sequence includes at least one standard string; the first encoding process on the first string sequence to obtain the first sample data includes: querying, from a pre-generated first encoding table, the first encoding values respectively corresponding to the standard strings in the first string sequence; and generating the first sample data based on the first encoding values respectively corresponding to the standard strings in the first string sequence.
[0010] In an alternative implementation, the generating the first sample data based on the first encoding values respectively corresponding to the standard strings in the first string sequence includes: constructing a first sparse matrix based on the first encoding values respectively corresponding to the standard strings in the first string sequence; converting the first sparse matrix into a first dense matrix with a preset dimension; and determining the first dense matrix as the first sample data.
[0011] In an alternative implementation, the second sample code includes: a second string sequence composed of the intermediate language; the second encoding process on the second sample code to obtain the second sample data includes: querying, from a pre-generated second encoding table, the second encoding values respectively corresponding to the strings in the second string sequence; and generating the second sample data based on the second encoding values respectively corresponding to the strings in the second string sequence.
[0012] In an alternative embodiment, generating the second sample data based on the second encoding values respectively corresponding to the strings in the second string sequence includes: constituting a second sparse matrix based on the second encoding values respectively corresponding to the strings in the second string sequence; converting the second sparse matrix into a second dense matrix with a preset dimension; and determining the second dense matrix as the second sample data.
[0013] In an alternative embodiment, generating the target sample data corresponding to the first sample code based on the first sample data and the second sample data includes: performing a fusion process on the first sample data and the second sample data to obtain the target sample data.
[0014] In this way, by performing a fusion process on the first sample data and the second sample data, the obtained target sample data not only contains relatively shallow-level feature information, but also contains code role feature information and code function feature information, making the feature information covered by the sample data more in-depth and comprehensive. Furthermore, the target sample data generated by using the sample data generation method provided in the embodiments of the present disclosure can meet the usage requirements in task complex scenarios. In addition, it is convenient to generate a target neural network with strong generalization ability based on the target sample data subsequently.
[0015] In a second aspect, an embodiment of the present disclosure further provides a method for generating a neural network, including: obtaining target sample data; the target sample data is generated by using the sample data generation method as described in any one of the first aspect; and using the target sample data to train a neural network to be trained to obtain a target neural network.
[0016] In this way, the target neural network generated by training with the target sample data has strong generalization ability and can meet the requirements of various types of tasks.
[0017] In a third aspect, an embodiment of the present disclosure further provides a sample data generation device, including: a first acquisition module, configured to acquire a first sample code written in a high-level language; a conversion module, configured to convert the first sample code into a second sample code written in a preset intermediate language; an encoding module, configured to perform a first encoding process on the first sample code to obtain first sample data; and perform a second encoding process on the second sample code to obtain second sample data; and a generation module, configured to generate target sample data corresponding to the first sample code based on the first sample data and the second sample data.
[0018] In an alternative embodiment, when the conversion module executes the conversion of the first sample code into a second sample code written in a preset intermediate language, it is specifically configured to: based on the conversion relationship between the high-level language and the preset intermediate language, convert the first sample code into the second sample code.
[0019] In an alternative embodiment, when the encoding module executes the first encoding process on the first sample code to obtain the first sample data, it is specifically configured to: determine a target string from the first sample code, and convert the target string into a standard string corresponding to the target string to obtain the converted first sample code; the target string includes: variable names and / or function names; perform word segmentation on the converted first sample code to generate a first string sequence corresponding to the converted first sample code; perform a first encoding process on the first string sequence to obtain the first sample data.
[0020] In an alternative embodiment, at least one standard string is included in the first string sequence; when the encoding module executes the first encoding process on the first string sequence to obtain the first sample data, it is specifically configured to: query, from a pre-generated first encoding table, the first encoding values respectively corresponding to the standard strings in the first string sequence; based on the first encoding values respectively corresponding to the standard strings in the first string sequence, generate the first sample data.
[0021] In an alternative embodiment, when the encoding module executes the generation of the first sample data based on the first encoding values respectively corresponding to the standard strings in the first string sequence, it is specifically configured to: based on the first encoding values respectively corresponding to the standard strings in the first string sequence, form a first sparse matrix; convert the first sparse matrix into a first dense matrix with a preset dimension; determine the first dense matrix as the first sample data.
[0022] In an alternative embodiment, a second string sequence composed of the intermediate language is included in the second sample code; when the encoding module executes the second encoding process on the second sample code to obtain the second sample data, it is specifically configured to: query, from a pre-generated second encoding table, the second encoding values respectively corresponding to the strings in the second string sequence; based on the second encoding values respectively corresponding to the strings in the second string sequence, generate the second sample data.
[0023] In an alternative embodiment, when generating the second sample data based on the second encoding values respectively corresponding to the individual strings in the second string sequence, the encoding module is specifically configured to: construct a second sparse matrix based on the second encoding values respectively corresponding to the individual strings in the second string sequence; convert the second sparse matrix into a second dense matrix with a preset dimension; and determine the second dense matrix as the second sample data.
[0024] In an alternative embodiment, when generating the target sample data corresponding to the first sample code based on the first sample data and the second sample data, the generating module is specifically configured to: perform a fusion process on the first sample data and the second sample data to obtain the target sample data.
[0025] In a fourth aspect, an embodiment of the present disclosure further provides a generating device for a neural network, including: a second obtaining module, configured to obtain target sample data; the target sample data is generated by using the sample data generating method according to any one of the first aspect; and a training module, configured to train a neural network to be trained by using the target sample data to obtain a target neural network.
[0026] In a fifth aspect, an alternative implementation of the present disclosure further provides a computer device, including a processor and a memory, where the memory stores machine-readable instructions executable by the processor, and the processor is configured to execute the machine-readable instructions stored in the memory. When the machine-readable instructions are executed by the processor, the machine-readable instructions are executed to perform the steps in the first aspect, or any possible implementation manner in the first aspect, or the steps in the implementation manner shown in the second aspect.
[0027] In a sixth aspect, an alternative implementation of the present disclosure further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run, it executes the steps in the first aspect, or any possible implementation manner in the first aspect, or the steps in the implementation manner shown in the second aspect.
[0028] For the effect descriptions of the above sample data generating device, neural network generating device, computer device, and computer-readable storage medium, please refer to the descriptions of the above sample data generating method and neural network generating method respectively, and details are not described herein again.
[0029] To make the above objects, features, and advantages of the present disclosure more obvious and understandable, the following specific embodiments are given, and detailed descriptions are made in conjunction with the accompanying drawings as follows. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the accompanying drawings required for use in the embodiments will be briefly introduced below. These accompanying drawings are incorporated into the specification and constitute a part of this specification. These accompanying drawings show embodiments that conform to the present disclosure and are used together with the specification to illustrate the technical solutions of the present disclosure. It should be understood that the following accompanying drawings only show some embodiments of the present disclosure and should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other relevant accompanying drawings can also be obtained based on these accompanying drawings.
[0031] Figure 1 The flowchart of a method for generating sample data provided by an embodiment of the present disclosure is shown;
[0032] Figure 2 The structural schematic diagram of a first sample data generator provided by an embodiment of the present disclosure is shown;
[0033] Figure 3 The structural schematic diagram of a second sample data generator provided by an embodiment of the present disclosure is shown;
[0034] Figure 4 The flowchart of a method for generating a neural network provided by an embodiment of the present disclosure is shown;
[0035] Figure 5 The schematic diagram of a device for generating sample data provided by an embodiment of the present disclosure is shown;
[0036] Figure 6 The schematic diagram of a device for generating a neural network provided by an embodiment of the present disclosure is shown;
[0037] Figure 7 The schematic diagram of a computer device provided by an embodiment of the present disclosure is shown. Detailed implementation manners
[0038] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only some embodiments of the present disclosure, rather than all embodiments. Usually, the components of the embodiments of the present disclosure described and shown here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present disclosure is not intended to limit the scope of the present disclosure to be protected, but only represents the selected embodiments of the present disclosure. Based on the embodiments of the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present disclosure.
[0039] It has been found through research that common ways to generate sample data from code include: regarding the code as a form similar to natural language, converting it into character tokens represented by numbers, and then performing embedding processing on the tokens to obtain the sample data of the code. The sample data obtained in this way only considers the shallow meaning of the code, so the feature information included in the code is less. Such sample data can only be applied to scenarios with relatively simple tasks, such as tasks like finding typos in code; while for scenarios with complex tasks, such as code summarization, comment generation, etc., it is required to be able to convert the code into sample data containing more feature information; in this way, a neural network model with stronger generalization ability can be trained using the sample data converted from the code, and the current ways to convert code into sample data cannot meet the usage requirements of the current scenarios with complex tasks.
[0040] Based on the above research, the present disclosure provides a method for generating sample data, a method for generating a neural network, an apparatus, a computer device, and a storage medium. The generated target sample data includes the first sample data after performing a first encoding process on the first sample code written in a high-level language. The method for generating sample data provided in the embodiments of the present disclosure will also convert the first sample code into a second sample code written in a preset intermediate language. Since the intermediate language is more concise and the function and role of the code can be inferred from the second sample code in this form, therefore, the target sample data also includes the second sample data after performing a second encoding process on the second sample code written in the preset intermediate language, making the feature information covered by the sample data deeper and more comprehensive. Furthermore, the target sample data generated by using the method for generating sample data provided in the embodiments of the present disclosure can meet the usage requirements in scenarios with complex tasks.
[0041] Regarding the defects existing in the above solutions, they are all the results obtained by the inventors after practice and careful research. Therefore, the process of discovering the above problems and the solutions proposed by the present disclosure for the above problems in the following text should all be the contributions made by the inventors during the process of the present disclosure.
[0042] It should be noted that: similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0043] To facilitate the understanding of this embodiment, first, a method for generating sample data disclosed in the embodiments of the present disclosure will be introduced in detail, and then a method for generating a neural network disclosed in the embodiments of the present disclosure will be introduced in detail. The execution subject of the method for generating sample data and the method for generating a neural network provided in the embodiments of the present disclosure is generally a computer device with certain computing capabilities. Such a computer device may include, for example: a terminal device, a server, or other processing devices. Among them, the terminal device may be a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. In some possible implementation manners, the generation of the sample data and the method for generating a neural network may be implemented by a processor invoking computer-readable instructions stored in a memory.
[0044] The method for generating sample data provided in the embodiments of the present disclosure performs a first encoding process on a first sample code written in a high-level language to obtain first sample data corresponding to the first sample code. The first sample data contains the shallow meaning corresponding to the first sample code and can be applied to the task scenario of finding typos in the code. In addition, the method for generating sample data provided in the embodiments of the present disclosure also converts the first sample code written in a high-level language into a second sample code written in a preset intermediate language. Here, since the preset intermediate language is more concise and is easier to analyze the role and function of the code subsequently; therefore, by performing a second encoding process on the second sample data written in the preset intermediate language code, second sample data corresponding to the second sample code is obtained. The second sample data contains deeper meanings corresponding to the second sample code, such as code role feature information, code function feature information, etc., and can be applied to relatively complex task scenarios, such as code summarization, comment generation, and other task scenarios.
[0045] The method for generating sample data provided in the embodiments of the present disclosure will be described below.
[0046] See Figure 1 As shown, it is a flowchart of a method for generating sample data provided in the embodiments of the present disclosure. The method includes steps S101 to S104, where:
[0047] S101. Obtain a first sample code written in a high-level language.
[0048] Among them, the high-level language may include, for example, but is not limited to computer programming languages, such as C language, Python language, etc.; the first sample code is marked with annotation information corresponding to the current task scenario.
[0049] Exemplarily, if the current scenario includes a task scenario of finding misspelled characters in code, the corresponding annotation information includes: the positions where there are misspelled characters in the first sample code, and / or the specific characters in the code; if the current scenario includes a code annotation scenario, the corresponding annotation information includes the annotation information for annotating the first sample code.
[0050] Exemplarily, the obtained first sample code includes: for different task scenarios, using a high-level language to write code corresponding to the task scenario, and through means such as manual annotation, annotating the annotation information corresponding to the task scenario in the code, and generating the first sample code.
[0051] Continuing from the above S101, the method for generating sample data provided by the embodiments of the present disclosure further includes:
[0052] S102. Convert the first sample code into a second sample code written in a preset intermediate language.
[0053] Among them, the intermediate language is a language that is easy to translate the first sample code into an equivalent internal representation, and it is easier to understand; it is used to represent the first sample code in a more understandable, relatively concise, and easy-to-analyze code form; the preset intermediate language may include but is not limited to the LLVM IR language; the LLVM IR language is a general intermediate language, which is generated by using the LLVM compiler to transform a high-level language. Among them, the LLVM compiler may include the following three parts: the front end, the optimizer, and the back end; the front end is responsible for analyzing the source code, can check syntax-level errors, and construct an abstract syntax tree; the optimizer is used to optimize the abstract syntax tree to obtain an optimized abstract syntax tree; the back end is used to process the optimized abstract syntax tree to generate executable machine code.
[0054] In a specific implementation, based on the conversion relationship between the high-level language and the preset intermediate language, the first sample code can be converted into the second sample code.
[0055] Exemplarily, the LLVM compiler can first analyze the obtained first sample code, check the syntax-level errors in the first sample code, and construct the abstract syntax tree of the first sample code; then optimize and process the abstract syntax tree to generate a second sample code written in the LLVM IR language.
[0056] Continuing from the above S102, the method for generating sample data provided by the embodiments of the present disclosure further includes:
[0057] S103. Perform a first encoding process on the first sample code to obtain first sample data; and perform a second encoding process on the second sample code to obtain second sample data.
[0058] In a specific implementation, the following method can be used to perform a first encoding process on the first sample code to obtain first sample data: Determine a target string from the first sample code, and convert the target string into a standard string corresponding to the target string to obtain the first sample code after conversion; perform word segmentation on the first sample code after conversion to generate a first string sequence corresponding to the first sample code after conversion; perform a first encoding process on the first string sequence to obtain the first sample data.
[0059] Among them, the target string may include, but is not limited to, variable names and / or function names in the first sample code, which are defined by the programmer himself or determined based on the functions of the corresponding high-level language. For example, for variable names, in different first sample codes, the variable names will be different; for user-defined functions, in different first sample codes, the function names may also be different; due to the randomness of the programmer's self-defined variable names and / or function names, it is difficult to set an encoding value for each variable name and / or function name in the first encoding table, resulting in limited use of the first encoding table, and further limiting the generation process of sample data within a certain code range.
[0060] The variable names and / or function names that are not common in different first sample codes can be converted into common standard strings to obtain the first sample code after conversion, and then the first encoding table is used to perform a first encoding process on the first sample code after conversion to enhance the generality of the sample data generation process.
[0061] The standard string is a string that can be found in the pre-generated first encoding table. The standard string can include a single character (which can also be called a word, that is, a token), or can include multiple characters, and can be specifically set according to actual needs, and no specific restrictions are made here.
[0062] Among them, the first encoding table is generated for the first sample code after conversion, and can convert the standard strings in the first sample code after conversion into numbers. In this first encoding table, different strings correspond to different encoding values composed of numbers.
[0063] For example, in the first coding table, if the standard string is "AAA", the corresponding code is "1234"; if the standard string is "BBB", the corresponding code is "2556"; if the standard string is "CCC", the corresponding code is "3647". If the string sequence obtained after word segmentation of the converted first sample code includes: "AAA", "CCC", "BBB", then after performing the first coding process on the above string sequence using the first coding table, the first sample data obtained includes: 123436472556.
[0064] Exemplarily, in order to facilitate encoding the first sample code using the first coding table, the variable names and / or function names in the first sample code can be determined as the target strings in the first sample code, and the target strings can be converted into the standard strings existing in the pre-generated first coding table to obtain the converted first sample code.
[0065] Among them, the pre-generated first coding table can include each standard string and the corresponding coding value of each standard string; the standard string can include at least one character. Specifically, the pre-generated first coding table can be set according to actual needs and is not specifically limited here. For example, the pre-generated first coding table can be as shown in Table 1:
[0066] Table 1
[0067] idx token idx token idx token idx token 1 A 5 , 9 c 13 } 2 int 6 b 10 = 14 10 3 ( 7 ) 11 ; 15 return 4 a 8 { 12 + .... ....
[0068] Among them, Table 1 includes multiple standard characters, where token is used to indicate each standard character, and idx is used to indicate the coding value corresponding to each standard character.
[0069] When converting the target string into the standard string corresponding to the target string, first, a conversion relationship between the target string and the standard string is established. Using this conversion relationship, after obtaining the target string, the target string can be converted into the standard string.
[0070] Exemplarily, the first sample code includes: int f(int num1, int num2) { int c = 10; return num1 + num2 + c;}, determining that the target strings in the first sample code include: f, mum1, num2; converting the target strings in the first sample code into standard strings corresponding to the target strings. For example, if the standard string corresponding to the target string f is A, then convert f to A; the standard string corresponding to the target string num1 is a, then convert num1 to a; the standard string corresponding to the target string mum2 is b, then convert mum2 to b, and the converted first sample code is: int A(int a, int b) { int c = 10; return a + b + c;}.
[0071] In a specific implementation, after obtaining the converted first sample code, use a lexical analyzer to perform word segmentation processing on the converted first sample code, represent the converted first sample code in the form of a string sequence, and generate a first string sequence corresponding to the converted first sample code.
[0072] Among them, the lexical analyzer may include but is not limited to: a program or function for converting a string into a character sequence; the first string sequence may include at least one standard string, where each standard string may be composed of a single character or multiple characters; in the case where each string in the first string sequence is composed of a single character, the first string sequence may also be referred to as a first character sequence, that is, a token sequence.
[0073] In a specific implementation, after generating the first string sequence corresponding to the converted first sample code, the following method can be used to perform a first encoding process on the first string sequence to obtain first sample data: query the first encoding values corresponding to each standard string in the first string sequence from a pre-generated first encoding table; based on the first encoding values corresponding to each standard string in the first string sequence, generate first sample data. Exemplarily, assuming that the pre-generated first encoding table is shown in Table 1, then according to Table 1, the first encoding values corresponding to each standard string in the first string sequence can be queried, that is, 2, 1, 3, 2, 4, 5, 2, 6, 7, 8, 2, 9, 10, 14, 11, 15, 4, 12, 6, 12, 9.
[0074] In practical applications, due to the large number of standard strings, when setting the first encoding value for each annotation string, there will be a large number of 0s in the first encoding value. Using the first encoding value to directly form the first sample data will result in a large number of 0s in the first sample data, that is, the first sparse matrix. For computer storage and access, more memory access time is required to access the sparse matrix. To reduce the dimension of the first sample data, in another embodiment of the present disclosure, after determining the first encoding value corresponding to each standard string in the first string sequence based on the pre-generated first encoding table, the following method can be used to generate the first sample data based on the first encoding value corresponding to each standard string in the first string sequence: Based on the first encoding value corresponding to each standard string in the first string sequence, form the first sparse matrix; convert the first sparse matrix into a first dense matrix with a preset dimension; determine the first dense matrix as the first sample data.
[0075] Among them, the dimension of the first dense matrix with a preset dimension can be set according to actual needs, and no specific limitation is made here.
[0076] Exemplarily, after generating the first sparse matrix based on the first encoding value corresponding to each standard string in the first string sequence, the embedding layer (i.e., the embedding layer) can be used to convert the first sparse matrix into a first dense matrix with a preset dimension to generate the first sample data.
[0077] Exemplarily, when using the embedding layer to convert the first sparse matrix into a first dense matrix with a preset dimension, in fact, the first sparse matrix is multiplied by a preset conversion matrix to obtain the first dense matrix.
[0078] Exemplarily, if the first sparse matrix is:
[0079] The conversion matrix is: Then multiplying the two matrices gives the first sample data as:
[0080] Exemplarily, the structural schematic diagram of the first sample data generator 20 for performing the first encoding process on the first sample code can be as Figure 2 shown, including: a preprocessing unit 21, a lexical analyzer 22, a first encoding value determination unit 23, and an embedding layer 24; wherein, the output end of the preprocessing unit 21 is connected to the input end of the lexical analyzer 22, the output end of the lexical analyzer 22 is connected to the input end of the first encoding value determination unit 23, and the output end of the first encoding value determination unit 23 is connected to the input end of the embedding layer 24.
[0081] Among them, the preprocessing unit 21 is used to determine a target string from the first sample code, convert the target string into a standard string to obtain the converted first sample code, and transmit the converted first sample code to the lexical analyzer 22; the lexical analyzer 22 is used to perform word segmentation on the received converted first sample code to obtain a first string sequence corresponding to the converted first sample code, and transmit the first string sequence to the first encoding value determination unit 23; the first encoding value determination unit 23, which can also be called a lookup table unit, is used to receive the first string sequence, and based on a pre-generated first encoding table, look up the first encoding values corresponding to each standard string in the first string sequence, and transmit the first encoding values corresponding to each standard string in the first string sequence to the embedding layer 23; the embedding layer 23 is used to receive the first encoding values corresponding to each standard string in the first string sequence, and based on the first encoding values corresponding to each standard string in the first string sequence, construct a first sparse matrix, convert the first sparse matrix into a first dense matrix with a preset dimension, and determine the first dense matrix as the first sample data.
[0082] In a specific implementation, the second sample code can be second-encoded in, but not limited to, the following manner to obtain second sample data: query the second encoding values corresponding to each string in the second string sequence from a pre-generated second encoding table; generate second sample data based on the second encoding values corresponding to each string in the second string sequence.
[0083] Among them, the second sample code includes: a second string sequence composed of intermediate language; the second string sequence includes at least one string, and each string can be composed of at least one character; the pre-generated second encoding table can include multiple strings composed of intermediate language and the second encoding values corresponding to each string. The specific pre-generated second encoding table can be set according to actual needs and will not be specifically limited here. For example, the pre-generated second encoding table can be as shown in Table 2:
[0084] Table 2
[0085] idx token idx token idx token idx token 1 define 5 # 9 3 13 } 2 i32 6 O 10 = 14 align 3 @ 7 % 11 alloca 15 4 4 A 8 { 12 ,... . ....
[0086] Among them, Table 2 includes multiple strings composed of intermediate language, where token is used to indicate each string, and idx is used to indicate the encoding value corresponding to each string.
[0087] Exemplarily, in Table 2, the second encoding values corresponding to each string in the second sample code written in the intermediate language LLVM IR can be looked up.
[0088] In a specific implementation, after determining the second encoding values corresponding to the respective strings included in the second string sequence, the second sample data may be generated in the following manner: Based on the second encoding values corresponding to the respective strings in the second string sequence, a second sparse matrix is constructed; the second sparse matrix is converted into a second dense matrix with a preset dimension; and the second dense matrix is determined as the second sample data.
[0089] Among them, the dimension corresponding to the second dense matrix with a preset dimension can be set according to actual needs, and no specific limitation is made here.
[0090] Exemplarily, after generating the second sparse matrix based on the second encoding values corresponding to the respective strings in the second string sequence, an embedding layer (i.e., the embedding layer) can be used to convert the second sparse matrix into a second dense matrix with a preset dimension to generate the second sample data.
[0091] Exemplarily, a schematic structural diagram of the second sample data generator 30 for performing second encoding processing on the second sample code may be as Figure 3 shown, including: an intermediate language conversion unit 31 (such as, but not limited to, LLVM IRBuilder), a second encoding value determination unit 32, and an embedding layer 33; wherein, the output end of the intermediate language conversion unit 31 is connected to the input end of the second encoding value determination unit 32, and the output end of the second encoding value determination unit 32 is connected to the input end of the embedding layer 33.
[0092] Among them, the intermediate language conversion unit 31 is used to convert the first sample code written in a high-level language into a second sample code written in an intermediate language based on the conversion relationship between the high-level language and the preset intermediate language, and transmit the second sample code to the second encoding value determination unit 32; the second encoding value determination unit 32, which may also be referred to as a lookup table unit, is used to receive the second sample code, look up the second encoding values corresponding to the respective strings in the second sample code based on the pre-generated second encoding table, and transmit the second encoding values corresponding to the respective strings in the second sample code to the embedding layer 33; the embedding layer 33 is used to receive the second encoding values corresponding to the respective strings in the second sample code, form a second sparse matrix with the second encoding values corresponding to the respective strings in the second sample code, convert the second sparse matrix into a second dense matrix with a preset dimension, and determine the second dense matrix as the second sample data.
[0093] In a specific implementation, after generating the first sample data and the second sample data based on the specific implementation manner shown in S103 above, the method for generating sample data provided by the embodiments of the present disclosure further includes:
[0094] S104. Generate target sample data corresponding to the first sample code based on the first sample data and the second sample data.
[0095] In a specific implementation, the first sample data and the second sample data can be fused to obtain the target sample data.
[0096] Among them, the fusion process can include but is not limited to at least one of the following: splicing process, parallel connection process, weighted summation process, etc.; the specific fusion process can be set according to actual needs and is not specifically limited here.
[0097] Exemplarily, the splicing process is, for example, splicing the first sample data and the second sample data in height or width. Among them, if splicing in height, the first sample data and the second sample data need to be consistent in width and data channels; if splicing in width, the first sample data and the second sample data need to be consistent in height and data channels. (If they are inconsistent, the sample data with a smaller size can be padded with 0 so that the padded sample data and the unpadded sample data meet the size requirements for splicing).
[0098] For example, splicing the first sample data and the second sample data in height, and the size of the first sample data is w (width) * h1 (height) * c (number of data channels), and the size of the second sample data is: w (width) * h2 (height) * c (number of data channels). After splicing the first sample data and the second sample data in height, the size of the obtained target sample data is: w * (h1 + h2) * c.
[0099] For example, splicing the first sample data and the second sample data in width, and the size of the first sample data is w1 (width) * h (height) * c (number of data channels), and the size of the second sample data is: w2 (width) * h (height) * c (number of data channels). After splicing the first sample data and the second sample data in width, the size of the obtained target sample data is: (w1 + w2) * h * c.
[0100] The parallel connection process is, for example, superimposing the first sample data and the second sample data on the data channels, which requires that the height and width of the first sample data and the second sample data are both the same. Similarly, if they are inconsistent in height and width, the sample data with a smaller size can be padded with 0 so that the padded sample data and the unpadded sample data meet the size requirements for splicing.
[0101] For example, if the size of the first sample data is w (width) * h (height) * c1 (number of data channels), and the size of the second sample data is: w (width) * h (height) * c2 (number of data channels), after superimposing the first sample data and the second sample data on the data channels, the size of the obtained target sample data is: w * h * (c1 + c2).
[0102] Exemplarily, the first dense matrix corresponding to the first sample data and the second dense matrix corresponding to the second sample data can be processed in parallel to obtain the target sample data.
[0103] In the embodiments of the present disclosure, the generated target sample data includes the first sample data after the first coding process of the first sample code written in a high-level language. The method for generating sample data provided by the embodiments of the present disclosure also converts the first sample code into a second sample code written in a preset intermediate language. Since the intermediate language is more concise and the function and role of the code can be inferred from the second sample code in this form, the target sample data also includes the second sample data after the second coding process of the second sample code written in the preset intermediate language, making the feature information covered by the sample data deeper and more comprehensive. Furthermore, the target sample data generated by using the method for generating sample data provided by the embodiments of the present disclosure can meet the usage requirements in complex task scenarios.
[0104] See Figure 4 As shown, it is a flowchart of a method for generating a neural network provided by the embodiments of the present disclosure. The method includes steps S401 to S402, where:
[0105] S401. Obtain target sample data.
[0106] Among them, the target sample data is generated by using the specific implementation manners shown in S101 to S104 in the method for generating sample data provided by the embodiments of the present disclosure.
[0107] S402. Use the target sample data to train the neural network to be trained to obtain a target neural network.
[0108] In a specific implementation, the target sample data is used as supervised data to perform supervised training on the neural network to be trained, and a target neural network capable of generating corresponding codes based on multiple task scenarios is generated.
[0109] In the embodiments of the present disclosure, sample data covering relatively rich and comprehensive feature information is used as supervised data to train the neural network to be trained to obtain a target neural network; the target neural network generated in this way has strong generalization ability and is suitable for generating codes corresponding to multiple task scenarios.
[0110] Those skilled in the art can understand that in the above methods of the specific embodiments, the writing order of each step does not mean a strict execution order and does not impose any limitation on the implementation process. The specific execution order of each step should be determined according to its function and possible internal logic.
[0111] Based on the same inventive concept, the embodiments of the present disclosure also provide a sample data generation device corresponding to the sample data generation method. Since the principle of solving problems by the device in the embodiments of the present disclosure is similar to the above sample data generation method of the embodiments of the present disclosure, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be elaborated.
[0112] Refer to Figure 5 As shown in the figure, it is a schematic diagram of a sample data generation device provided by an embodiment of the present disclosure. The device includes: a first sample module 501, a conversion module 502, an encoding module 503, and a generation module 504; wherein,
[0113] A first acquisition module 501 is configured to acquire a first sample code written in a high-level language; a conversion module 502 is configured to convert the first sample code into a second sample code written in a preset intermediate language; an encoding module 503 is configured to perform a first encoding process on the first sample code to obtain first sample data; and perform a second encoding process on the second sample code to obtain second sample data; a generation module 504 is configured to generate target sample data corresponding to the first sample code based on the first sample data and the second sample data.
[0114] In an optional implementation manner, when the conversion module 502 executes the conversion of the first sample code into a second sample code written in a preset intermediate language, it is specifically configured to: convert the first sample code into the second sample code based on the conversion relationship between the high-level language and the preset intermediate language.
[0115] In an optional implementation manner, when the encoding module 503 executes the first encoding process on the first sample code to obtain first sample data, it is specifically configured to: determine a target string from the first sample code, and convert the target string into a standard string corresponding to the target string to obtain a converted first sample code; the target string includes: variable names and / or function names; perform word segmentation processing on the converted first sample code to generate a first string sequence corresponding to the converted first sample code; and perform a first encoding process on the first string sequence to obtain the first sample data.
[0116] In an alternative embodiment, the first string sequence includes at least one standard string; when the encoding module 503 performs the first encoding process on the first string sequence to obtain the first sample data, it is specifically configured to: query, from a pre-generated first encoding table, the first encoding values respectively corresponding to the standard strings in the first string sequence; and generate the first sample data based on the first encoding values respectively corresponding to the standard strings in the first string sequence.
[0117] In an alternative embodiment, when the encoding module 503 generates the first sample data based on the first encoding values respectively corresponding to the standard strings in the first string sequence, it is specifically configured to: construct a first sparse matrix based on the first encoding values respectively corresponding to the standard strings in the first string sequence; convert the first sparse matrix into a first dense matrix with a preset dimension; and determine the first dense matrix as the first sample data.
[0118] In an alternative embodiment, the second sample code includes: a second string sequence composed of the intermediate language; when the encoding module 503 performs a second encoding process on the second sample code to obtain second sample data, it is specifically configured to: query, from a pre-generated second encoding table, the second encoding values respectively corresponding to the strings in the second string sequence; and generate the second sample data based on the second encoding values respectively corresponding to the strings in the second string sequence.
[0119] In an alternative embodiment, when the encoding module 503 generates the second sample data based on the second encoding values respectively corresponding to the strings in the second string sequence, it is specifically configured to: construct a second sparse matrix based on the second encoding values respectively corresponding to the strings in the second string sequence; convert the second sparse matrix into a second dense matrix with a preset dimension; and determine the second dense matrix as the second sample data.
[0120] In an alternative embodiment, when the generation module 504 generates the target sample data corresponding to the first sample code based on the first sample data and the second sample data, it is specifically configured to: perform a fusion process on the first sample data and the second sample data to obtain the target sample data.
[0121] Based on the same inventive concept, an embodiment of the present disclosure further provides a neural network generation device corresponding to the neural network generation method. Since the principle of solving problems by the device in the embodiment of the present disclosure is similar to that of the neural network generation method in the above embodiment of the present disclosure, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.
[0122] Referring to Figure 6 as shown, which is a schematic diagram of a generation device of a neural network provided by an embodiment of the present disclosure. The device includes: a second acquisition module 601 and a training module 602; wherein,
[0123] The second acquisition module 601 is configured to acquire target sample data; the target sample data is generated by using the generation method described in any one of the first aspect; the training module 602 is configured to use the target sample data to train a neural network to be trained, and obtain a target neural network.
[0124] For the description of the processing flow of each module in the device and the interaction flow between the modules, reference can be made to the relevant descriptions in the above method embodiments, which will not be elaborated here.
[0125] Based on the same inventive concept, an embodiment of the present application further provides a computer device. Referring to Figure 7 as shown, which is a schematic diagram of the structure of a computer device 700 provided by an embodiment of the present application, including a processor 701, a memory 702, and a bus 703. Among them, the memory 702 is used to store execution instructions, including an internal memory 7021 and an external memory 7022; the internal memory 7021 here is also called the main memory, which is used to temporarily store the operation data in the processor 701 and the data exchanged with the external memory 7022 such as a hard disk. The processor 701 exchanges data with the external memory 7022 through the internal memory 7021. When the computer device 700 runs, the processor 701 communicates with the memory 702 through the bus 703, so that the processor 701 executes the following instructions:
[0126] Acquire a first sample code written in a high-level language; convert the first sample code into a second sample code written in a preset intermediate language; perform a first encoding process on the first sample code to obtain first sample data; and perform a second encoding process on the second sample code to obtain second sample data; generate target sample data corresponding to the first sample code based on the first sample data and the second sample data.
[0127] Or the processor 701 executes the following instructions:
[0128] Acquire target sample data; the target sample data is generated by using a generation method of sample data; use the target sample data to train a neural network to be trained, and obtain a target neural network.
[0129] Among them, the specific processing flow of the processor 701 can be referred to the description in the above method embodiments, which will not be repeated here.
[0130] Embodiments of the present disclosure also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of the sample data generation method and the neural network generation method described in the above method embodiments. Among them, the storage medium may be a volatile or non-volatile computer-readable storage medium.
[0131] Embodiments of the present disclosure also provide a computer program product, which carries program code. The instructions included in the program code can be used to execute the steps of the sample data generation method and the neural network generation method described in the above method embodiments. For details, reference can be made to the above method embodiments, which will not be elaborated here.
[0132] Among them, the above computer program product can be specifically implemented in a manner of hardware, software, or a combination thereof. In an alternative embodiment, the computer program product is specifically embodied as a computer storage medium. In another alternative embodiment, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), etc.
[0133] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described systems and devices can refer to the corresponding processes in the foregoing method embodiments, which will not be elaborated here. In several embodiments provided by the present disclosure, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some communication interfaces. The indirect coupling or communication connection of the devices or units can be in an electrical, mechanical, or other form.
[0134] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0135] In addition, in each embodiment of the present disclosure, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0136] When the above-mentioned functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on such understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present disclosure. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.
[0137] Finally, it should be noted that the above-described embodiments are only specific implementation manners of the present disclosure, used to illustrate the technical solutions of the present disclosure, rather than limiting them. The protection scope of the present disclosure is not limited thereto. Although the present disclosure has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present disclosure can still modify the technical solutions described in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes, or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should all be covered by the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.
Claims
1. A method for generating sample data, characterized in that, Including: Obtain a first sample code written in a high-level language; Convert the first sample code into a second sample code written in a preset intermediate language; The second sample code includes: a second string sequence composed of the intermediate language; Perform a first encoding process on the first sample code to obtain first sample data; wherein, performing the first encoding process on the first sample code to obtain first sample data includes: determining a target string from the first sample code, and converting the target string into a standard string corresponding to the target string to obtain a converted first sample code; performing word segmentation processing on the converted first sample code to generate a first string sequence corresponding to the converted first sample code; performing a first encoding process on the first string sequence to obtain the first sample data; the target string includes: variable names and / or function names; the first string sequence includes at least one standard string; Perform a second encoding process on the second sample code to obtain second sample data; wherein, performing the second encoding process on the second sample code to obtain second sample data includes: querying, from a pre-generated second encoding table, second encoding values respectively corresponding to each string in the second string sequence; generating the second sample data based on the second encoding values respectively corresponding to each string in the second string sequence; Fuse the first sample data and the second sample data to generate target sample data corresponding to the first sample code; the fusion process includes but is not limited to at least one of the following: splicing process, parallel connection process, weighted summation process; Wherein, performing the first encoding process on the first string sequence to obtain the first sample data includes: querying, from a pre-generated first encoding table, first encoding values respectively corresponding to each standard string in the first string sequence; generating the first sample data based on the first encoding values respectively corresponding to each standard string in the first string sequence.
2. The method according to claim 1, wherein The converting the first sample code into a second sample code written in a preset intermediate language includes: Based on the conversion relationship between the high-level language and the preset intermediate language, convert the first sample code into the second sample code.
3. The method according to claim 1, wherein The generating the first sample data based on the first encoding values respectively corresponding to each standard string in the first string sequence includes: Based on the first encoding values respectively corresponding to each standard string in the first string sequence, form a first sparse matrix; Convert the first sparse matrix into a first dense matrix with a preset dimension; Determine the first dense matrix as the first sample data.
4. The method according to claim 1, characterized in that, The generating the second sample data based on the second encoding values respectively corresponding to each string in the second string sequence includes: Based on the second encoding values respectively corresponding to each string in the second string sequence, form a second sparse matrix; Convert the second sparse matrix into a second dense matrix with a preset dimension; Determine the second dense matrix as the second sample data.
5. The method according to any one of claims 1 to 3, characterized in that Generating the target sample data corresponding to the first sample code based on the first sample data and the second sample data includes: Performing a fusion process on the first sample data and the second sample data to obtain the target sample data.
6. A method for generating a neural network, characterized in that, Including: Obtain target sample data; The target sample data is generated by using the sample data generation method according to any one of claims 1-5; Using the target sample data to train a neural network to be trained to obtain a target neural network.
7. A generating device for sample data, characterized in that, Including: A first acquisition module for acquiring a first sample code written in a high-level language; A conversion module for converting the first sample code into a second sample code written in a preset intermediate language; The second sample code includes: a second string sequence composed of the intermediate language; An encoding module for performing a first encoding process on the first sample code to obtain first sample data; wherein, performing the first encoding process on the first sample code to obtain first sample data includes: determining a target string from the first sample code, and converting the target string into a standard string corresponding to the target string to obtain a converted first sample code; performing a word segmentation process on the converted first sample code to generate a first string sequence corresponding to the converted first sample code; performing a first encoding process on the first string sequence to obtain the first sample data; the target string includes: a variable name and / or a function name; the first string sequence includes at least one standard string; The encoding module is further configured to perform a second encoding process on the second sample code to obtain second sample data; wherein, performing the second encoding process on the second sample code to obtain second sample data includes: querying second encoding values respectively corresponding to each string in the second string sequence from a pre-generated second encoding table; generating the second sample data based on the second encoding values respectively corresponding to each string in the second string sequence; A generation module for performing a fusion process on the first sample data and the second sample data to generate target sample data corresponding to the first sample code; the fusion process includes but is not limited to at least one of the following: splicing process, parallel connection process, weighted summation process; When performing the first encoding process on the first string sequence to obtain the first sample data, the encoding module is specifically configured to: query first encoding values respectively corresponding to each standard string in the first string sequence from a pre-generated first encoding table; generate the first sample data based on the first encoding values respectively corresponding to each standard string in the first string sequence.
8. A generation device of a neural network, characterized in that, Including: A second acquisition module for acquiring target sample data; The target sample data is generated by using the sample data generation method according to any one of claims 1-5; A training module for using the target sample data to train a neural network to be trained to obtain a target neural network.
9. A computer device, characterized in that, Comprising: A processor and a memory, the memory storing machine-readable instructions executable by the processor, the processor being configured to execute the machine-readable instructions stored in the memory, and when the machine-readable instructions are executed by the processor, the processor executes the steps of the method for generating sample data according to any one of claims 1 to 5, or executes the steps of the method for generating a neural network according to claim 6.
10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is run by a computer device, the computer device executes the steps of the method for generating sample data according to any one of claims 1 to 5, or executes the steps of the method for generating a neural network according to claim 6.
Citation Information
Patent Citations
Language transformation model training method and device, language transformation method and device, equipment and medium
CN113609157A
Systems, methods, and apparatuses for heterogeneous computing
JP2021064378A