Training sample generation method and device, equipment and storage medium
By splitting the instruction data and semantic coding feature splicing, and combining the clustering algorithm to generate high-quality training samples, the problem of poor quality training samples in the existing technology is solved, and the training effect and application performance of the model are improved.
Patent Information
- Application Number
- CN202510433825.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-11
AI Technical Summary
The existing text data augmentation technology cannot effectively identify and process complex structured information in system operation and maintenance instructions, resulting in poor quality of the generated training samples, affecting the model training effect and practical application performance.
By splitting the instruction data, obtaining the instruction template and self-assignment data, determining the semantic encoding and performing feature splicing, the augmented instruction encoding is clustered using clustering algorithms and preset clustering cluster centers to generate high-quality and semantically consistent training samples.
It realizes high-quality training sample generation, improves the data information and performance of the model, and enhances the training effect of the data detection model.
Smart Images

Figure CN120296424A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of model training, and in particular, to a method, device, equipment and storage medium for generating training samples. Background Art
[0002] In the daily operation of current database systems, how to effectively identify special data (such as abnormal data) in system instructions has become an urgent problem to be solved. The instruction records of normal operation account for the vast majority, but only a very small amount of special information is contained in these massive instructions. It is often difficult to obtain enough samples for effective model training due to the scarcity of special data, which poses a huge challenge to traditional data detection methods.
[0003] Currently, text data augmentation techniques aim to improve the generalization ability and robustness of models by augmenting limited text data. Existing text data augmentation techniques, such as rule-based methods (synonym replacement, random insertion, deletion, and swapping), language model methods (back translation and Transformer-based generation, etc.), and Seq2Seq models, mainly focus on changes in the surface information of text. For example, in system instructions, key parameters such as operation types, timestamps, and user information have specific semantic meanings, but existing augmentation methods cannot accurately identify and process these semantic information.
[0004] However, some instructions (such as system operation and maintenance instructions) usually contain complex structured information, such as command sequences, parameter configurations, and status feedback, etc. When processing instructions of types such as system operation and maintenance, existing data augmentation techniques often fail to fully consider the characteristics and requirements of these instructions. Only by simply replacing synonyms or performing random operations, etc., to augment data may disrupt this inherent semantic connection, and the generated new samples may be similar to the original samples on the surface. It is difficult for existing technologies to generate high-quality augmented data that meets the actual application scenarios, thus affecting the model training effect and actual application performance. Summary of the Invention
[0005] The present invention provides a method, device, equipment and storage medium for generating training samples to solve the problem of poor quality of the generated training samples.
[0006] In a first aspect, the present invention provides a method for generating training samples, including:
[0007] Obtain instruction data, and split each piece of the instruction data to obtain an instruction template and self-assigned data in the instruction data;
[0008] Determine a first semantic encoding of the instruction template and a second semantic encoding of the self-assigned data, and perform feature splicing on the first semantic encoding and the second semantic encoding to obtain an instruction encoding;
[0009] Amplify the instruction encoding to obtain the encoding to be processed, and perform clustering processing on the encoding to be processed based on the preset cluster center to obtain the target instruction sample, where the preset cluster center is the instruction encoding of the preset standard instruction.
[0010] In a second aspect, the present invention provides a training sample generation device, including:
[0011] An instruction splitting module, configured to obtain instruction data and split each piece of the instruction data to obtain an instruction template and the self-assigned data in the instruction data;
[0012] An instruction encoding determination module, configured to determine the first semantic encoding of the instruction template and the second semantic encoding of the self-assigned data, and perform feature splicing on the first semantic encoding and the second semantic encoding to obtain an instruction encoding;
[0013] A sample generation module, configured to amplify the instruction encoding to obtain the encoding to be processed, and perform clustering processing on the encoding to be processed based on the preset cluster center to obtain the target instruction sample, where the preset cluster center is the instruction encoding of the preset standard instruction.
[0014] In a third aspect, the present invention provides an electronic device, which includes:
[0015] At least one processor;
[0016] And a memory communicatively connected to at least one processor;
[0017] Wherein, the memory stores a computer program executable by at least one processor, and the computer program is executed by at least one processor so that at least one processor can execute the training sample generation method in the first aspect above.
[0018] In a fourth aspect, the present invention provides a computer-readable storage medium, which stores computer instructions for causing a processor to implement the training sample generation method in the first aspect above when executed.
[0019] The training sample generation solution provided by the present invention splits out the instruction template and variable self-assigned data according to the instruction data, and performs in-depth semantic parsing on the instruction through semantic encoding and feature splicing to obtain the semantic information encoding of the instruction. Then, the clustering algorithm and the preset cluster center are used to cluster the augmented instruction encoding to obtain high-quality and semantically consistent training samples, realizing the augmentation of few samples. Thus, when training the data detection model, the data information can be effectively enhanced to improve the model performance.
[0020] It should be understood that the content described in this section is not intended to identify the key or important features of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0022] Figure 1 is a flowchart of a training sample generation method provided according to Embodiment 1 of the present invention;
[0023] Figure 2 is a flowchart of a training sample generation method provided according to Embodiment 2 of the present invention;
[0024] Figure 3 is a schematic structural diagram of a training sample generation device provided according to Embodiment 3 of the present invention;
[0025] Figure 4 is a schematic structural diagram of an electronic device provided according to Embodiment 4 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0026] In order to enable those skilled in the art to better understand the solution of the present invention, the following clearly and completely describes the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0027] It should be noted that the terms "first", "second", etc. in the description, claims and above-mentioned drawings of the present invention are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In the description of the present invention, unless otherwise specified, "a plurality of" means two or more. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally means that the associated objects before and after are in an "or" relationship. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0028] Embodiment 1
[0029] Figure 1 The figure is a flowchart of a training sample generation method provided in Embodiment 1 of the present invention. This embodiment is applicable to the situation of generating instruction training samples for model training. This method can be executed by a training sample generation device, which can be implemented in the form of hardware and / or software. The training sample generation device can be configured in an electronic device, which can be composed of two or more physical entities or one physical entity.
[0030] As Figure 1 shown, a training sample generation method provided in Embodiment 1 of the present invention specifically includes the following steps:
[0031] S101. Obtain instruction data, and split each piece of the instruction data to obtain an instruction template and self-assigned data in the instruction data.
[0032] In this embodiment, the acquired instruction data may include key parameters such as operation type, timestamp, and user information, where the type of instruction may be an operation and maintenance instruction, etc. For example, the operation and maintenance instructions already existing in the power grid business data can be exported through the database. The instruction may specifically include a query statement and a description of the platform and risk level. The instruction data may then be subjected to preprocessing operations such as data cleaning to remove highly repetitive information in the instruction data. Each preprocessed instruction data is then parsed and split using a preset algorithm or model to obtain an instruction template and self-assigned data in each instruction data, where the self-assigned data in the instruction data is variable content in the instruction data, such as an IP address and an operation database. For example, the Drain model may be used to split instruction data, and the specific process includes:
[0033] First, a predefined regular expression is used to extract self-assignment information from each instruction data, and then the instruction data is classified according to the instruction length and starting word. Assuming that each instruction with the same event has the same length, the instructions can be further classified according to the starting word to form leaf nodes. If the currently judged word is a number, it is assigned to the wildcard node. The data in the instruction template is compared bit by bit with the similarity of the historical instruction template. If the highest similarity is less than the preset parameter, a new instruction template is created, otherwise, the historical instruction template is determined as the instruction template for the current instruction data. This process decomposes the instruction into an instruction template and a self-assignment part, ensuring that semantically similar instructions are correctly classified, thereby improving the accuracy of instruction parsing and the effectiveness of subsequent data analysis, and providing a solid foundation for downstream tasks such as instruction data detection.
[0034] S102: Determine a first semantic code of the instruction template and a second semantic code of the self-assigned data, and perform feature concatenation on the first semantic code and the second semantic code to obtain an instruction code.
[0035] In this embodiment, the instruction template and the self-assigned data can be semantically encoded using a preset model or network to obtain a first semantic encoding and a second semantic encoding, wherein the encoding can be expressed as an encoding vector. Then, the first semantic encoding and the second semantic encoding can be feature concatenated to obtain the semantic encoding of the instruction, that is, the instruction encoding.
[0036] S103, amplifying the instruction code to obtain a code to be processed, and clustering the code to be processed based on a preset cluster center to obtain a target instruction sample, wherein the preset cluster center is an instruction code of a preset standard instruction.
[0037] In this embodiment, the instruction encoding can be amplified by a preset method to obtain the encoding to be processed. Then, taking the instruction encoding of the preset standard instruction as the center of the clustering cluster, the encoding to be processed is clustered by a clustering algorithm to obtain multiple clustering clusters. Outliers and smaller clusters can be removed, etc., so as to delete the augmented instruction encodings with low quality and achieve refined data augmentation. The result after removal is the encoding of the target instruction sample. Among them, for few-shot learning, a standard sample instruction is preset in advance, and this instruction is the preset clustering cluster center.
[0038] The technical solution of the embodiment of the present invention splits the instruction template and variable self-assigned data from the instruction data, and through semantic encoding and feature splicing, deeply semantically analyzes the instruction to obtain the semantic information encoding of the instruction. Then, using the clustering algorithm and the preset clustering cluster center to cluster the augmented instruction encoding, high-quality and semantically consistent training samples are obtained. Therefore, when training the data detection model, the data information can be effectively enhanced to improve the model performance.
[0039] Optionally, the feature splicing of the first semantic encoding and the second semantic encoding to obtain an instruction encoding includes: using a preset multi-head attention module to perform feature enhancement and feature splicing on the first semantic encoding and the second semantic encoding to obtain an instruction encoding.
[0040] Specifically, the multi-head attention module can calculate the similarity between the query (Query), key (Key), and value (Value) to obtain the attention score, so as to weight different parts of the input sequence. Each attention head can focus on different aspects of the input sequence and capture different features. Finally, the outputs of multiple attention heads are spliced together and transformed through a linear layer to obtain the final multi-head attention output, that is, the instruction encoding.
[0041] Further, the preset multi-head attention module includes a scaled dot-product attention module.
[0042] Specifically, the scaled dot-product attention is an implementation of the self-attention mechanism. The scaled dot-product attention calculates the similarity between the Query and the Key to generate a normalized attention weight, so as to achieve weighted summation of different parts of the input sequence. This method not only improves the calculation efficiency but also enhances the stability and performance.
[0043] Optionally, the instruction data includes instruction data related to power grid operation.
[0044] Specifically, the instruction data can be instruction data related to power grid operation, such as power grid operation and maintenance instruction data and power grid control instruction data, etc.
[0045] Example 2
[0046] Figure 2 As shown in the flowchart of a training sample generation method provided in Example 2 of the present invention, the technical solution of the embodiment of the present invention is further optimized on the basis of the above optional technical solutions, and a specific method for generating an instruction training sample for model training is given.
[0047] Optionally, the using the preset multi-head attention module to perform feature enhancement and feature splicing on the first semantic encoding and the second semantic encoding to obtain an instruction encoding includes: using a scaled dot-product attention module to perform attention calculation on the second semantic encoding to obtain an encoded feature-enhanced encoding, and performing feature splicing on the feature-enhanced encoding and the first semantic encoding to obtain a splicing result; performing feature encoding on the splicing result to obtain an instruction encoding. The advantage of this setting is that the scaled dot-product attention module performs feature enhancement and feature splicing on the semantic encoding of self-assigned data with business meanings, realizing the incorporation of variable self-assigned information into the learning and augmentation process, which is equivalent to incorporating professional knowledge rather than just instruction templates, so as to be able to more accurately grasp the generation of instructions, obtain high-quality instruction semantic feature encodings, and improve the generation quality of subsequent target instruction samples.
[0048] Optionally, the amplifying the instruction encoding to obtain a to-be-processed encoding includes: determining a target encoding whose similarity to the self-assigned encoding in the instruction encoding is greater than a preset threshold; using the target encoding to replace the self-assigned encoding in the instruction encoding to obtain a to-be-processed encoding. The advantage of this setting is that it realizes coarse-grained augmentation of data that conforms to the instruction semantic similarity.
[0049] Optionally, the clustering the to-be-processed encoding based on a preset clustering cluster center to obtain a target instruction sample includes: performing DBSCAN clustering processing on the to-be-processed encoding based on a plurality of preset clustering cluster centers to obtain a plurality of clustering clusters, where the preset clustering cluster center is the instruction encoding of a preset abnormal instruction; screening the clustering clusters to obtain a target instruction encoding, and generating a target instruction sample according to the target instruction encoding. The advantage of this setting is that in the scenario of dealing with instructions with strong context semantic associations, after feature splicing and data coarse augmentation of semantic encodings, DBSCAN clustering can generate a fine augmentation result, and the generated data quality is more in line with the application scenario and of higher quality than existing methods.
[0050] As Figure 2 shown, a training sample generation method provided in Example 2 of the present invention specifically includes the following steps:
[0051] S201. Obtain instruction data, and split each piece of the instruction data to obtain an instruction template and the self-assigned data in the instruction data.
[0052] S202. Determine the first semantic encoding of the instruction template and the second semantic encoding of the self-assigned data.
[0053] S203. Use the scaled dot-product attention module to perform attention calculation on the second semantic encoding to obtain an encoding with enhanced features, and perform feature splicing on the encoding with enhanced features and the first semantic encoding to obtain a splicing result.
[0054] Specifically, the calculation process of the scaled dot-product attention module can be expressed as:
[0055]
[0056] where h is the number of attention heads, d represents the dimension of the input, q, k, and v represent the query, key, and value, and φ p is the first semantic encoding processed by the feed-forward layer. The scaled dot-product attention module can perform attention calculation on the second semantic encoding of the self-assigned data, calculate q, k, and v to obtain an encoding with enhanced features. Then, perform feature splicing on the first semantic encoding processed by the feed-forward layer and the encoding with enhanced features to obtain a splicing result. The scaled dot-product attention module enhances the representation ability of the instruction sequence modeling and reduces the loss of instruction information caused by splitting the instruction template. Among them, the processing of the feed-forward layer and the feature enhancement can be parallel, which can enhance the generalization ability of this solution for different instruction sources.
[0057] S204. Perform feature encoding on the splicing result to obtain an instruction encoding.
[0058] Specifically, the splicing result can be subjected to feature encoding to obtain an instruction encoding.
[0059] S205. Determine a target encoding whose similarity to the self-assigned encoding in the instruction encoding is greater than a preset threshold.
[0060] Specifically, it can be determined from a preset statement encoding library that the encoding whose similarity to the self-assigned (encoding) part in the instruction encoding is greater than the preset threshold, and this encoding is the target encoding. The target encoding is generally the encoding of a near-synonym or synonym of the self-assigned value.
[0061] S206. Use the target encoding to replace the self-assigned encoding in the instruction encoding to obtain a to-be-processed encoding.
[0062] Specifically, by replacing the self-assigned encoding in the instruction encoding with the target encoding, data coarse-grained augmentation that conforms to the instruction semantic similarity is achieved.
[0063] S207. Perform DBSCAN clustering on the to-be-processed encoding based on multiple preset cluster centers to obtain multiple clusters, where the preset cluster centers are instruction encodings of preset abnormal instructions.
[0064] Specifically, in order to improve the quality of the generated abnormal instruction encodings, during clustering, the instruction encodings of the preset abnormal instructions can be set as the benchmark points of the clusters, and DBSCAN clustering is performed on the to-be-processed encoding to obtain multiple clusters.
[0065] S208. Screen the clusters to obtain target instruction encodings, and generate target instruction samples according to the target instruction encodings.
[0066] Specifically, screening conditions can be preset, such as removing smaller clusters, etc. The clusters are screened according to the preset screening conditions to obtain target instruction encodings. The target instruction encodings can be used as target instruction samples, or the target instruction encodings can be restored, and the restored text is the target instruction sample.
[0067] The training sample generation method provided by the embodiments of the present invention uses a scaled dot product attention module to perform feature enhancement and feature splicing on the semantic encoding of self-assigned data with business meaning, realizing the incorporation of variable self-assigned information into the learning and augmentation process, which is equivalent to incorporating professional knowledge rather than just instruction templates, so as to be able to more accurately grasp the generation of instructions, obtain high-quality instruction semantic feature encodings, improve the generation quality of subsequent target instruction samples, and realize coarse-grained augmentation of data that conforms to the instruction semantic similarity. In the scenario of dealing with instructions with strong context semantic associations, after feature splicing and data coarse augmentation of the semantic encoding, DBSCAN clustering can be used to generate fine-grained augmentation results, and the generated data quality is more suitable for the application scenario and of higher quality than existing methods. In the scenario of dealing with less labeled abnormal data or data without labels, the training samples generated by this method can help the training of the anomaly detection model, improve the richness of the model training data, and thus improve the generalization of the model.
[0068] Embodiment III
[0069] Figure 3 It is a structural schematic diagram of a training sample generation device provided by Embodiment III of the present invention. As Figure 3 shown, the device includes: an instruction splitting module 301, an instruction encoding determination module 302, and a sample generation module 303, where:
[0070] The instruction splitting module is used to obtain instruction data and split each piece of the instruction data to obtain an instruction template and self-assigned data in the instruction data;
[0071] An instruction encoding determination module, configured to determine a first semantic encoding of the instruction template and a second semantic encoding of the self-assigned data, and perform feature splicing on the first semantic encoding and the second semantic encoding to obtain an instruction encoding;
[0072] A sample generation module, configured to amplify the instruction encoding to obtain a to-be-processed encoding, and perform clustering processing on the to-be-processed encoding based on a preset clustering cluster center to obtain a target instruction sample, where the preset clustering cluster center is an instruction encoding of a preset standard instruction.
[0073] The training sample generation device provided by the embodiment of the present invention splits the instruction data into an instruction template and variable self-assigned data, and through semantic encoding and feature splicing, deeply semantically analyzes the instruction to obtain a semantic information encoding of the instruction, and then uses a clustering algorithm and a preset clustering cluster center to cluster the augmented instruction encoding to obtain high-quality and semantically consistent training samples, so that when training a data detection model, the data information can be effectively enhanced to improve the model performance.
[0074] Optionally, the instruction encoding determination module includes:
[0075] An instruction encoding determination unit, configured to perform feature enhancement and feature splicing on the first semantic encoding and the second semantic encoding by using a preset multi-head attention module to obtain an instruction encoding.
[0076] Further, the preset multi-head attention module includes a scaled dot product attention module.
[0077] Further, the performing feature enhancement and feature splicing on the first semantic encoding and the second semantic encoding by using a preset multi-head attention module to obtain an instruction encoding includes: performing attention calculation on the second semantic encoding by using a scaled dot product attention module to obtain an encoding with enhanced features, and performing feature splicing on the encoding with enhanced features and the first semantic encoding to obtain a splicing result; performing feature encoding on the splicing result to obtain an instruction encoding.
[0078] Optionally, the sample generation module includes:
[0079] A clustering cluster determination unit, configured to perform DBSCAN clustering processing on the to-be-processed encoding based on multiple preset clustering cluster centers to obtain multiple clustering clusters, where the preset clustering cluster center is an instruction encoding of a preset abnormal instruction;
[0080] A sample generation determination unit, configured to screen the clustering clusters to obtain a target instruction encoding, and generate a target instruction sample according to the target instruction encoding.
[0081] Optionally, the instruction data includes instruction data related to power grid operation.
[0082] Optionally, the sample generation module includes:
[0083] A target encoding determination unit, configured to determine a target encoding whose similarity to the self-assigned encoding in the instruction encoding is greater than a preset threshold;
[0084] A to-be-processed encoding determination unit, configured to replace the self-assigned encoding in the instruction encoding with the target encoding to obtain a to-be-processed encoding.
[0085] The training sample generation device provided by the embodiments of the present invention can execute the training sample generation method provided by any embodiment of the present invention, and has functional modules and beneficial effects corresponding to the execution of the method.
[0086] Embodiment 4
[0087] Figure 4 FIG. shows a schematic structural diagram of an electronic device 40 that can be used to implement the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device (such as a helmet, glasses, a watch, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0088] As Figure 4 shown, the electronic device 40 includes at least one processor 41, and a memory communicatively connected to at least one processor 41, such as a read-only memory (ROM) 42, a random access memory (RAM) 43, etc. Among them, the memory stores a computer program executable by at least one processor. The processor 41 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 42 or the computer program loaded from the storage unit 48 into the random access memory (RAM) 43. In the RAM 43, various programs and data required for the operation of the electronic device 40 can also be stored. The processor 41, the ROM 42, and the RAM 43 are connected to each other through a bus 44. The input / output (I / O) interface 45 is also connected to the bus 44.
[0089] Multiple components in the electronic device 40 are connected to the I / O interface 45, including: an input unit 46, such as a keyboard, a mouse, etc.; an output unit 47, such as various types of displays, speakers, etc.; a storage unit 48, such as a disk, an optical disc, etc.; and a communication unit 49, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 49 allows the electronic device 40 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0090] The processor 41 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 41 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 41 executes the various methods and processes described above, such as the training sample generation method.
[0091] In some embodiments, the training sample generation method can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 48. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 40 via the ROM 42 and / or the communication unit 49. When the computer program is loaded into the RAM 43 and executed by the processor 41, one or more steps of the training sample generation method described above can be executed. Alternatively, in other embodiments, the processor 41 can be configured to execute the training sample generation method by any other suitable means (e.g., by means of firmware).
[0092] The various embodiments of the systems and technologies described above in this article can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs, which can be executed and / or interpreted on a programmable system including at least one programmable processor, and the programmable processor can be a special or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0093] A computer program for implementing the method of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when executed by the processor, the functions / operations specified in the flowchart and / or block diagram are implemented. The computer program can be executed entirely on the machine, partially on the machine, partially on the machine as an independent software package and partially on a remote machine, or entirely on a remote machine or server.
[0094] The computer device provided above can be used to execute the training sample generation method provided in any of the above embodiments, and has corresponding functions and beneficial effects.
[0095] Embodiment Five
[0096] In the context of the present invention, a computer-readable storage medium can be a tangible medium, and the computer-executable instructions are used to execute a training sample generation method when executed by a computer processor. The method includes:
[0097] Obtain instruction data, and split each piece of the instruction data to obtain an instruction template and self-assigned data in the instruction data;
[0098] Determine a first semantic encoding of the instruction template and a second semantic encoding of the self-assigned data, and perform feature splicing on the first semantic encoding and the second semantic encoding to obtain an instruction encoding;
[0099] Amplify the instruction encoding to obtain a to-be-processed encoding, and perform clustering processing on the to-be-processed encoding based on a preset clustering cluster center to obtain a target instruction sample, where the preset clustering cluster center is the instruction encoding of a preset standard instruction.
[0100] In the context of the present invention, a computer-readable storage medium can be a tangible medium, which can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the above. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium will include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above.
[0101] The computer device provided above can be used to execute the training sample generation method provided in any of the above embodiments, and has corresponding functions and beneficial effects.
[0102] It should be noted that in the embodiments of the above training sample generation device, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of the functional units are only for the convenience of mutual distinction and do not limit the protection scope of the present invention.
[0103] Note that the above is only a preferred embodiment of the present invention and the applied technical principle. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, re-adjustments and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments. Without departing from the concept of the present invention, more other equivalent embodiments can be included, and the scope of the present invention is determined by the scope of the appended claims.
Claims
1. A method for generating training samples, characterized in that, Including: Obtain instruction data, and split each piece of the instruction data to obtain an instruction template and self-assigned data in the instruction data; Determine a first semantic encoding of the instruction template and a second semantic encoding of the self-assigned data, and perform feature splicing on the first semantic encoding and the second semantic encoding to obtain an instruction encoding; Amplify the instruction encoding to obtain a to-be-processed encoding, and perform clustering processing on the to-be-processed encoding based on a preset clustering cluster center to obtain a target instruction sample, where the preset clustering cluster center is the instruction encoding of a preset standard instruction.
2. The method according to claim 1, characterized in that, The performing feature splicing on the first semantic encoding and the second semantic encoding to obtain an instruction encoding includes: Using a preset multi-head attention module to perform feature enhancement and feature splicing on the first semantic encoding and the second semantic encoding to obtain an instruction encoding.
3. The method according to claim 2, characterized in that, The preset multi-head attention module includes a scaled dot product attention module.
4. The method according to claim 2 or 3, characterized in that, The using a preset multi-head attention module to perform feature enhancement and feature splicing on the first semantic encoding and the second semantic encoding to obtain an instruction encoding includes: Using a scaled dot product attention module to perform attention calculation on the second semantic encoding to obtain an encoding with enhanced features, and performing feature splicing on the encoding with enhanced features and the first semantic encoding to obtain a splicing result; Perform feature encoding on the splicing result to obtain an instruction encoding.
5. The method according to claim 1, wherein The performing clustering processing on the to-be-processed encoding based on a preset clustering cluster center to obtain a target instruction sample includes: Based on a plurality of preset clustering cluster centers, perform DBSCAN clustering processing on the to-be-processed encoding to obtain a plurality of clustering clusters, where the preset clustering cluster center is the instruction encoding of a preset abnormal instruction; Screen the clustering clusters to obtain a target instruction encoding, and generate a target instruction sample according to the target instruction encoding.
6. The method according to claim 1, wherein The instruction data includes instruction data related to power grid operation.
7. The method according to claim 1, wherein The amplifying the instruction encoding to obtain a to-be-processed encoding includes: Determine a target encoding whose similarity to the self-assigned encoding in the instruction encoding is greater than a preset threshold; Use the target encoding to replace the self-assigned encoding in the instruction encoding to obtain a to-be-processed encoding.
8. A training sample generation device, characterized in that Including: An instruction splitting module, configured to obtain instruction data, and split each piece of the instruction data to obtain an instruction template and self-assigned data in the instruction data; An instruction encoding determining module, configured to determine a first semantic encoding of the instruction template and a second semantic encoding of the self-assigned data, and perform feature splicing on the first semantic encoding and the second semantic encoding to obtain an instruction encoding; A sample generating module, configured to amplify the instruction encoding to obtain a to-be-processed encoding, and perform clustering processing on the to-be-processed encoding based on a preset clustering cluster center to obtain a target instruction sample, where the preset clustering cluster center is the instruction encoding of a preset standard instruction.
9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; where, The memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, the at least one processor is enabled to execute the training sample generation method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for implementing the training sample generation method according to any one of claims 1-7 when executed by a processor.