A sample generation method, device, apparatus and storage medium
By generating sample templates that do not involve user privacy and augmenting real-world data, the problem of missing samples caused by user privacy is solved, thereby improving the training accuracy and robustness of the identity recognition model.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NEUSOFT CORP
- Filing Date
- 2023-05-11
- Publication Date
- 2026-05-15
AI Technical Summary
In the training of identity recognition models, the lack of samples due to user privacy issues affects the accuracy of model training.
By generating sample templates, ensuring that the data items under each header are empty, and filling in simulated identity information from a pre-defined header database, combined with data augmentation processing under real collection conditions, a training sample set is generated.
Ensure that training samples do not involve the leakage of user privacy and conform to the actual collection logic, avoid missing samples, and improve the accuracy and robustness of model training.
Smart Images

Figure CN116561578B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, specifically to a method, apparatus, device, and storage medium for sample generation. Background Technology
[0002] In certain scenarios where security is required, it is often necessary to scan a user's identity image to identify their true identity information, such as by scanning a user's ID card.
[0003] Typically, a large number of identity images taken in different environments can be used as training samples to train an identity recognition model, thereby accurately identifying the real identity of each user. However, because user identity images (such as ID card images) may involve user privacy issues, users usually do not easily provide their identity images to third parties as training samples. This means that the identity recognition model may not have a large number of real identity samples to support accurate training, resulting in a sample shortage problem, which makes it impossible to ensure the accuracy of model training. Summary of the Invention
[0004] This application provides a method, apparatus, device, and storage medium for generating samples, which, while preventing the leakage of user privacy, ensures the effective generation of training samples, avoids the loss of training samples, and improves the accuracy of model training.
[0005] In a first aspect, embodiments of this application provide a method for generating samples, the method comprising:
[0006] A sample template is determined, wherein the data items under each header in the sample template are empty;
[0007] Based on the sample template and the pre-defined header database, generate the corresponding initial sample;
[0008] The initial samples are subjected to data augmentation processing under a real acquisition environment to obtain the corresponding training sample set.
[0009] Secondly, embodiments of this application provide a sample generation apparatus, the apparatus comprising:
[0010] The sample template determination module is used to determine the sample template, wherein the data items under each header in the sample template are empty.
[0011] The initial sample generation module is used to generate corresponding initial samples based on the sample template and the preset header database;
[0012] The sample augmentation module is used to perform data augmentation processing on the initial samples under a real acquisition environment to obtain the corresponding training sample set.
[0013] Thirdly, embodiments of this application provide an electronic device, which includes:
[0014] A processor and a memory, the memory being used to store a computer program, and the processor being used to invoke and run the computer program stored in the memory to perform the sample generation method provided in the first aspect of this application.
[0015] Fourthly, embodiments of this application provide a computer-readable storage medium for storing a computer program that causes a computer to perform the sample generation method provided in the first aspect of this application.
[0016] Fifthly, embodiments of this application provide a computer program product, including a computer program / instructions that, when executed by a processor, implement the sample generation method provided in the first aspect of this application.
[0017] This application provides a method, apparatus, device, and storage medium for sample generation. First, a sample template is determined where each header field is empty. Then, based on this sample template and a pre-defined header database, a corresponding initial sample is generated. This initial sample ensures that each header field in the initial sample originates from the header database and does not involve the user's actual privacy information, thus preventing user privacy leaks. Furthermore, by performing data augmentation processing on the initial sample under a realistic data collection environment, a corresponding training sample set is obtained. This ensures that the training samples, while preventing user privacy leaks, conform to the real-world data collection logic of actual samples, thereby ensuring the effective generation of training samples, avoiding missing training samples, and improving the accuracy and robustness of model training. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a flowchart illustrating a sample generation method according to an embodiment of this application;
[0020] Figure 2 This is a schematic diagram illustrating the principle of the user identification process in an embodiment of this application;
[0021] Figure 3 A flowchart illustrating another method for generating samples according to an embodiment of this application;
[0022] Figure 4This is a schematic block diagram of a sample generation device shown in an embodiment of this application;
[0023] Figure 5 This is a schematic block diagram of an electronic device shown in an embodiment of this application. Detailed Implementation
[0024] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0025] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.
[0026] To address the issue of inaccurate model training due to a large number of missing training samples involving user privacy, this application proposes a sample generation scheme. A sample template is defined, in which each data item under each header is empty. By mimicking this template, corresponding initial samples can be generated from a pre-defined header database. This ensures that the data items under each header in the initial samples originate from the header database and do not involve real user privacy information, thus preventing user privacy leaks. Furthermore, by performing data augmentation processing on the initial samples under a realistic data collection environment, a corresponding training sample set is obtained. This ensures that the training samples, while preventing user privacy leaks, conform to the real-world data collection logic of actual samples, thereby ensuring the effective generation of training samples, avoiding missing training samples, and improving the accuracy and robustness of model training.
[0027] Figure 1 This is a flowchart illustrating a sample generation method according to an embodiment of this application. (Refer to...) Figure 1 The method may include the following steps:
[0028] S110, Determine the sample template, where the data items under each header in the sample template are empty.
[0029] Because users' real identity images may involve privacy issues, users typically do not easily provide their real identity images to third parties as training samples, leading to a lack of training samples in model training. Therefore, to obtain a large number of samples sufficient to support accurate model training and avoid this lack of training samples, this application requires that the training samples do not contain users' real privacy information to prevent privacy leaks. Furthermore, to ensure the accuracy of model training, the training samples must mimic users' real identity images, retaining the same background style, so that the model trained using these samples can be accurately applied to various real-world scenarios involving user identity recognition based on their real identity images.
[0030] In order to meet the requirements of the training samples mentioned above, this application will perform corresponding feature analysis based on the user's real identity image to delete the user's privacy information in the real identity image, while retaining the specific header information where the user's privacy information is located, thereby generating a sample template.
[0031] The header in the sample template can be a descriptive title used to categorize the nature of various types of user identity and privacy information, such as the user's name, gender, and address.
[0032] Through the above processing, the data items under each header in the sample template of this application can be empty.
[0033] As an optional implementation scheme in this application, for the sample template, this application can perform fuzzing on the data items under each header in the reference sample to obtain the sample template.
[0034] In other words, this application can capture a frame of a user's real identity image, such as a real ID card image, as a reference sample in this application. Then, by performing corresponding feature analysis on the reference sample, the data items under each header in the reference sample can be identified. These data items may contain the user's private information. Therefore, this application can blur the data items under each header in the reference sample, making each data item under each header empty, thus obtaining the corresponding sample template.
[0035] S120: Generate the corresponding initial sample based on the sample template and the pre-defined header database.
[0036] After determining the sample template, in order to make the training samples resemble the user's real identity image so as to train a network model that can recognize the user's identity, the corresponding identity information needs to be filled in under each header of the sample template to generate a simulated user identity information without involving the user's real privacy information.
[0037] In this application, for each header in the sample template, a corresponding header database is pre-defined. This header database can mimic the specific format of the real identity information under that header and pre-store a large amount of identity information that conforms to the requirements of that header. For example, in the header database under the address header, complete address information can be constructed by combining elements such as the city, county, town, village, district, community, street, and house number of each province.
[0038] Therefore, after determining the sample template, this application can select an identity information item from the corresponding header database for each header in the sample template and fill it into that header in the sample template, thus obtaining an initial sample containing the user's simulated identity information. By continuously performing the above operations, a large number of initial samples can be obtained, enabling the initial samples to support the accurate training of the model without involving the user's real privacy information.
[0039] S130, perform data augmentation processing on the initial samples under a real acquisition environment to obtain the corresponding training sample set.
[0040] When identifying users' identities in real-world scenarios by capturing their real-life images, these images are often affected by various factors such as lighting, shooting angle, tilt, and background. Initial samples, however, are obtained by directly filling in simulated user identity information into a sample template, thus avoiding these real-world shooting interferences. Therefore, directly using these initial samples to train a model would render the trained model unsuitable for use with real-life identity images captured in actual shooting environments.
[0041] Therefore, to ensure the applicability of the trained model in real-world data acquisition environments, this application identifies various shooting interference factors that users encounter when capturing real identity images in such environments. Then, at least one of these interference factors is used to perform corresponding data augmentation on the initial samples, enabling the augmented initial samples to simulate real identity images captured in real-world environments as closely as possible. These augmented initial samples are then combined into a training sample set. Subsequently, when using this training sample set to train the model, the robustness of the model can be greatly improved.
[0042] In some feasible implementations, in order to ensure the effective identification of user identity information, this application may set two types of networks in the target model when training the corresponding target model using the training sample set: a text detection network and a text recognition network.
[0043] The text detection network can be the DBNet network, which uses a segmentation-based differentiable binarization algorithm to quickly detect text box regions in the training samples.
[0044] The text recognition network can be a convolutional recurrent neural network (CRNN), which can analyze the pixel features in each text box region and use the long short-term memory (LSTM) algorithm to extract the corresponding temporal features from the pixel features to process text sequences of arbitrary length, thereby recognizing the identity information in each text box region.
[0045] For the training sample set {Sample _1 Sample _2 Sample _n Each training sample in} _i It can be input into the text detection network of the target model. Assume the training sample is Sample. _i If a text contains k lines of text, then the text detection network can detect k lines of text {Row}. _1 Row _2 Row _k}, then k text lines {Row _1 Row _2 Row _k The output is k text box regions {Img} _1 1mg _2 , ..., Img _k}. Furthermore, the k text box regions {Img _1 1mg _2 , ..., Img _k} This input is then fed into the text recognition network of the target model to obtain the text label sequence {Label} within k text box regions. _1 Label _2 , ..., Label _k}, thus obtaining the training sample Sample _i Identity information in the system.
[0046] By using the above process, the text detection network and text recognition network in the target model can be jointly trained using the training sample set, thus training an accurate target model.
[0047] After training the target model, it can be applied to real-world user identification scenarios. Since users may capture other objects in the background when taking real-world identity images in a real-world environment, or the text in the image may be inverted, the target model in this application can be combined with a contour detection model and an orientation adjustment model to jointly identify identity information in order to ensure the accuracy of user identification.
[0048] For a user-captured image of their verified identity, this image may contain background objects, some of which may have text. Therefore, to prevent the target model from misidentifying the text content on these background objects, such as... Figure 2 As shown, this application first inputs the real identity image into a pre-built contour detection model. The contour detection model can detect the specific contour of the target identity object in the real identity image and extract the target identity image after excluding the environmental background, such as the contour of a real ID card.
[0049] Then, in order to avoid the problem of text recognition difficulties caused by the inversion of text in the real identity image, this application can input the target identity image extracted by the contour detection model into the pre-constructed orientation adjustment model. The orientation adjustment model can identify the text orientation in the target identity image and adjust the target identity image to the orientation of the text.
[0050] Furthermore, the positive target identity image can be input into the text detection network in the target model to detect various text box regions within the target identity image. These text box regions are then further input into the text recognition network in the target model to identify the actual text content within each text box region, thereby accurately identifying the user's identity information.
[0051] The technical solution provided in this application first determines a sample template where each header has an empty data item. Then, based on this sample template and a pre-defined header database, a corresponding initial sample is generated. This initial sample ensures that each header's data item originates from the header database and does not involve the user's actual privacy information, thus preventing user privacy leaks. Furthermore, by performing data augmentation processing on the initial sample under a realistic data collection environment, a corresponding training sample set is obtained. This ensures that the training samples, while preventing user privacy leaks, conform to the real-world data collection logic of actual samples, thereby ensuring the effective generation of training samples, avoiding missing training samples, and improving the accuracy and robustness of model training.
[0052] As an optional implementation scheme in this application, in order to ensure the effectiveness of the training samples, this application can provide a detailed description of the specific generation process of the training samples from two aspects: the header database and the data augmentation method.
[0053] Figure 3 A flowchart illustrating another sample generation method shown in an embodiment of this application is provided, as follows: Figure 3 As shown, the method may include the following steps:
[0054] S310, Determine the sample template, where the data items under each header in the sample template are empty.
[0055] S320, Generate virtual data items under the first type of header based on the header database corresponding to the first type of header in the sample template.
[0056] Considering that the data items under certain headers in the real identity images may have certain relationships and are not randomly filled in, this application can perform correlation analysis on the headers in the sample template to ensure the realism of the training samples, and divide the headers in the sample template into two types: the first type of header and the second type of header.
[0057] The first type of header can be a partial header that allows for random filling of data items to simulate user identity information, such as name, gender, and address when an ID card image is used as a sample template. The second type of header can be a header that requires specific data items to be specified by some data items under the first type of header to simulate user identity information, such as account number when an ID card image is used as a sample template.
[0058] Therefore, for each first-type header in the sample template, this application can set a corresponding header database for the first-type header. The header database can be modeled after the specific format of the real identity information under the header and pre-store a large amount of identity information that meets the requirements of the header.
[0059] When generating training samples using a sample template, the first step is to determine each first-class header and its corresponding header database within the sample template. Then, for each first-class header, an identity information item can be randomly selected from the corresponding header database and used as a virtual data item under that first-class header.
[0060] As an optional implementation scheme in this application, multiple initial samples can be generated by repeatedly selecting virtual data items under each first-class header. To avoid duplication among the initial samples, this application can use each pre-stored identity information as a corresponding candidate data item in the header database corresponding to each first-class header, and set a common attribute for each candidate data item to indicate the degree of commonness of the candidate data item.
[0061] Therefore, for virtual data items under the first type of header, this application can determine them by the following steps: setting the selection probability of the candidate data items based on the common attributes of the candidate data items in the header database corresponding to the first type of header in the sample template; and selecting virtual data items under the first type of header from the header database based on the selection probability of the candidate data items.
[0062] In other words, for each first-class header, when selecting virtual data items, the common attributes of each candidate data item already stored in the header database corresponding to that first-class header can be determined first, thus judging the commonality of each candidate data item. Then, according to the commonality of each candidate data item, a corresponding selection probability is assigned to each candidate data item. The more common the candidate data item, the lower the selection probability. Therefore, based on the selection probability of each candidate data item, virtual data items under that first-class header can be selected from the header database corresponding to that first-class header, thus avoiding the selection of common candidate data items to a certain extent, thereby reducing the duplication among training samples.
[0063] S330, Based on the virtual data items associated with the first type of header in the second type of header within the sample template, generate virtual data items under the second type of header.
[0064] After identifying the virtual data items under each first-class header in the sample template, for each second-class header in the sample template, the associated first-class headers can be determined through header relationships. Then, by performing corresponding text transformations on the virtual data items under each associated first-class header, the virtual data items under that second-class header can be obtained.
[0065] Using an ID card image as a sample template, we can use name, gender, date of birth, and address as the first type of header, and ID number as the second type of header. After determining the virtual data items under each of the first type of headers, since the data items under the second type of header represented by the ID number consist of a six-digit address code, an eight-digit date of birth code, a three-digit sequence code, and a one-digit check code, and the sequence code is related to the user's gender, we can determine that the first type of header associated with the second type of header is the first type of header represented by gender, date of birth, and address. Then, by encoding and transforming the virtual data items under gender, date of birth, and address, we obtain the virtual data items under the second type of header.
[0066] For example, the virtual data item under the second type of header represented by the ID number can be obtained by the following formula: And W i =2 (i-1) (mod11).
[0067] Where A is the check code to be solved, i represents the position number of the number character in the identity account from left to right, and a i W represents the value of the number character at the i-th position. i is the weighting factor at the i-th position.
[0068] S340, for each header in the first type of header and the second type of header, according to the position of the header in the sample template, add the virtual data item under the header to the sample template to obtain the corresponding initial sample.
[0069] After determining the virtual data items under the first and second type of headers, in order to generate accurate training samples, this application can determine the position of each first type and each second type of header within the sample template. Then, to ensure a one-to-one correspondence between headers and data items, this application can set a relative positional relationship between headers and data items.
[0070] Furthermore, for each header in both the first and second types, its position within the sample template can be determined. Then, based on the relative positions of the headers and data items, the position of the virtual data items under those headers within the sample template can be determined. Therefore, the virtual data items under those headers can be added to their corresponding positions within the sample template. By adding corresponding virtual data items under each header within the sample template, the user's identity information can be simulated, resulting in the corresponding initial sample.
[0071] Considering that the real identity information of a certain type of user may contain bilingual identity information, in order to achieve the recognition of bilingual identity information, it is also necessary to generate training samples containing bilingual identity information.
[0072] In this application, a bilingual positional relationship can be set for each virtual data item under a header. After adding the virtual data item under each header to its position in the sample template, the virtual data under each header can be translated using a pre-built translation model to obtain virtual data items in another language. Then, based on the bilingual positional relationship corresponding to each virtual data item under a header and its position in the sample template, the position of the virtual data item in the other language under that header in the sample template is determined, thereby adding the virtual data item in the other language to the sample template accordingly, generating an initial sample containing bilingual identity information.
[0073] S350 determines the corresponding data augmentation method based on the sample interference factors in the actual acquisition environment.
[0074] To ensure the applicability of the trained model in real-world data acquisition environments, this application analyzes various shooting interferences that users may encounter when capturing real identity images in such environments to identify sample interference factors. Then, for each sample interference factor, this application sets a data augmentation method so that, after processing the sample using this method, it can simulate as closely as possible the real impact of that interference factor on image capture.
[0075] For example, if a user's real identity image is interfered with by factors such as lighting, shooting angle, tilt, and background in a real acquisition environment, this application can set six data augmentation methods: adding background, adding noise, affine transformation (image tilt), perspective transformation (affine transformation), adjusting brightness (changing light intensity), and adjusting chroma (changing image color), so that the training samples can approximate real identity images in a real acquisition environment.
[0076] In the case of adding a background as a data augmentation method, this application can randomly select different background images to perform data augmentation processing on the initial sample, and there can be 50 background images.
[0077] For Gaussian noise blurring, a number can be randomly selected from [1, 10] to represent the size of the convolution kernel, and the initial sample can be data augmented to change the degree of noise added to the initial sample.
[0078] For affine transformation as a data augmentation method, a number can be randomly selected from [-45, 45] to represent the rotation angle of the image, and the initial sample can be data augmented to increase the number of training samples taken from multiple angles.
[0079] For perspective transformation as a data augmentation method, a number can be randomly selected from [-30, 30] to represent the tilt of the image, and the initial sample can be data augmented to increase the number of training samples taken from multiple directions.
[0080] For the data augmentation method of adjusting brightness, a number can be randomly selected from [-5,5] to represent the brightness of the image, and the initial sample can be data augmented to increase the number of training samples taken under different lighting intensities.
[0081] For the data augmentation method of adjusting chroma, a number can be randomly selected from [1,5] to add different colors of background light to the training samples.
[0082] S360 performs full combination and full permutation of data augmentation methods to obtain the corresponding target augmentation method set.
[0083] Since different combinations and permutations of data augmentation methods can produce different augmentation effects on training samples, this application can perform full combination and full permutation of multiple data augmentation methods to obtain a target augmentation method set containing at least one of the multiple data augmentation methods. This target augmentation method set can include all combinations and permutations of multiple data augmentation methods.
[0084] As an optional implementation scheme in this application, this application can perform a full combination of data augmentation methods to obtain a corresponding data augmentation combination pattern; and perform a full permutation of the data augmentation methods in each data augmentation combination pattern to obtain a corresponding target augmentation method set.
[0085] In other words, for multiple data augmentation methods, a full combination approach can be used to combine each data augmentation method with a certain number of combinations, resulting in all possible data augmentation combination patterns. Taking data augmentation methods a, b, and c as an example, by performing a full combination of a, b, and c, eight data augmentation combination patterns can be obtained: empty, a, b, c, ab, ac, bc, and abc. Different data augmentation combination patterns can produce different augmentation effects after data augmentation processing of the initial sample.
[0086] Furthermore, if a data augmentation combination pattern contains two or more data augmentation methods, then when the initial sample is augmented sequentially using each data augmentation method in the data augmentation combination pattern, the different order of the data augmentation methods in the data augmentation combination pattern will also produce different augmentation effects on the training sample.
[0087] Therefore, for each data augmentation combination pattern, this application will also perform a full permutation of each data augmentation method in that combination pattern to obtain the corresponding target augmentation method set. This target augmentation method set can ensure that the initial samples can undergo comprehensive data augmentation processing under various augmentation effects, thereby ensuring the diversity of training samples.
[0088] Taking the eight data augmentation combinations (empty, a, b, c, ab, ac, bc, and abc) as an example, for the four data augmentation combinations ab, ac, bc, and abc, we need to perform full permutations. Therefore, ab can result in two augmentation methods: ab and ba; ac can result in two augmentation methods: ac and ca; bc can result in two augmentation methods: bc and cb; and abc can result in six augmentation methods: abc, acb, bac, bca, cab, and cba. Thus, for the three data augmentation methods a, b, and c, we can obtain sixteen augmentation methods: empty, a, b, c, ab, ba, ac, ca, bc, cb, abc, acb, bac, bca, cab, and cba.
[0089] S370: Use each target augmentation method in the target augmentation method set to perform data augmentation processing on the initial sample to obtain the corresponding training sample set.
[0090] For each target augmentation method in the target augmentation method set, the initial samples can be augmented sequentially using the various data augmentation methods contained within that target augmentation method, thereby obtaining the corresponding training sample set and ensuring the diversity of the training samples.
[0091] The technical solution provided in this application first determines a sample template where each header has an empty data item. Then, based on this sample template and a pre-defined header database, a corresponding initial sample is generated. This initial sample ensures that each header's data item originates from the header database and does not involve the user's actual privacy information, thus preventing user privacy leaks. Furthermore, by performing data augmentation processing on the initial sample under a realistic data collection environment, a corresponding training sample set is obtained. This ensures that the training samples, while preventing user privacy leaks, conform to the real-world data collection logic of actual samples, thereby ensuring the effective generation of training samples, avoiding missing training samples, and improving the accuracy and robustness of model training.
[0092] Figure 4 This is a schematic block diagram illustrating a sample generation device according to an embodiment of this application. Figure 4 As shown, the device 400 may include:
[0093] The sample template determination module 410 is used to determine the sample template, wherein the data items under each header in the sample template are empty.
[0094] The initial sample generation module 420 is used to generate corresponding initial samples based on the sample template and the preset header database;
[0095] The sample augmentation module 430 is used to perform data augmentation processing on the initial sample under a real acquisition environment to obtain the corresponding training sample set.
[0096] In some implementations, the initial sample generation module 420 may include:
[0097] The first type of header processing unit is used to generate virtual data items under the first type of header based on the header database corresponding to the first type of header in the sample template.
[0098] The second type of header processing unit is used to generate virtual data items under the second type of header based on the virtual data items associated with the first type of header in the sample template.
[0099] An initial sample generation unit is used to add virtual data items under each of the first type of headers and the second type of headers to the sample template according to the position of the header in the sample template, so as to obtain the corresponding initial sample.
[0100] In some possible implementations, the first type of header processing unit can be specifically used for:
[0101] Based on the common attributes of the candidate data items in the header database corresponding to the first type of header in the sample template, the selection probability of the candidate data items is set;
[0102] Based on the selection probability of the candidate data items, virtual data items under the first type of header are selected from the header database.
[0103] In some implementations, the sample augmentation module 430 may include:
[0104] The augmentation method determination unit is used to determine the corresponding data augmentation method based on the sample interference factors in the real acquisition environment;
[0105] An enhancement method optimization unit is used to perform full combination and full permutation of the data enhancement methods to obtain a corresponding target enhancement method set;
[0106] The data augmentation unit is used to perform data augmentation processing on the initial sample using each target augmentation method in the target augmentation method set to obtain the corresponding training sample set.
[0107] In some implementations, the enhancement optimization unit can be specifically used for:
[0108] The data augmentation methods are fully combined to obtain the corresponding data augmentation combination mode;
[0109] By performing a full permutation of the data augmentation methods in each of the aforementioned data augmentation combination patterns, the corresponding target augmentation method set is obtained.
[0110] In some implementations, the sample template determination module 410 can be specifically used for:
[0111] The data items under each header in the reference sample are blurred to obtain the sample template.
[0112] In this embodiment, a sample template is first determined where each header has an empty data item. Then, based on this sample template and a pre-defined header database, a corresponding initial sample is generated. This initial sample ensures that each header's data item originates from the header database and does not involve the user's actual privacy information, thus preventing user privacy leaks. Furthermore, by performing data augmentation processing on the initial sample under a realistic data collection environment, a corresponding training sample set is obtained. This ensures that the training samples, while preventing user privacy leaks, conform to the real-world data collection logic of actual samples, thereby ensuring the effective generation of training samples, avoiding missing training samples, and improving the accuracy and robustness of model training.
[0113] It should be understood that the device embodiments and method embodiments can correspond to each other, and similar descriptions can be referred to the method embodiments. To avoid repetition, further details will not be provided here. Specifically, Figure 4 The apparatus 400 shown can execute any of the method embodiments in this application, and the foregoing and other operations and / or functions of each module in the apparatus 400 are respectively for implementing the corresponding processes in the various methods in the embodiments of this application. For the sake of brevity, they will not be described in detail here.
[0114] The apparatus 400 of this application embodiment has been described above from the perspective of functional modules in conjunction with the accompanying drawings. It should be understood that this functional module can be implemented in hardware, in software instructions, or in a combination of hardware and software modules. Specifically, the steps of the method embodiments in this application can be completed by integrated logic circuits in the processor's hardware and / or by software instructions. The steps of the method disclosed in this application embodiment can be directly embodied as being executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. Optionally, the software module can be located in a mature storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. This storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps in the above method embodiments.
[0115] Figure 5 This is a schematic block diagram of an electronic device shown in an embodiment of this application.
[0116] like Figure 5 As shown, the electronic device 500 may include:
[0117] The system includes a memory 510 and a processor 520. The memory 510 stores computer programs and transfers the program code to the processor 520. In other words, the processor 520 can retrieve and run the computer program from the memory 510 to implement the methods described in the embodiments of this application.
[0118] For example, the processor 520 can be used to execute the above-described method embodiments according to instructions in the computer program.
[0119] In some embodiments of this application, the processor 520 may include, but is not limited to:
[0120] General-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0121] In some embodiments of this application, the memory 510 includes, but is not limited to:
[0122] Volatile memory and / or non-volatile memory. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static RAM (SRAM), Dynamic RAM (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDRSDRAM), Enhanced SDRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), and Direct Rambus RAM (DR RAM).
[0123] In some embodiments of this application, the computer program may be divided into one or more modules, which are stored in the memory 510 and executed by the processor 520 to perform the method provided in this application. The one or more modules may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program in the electronic device.
[0124] like Figure 5 As shown, the electronic device may also include:
[0125] Transceiver 530, which can be connected to processor 520 or memory 510.
[0126] The processor 520 can control the transceiver 530 to communicate with other devices; specifically, it can send information or data to other devices or receive information or data sent by other devices. The transceiver 530 may include a transmitter and a receiver. The transceiver 530 may further include antennas, and the number of antennas may be one or more.
[0127] It should be understood that the various components in the electronic device are connected through a bus system, which includes a data bus, a power bus, a control bus, and a status signal bus.
[0128] This application also provides a computer storage medium storing a computer program thereon, which, when executed by a computer, enables the computer to perform the methods of the above-described method embodiments. Alternatively, embodiments of this application also provide a computer program product containing instructions that, when executed by a computer, cause the computer to perform the methods of the above-described method embodiments.
[0129] When implemented using software, it can be implemented entirely or partially as a computer program product. This computer program product includes one or more computer instructions. When these computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., digital video disc (DVD)), or a semiconductor medium (e.g., solid-state disk (SSD)).
[0130] Those skilled in the art will recognize that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0131] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or modules may be electrical, mechanical, or other forms.
[0132] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. For example, the functional modules in the various embodiments of this application may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.
[0133] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for generating samples, characterized in that, include: A sample template is determined, wherein the data items under each header in the sample template are empty; Based on the sample template and the pre-defined header database, a corresponding initial sample is generated; wherein, each header in the sample template has a corresponding header database, and the header database stores identity information that conforms to the requirements of the header according to the specific format of the real identity information under the header; The initial samples are subjected to data augmentation processing under a real acquisition environment to obtain the corresponding training sample set; The step of generating corresponding initial samples based on the sample template and the pre-defined header database includes: For each header in the sample template, select an identity information item from the header database corresponding to the header and fill it into the header of the sample template to obtain the corresponding initial sample.
2. The method according to claim 1, characterized in that, The step of generating corresponding initial samples based on the sample template and the pre-defined header database includes: Based on the header database corresponding to the first type of header in the sample template, generate virtual data items under the first type of header; Based on the virtual data items associated with the first type of table header in the second type of table header in the sample template, generate virtual data items under the second type of table header; For each header in the first type and the second type, according to the position of the header in the sample template, the virtual data item under the header is added to the sample template to obtain the corresponding initial sample.
3. The method according to claim 2, characterized in that, The step of generating virtual data items under the first type of header based on the header database corresponding to the first type of header in the sample template includes: Based on the common attributes of the candidate data items in the header database corresponding to the first type of header in the sample template, the selection probability of the candidate data items is set; Based on the selection probability of the candidate data items, virtual data items under the first type of header are selected from the header database.
4. The method according to claim 1, characterized in that, The process of performing data augmentation on the initial samples under a real-world acquisition environment to obtain the corresponding training sample set includes: Based on the interference factors of samples in the real collection environment, determine the corresponding data augmentation method; The data augmentation methods are fully combined and fully permuted to obtain the corresponding target augmentation method set; By using each target augmentation method in the target augmentation method set, the initial sample is subjected to data augmentation processing to obtain the corresponding training sample set.
5. The method according to claim 4, characterized in that, The process of combining and permuting the data augmentation methods to obtain the corresponding target augmentation method set includes: The data augmentation methods are fully combined to obtain the corresponding data augmentation combination mode; By performing a full permutation of the data augmentation methods in each of the aforementioned data augmentation combination patterns, the corresponding target augmentation method set is obtained.
6. The method according to claim 1, characterized in that, The determination of the sample template includes: The data items under each header in the reference sample are blurred to obtain the sample template.
7. A sample generation apparatus, characterized in that, include: The sample template determination module is used to determine the sample template, wherein the data items under each header in the sample template are empty. The initial sample generation module is used to generate corresponding initial samples based on the sample template and the pre-set header database; wherein, each header in the sample template has a corresponding header database, and the header database stores identity information that conforms to the requirements of the header according to the specific format of the real identity information under the header; The sample augmentation module is used to perform data augmentation processing on the initial sample under a real acquisition environment to obtain the corresponding training sample set; The initial sample generation module is specifically used to select an identity information item from the header database corresponding to each header in the sample template and fill it into the header of the sample template to obtain the corresponding initial sample.
8. An electronic device, characterized in that, include: A processor and a memory, the memory being used to store a computer program, the processor being used to invoke and run the computer program stored in the memory to perform the sample generation method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, Used to store a computer program that causes a computer to perform the sample generation method as described in any one of claims 1-6.
10. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the method for generating samples as described in any one of claims 1-6.