Sample generation method and device, equipment and medium

By generating new label text and target objects to replace the original object, the problems of low sample quality and poor adaptability in the prior art are solved, and the diversity of samples and the generalization ability of the model are improved.

CN120199512APending Publication Date: 2025-06-24PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510266115.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

In the prior art, the sample generation method relies on a static, predefined sample set, resulting in the low quality of the generated samples and cannot adapt to different application scenarios, which limits the generalization ability and robustness of the diagnostic model.

Method used

By obtaining the object and label text in the original sample, a new label text is generated using a large language model, and then a target object is generated using the literary graph model, the object features are extracted, and the similarity is calculated to determine whether the target object meets the replacement requirements. If it is satisfied, the original object will be replaced and the target sample will be generated.

Benefits of technology

It improves the quality and diversity of samples, ensures the context consistency of the target samples, enhances the generalization ability and robustness of the model, and adapts to different application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120199512A_ABST
    Figure CN120199512A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, in particular to a sample generation method and device, equipment and a medium. The method is applied to medical scenes, and comprises the following steps: generating a plurality of new label texts according to label texts of objects in samples, enriching the objects in the samples, generating corresponding objects according to the new label texts, replacing original objects with the objects generated by the new label texts, and generating a plurality of target samples different from original samples. Other objects in the target sample and the original sample are not changed, contexts of the target sample are ensured to be consistent, interference in the generation process is reduced, the target sample can be compatible with a scene in the original sample, an abnormal object is added to the scene in the original sample, the diversity of the sample is increased, and the quality of the generated sample is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method, device, equipment and medium for sample generation. Background Art

[0002] With the further development of artificial intelligence technology, applying artificial intelligence to solve problems in the medical field has become a hot topic. Artificial intelligence technology can learn new medical methods in a short time and apply them in practice, which can, to a certain extent, make up for the shortage of doctors caused by the long training cycle, assist doctors in diagnosis, reduce the misdiagnosis rate and missed diagnosis rate, and improve the quality of medical diagnosis. However, the feature distribution shift between sample data and test data in medical assisted diagnosis is very likely to lead to disease misdiagnosis and bring great risks to patients.

[0003] In the prior art, an out-of-distribution detection method is used to generate sample data in medical assisted diagnosis. This method relies on a static and pre-defined sample set, and the generated sample quality is low, which cannot be applied to different application scenarios, restricting the generalization ability and robustness of the diagnosis model.

[0004] Therefore, in the process of sample generation, how to generate high-quality samples to adapt to different application scenarios has become an urgent problem to be solved. Summary of the Invention

[0005] In view of this, embodiments of the present invention provide a method, device, equipment and medium for sample generation to solve the problem of low quality of generated samples in the process of sample generation.

[0006] In a first aspect, an embodiment of the present invention provides a method for sample generation, and the method for sample generation includes: Obtain an original sample, an original object in the original sample, and the label text of the original object; Generate N new label texts corresponding to the label text according to the label text and a preset large language model, where N is an integer greater than zero; Generate N target objects corresponding to the new label texts according to the new label texts and a preset text-to-image model; Extract the image features of the original object and each target object respectively to obtain the original features of the original object and the target features of each target object; Calculate the similarity between the original features and the target features, and judge whether the target object meets the replacement requirement according to the similarity. If the replacement requirement is met, use the target object to replace the original object to obtain a target sample.

[0007] In a second aspect, an embodiment of the present invention provides a sample generation device, and the sample generation device includes: An acquisition module for acquiring original samples, original objects in the original samples, and label texts of the original objects; A first generation module for generating N new label texts corresponding to the label texts according to the label texts and a preset large language model, where N is an integer greater than zero; A second generation module for generating N target objects corresponding to the new label texts according to the new label texts and a preset text-to-image model; An extraction module for respectively extracting image features of the original object and each target object to obtain the original features of the original object and the target features of each target object; A replacement module for calculating the similarity between the original features and the target features, and judging whether the target object meets the replacement requirement according to the similarity. If the replacement requirement is met, the target object is used to replace the original object to obtain a target sample.

[0008] In a third aspect, an embodiment of the present invention provides a computer device, which includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the sample generation method described in the first aspect is implemented.

[0009] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the sample generation method described in the first aspect is implemented.

[0010] The beneficial effects of the present invention compared with the prior art are as follows: In the present invention, multiple new label texts are generated according to the label texts of the objects in the sample to enrich the objects in the sample. Corresponding objects are generated according to the new label texts, and the objects generated by the new label texts are used to replace the original objects to generate multiple target samples different from the original samples. The other objects in the target samples are the same as those in the original samples, ensuring the consistency of the context of the target samples, reducing interference in the generation process, and the target samples can be compatible with the scenarios in the original samples, adding abnormal objects to the scenarios in the original samples to increase the diversity of the samples and improve the quality of the generated samples. Description of the Drawings

[0011] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0012] Figure 1 It is a schematic diagram of the application environment of a sample generation method provided by an embodiment of the present invention; Figure 2 It is a schematic flowchart of a sample generation method provided by an embodiment of the present invention; Figure 3 It is a schematic structural diagram of a sample generation device provided by an embodiment of the present invention; Figure 4 It is a schematic structural diagram of a computer device provided by an embodiment of the present invention. Detailed implementation manners

[0013] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0014] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present invention. However, those skilled in the art should clearly understand that the present invention can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present invention.

[0015] It should be understood that when used in the specification and appended claims of the present invention, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0016] It should also be understood that the term "and / or" as used in the specification and appended claims of the present invention refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0017] As used in the specification and appended claims of the present invention, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" according to the context. Similarly, the phrase "if determined" or "if detecting [the described condition or event]" can be interpreted as meaning "once determined", "in response to determining", "once detecting [the described condition or event]", or "in response to detecting [the described condition or event]" according to the context.

[0018] In addition, in the description of the specification and the appended claims of the present invention, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and should not be construed as indicating or implying relative importance.

[0019] The reference to "one embodiment" or "some embodiments" in the description of the present invention means that a specific feature, structure, or characteristic described in connection with the embodiment is included in one or more embodiments of the present invention. Thus, the statements "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized.

[0020] The embodiments of the present invention can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is a theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.

[0021] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, mechatronics, etc. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0022] It should be understood that the magnitudes of the sequence numbers of the steps in the following embodiments do not mean the order of execution is prior or subsequent. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0023] In order to illustrate the technical solution of the present invention, specific embodiments are used for illustration below.

[0024] A sample generation method provided by an embodiment of the present invention can be applied, for example, in Figure 1In the application environment, the client communicates with the server. The client includes, but is not limited to, computer devices such as a personal digital assistant (PDA), a palmtop computer, a desktop computer, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a cloud computer device, and a personal digital assistant (PDA). The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, a content delivery network (CDN), and big data and artificial intelligence platforms.

[0025] See Figure 2 , which is a schematic flowchart of a sample generation method provided by an embodiment of the present invention. The above sample generation method can be applied to Figure 1 the server in Figure 2 As shown in

[0026] S201: Obtain an original sample, an original object in the original sample, and a label text of the original object.

[0027] In step S201, the original sample is sample data common in the corresponding scenario, and the original object is an object in the corresponding scenario of the original sample. Among them, the original object can be one or more. The label text is the label of the original object represented in text form.

[0028] In this embodiment, an original sample, an original object in the original sample, and a label text of the original object are obtained, where the original sample is corresponding image data, such as a medical image, and the original object is an object that appears in the image and can be any object in the image.

[0029] It should be noted that the original sample in this embodiment is the image data of a certain disease corresponding to a medical center. The disease can be pneumonia, and the medical center can be a hospital, a pharmaceutical company, etc. When it is necessary to use a prediction model to predict pneumonia for the image data, a large number of different pneumonia image data need to be used as sample data to train the prediction model. However, the pneumonia image data in the medical center may only contain the image data of common pneumonia disease symptoms. In practical applications, the prediction model predicts the pneumonia disease symptoms included in the pneumonia image data. If new symptoms appear, which are different from the pneumonia image data in the sample data, when the prediction model processes the pneumonia disease prediction of the new symptoms, the prediction results may not be satisfactory, resulting in false detection or missed detection of the new symptoms. Therefore, it is necessary to expand the pneumonia image data in the medical center and add the pneumonia image data corresponding to the new symptoms that may appear to help the prediction model better identify the new symptoms.

[0030] In this example, the original sample, the original object in the original sample, and the label text of the original object are obtained, so as to generate the corresponding new label text according to the corresponding label text, and then use the object corresponding to the new label text to replace the original object, enrich the objects included in the original sample, and increase the diversity of the original sample.

[0031] S202: According to the label text and a preset large language model, generate N new label texts corresponding to the label text, where N is an integer greater than zero.

[0032] In step S202, the preset large language model is a model that can generate corresponding text, and the new label text is a text similar to the label text of the original object.

[0033] In this embodiment, the preset large language model is the GPT-4 model. The label text is input into the GPT-4 model, and the GPT-4 model is used to generate N new labels corresponding to the label text, where N is an integer greater than zero and N is a pre-set value. Other large language models can also be used, which is not limited in this embodiment.

[0034] It should be noted that the N new label texts are texts similar to the label text, so that the object generated by the new label text can be used to replace the original object.

[0035] In this embodiment, according to the label text and a preset large language model, N new label texts corresponding to the label text are generated. The new label texts are diverse, enriching the diversity of the objects in the original sample, thereby increasing the diversity of the sample and providing a broad learning basis for training the corresponding model using the sample data.

[0036] S203: According to the new label text and a preset text-to-image model, generate N target objects corresponding to the new label text.

[0037] In step S203, the preset text-to-image model is a model that can generate an image corresponding to the text semantics according to the text. Each new label text can generate a corresponding target object, where the target object is the object corresponding to the new label text.

[0038] In this embodiment, the preset text-to-image model is the Stable Diffusion Inpainting model. According to the new label text and the preset text-to-image model, N target objects corresponding to the new label text are generated, and each target object can be represented by the corresponding new label text. That is, the input of the preset text-to-image model is the corresponding new label text, and the output is the corresponding target object. Among them, the preset text-to-image model can also be other text-to-image models, which are not limited in this embodiment. Each new label text can generate a target object, and N new labels correspond to N target objects.

[0039] It should be noted that the generated target object is an image containing the target object, where the size of the image containing the target object can be determined according to the pre-set image size.

[0040] In this embodiment, according to the new label text and the preset text-to-image model, N target objects corresponding to the new label text are generated, and the corresponding target object is directly generated according to the new label text, so as to replace the original object with the target object, improve the diversity of objects in the corresponding scenario, increase the sample data, enrich the diversity of samples, and increase the possibility of abnormal objects in the corresponding scenario, so that the corresponding model can learn the abnormal situations in the corresponding scenario, thereby improving the ability of the corresponding model to recognize abnormal situations and enhancing the generalization ability and robustness of the model.

[0041] Optionally, generating N target objects corresponding to the new label text according to the new label text and the preset text-to-image model includes: For any new label text, generating an object sequence corresponding to the new label text according to the new label text and the preset text-to-image model; Extracting features of each object in the object sequence to obtain the object features of each object; Filtering the object sequence according to the object features of each object to obtain the target object corresponding to the new label text.

[0042] In this embodiment, for any new label text, an object sequence corresponding to the new label text is generated according to the new label text and a pre-set text-to-image model. That is, for a new label text, multiple objects corresponding to the new label text can be generated. Each object is generated in the form of an image. The object sequence can be an image sequence of the same target with different features. For example, in the process of disease diagnosis, the new label text is the lungs with a disease. That is, an object sequence of the lungs with a disease is generated according to the corresponding new label text. Different pathologies have different effects on the lungs, and the size, shape, and color of each lung object in the generated object sequence may be different. The pre-set text-to-image model can be the Stable Diffusion Inpainting model or other text-to-image models, which is not limited in this embodiment.

[0043] Feature extraction is performed on each object in the object sequence to obtain the object features of each object. That is, image feature extraction is performed on the image corresponding to each object. When performing feature extraction, a pre-set image encoder can be used for feature extraction. For example, the image encoder of the Clip (Contrastive Language-Image Pre-Training) model is used to perform feature extraction on each object in the object sequence to obtain the object features of each object. It can also be other encoder models, which is not limited in this embodiment.

[0044] According to the object features of each object, the object sequence is screened to obtain the target object corresponding to the new label text. When screening the object sequence, the object sequence can be screened according to the similarity between the object features of each object and the mean of the object features of all objects in the object sequence. The object corresponding to the maximum similarity value between the object features and the mean of the object features of all objects in the object sequence is determined as the target object. That is, the object when the object features are closest to the mean of the object features of all objects in the object sequence is determined as the target object. That is, according to the object features of each object, the mean of the object features of all objects in the object sequence is calculated to obtain the mean feature, and the object corresponding to the object feature closest to the mean feature is determined as the target object. When determining the object feature closest to the mean feature, the cosine value between the object feature and the mean feature can be calculated. The larger the cosine value, the closer it is to the mean feature.

[0045] In this embodiment, according to the object features of each object, the object sequence is screened to select the corresponding target object. When screening, the object when the object features are closest to the mean of the object features of all objects in the object sequence is determined as the target object, so that the object features of the selected target object are similar to the object features of other objects, avoiding the risk that the selected target object is quite different from the object represented by the new label text.

[0046] Optionally, according to the object features of each object, the object sequence is filtered to obtain the target object corresponding to the new label text, including: Feature extraction is performed on the original sample to obtain the sample features of the original sample; Calculate the similarity between the sample features and the object features of each object in the object sequence, and select the object corresponding to the maximum similarity value as the target object corresponding to the new label text.

[0047] In this embodiment, feature extraction is performed on the original sample to obtain the sample features of the original sample. Among them, the sample features of the original sample include the features of the original object and the features of the context corresponding to the original object. According to the sample features and the object features, the object sequence is filtered to obtain the target object corresponding to the new label text. Among them, when filtering the object sequence, the similarity between the sample features and the object features of each object in the object sequence can be calculated, and the object corresponding to the maximum similarity value is selected as the target object corresponding to the new label text.

[0048] In this embodiment, the object corresponding to the maximum similarity value is selected as the target object corresponding to the new label text. When filtering, the sample features of the original sample are considered, and filtering is performed according to the similarity between the sample features of the original sample and the object features of each object in the object sequence, so that the selected target object is more in line with the corresponding scenario of the original sample. After using the target object to replace the original object, the consistency of the text on the target sample can be maintained, and the rationality of the target sample can be improved.

[0049] S204: Respectively extract the image features of the original object and each target object to obtain the original features of the original object and the target features of each target object.

[0050] In step S204, the image features of the original object and each target object are respectively extracted to obtain the original features of the original object and the target features of each target object, so as to determine whether to use the target object to replace the original object according to the original features and the target features.

[0051] In this embodiment, the image features of the original object and each target object are respectively extracted. Among them, when extracting the image features of the original object and each target object, a preset image encoder can be used for feature extraction. For example, the image encoder of the Clip (Contrastive Language-Image Pre-Training) model can be used for extraction to obtain the original features of the original object and the target features of each target object. It can also be other encoder models, which are not limited in this embodiment.

[0052] In this embodiment, the image features of the original object and each target object are extracted respectively, so as to determine whether to replace the original object with the target object according to the original features and the target features, avoiding that when the target object is too similar to the original object, the obtained target sample is meaningless for the training of the model, and when the target object is too different from the original object, the obtained target sample is invalid.

[0053] S205: Calculate the similarity between the original features and the target features. According to the similarity, determine whether the target object meets the replacement requirement. If the replacement requirement is met, use the target object to replace the original object to obtain a target sample.

[0054] In step S205, calculate the similarity between the original features and the target features. According to the similarity, determine whether the target object meets the replacement requirement. Among them, the similarity between the original features and the target features represents the similarity between the original object and the target object. The greater the similarity, the more similar the original object and the target object are. The smaller the similarity, the less similar the original object and the target object are. If the replacement requirement is met, use the target object to replace the original object to obtain a target sample, that is, remove the original object in the original sample and place the target object at the position of the original object to obtain a target sample.

[0055] In this embodiment, when calculating the similarity between the original features and the target features, the distance formula can be used for calculation. The original features and the target features are represented by corresponding vectors, and the distance between the corresponding vectors of the original features and the target features is calculated. The corresponding distance is used as the similarity between the original features and the target features. Calculate the similarity between the original features and the target features. According to the similarity, determine whether the target object meets the replacement requirement. Among them, the replacement condition can be that when the similarity between the original features and the target features is less than a preset threshold, it is determined that the target object meets the replacement requirement.

[0056] It should be noted that when the target object meets the replacement requirement, use the target object to replace the original object. When replacing, obtain the detection frame of the original object, and align the center position of the target object with the center position of the detection frame of the original object to obtain a target sample.

[0057] It should be noted that there are multiple new label texts generated according to the label text in the original object, and there are multiple corresponding target objects. That is, when replacing one of the original objects in an original sample, if there are multiple target objects that meet the replacement condition, use the multiple target objects that meet the replacement condition to replace the original object respectively, and multiple target samples can be obtained.

[0058] In this embodiment, when calculating the similarity between the original feature and the target feature and determining whether the target object meets the replacement requirement according to the similarity, when the similarity between the original feature and the target feature is less than the preset threshold, it is determined that the target object meets the replacement requirement. Among them, the smaller the similarity between the original feature and the target feature, the greater the difference between the target object and the original object. After replacing the original object with the target object, the diversity of the objects in the original sample is increased. Multiple target samples can be obtained from one original sample, enriching the sample data, so that when using the target samples to train the corresponding model, it can help the model better identify newly emerging anomalies, so that when encountering various unknown samples that may occur in actual applications, the model can perform more stably.

[0059] Optionally, calculating the similarity between the original feature and the target feature includes: Calculating the cosine value between the original feature and the target feature, and determining the cosine value as the similarity between the original feature and the target feature.

[0060] In this embodiment, the original feature and the target feature are represented by corresponding vectors. Calculate the cosine value between the vector corresponding to the original feature and the vector corresponding to the target feature, and determine the cosine value as the similarity between the original feature and the target feature.

[0061] In this embodiment, the cosine formula is used to calculate the corresponding similarity, only considering the directions of the vectors corresponding to the original feature and the target feature, and ignoring their lengths. Therefore, even if the lengths of the vectors are different, as long as their directions are the same, the cosine similarity can be very high. The accuracy of calculating the similarity is improved.

[0062] Optionally, determining whether the target object meets the replacement requirement according to the similarity includes: Obtain the preset similarity range. If the similarity is within the preset similarity range, it is determined that the target object meets the replacement requirement. If the similarity is not within the preset similarity range, it is determined that the target object does not meet the replacement requirement.

[0063] In this embodiment, when determining whether the target object meets the replacement requirement, it is judged according to the similarity between the original feature and the target feature. The similarity between the original feature and the target feature characterizes the similarity between the original object and the target object. When the similarity is greater, the original object and the target object are more similar. When the similarity is smaller, the original object and the target object are less similar. If the similarity is greater, the difference between the target object and the original object is smaller, and the significance of replacing the original object with the target object is not great, and it cannot improve the recognition of abnormal objects by the corresponding model. If the similarity is smaller, the difference between the target object and the original object is greater. After replacing the original object with the target object, the context difference between the target object and the original sample is greater, and the authenticity of the appearance of the abnormal object is smaller.

[0064] To increase the validity and authenticity of the target sample, a preset similarity range is obtained. If the similarity is within the preset similarity range, it is determined that the target object meets the replacement requirement. If the similarity is not within the preset similarity range, it is determined that the target object does not meet the replacement requirement. Among them, when the similarity is within the preset similarity range, that is, the similarity is greater than the first preset threshold and less than the second preset threshold, and the first preset threshold is less than the second preset threshold. When the similarity is greater than the first preset threshold, it is considered that there are corresponding differences between the selected target object and the original object, preventing the selected target object from being invalid. When the similarity is less than the second preset threshold, it is considered that the difference between the selected target object and the original object cannot be too large, preventing the selected target object from lacking authenticity.

[0065] Optionally, after determining whether the target object meets the replacement requirement, it further includes: If it does not meet the replacement requirement, a new target object that meets the replacement requirement is obtained, and the original object is replaced with the new target object to obtain the target sample.

[0066] In this embodiment, when using N target objects to replace the corresponding original objects respectively, each target object is judged. When the replacement condition is met, the corresponding target object is used to replace the original object. If one of the target objects does not meet the replacement condition, for example, the similarity between the target object and the original object is too high, a new target object that meets the replacement requirement is obtained, and the original object is replaced with the new target object to obtain the target sample, that is, this target object is discarded until all N target objects are judged.

[0067] In this embodiment, when judging each target object, if the target object does not meet the replacement requirement, the target object is not used to replace the original object, and a new target object that meets the replacement requirement is continuously used for replacement to ensure that all obtained target samples are valid and real.

[0068] Optionally, using the target object to replace the original object to obtain the target sample includes: Obtain the detection frame of the original object; According to the detection frame of the original object, adjust the size of the target object to obtain the adjusted target object, so that the size of the adjusted target object is equal to the size of the detection frame of the original object; In the detection frame, use the adjusted target object to replace the original object to obtain the target sample.

[0069] In this embodiment, when replacing the original object with the target object, the detection frame of the original object is obtained, and according to the detection frame of the original object, the size of the target object is adjusted to obtain the adjusted target object, so that the size of the adjusted target object is equal to the size of the detection frame of the original object. That is, when the target object and the detection frame of the original object cannot completely coincide, the size of the target object is adjusted so that the size of the adjusted target object is equal to the size of the detection frame of the original object, so that the adjusted target object can completely coincide with the original object. In the detection frame, the original object is replaced with the adjusted target object to obtain the target sample. Among them, the representation of the target sample is as follows: Among them, is the target sample, is the identifier of the original sample, is the detection frame of the original object, that is, the detection frame of the target object, is the target object. That is, the target sample includes the identifier of the original sample, the detection frame of the target object and the target object.

[0070] In the present invention, multiple new label texts are generated according to the label text of the object in the sample to enrich the objects in the sample. Corresponding objects are generated according to the new label texts, and the objects generated by the new label texts are used to replace the original objects to generate multiple target samples different from the original sample. The other objects in the target sample and the original sample remain unchanged to ensure the consistency of the context of the target sample, reduce the interference in the generation process, and the target sample can be compatible with the scene in the original sample, adding abnormal objects to the scene in the original sample to increase the diversity of the sample and improve the quality of the generated sample.

[0071] Please refer to Figure 3 , Figure 3 which is a schematic structural diagram of a sample generation device provided by an embodiment of the present invention. This sample generation device corresponds one-to-one with the sample generation method in the above embodiment. For details, please refer to Figure 2 and Figure 2 the relevant descriptions in the corresponding embodiments. For the sake of convenience of description, only the parts related to this embodiment are shown. See Figure 3 , the sample generation device 30 includes: an acquisition module 31, a first generation module 32, a second generation module 33, an extraction module 34, and a replacement module 35.

[0072] The acquisition module 31 is used to acquire the original sample, the original object in the original sample, and the label text of the original object.

[0073] The first generation module 32 is used to generate N new label texts corresponding to the label text according to the label text and a preset large language model, where N is an integer greater than zero.

[0074] The second generation module 33 is used to generate N target objects corresponding to the new label text according to the new label text and a preset text-to-image model.

[0075] The extraction module 34 is used to extract the image features of the original object and each target object respectively, so as to obtain the original features of the original object and the target features of each target object.

[0076] The replacement module 35 is used to calculate the similarity between the original features and the target features, and judge whether the target object meets the replacement requirement according to the similarity. If the replacement requirement is met, the target object is used to replace the original object to obtain a target sample.

[0077] Optionally, the above-mentioned second generation module 33 includes: The first generation unit is used to generate an object sequence corresponding to the new label text according to the new label text and a preset text-to-image model for any new label text.

[0078] The obtaining unit is used to perform feature extraction on each object in the object sequence to obtain the object features of each object.

[0079] The screening unit is used to screen the object sequence according to the object features of each object to obtain the target object corresponding to the new label text.

[0080] Optionally, the above-mentioned screening unit includes: The obtaining subunit is used to perform feature extraction on the original sample to obtain the sample features of the original sample.

[0081] The calculation subunit is used to calculate the similarity between the sample features and the object features of each object in the object sequence, and select the object corresponding to the maximum similarity value as the target object corresponding to the new label text.

[0082] Optionally, the above-mentioned replacement module 35 includes: The calculation unit is used to calculate the cosine value between the original features and the target features, and determine the cosine value as the similarity between the original features and the target features.

[0083] Optionally, the above-mentioned replacement module 35 further includes: The judgment unit is used to obtain a preset similarity range. If the similarity is within the preset similarity range, it is determined that the target object meets the replacement requirement. If the similarity is not within the preset similarity range, it is determined that the target object does not meet the replacement requirement.

[0084] Optionally, the above-mentioned replacement module 35 further includes: The first obtaining unit is used to, if the replacement requirement is not met, obtain a new target object that meets the replacement requirement, and use the new target object to replace the original object to obtain a target sample.

[0085] Optionally, the above replacement module 35 further includes: A second acquisition unit, configured to acquire the detection frame of the original object.

[0086] An adjustment unit, configured to adjust the size of the target object according to the detection frame of the original object, so as to obtain an adjusted target object, and make the size of the adjusted target object equal to the size of the detection frame of the original object.

[0087] A replacement unit, configured to replace the original object with the adjusted target object in the detection frame, so as to obtain a target sample.

[0088] It should be noted that, for the information interaction, execution process, etc. among the above units, since they are based on the same concept as the method embodiment of the present invention, for their specific functions and the technical effects brought, reference can be specifically made to the method embodiment part, and details are not described herein again.

[0089] Figure 4 is a schematic structural diagram of a computer device provided by an embodiment of the present invention. As Figure 4 shown, the computer device of this embodiment includes: at least one processor ( Figure 4 only one is shown in the figure), a memory, and a computer program stored in the memory and executable on at least one processor. When the processor executes the computer program, the steps in any of the above method embodiments for generating samples are implemented.

[0090] The computer device may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that Figure 4 merely examples of computer devices are given, which do not constitute a limitation on computer devices. A computer device may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, it may also include a network interface, a display screen, and an input device, etc.

[0091] The so-called processor may be a CPU, and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0092] The memory includes a readable storage medium, an internal memory, etc. Among them, the internal memory can be the memory of a computer device, and the internal memory provides an environment for the operation of the operating system and computer-readable instructions in the readable storage medium. The readable storage medium can be the hard disk of a computer device, and in some other embodiments, it can also be an external storage device of the computer device, for example, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the computer device. Further, the memory can also include both the internal storage unit of the computer device and the external storage device. The memory is used to store the operating system, application programs, a BootLoader, data, and other programs, etc., and the other programs such as the program code of a computer program. The memory can also be used to temporarily store the data that has been output or will be output.

[0093] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above device can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above method embodiments of this application, a computer program can be used to instruct the relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can at least include: any entity or device capable of carrying the computer program code, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.

[0094] All or part of the processes in the above method embodiments of this application can also be completed by a computer program product. When the computer program product runs on a computer device, it enables the computer device to execute and implement the steps in the above method embodiments.

[0095] In the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0096] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.

[0097] In the embodiments provided in this application, it should be understood that the disclosed device / computer equipment and method can be implemented in other ways. For example, the device / computer equipment embodiments described above are only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.

[0098] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0099] The above embodiments are only used to illustrate the technical solutions of this application, rather than to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of each embodiment of this application, and should all be included in the protection scope of this application.

Claims

1. A sample generation method, characterized in that: The sample generation method comprises: Obtaining an original sample, an original object in the original sample, and a label text of the original object; Generate N new label texts corresponding to the label text according to the label text and a preset large language model, where N is an integer greater than zero; Generate N target objects corresponding to the new label text according to the new label text and a preset text graph model; Extracting image features of the original object and each target object respectively to obtain original features of the original object and target features of each target object; The similarity between the original feature and the target feature is calculated, and based on the similarity, it is determined whether the target object meets the replacement requirement. If the replacement requirement is met, the target object is used to replace the original object to obtain a target sample.

2. The sample generation method according to claim 1, characterized in that: The step of generating N target objects corresponding to the new label text according to the new label text and a preset text graph model includes: For any new label text, generating an object sequence corresponding to the new label text according to the new label text and a preset text graph model; Extracting features from each object in the object sequence to obtain object features of each object; The object sequence is screened according to the object feature of each object to obtain the target object corresponding to the new label text.

3. The sample generation method according to claim 2, characterized in that: The object sequence is screened according to the object feature of each object to obtain the target object corresponding to the new label text, including: Performing feature extraction on the original sample to obtain sample features of the original sample; The similarity between the sample feature and the object feature of each object in the object sequence is calculated, and the object corresponding to the maximum similarity value is selected as the target object corresponding to the new label text.

4. The sample generation method according to claim 1, characterized in that: The calculating the similarity between the original feature and the target feature includes: A cosine value between the original feature and the target feature is calculated, and the cosine value is determined as the similarity between the original feature and the target feature.

5. The sample generation method according to claim 1, characterized in that: The determining, based on the similarity, whether the target object meets the replacement requirement includes: A preset similarity range is obtained. If the similarity is within the preset similarity range, it is determined that the target object meets the replacement requirement. If the similarity is not within the preset similarity range, it is determined that the target object does not meet the replacement requirement.

6. The sample generation method according to claim 1, characterized in that: After determining whether the target object meets the replacement requirement, the method further includes: If the replacement requirement is not met, a new target object that meets the replacement requirement is obtained, and the new target object is used to replace the original object to obtain a target sample.

7. The sample generation method according to claim 1, characterized in that: The step of replacing the original object with the target object to obtain a target sample includes: Obtaining a detection frame of the original object; According to the detection frame of the original object, the size of the target object is adjusted to obtain an adjusted target object, so that the size of the adjusted target object is equal to that of the detection frame of the original object; In the detection frame, the original object is replaced with the adjusted target object to obtain a target sample.

8. A sample generating device, characterized in that: The sample generating device comprises: An acquisition module, used to acquire an original sample, an original object in the original sample, and a label text of the original object; A first generating module, configured to generate N new label texts corresponding to the label text according to the label text and a preset large language model, where N is an integer greater than zero; A second generating module is used to generate N target objects corresponding to the new label text according to the new label text and a preset text graph model; An extraction module, used to extract the image features of the original object and each target object respectively, to obtain the original features of the original object and the target features of each target object; The replacement module is used to calculate the similarity between the original feature and the target feature, and judge whether the target object meets the replacement requirements according to the similarity. If the replacement requirements are met, the target object is used to replace the original object to obtain a target sample.

9. A computer device, characterized in that: The computer device includes a processor, a memory, and a computer program stored in the memory and executable on the processor, and the processor implements the sample generation method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the sample generation method according to any one of claims 1 to 7 is implemented.