Data processing method and device of machine learning model, electronic equipment, computer readable storage medium and computer program product
By identifying and combining image elements and generating diverse training data, the problem of machine learning models' low understanding of unconventional data is solved, and the accuracy and generalization ability of the model are improved.
Patent Information
- Application Number
- CN202510322768.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-08
AI Technical Summary
The prior art uses conventional text image pairs when training machine learning models, resulting in low understanding of unconventional data and low accuracy.
By identifying and combining elements on the original image samples, extracting subsets of high- and low-frequency elements, generating text samples with negative and positive descriptions, combining image generation, diverse training data are formed to eliminate the model's bias against unconventional data.
It improves the understanding and accuracy of machine learning models for unconventional data, and enhances the generalization ability and comparative learning effect of the model.
Smart Images

Figure CN120278213A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a data processing method, apparatus, electronic device, computer-readable storage medium, and computer program product for a machine learning model. Background Art
[0002] In the field of machine learning, especially in tasks involving image and text processing, how to effectively generate targeted training data to improve the performance of machine learning models is an important research direction. Different training sets often lead to different training effects. In the related art, when training a machine learning model, conventional text-image pairs are often used as the training set, resulting in a low understanding ability of the machine learning model for unconventional data and a low accuracy of the machine learning model. Summary of the Invention
[0003] Embodiments of this application provide a data processing method, apparatus, electronic device, computer-readable storage medium, and computer program product for a machine learning model, which can improve the accuracy of the machine learning model.
[0004] The technical solution of the embodiments of this application is implemented as follows:
[0005] Embodiments of this application provide a data processing method for a machine learning model, the method comprising:
[0006] Performing object recognition on a plurality of original image samples in a first training set, and combining the recognized elements into an element set;
[0007] Extracting a plurality of element subsets from the element set, and determining the number of occurrences of the element subsets in the plurality of original image samples;
[0008] Based on the number of occurrences, identifying at least one of a high-frequency element subset and a low-frequency element subset from the plurality of element subsets;
[0009] Generating text based on at least one of the high-frequency element subset and the low-frequency element subset to obtain a first generated text sample, wherein the first generated text sample includes at least one of a negative description of the high-frequency element subset and a positive description of the low-frequency element subset;
[0010] Generating an image based on the first generated text sample to obtain a generated image sample, wherein the first generated text sample and the generated image sample are used to train the machine learning model.
[0011] Embodiments of this application provide a data processing method for a machine learning model, the method comprising:
[0012] Obtain a first training set, where the first training set includes a plurality of original text samples and a plurality of original image samples, and one of the original text samples is used to describe one of the original image samples;
[0013] For each of the original image samples described by the original text samples, perform the following processing:
[0014] Identify first element samples from the original image samples;
[0015] Generate second element samples that do not appear in the original image samples based on the first element samples;
[0016] Rewrite the original text samples based on the second element samples to obtain second generative text samples for negatively describing the second element samples, where the second generative text samples and the original image samples are used to train the machine learning model.
[0017] In the above solution, the first element samples include main element samples and secondary element samples; the identifying first element samples from the original image samples includes:
[0018] Perform feature recognition on the original image samples to obtain a plurality of candidate element samples;
[0019] Use the candidate element samples that meet the preset conditions as the main element samples, where the preset conditions include at least one of the following: the distance between the position of the candidate element sample and the center of the original image sample is less than a distance threshold, the ratio of the size of the candidate element sample to the size of the original image sample is greater than a size threshold, and the similarity between the candidate features of the candidate element sample and the element features of other element samples in the original image sample except the candidate element sample is less than a similarity threshold;
[0020] Use the candidate element samples other than the main element samples among the plurality of candidate element samples as the secondary element samples.
[0021] In the above solution, the generating second element samples that do not appear in the original image samples based on the first element samples includes:
[0022] Generate at least one of similar elements of the first element samples and related elements of the first element samples, where the similar elements are elements with a similarity greater than a similarity threshold to the first element samples, and the related elements are elements with a co-occurrence frequency greater than a frequency threshold with the first element samples in image samples other than the original image sample;
[0023] Use the similar elements and the related elements as the second element samples.
[0024] In the above solution, the original text sample includes a positive description of the first element sample; the rewriting of the original text sample based on the second element sample to obtain a second generative text sample for negatively describing the second element sample includes:
[0025] Generating a negative description of the second element sample;
[0026] Combining the negative description of the second element sample and the positive description of the first element sample to obtain a second generative text sample.
[0027] An embodiment of the present application provides a data processing method for a machine learning model, and the method includes:
[0028] Performing a first rewriting on the original text sample in the first training set to obtain a first generative text sample, where the original text sample is used to describe a first element sample included in the original image sample in the first training set, and the first generative text sample is used to describe high-frequency elements and low-frequency elements co-occurring with the first element in a plurality of the original image samples;
[0029] Generating a generative image sample based on the first generative text sample;
[0030] Performing a second rewriting on the original text sample to obtain a second generative text sample, where the second generative text sample is used to describe the first element sample and a second element sample not present in the original image sample, and the first training set composed of the original text sample and the original image sample, the second training set composed of the first generative text sample and the generative image sample, and the third training set composed of the second generative text sample and the original image sample are used to train the machine learning model.
[0031] In the above solution, when the number of the original image samples is multiple, the performing a first rewriting on the original text sample in the first training set to obtain a first generative text sample includes:
[0032] Performing target recognition on the multiple original image samples, and combining the recognized elements into an element set;
[0033] Extracting multiple element subsets from the element set, and determining the number of occurrences of the element subsets in the multiple original image samples;
[0034] Based on the number of occurrences, identifying at least one of a high-frequency element subset and a low-frequency element subset from the multiple element subsets;
[0035] Perform a first rewrite on the original text sample based on at least one of the high-frequency element subset and the low-frequency element subset to obtain a first generated text sample, where the first generated text sample includes at least one of a negative description of the high-frequency element subset and a positive description of the low-frequency element subset.
[0036] In the above solution, the second rewrite of the original text sample to obtain a second generated text sample includes:
[0037] For each original image sample described by the original text sample, perform the following processing:
[0038] Identify the first element sample from the original image sample;
[0039] Generate a second element sample that does not appear in the original image sample based on the first element sample;
[0040] Rewrite the original text sample based on the second element sample to obtain a second generated text sample for negatively describing the second element sample.
[0041] An embodiment of the present application provides a data processing device for a machine learning model. The device includes:
[0042] An element recognition module, configured to perform target recognition on a plurality of original image samples in a first training set, and combine the recognized elements into an element set;
[0043] A subset extraction module, configured to extract a plurality of element subsets from the element set, and determine the number of occurrences of the element subsets in the plurality of original image samples;
[0044] A subset recognition module, configured to recognize at least one of a high-frequency element subset and a low-frequency element subset from the plurality of element subsets based on the number of occurrences;
[0045] A first generation module, configured to perform text generation based on at least one of the high-frequency element subset and the low-frequency element subset to obtain a first generated text sample, where the first generated text sample includes at least one of a negative description of the high-frequency element subset and a positive description of the low-frequency element subset;
[0046] A second generation module, configured to perform image generation based on the first generated text sample to obtain a generated image sample, where the first generated text sample and the generated image sample are used to train the machine learning model.
[0047] An embodiment of the present application provides a data processing device for a machine learning model. The device includes:
[0048] A data acquisition module for acquiring a first training set, where the first training set includes a plurality of original text samples and a plurality of original image samples, and one of the original text samples is used to describe one of the original image samples;
[0049] Perform the following processing on the original image sample described by each of the original text samples:
[0050] A sample recognition module for recognizing first element samples from the original image samples;
[0051] A third generation module for generating second element samples that do not appear in the original image samples based on the first element samples;
[0052] A sample rewriting module for rewriting the original text sample based on the second element samples to obtain second generative text samples for negatively describing the second element samples, where the second generative text samples and the original image samples are used to train the machine learning model.
[0053] An embodiment of the present application provides a data processing device for a machine learning model, the device includes:
[0054] A first rewriting module for performing a first rewrite on the original text samples in the first training set to obtain first generative text samples, where the original text samples are used to describe the first element samples included in the original image samples in the first training set, and the first generative text samples are used to describe high-frequency elements and low-frequency elements that co-occur with the first element in a plurality of the original image samples;
[0055] A fourth generation module for generating generative image samples based on the first generative text samples;
[0056] A second rewriting module for performing a second rewrite on the original text samples to obtain second generative text samples, where the second generative text samples are used to describe the first element samples and second element samples that do not appear in the original image samples, and the first training set composed of the original text samples and the original image samples, the second training set composed of the first generative text samples and the generative image samples, and the third training set composed of the second generative text samples and the original image samples are used to train the machine learning model.
[0057] An embodiment of the present application provides an electronic device, the electronic device includes:
[0058] A memory for storing computer-executable instructions or computer programs;
[0059] A processor, when executing computer-executable instructions or computer programs stored in the memory, implements the data processing method of the machine learning model provided by the embodiments of the present application.
[0060] The embodiments of the present application provide a computer-readable storage medium storing a computer program or computer-executable instructions, which are used to implement the data processing method of the machine learning model provided by the embodiments of the present application when being executed by a processor.
[0061] The embodiments of the present application provide a computer program product including a computer program or computer-executable instructions, which implement the data processing method of the machine learning model provided by the embodiments of the present application when the computer program or computer-executable instructions are executed by a processor.
[0062] The embodiments of the present application have the following beneficial effects:
[0063] By combining the element combinations identified from multiple original image samples into an element set, extracting multiple element subsets from the element set, identifying high-frequency element subsets and low-frequency element subsets from the element subsets based on the occurrence times, generating a first generative text sample based on at least one of the negative description of the high-frequency element subset and the positive description of the low-frequency element subset, and obtaining a generative image sample based on the first generative text sample. Since the negative description of the high-frequency element subset or the positive description of the low-frequency element subset is text contrary to conventional text, the generative image sample obtained through the first generative text sample is also an unconventional image. Compared with training a machine learning model only using conventional data, the bias of the machine learning model against unconventional data is eliminated, and the accuracy of the machine learning model is improved. Description of the Drawings
[0064] Figure 1 is a schematic architecture diagram of a data processing system 100 of a machine learning model provided by the embodiments of the present application;
[0065] Figure 2A is a schematic structural diagram of a server 200-1 provided by the embodiments of the present application;
[0066] Figure 2B is a schematic structural diagram of a server 200-2 provided by the embodiments of the present application;
[0067] Figure 2C is a schematic structural diagram of a server 200-3 provided by the embodiments of the present application;
[0068] Figure 3A is a first flowchart of the data processing method of the machine learning model provided by the embodiments of the present application;
[0069] Figure 3B It is the second process schematic diagram of the data processing method of the machine learning model provided by the embodiment of the present application;
[0070] Figure 3C It is the third process schematic diagram of the data processing method of the machine learning model provided by the embodiment of the present application;
[0071] Figure 3D It is the fourth process schematic diagram of the data processing method of the machine learning model provided by the embodiment of the present application;
[0072] Figure 3E It is the fifth process schematic diagram of the data processing method of the machine learning model provided by the embodiment of the present application;
[0073] Figure 3F It is the sixth process schematic diagram of the data processing method of the machine learning model provided by the embodiment of the present application;
[0074] Figure 3G It is the seventh process schematic diagram of the data processing method of the machine learning model provided by the embodiment of the present application;
[0075] Figure 3H It is the eighth process schematic diagram of the data processing method of the machine learning model provided by the embodiment of the present application;
[0076] Figure 3I It is the ninth process schematic diagram of the data processing method of the machine learning model provided by the embodiment of the present application;
[0077] Figure 3J It is the tenth process schematic diagram of the data processing method of the machine learning model provided by the embodiment of the present application;
[0078] Figure 3K It is the eleventh process schematic diagram of the data processing method of the machine learning model provided by the embodiment of the present application;
[0079] Figure 3L It is the twelfth process schematic diagram of the data processing method of the machine learning model provided by the embodiment of the present application;
[0080] Figure 3M It is the thirteenth process schematic diagram of the data processing method of the machine learning model provided by the embodiment of the present application;
[0081] Figure 4 It is the principle schematic diagram of training the machine learning model provided by the embodiment of the present application;
[0082] Figure 5 It is the process schematic diagram of training the machine learning model provided by the embodiment of the present application;
[0083] Figure 6It is a schematic structural diagram of the multi-modal pre-trained neural network model provided by an embodiment of the present application;
[0084] Figure 7 It is a schematic principle diagram of training a multi-modal model provided by an embodiment of the present application;
[0085] Figure 8 It is a schematic diagram of constructing a negative semantic text provided by an embodiment of the present application;
[0086] Figure 9 It is a schematic diagram of constructing a complex image through an image generation model provided by an embodiment of the present application.
[0087] It should be noted that the above-mentioned "first" and "second" are only used to distinguish different solutions, and do not represent the distinction of the quality of the solutions or the priority in the implementation process. Detailed implementation manners
[0088] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations of the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present application.
[0089] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0090] In the following description, the terms "first / second / third" are only used to distinguish similar objects, and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged in a specific order or sequence when allowed, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0091] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.
[0092] Unless otherwise specified, at least one as described below refers to a case of one or more, and "multiple" can refer to a case of two or more.
[0093] Unless otherwise defined, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by those of ordinary skill in the art to which this application pertains. The terms used in the embodiments of this application are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0094] In the practical application of data collection and processing in the embodiments of this application, it should strictly comply with the requirements of relevant laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing behaviors within the scope authorized by laws and regulations and the personal information subject.
[0095] Before further elaborating on the embodiments of this application, the nouns and terms involved in the embodiments of this application are explained. The nouns and terms involved in the embodiments of this application are applicable to the following explanations.
[0096] 1) In response to: used to indicate the conditions or states upon which the executed operations depend. When the dependent conditions or states are met, one or more of the executed operations can be real-time or can have a set delay; without special instructions, there is no limitation on the execution order of multiple executed operations.
[0097] 2) Human-computer interaction interface, an interface for providing human-computer interaction functions / an interface for displaying images.
[0098] For example, a graphical user interface (GUI) display, such as an augmented reality (AR) interface, a virtual reality (VR) interface, a voice user interface (VUI), an interactive projection interface (using projection technology to display information on a plane), an eye movement detection interface (an interface controlled by detecting the user's line of sight), a holographic interface (a three-dimensional hologram formed by projecting an image through holographic projection technology, allowing a stereoscopic image to be seen without wearing special glasses), a multimodal interface (an interactive interface that combines multiple interaction methods such as touch, vision, and hearing), a brain-machine interface (BMI) interface, etc.
[0099] 3) Element, refers to an object or entity in the original image sample that has specific semantics and features. An element is the basic component that constitutes the image content. For example, for an original image sample of a cream cake, the elements include the cake, the cream, etc.
[0100] 4) High-frequency element subset, refers to the element subset of a preset proportion starting from the top when the element subsets are sorted in descending order according to the number of occurrences.
[0101] 5) A low-frequency element subset refers to the element subsets other than the high-frequency element subsets in the sorting result when sorting the element subsets in descending order of the number of occurrences.
[0102] 6) A positive description is a direct and affirmative statement of an element or a combination of elements, emphasizing its existence or characteristics. For example, "This is a football cake" clearly indicates the combination of the cake and the football exists in the described situation.
[0103] 7) A negative description is made by negating or excluding certain elements. It does not directly state the existence of elements but emphasizes the non-existence of certain elements or the difference from the normal situation. For example, "This picture is not a cream cake" negates the overall existence of the cream cake, and "This is a cake without cream" negates the situation of cream in the cake.
[0104] 8) A main element sample refers to an object or area in the original image sample that undertakes the core narrative function, has high visual saliency, and is directly related to the image theme. The characteristics of the main element sample include at least one of the following: occupying a relatively large proportion of the picture (greater than a size threshold, such as greater than 20%), or being located in the visual center area (such as the intersection of the rule of thirds), carrying key semantic information (such as a face, the main body of a commodity, pathological features), and corresponding to the core keywords in the text description in cross-modal association.
[0105] 9) A secondary element sample refers to the visual content in the original image sample that plays an auxiliary role in expressing the theme or has a logical association with the main element sample but is not necessarily present. The characteristics of the secondary element sample include at least one of the following: having a weaker color or light and shadow contrast than the main element (for example, at least 30% lower saturation or brightness), having a low information-bearing density (such as repetitive textures, background ornaments), and being weakened or implicitly mentioned in the text description (such as the roadside trees in a "street scene").
[0106] 10) The first element sample refers to the set of the main element sample and the secondary element sample identified in the original image sample. For example, if the cake is the main element sample and the cream is the secondary element sample, then both the cake and the cream are the first element samples.
[0107] 11) The second element sample is a set composed of similar elements and related elements of the first element sample, where the similar elements refer to the elements with a similarity greater than the similarity threshold to the first element sample, and the related elements refer to the elements with a co-occurrence frequency greater than the frequency threshold in image samples other than the original image sample.
[0108] 12) An anchor sample is a fixed reference sample selected in contrastive learning or related tasks. It is used to compare with other samples (positive samples and negative samples) to help the machine learning model learn the similarities and differences between samples.
[0109] In related technologies, when training a machine learning model, conventional text-image pairs are often used as the training set, resulting in low understanding ability of the machine learning model for unconventional data and low accuracy of the machine learning model.
[0110] Based on the above analysis, the applicant found that the data processing method of the machine learning model in related technologies cannot eliminate the bias against unconventional data. In view of the above problems, the embodiments of the present application provide a data processing method for a machine learning model, which can improve the accuracy of the machine learning model.
[0111] The following describes the exemplary applications of the electronic device provided by the embodiments of the present application. The electronic device provided by the embodiments of the present application can be implemented as various types of terminals such as a laptop computer, a tablet computer, a desktop computer, a set-top box, a smart phone, a smart speaker, a smart watch, a smart TV, a vehicle-mounted terminal, etc., or can be implemented as a server. Below, the exemplary application when the electronic device is implemented as a server will be described.
[0112] See Figure 1 , Figure 1 is a schematic diagram of the architecture of the data processing system 100 of the machine learning model provided by the embodiments of the present application. To implement a data processing application that supports a machine learning model, the terminal 400 is connected to the server 200 through the network 300. The network 300 can be a wide area network, a local area network, or a combination of the two.
[0113] The server 200 is used to train the machine learning model with the first generative text sample and the generative image sample to obtain a pre-trained machine learning model, call the pre-trained machine learning model based on the text sent by the terminal 400 to generate an image corresponding to the text, or identify an image corresponding to the text and send the image to the terminal 400 for display on the human-computer interaction interface 410.
[0114] The machine learning model trained with the first generative text sample and the generative image sample has the following performance:
[0115] Cross-modal retrieval ability: The pre-trained machine learning model can accurately identify an image including the elements described in the text according to the text.
[0116] Efficient feature extraction ability: The pre-trained machine learning model can efficiently extract key features from the image for identifying and classifying elements.
[0117] Good generalization ability: The pre-trained machine learning model learns features and patterns during training, enabling it to maintain a high recognition accuracy when faced with new and unseen images.
[0118] Fast inference speed: After being trained, the machine learning model can achieve fast inference on the terminal, providing real-time element recognition services.
[0119] Semantic understanding ability: The pre-trained machine learning model can understand the semantic information in the text description, including negative descriptions of high-frequency element subsets and positive descriptions of low-frequency element subsets, thus more accurately identifying elements in the image.
[0120] Taking the video clip recognition scenario as an example, in response to receiving a natural language query statement, the terminal 400 extracts keywords from the natural language query statement and sends them to the server 200. The server 200 calls the pre-trained machine learning model through the keywords, obtains multiple consecutive image frames corresponding to the keywords in the video, and sends the video clip composed of the image frames to the terminal 400 for display on the human-computer interaction interface 410.
[0121] Taking the intelligent search scenario as an example, the terminal 400 sends the search text to the server 200. The server 200 trains the machine learning model through the first generative text sample and generative image sample to obtain the pre-trained machine learning model, calls the pre-trained machine learning model based on the search text, obtains the image corresponding to the search text, and sends the image to the terminal 400 for display on the human-computer interaction interface 410.
[0122] Taking the shopping scenario as an example, the terminal 400 sends the description text of the commodity to the server 200. The server 200 trains the machine learning model through the first generative text sample and generative image sample to obtain the pre-trained machine learning model, calls the pre-trained machine learning model based on the description text of the commodity, obtains the commodity image corresponding to the description text, and sends the commodity image to the terminal 400 for display on the human-computer interaction interface 410.
[0123] Taking the educational courseware generation scenario as an example, the terminal 400 sends the courseware text to the server 200. The server 200 calls the pre-trained machine learning model based on the courseware text, obtains the image corresponding to the courseware text, combines the images into a file or presentation, and sends it to the terminal 400 for display on the human-computer interaction interface 410.
[0124] Taking the game development scenario as an example, the terminal 400 sends a description of the game screen or a description of the key elements in the game scene to the server 200. The server 200 is used to call a pre-trained machine learning model based on the description to obtain a screen corresponding to the description, or a screen including the key elements, and send the screen to the terminal 400 for display on the human-computer interaction interface 410.
[0125] In some embodiments, the server 200 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It may also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The terminal and the server may be directly or indirectly connected through wired or wireless communication methods, which are not limited in the embodiments of the present application.
[0126] See Figure 2A , Figure 2A is a schematic structural diagram of the server 200-1 provided in the embodiments of the present application. The server 200-1 is an implementation manner in which the above-mentioned server 200 is used to train a machine learning model through a first generative text sample and a generative image sample. Figure 2A The server 200-1 shown includes at least one processor 210, a memory 230, and at least one network interface 220. Each component in the server 200-1 is coupled together through a bus system 240. It can be understood that the bus system 240 is used to implement the connection and communication between these components. In addition to the data bus, the bus system 240 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 2A all kinds of buses are labeled as the bus system 240.
[0127] The processor 210 may be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a Digital Signal Processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or any conventional processor, etc.
[0128] The memory 230 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memories, hard disk drives, optical disc drives, etc. The memory 230 optionally includes one or more storage devices that are physically located far from the processor 210.
[0129] The memory 230 includes volatile memory, non-volatile memory, or both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 230 described in the embodiments of the present application is intended to include any suitable type of memory.
[0130] In some embodiments, the memory 230 is capable of storing data to support various operations. Examples of such data include programs, modules, and data structures, or subsets or supersets thereof, which are exemplarily described below.
[0131] The operating system 231 includes system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, the core library layer, the driver layer, etc., for implementing various basic services and processing hardware-based tasks;
[0132] The network communication module 232 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 220. Exemplary network interfaces 220 include: Bluetooth, Wi-Fi (Wireless Fidelity), and Universal Serial Bus (USB), etc.;
[0133] In some embodiments, the device provided in the embodiments of the present application may be implemented in software. Figure 2A The data processing device 233 storing the machine learning model in the memory 230 is shown. It may be software in the form of programs and plugins, etc., and includes the following software modules: the element recognition module 2331, the subset extraction module 2332, the subset recognition module 2333, the first generation module 2334, and the second generation module 2335. These modules are logical, so they can be arbitrarily combined or further split according to the functions to be implemented. The functions of each module will be described below.
[0134] See Figure 2B , Figure 2B which is a schematic structural diagram of the server 200-2 provided in the embodiments of the present application. The server 200-2 is an implementation manner of the above-mentioned server 200 for training a machine learning model by using a second generated text sample and a target original image sample. Figure 2BThe server 200-2 shown includes: at least one processor 250, a memory 270, and at least one network interface 260. Each component in the server 200-2 is coupled together through a bus system 280. It can be understood that the bus system 280 is used to implement connection communication between these components. In addition to a data bus, the bus system 280 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 2B all kinds of buses are labeled as the bus system 280.
[0135] The processor 250 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0136] The memory 270 can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disc drives, etc. Optionally, the memory 270 includes one or more storage devices that are physically remote from the processor 250.
[0137] The memory 270 includes volatile memory or non-volatile memory, and can also include both volatile and non-volatile memory. The non-volatile memory can be a read-only memory (ROM, Read Only Memory), and the volatile memory can be a random access memory (RAM, Random Access Memory). The memory 270 described in the embodiments of the present application is intended to include any suitable type of memory.
[0138] In some embodiments, the memory 270 is capable of storing data to support various operations. Examples of such data include programs, modules, and data structures, or subsets or supersets thereof, which are described below by way of example.
[0139] An operating system 271, including system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks;
[0140] A network communication module 272, for reaching other electronic devices via one or more (wired or wireless) network interfaces 260. Exemplary network interfaces 260 include: Bluetooth, wireless compatibility certification (WiFi), and universal serial bus (USB), etc.;
[0141] In some embodiments, the device provided by the embodiments of the present application may be implemented in software. Figure 2B The data processing device 273 storing the machine learning model in the memory 270 is shown, which may be software in the form of a program, a plug-in, etc., and includes the following software modules: a data acquisition module 2731, a sample recognition module 2732, a third generation module 2733, and a sample rewriting module 2734. These modules are logical, so they can be combined arbitrarily or further split according to the implemented functions. The functions of each module will be described below.
[0142] See Figure 2C , Figure 2C FIG. is a schematic structural diagram of the server 200-3 provided by the embodiments of the present application. The server 200-3 is an implementation manner of the above-mentioned server 200 for training a machine learning model through the first training set, the second training set, and the third training set. Figure 2C The server 200-3 shown includes: at least one processor 290, a memory 292, and at least one network interface 291. Each component in the server 200-3 is coupled together through a bus system 293. It can be understood that the bus system 293 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 293 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 2C all kinds of buses are labeled as the bus system 293.
[0143] The processor 290 may be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or any conventional processor, etc.
[0144] The memory 292 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memories, hard disk drives, optical disc drives, etc. The memory 292 optionally includes one or more storage devices that are physically located away from the processor 290.
[0145] The memory 292 includes volatile memory or non-volatile memory, and may also include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM, Read Only Memory), and the volatile memory may be a random access memory (RAM, Random Access Memory). The memory 292 described in the embodiments of the present application is intended to include any suitable type of memory.
[0146] In some embodiments, the memory 292 is capable of storing data to support various operations. Examples of such data include programs, modules, and data structures, or subsets or supersets thereof, which are illustrated below.
[0147] The operating system 2921 includes system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, the core library layer, the driver layer, etc., for implementing various basic services and processing hardware-based tasks;
[0148] The network communication module 2922 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 291. Exemplary network interfaces 291 include: Bluetooth, Wireless Fidelity (WiFi), and Universal Serial Bus (USB), etc.;
[0149] In some embodiments, the device provided by the embodiments of the present application can be implemented in software. Figure 2C The data processing device 2923 for the machine learning model stored in the memory 292 is shown. It can be software in the form of programs and plugins, etc., including the following software modules: the first rewriting module 29231, the fourth generating module 29232, and the second rewriting module 29233. These modules are logical, so they can be combined or further split arbitrarily according to the functions to be implemented. The functions of each module will be described below.
[0150] In some embodiments, the terminal or the server can implement the data processing method provided by the embodiments of the present application by running various computer-executable instructions or computer programs. For example, the computer-executable instructions can be microprogram-level commands, machine instructions, or software instructions. The computer program can be a native program or a software module in the operating system; it can be a local (Native) application (APPlication, APP), that is, a program that needs to be installed in the operating system to run, such as a video APP or a shopping APP; it can also be a small program that can be embedded in any APP, that is, a program that only needs to be downloaded to the browser environment to run. In short, the above computer-executable instructions can be instructions in any form, and the above computer programs can be application programs, modules, or plugins in any form.
[0151] The data processing method for the machine learning model provided by the embodiments of the present application will be described in combination with the exemplary applications and implementations of the server provided by the embodiments of the present application.
[0152] See Figure 3A , Figure 3A which is the first process schematic diagram of the data processing method for the machine learning model provided by the embodiments of the present application. Taking the server as the main body, it will be combined with Figure 3AThe steps shown will be described.
[0153] In step 101, target recognition is performed on multiple original image samples in the first training set, and the recognized elements are combined into an element set.
[0154] In some embodiments, target recognition can be performed on multiple original image samples in the first training set by any one of the following algorithms: You Only Look Once (YOLO), Fast Region-based Convolutional Neural Network (Fast R-CNN), Single Shot MultiBox Detector (SSD), Region Proposal Network (RPN).
[0155] Exemplarily, taking YOLO as an example, for each original image sample in the first training set, the following processing is performed: preprocess the original image sample such as scaling and normalizing; extract image features from the preprocessed original image sample through a Convolutional Neural Network (CNN) to obtain a feature map; divide the feature map into multiple grids; for each grid, predict multiple bounding boxes and the probabilities of the elements included in these bounding boxes to obtain the elements in the original image sample. Combine the elements recognized in multiple original image samples into an element set.
[0156] In step 102, multiple element subsets are extracted from the element set, and the occurrence times of the element subsets in multiple original image samples are determined.
[0157] In some embodiments, the number of elements included in each element subset is greater than or equal to 2, and different element subsets may include the same elements.
[0158] Exemplarily, if the element set is {cake, cream, bread, chocolate, football}, the element subsets can be {cake, cream}, {cake, football}, {cream, bread}, {bread, chocolate}, etc. Determining the occurrence times of the element subsets in multiple original image samples means determining the number of times the elements in each element subset co-occur in the same original image sample. For example, if {cake, cream} appears simultaneously in original image sample A, original image sample B, and original image sample C, then the occurrence times of {cake, cream} is 3. If {cake, football} does not appear simultaneously in any original image sample, then the occurrence times of {cake, football} is 0.
[0159] In step 103, at least one of a high-frequency element subset and a low-frequency element subset is identified from multiple element subsets based on the occurrence times.
[0160] Here, based on the occurrence times, at least one high-frequency element subset can be identified only from multiple element subsets, or at least one low-frequency element subset can be identified only from multiple element subsets, or at least one high-frequency element subset and at least one low-frequency element subset can be identified from multiple element subsets.
[0161] In some embodiments, referring to Figure 3B , Figure 3B is the second process schematic diagram of the data processing method of the machine learning model provided by the embodiments of the present application. Figure 3A Step 103 of Figure 3B “Based on the occurrence times, at least one of a high-frequency element subset and a low-frequency element subset is identified from multiple element subsets” can be implemented through
[0162] Steps 1031 to 1033 of
[0163] In step 1031, multiple element subsets are sorted in descending order based on the occurrence times to obtain a sorting result.
[0164] Continuing with the example of the above step 102, if the occurrence times of the element subset {cake, cream} is 3, the occurrence times of {cake, football} is 0, the occurrence times of {cream, bread} is 2, and the occurrence times of {bread, chocolate} is 1, then multiple element subsets are sorted in descending order based on the occurrence times, and the obtained sorting result is: {cake, cream}, {cream, bread}, {bread, chocolate}, {cake, football}.
[0165] In some embodiments, the preset ratio can be set according to business requirements, and a high-frequency element subset includes at least two high-frequency elements.
[0166] Continuing with the example of the above step 1031, if the preset ratio is 50%, then {cake, cream} and {cream, bread} are used as high-frequency element subsets.
[0167] In step 1033, the element subsets other than the high-frequency element subsets among the multiple element subsets are used as low-frequency element subsets.
[0168] Continuing with the example in step 1032 above, among the multiple element subsets, the element subsets other than the high-frequency element subsets {cake, cream} and {cream, bread} are {bread, chocolate} and {cake, football}. Take {bread, chocolate} and {cake, football} as low-frequency element subsets, where one low-frequency element subset includes at least two low-frequency elements.
[0169] In the embodiments of this application, by classifying the element subsets into two categories: high-frequency element subsets and low-frequency element subsets, the data distribution can be better balanced. Different strategies can be adopted for high-frequency element subsets and low-frequency element subsets, so as to more comprehensively consider the diversity of data. Taking the element subsets other than the high-frequency element subsets as low-frequency element subsets can make the model focus on those less common element combinations, which helps to discover some rare patterns and features, and improve the generalization ability of the model and the ability to handle abnormal situations.
[0170] In some embodiments, step 103 "Based on the occurrence times, identify at least one of the high-frequency element subsets and low-frequency element subsets from multiple element subsets" can also be implemented by performing the following processing: Take the element subset with the occurrence times greater than the first occurrence threshold as the high-frequency element subset; Take the element subset with the occurrence times less than the second occurrence threshold as the low-frequency element subset, where the first occurrence threshold is greater than or equal to the second occurrence threshold.
[0171] Continuing with the example in step 1031 above, taking the first occurrence threshold greater than the second occurrence threshold as an example, if the first occurrence threshold is 2 and the second occurrence threshold is 1, and the element subset with the occurrence times greater than the first occurrence threshold is {cake, cream}, then take {cake, cream} as the high-frequency element subset, and the element subset with the occurrence times less than the second occurrence threshold is {cake, football}, then take {cake, football} as the low-frequency element subset.
[0172] In the embodiments of this application, by setting the first occurrence threshold and the second occurrence threshold, the element subsets can be clearly divided into two categories: high-frequency element subsets and low-frequency element subsets, which helps to perform targeted processing for different types of element subsets subsequently. The values of the occurrence thresholds can be flexibly adjusted according to the specific data set and task requirements to better distinguish high-frequency element subsets and low-frequency element subsets, and the threshold setting has a certain interpretability.
[0173] Continue to refer to Figure 3A , in step 104, text generation is performed based on at least one of the high-frequency element subsets and low-frequency element subsets to obtain the first generative text sample, where the first generative text sample includes at least one of the negative description of the high-frequency element subset and the positive description of the low-frequency element subset.
[0174] Here, text generation can be performed based only on the high-frequency element subset to obtain a first generated text sample, where the first generated text sample includes only negative descriptions of the high-frequency element subset; text generation can be performed based only on the low-frequency element subset to obtain a first generated text sample, where the first generated text sample includes only positive descriptions of the low-frequency element subset; text generation can also be performed based on the high-frequency element subset and the low-frequency element subset to obtain a first generated text sample, where the first generated text sample includes negative descriptions of the high-frequency element subset and positive descriptions of the low-frequency element subset.
[0175] In some embodiments, referring to Figure 3C , Figure 3C is the third process schematic diagram of the data processing method of the machine learning model provided by the embodiments of the present application. Figure 3A Step 104 “Generate a first generated text sample based on at least one of the high-frequency element subset and the low-frequency element subset” of Figure 3C can be implemented through steps 1041 to 1043 of
[0176] In step 1041, for at least some high-frequency elements in the high-frequency element subset, generate negative descriptions of at least some high-frequency elements, where the negative descriptions of at least some high-frequency elements indicate that the generated image sample does not include at least some high-frequency elements.
[0177] In some embodiments, select at least some high-frequency elements from the high-frequency element subset, construct a prompt based on at least some high-frequency elements, where the prompt is used to indicate generating negative descriptions of at least some high-frequency elements; call a pre-trained language understanding model based on the prompt to generate negative descriptions of at least some high-frequency elements.
[0178] For example, if the high-frequency element subset is {cake, cream}, if all high-frequency elements in the high-frequency element subset are selected, that is, cake and cream, the generated negative description of the high-frequency elements can be “This picture is not a cream cake”; if some high-frequency elements in the high-frequency element subset are selected, the generated negative description of some high-frequency elements can be “This is a cake without cream”.
[0179] In step 1042, for each low-frequency element in the low-frequency element subset, generate a positive description of the low-frequency element, where the positive description indicates that the generated image sample includes the low-frequency element.
[0180] In some embodiments, construct a prompt based on all low-frequency elements in the low-frequency element subset, where the prompt is used to indicate generating positive descriptions of the low-frequency elements; call a pre-trained language understanding model based on the prompt to generate positive descriptions of the low-frequency elements.
[0181] Exemplarily, the low-frequency element subset is {cake, football}, the low-frequency elements include a cake and a football, and the positive description of the low-frequency elements can be "This is a football cake".
[0182] In step 1043, based on at least one of the negative descriptions of at least some of the high-frequency elements and the positive descriptions of the low-frequency elements, a first generative text sample is generated.
[0183] In some embodiments, a first generative text sample can be generated only based on the negative descriptions of at least some of the high-frequency elements; a first generative text sample can be generated only based on the positive descriptions of the low-frequency elements; or a first generative text sample can be generated based on the negative descriptions of at least some of the high-frequency elements and the positive descriptions of the low-frequency elements.
[0184] Continuing with the examples of steps 1041 and 1042 above, taking the generation of a first generative text sample based on the negative descriptions of at least some of the high-frequency elements and the positive descriptions of the low-frequency elements as an example, the negative description of some of the high-frequency elements can be "This is a cake without cream", and the positive description of the low-frequency elements can be "This is a football cake", then the first generative text sample can be "This is a football cake without cream".
[0185] In the embodiments of the present application, more diverse text content is generated through the negative descriptions of high-frequency elements and the positive descriptions of low-frequency elements, increasing the diversity of the training data, enabling the model to learn a wider range of semantic representations and conceptual relationships. The positive description of low-frequency elements highlights the elements that are less common in the original data, enabling the model to pay more attention to these rare information, improving the understanding and processing ability of uncommon situations, thereby enhancing the generalization ability of the model and improving the accuracy and diversity of the generated images. Combining the negative descriptions of high-frequency elements and the positive descriptions of low-frequency elements can, to a certain extent, balance the occurrence frequencies of different elements in the data, avoiding the model being overly biased towards common elements while ignoring other important but less frequently occurring elements.
[0186] In some embodiments, in the case of at least some high-frequency elements that are the first part of the high-frequency element subset, when performing step 1041 "generate negative descriptions of at least some high-frequency elements", the following processing is performed: generate positive descriptions of the second part of the high-frequency elements in the high-frequency element subset, where the second part of the high-frequency elements is different from the first part of the high-frequency elements. Step 1043 "generate negative descriptions of at least some high-frequency elements for at least some high-frequency elements in the high-frequency element subset" can be implemented by performing the following processing: based on the negative descriptions corresponding to the first part of the high-frequency elements, the positive descriptions corresponding to the second part of the high-frequency elements, and the positive descriptions of the low-frequency elements, generate a first generative text sample.
[0187] Continuing with the examples in steps 1041 and 1043 above, for the high-frequency element subset {cake, cream}, if the high-frequency element in the first part is "cream", then the high-frequency element in the second part is "cake". The negative description of the high-frequency element in the first part is "without cream", and the positive description of the high-frequency element in the second part is "this is a cake". The positive descriptions of the low-frequency elements cake and football in the low-frequency element subset {cake, football} are "this is a football cake". Then, based on the negative description corresponding to the high-frequency element in the first part, the positive description corresponding to the high-frequency element in the second part, and the positive description of the low-frequency elements, the first generated text sample generated is "this is a football cake, without cream".
[0188] In the embodiments of the present application, the first generated text sample includes both the negative description of the high-frequency element in the first part of the high-frequency element subset and the positive description of the high-frequency element in the second part as well as the positive description of the low-frequency elements, which can more comprehensively express the element information in the image and avoid information loss caused by only focusing on a single type of description. By simultaneously generating the negative description of the high-frequency element in the first part and the positive description of the high-frequency element in the second part, the description of the high-frequency element is made more balanced, which helps the model better understand the characteristics and meanings of the high-frequency elements in different situations and improves the model's understanding and processing ability of the high-frequency elements. The combination of the positive and negative descriptions strengthens the effect of contrast learning, and the model can better learn the differences between the existence and non-existence of elements, thereby better understanding the semantic content of the image.
[0189] Continue to refer to Figure 3A , in step 105, image generation is performed based on the first generated text sample to obtain a generated image sample, where the first generated text sample and the generated image sample are used to train a machine learning model.
[0190] In some embodiments, the first training set further includes a plurality of original text samples, and one original text sample is used to describe one original image sample. Refer to Figure 3D , Figure 3D is the fourth process schematic diagram of the data processing method of the machine learning model provided by the embodiments of the present application. Before training the machine learning model, for the original image sample described by each original text sample, perform Figure 3D steps 201 to 203 as follows for specific description.
[0191] In step 201, a first element sample is identified from the original image sample.
[0192] In some embodiments, the first element sample includes a main element sample and a secondary element sample. Refer to Figure 3E , Figure 3E is the fifth process schematic diagram of the data processing method of the machine learning model provided by the embodiments of the present application.Figure 3D Step 201, "Identify the first element sample from the original image sample" of Figure 3E can be implemented through steps 2011 to 2013 of
[0193] and the following is a specific description.
[0194] In step 2011, feature recognition is performed on the original image sample to obtain multiple candidate element samples.
[0195] In some embodiments, step 2011, "Perform feature recognition on the original image sample to obtain multiple candidate element samples", can be implemented by performing the following processing: determine the features to be recognized and convert them into specific descriptive prompt words; vectorize the original image sample to obtain the original image vector; based on the original image vector and the prompt words, call the pre-trained image recognition model to obtain multiple candidate element samples corresponding to the prompt words.
[0196] For example, if the original image sample is a birthday cake image, the prompt words may include "Find the cake part in the image", "Identify the cream part in the image", "Determine the background color in the image", etc. According to the prompt word "Find the cake part in the image", the pre-trained image recognition model will analyze the parts in the image related to the features such as the shape and color of the cake and determine the cake as a candidate element sample.
[0197] Here, step 2011, "Perform feature recognition on the original image sample to obtain multiple candidate element samples", can also be implemented by performing the following processing: obtain multiple training image samples, label the elements in these training image samples, label elements such as cakes, cream, backgrounds, etc., and form a training set with the labeled training image samples and the corresponding element labels; determine the features to be learned by the image recognition model, and these features can be shape, color, texture, etc.; based on the training set, train the initialized image recognition model to obtain the pre-trained image recognition model, where, during the training process, the initialized image recognition model is used to learn the relationship between the features in the training image samples and the element labels; based on the original image sample, call the pre-trained image recognition model to obtain multiple candidate element samples.
[0198] In some embodiments, candidate features (such as color, shape, texture, brightness) are extracted from candidate element samples, and the candidate features are vectorized to obtain a first vector; element features are extracted from other element samples, and the element features are vectorized to obtain a second vector. The similarity between the candidate features and the element features is determined by any one of the following algorithms: cosine similarity, Euclidean distance, Manhattan distance, and Pearson correlation coefficient. Taking cosine similarity as an example, the dot product of the first vector and the second vector is determined; the product of the lengths of the first vector and the second vector is determined; the ratio of the dot product and the length product is taken as the similarity.
[0199] For example, the distance threshold is 5. If the position of candidate element sample A is (x1, y1) and the center of the original image sample is (x2, y2), then the distance between the position of candidate element sample A and the center of the original image sample is If the distance between the position of candidate element sample A and the center of the original image sample is 4, then candidate element sample A is a main element sample; if the size threshold is 1 / 3, the size of candidate element sample B is 3 cm * 4 cm, and the size of the original image sample is 3 cm * 8 cm, and the ratio of the size of candidate element sample B to the size of the original image sample is 1 / 2, then candidate element sample B is a main element sample; if the similarity threshold is 0.7, the first vector corresponding to candidate element sample C is m and the second vector is n, then the dot product is m * n, the length of the first vector is |m|, the length of the second vector is |n|, then the length product is |m| * |n|, and the ratio of the dot product and the length product is (m * n) / (|m| * |n|) as the similarity. If the similarity between the first vector and the second vector is 0.8, then candidate element sample C is taken as a main element sample.
[0200] In step 2013, the candidate element samples other than the main element samples among the multiple candidate element samples are used as secondary element samples.
[0201] For example, the multiple candidate element samples are candidate element sample A, candidate element sample B, candidate element sample C, candidate element sample D, and candidate element sample E. If candidate element sample A, candidate element sample B, and candidate element sample C are main element samples, then candidate element sample D and candidate element sample E are used as secondary element samples.
[0202] By setting preset conditions such as a distance threshold, a size threshold, and a similarity threshold, the embodiments of the present application can determine candidate element samples located near the center of the image, having a larger size, or being significantly different from other elements as main element samples, which can highlight the key content in the original image sample, help to pay more attention to these important elements in subsequent processing and analysis, and can avoid misjudging candidate element samples that are too similar to other elements as main element samples, contributing to improving the robustness of the model so that it can more accurately identify and distinguish main elements and secondary elements in the face of various complex image scenes. By adjusting parameters such as the distance threshold, the size threshold, and the similarity threshold, the main element samples and secondary element samples can be flexibly determined according to different application scenarios and requirements.
[0203] Continue to refer to Figure 3D , in step 202, a second element sample that does not appear in the original image sample is generated based on the first element sample.
[0204] In some embodiments, refer to Figure 3F , Figure 3F is the sixth process schematic diagram of the data processing method of the machine learning model provided by the embodiments of the present application. Figure 3D Step 202 of Figure 3F “generating a second element sample that does not appear in the original image sample based on the first element sample” can be implemented through Steps 2021 to 2022 of
[0205] In step 2021, at least one of a similar element of the first element sample and a related element of the first element sample is generated, where the similar element is an element whose similarity to the first element sample is greater than the similarity threshold, and the related element is an element whose co-occurrence times with the first element sample in an image sample other than the original image sample is greater than the times threshold.
[0206] In some embodiments, a third feature of the first element sample is extracted, a fourth feature of other elements is extracted, and the similarity between the third feature and the fourth feature is determined as the similarity between the first element sample and the other elements. The other elements are elements that do not appear in the original image sample, and the similarity can be determined in the following ways: cosine similarity, Euclidean distance, Manhattan distance, and Pearson correlation coefficient.
[0207] Exemplarily, taking the Euclidean distance as an example, if the third feature is (x3, y3) and the fourth feature is (x4, y4), the distance between the third feature and the fourth feature is determined, that is The similarity between the first element sample and other elements. If the similarity threshold is 0.7, the similarity between other element A and the first element sample is 0.6, and the similarity between other element B and the first element sample is 0.8, then other element A is taken as the similar element of the first element sample. If the frequency threshold is 3, the co-occurrence frequency of other element C and the first element sample in image samples other than the original image sample is 5, and the co-occurrence frequency of other element D and the first element sample in image samples other than the original image sample is 1, then other element D is taken as the relevant element of the first element sample.
[0208] In step 2022, the similar elements and relevant elements are used as the second element sample.
[0209] Continuing with the example in step 2021 above, the similar element (other element A) and the relevant element (other element D) are used as the second element sample.
[0210] In the embodiments of the present application, by generating the similar elements and relevant elements of the first element sample to be used as the second element sample, it can provide richer information for subsequent analysis and processing. The similar elements are screened based on the similarity threshold, and the relevant elements are determined based on the co-occurrence frequency threshold, expanding the range of elements from different perspectives. Introducing similar elements and relevant elements allows the model to learn more information related to the original element sample, thereby improving the generalization ability of the model. The model can better understand the relationships between elements, not only limited to the content in the original image sample, but also extended to the information in other relevant image samples.
[0211] In some embodiments, step 202 "generating a second element sample that does not appear in the original image sample based on the first element sample" can also be implemented by performing the following processing: constructing a prompt word based on the first element sample, where the prompt word is used to indicate generating a second element sample that does not appear in the original image sample; and calling a pre-trained image recognition model based on the prompt word to generate a second element sample that does not appear in the original image sample.
[0212] For example, if the first element sample is a cake and cream, the prompt word can be "The first element sample included in this original image sample is a cake and cream. List the elements that do not appear in this image and are similar to or relevant to the first element sample as the second element sample."
[0213] In the embodiments of the present application, by constructing a prompt based on the first element sample, it is possible to clearly instruct the model to generate an element that is related to the first element sample but does not appear in the original image sample, making the generated second element sample more targeted and meeting specific requirements. By invoking the pre-trained image recognition model, the knowledge and features learned by the model from large-scale data are fully utilized. These pre-trained models usually have strong image understanding and generation capabilities, and can better generate reasonable second element samples according to the prompt.
[0214] Continue to refer to Figure 3D , in step 203, the original text sample is rewritten based on the second element sample to obtain a second generative text sample for negatively describing the second element sample, where the second generative text sample and the original image sample are used to train the machine learning model.
[0215] Here, the machine learning model can be used to generate images according to text, and can also be used to identify the image corresponding to the text from multiple images according to the text.
[0216] In some embodiments, the original text sample includes a positive description of the first element sample, refer to Figure 3G , Figure 3G is the seventh process schematic diagram of the data processing method of the machine learning model provided by the embodiments of the present application. Figure 3D Step 203 of Figure 3G "rewrite the original text sample based on the second element sample to obtain a second generative text sample for negatively describing the second element sample" can be implemented through
[0217] In step 2031, generate a negative description of the second element sample.
[0218] In some embodiments, a prompt is constructed based on the second element sample, where the prompt is used to indicate the generation of a negative description of the second element sample; the pre-trained language understanding model is invoked based on the prompt to obtain a negative description of the second element sample.
[0219] For example, the first element sample includes a cake and cream, the similar elements of the cake are bread, the similar elements of the cream are jam, the related elements of the cake are chocolate, and the related elements of the cream are candles, that is, the second element sample includes bread, jam, chocolate and candles. The prompt can be "Please generate a negative description of bread, jam, chocolate and candles", and the obtained negative description of the second element sample can be "This is not bread, there is no jam and chocolate on it, and there are no candles inserted".
[0220] In step 2032, the negative description of the second element sample and the positive description of the first element sample are combined to obtain a second generative text sample.
[0221] Continuing with the example in step 2031 above, the positive description of the first element sample is "This is a cream cake". Combining the negative description of the second element sample and the positive description of the first element sample, the obtained second generative text sample is "This is a cream cake, this is not bread, there is no jam or chocolate on it, and there are no candles inserted".
[0222] In the embodiment of the present application, the negative description of the second element sample and the positive description of the first element sample are combined into a second generative text sample, which can provide richer semantic information. The combination of the positive and negative descriptions forms a distinct contrast, which helps to highlight the differences and characteristics between the first element sample and the second element sample. This contrast effect can make it easier for the machine learning model to understand and distinguish different elements, improving the ability to understand and analyze the image content.
[0223] In some embodiments, refer to Figure 3H , Figure 3H is the eighth process schematic diagram of the data processing method of the machine learning model provided by the embodiment of the present application. Before step 203 "Rewriting the original text sample based on the second element sample to obtain a second generative text sample for negatively describing the second element sample", execute Figure 3H steps 204 to 207, which are specifically described below.
[0224] In step 204, a second training set is constructed based on the first generative text sample and the generative image sample, and a third training set is constructed based on the second generative text sample and the original image sample.
[0225] In some embodiments, for each first generative text sample, a first text-image sample pair is constructed by combining the first generative text sample and the generative image sample described by the first generative text sample; multiple first text-image sample pairs are combined into a second training set. For each second generative text sample, a second text-image sample pair is constructed by combining the second generative text sample and the original image sample; multiple second text-image sample pairs are combined into a third training set.
[0226] Exemplarily, the first generative text sample A is "This is a football cake without cream", and the first generative text sample A is used to describe the generative image sample A. The first generative text sample A and the generative image sample A are combined into the first text-image sample pair A. The first generative text sample B is "There is a football and a badminton racket in the corner without a badminton", and the first generative text sample B is used to describe the generative image sample B. The first generative text sample B and the generative image sample B are combined into the first text-image sample pair B. The first text-image sample pair A and the first text-image sample pair B are combined into the second training set. The second generative text sample A is "This is a cream cake rather than bread", and the second generative text sample B is "This is a cream cake without a candle on it". The second generative text sample A and the original image sample are combined into the second text-image sample pair A, and the second generative text sample B and the original image sample are combined into the second text-image sample pair B. The second text-image sample pair A and the second text-image sample pair B are combined into the third training set.
[0227] In step 205, multiple text-image sample pairs are collected from the first training set, the second training set, and the third training set.
[0228] In some embodiments, multiple text-image sample pairs are collected for each training set, and each text-image sample pair includes an image sample and a text sample for describing the image sample.
[0229] Continuing with the example in step 204 above, the text-image sample pair in the first training set can be <cream cake image, "This is a cream cake">, the text-image sample pair in the second training set can be <football cake image, "This is a football cake without cream">, and the text-image sample pair in the third training set can be <cream cake image, "This is a cream cake rather than bread">.
[0230] In step 206, the text-image matching degree of each text-image sample pair is determined. The text-image sample pairs with a text-image matching degree greater than the matching degree threshold are used as positive samples, and the text-image sample pairs with a text-image matching degree less than the matching degree threshold are used as negative samples.
[0231] In some embodiments, for each text-image sample pair, the image features of the image sample in the text-image sample pair are extracted, the text features of the text sample in the text-image sample pair are extracted, and the similarity between the image features and the text features is determined as the text-image matching degree.
[0232] Exemplarily, image features can be extracted by any of the following methods: CNN, Scale-Invariant Feature Transform (SIFT), Histogram of Oriented Gradients (HOG). Text features can be extracted by any of the following methods: Bag of Words (BoW), Term Frequency-Inverse Document Frequency (TF-IDF), Word Embedding, Long Short-Term Memory (LSTM). The similarity can be determined by any of the following methods: cosine similarity, Euclidean distance, Manhattan distance, and Pearson correlation coefficient. If the matching degree threshold is 0.8, the image-text matching degree of image-text sample pair A is 0.9, and the image-text matching degree of image-text sample pair B is 0.3, then image-text sample pair A is taken as the positive sample, and image-text sample pair B is taken as the negative sample.
[0233] In step 207, based on a preset loss function, positive samples, and negative samples, the loss is determined, and the initialized machine learning model is updated based on the loss to obtain a pre-trained machine learning model.
[0234] In some embodiments, the loss is backpropagated to update the parameters of the initialized machine learning model. The process of calculating the loss and updating the parameters is iterated multiple times until the global loss converges, and the iteration process is stopped to form a pre-trained machine learning model.
[0235] Exemplarily, refer to Figure 4 , Figure 4 is a schematic diagram of the principle of training a machine learning model provided by an embodiment of the present application. In Figure 4 , based on the positive and negative samples collected from the first training set, the second training set, and the third training set, and a preset loss function, the parameters of the initialized machine learning model are updated to obtain a pre-trained machine learning model. Backpropagation is implemented by the backpropagation algorithm. The gradient of each neuron is calculated from the output layer to the input layer, and the weights and biases of the neurons are updated according to the gradient. The parameters are continuously updated in a gradient descent manner to reduce the loss value. Gradient descent can adopt various gradient descent algorithms, such as batch gradient descent algorithm, stochastic gradient descent algorithm, adaptive gradient descent algorithm, and momentum gradient descent algorithm.
[0236] In the embodiments of the present application, by constructing the second training set and the third training set, the diversity and richness of the training data are increased, different types of samples are provided for the model, which helps the model learn a wider range of features and patterns. Using multiple training sets and combining the text-image matching degree to determine positive and negative samples for training can enable the model to better understand the relationship between images and texts and improve the generalization ability of the model.
[0237] In some embodiments, referring to Figure 3I , Figure 3I FIG. is the ninth process schematic diagram of the data processing method of the machine learning model provided by the embodiments of the present application. Before step 207 "determine the loss based on a preset loss function, positive samples, and negative samples", execute Figure 3I steps 301 to 305 of
[0238] In step 301, for each text-image sample pair, perform a first encoding on the image sample in the text-image sample pair to obtain an image vector, and perform a second encoding on the text sample in the text-image sample pair to obtain a text vector.
[0239] In some embodiments, the image sample in the text-image sample pair can be first encoded by a CNN to obtain an image vector. The text sample can be second encoded by a bag-of-words model or word vector encoding to obtain a text vector.
[0240] Exemplarily, the convolutional kernel in the convolutional layer of the CNN slides on the image sample to extract local features; the pooling layer is used to extract features from the image sample to obtain main features; the fully connected layer maps the local features and the main features to an image vector. Taking word vector encoding as an example, the text sample is tokenized to obtain multiple words. For each word, a pre-trained word vector model is called for encoding, such as a word vector model (Word to Vector, Word2Vec) or a global word frequency statistics-based word embedding method (Global Vectors for word representation, GloVe), to obtain the word vector corresponding to each word, and the word vectors of all words are averaged or summed to obtain a text vector.
[0241] In step 302, determine the first similarity between the image vector and the text vector, and construct a first loss function based on the similarity.
[0242] In some embodiments, the first similarity can be determined by any one of the following algorithms: cosine similarity, Euclidean distance, Manhattan distance, and Pearson correlation coefficient. For positive samples, the first log-likelihood function is determined based on cosine similarity; for negative samples, the second log-likelihood function is determined based on cosine similarity; the sum of the first log-likelihood function and the second log-likelihood function is used as the first loss function.
[0243] Exemplarily, the first log-likelihood function is shown in the following formula (1):
[0244]
[0245] Wherein, represents the first log-likelihood function, N represents the number of image-text sample pairs, s i,i represents the cosine similarity between the i-th image vector and its corresponding i-th text vector, τ is a temperature parameter used to adjust the distribution of similarities, log is the logarithmic function, represents normalizing the similarities between all text vectors and the current image vector, ensuring that the cosine similarity of positive sample pairs (i.e., matching image-text sample pairs) is as large as possible. Similarly, for negative samples, the second log-likelihood function is determined based on cosine similarity, as shown in the following formula (2):
[0246]
[0247] Wherein, represents the second log-likelihood function, represents normalizing the similarities between all image vectors and the current text vector, ensuring that the cosine similarity of negative sample pairs (i.e., non-matching image-text sample pairs) is as small as possible, and the meanings of the remaining letters are the same as those in formula (1).
[0248] The sum of the first log-likelihood function and the second log-likelihood function is used as the first loss function, as shown in the following formula (3):
[0249]
[0250] Wherein, L cLIP represents the first loss function, represents the first log-likelihood function, represents the second log-likelihood function.
[0251] In step 303, determine the second similarity between the preset anchor sample and the positive sample, and determine the third similarity between the anchor sample and the negative sample.
[0252] In some embodiments, extract the anchor features of the anchor samples, extract the positive features of the positive samples, and extract the negative sample features, determine the second similarity between the anchor features and the positive features, and determine the third similarity between the anchor features and the negative features.
[0253] Exemplarily, both the second similarity and the third similarity can be implemented by any one of the following algorithms: cosine similarity, Euclidean distance, Manhattan distance, and Pearson correlation coefficient.
[0254] In step 304, construct a second loss function based on the second similarity and the third similarity.
[0255] In some embodiments, for each anchor sample, determine the difference between the second similarity and the third similarity, determine the sum of the difference and a preset value, and determine the maximum value between the sum and zero; sum up the maximum values corresponding to multiple anchor samples respectively as the second loss function.
[0256] Exemplarily, the second loss function can refer to the following formula (4):
[0257]
[0258] where s a 、s + and s- represent the anchor sample, the positive sample, and the negative sample respectively, L triplet is the second loss function, θ is a mapping function for mapping a sample to a vector, max is used to determine the maximum value among multiple values, ‖·‖2 represents the L2 norm for calculating the distance between two feature vectors, that is, the similarity, and ε is a positive constant for controlling the interval between positive and negative samples to avoid trivial solutions (i.e., the features of all samples are mapped to the same point).
[0259] In step 305, perform a weighted sum of the first loss function and the second loss function to obtain a preset loss function.
[0260] In some embodiments, the weights can be adjusted according to specific requirements and experimental results, or can be dynamically adjusted according to the characteristics of the data and the performance of the model.
[0261] Continuing with the examples in steps 302 and 304 above, taking the adjustment of weights according to requirements as an example, the sum of the first weight of the first loss function and the second weight of the second loss function is 1. Construct multiple combinations of the first weight and the second weight, determine the performance of the machine learning model under multiple combinations, determine the first weight and the second weight corresponding to the optimal performance, determine the first product of the first loss function and the first weight, determine the second product of the second loss function and the second weight, and determine the sum of the first product and the second product as the preset loss function. Taking the dynamic adjustment of weights as an example, if the machine learning model performs poorly in the alignment of images and texts, the first weight of the first loss function can be increased; if the model performs poorly in differentiating positive samples and negative samples, the second weight of the second loss function can be increased.
[0262] In the embodiments of this application, by respectively determining the first similarity between the image vector and the text vector, the second similarity between the anchor sample and the positive sample, and the third similarity between the anchor sample and the negative sample, the similarity of the data is considered from multiple dimensions, the relationships in the data are comprehensively captured, and the learning effect of the model is improved. Based on different similarities, the first loss function and the second loss function are constructed. The first loss function focuses on the matching degree between the image and the text in the image-text sample pair, and the second loss function focuses on the relationships between the anchor sample and the positive sample and the negative sample, enabling the model to be optimized from different aspects, avoiding the influence of a single factor on model training, and improving the performance of the model.
[0263] In some embodiments, after step 105, "generating an image based on the first generated text sample to obtain a generated image sample", the following processing is performed: filtering out the generated image samples that do not meet the preset screening conditions from multiple generated image samples, where the screening conditions include: any high-frequency element in the high-frequency element subset does not appear in the generated image sample, and the generated image sample includes at least one low-frequency element in the low-frequency element subset.
[0264] Continuing with the example in step 1033 above, the high-frequency elements in the high-frequency element subset {cream, bread} include cream and bread, and the low-frequency elements in the low-frequency element subset {cake, football} include cake and football. If the generated image sample is a football cake and there is no cream on the cake and no bread in the image, the generated image sample meets the preset screening conditions; if the generated image sample is a cream cake, or if the football does not appear in the generated image sample, the generated image sample does not meet the preset screening conditions, and the generated image sample is filtered.
[0265] See Figure 3J , Figure 3J is the tenth process schematic diagram of the data processing method of the machine learning model provided by the embodiments of this application. Taking the server as the main body, it will be combined with Figure 3JThe steps shown will be described.
[0266] In step 401, a first training set is obtained, where the first training set includes a plurality of original text samples and a plurality of original image samples, and one original text sample is used to describe one original image sample.
[0267] Exemplarily, the original text sample A "This is a cream cake" is used to describe the original image sample A; the original text sample B "There is a football on this lawn" is used to describe the original image sample B.
[0268] Here, for the original image sample described by each original text sample, the following steps 402 to 404 are performed.
[0269] In step 402, a first element sample is identified from the original image sample.
[0270] In some embodiments, the first element sample includes a main element sample and a secondary element sample. The main element sample includes at least one of the following: the distance between the position of the main element sample and the center of the original image sample is less than a distance threshold, the ratio of the size of the main element sample to the size of the original image sample is greater than a size threshold, and the similarity between the candidate features of the main element sample and the element features of other element samples in the original image sample except the main element sample is less than a similarity threshold. The element samples in the first element sample other than the main element sample are used as secondary element samples. For details not described, reference may be made to the embodiments of step 201 above.
[0271] Continuing with the example of step 401 above, for the original text sample A, a plurality of element samples are identified from the original image sample A, such as {cream, cake, colored sugar crumbs}, where the cake is the main element sample, and the cream and colored sugar crumbs are secondary element samples.
[0272] In step 403, a second element sample that does not appear in the original image sample is generated based on the first element sample.
[0273] In some embodiments, at least one of a similar element of the first element sample and a related element of the first element sample is generated, where the similar element is an element with a similarity greater than the similarity threshold to the first element sample, and the related element is an element with a co-occurrence frequency greater than a frequency threshold in image samples other than the original image sample; the similar element and the related element are used as the second element sample. For details not described, reference may be made to the embodiments of step 202 above.
[0274] For example, if the similarity threshold is 0.7, the similarity between other element A and the first element sample is 0.6, and the similarity between other element B and the first element sample is 0.8, then other element A is taken as the similar element of the first element sample. If the occurrence threshold is 3, the co-occurrence times of other element C and the first element sample in image samples other than the original image sample is 5, and the co-occurrence times of other element D and the first element sample in image samples other than the original image sample is 1, then other element D is taken as the related element of the first element sample. The similar element (other element A) and the related element (other element D) are taken as the second element sample.
[0275] In step 404, the original text sample is rewritten based on the second element sample to obtain a second generative text sample for negatively describing the second element sample, where the second generative text sample and the original image sample are used to train a machine learning model.
[0276] In some embodiments, a negative description of the second element sample is generated, where the original text sample further includes a positive description of the first element sample. The negative description of the second element sample and the positive description of the first element sample are combined to obtain a second generative text sample. For the details not described, reference may be made to the embodiments of step 203 above.
[0277] For example, the first element sample includes cake and cream. The similar element of the cake is bread, the similar element of the cream is jam, the related element of the cake is chocolate, and the related element of the cream is candle. That is, the second element sample includes bread, jam, chocolate, and candle. The negative description of the second element sample can be "This is not bread, there is no jam and chocolate on it, and there is no candle inserted". If the positive description of the first element sample is "This is a cream cake", then the second generative text sample is "This is a cream cake, this is not bread, there is no jam and chocolate on it, and there is no candle inserted".
[0278] In the embodiment of the present application, the negative description of the second element sample and the positive description of the first element sample are combined into a second generative text sample, which can provide richer semantic information. The combination of the positive and negative descriptions forms a sharp contrast, which helps to highlight the differences and characteristics between the first element sample and the second element sample. This contrast effect can make the machine learning model easier to understand and distinguish different elements, and improve the understanding and analysis ability of the image content.
[0279] See Figure 3K , Figure 3K is the eleventh process schematic diagram of the data processing method of the machine learning model provided by the embodiment of the present application. Taking the server as the main body, it will be described in combination with Figure 3K the steps shown.
[0280] In step 501, a first rewriting is performed on the original text samples in the first training set to obtain first generative text samples, where the original text samples are used to describe the first element samples included in the original image samples in the first training set, and the first generative text samples are used to describe the high-frequency elements and low-frequency elements that co-occur with the first element in multiple original image samples.
[0281] Here, one original text sample is used to describe one original image sample, and the original text sample includes the name of the first element sample in the described original image sample.
[0282] In some embodiments, when the number of original image samples is multiple, refer to Figure 3L , Figure 3L which is the twelfth process schematic diagram of the data processing method of the machine learning model provided by the embodiments of this application. Figure 3K The step 501 of "<perform a first rewriting on the original text samples in the first training set to obtain first generative text samples>" in Figure 3L can be implemented by steps 5011 to 5014 in
[0283] In step 5011, target recognition is performed on multiple original image samples, and the recognized elements are combined into an element set.
[0284] In some embodiments, target recognition can be performed on multiple original image samples by any one of the following algorithms: YOLO, Fast R-CNN, SSD, RPN. For details not described, reference can be made to the embodiments of step 101 above.
[0285] Exemplarily, taking Fast R-CNN as an example, for each original image sample, a convolutional neural network is used to extract features from the input original image sample to obtain a feature map; multiple candidate regions in the original image sample are generated through selective search; each candidate region is mapped to the feature map through a pooling operation to obtain a feature vector corresponding to each candidate region; the feature vector is input into a fully connected layer to fine-tune the candidate region, and the elements in the fine-tuned candidate region are used as the recognized elements; the recognized elements are combined into an element set.
[0286] In step 5012, multiple element subsets are extracted from the element set, and the number of occurrences of the element subsets in multiple original image samples is determined.
[0287] Here, the number of elements included in each element subset is greater than or equal to 2, and different element subsets may include the same elements. For details not described, reference can be made to the embodiments of step 102 above, which will not be elaborated here.
[0288] In step 5013, at least one of a high-frequency element subset and a low-frequency element subset is identified from multiple element subsets based on the occurrence times.
[0289] Here, the multiple element subsets are sorted in descending order based on the occurrence times to obtain a sorting result; a preset proportion of element subsets are selected starting from the first element subset in the sorting result as the high-frequency element subset; and the element subsets other than the high-frequency element subset among the multiple element subsets are used as the low-frequency element subset. For the details not described, reference can be made to the embodiment of step 103 above.
[0290] In step 5014, a first rewriting of the original text sample is performed based on at least one of the high-frequency element subset and the low-frequency element subset to obtain a first generative text sample, where the first generative text sample includes at least one of a negative description of the high-frequency element subset and a positive description of the low-frequency element subset.
[0291] Here, text generation can be performed only based on the high-frequency element subset to obtain a first generative text sample, where the first generative text sample only includes a negative description of the high-frequency element subset; text generation can be performed only based on the low-frequency element subset to obtain a first generative text sample, where the first generative text sample only includes a positive description of the low-frequency element subset; or text generation can be performed based on both the high-frequency element subset and the low-frequency element subset to obtain a first generative text sample, where the first generative text sample includes a negative description of the high-frequency element subset and a positive description of the low-frequency element subset. For the details not described, reference can be made to the embodiment of step 104 above.
[0292] In the embodiments of the present application, the first generative text sample includes at least one of a negative description of the high-frequency element subset and a positive description of the low-frequency element subset. The diverse description methods can provide more learning signals for the model. The model can better understand the distribution and characteristics of the elements in the image by learning these text samples, thereby improving the performance and generalization ability of the model. By analyzing and rewriting the original image, a first generative text sample different from the original text sample is generated, increasing the data diversity and improving the robustness and adaptability of the model.
[0293] Continue to refer to Figure 3K In step 502, image generation is performed based on the first generative text sample to obtain a generative image sample.
[0294] In some embodiments, natural language processing techniques are used to convert the first generative text sample into a first text sample vector; key features are extracted from the first text sample vector; based on the key features, a pre-trained image generation model is called. The generator of the image generation model generates candidate images according to the input key features, and the discriminator of the image generation model determines whether the generated candidate images are real. Through the adversarial training of the generator and the discriminator, the generator gradually generates images that conform to the description of the first generative text sample, that is, generative text images.
[0295] Exemplarily, the image generation model can be any one of the following: Generative Adversarial Networks (GAN), Diffusion Model, Autoregressive Model, Variational AutoEncoder (VAE).
[0296] In step 503, the original text sample is secondarily rewritten to obtain a second generative text sample. The second generative text sample is used to describe the first element sample and the second element sample that does not appear in the original image sample. The first training set composed of the original text sample and the original image sample, the second training set composed of the first generative text sample and the generative image sample, and the third training set composed of the second generative text sample and the original image sample are used to train the machine learning model.
[0297] In some embodiments, refer to Figure 3M , Figure 3M is a schematic flowchart of the data processing method of the machine learning model provided by the embodiments of the present application. Figure 3K For step 503 of “secondarily rewrite the original text sample to obtain a second generative text sample” of Figure 3M , for each original image sample described by the original text sample, it can be implemented by executing
[0298] In step 5031, the first element sample is identified from the original image sample.
[0299] Here, feature recognition is performed on the original image sample to obtain multiple candidate element samples; the candidate element samples that meet the preset conditions are used as the main element samples, where the preset conditions include at least one of the following: the distance between the position of the candidate element sample and the center of the original image sample is less than the distance threshold, the ratio of the size of the candidate element sample to the size of the original image sample is greater than the size threshold, and the similarity between the candidate feature of the candidate element sample and the element features of other element samples in the original image sample except the candidate element sample is less than the similarity threshold; the candidate element samples other than the main element samples among the multiple candidate element samples are used as the secondary element samples. Specific examples can refer to the examples in steps 2011 to 2013 above, which will not be elaborated here.
[0300] In step 5032, a second element sample that does not appear in the original image sample is generated based on the first element sample.
[0301] Here, at least one of a similar element and a related element of the first element sample is generated, where the similar element is an element whose similarity to the first element sample is greater than the similarity threshold, and the related element is an element whose co-occurrence times in image samples other than the original image sample is greater than the times threshold; the similar element and the related element are used as the second element samples. Specific examples can refer to the examples in steps 2021 to 2022 above, which will not be elaborated here.
[0302] In step 5033, the original text sample is rewritten based on the second element sample to obtain a second generative text sample for negatively describing the second element sample.
[0303] Here, a negative description of the second element sample is generated, where the original text sample also includes a positive description of the first element sample; the negative description of the second element sample and the positive description of the first element sample are combined to obtain the second generative text sample. Specific examples can refer to the examples in steps 2031 to 2032 above, which will not be elaborated here.
[0304] As an example of steps 501 to 503, refer to Figure 5 , Figure 5 is a schematic flowchart of the process for training a machine learning model provided by an embodiment of the present application. In Figure 5In [the method], the original text sample and the original image sample described by the original text sample are combined into a first training set; a high-frequency element subset and a low-frequency element subset are identified from the original image sample, a first generative text sample is generated based on the negative description of the high-frequency element subset and the positive description of the low-frequency element subset, and the first generative text sample and the generative image sample described by the first generative text sample are combined into a second training set; a first element sample and a second element sample are identified from the original image sample, a second generative text sample is generated based on the positive description of the first element sample and the negative description of the second element sample, and the second generative text sample and the original image sample are combined into a third training set; a machine learning model is trained based on the first training set, the second training set, and the third training set.
[0305] In the embodiment of this application, by identifying a first element sample from the original image sample and generating a second element sample that does not appear in the original image, the content range expressed by the image can be expanded, enriching the diversity of the data. Based on the second element sample, the original text sample is rewritten to obtain a second generative text sample for negatively describing the second element sample, enhancing the semantic expression of the text and making the text more richly and accurately reflect the content of the image. Introducing the second element sample that does not appear in the original image allows the model to learn more different element combinations and semantic relationships, thereby improving the generalization ability of the model.
[0306] Next, an exemplary application of the embodiment of this application in an application scenario of a video platform will be described.
[0307] Video moment retrieval aims to find the target video segment in a long video that best matches the description of a natural language query statement given by a user. Related technologies directly calculate the similarity between a preset candidate segment and the text, sort the similarities, and obtain video segments with similarities higher than the similarity threshold, or take the entire video as the processing unit and directly use the segment time points as the prediction targets. With the development of multimodal large language models, related technologies use large language models for video moment retrieval. For example, the large language model is used to describe the video shots or time windows, and the moment retrieval is performed by calculating the similarity between the matching segment description and the query or directly requesting the large language model to predict the edges of the moment.
[0308] Although vision-language models (i.e., cross-modal matching models) have made progress under large-scale training, the ability to understand complex texts containing negation has not been fully studied. Related technologies assist training and improve the effect of cross-modal matching by constructing negative samples of texts and segments in video temporal localization, or by decomposing texts into corresponding subjects, predicates, and objects and then matching them with sub-regions of images to improve the degree of detail matching. However, these methods are not optimized for the problem of the model understanding complex texts. The Contrastive Language-Image Pre-training (CLIP) model supplements the negation logic by separately introducing learnable no prompts and a negation text encoder. See Figure 6 , Figure 6 which is a schematic structural diagram of the multi-modal pre-training neural network model provided by an embodiment of the present application.
[0309] Related technologies only use samples between different concepts in the dataset as negative samples for training, which can neither improve the model's understanding of negation nor handle the situation of uncommon concept combinations, resulting in the model being unable to understand the complex query semantics of partial positive and partial negative when performing retrieval. At the same time, since the constructed training data is negative sentences generated based on templates and the sentence patterns are single, the robustness of the model is not strong enough and it is easy to fall into bias.
[0310] Embodiments of the present application add text expressions containing negation concepts to the current image-text matching dataset, introduce a large language model to improve the expression diversity of negation meanings, and optimize the comprehensiveness of the dataset; construct descriptive texts containing negation words and use text-image generation technology to directionally generate rare scene images and automatically construct a large number of negative scene training data to reduce the bias of the current training data; through a contrastive learning framework, introduce positive and negative concepts to increase the understanding ability of the image and text for negation concepts, so that the model can better support queries containing negation meanings. It solves the deficiencies of related technologies in understanding and processing negation in texts and improves the adaptation ability in the retrieval scenario. When performing advertisement point retrieval for movie and TV drama scenes, customers will have relatively complex requirements. For example, as a beverage advertisement, the customer hopes to find scenes related to the product but avoid the appearance of other competing products in the picture, such as locating the picture of "a person eating, but not containing any beverages" in a long video. The data processing method of the machine learning model provided by the embodiments of the present application can also be applied to other video clip retrieval scenarios, such as an auxiliary video editing retrieval library. Complex text queries containing negation concepts can support users to better express and retrieve the target clip, improving efficiency.
[0311] The data processing method of the machine learning model provided by the embodiments of this application optimizes the cross-modal matching model (i.e., the above-mentioned machine learning model) in video clip retrieval, enabling it to have better understanding ability for processing complex texts containing negative semantics and supporting the matching retrieval of complex queries. Specifically, for a cross-modal matching model (taking the CLIP model as an example), the dataset is expanded by adding text expression forms of negative concepts, the negative semantics dataset is expanded by using text-to-image technology to construct unconventional pictures, and the loss function in model training is optimized to actively capture negative semantics, so as to improve the model's understanding of negative meanings.
[0312] CLIP is a multi-modal model proposed by OpenAI, which realizes cross-modal understanding between images and texts. The working principle of the model is to embed (map) images and texts into a common semantic space. Through contrastive learning, relevant images and texts are brought closer to each other, while irrelevant ones are kept away from each other. When training the image model, the images used are not necessarily from videos, but must be matching text-image pairs; the model trained on images can be supplemented with additional video training data or directly transferred to videos to recognize the video frames. See Figure 7 , Figure 7 is a schematic diagram of the principle for training a multi-modal model provided by the embodiments of this application. The CLIP model consists of an image encoder and a text encoder based on the Transformer architecture. The Transformer architecture can handle long-range dependencies and is trained on large-scale data. Cross-modal models such as CLIP align the visual and text feature spaces. By using similarity measurement methods (such as cosine similarity, Euclidean distance, etc.), the similarity between the text vector and the visual feature vector of each preset video clip (such as split clips based on storyboards or fixed-duration segments) (the video features can be obtained by using average pooling, etc.) is calculated. According to the calculated similarity, the video clips are sorted from high to low similarity to determine whether there are matching query results for video clip retrieval.
[0313] See Figure 8 , Figure 8It is a schematic diagram for constructing negative semantic text provided by an embodiment of this application. For the pictures in the image-text matching training dataset, a Visual Language Large Model (VLLM) is used for object detection and description to identify the specific elements present in the image, such as a cake, cream, colored sprinkles, and a blue background. Among them, the cake is the main element, and the other elements are positive concepts (w_r) related to the picture. Prompt words are constructed based on the elements in the image, and the output is obtained by calling the large model based on the prompt words. For example, the prompt word can be: "Please describe this picture in detail, identify the main element and auxiliary elements in the picture, and return them in JSON format, such as {"main_element": "cake", "relevant_elements": ["cream", "sprinkles", "blue background"]}."
[0314] Next, similar to the above, by constructing prompt words and inputting the prompt words into the VLLM, elements that are similar to but absent from the main element in the image (w_d) are listed, such as bread and ice cream, and elements that are related to the main element but absent from the image (w_n) are listed, such as candles, fruits, chocolates, and flowers that may be on the cake. The prompt word corresponding to the similar elements can be: "List at least three elements similar to the main element 'cake' in this picture and return them in JSON format, such as {"parallel_elements": ["bread", "ice cream", "pudding"]}." The prompt word corresponding to the relevant elements can be "Please list the elements that are related to the main element 'cake' in this picture but are not in the picture, listing at least three. For example, if the picture is of a cake and there may be candles, fruits, etc. on the cake, but in fact these elements are not in the picture, list other conforming elements and return them in JSON format, such as {"negative_elements": ["candles", "fruits", "chocolates", "flowers"]}."
[0315] Finally, based on the detection and recognition results, negative descriptions related to the main body of the image can be generated. For example, "This is not bread, but a cream cake instead of an ice cream." The contrastive learning framework of related technologies can master the differences between different subjects through the comparison of positive and negative samples in the entire training set, but the semantic expression learning of negative texts is not sufficient. Therefore, actively constructing more data of negative texts can improve the training effect of the model. For each negative description, there will be a set of annotation results of <w_d, w_r, image, caption>. Input the constructed prompt words into the large language model to obtain rich negative descriptions. For example, the prompt words can be: "Please create a sentence according to each of the above negative elements, which needs to conform to the content of the picture and contain the negative elements, and return it in JSON format." The output text can be: "{"Bread": "This is not bread, but a cake decorated with cream and colored sugar sprinkles on the surface.", "Mousse": "This is not a mousse with a light texture, but a cream cake with multiple layers.", "Pudding": "This is not a pudding with a smooth texture, but a round and beautifully decorated cream cake."}"
[0316] For the related but non-existent elements (w_n) in the image, use the large language model to generate complex combined sentences that contain both correct and negative concepts. For example, "This is a cream cake, but there are no candles on the cake, no flowers on the top, and it is not chocolate flavored, etc." For each description, there will be a set of annotation results of <w_n, w_r, image, caption>. The prompt words can be: "Please add the above elements that do not exist in the picture to the description of the picture. Describe each element in a separate sentence and return it in JSON format. For example, [{"Candle": "This is a beautifully decorated round cream cake, and there are no candles on it."}]"
[0317] Although the images in the large-scale training data reflect the dependency relationships of concepts in the real world to a certain extent, they also limit the model's understanding of diverse concept combinations and negative concepts. For example, a person wearing a suit is very likely to appear in an office and less likely to appear on the grass. To untie the "inherent impression" of concept combinations in the model, the data processing method of the machine learning model provided by the embodiments of this application extends the model's understanding of the negative concept space. See Figure 9 , Figure 9It is a schematic diagram of constructing a complex image through an image generation model provided by an embodiment of the present application. First, for the elements existing in the training data (i.e., the first training set), the frequency (i.e., the number of occurrences) of any two-element pairs (i.e., element subsets) that appear in the same image is counted to obtain a list of common concept combinations (i.e., high-frequency element subsets). Secondly, for each main element (i.e., the main element sample), m high-frequency words w_n are selected as negative concepts (i.e., negative descriptions) and m elements w_r that are not related to it (elements that do not appear simultaneously with the main element) and are not in the top K two-element pairs are selected as positive concepts (i.e., positive descriptions). m text descriptions are generated using a large language model (LLM), and these text descriptions are input into the image generation model to obtain corresponding pictures. For example, for "cake", "cream" is selected as the negative concept w_n, and "football" is selected as the positive concept w_r. By constructing a prompt and inputting the prompt into the LLM, a description of a picture that does not contain cream but has a football and a cake is generated, and the description is input into the image generation model to generate the corresponding picture. Additionally, the generated pictures can be input into the VLLM to determine whether the generated pictures conform to the corresponding positive and negative element rules (i.e., cream is the negative concept w_n and should not appear, and football is the positive concept / positive element w_r and should appear), and only the valid pictures (i.e., a cake picture in the shape of a football and without cream) are retained. Then, m <w_n, w_r, caption, image> combination training data can be obtained for each main concept.
[0318] When CLIP is trained, a multi-class N-pair loss function is used, that is, when each anchor has multiple positive and negative examples. Given N image-text pairs, CLIP learns a multi-modal embedding by jointly training a visual encoder and a text encoder to maximize the cosine similarity between the image embedding and the text embedding of the N true image-text pairs in a batch, while minimizing the cosine similarity between the embeddings of N 2 - N incorrectly paired embeddings, where the loss function is a symmetric cross-entropy loss.
[0319] Through the above two training data augmentation methods, a dataset D = {w_n / w_d, w_r, caption, image} for negative concepts can be obtained. Through contrastive learning of text-image pairs with positive and negative elements respectively, it is made to be close to the positive concept and far from the negative concept. Using the triplet loss (i.e., the second loss function), when calculating the loss, a positive sample, a negative sample, and an anchor sample are used as inputs simultaneously. The triplet loss L triplet can be seen in the above formula (4).
[0320] L triplet After weighting with the original loss function of CLIP (i.e., the first loss function), the complete training loss function (i.e., the preset loss function) is obtained. Among them, the loss function weighting is empirically weighted in equal proportion. For example, (L1 + L triplet ) / 2, where L1 is the original loss function of CLIP. Considering that the model has insufficient understanding of negative words in the text, it is necessary to train the embedding layer and attention layer of the text encoder to endow the model with new knowledge about how negative words affect the semantics of a given scenario. Therefore, the image encoder is frozen, and the text encoder of CLIP is fine-tuned based on the final loss function.
[0321] By using the large language model and the image generation model in the embodiments of the present application, large-scale high-quality cross-modal matching training data containing negative complex semantics can be effectively constructed, which can enhance the generalization ability of the model and the understanding ability of negative concepts. By designing a contrastive learning framework that simultaneously considers positive and negative concepts, the model can explicitly understand the complex semantics containing negation during the training process, improve the model's understanding of negation, and support more robust cross-modal retrieval. The progress of the cross-modal matching model in understanding negative concepts helps to support more complex video retrieval applications and enhance the adaptability during retrieval. For example, users can search for images or videos that contain certain objects but do not contain other objects through natural language.
[0322] Next, the implementation of the data processing device 233 of the machine learning model provided in the embodiments of the present application as a software module will be continued. In some embodiments, as Figure 2A shown, the software module stored in the data processing device 233 of the machine learning model in the memory 230 may include:
[0323] The element recognition module 2331 is used to perform target recognition on multiple original image samples in the first training set and combine the recognized elements into an element set.
[0324] The subset extraction module 2332 is used to extract multiple element subsets from the element set and determine the number of occurrences of the element subsets in multiple original image samples.
[0325] The subset recognition module 2333 is used to recognize at least one of the high-frequency element subsets and the low-frequency element subsets from multiple element subsets based on the number of occurrences.
[0326] The first generation module 2334 is used to generate text based on at least one of the high-frequency element subsets and the low-frequency element subsets to obtain a first generative text sample, where the first generative text sample includes at least one of the negative description of the high-frequency element subset and the positive description of the low-frequency element subset.
[0327] A second generation module 2335 is configured to generate a generative image sample based on a first generative text sample, where the first generative text sample and the generative image sample are used to train a machine learning model.
[0328] In some embodiments, the subset recognition module 2333 is further configured to perform a descending order sorting on multiple element subsets based on the occurrence times to obtain a sorting result; select a preset proportion of element subsets starting from the first element subset in the sorting result as high-frequency element subsets; and use the element subsets other than the high-frequency element subsets in the multiple element subsets as low-frequency element subsets.
[0329] In some embodiments, the first generation module 2334 is further configured to generate negative descriptions of at least some high-frequency elements in the high-frequency element subsets, where the negative descriptions of the at least some high-frequency elements represent that the generative image sample does not have the at least some high-frequency elements; generate positive descriptions of each low-frequency element in the low-frequency element subsets, where the positive descriptions represent that the generative image sample includes the low-frequency elements; and generate the first generative text sample based on at least one of the negative descriptions of the at least some high-frequency elements and the positive descriptions of the low-frequency elements.
[0330] In some embodiments, when at least some high-frequency elements are the high-frequency elements in the first part of the high-frequency element subsets, the first generation module 2334 is further configured to generate positive descriptions of the high-frequency elements in the second part of the high-frequency element subsets, where the high-frequency elements in the second part are different from those in the first part; and generate the first generative text sample based on the negative descriptions corresponding to the high-frequency elements in the first part, the positive descriptions corresponding to the high-frequency elements in the second part, and the positive descriptions of the low-frequency elements.
[0331] In some embodiments, the first training set further includes multiple original text samples, where one original text sample is used to describe one original image sample. The second generation module 2335 is further configured to perform the following processing on the original image sample described by each original text sample: identify a first element sample from the original image sample; generate a second element sample that does not appear in the original image sample based on the first element sample; and rewrite the original text sample based on the second element sample to obtain a second generative text sample for negatively describing the second element sample, where the second generative text sample and the original image sample are used to train a machine learning model.
[0332] In some embodiments, the first element sample includes a main element sample and a secondary element sample. The second generation module 2335 is further configured to perform feature recognition on the original image sample to obtain a plurality of candidate element samples; use the candidate element samples that meet the preset conditions as the main element samples, where the preset conditions include at least one of the following: the distance between the position of the candidate element sample and the center of the original image sample is less than a distance threshold, the ratio of the size of the candidate element sample to the size of the original image sample is greater than a size threshold, and the similarity between the candidate features of the candidate element sample and the element features of other element samples in the original image sample except the candidate element sample is less than a similarity threshold; use the candidate element samples other than the main element samples among the plurality of candidate element samples as the secondary element samples.
[0333] In some embodiments, the second generation module 2335 is further configured to generate at least one of a similar element of the first element sample and a related element of the first element sample, where the similar element is an element whose similarity to the first element sample is greater than a similarity threshold, and the related element is an element whose co-occurrence times with the first element sample in an image sample other than the original image sample is greater than a frequency threshold; use the similar element and the related element as the second element samples.
[0334] In some embodiments, the original text sample includes a positive description of the first element sample. The second generation module 2335 is further configured to generate a negative description of the second element sample; combine the negative description of the second element sample and the positive description of the first element sample to obtain a second generative text sample.
[0335] In some embodiments, the second generation module 2335 is further configured to construct a second training set based on the first generative text sample and the generative image sample, and construct a third training set based on the second generative text sample and the original image sample; collect a plurality of text-image sample pairs from the first training set, the second training set, and the third training set; determine the text-image matching degree of each text-image sample pair, use the text-image sample pairs with a text-image matching degree greater than a matching threshold as positive samples, and use the text-image sample pairs with a text-image matching degree less than a matching threshold as negative samples; determine a loss based on a preset loss function, positive samples, and negative samples, and update the initialized machine learning model based on the loss to obtain a pre-trained machine learning model.
[0336] In some embodiments, the second generation module 2335 is further configured to, for each pair of graphic and text samples, perform a first encoding on the image sample in the pair of graphic and text samples to obtain an image vector, and perform a second encoding on the text sample in the pair of graphic and text samples to obtain a text vector; determine a first similarity between the image vector and the text vector, and construct a first loss function based on the similarity; determine a second similarity between a preset anchor sample and a positive sample, and determine a third similarity between the anchor sample and a negative sample; construct a second loss function based on the second similarity and the third similarity; and perform a weighted sum on the first loss function and the second loss function to obtain a preset loss function.
[0337] In some embodiments, the second generation module 2335 is further configured to filter out generative image samples that do not meet the preset screening conditions from multiple generative image samples, where the screening conditions include: any high-frequency element in the high-frequency element subset does not appear in the generative image sample, and the generative image sample includes at least one low-frequency element in the low-frequency element subset.
[0338] Next, the exemplary structure of the data processing device 273 of the machine learning model provided in the embodiments of the present application implemented as software modules will be continued. In some embodiments, as Figure 2B shown, the software modules stored in the data processing device 273 of the machine learning model in the memory 270 may include:
[0339] The data acquisition module 2731 is configured to acquire a first training set, where the first training set includes multiple original text samples and multiple original image samples, and one original text sample is used to describe one original image sample.
[0340] The sample recognition module 2732 is configured to, for the original image sample described by each original text sample, perform the following processing: identify a first element sample from the original image sample.
[0341] The third generation module 2733 is configured to generate a second element sample that does not appear in the original image sample based on the first element sample.
[0342] The sample rewriting module 2734 is configured to rewrite the original text sample based on the second element sample to obtain a second generative text sample for negatively describing the second element sample, where the second generative text sample and the original image sample are used to train the machine learning model.
[0343] Next, the exemplary structure of the data processing device 2923 of the machine learning model provided in the embodiments of the present application implemented as software modules will be continued. In some embodiments, as Figure 2C shown, the software modules stored in the data processing device 2923 of the machine learning model in the memory 292 may include:
[0344] The first rewriting module 29231 is configured to perform a first rewriting on the original text samples in the first training set to obtain first generative text samples, where the original text samples are used to describe the first element samples included in the original image samples in the first training set, and the first generative text samples are used to describe the high-frequency elements and low-frequency elements co-occurring with the first element among multiple original image samples.
[0345] The fourth generation module 29232 is configured to perform image generation based on the first generative text samples to obtain generative image samples.
[0346] The second rewriting module 29233 is configured to perform a second rewriting on the original text samples to obtain second generative text samples, where the second generative text samples are used to describe the first element samples and the second element samples not appearing in the original image samples. The first training set composed of the original text samples and the original image samples, the second training set composed of the first generative text samples and the generative image samples, and the third training set composed of the second generative text samples and the original image samples are used to train a machine learning model.
[0347] In some embodiments, when the number of original image samples is multiple, the first rewriting module 29231 is further configured to perform object recognition on the multiple original image samples, combine the recognized elements into an element set; extract multiple element subsets from the element set, and determine the number of occurrences of the element subsets in the multiple original image samples; based on the number of occurrences, recognize at least one of the high-frequency element subsets and the low-frequency element subsets from the multiple element subsets; perform a first rewriting on the original text samples based on at least one of the high-frequency element subsets and the low-frequency element subsets to obtain first generative text samples, where the first generative text samples include at least one of a negative description of the high-frequency element subset and a positive description of the low-frequency element subset.
[0348] In some embodiments, the second rewriting module 29233 is further configured to perform the following processing on the original image samples described by each original text sample: recognize the first element samples from the original image samples; generate second element samples that do not appear in the original image samples based on the first element samples; rewrite the original text samples based on the second element samples to obtain second generative text samples for negatively describing the second element samples.
[0349] An embodiment of the present application provides a computer program product, which includes a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium, and the processor executes the computer-executable instructions, so that the electronic device executes the data processing method of the machine learning model described above in the embodiment of the present application.
[0350] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, in which computer-executable instructions or a computer program are stored. When the computer-executable instructions or the computer program are executed by a processor, the processor will be caused to execute the data processing method of the machine learning model provided by the embodiment of the present application. For example, as Figure 3A the data processing method of the machine learning model shown.
[0351] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disc, or CD-ROM; or it may be various devices including one or any combination of the above memories.
[0352] In some embodiments, the computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including being deployed as an independent program or being deployed as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0353] As an example, the computer-executable instructions may or may not correspond to a file in the file system, and may be stored as part of a file that stores other programs or data. For example, they may be stored in one or more scripts in a Hyper Text Markup Language (HTML) document, stored in a single file dedicated to the program being discussed, or stored in multiple cooperating files (for example, files that store one or more modules, subroutines, or code portions).
[0354] As an example, the computer-executable instructions may be deployed to execute on one electronic device, or on multiple electronic devices located at one location, or on multiple electronic devices distributed at multiple locations and interconnected by a communication network.
[0355] In summary, in the embodiment of the present application, by combining the element combinations identified in multiple original image samples into an element set, extracting multiple element subsets from the element set, identifying high-frequency element subsets and low-frequency element subsets from the element subsets based on the occurrence times, generating a first generative text sample based on at least one of the negative descriptions of the high-frequency element subsets and the positive descriptions of the low-frequency element subsets, and obtaining a generative image sample based on the first generative text sample. Since the negative descriptions of the high-frequency element subsets or the positive descriptions of the low-frequency element subsets are texts opposite to the conventional texts, the generative image sample obtained from the first generative text sample is also an unconventional image. Compared with only using conventional data to train a machine learning model, the bias of the machine learning model against unconventional data is eliminated, and the accuracy of the machine learning model is improved.
[0356] As mentioned above, the above are only embodiments of the present application and are not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and scope of the present application are all included in the protection scope of the present application.
Claims
1. A data processing method for a machine learning model, characterized in that, The method includes: Performing object recognition on multiple original image samples in a first training set, and combining the recognized elements into an element set; Extracting multiple element subsets from the element set, and determining the number of occurrences of the element subsets in the multiple original image samples; Based on the number of occurrences, identifying at least one of a high-frequency element subset and a low-frequency element subset from the multiple element subsets; Performing text generation based on at least one of the high-frequency element subset and the low-frequency element subset to obtain a first generated text sample, where the first generated text sample includes at least one of a negative description of the high-frequency element subset and a positive description of the low-frequency element subset; Performing image generation based on the first generated text sample to obtain a generated image sample, where the first generated text sample and the generated image sample are used to train the machine learning model.
2. The method according to claim 1, wherein The identifying at least one of a high-frequency element subset and a low-frequency element subset from the multiple element subsets based on the number of occurrences includes: Performing a descending order sorting on the multiple element subsets based on the number of occurrences to obtain a sorting result; Selecting a preset proportion of the element subsets starting from the first element subset in the sorting result as the high-frequency element subset; Taking the element subsets among the multiple element subsets other than the high-frequency element subset as the low-frequency element subset.
3. The method according to claim 1, characterized in that, The performing text generation based on at least one of the high-frequency element subset and the low-frequency element subset to obtain a first generated text sample includes: Generating a negative description of at least some of the high-frequency elements in the high-frequency element subset, where the negative description of the at least some of the high-frequency elements represents that the at least some of the high-frequency elements do not appear in the generated image sample; Generating a positive description of each low-frequency element in the low-frequency element subset, where the positive description represents that the generated image sample includes the low-frequency element; Generating a first generated text sample based on at least one of the negative description of the at least some of the high-frequency elements and the positive description of the low-frequency elements.
4. The method according to claim 3, characterized in that, When the at least some of the high-frequency elements are the high-frequency elements in the first part of the high-frequency element subset, when generating the negative description of the at least some of the high-frequency elements, the method further includes: Generating a positive description of the high-frequency elements in the second part of the high-frequency element subset, where the high-frequency elements in the second part are different from the high-frequency elements in the first part; The generating a first generated text sample based on at least one of the negative description of the at least some of the high-frequency elements and the positive description of the low-frequency elements includes: Generating a first generated text sample based on the negative description corresponding to the high-frequency elements in the first part, the positive description corresponding to the high-frequency elements in the second part, and the positive description of the low-frequency elements.
5. The method according to any one of claims 1 to 4, wherein The first training set further includes a plurality of original text samples, and one of the original text samples is used to describe one of the original image samples; The method further includes: For each of the original image samples described by the original text samples, perform the following processing: Identify a first element sample from the original image sample; Generate a second element sample that does not appear in the original image sample based on the first element sample; Rewrite the original text sample based on the second element sample to obtain a second generative text sample for negatively describing the second element sample, wherein the second generative text sample and the original image sample are used to train the machine learning model.
6. The method according to claim 5, wherein The first element sample includes a main element sample and a secondary element sample; The identifying a first element sample from the original image sample includes: Performing feature recognition on the original image sample to obtain a plurality of candidate element samples; Taking the candidate element samples that meet the preset conditions as the main element samples, wherein the preset conditions include at least one of the following: the distance between the position of the candidate element sample and the center of the original image sample is less than a distance threshold, the ratio of the size of the candidate element sample to the size of the original image sample is greater than a size threshold, and the similarity between the candidate feature of the candidate element sample and the element features of other element samples in the original image sample except the candidate element sample is less than a similarity threshold; Taking the candidate element samples other than the main element samples among the plurality of candidate element samples as the secondary element samples.
7. The method according to claim 5, wherein The generating a second element sample that does not appear in the original image sample based on the first element sample includes: Generating at least one of a similar element of the first element sample and a related element of the first element sample, wherein the similar element is an element whose similarity to the first element sample is greater than a similarity threshold, and the related element is an element whose co-occurrence times with the first element sample in image samples other than the original image sample is greater than a times threshold; Taking the similar element and the related element as the second element samples.
8. The method according to claim 5, wherein The original text sample includes a positive description of the first element sample; The rewriting the original text sample based on the second element sample to obtain a second generative text sample for negatively describing the second element sample includes: Generating a negative description of the second element sample; Combining the negative description of the second element sample and the positive description of the first element sample to obtain a second generative text sample.
9. The method according to claim 5, wherein After the rewriting the original text sample based on the second element to obtain a second generative text sample, the method further includes: Constructing a second training set based on the first generative text sample and the generative image sample, and constructing a third training set based on the second generative text sample and the original image sample; Collect a plurality of text-image sample pairs from the first training set, the second training set, and the third training set; Determine the text-image matching degree of each text-image sample pair, and use the text-image sample pairs with a text-image matching degree greater than the matching degree threshold as positive samples, and use the text-image sample pairs with a text-image matching degree less than the matching degree threshold as negative samples; Determine the loss based on a preset loss function, the positive samples, and the negative samples, and update the initialized machine learning model based on the loss to obtain the pre-trained machine learning model.
10. The method according to claim 9, wherein Before determining the loss between the positive samples and the negative samples based on the preset loss function, the method further includes: For each text-image sample pair, perform a first encoding on the image sample in the text-image sample pair to obtain an image vector, and perform a second encoding on the text sample in the text-image sample pair to obtain a text vector; Determine the first similarity between the image vector and the text vector, and construct a first loss function based on the similarity; Determine the second similarity between a preset anchor sample and the positive sample, and determine the third similarity between the anchor sample and the negative sample; Construct a second loss function based on the second similarity and the third similarity; Perform a weighted sum on the first loss function and the second loss function to obtain the preset loss function.
11. The method according to any one of claims 1 to 10, characterized in that After generating a generated image sample based on the first generated text sample, the method further includes: Filter out the generated image samples that do not meet the preset screening conditions from the multiple generated image samples, where the screening conditions include: any high-frequency element in the high-frequency element subset does not appear in the generated image sample, and the generated image sample includes at least one low-frequency element in the low-frequency element subset.
12. A method for processing data of a machine learning model, characterized in that, The method includes: Obtain a first training set, where the first training set includes a plurality of original text samples and a plurality of original image samples, and one original text sample is used to describe one original image sample; For the original image sample described by each original text sample, perform the following processing: Identify a first element sample from the original image sample; Generate a second element sample that does not appear in the original image sample based on the first element sample; Rewrite the original text sample based on the second element sample to obtain a second generated text sample for negatively describing the second element sample, where the second generated text sample and the original image sample are used to train the machine learning model.
13. A data processing method for a machine learning model, characterized in that, The method includes: Perform a first rewrite on the original text samples in the first training set to obtain first generated text samples, where the original text samples are used to describe the first element samples included in the original image samples in the first training set, and the first generated text samples are used to describe the high-frequency elements and low-frequency elements that co-occur with the first element in the multiple original image samples; Generate a generated image sample based on the first generated text sample; Perform a second rewrite on the original text sample to obtain a second generative text sample, where the second generative text sample is used to describe the first element sample and a second element sample that does not appear in the original image sample. The first training set composed of the original text sample and the original image sample, the second training set composed of the first generative text sample and the generative image sample, and the third training set composed of the second generative text sample and the original image sample are used to train the machine learning model.
14. The method according to claim 13, wherein When the number of the original image samples is multiple, the performing a first rewrite on the original text sample in the first training set to obtain a first generative text sample includes: Perform object recognition on the multiple original image samples, and combine the recognized elements into an element set; Extract multiple element subsets from the element set, and determine the number of occurrences of the element subsets in the multiple original image samples; Based on the number of occurrences, recognize at least one of a high-frequency element subset and a low-frequency element subset from the multiple element subsets; Perform a first rewrite on the original text sample based on at least one of the high-frequency element subset and the low-frequency element subset to obtain a first generative text sample, where the first generative text sample includes at least one of a negative description of the high-frequency element subset and a positive description of the low-frequency element subset.
15. The method according to claim 13, wherein The performing a second rewrite on the original text sample to obtain a second generative text sample includes: For each original image sample described by the original text sample, perform the following processing: Identify the first element sample from the original image sample; Generate a second element sample that does not appear in the original image sample based on the first element sample; Rewrite the original text sample based on the second element sample to obtain a second generative text sample for negatively describing the second element sample.
16. A data processing device for a machine learning model, characterized in that, The device includes: An element recognition module, configured to perform object recognition on multiple original image samples in a first training set, and combine the recognized elements into an element set; A subset extraction module, configured to extract multiple element subsets from the element set, and determine the number of occurrences of the element subsets in the multiple original image samples; A subset recognition module, configured to recognize at least one of a high-frequency element subset and a low-frequency element subset from the multiple element subsets based on the number of occurrences; A first generation module, configured to perform text generation based on at least one of the high-frequency element subset and the low-frequency element subset to obtain a first generative text sample, where the first generative text sample includes at least one of a negative description of the high-frequency element subset and a positive description of the low-frequency element subset; A second generation module, configured to perform image generation based on the first generative text sample to obtain a generative image sample, where the first generative text sample and the generative image sample are used to train the machine learning model.
17. A data processing device for a machine learning model, characterized in that, The device includes: A data acquisition module, configured to acquire a first training set, where the first training set includes a plurality of original text samples and a plurality of original image samples, and one of the original text samples is used to describe one of the original image samples; For each of the original image samples described by the original text samples, perform the following processing: A sample recognition module, configured to recognize first element samples from the original image samples; A third generation module, configured to generate second element samples that do not appear in the original image samples based on the first element samples; A sample rewriting module, configured to rewrite the original text samples based on the second element samples to obtain second generative text samples for negatively describing the second element samples, where the second generative text samples and the original image samples are used to train the machine learning model.
18. An electronic device, characterized in that, The electronic device includes: A memory, configured to store computer-executable instructions or computer programs; A processor, configured to implement the data processing method of the machine learning model according to any one of claims 1 to 15 when executing the computer-executable instructions or computer programs stored in the memory.
19. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that, The computer-executable instructions or computer programs, when executed by the processor, implement the data processing method of the machine learning model according to any one of claims 1 to 15.
20. A computer program product, comprising computer-executable instructions or a computer program, characterized in that, The computer-executable instructions or computer programs, when executed by the processor, implement the data processing method of the machine learning model according to any one of claims 1 to 15.
Citation Information
Cited By
Model training and image-text matching method and device, electronic equipment and storage medium
CN121436092A