Data enhancement method, device, equipment and storage medium for graphic and text cross-modal model
By performing business scenario classification image enhancement and mosaic processing on the training data set of the cross-modal double tower model of graphics and text, as well as back-translation of text data and repeated statement operations, the problem of insufficient generalization capabilities of the model is solved, and better data enhancement effects and model performance are achieved.
Patent Information
- Application Number
- CN202210898897.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-28
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-07-28
AI Technical Summary
The generalization capability of the cross-modal dual-tower model of graphics and text is insufficient. The existing technology cannot meet complex business scenarios in terms of data enhancement, and traditional EDA data enhancement technology may lead to semantic changes.
By classifying the image data set by business scenarios, image enhancement processing is performed, and image data sets are enriched by combining mosaic data enhancement processing; text data is back-translated and sentence repetitive operations are performed to generate preprocessed text data to enhance text data sets.
It improves the generalization capability of the cross-modal dual-tower model of graphics and text, meets the needs of complex business scenarios, enriches the data set, reduces the dependence on batch_size, and improves the recognition efficiency of the model.
Smart Images

Figure CN115203375B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a data enhancement method, device, electronic device and computer-readable storage medium for a cross-modal model of graphics and text. Background Art
[0002] With the development of image and text recognition technology, image and text cross-modal pre-training models are crucial technical means in image and text search, image captioning, and visual question answering (VAR) tasks. The learning process of image and text cross-modal pre-training models requires a large amount of images and texts as training data, and data enhancement processing is required for the training data.
[0003] Currently, pre-trained models are mainly divided into single encoders and image-text cross-modal dual-tower models. For the single encoder, data enhancement is performed by increasing the number and diversity of training data, and better generalization capabilities are obtained on small data sets. However, for massive data, the model recognition efficiency will be low. For the image-text cross-modal dual-tower model, single-point data enhancement is performed on images, which cannot meet complex business scenarios. The use of traditional EDA data enhancement technology for text data may bring about semantic changes, resulting in insufficient generalization capabilities of the image-text cross-modal dual-tower model. Summary of the invention
[0004] The present invention provides a data enhancement method, device and computer-readable storage medium for a cross-modal image-text model, the main purpose of which is to solve the problem of insufficient generalization ability of the cross-modal image-text dual-tower model.
[0005] To achieve the above object, the present invention provides a data enhancement method for a cross-modal model of graphics and text, comprising:
[0006] Acquire a training data set for a cross-modal image-text model, wherein the training data set includes an image data set and a text data set;
[0007] Classifying the image data set in the training data set according to business scenarios to obtain image categories, and performing image enhancement processing on the image data set based on the image categories to obtain a first image enhanced data set;
[0008] Dividing the first image enhancement data set into a plurality of preliminary image enhancement data subsets according to a preset rule, selecting a group of four images from each of the plurality of preliminary image enhancement data subsets, performing mosaic data enhancement processing on the four images to obtain a spliced image, and adding the spliced image of each preliminary image enhancement data subset to the first image enhancement data set to obtain an image enhancement data set;
[0009] A preset amount of text data is selected from the text data set in the training data set, back-translated and sentence repeated operations are performed on the preset amount of text data to obtain a preset amount of preprocessed text data, and the preset amount of preprocessed text data is added to the text data set to obtain a text enhancement data set.
[0010] Optionally, performing back-translation and sentence repeating operations on the preset amount of text data to obtain a preset amount of preprocessed text data includes:
[0011] Using a preset machine translation model, translating the preset amount of text data into first language text data, and then translating the first language text data into original language text data;
[0012] Randomly selecting a preset number of words from the preset number of text data, and backfilling the preset number of words into the preset number of text data to obtain a preset number of first preprocessed text data;
[0013] Acquire text data corresponding to each group of the four randomly scaled images from the text data set, and concatenate the corresponding text data to obtain a plurality of second preprocessed text data;
[0014] The original language text data, the preset number of first preprocessed text data and the plurality of second preprocessed text data are combined to obtain a preset number of preprocessed text data.
[0015] Optionally, performing image enhancement processing on the image dataset based on the image category to obtain a first image enhanced dataset includes:
[0016] Based on the image category, an image enhancement processing algorithm is selected from a preset algorithm library, and image enhancement processing is performed on the image data set in the training data set in the spatial domain to obtain a grayscale image set;
[0017] Performing Gaussian filtering on the grayscale image set in the frequency domain to obtain a smoothed grayscale image set;
[0018] The brightness, contrast, saturation and hue of the smoothed grayscale image set are randomly changed to obtain a preliminary enhanced data set.
[0019] Optionally, based on the image category, selecting an image enhancement processing algorithm from a preset algorithm library, performing image enhancement processing on the image data set in the training data set in the spatial domain to obtain a grayscale image set, includes:
[0020] According to the image category selection, a grayscale algorithm is selected from a preset algorithm library to perform grayscale transformation on the image data set to obtain a preliminary grayscale image set;
[0021] According to the image category selection, a sharpening algorithm is selected from a preset algorithm library, and the preliminary grayscale image set is sharpened to obtain a grayscale image set.
[0022] Optionally, performing Gaussian filtering on the grayscale image set in the frequency domain to obtain a smoothed grayscale image set includes:
[0023] According to a preset rule, different weights are assigned to pixels at different positions of the grayscale image in the grayscale image set to obtain position pixel weights;
[0024] Using a preset convolution template and based on the position pixel weights, weighted averaging is performed on pixels in the neighborhood of the grayscale image set to obtain a smoothed grayscale image set.
[0025] Optionally, performing mosaic data enhancement processing on the four images to obtain a spliced image includes:
[0026] Randomly scaling the four images respectively to obtain four randomly scaled images;
[0027] Randomly select a stitching center coordinate in a preset area, stitch according to the stitching center coordinate, and stitch the four randomly scaled images into the preset area;
[0028] When the four randomly scaled images exceed the preset area, the exceeded area is cropped to obtain a spliced image;
[0029] When the four randomly scaled images do not fill up the preset area, the unfilled area is filled to obtain a spliced graphic.
[0030] Optionally, the classifying the image data set in the training data set according to business scenarios to obtain image categories includes:
[0031] Extracting a feature vector set of an image data set in the training data set;
[0032] Matching the feature vector set with a business scene picture set in a preset business scene library to obtain a matching similarity set;
[0033] A matching similarity satisfying a similarity threshold is selected from the matching similarity set, and the business scene category marked by the business scene icon of the matching similarity satisfying the similarity threshold is used as the image category of the corresponding image data.
[0034] In order to solve the above problems, the present invention also provides a data enhancement device for a graphic-text cross-modal model, the device comprising:
[0035] A training set acquisition module, used to acquire a training data set for the image-text cross-modal model, wherein the training data set includes an image data set and a text data set;
[0036] an image enhancement processing module, configured to classify the image data set in the training data set according to the business scenario to obtain image categories, and based on the image categories, perform image enhancement processing on the image data set to obtain a first image enhancement data set; divide the first image enhancement data set into a plurality of preliminary image enhancement data subsets according to a preset rule, select a group of four images from the plurality of preliminary image enhancement data subsets, perform mosaic data enhancement processing on the four images to obtain a spliced image, and add the spliced image of each preliminary image enhancement data subset to the first image enhancement data set to obtain an image enhancement data set;
[0037] The text data enhancement processing module is used to select a preset amount of text data from the text data set in the training data set, perform back translation and sentence repeat operations on the preset amount of text data to obtain a preset amount of preprocessed text data, and add the preset amount of preprocessed text data to the text data set to obtain a text enhancement data set.
[0038] In order to solve the above problem, the present invention further provides an electronic device, the electronic device comprising:
[0039] at least one processor; and,
[0040] a memory communicatively connected to the at least one processor; wherein,
[0041] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the data enhancement method of the graphic-text cross-modal model described above.
[0042] In order to solve the above problems, the present invention also provides a computer-readable storage medium, in which at least one computer program is stored. The at least one computer program is executed by a processor in an electronic device to implement the data enhancement method of the above-mentioned graphic and text cross-modal model.
[0043] In the embodiment of the present invention, the image data set in the training data set is classified according to the business scenario to obtain image categories, and based on the image categories, the image data set is subjected to image enhancement processing to obtain a first image enhancement data set, which can meet various different business scenarios and improve the generalization ability of the image-text cross-modal model; the first image enhancement data set is subjected to mosaic data enhancement processing to obtain a spliced image, and the spliced image of each preliminary image enhancement data subset is added to the first image enhancement data set to obtain an image enhancement data set, which enriches the image data set and indirectly increases the batch_size of the training set, reduces the dependence on batch_size, and improves the generalization ability of the image-text cross-modal dual-tower model; a preset number of text data is selected from the text data set in the training data set, and the preset number of text data is back-translated and sentence repeated to obtain a preset number of pre-processed text data, and the preset number of pre-processed text data is added to the text data set to obtain a text enhancement data set, which does not change the semantics of the text data set, but changes the expression method, so that the text data set has multiple perspectives, and improves the generalization ability of the image-text cross-modal dual-tower model. Therefore, the data enhancement method, device, electronic device and computer-readable storage medium of the cross-modal image-text model proposed in the present invention can solve the problem of insufficient generalization ability of the cross-modal image-text dual-tower model. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 A schematic diagram of a flow chart of a data enhancement method for a graphic-text cross-modal model provided by an embodiment of the present invention;
[0045] Figure 2 for Figure 1 A schematic diagram of the detailed implementation process of one of the steps in the data enhancement method for the cross-modal model of text and images shown;
[0046] Figure 3 for Figure 1 A schematic diagram of a detailed implementation process of another step in the data enhancement method for the cross-modal model of text and images shown;
[0047] Figure 4 A functional module diagram of a data enhancement device for a graphic-text cross-modal model provided by an embodiment of the present invention;
[0048] Figure 5 A schematic diagram of the structure of an electronic device for implementing the data enhancement method of the graphic-text cross-modal model provided by one embodiment of the present invention.
[0049] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0050] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0051] The embodiment of the present application provides a data enhancement method for a graphic cross-modal model. The execution subject of the data enhancement method for the graphic cross-modal model includes but is not limited to at least one of the electronic devices such as a server, a terminal, etc. that can be configured to execute the method provided by the embodiment of the present application. In other words, the data enhancement method for the graphic cross-modal model can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes but is not limited to: a single server, a server cluster, a cloud server or a cloud server cluster, etc. The server can be an independent server, or it can be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks (Content Delivery Network, CDN), and big data and artificial intelligence platforms.
[0052] Reference Figure 1 FIG. 1 is a flow chart of a data enhancement method for a cross-modal model of images and texts provided by an embodiment of the present invention. In this embodiment, the data enhancement method for the cross-modal model of images and texts includes:
[0053] S1. Obtain a training data set for a cross-modal image-text model, wherein the training data set includes an image data set and a text data set.
[0054] In an embodiment of the present invention, the image-text cross-modal model may be a dual-tower model such as CLIP (Contrastive Language–Image Pre-training) and ALBEF (ALign Before Fuse).
[0055] The training data set described in the embodiment of the present invention includes an image data set and a text data set. The image data set can be images of the same object or person in different angles, different pixel colors, etc., or it can be images of different objects or people in different angles, different pixel colors, etc. The amount of image data is usually relatively large, reaching the level of millions.
[0056] In the embodiment of the present invention, the text data set may be text information associated with the image data set.
[0057] S2. Classify the image data set in the training data set according to business scenarios to obtain image categories, and perform image enhancement processing on the image data set based on the image categories to obtain a first image enhanced data set.
[0058] In detail, the image data set in the training data set is classified according to business scenarios in S2 to obtain image categories, including:
[0059] Extracting a feature vector set of an image data set in the training data set;
[0060] Matching the feature vector set with a business scene picture set in a preset business scene library to obtain a matching similarity set;
[0061] A matching similarity satisfying a similarity threshold is selected from the matching similarity set, and the business scene category marked by the business scene icon of the matching similarity satisfying the similarity threshold is used as the image category of the corresponding image data.
[0062] In the embodiment of the present invention, the HOG algorithm, the SIF algorithm or the convolutional neural network may be used to extract the feature vector set of the image data set.
[0063] Further, see Figure 2 As shown, the step of performing image enhancement processing on the image dataset based on the image category in S2 to obtain a first image enhanced dataset includes:
[0064] S21, based on the image category, selecting an image enhancement processing algorithm from a preset algorithm library, performing image enhancement processing on the image data set in the training data set in the spatial domain, and obtaining a grayscale image set;
[0065] S22, performing Gaussian filtering on the grayscale image set in the frequency domain to obtain a smoothed grayscale image set;
[0066] S23, randomly changing the brightness, contrast, saturation and hue of the smoothed grayscale image set to obtain a preliminary enhanced data set.
[0067] In the embodiment of the present invention, the brightness refers to the brightness of the light shining on the image. When the brightness of the image increases, it will appear dazzling or glaring, and the smaller the brightness, the image will appear gray; the contrast refers to the difference between different colors. The greater the contrast, the greater the contrast between different colors, and the smaller the contrast, the smaller the contrast between different colors; the saturation refers to the concentration of the image color. The higher the saturation, the richer the color, and the lower the saturation, the older the color; the hue refers to the brightness of the primary color in various image color modes. For a grayscale image, when the hue level is 255, it is white, and when the level is 0, it is black.
[0068] Further, the S21 includes:
[0069] According to the image category selection, a grayscale algorithm is selected from a preset algorithm library to perform grayscale transformation on the image data set to obtain a preliminary grayscale image set;
[0070] According to the image category selection, a sharpening algorithm is selected from a preset algorithm library, and the preliminary grayscale image set is sharpened to obtain a grayscale image set.
[0071] In the embodiment of the present invention, the grayscale transformation is to convert a color image into a grayscale image, and the grayscale conversion algorithm includes a component method, a maximum value method, an average value method, a weighted average method, and the like.
[0072] In an embodiment of the present invention, the sharpening algorithm includes a differential edge detection algorithm and a convolution edge enhancement algorithm.
[0073] In an embodiment of the present invention, different grayscale algorithms are selected according to the image category, and the image data set is grayscale transformed. Different sharpening algorithms can be selected according to the image category. The obtained grayscale image set is more reasonable and can meet various business scenarios. The generalization ability of the image-text cross-modal model is improved. The image data set is converted into a grayscale image set, which reduces the operation time of the image-text cross-modal dual-tower model and improves the recognition efficiency of the image-text cross-modal dual-tower model.
[0074] Further, the S22 includes:
[0075] According to a preset rule, different weights are assigned to pixels at different positions of the grayscale image in the grayscale image set to obtain position pixel weights;
[0076] Using a preset convolution template and based on the position pixel weights, weighted averaging is performed on pixels in the neighborhood of the grayscale image set to obtain a smoothed grayscale image set.
[0077] In the embodiment of the present invention, the preset rule is that the weight of the pixel closer to the center of the neighborhood is higher, so that when the image details are blurred, the overall grayscale feature distribution of the image can be more retained.
[0078] In the embodiment of the present invention, the preset convolution template is a convolution smoothing filter with convolution kernels of different sizes.
[0079] In an embodiment of the present invention, Gaussian filtering is performed on the grayscale image set in the frequency domain to obtain a smoothed grayscale image set, which can better retain the edge characteristics of the image data set, better smooth the local area, ensure the accuracy of the training data, and improve the accuracy of the image-text cross-modal dual-tower model.
[0080] In an embodiment of the present invention, the image data set is classified according to different business scenarios, and different image enhancement processing algorithms are selected according to the image category, which can meet various business scenarios. The obtained grayscale image set is more reasonable, and the generalization ability of the image-text cross-modal model is improved.
[0081] In an embodiment of the present invention, the image data set is subjected to image enhancement processing in the spatial domain and the frequency domain, and the brightness, contrast, saturation and hue of the data image set are randomly changed, so that the image data has multiple perspectives, thereby improving the generalization ability of the image-text cross-modal model.
[0082] S3. Divide the first image enhancement data set into multiple preliminary image enhancement data subsets according to preset rules, select a group of four images from the multiple preliminary image enhancement data subsets, perform mosaic data enhancement processing on the four images to obtain a spliced image, and add the spliced image of each preliminary image enhancement data subset to the first image enhancement data set to obtain an image enhancement data set.
[0083] In the embodiment of the present invention, the preset rule is a batchsize setting rule, which can be obtained based on historical training data and cannot be too large or too small. Usually, the batchsize is a multiple of 2.
[0084] In detail, the mosaic data enhancement process is performed on the four images in S3 to obtain a spliced image, including:
[0085] Randomly scaling the four images respectively to obtain four randomly scaled images;
[0086] Randomly select a stitching center coordinate in a preset area, stitch according to the stitching center coordinate, and stitch the four randomly scaled images into the preset area;
[0087] When the four randomly scaled images exceed the preset area, the exceeded area is cropped to obtain a spliced image;
[0088] When the four randomly scaled images do not fill up the preset area, the unfilled area is filled to obtain a spliced graphic.
[0089] In the embodiment of the present invention, the mosaic data enhancement is a new data enhancement method for mixing four images.
[0090] In the embodiment of the present invention, the four images may also be annotated, and when random scaling is performed, the corresponding annotation boxes are also scaled accordingly.
[0091] In the embodiment of the present invention, the image data set is enriched through Mosaic data enhancement, and the batch_size of the training set is indirectly increased, the dependence on batch_size is reduced, and the generalization ability of the image-text cross-modal dual-tower model is improved.
[0092] S4. Select a preset number of text data from the text data set in the training data set, perform back translation and sentence repeat operations on the preset number of text data to obtain a preset number of preprocessed text data, and add the preset number of preprocessed text data to the text data set to obtain a text enhancement data set.
[0093] For details, see Figure 3 As shown, the back-translation and sentence repetition operations on the preset amount of text data in S4 to obtain a preset amount of preprocessed text data include:
[0094] S41, using a preset machine translation model, translating the preset amount of text data into first language text data, and then translating the first language text data into original language text data;
[0095] S42, randomly selecting a preset number of words from the preset number of text data, and backfilling the preset number of words into the preset number of text data to obtain a preset number of first preprocessed text data;
[0096] S43, acquiring text data corresponding to each group of the four randomly scaled images from the text data set, and splicing the corresponding text data to obtain a plurality of second preprocessed text data;
[0097] S44: Merge the original language text data, the preset number of first preprocessed text data, and the plurality of second preprocessed text data to obtain a preset number of preprocessed text data.
[0098] In an embodiment of the present invention, a preset machine translation model is used to translate the preset amount of text data into text data in a first language, and then the text data in the first language is translated into text data in the original language. For example, if the preset amount of text data is in Chinese, the preset amount of text data can be translated into text data in English format, and then the text data in English format can be back translated into text data in Chinese format.
[0099] In an embodiment of the present invention, a preset number of words are randomly selected from the preset number of text data, and the preset number of words are backfilled into the preset number of text data to obtain a preset number of first preprocessed text data. For example, if the original language text data is "three young men and a young woman wearing sneakers are leaping in midair at the top of a flight of concrete stairs", "a", "in" and "midair" are selected and backfilled into the original language text data to obtain "three young men and a a young woman wearing sneakers are leaping in in midair midair at the top of a flight of concrete stairs".
[0100] In the embodiment of the present invention, the text dataset is back-translated and sentence repeated, which does not change the semantics of the text dataset, but changes the expression method, so that the text dataset has multiple perspectives, thereby improving the generalization ability of the text-image cross-modal dual-tower model.
[0101] In the embodiment of the present invention, the image data set in the training data set is classified according to the business scenario to obtain image categories, and based on the image categories, the image data set is subjected to image enhancement processing to obtain a first image enhancement data set, which can meet various different business scenarios and improve the generalization ability of the image-text cross-modal model; the first image enhancement data set is subjected to mosaic data enhancement processing to obtain a spliced image, and the spliced image of each preliminary image enhancement data subset is added to the first image enhancement data set to obtain an image enhancement data set, which enriches the image data set and indirectly increases the batch_size of the training set, reduces the dependence on batch_size, and improves the generalization ability of the image-text cross-modal dual-tower model; a preset number of text data is selected from the text data set in the training data set, and the preset number of text data is back-translated and sentence repeated to obtain a preset number of pre-processed text data, and the preset number of pre-processed text data is added to the text data set to obtain a text enhancement data set, which does not change the semantics of the text data set, but changes the expression method, so that the text data set has multiple perspectives, and improves the generalization ability of the image-text cross-modal dual-tower model. Therefore, the data enhancement method of the image-text cross-modal model proposed in the present invention can solve the problem of insufficient generalization ability of the image-text cross-modal dual-tower model.
[0102] like Figure 4 , which is a functional module diagram of a data enhancement device for a graphic-text cross-modal model provided by an embodiment of the present invention.
[0103] The data enhancement device 100 of the cross-modal model of text and images of the present invention can be installed in an electronic device. According to the functions to be implemented, the data enhancement device 100 of the cross-modal model of text and images can include a training set acquisition module 101, an image enhancement processing module 102 and a text data enhancement processing module 103. The module of the present invention can also be referred to as a unit, which refers to a series of computer program segments that can be executed by a processor of an electronic device and can complete fixed functions, which are stored in the memory of the electronic device.
[0104] In this embodiment, the functions of each module / unit are as follows:
[0105] The training set acquisition module 101 is used to acquire a training data set of the image-text cross-modal model, wherein the training data set includes an image data set and a text data set;
[0106] The image enhancement processing module 102 is used to classify the image data set in the training data set according to the business scenario to obtain image categories, and based on the image categories, perform image enhancement processing on the image data set to obtain a first image enhancement data set; divide the first image enhancement data set into multiple preliminary image enhancement data subsets according to a preset rule, select a group of four images from the multiple preliminary image enhancement data subsets, perform mosaic data enhancement processing on the four images to obtain a spliced image, and add the spliced image of each preliminary image enhancement data subset to the first image enhancement data set to obtain an image enhancement data set;
[0107] The text data enhancement processing module 103 is used to select a preset number of text data from the text data set in the training data set, perform back translation and sentence repetition operations on the preset number of text data to obtain a preset number of preprocessed text data, and add the preset number of preprocessed text data to the text data set to obtain a text enhancement data set.
[0108] In detail, each module described in the data enhancement device 100 of the graphic-text cross-modal model in the embodiment of the present invention is used in the same manner as described above. Figures 1 to 3 The data enhancement method of the cross-modal model of images and texts described in is the same technical means and can produce the same technical effects, so I will not go into details here.
[0109] like Figure 5 , is a schematic diagram of the structure of an electronic device for implementing a data enhancement method for a graphic-text cross-modal model provided by an embodiment of the present invention.
[0110] The electronic device 1 may include a processor 10, a memory 11, a communication bus 12, and a communication interface 13, and may also include a computer program stored in the memory 11 and executable on the processor 10, such as a data enhancement program for a graphic-text cross-modal model.
[0111] Among them, the processor 10 can be composed of an integrated circuit in some embodiments, for example, it can be composed of a single packaged integrated circuit, or it can be composed of multiple integrated circuits with the same function or different functions, including one or more central processing units (CPU), microprocessors, digital processing chips, graphics processors and various control chips. The processor 10 is the control core (ControlUnit) of the electronic device, and uses various interfaces and lines to connect various components of the entire electronic device. It runs or executes programs or modules stored in the memory 11 (for example, executes data enhancement programs for graphic cross-modal models, etc.), and calls data stored in the memory 11 to execute various functions of the electronic device and process data.
[0112] The memory 11 includes at least one type of readable storage medium, and the readable storage medium includes a flash memory, a mobile hard disk, a multimedia card, a card-type memory (for example, SD or DX memory, etc.), a magnetic memory, a disk, an optical disk, etc. In some embodiments, the memory 11 may be an internal storage unit of an electronic device, such as a mobile hard disk of the electronic device. In other embodiments, the memory 11 may also be an external storage device of an electronic device, such as a plug-in mobile hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the electronic device. Further, the memory 11 may also include both an internal storage unit of the electronic device and an external storage device. The memory 11 can not only be used to store application software and various types of data installed in the electronic device, such as the code of the data enhancement program of the graphic cross-modal model, but also can be used to temporarily store data that has been output or is to be output.
[0113] The communication bus 12 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. The bus is configured to realize connection and communication between the memory 11 and at least one processor 10, etc.
[0114] The communication interface 13 is used for communication between the above-mentioned electronic device and other devices, including a network interface and a user interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is generally used to establish a communication connection between the electronic device and other electronic devices. The user interface may be a display (Display), an input unit (such as a keyboard (Keyboard)), and optionally, the user interface may also be a standard wired interface, a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, and an OLED (Organic Light-Emitting Diode, organic light-emitting diode) touch device, etc. Among them, the display may also be appropriately referred to as a display screen or a display unit, which is used to display information processed in the electronic device and to display a visual user interface.
[0115] Figure 5 Only an electronic device with components is shown, and those skilled in the art will understand that Figure 5 The structure shown does not constitute a limitation on the electronic device 1, and may include fewer or more components than shown in the figure, or combine certain components, or arrange the components differently.
[0116] For example, although not shown, the electronic device may also include a power source (such as a battery) for supplying power to each component. Preferably, the power source may be logically connected to the at least one processor 10 through a power management device, so that the power management device can realize functions such as charging management, discharging management, and power consumption management. The power source may also include one or more DC or AC power sources, recharging devices, power failure detection circuits, power converters or inverters, power status indicators, and other arbitrary components. The electronic device may also include a variety of sensors, Bluetooth modules, Wi-Fi modules, etc., which will not be repeated here.
[0117] It should be understood that the embodiment is for illustration only and the scope of the patent application is not limited to this structure.
[0118] The data enhancement program of the graphic-text cross-modal model stored in the memory 11 of the electronic device 1 is a combination of multiple instructions. When running in the processor 10, it can achieve:
[0119] Acquire a training data set for a cross-modal image-text model, wherein the training data set includes an image data set and a text data set;
[0120] Classifying the image data set in the training data set according to business scenarios to obtain image categories, and performing image enhancement processing on the image data set based on the image categories to obtain a first image enhanced data set;
[0121] Dividing the first image enhancement data set into a plurality of preliminary image enhancement data subsets according to a preset rule, selecting a group of four images from each of the plurality of preliminary image enhancement data subsets, performing mosaic data enhancement processing on the four images to obtain a spliced image, and adding the spliced image of each preliminary image enhancement data subset to the first image enhancement data set to obtain an image enhancement data set;
[0122] A preset amount of text data is selected from the text data set in the training data set, back-translated and sentence repeated operations are performed on the preset amount of text data to obtain a preset amount of preprocessed text data, and the preset amount of preprocessed text data is added to the text data set to obtain a text enhancement data set.
[0123] Specifically, the specific implementation method of the processor 10 for the above instructions can refer to the description of the relevant steps in the corresponding embodiment of the accompanying drawings, which will not be repeated here.
[0124] Furthermore, if the module / unit integrated in the electronic device 1 is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. The computer-readable storage medium can be volatile or non-volatile. For example, the computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disk, a computer memory, and a read-only memory (ROM).
[0125] The present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor of an electronic device, the computer program can implement:
[0126] Acquire a training data set for a cross-modal image-text model, wherein the training data set includes an image data set and a text data set;
[0127] Classifying the image data set in the training data set according to business scenarios to obtain image categories, and performing image enhancement processing on the image data set based on the image categories to obtain a first image enhanced data set;
[0128] Dividing the first image enhancement data set into a plurality of preliminary image enhancement data subsets according to a preset rule, selecting a group of four images from each of the plurality of preliminary image enhancement data subsets, performing mosaic data enhancement processing on the four images to obtain a spliced image, and adding the spliced image of each preliminary image enhancement data subset to the first image enhancement data set to obtain an image enhancement data set;
[0129] A preset amount of text data is selected from the text data set in the training data set, back-translated and sentence repeated operations are performed on the preset amount of text data to obtain a preset amount of preprocessed text data, and the preset amount of preprocessed text data is added to the text data set to obtain a text enhancement data set.
[0130] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative, for example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation.
[0131] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0132] In addition, each functional module in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of hardware plus software functional modules.
[0133] It is obvious to those skilled in the art that the present invention is not limited to the details of the above exemplary embodiments, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0134] Therefore, no matter from which point of view, the embodiments should be regarded as illustrative and non-restrictive, and the scope of the present invention is limited by the appended claims rather than the above description, so it is intended that all changes falling within the meaning and scope of the equivalent elements of the claims are included in the present invention. Any attached figure mark in the claims should not be regarded as limiting the claims involved.
[0135] The blockchain referred to in this invention is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, encryption algorithm, etc. Blockchain is essentially a decentralized database, a string of data blocks generated by cryptographic methods. Each data block contains a batch of network transaction information, which is used to verify the validity of its information (anti-counterfeiting) and generate the next block. Blockchain can include the underlying blockchain platform, platform product service layer, and application service layer.
[0136] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.
[0137] In addition, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices stated in the system claim can also be implemented by one unit or device through software or hardware. The words first, second, etc. are used to indicate names, and do not indicate any particular order.
[0138] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the spirit and scope of the technical solution of the present invention.
Claims
1. A data enhancement method for a cross-modal model of images and texts, characterized in that: The method comprises: Acquire a training data set for a cross-modal image-text model, wherein the training data set includes an image data set and a text data set; Classifying the image data set in the training data set according to business scenarios to obtain image categories, and performing image enhancement processing on the image data set based on the image categories to obtain a first image enhanced data set; The first image enhancement data set is divided into a plurality of preliminary image enhancement data subsets according to a preset rule, a group of four images are selected from each of the plurality of preliminary image enhancement data subsets, mosaic data enhancement processing is performed on the four images to obtain a spliced image, and the spliced image of each preliminary image enhancement data subset is added to the first image enhancement data set to obtain an image enhancement data set; Selecting a preset amount of text data from the text data set in the training data set, performing back translation and sentence repeating operations on the preset amount of text data to obtain a preset amount of preprocessed text data, and adding the preset amount of preprocessed text data to the text data set to obtain a text enhancement data set; Among them, the back-translation and sentence repetition operations on the preset number of text data to obtain a preset number of pre-processed text data include: using a preset machine translation model to translate the preset number of text data into first language text data, and then translating the first language text data into original language text data; randomly selecting a preset number of words from the preset number of text data, and backfilling the preset number of words into the preset number of text data to obtain a preset number of first pre-processed text data; obtaining text data corresponding to each group of four randomly scaled images from the text data set, splicing the corresponding text data to obtain a plurality of second pre-processed text data; merging the original language text data, the preset number of first pre-processed text data and the plurality of second pre-processed text data to obtain a preset number of pre-processed text data.
2. The data enhancement method for the graphic-text cross-modal model according to claim 1, characterized in that: The step of performing image enhancement processing on the image dataset based on the image category to obtain a first image enhanced dataset includes: Based on the image category, an image enhancement processing algorithm is selected from a preset algorithm library, and image enhancement processing is performed on the image data set in the training data set in the spatial domain to obtain a grayscale image set; Performing Gaussian filtering on the grayscale image set in the frequency domain to obtain a smoothed grayscale image set; The brightness, contrast, saturation and hue of the smoothed grayscale image set are randomly changed to obtain a preliminary enhanced data set.
3. The data enhancement method for the graphic-text cross-modal model according to claim 2, characterized in that: The method of selecting an image enhancement processing algorithm from a preset algorithm library based on the image category, performing image enhancement processing on the image data set in the training data set in the spatial domain, and obtaining a grayscale image set includes: According to the image category, a grayscale algorithm is selected from a preset algorithm library to perform grayscale transformation on the image data set to obtain a preliminary grayscale image set; According to the image category, a sharpening algorithm is selected from a preset algorithm library, and the preliminary grayscale image set is sharpened to obtain a grayscale image set.
4. The data enhancement method for the graphic-text cross-modal model according to claim 2, characterized in that: The step of performing Gaussian filtering on the grayscale image set in the frequency domain to obtain a smooth grayscale image set includes: According to a preset rule, different weights are assigned to pixels at different positions of the grayscale image in the grayscale image set to obtain position pixel weights; Using a preset convolution template and based on the position pixel weights, weighted averaging is performed on pixels in the neighborhood of the grayscale image set to obtain a smoothed grayscale image set.
5. The data enhancement method for the graphic-text cross-modal model according to claim 1, characterized in that: The mosaic data enhancement process is performed on the four images to obtain a spliced image, including: Randomly scaling the four images respectively to obtain four randomly scaled images; Randomly select a stitching center coordinate in a preset area, stitch according to the stitching center coordinate, and stitch the four randomly scaled images into the preset area; When the four randomly scaled images exceed the preset area, the exceeded area is cropped to obtain a spliced image; When the four randomly scaled images do not fill up the preset area, the unfilled area is filled to obtain a spliced graphic.
6. The data enhancement method for the graphic-text cross-modal model according to claim 1, characterized in that: The classifying the image data set in the training data set according to the business scenario to obtain the image category includes: Extracting a feature vector set of an image data set in the training data set; Matching the feature vector set with a business scene picture set in a preset business scene library to obtain a matching similarity set; A matching similarity satisfying a similarity threshold is selected from the matching similarity set, and the business scene category marked by the business scene icon of the matching similarity satisfying the similarity threshold is used as the image category of the corresponding image data.
7. A data enhancement device for a cross-modal model of graphics and text, used to implement the data enhancement method for a cross-modal model of graphics and text as described in any one of claims 1 to 6, characterized in that: The device comprises: A training set acquisition module, used to acquire a training data set for the image-text cross-modal model, wherein the training data set includes an image data set and a text data set; an image enhancement processing module, configured to classify the image data set in the training data set according to the business scenario to obtain image categories, and based on the image categories, perform image enhancement processing on the image data set to obtain a first image enhancement data set; divide the first image enhancement data set into a plurality of preliminary image enhancement data subsets according to a preset rule, select a group of four images from the plurality of preliminary image enhancement data subsets, perform mosaic data enhancement processing on the four images to obtain a spliced image, and add the spliced image of each preliminary image enhancement data subset to the first image enhancement data set to obtain an image enhancement data set; The text data enhancement processing module is used to select a preset amount of text data from the text data set in the training data set, perform back translation and sentence repeat operations on the preset amount of text data to obtain a preset amount of preprocessed text data, and add the preset amount of preprocessed text data to the text data set to obtain a text enhancement data set.
8. An electronic device, characterized in that: The electronic device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the data enhancement method of the graphic-text cross-modal model as described in any one of claims 1 to 6.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the data enhancement method for the graphic-text cross-modal model is implemented.
Citation Information
Patent Citations
Cross-language description-oriented adversarial data enhancement method, system and storage medium
CN112819091A
Systems and methods for identifying non-compliant images using neural network architectures
US20210232620A1