Method and processor for processing graphics data
By extracting contextual features from images and text and then concatenating and encoding them, the problem of low efficiency in image and text data processing is solved, implicit alignment between images and text is achieved, processing efficiency is improved, and the application scope is expanded.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ALIBABA (CHINA) CO LTD
- Filing Date
- 2022-09-20
- Publication Date
- 2026-05-12
AI Technical Summary
Existing technologies tend to lose data when compressing images and text into global feature vectors, resulting in low efficiency in image and text data processing.
By extracting image context features and text context features from the original image and text, concatenating them and performing feature encoding, a target image or text is generated to achieve implicit alignment between the image and text and reduce the impact of data noise.
It improves the efficiency of image and text data processing, enables the generation of text from images or images from text, and expands the scope of applications.
Smart Images

Figure CN115563334B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing, and more specifically, to a method and processor for processing image and text data. Background Technology
[0002] Currently, information on the internet is mainly in the form of videos, pictures, and text, while most user requests are text-based. Therefore, establishing connections between different modalities is of great significance.
[0003] When establishing connections between different modalities, a feature model is typically learned to provide a unified feature representation for images and text. However, this unified feature representation requires compressing the images and text into a global feature vector. This compression process often results in significant data loss of both images and text, leading to inefficient image and text processing.
[0004] There is currently no effective solution to the above problems. Summary of the Invention
[0005] This invention provides a method and processor for processing graphic and textual data, thereby at least addressing the technical problem of low processing efficiency for graphic and textual data.
[0006] According to one aspect of the present invention, a method for processing image and text data is provided, comprising: acquiring an original image and original text to be processed, wherein the original image and original text are used to describe at least one identical object; extracting image context features from the original image and extracting text context features from the original text; concatenating the image context features and text context features to obtain an original feature vector; and performing feature encoding on the original feature vector to obtain target text corresponding to the original image and / or a target image corresponding to the original text, wherein the target text and / or the target image include the same object.
[0007] According to another aspect of the present invention, a method for processing image and text data is also provided, comprising: acquiring original image samples and original text samples to be processed, wherein the original image samples and original text samples are used to describe at least one identical object; extracting image context feature samples from the original image samples and extracting text context feature samples from the original text; concatenating the image context feature samples and text context feature samples to obtain original feature vector samples; outputting the original feature vector samples, wherein the text used to describe the original image samples obtained by feature encoding the original feature vector samples, and / or the image used to describe the original text samples, are used as training samples to train an image and text processing model; wherein the image and text processing model is used to convert an input image into target text used to describe the input image, the target text including the identical object, and / or the image and text processing model is used to convert the input text into a target image used to describe the input text, the target image including the identical object.
[0008] According to another aspect of the present invention, another method for processing graphic and text data is also provided, comprising: displaying an original image and original text to be processed on the presentation screen of a virtual reality (VR) device or an augmented reality (AR) device, wherein the original image and original text are used to describe at least one identical object; the VR device or AR device extracts image context features from the original image and extracts text context features from the original text; after concatenating the image context features and text context features to obtain an original feature vector, driving the VR device or AR device to render and display target text corresponding to the original image obtained by feature encoding the original feature vector, the target text including the same object; and / or driving the VR device or AR device to render and display target image corresponding to the original text obtained by feature encoding the original feature vector, the target image including the same object.
[0009] According to another aspect of the present invention, another method for processing image and text data is also provided, comprising: obtaining an original image and original text to be processed by calling a first interface, wherein the original image and original text are used to describe at least one identical object; extracting image context features from the original image and extracting text context features from the original text, wherein the image context features include image features in different image regions of the original image, and the text context features include text features at different text positions in the original text; concatenating the image context features and the text context features to obtain an original feature vector, wherein the features in the original feature vector are image context features or text context features; performing feature encoding on the original feature vector to obtain target text corresponding to the original image, wherein the target text includes the same object; and / or performing feature encoding on the original feature vector to obtain a target image corresponding to the original text, wherein the target image includes the same object; and outputting the target text and / or the target image by calling a second interface, wherein the second interface includes a second parameter, the value of which is the target text and / or the target image.
[0010] According to one aspect of the present invention, an apparatus for processing image and text data is provided, comprising: a first acquisition unit, configured to acquire an original image and original text to be processed, wherein the original image and original text are used to describe at least one identical object; a first extraction unit, configured to extract image context features from the original image and extract text context features from the original text; a first processing unit, configured to concatenate the image context features and the text context features to obtain an original feature vector; and a second processing unit, configured to perform feature encoding on the original feature vector to obtain target text corresponding to the original image and / or a target image corresponding to the original text, wherein the target text and / or the target image include the same object.
[0011] According to another aspect of the present invention, a processing apparatus for image and text data is also provided, comprising: a second acquisition unit, configured to acquire original image samples and original text samples to be processed, wherein the original image samples and original text samples are used to describe at least one identical object; a second extraction unit, configured to extract image context feature samples from the original image samples and extract text context feature samples from the original text; a third processing unit, configured to concatenate the image context feature samples and the text context feature samples to obtain original feature vector samples; and a first output unit, configured to output the original feature vector samples, wherein text used to describe the original image samples and / or images used to describe the original text samples, obtained by feature encoding of the original feature vector samples, are used as training samples to train an image and text processing model; wherein the image and text processing model is used to convert an input image into target text used to describe the input image, the target text including identical objects, and / or the image and text processing model is used to convert the input text into target images used to describe the input text, the target images including identical objects.
[0012] According to another aspect of the present invention, another image and text data processing apparatus is also provided, comprising: a presentation unit, configured to display an original image and original text to be processed on a presentation screen of a virtual reality (VR) device or an augmented reality (AR) device, wherein the original image and original text are used to describe at least one identical object; a third extraction unit, configured to extract image context features from the original image and extract text context features from the original text by the VR device or AR device; and a driving unit, configured to, after concatenating the image context features and text context features to obtain an original feature vector, drive the VR device or AR device to render and display target text corresponding to the original image obtained by feature encoding the original feature vector, wherein the target text includes the same object, and / or drive the VR device or AR device to render and display target image corresponding to the original text obtained by feature encoding the original feature vector, wherein the target image includes the same object.
[0013] According to another aspect of the present invention, another image and text data processing apparatus is also provided, comprising: a third acquisition unit, configured to acquire an original image and original text to be processed by calling a first interface, wherein the original image and original text are used to describe at least one identical object; a fourth extraction unit, configured to extract image context features from the original image and extract text context features from the original text, wherein the image context features include image features in different image regions of the original image, and the text context features include text features at different text positions in the original text; a fourth processing unit, configured to concatenate the image context features and the text context features to obtain an original feature vector, wherein the features in the original feature vector are image context features or text context features; a fifth processing unit, configured to perform feature encoding on the original feature vector to obtain target text corresponding to the original image, wherein the target text includes the same object; and / or, perform feature encoding on the original feature vector to obtain a target image corresponding to the original text, wherein the target image includes the same object; and a second output unit, configured to output the target text and / or the target image by calling a second interface, wherein the second interface includes a second parameter, the value of which is the target text and / or the target image.
[0014] According to another aspect of the present invention, a computer-readable storage medium is also provided, the computer-readable storage medium including a stored program, wherein, when the program is running, it controls the device where the storage medium is located to execute the image and text data processing method of any of the above-mentioned methods.
[0015] According to another aspect of the present invention, a processor is also provided, which is used to run a program, wherein the image and text data processing method described above is executed during program execution.
[0016] In this embodiment of the invention, an original image and original text to be processed are obtained, wherein the original image and original text are used to describe at least one identical object; image context features are extracted from the original image, and text context features are extracted from the original text; the image context features and text context features are concatenated to obtain an original feature vector; the original feature vector is feature-encoded to obtain target text corresponding to the original image, and / or target image corresponding to the original text, wherein the target text and / or target image include the same object. In other words, this embodiment of the invention obtains the image context features of the original image and the text context features of the original text, concatenates the image context features and text context features for unified feature encoding processing, thereby implicitly aligning the original image and original text, reducing the impact of data noise, and realizing the ability to generate text from images or images from text, thus achieving the technical effect of improving the processing efficiency of image and text data and solving the technical problem of low processing efficiency of image and text data. Attached Figure Description
[0017] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this invention, illustrate exemplary embodiments of the invention and are used to explain the invention, but do not constitute an undue limitation of the invention. In the drawings:
[0018] Figure 1 This is a hardware structure block diagram of a computer terminal (or mobile device) for a method of processing graphic data according to an embodiment of the present invention.
[0019] Figure 2 This is a flowchart of a method for processing graphic data according to an embodiment of the present invention;
[0020] Figure 3 This is a flowchart of another method for processing graphic and textual data according to an embodiment of the present invention;
[0021] Figure 4 This is a schematic diagram of the hardware environment of a virtual reality device according to an embodiment of the present invention for processing graphic data;
[0022] Figure 5 This is a flowchart of another method for processing graphic and textual data according to an embodiment of the present invention;
[0023] Figure 6 This is a schematic diagram of the processing result of a method for processing graphic data according to an embodiment of the present invention;
[0024] Figure 7 This is a flowchart of another method for processing graphic and textual data according to an embodiment of the present invention;
[0025] Figure 8 This is a schematic diagram of image processing by a computer device according to an embodiment of the present invention;
[0026] Figure 9 This is a schematic diagram illustrating the processing of graphic and textual data according to an embodiment of the present invention;
[0027] Figure 10 This is a structural block diagram of a computing environment according to an embodiment of the present invention;
[0028] Figure 11 This is a structural block diagram of a service grid for a method of processing graphic data according to an embodiment of the present invention;
[0029] Figure 12 This is a schematic diagram of a graphic data processing apparatus according to an embodiment of the present invention;
[0030] Figure 13 This is a schematic diagram of another image and text data processing device according to an embodiment of the present invention;
[0031] Figure 14 This is a schematic diagram of another image and text data processing device according to an embodiment of the present invention;
[0032] Figure 15 This is a schematic diagram of another image and text data processing device according to an embodiment of the present invention;
[0033] Figure 16 This is a structural block diagram of a computer terminal according to an embodiment of the present invention. Detailed Implementation
[0034] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0035] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0036] First, some nouns or terms that appear in the description of the embodiments of the present invention shall be interpreted as follows:
[0037] Multimodal feature learning can be used to perform machine learning on large amounts of image and text data, and train a feature encoding model (such as a neural network). This model can map user-input image or text data into a feature vector, and can be used as a pre-trained model for tasks such as image recognition, image and text retrieval, and image description generation.
[0038] The feature encoding model (Transformer) is a model that can be used to extract features from input images or text;
[0039] InfoNoise Contrastive Estimation (InfoNCE) is a loss function that trains a feature model using the similarity between positive sample pairs and the similarity between negative sample pairs.
[0040] Natural language image captioning is a key technology for multimodal feature learning and processing. It can be used to perform multimodal transformation from images to text, and to generate corresponding text descriptions from input images.
[0041] Example 1
[0042] According to an embodiment of the present invention, an embodiment of a method for processing graphic and textual data is also provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0043] The method embodiment provided in Embodiment 1 of the present invention can be executed in a mobile terminal, computer terminal or similar computing device. Figure 1 This is a hardware structure block diagram of a computer terminal (or mobile device) according to an embodiment of the present invention for processing graphic data. Figure 1 As shown, the computer terminal 10 (or mobile device 10) may include one or more processors 102 (shown as 102a, 102b, ..., 102n in the figure) (processor 102 may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.), a memory 104 for storing data, and a transmission module 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of a BUS bus), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, computer terminal 10 may also include... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.
[0044] It should be noted that the aforementioned one or more processors 102 and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within the computer terminal 10 (or mobile device). As involved in the embodiments of the present invention, the data processing circuits serve as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).
[0045] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the image and text data processing method in this embodiment of the invention. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, thereby implementing the above-mentioned application vulnerability detection method. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the computer terminal 10 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0046] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the computer terminal 10. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module, used for wireless communication with the Internet.
[0047] The display can be, for example, a touchscreen liquid crystal display (LCD) that allows the user to interact with the user interface of the computer terminal 10 (or mobile device).
[0048] It should be noted here that, in some optional embodiments, the above... Figure 1 The computer device (or mobile device) shown may include hardware elements (including circuitry), software elements (including computer code stored on a computer-readable medium), or a combination of both hardware and software elements. It should be noted that... Figure 1This is only one instance of a specific particular instance and is intended to illustrate the types of components that may exist in the aforementioned computer device (or mobile device).
[0049] exist Figure 1 In the operating environment shown, this invention provides a method for application on the active protection side, such as... Figure 2 The illustrated method for processing text and image data. It should be noted that the text and image data processing method in this embodiment can be derived from... Figure 1 The mobile terminal in the illustrated embodiment is executed.
[0050] Figure 2 This is a flowchart of a method for processing graphic and textual data according to an embodiment of the present invention. Figure 2 As shown, the method may include the following steps:
[0051] Step S202: Obtain the original image and original text to be processed, wherein the original image and original text are used to describe at least one identical object.
[0052] In the technical solution provided by step S202 of the present invention, the original image and original text to be processed can be obtained. The original image and original text can be used to describe at least one identical object. For example, at least one object in the original image is the same as the object described in the original text. For example, if the original image is two dogs running in a field, the original text can be "two dogs running in a field". The original image can be an artistic image, material, or comic, etc. This is only an example and no specific limitation is made on the source and type of the original image. The original text can be text entered by the user in the search engine, which can include characters in the text, etc. This is only an example and no specific limitation is made on the source and type of the original text.
[0053] For example, the original image and original text to be processed can be obtained. The original image and original text can be user input, or they can be pre-set query information obtained when a certain trigger condition is met. This is just an example and no specific restrictions are placed on the way the original image and original text are obtained.
[0054] Step S204: Extract image context features from the original image and extract text context features from the original text.
[0055] In the technical solution provided by step S204 of the present invention, the image context features include image features in different image regions of the original image, and the text context features include text features at different text positions in the original text.
[0056] In this embodiment, image context features can be extracted from the acquired original image, and text context features can be extracted from the original text. The image context features can be image context feature vectors, which can be used to represent objects in the image. They can be obtained by combining image features from different regions in the original image. For example, if the original image has "clouds" drawn in the upper left corner, "puppy" drawn in the lower right corner, and "flowers" drawn in the middle, then image context features representing "clouds," "puppy," and "flowers" can be obtained. The text context features can be text context feature vectors, which can be used to represent objects described in the original text. They can be obtained by combining text features from different text positions in the original text. For example, if the original text is "two dogs running in a field," then text context features representing "dogs," "running," and "field" can be obtained.
[0057] In related technologies, only the overall features of the original image and original text are extracted to obtain global features. However, processing only global features can easily lead to the problem of missing information in the original image or original text. In contrast, the embodiments of the present invention can train the ability to extract contextual features of each region in the original image, obtaining image contextual features and text contextual features. This allows for multimodal feature learning for multiple regions, enabling more accurate identification of the original image and original text, thereby improving the accuracy of prediction of the original image or original text.
[0058] Optionally, feature vectors can be extracted from different regions of the original image. Multiple image feature vectors extracted from an original image can be combined to obtain the image context feature vector of the original image. For example, feature vectors from different regions of the original image can be extracted to obtain a feature vector that describes the sky background of the original image, a feature vector that describes the ground in the original image, and a feature vector that describes the color in the original image. All the obtained feature vectors can be combined to obtain the image context features of the original image.
[0059] Optionally, the text feature vectors corresponding to different positions in the sentence sequence of the original text can be combined. In the combination, a certain feature vector may describe the subject noun of the sentence, a certain feature vector may describe the action, a certain feature vector may describe the state, etc. All the obtained feature vectors are combined together to obtain the text context features of the original text.
[0060] For example, the original image mentioned above can be a raw picture. Before inputting the raw image, it can be divided into multiple non-overlapping patches, such as 10*10 non-overlapping patches. Each patch can be used to represent an image token. Similarly, before inputting the raw text, it can be divided into multiple text patches (tokens). Each text token can be approximated as a word. The feature vectors of each image patch and each text patch are determined to obtain the image context features and text context features.
[0061] Step S206: Concatenate the image context features and the text context features to obtain the original feature vector.
[0062] In the technical solution provided by step S206 of the present invention, the features in the original feature vector are image context features or text context features.
[0063] In this embodiment, image context features and text context features are obtained, and the image context features and text context features can be concatenated to obtain the original feature vector.
[0064] Optionally, the image context features output by the image encoder and the text context features output by the text encoder can be obtained, and the image context feature vector and the text context feature vector can be uniformly concatenated to obtain the original feature vector.
[0065] For example, the M image context features of the original image and the N text context features of the original text can be concatenated together in the order of image context features first and text context features last to obtain the original feature vector containing M+N context features.
[0066] In related technologies, only the text context feature vector is input into the encoder for processing, while the image context feature vector is only used as an additional reference. Alternatively, the image or text information is predicted by weighted summation of the two vectors when encoding the text context feature vector. However, in this embodiment of the invention, the image context feature vector and the text context feature vector are uniformly concatenated, and the concatenated feature vector is uniformly processed to generate the corresponding image or text description. The text context feature vector and the image context feature vector are not treated differently, thereby achieving the technical effect of improving the processing efficiency of image and text data and solving the technical problem of low processing efficiency of image and text data.
[0067] Step S208: Perform feature encoding on the original feature vector to obtain the target text corresponding to the original image, and / or the target image corresponding to the original text, wherein the target text and / or the target image include the same object.
[0068] In the technical solution provided by step S208 of the present invention, the original feature vector can be processed by feature encoding. The original feature vector that has been encoded can be processed by feature mapping, and then the mapped feature vector can be processed to obtain the target text corresponding to the original image, or to obtain the target image corresponding to the original text. The object described by the target text may include the same object described by the original text in the original image; the object described by the target image may include the same object described by the original text in the original image.
[0069] Optionally, the original feature vector obtained after concatenation can be used for feature encoding and layer-by-layer mapping. For example, in the first layer of mapping, the similarity between the first feature vector and other feature vectors in the original feature vector can be determined. The two vectors with the highest similarity can be weighted and averaged to obtain the updated first feature vector. It should be noted that the mapping method here is only for illustration and does not impose specific restrictions on the mapping method. By obtaining new feature vectors through layer-by-layer mapping, the new feature vectors can be decoded to obtain the target text corresponding to the original image and / or the target image corresponding to the original text.
[0070] For example, concatenating M image context feature vectors and N text context feature vectors yields an original feature vector containing M+N context features. An encoder (Transformer encoder) can then be used to map the original feature vector layer by layer to obtain the target text corresponding to the original image, and / or the target image corresponding to the original text.
[0071] This invention, through feature encoding of the original feature vector, obtains the target text corresponding to the original image and / or the target image corresponding to the original text. This enables the generation of text from images and images from text, thus expanding the scope of application of this invention beyond the task of searching for text from images or searching for images from text.
[0072] Through steps S202 to S208 of the present invention, an original image and original text to be processed are obtained to describe at least one identical object. Image context features containing different image regions are extracted from the original image, and context features containing different text positions in the original text are extracted from the original text. The image context features and text context features are concatenated to obtain an original feature vector. Feature encoding is performed on the original feature vector to obtain the target text corresponding to the original image and / or the target image corresponding to the original text, wherein the target text and / or the target image include the same object. By uniformly performing feature encoding processing on the original feature vector, the influence of data noise is reduced, thereby achieving the technical effect of improving the processing efficiency of image and text data and solving the technical problem of low processing efficiency of image and text data.
[0073] The method described in this embodiment will be further described below.
[0074] As an optional implementation, step S208, which involves feature encoding the original feature vector to obtain target text corresponding to the original image and / or target image corresponding to the original text, includes: mapping the original feature vector to a target feature vector, wherein the target feature vector includes an image feature vector and a text feature vector, wherein the similarity between the object represented by the features in the image feature vector and the object represented by the text context features is greater than a first similarity threshold, and the similarity between the object represented by the features in the text feature vector and the object represented by the image context features is greater than a second similarity threshold; generating the image feature vector into a target image corresponding to the original text, and / or generating the text feature vector into target text corresponding to the original image.
[0075] In this embodiment, the original feature vector can be mapped to obtain a target feature vector including text feature vectors and image feature vectors. The target feature vector can be processed using a decoder to generate a target image corresponding to the original text from the image feature vector in the target feature vector, and / or to generate target text corresponding to the original image from the text feature vector in the target feature vector. The similarity between the object represented by the feature in the image feature vector and the object represented by the text context feature can be greater than a first similarity threshold; the similarity between the object represented by the feature in the text feature vector and the object represented by the image context feature can be greater than a second similarity threshold; the first similarity threshold and the second similarity threshold can be preset values, and the magnitudes of the first similarity threshold and the second similarity threshold can be the same or different.
[0076] Alternatively, an encoder (Transformer encoder) can be used to map the concatenated original feature vector layer by layer. Each layer can be used to determine the similarity between each vector and other vectors, and vectors with high similarity are processed to finally obtain image feature vectors with a similarity greater than the similarity threshold with the text context features.
[0077] For example, we can obtain an original feature vector containing M+N contextual features after concatenation. This original feature vector can then be mapped layer by layer using a Transformer encoder. The first M feature vectors are mapped to the target feature vector through the last layer, resulting in M image feature vectors. This layer-by-layer mapping reduces the distance between the M image feature vectors and the corresponding M image feature vectors from the original image segmentation. Similarly, the last N text feature vectors are mapped to the target feature vector through the last layer, resulting in N text feature vectors representing the text. This layer-by-layer mapping reduces the distance between the N text feature vectors and the vectors corresponding to the N original text words, ultimately achieving implicit alignment between the image feature vectors and the text feature vectors. During model training, the model parameters are adjusted based on the target feature vectors, thereby improving the model's prediction accuracy.
[0078] As an optional implementation, the original feature vector includes multiple features. Mapping the original feature vector to the target feature vector includes: comparing the similarity of each feature among the multiple features with the features in the original feature vector other than each feature to obtain the features in the image feature vector or the text feature vector corresponding to each feature; and generating the target feature vector based on the features in the image feature vector or the text feature vector corresponding to each feature.
[0079] In this embodiment, the original feature vector includes multiple features. Each feature can be compared with the features other than each feature in the original feature vector to obtain the features in the image feature vector or text feature vector corresponding to each feature. The target feature vector can be generated based on the features in the image feature vector or text feature vector corresponding to each feature.
[0080] In the process of generating text from the original image, it is necessary to align the image context features in the original image with the corresponding text context features. In related technologies, the nth image context feature and the nth text context feature are usually forcibly aligned one by one to achieve the purpose of aligning the image context features with the text context features. However, when there is noisy data, for example, when there is no feature vector in the image to represent "car" but there is a feature vector in the text to represent "car", the noisy data cannot be aligned successfully. This method is prone to interfering with the alignment results, resulting in inaccurate prediction or retrieval results of the model.
[0081] In this embodiment of the invention, instead of forcibly aligning the nth image context feature and the nth text context feature one by one, the similarity between each feature in the original feature vector and all other features in the original feature vector is determined. The similarity between features can be determined by comparing the objects represented by the features, and the features in the image feature vector or text feature vector corresponding to each feature are obtained. The target feature vector can be generated based on the features in the image feature vector or text feature vector corresponding to each feature. That is, if the label is incorrect, it is equivalent to not participating in feature training, thereby achieving the purpose of implicitly determining the correspondence between text and image, and completing the implicit alignment of the semantic content corresponding to the image context features and text context features. In this way, the efficiency of the model in processing image and text data can be improved, thereby reducing the impact of noisy data.
[0082] As an optional implementation, the features in the original feature vector are compared with the features in the original feature vector excluding the feature itself to obtain features with a similarity greater than a similarity threshold. This includes: in the feature encoder, comparing the features in the original feature vector with the features in the original feature vector excluding the feature itself to obtain features with a similarity greater than a similarity threshold, and generating a target feature vector based on the features with a similarity greater than the similarity threshold. The feature encoder is used to convert the input data into a feature vector.
[0083] In this embodiment, the original feature vector can be obtained by a feature encoder. In the feature encoder, the features in the original feature vector can be compared with the features in the original feature vector except for the features themselves. For example, the comparison can be performed on the object represented by each feature to complete the similarity comparison between each feature and the features in the original feature vector except for each feature. The feature with a similarity greater than the similarity threshold is obtained, and the target feature vector is generated. The feature encoder can be a self-attention mechanism encoder or other encoders that can perform multimodal conversion, and all such encoders should be within the protection scope of this invention. No specific restrictions are made on the type of encoder here. The feature encoder is used to convert the input data into a feature vector. For example, the original feature vector can be converted into a new feature vector.
[0084] Optionally, image context features and text context features can be concatenated to obtain an original feature vector. The obtained original feature vector can be input into a feature encoder. In the feature encoder, the features in the original feature vector are compared with the features in the original feature vector except for the image features. The image features and text features that describe the same object have a high similarity. For example, if the similarity between two features is 100%, it can be determined that the two features correspond to the same object. Based on the features with a similarity greater than the similarity threshold, a target feature vector is generated. In this embodiment of the invention, the image context features and text context features are treated uniformly by the feature encoder to obtain the target feature vector.
[0085] For example, a feature encoder can perform layer-by-layer mapping on the original feature vector containing M+N contextual features. After L layers of mapping, a new target feature vector containing M+N contextual features can be generated. The first M image contextual features are mapped to M feature vectors representing the image through the last layer of mapping, and the last N text contextual features are mapped to N feature vectors representing the text through the last layer of mapping. The target feature vector is composed of the M image feature vectors representing the image and the N text feature vectors representing the text.
[0086] As an optional implementation, generating a target image corresponding to the original text from the image feature vector in the target feature vector includes: generating a target image corresponding to the original text from the image feature vector in the target feature vector based on an image generation model, wherein the parameters of the image generation model are adjusted by a loss function between the image feature vector and the original image feature vector of the original image, and the image generation model is a machine learning model; and / or generating target text corresponding to the original image from the text feature vector in the target feature vector based on a text generation model, wherein the parameters of the text generation model are adjusted by a loss function between the text feature vector and the original text feature vector of the original text, and the text generation model is a machine learning model.
[0087] In this embodiment, the image feature vector in the target feature vector can be processed by an image generation model to obtain a target image corresponding to the original text. The image generation model can be used to generate an image from text information. The parameters of the image generation model can be adjusted by a loss function between the image feature vector and the original image feature vector. The loss function can be cross-entropy loss, logistic regression loss function, etc., and no specific restrictions are placed on the type of loss function here. The original image feature vector can be a vector obtained by mapping the pixel values of the original image, and can be used to represent the content in the original image. No specific restrictions are placed on the method of determining the original image here. The image generation model can be a machine learning model, such as a neural network model, linear regression model, etc. This is only for illustrative purposes and no specific restrictions are placed on the type of machine learning model.
[0088] Optionally, the original image can be segmented to obtain original image blocks, and the pixel values of the original image blocks or the original image blocks can be mapped to obtain original image vectors. The image generation model processes the image feature vector in the target feature vector to generate a target image corresponding to the original text. The loss can be calculated based on the image feature vector in the target feature vector and the original image feature vector. The parameters of the image generation model can be adjusted based on the calculation result of the loss function to reduce the distance between the image feature vector and the original image feature vector, thereby improving the accuracy of the image generation model in generating images from text.
[0089] Optionally, the application scenarios for image-to-text conversion can include: helping blind people describe the current actual scene in words; assisting image retrieval, such as better searching image content using text; and being used in children's education, such as generating language descriptions from images in books to help children understand the world. The application scenarios here are only illustrative and do not impose specific limitations on the application scenarios.
[0090] In this embodiment, the text feature vector in the target feature vector can be processed by a text generation model to obtain the target text corresponding to the original image. The text generation model can be used to generate text from image information. The parameters of the text generation model can be adjusted by a loss function between the text feature vector and the original text feature vector. The loss function can be cross-entropy loss, logistic regression loss function, etc., and no specific restrictions are placed on the type of loss function here. The original text feature vector can be a vector corresponding to each word or phrase predefined in a dictionary. For example, if the dictionary records the vector x1 corresponding to the word w1, then the original text word or vector can be obtained by looking up the dictionary table. No specific restrictions are placed on the method of determining the original text here. The text generation model can be a machine learning model, such as a neural network model, linear regression model, etc. This is only for illustrative purposes and no specific restrictions are placed on the type of machine learning model.
[0091] Optionally, the original text can be segmented into multiple words / phrases / characters, and the vectors corresponding to the segmented text can be determined to obtain the original text vector. The text generation model processes the text feature vector in the target feature vector to generate the target text corresponding to the original image. The loss can be calculated based on the text feature vector in the target feature vector and the original text feature vector. The parameters of the text generation model can be adjusted based on the calculation result of the loss function to reduce the distance between the text feature vector and the original text feature vector, thereby improving the accuracy of the text generation model in generating text from the image.
[0092] Optionally, the application scenarios for text-to-image generation can include: in the process of artistic creation, helping to generate new paintings, materials or comics, or advertising design, etc.; it can be used for image editing to help replace and generate new backgrounds or foregrounds, etc.; it can be used for suspect portrait generation to help the police track down suspects, etc. For example, a suspect portrait can be generated based on the description of the suspect, etc. The application scenarios here are only examples and do not impose specific limitations on the application scenarios.
[0093] Since the process of retrieving images or text requires extracting feature vectors from the images and text, and then determining the distance between the image feature vector and the original image feature vector, and the distance between the text feature vector and the original text feature vector, in order to generate text from images or images from text, this embodiment of the invention can extract image context features and original image feature vectors from the original image through an image encoder, and extract text context features and original text feature vectors from the original text through a text encoder. Based on the extracted image context features and original image feature vectors, the parameters of the image generation model are adjusted, and based on the extracted text context features and original text feature vectors, the parameters of the text generation model are adjusted, thereby improving the accuracy of image or text prediction in this embodiment of the invention.
[0094] As an optional implementation, step S206 involves concatenating the image context features and the text context features to obtain the original feature vector, including: connecting the image context features and the text context features in the target order to obtain the original feature vector.
[0095] In this embodiment, the image context features and text context features can be concatenated according to the target order to obtain the original feature vector. The target order can be a pre-defined order, such as an image first and text last.
[0096] Optionally, the image context feature vector and the text context feature vector can be concatenated in the order of image context feature vector first and text context feature vector last to obtain the original feature vector. For example, M image context feature vectors and N text context feature vectors can be concatenated in the order of image first and text last to obtain M+N original feature vectors.
[0097] In this embodiment, by uniformly concatenating the image context feature vector and the text context feature vector, the problem of image and text information loss caused by only obtaining the context feature vector of the original image or the context feature vector of the original text in related technologies is solved. Furthermore, this embodiment of the invention concatenates the image context features and the text context features to obtain the original feature vector, and processes the original feature vector, thereby reducing the impact of data noise and improving the effect of feature training.
[0098] As an optional implementation, step S204, extracting image context features from the original image and extracting text context features from the original text, includes: extracting image context features from the original image based on an image feature extraction model and extracting text context features from the original text based on a text feature extraction model.
[0099] In this embodiment, the original image can be processed based on an image feature extraction model to obtain image context features; the original text can be processed based on a text feature extraction model to obtain text context features. The image feature extraction model can be an image encoder; the text feature extraction model can be a text encoder.
[0100] For example, an image encoder (which can be a Transformer) can be used to process the original image to obtain an image context feature vector, and a text encoder (which can be a Transformer) can be used to process the original text to obtain a text context feature vector.
[0101] As an optional implementation, global image features are extracted from the original image based on an image feature extraction model, wherein the parameters of the image processing model are adjusted by a loss function between the image feature vector and the original image feature vector of the original image; the output text with the highest similarity to the original image is retrieved from a text database based on the global image features; and / or, global text features are extracted from the original text based on a text feature extraction model, wherein the parameters of the text feature extraction model are adjusted by a loss function between the text feature vector and the original text feature vector of the original text, and the output image with the highest similarity to the original text is retrieved from the image database based on the global text features.
[0102] In this embodiment, for the input original image and original text, global features can be extracted from the original image using an image feature extraction model to obtain global image features, and global features can be extracted from the original text using a text encoder to obtain global text features. The parameters of the image feature extraction model and the text feature extraction model can be adjusted based on the loss function between the global image features and the global text features, such as a contrastive learning loss function or a cross-entropy loss function. No specific restrictions are placed on the type of loss function here. The global image features can be used to retrieve the output text with the highest similarity to the original image; the global text features are used to retrieve the output image with the highest similarity to the original text.
[0103] Optionally, the parameters of the image feature extraction model can be adjusted using a loss function to reduce the distance between the image feature vector and the correct text feature vector, and increase the distance between the image feature vector and the incorrect text feature vector. For example, the contrastive learning loss function (InfoNCE) can be used to reduce the distance between the image feature vector and the correct text feature vector, thereby improving the accuracy of the model in the prediction process.
[0104] Optionally, the loss function can be in the form of positive sample distance divided by negative sample distance, where the positive sample distance in the numerator is the Euclidean distance between the image feature vector and the corresponding correct text feature vector, and the negative sample distance in the denominator is the sum of the Euclidean distances between the image feature vector and all other text feature vectors. During training, the loss function can be minimized to reduce the positive sample distance and increase the negative sample distance.
[0105] It should be noted that the distances between positive and negative samples mentioned above can be Euclidean distances, cosine distances, etc., and no specific restrictions are placed on the form of the distances here.
[0106] In related technologies, image feature vectors are extracted using an image encoder, and text feature vectors are extracted using a text encoder. However, the image feature vectors extracted by the image encoder are only global image feature vectors and cannot describe detailed information about different regions of the image. Similarly, the text feature vectors extracted by the text encoder are only global text feature vectors and can only be used to describe the general meaning of the entire sentence. The above methods are relatively coarse in terms of information extraction and have inaccuracy issues. To solve the above problems, this invention extracts contextual feature vectors from both the image and the text, thereby achieving the goal of more accurately determining detailed information about different regions of the image.
[0107] In this embodiment of the invention, by training the model's ability to extract contextual features, the model can learn multimodal features for each region in the image, thereby more accurately determining global image features and text features.
[0108] In this embodiment of the invention, the image context features of the original image and the text context features of the original text are obtained, the image context features and the text context features are concatenated and feature encoding is performed, thereby implicitly aligning the original image and the original text and processing them uniformly, thereby reducing the impact of data noise, realizing the ability to generate text from images or images from text, and thus achieving the technical effect of improving the processing efficiency of image and text data, solving the technical problem of low processing efficiency of image and text data.
[0109] The following describes the method for processing image and text data used for model training according to an embodiment of the present invention.
[0110] Figure 3 This is a flowchart of another method for processing graphic and textual data according to an embodiment of the present invention. Figure 3 As shown, the method may include the following steps:
[0111] Step S302: Obtain the original image sample and the original text sample to be processed, wherein the original image sample and the original text sample are used to describe at least one identical object.
[0112] In the technical solution provided by step S302 of the present invention, original image samples and original text samples to be processed can be obtained. The original image samples and original text samples can be training samples for model training, or pre-provided samples. They can be used to describe at least one identical object. For example, at least one object in the original image sample is the same as the object described in the original text sample. For example, if the original image sample is two dogs running in a field, then the original text sample can be "two dogs running in a field". The original image sample can be an artistically created image, or it can be a source material or a comic, etc. This is only an example and no specific limitation is made on the source and type of the original image sample. The original text sample can be text entered by the user in a search engine, which can include characters in the text, etc. This is only an example and no specific limitation is made on the source and type of the original text.
[0113] Step S304: Extract image context feature samples from the original image samples and extract text context feature samples from the original text.
[0114] In the technical solution provided by step S304 of the present invention, the image context feature sample includes image features in different image regions of the original image sample, and the text context feature sample includes text features at different text positions in the original text sample.
[0115] Step S306: Concatenate the image context feature samples and the text context feature samples to obtain the original feature vector samples.
[0116] In the technical solution provided by step S306 of the present invention, the features in the original feature vector sample are image context feature samples or text context feature samples.
[0117] Step S308: Output the original feature vector sample, wherein the text used to describe the original image sample obtained by feature encoding the original feature vector sample, and / or the image used to describe the original text sample, are used as training samples to train the image-text processing model; wherein the image-text processing model is used to convert the input image into target text used to describe the input image, the target text including the same object, and / or the image-text processing model is used to convert the input text into target image used to describe the input text, the target image including the same object.
[0118] In the technical solution provided by step S308 of the present invention, the original feature vector samples are output as training samples to train the model and obtain the image and text processing model. The image and text processing model can be an image and text multimodal feature model, which can be used to convert the input image into target text for describing the input image, and / or convert the input text into target image for describing the input text.
[0119] Optionally, the original feature vector samples can be feature-encoded to obtain text describing the original image samples, and / or images describing the original text samples. The model parameters can be adjusted based on the obtained text of the original image samples and the original text samples, and / or based on the obtained images describing the original text samples and the original image samples, to obtain an image-text processing model.
[0120] The method described in this embodiment will be further described below.
[0121] As an optional implementation, the image to be retrieved is obtained; the image to be retrieved is converted into the corresponding target text based on the image-text processing model, and global image features are extracted from the image to be retrieved based on the target text corresponding to the image to be retrieved; based on the global image features, the output text with the highest similarity to the image to be retrieved is retrieved from the text database.
[0122] In this embodiment, an image to be retrieved can be acquired, and the image to be retrieved can be converted into corresponding target text based on an image-text processing model. Global image features can be extracted from the image to be retrieved based on the target text corresponding to the image to be retrieved. The output text with the highest similarity to the image to be retrieved can be retrieved from the text database based on the global image features, thus obtaining the text used to describe the original image sample. The text database can be a pre-acquired database, such as a pre-stored database obtained from the Internet. No specific restrictions are placed on the acquisition method and location of the text database here.
[0123] Optionally, in the scenario of searching for text using images, the image to be retrieved can be obtained, and the image to be retrieved can be converted into the corresponding target text based on the image-text processing model. Furthermore, global image features can be extracted from the image to be retrieved based on the target text corresponding to the image to be retrieved. Based on the global image features, the output text with the highest similarity among the images to be retrieved can be retrieved from the text database to obtain the text used to describe the original image sample.
[0124] As an optional implementation, the text to be retrieved is obtained; the text to be retrieved is converted into a corresponding target image based on the image processing model, and global text features are extracted from the text to be retrieved based on the target image corresponding to the text to be retrieved; based on the global text features, the output image with the highest similarity to the text to be retrieved is retrieved from the image database.
[0125] In this embodiment of the invention, the text to be retrieved can be obtained, and the text to be retrieved can be converted into a corresponding target image based on the image processing model. Furthermore, global text features can be extracted from the text to be retrieved based on the target image corresponding to the text to be retrieved. Based on the global text features, the output image with the highest similarity to the text to be retrieved can be retrieved from the image database to obtain an image used to describe the original text sample. The image database can be a pre-acquired database, such as a pre-stored database obtained from the Internet. No specific restrictions are placed on the acquisition method and location of the image database here.
[0126] Optionally, in the scenario of searching for images with text, the text to be searched can be obtained, and the text to be searched can be converted into the corresponding target image based on the text processing model. Furthermore, global text features can be extracted from the text to be searched based on the target image corresponding to the text to be searched. Based on the global text features, the output image with the highest similarity among the texts to be searched can be retrieved from the image database to obtain an image used to describe the original text sample.
[0127] In this embodiment of the invention, original image samples and original text samples to be processed are obtained to describe at least one identical object. Image context feature samples containing different image regions are extracted from the original image samples, and context feature samples containing different text positions in the original text are extracted from the original text samples. The image context feature samples and text context feature samples are concatenated to obtain original feature vector samples. Feature encoding is performed on the original feature vector samples to obtain target text used to describe the input image sample, and / or an image used to describe the original text sample. The obtained images and / or samples can be used as training samples to train an image-text processing model. The image-text processing model performs unified processing between the original image and the original text, thereby reducing the impact of data noise and achieving the technical effect of improving the processing efficiency of image-text data, thus solving the technical problem of low processing efficiency of image-text data.
[0128] As another alternative embodiment, Figure 4 This is a schematic diagram of the hardware environment of a virtual reality device according to an embodiment of the present invention for processing graphic data. Figure 4 As shown, the virtual reality device 404 is connected to the terminal 406, and the terminal 406 is connected to the server 402 via a network. The virtual reality device 404 is not limited to: virtual reality headsets, virtual reality glasses, virtual reality all-in-one machines, etc. The terminal 406 is not limited to PCs, mobile phones, tablets, etc. The server 402 can be a server corresponding to a media file operator. The network includes, but is not limited to: wide area network, metropolitan area network, or local area network.
[0129] Optionally, the virtual reality device 404 in this embodiment includes a memory, a processor, and a transmission device. The memory stores an application program that can be used to perform: acquiring an original image and original text to be processed; extracting image context features from the original image and text context features from the original text; concatenating the image context features and text context features to obtain an original feature vector; and performing feature encoding on the original feature vector to obtain target text corresponding to the original image, and / or a target image corresponding to the original text.
[0130] Optionally, the terminal in this embodiment can be used to display the original image and original text to be processed on the presentation screen of a virtual reality (VR) device or an augmented reality (AR) device. The VR device or AR device extracts image context features from the original image and text context features from the original text. After concatenating the image context features and text context features to obtain the original feature vector, the VR device or AR device is driven to render and display the target text corresponding to the original image obtained by feature encoding the original feature vector.
[0131] Optionally, the virtual reality device 404 in this embodiment includes an eye-tracking head-mounted display (HMD) and an eye-tracking module that function the same as in the embodiments described above. That is, the screen in the HMD displays real-time images, and the eye-tracking module in the HMD acquires the real-time movement path of the user's eyes. In this embodiment, the terminal acquires the user's position and movement information in real three-dimensional space through a tracking system, and calculates the three-dimensional coordinates of the user's head in virtual three-dimensional space, as well as the user's field of vision orientation in virtual three-dimensional space.
[0132] Figure 4 The hardware structure block diagram shown can serve not only as an exemplary block diagram of the aforementioned AR / VR device (or mobile device), but also as an exemplary block diagram of the aforementioned server. In the operating environment described above, the present invention also provides, for example... Figure 5 The illustrated method for processing image and text data can be applied to virtual reality (VR) devices or augmented reality (AR) devices, and the model can be used to analyze video segments within VR or AR devices. It should be noted that the image and text data processing method in this embodiment can be... Figure 5 The mobile terminal in the illustrated embodiment is executed.
[0133] Figure 5 This is a flowchart of another method for processing graphic and textual data according to an embodiment of the present invention, such as... Figure 5As shown, the method may include the following steps.
[0134] Step S502: Display the original image and original text to be processed on the presentation screen of the virtual reality (VR) device or augmented reality (AR) device, wherein the original image and original text are used to describe at least one identical object.
[0135] In step S504, the VR device or AR device extracts image context features from the original image and text context features from the original text.
[0136] In the technical solution provided by step S504 of the present invention, the image context features include image features in different image regions of the original image, and the text context features include text features at different text positions in the original text.
[0137] Step S506: After concatenating the image context features and the text context features to obtain the original feature vector, drive the VR device or AR device to render and display the target text corresponding to the original image obtained by feature encoding the original feature vector, the target text including the same object, and / or drive the VR device or AR device to render and display the target image corresponding to the original text obtained by feature encoding the original feature vector, the target image including the same object.
[0138] Optionally, in this embodiment, the above-described method for processing image and text data can be applied to a hardware environment consisting of a server and a virtual reality device. The original image and text to be processed are displayed on the screen of the virtual reality device or augmented reality device. The server can be a server corresponding to a media file operator. The aforementioned network includes, but is not limited to, a wide area network (WAN), a metropolitan area network (MAN), or a local area network (LAN). The aforementioned virtual reality device is not limited to, for example, a virtual reality headset, virtual reality glasses, or a standalone virtual reality device.
[0139] It should be noted that the above-described method for processing graphic data in VR or AR devices may include... Figure 5 The method of the illustrated embodiment is intended to drive a VR device or AR device to display target text corresponding to the original image obtained by feature encoding the original feature vector.
[0140] Optionally, the processor in this embodiment can invoke the application stored in the memory via the transmission device to perform the above steps. The transmission device can receive media files sent by the server via a network, and can also be used for data transmission between the processor and the memory.
[0141] Optionally, in a virtual reality device, there is a head-mounted display with eye tracking. The screen in the HMD is used to display the video footage. The eye-tracking module in the HMD is used to acquire the real-time movement path of the user's eyes. The tracking system is used to track the user's position and movement information in real three-dimensional space. The computing and processing unit is used to acquire the user's real-time position and movement information from the tracking system and calculate the three-dimensional coordinates of the user's head in the virtual three-dimensional space, as well as the user's field of vision orientation in the virtual three-dimensional space.
[0142] In this embodiment of the invention, the virtual reality device can be connected to a terminal, and the terminal and the server are connected through a network. The virtual reality device is not limited to: virtual reality helmet, virtual reality glasses, virtual reality all-in-one machine, etc. The terminal is not limited to PC, mobile phone, tablet computer, etc. The server can be the server corresponding to the media file operator. The network includes, but is not limited to: wide area network, metropolitan area network or local area network.
[0143] Figure 6 This is a schematic diagram of the processing result of a method for processing graphic and textual data according to an embodiment of the present invention, as shown below. Figure 6 As shown, the VR or AR device renders and displays the target text "Big-eyed, short-haired woman" obtained by feature encoding the original feature vector, which corresponds to the original image.
[0144] This invention displays the original image and original text to be processed on the presentation screen of a virtual reality (VR) device or an augmented reality (AR) device. The VR or AR device extracts image context features from the original image and text context features from the original text. After concatenating the image context features and text context features to obtain the original feature vector, the VR or AR device is driven to render and display the target text corresponding to the original image obtained by feature encoding the original feature vector. This achieves the technical effect of improving the processing efficiency of image and text data and solves the technical problem of low processing efficiency of image and text data.
[0145] This invention also provides another method for processing graphic data, which can be applied to the software-as-a-Service (SaaS) side.
[0146] Figure 7 This is a flowchart of another method for processing graphic and textual data according to an embodiment of the present invention, such as... Figure 7 As shown, the method may include the following steps.
[0147] Step S702: Obtain the original image and original text to be processed by calling the first interface, wherein the original image and original text are used to describe at least one identical object.
[0148] In the technical solution provided by step S702 of the present invention, the first interface can be an interface for data interaction between the server and the client. The client can use the original image and original text to be processed as a first parameter of the first interface to achieve the purpose of obtaining the original image and original text to be processed.
[0149] Step S704: Extract image context features from the original image and extract text context features from the original text. The image context features include image features in different image regions of the original image, and the text context features include text features at different text positions in the original text.
[0150] Step S706: Concatenate the image context features and the text context features to obtain the original feature vector, wherein the features in the original feature vector are either image context features or text context features.
[0151] Step S708: Perform feature encoding on the original feature vector to obtain target text corresponding to the original image, wherein the target text includes the same object; and / or: Perform feature encoding on the original feature vector to obtain target image corresponding to the original text, wherein the target image includes the same object.
[0152] Step S710: Output target text and / or target image by calling the second interface, wherein the second interface includes a second parameter, the value of which is the target text and / or target image.
[0153] In the technical solution provided by step S710 of the present invention, the second interface can be an interface for data interaction between the server and the client. The server can transmit the target image into the second interface as a parameter of the second interface to achieve the purpose of sending the target image to the client.
[0154] Figure 8 This is a schematic diagram of image processing using a computer device according to an embodiment of the present invention, such as... Figure 8 As shown, the original image and original text to be detected can be obtained by calling the first interface. The computer device extracts image context features from the original image and text context features from the original text. The image context features and text context features are concatenated to obtain the original feature vector. The original feature vector is then feature-encoded to obtain the target text corresponding to the original image. The target text can be output by calling the second interface.
[0155] Optionally, the platform can output the target text by calling a second interface, which can be used to deploy the original image via the Internet and connect it to the system to be measured, thereby outputting the target text.
[0156] In this embodiment of the invention, image context features of the original image and text context features of the original text are obtained, and the image context features and text context features are concatenated and feature encoded. This allows for implicit alignment and unified processing between the original image and the original text. By unifying the processing between the original image and the original text, the impact of data noise is reduced, enabling the generation of text from images or images from text. This achieves the technical effect of improving the processing efficiency of image and text data and solves the technical problem of low processing efficiency of image and text data.
[0157] Example 2
[0158] The preferred implementation of the method described above in this embodiment will be further introduced below, specifically using a multimodal feature learning method for image and text generation that can train a multimodal feature model from a large amount of image and text data.
[0159] Currently, information on the internet is mainly in the form of videos, images, and text, while most user requests are text-based. For example, search engine queries are primarily based on text. Therefore, establishing connections between different modalities is of great significance in information retrieval. This is usually achieved through multimodal feature learning, which learns a feature model and uses this model to provide a unified feature representation for images and text.
[0160] In related technologies, a multimodal (CLIP) model is proposed. After inputting image or text data, the CLIP model uses a neural network, such as a convolutional neural network (CNN), to represent the input data as a global feature vector. In the retrieval stage, it is only necessary to extract the feature vector and compare it. However, after compressing image and text information into a global feature vector, a lot of detailed information is lost, which affects the model performance and still has the technical problem of low processing efficiency for image and text data.
[0161] Another related technique proposes a super-powerful pre-trained (CoCa) model. This model, after acquiring input image or text data, extracts feature vectors from the image and text data. It then cross-compares these feature vectors to generate image descriptions. This method can be applied not only to image and text retrieval tasks but also to image description subtasks. However, the constraints on cross-comparison of image and text features are too strong, negatively impacting training performance when there is significant data noise. Therefore, it still suffers from the technical problem of low processing efficiency for image and text data.
[0162] To address the aforementioned issues, this invention proposes a multimodal feature learning method for image and text generation. By unifying the image and text context feature vectors into a single encoder, such as a Transformer encoder with a self-attention mechanism, the image feature vectors and text feature vectors can be implicitly aligned. This also reduces the impact of data noise, further improving the training effect of feature vectors and achieving the technical effect of improving the processing efficiency of image and text data, thus solving the technical problem of low processing efficiency of image and text data.
[0163] In an embodiment of the present invention, Figure 9 This is a schematic diagram of processing graphic data according to an embodiment of the present invention, such as... Figure 9 As shown, the process involves acquiring the image to be detected (which could be an image of a "woman with big eyes and short hair") and the text (the text of "woman with big eyes and short hair"). The image is processed by an image encoder to obtain an image context feature vector. The text is processed by a text encoder to obtain a text context feature vector. The image context feature vector and the text context feature vector are concatenated to obtain the original feature vector. A Transformer encoder based on a self-attention mechanism is used to encode the original feature vector to obtain the target feature vector. The target feature vector is then processed, as follows: Figure 9 As shown, text corresponding to an image or an image corresponding to text can be obtained, thereby achieving the technical effect of improving the processing efficiency of image and text data and solving the technical problem of low processing efficiency of image and text data. It should be noted that in the embodiments of the present invention, the image in the training process can also be two dogs running in a field, and the text can be "two dogs running in a field". This is only an example and does not impose specific restrictions on the content and type of the image and text.
[0164] Optionally, the training of the model includes a graph-text feature vector extraction module and an alignment and unified image-text generation module, wherein the graph-text feature vector extraction module may include an image encoder and a text encoder.
[0165] As an optional embodiment, feature vectors of the input image and text information are obtained.
[0166] In this embodiment, for the input original image and original text, the image encoder can extract global features from the original image, and the text encoder can extract global features from the original text, to obtain the image feature vector of the input image and the text feature vector of the text information.
[0167] Alternatively, a loss function can be used to reduce the distance between the image feature vector and the correct text feature vector, and increase the distance between the image feature vector and the incorrect text feature vector. For example, the contrastive learning loss (InfoNCE) function can be used to reduce the distance between the image feature vector and the correct text feature vector.
[0168] Optionally, the loss function can be in the form of positive sample distance divided by negative sample distance, where the positive sample distance in the numerator is the Euclidean distance between the image feature vector and the corresponding correct text feature vector, and the negative sample distance in the denominator is the sum of the Euclidean distances between the image feature vector and all other text feature vectors. During training, the loss function can be minimized to reduce the positive sample distance and increase the negative sample distance.
[0169] It should be noted that the distances between positive and negative samples mentioned above can be Euclidean distances, cosine distances, etc., and no specific restrictions are placed on the form of the distances here.
[0170] As an optional embodiment, in addition to the image feature vector extracted by the image encoder and the text feature vector extracted by the text encoder, the corresponding image context feature vector and text context feature vector are also extracted.
[0171] Optionally, the image context feature vector can be obtained by extracting feature vectors from different regions of the image and combining multiple image feature vectors extracted from an image. For example, feature vectors from different regions of the image can be extracted to obtain a feature vector that describes the sky and background, a feature vector that describes the ground, and a feature vector that describes the color. All the obtained feature vectors can be combined to obtain the image context feature vector.
[0172] Optionally, the text context feature vector can be a combination of text feature vectors corresponding to different positions in the text sentence sequence. In the combination, a certain feature vector may describe the subject noun of the sentence, a certain feature vector may describe the action, a certain feature vector may describe the state, etc. All the obtained feature vectors are combined together to obtain the text context feature vector.
[0173] In related technologies, image feature vectors are extracted using an image encoder, and text feature vectors are extracted using a text encoder. However, the image feature vectors extracted by the image encoder are only global image feature vectors and cannot describe detailed information about different regions of the image. Similarly, the text feature vectors extracted by the text encoder are only global text feature vectors and can only be used to describe the general meaning of the entire sentence. The above methods are relatively coarse in terms of information extraction and have inaccuracy issues. To solve the above problems, this invention extracts contextual feature vectors from both the image and the text, thereby achieving the goal of more accurately determining detailed information about different regions of the image.
[0174] As an optional embodiment, such as Figure 9 As shown, the self-attention mechanism encoder (Transformer) can be used to uniformly view the obtained image context features and text context feature vectors, thus completing the task of generating text from images or images from text.
[0175] In this embodiment, the obtained image context feature vector and text context feature vector can be used to generate text from an image and vice versa.
[0176] Since the model's retrieval task requires extracting feature vectors from images and text, and then determining the distance between the image feature vector and the text feature vector, but the context feature vector is a combination of multiple feature vectors, it cannot meet the requirement of extracting one feature vector from the image and one feature vector from the text. Therefore, in this embodiment of the invention, in addition to the image encoder extracting image features and the text encoder extracting text features, corresponding image context features and text context features are also extracted. By training the model's ability to extract context features, the model can perform multimodal feature learning for each region in the image, thereby more accurately determining global image features and text features. Furthermore, this embodiment can generate text from images and images from text, thus making the invention not limited to either image-to-text or text-to-image retrieval tasks, thereby expanding the scope of application of the model.
[0177] In this embodiment, by performing subsequent processing on the context features, the problem of image and text information loss caused by only obtaining global feature vectors in related technologies is solved. Furthermore, this embodiment of the invention unifies the image and text context feature vectors into a single self-attention mechanism encoder, so that implicit alignment can be performed between images and text, thereby reducing the impact of data noise and improving the effect of feature training.
[0178] Optionally, the expressive content between the image context feature vector and the text context feature vector can be aligned one by one to achieve implicit alignment between the image and the text.
[0179] For example, in image-to-text generation, the image context features of a dog or a running dog in the image need to be aligned with the corresponding text context features. Instead of explicitly aligning the nth image context feature with the nth text context feature one by one, the correspondence is implicitly determined using image-to-text and text-to-image mapping. This achieves the goal of implicitly determining the correspondence, thus representing an implicit alignment of the semantic content between image context features and text context features. If there is an incorrect label, such as an image without a car but the text mentions a car, the noisy data cannot be aligned successfully, meaning it is essentially not involved in feature training. This approach improves the model's efficiency in processing image and text data, thereby mitigating the impact of noisy data.
[0180] In this embodiment, such as Figure 9 As shown, in this embodiment of the invention, an image is processed by an image encoder (which can be a Transformer), and text information is processed by a text encoder (which can be a Transformer) to obtain image feature vectors, text feature vectors, image context feature vectors, and text context feature vectors. A loss function can be used to reduce the distance between the image feature vectors and the correct text feature vectors, thereby achieving the purpose of aligning the obtained global feature vectors.
[0181] Optionally, the image context feature vector output by the image encoder and the text context feature vector output by the text encoder are obtained. After the image context feature vector and the text context feature vector are uniformly concatenated, they are input into the self-attention mechanism encoder to generate the corresponding image or text description. The concatenation of the image context feature vector and the text context feature vector can be done by concatenating the image context feature vector and the text context feature vector together in the order of image first and text last. For example, M image context feature vectors and N text context feature vectors can be concatenated together in the order of image first and text last to obtain M+N context feature vectors.
[0182] In related technologies, only the text context feature vector is input into the encoder for processing, while the image context feature vector is only used as an additional reference. Alternatively, when encoding the text context feature vector, the two vectors are weighted and summed to complete the prediction of image or text information. However, in the embodiments of this invention, the text context feature vector and the image context feature vector are not treated differently. The image context feature vector and the text context feature vector are uniformly concatenated, and the concatenated feature vector is processed to generate the corresponding image or text description. This achieves the technical effect of improving the processing efficiency of image and text data and solves the technical problem of low processing efficiency of image and text data.
[0183] It should be noted that the self-attention mechanism encoder can be trained only during the model training process, which can help the model learn features better in noisy image and text data. After training, in actual use, you can choose to use an image encoder or a text encoder depending on the actual use case.
[0184] Optionally, the application scenarios of text-to-image generation in this embodiment of the invention can be: in the process of artistic creation, it can help generate new paintings, materials or comics, or advertising designs, etc.; it can be used for image editing to help replace and generate new backgrounds or foregrounds, etc.; it can be used for suspect portrait generation to help the police track down suspects, etc. For example, a suspect portrait can be generated based on the description of the suspect, etc. The application scenarios here are only examples and do not impose specific limitations on the application scenarios.
[0185] Optionally, the application scenarios for image-generated text in this embodiment of the invention can be: to help blind people describe the current actual scene in language; to help with image retrieval, such as better retrieving image content using text; or for children's education, such as generating language descriptions from images in books to help children understand the world. The application scenarios here are only illustrative examples and do not impose specific limitations on the application scenarios.
[0186] In this embodiment of the invention, by introducing unified generative learning of images and text, the contextual feature vectors of images and text are unified into one encoder. This enables implicit alignment between images and text, while reducing the impact of data noise, thereby improving the effect of feature training. This achieves the technical effect of improving the processing efficiency of image and text data and solves the technical problem of low processing efficiency of image and text data.
[0187] In another alternative embodiment, Figure 10 The use of the above is illustrated in a block diagram. Figure 1 The computer terminal 30 (or mobile device) shown is an embodiment of a computing node in computing environment 1001. Figure 10This is a structural block diagram of a computing environment according to an embodiment of the present invention, such as... Figure 10 As shown, computing environment 1001 includes multiple computing nodes (such as servers) running on a distributed network (represented in the figure as 1010-1, 1010-2, ...,). Each computing node contains local processing and memory resources, and end user 1002 can remotely run applications or store data within computing environment 1001. Applications can be provided as multiple services 1020-1, 1020-2, 1020-3, and 1020-4 within computing environment 1001, representing services "A", "D", "E", and "H", respectively.
[0188] End user 1002 can provide and access services through a web browser or other software application on the client. In some embodiments, the provisioning and / or requests of end user 1002 can be provided to ingress gateway 1030. Ingress gateway 1030 may include a corresponding agent to handle provisioning and / or requests for service 1020 (one or more services provided in computing environment 1001).
[0189] Service 1020 is provided or deployed based on various virtualization technologies supported by computing environment 1001. In some embodiments, service 1020 may be provided based on virtual machine (VM)-based virtualization, container-based virtualization, and / or similar methods. Virtual machine-based virtualization may involve simulating a real computer by initializing a virtual machine, executing programs and applications without directly accessing any actual hardware resources. While the machine is virtualized by a virtual machine, container-based virtualization may launch containers to virtualize an entire operating system (OS), allowing multiple workloads to run on a single OS instance.
[0190] In one embodiment based on container virtualization, several containers of service 1020 can be assembled into a POD (e.g., a Kubernetes POD). For example, such as Figure 10 As shown, service 1020-2 can be equipped with one or more PODs 1040-1, 1040-2, ..., 1040-N (collectively referred to as POD1040). Each POD1040 can include a proxy 1045 and one or more containers 1042-1, 1042-2, ..., 1042-M (collectively referred to as container 1042). One or more containers 1042 in POD1040 handle requests related to one or more corresponding functions of the service, and the proxy 1045 typically controls service-related network functions such as routing and load balancing. Other services 1020 can also be accompanied by PODs similar to POD1040.
[0191] During operation, executing a user request from end user 1002 may require invoking one or more services 1020 in computing environment 1001, and executing one or more functions of one service 1020 may require invoking one or more functions of another service 1020. For example... Figure 10 As shown, service "A" 1020-1 receives user requests from terminal user 1002 from ingress gateway 1030. Service "A" 1020-1 can call service "D" 1020-2, and service "D" 1020-2 can request service "E" 1020-3 to perform one or more functions.
[0192] The aforementioned computing environment can be a cloud computing environment, where resource allocation is managed by cloud services, allowing functionality development without needing to consider implementation, adjustment, or server scaling. This computing environment allows developers to execute event-responsive code without building or maintaining complex infrastructure. Services can be partitioned into a set of functions that can automatically and independently scale, rather than scaling a single hardware device to handle potential loads.
[0193] In another alternative embodiment, Figure 11 The use of the above is illustrated in a block diagram. Figure 1 The computer terminal (or mobile device) shown is an example of a service mesh. Figure 11 This is a structural block diagram of a service mesh for a method of processing graphic and textual data according to an embodiment of the present invention, such as... Figure 11 As shown, the service mesh 1100 is mainly used to facilitate secure and reliable communication between multiple microservices. Microservices refer to the decomposition of an application into multiple smaller services or instances, which are distributed across different clusters / machines to run.
[0194] like Figure 11 As shown, a microservice may include application service instance A and application service instance B, which together form the functional application layer of service mesh 1100. In one implementation, application service instance A runs as a container / process 1108 on machine / workload container group 1114 (POD), and application service instance B runs as a container / process 1110 on machine / workload container group 1116 (POD).
[0195] In one implementation, application service instance A can be a product query service, and application service instance B can be a product order placement service.
[0196] like Figure 11As shown, application service instance A and grid agent (sidecar) 1103 coexist in machine workload container group 1114, and application service instance B and grid agent 1105 coexist in machine workload container 1114. Grid agents 1103 and 1105 form the data plane layer of service mesh 1100. Grid agents 1103 and 1105 run as container / process 1104, which can receive requests 1112 for product query services, and as grid agent 1106. Grid agent 1103 and application service instance A can communicate bidirectionally, as can grid agent 1105 and application service instance B. Furthermore, grid agents 1103 and 1105 can also communicate bidirectionally with each other.
[0197] In one implementation, all traffic from application service instance A is routed to the appropriate destination via mesh proxy 1103, and all network traffic from application service instance B is routed to the appropriate destination via mesh proxy 1105. It should be noted that the network traffic mentioned herein includes, but is not limited to, Hypertext Transfer Protocol (HTTP), Representational State Transfer (REST), high-performance, general-purpose open-source frameworks (gRPC), and open-source in-memory data structure storage systems (Redis).
[0198] In one implementation, the functionality of the extended data plane layer can be achieved by writing custom filters for the agents (Envoy) in service mesh 1100. The service mesh agent configuration can enable the service mesh to correctly proxy service traffic, achieving service interoperability and service governance. Mesh agents 1103 and 1105 can be configured to perform at least one of the following functions: service discovery, health checking, routing, load balancing, authentication and authorization, and observability.
[0199] like Figure 11 As shown, the service mesh 1100 also includes a control plane layer. This control plane layer can consist of a set of services running in a dedicated namespace, hosted by a managed control plane component 1101 within machine / workload container groups (machine / Pods) 1102. For example... Figure 11 As shown, the managed control plane component 1101 communicates bidirectionally with grid agents 1103 and 1105. The managed control plane component 1101 is configured to perform several control and management functions. For example, the managed control plane component 1101 receives telemetry data transmitted by grid agents 1103 and 1105 and can further aggregate this telemetry data. In addition to these services, the managed control plane component 1101 can also provide a user-facing application programming interface (API) to facilitate manipulation of network behavior and provision of configuration data to grid agents 1103 and 1105. It should be noted that, for the foregoing method embodiments, for the sake of simplicity, they are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, because according to the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.
[0200] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, because according to the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.
[0201] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0202] Example 3
[0203] According to embodiments of the present invention, a method for implementing the above is also provided. Figure 2 The illustrated method for processing graphic and textual data is a graphic and textual data processing device.
[0204] Figure 12 This is a schematic diagram of a graphic data processing apparatus according to an embodiment of the present invention. Figure 12As shown, the image and text data processing device 1200 may include: a first acquisition unit 1202, a first extraction unit 1204, a first processing unit 1206, and a second processing unit 1208.
[0205] The first acquisition unit 1202 is used to acquire the original image and original text to be processed, wherein the original image and original text are used to describe at least one identical object.
[0206] The first extraction unit 1204 is used to extract image context features from the original image and text context features from the original text.
[0207] The first processing unit 1206 is used to concatenate the image context features and the text context features to obtain the original feature vector.
[0208] The second processing unit 1208 is used to perform feature encoding on the original feature vector to obtain target text corresponding to the original image, and / or target image corresponding to the original text, wherein the target text and / or target image include the same object.
[0209] It should be noted that the first acquisition unit 1202, the first extraction unit 1204, the first processing unit 1206, and the second processing unit 1208 mentioned above correspond to steps S202 to S208 in Embodiment 1. The four units and the corresponding steps implement the same examples and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above units, as part of the device, can run in the computer terminal provided in Embodiment 1.
[0210] According to embodiments of the present invention, a method for implementing the above is also provided. Figure 3 The illustrated method for processing graphic and textual data is a graphic and textual data processing device.
[0211] Figure 13 This is a schematic diagram of another image and text data processing apparatus according to an embodiment of the present invention, such as... Figure 13 As shown, the image and text data processing device 1300 may include: a second acquisition unit 1302, a second extraction unit 1304, a third processing unit 1306, and a first output unit 1308.
[0212] The second acquisition unit 1302 is used to acquire the original image sample and the original text sample to be processed, wherein the original image sample and the original text sample are used to describe at least one identical object.
[0213] The second extraction unit 1304 is used to extract image context feature samples from the original image samples and text context feature samples from the original text.
[0214] The third processing unit 1306 is used to concatenate the image context feature samples and the text context feature samples to obtain the original feature vector samples.
[0215] The first output unit 1308 is used to output original feature vector samples, wherein the text used to describe the original image sample is obtained by feature encoding the original feature vector samples, and / or the image used to describe the original text sample is used as training samples to train the image-text processing model; wherein the image-text processing model is used to convert the input image into target text used to describe the input image, the target text including the same object, and / or the image-text processing model is used to convert the input text into target image used to describe the input text, the target image including the same object.
[0216] It should be noted that the second acquisition unit 1302, the second extraction unit 1304, the third processing unit 1306, and the first output unit 1308 mentioned above correspond to steps S302 to S308 in Embodiment 1. The four units and the corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above units, as part of the device, can run in the computer terminal provided in Embodiment 1.
[0217] According to embodiments of the present invention, a method for implementing the above is also provided. Figure 5 The illustrated image and text data processing method includes an image and text data processing apparatus that can be applied to virtual reality (VR) devices or augmented reality (AR) devices, and the model can be used to perform predictive analysis on images to be analyzed in VR devices or AR devices.
[0218] Figure 14 This is a schematic diagram of another image and text data processing apparatus according to an embodiment of the present invention. Figure 14 As shown, the image and text data processing device 1400 may include: a presentation unit 1402, a third extraction unit 1404, and a driving unit 1406.
[0219] The presentation unit 1402 is used to display the original image and original text to be processed on the presentation screen of a virtual reality (VR) device or an augmented reality (AR) device, wherein the original image and original text are used to describe at least one identical object.
[0220] The third extraction unit 1404 is used by VR devices or AR devices to extract image context features from the original image and text context features from the original text.
[0221] The driving unit 1406 is used to, after concatenating image context features and text context features to obtain an original feature vector, drive a VR device or AR device to render and display target text corresponding to the original image obtained by feature encoding the original feature vector, the target text including the same object, and / or drive a VR device or AR device to render and display target image corresponding to the original text obtained by feature encoding the original feature vector, the target image including the same object.
[0222] It should be noted that the aforementioned presentation unit 1402, third extraction unit 1404, and driving unit 1406 correspond to steps S502 to S506 in Embodiment 1. The three units and their corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should be noted that the aforementioned units, as part of the device, can run in the computer terminal provided in Embodiment 1.
[0223] According to embodiments of the present invention, a method for implementing the above is also provided. Figure 7 The illustrated method for processing graphic and textual data is a graphic and textual data processing device.
[0224] Figure 15 This is a schematic diagram of another image and text data processing apparatus according to an embodiment of the present invention. Figure 15 As shown, the image and text data processing device 1500 may include: a third acquisition unit 1502, a fourth extraction unit 1504, a fourth processing unit 1506, a fifth processing unit 1508, and a second output unit 1510.
[0225] The third acquisition unit 1502 is used to acquire the original image and original text to be processed by calling the first interface, wherein the original image and original text are used to describe at least one identical object.
[0226] The fourth extraction unit 1504 is used to extract image context features from the original image and text context features from the original text. The image context features include image features in different image regions of the original image, and the text context features include text features at different text positions in the original text.
[0227] The fourth processing unit 1506 is used to concatenate the image context features and the text context features to obtain the original feature vector, wherein the features in the original feature vector are image context features or text context features.
[0228] The fifth processing unit 1508 is used to perform feature encoding on the original feature vector to obtain target text corresponding to the original image, wherein the target text includes the same object; and / or to perform feature encoding on the original feature vector to obtain target image corresponding to the original text, wherein the target image includes the same object.
[0229] The second output unit 1510 is used to output target text and / or target image by calling a second interface, wherein the second interface includes a second parameter, the value of which is the target text and / or target image.
[0230] It should be noted that the third acquisition unit 1502, the fourth extraction unit 1504, the fourth processing unit 1506, the fifth processing unit 1508, and the second output unit 1510 mentioned above correspond to steps S702 to S710 in Embodiment 1. The five units and the corresponding steps implement the same examples and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above units, as part of the device, can run in the computer terminal provided in Embodiment 1.
[0231] In the image and text data processing apparatus of this embodiment, by acquiring the image context features of the original image and the text context features of the original text, the image context features and the text context features are concatenated for unified feature encoding processing, thereby implicitly aligning the original image and the original text, reducing the impact of data noise, and realizing the ability to generate text from images or images from text, thereby achieving the technical effect of improving the processing efficiency of image and text data and solving the technical problem of low processing efficiency of image and text data.
[0232] Example 4
[0233] Embodiments of the present invention may provide a processor, which may include a computer terminal, which may be any one of a group of computer terminals. Optionally, in this embodiment, the computer terminal may also be replaced by a mobile terminal or other terminal device.
[0234] Optionally, in this embodiment, the computer terminal may be located in at least one of a plurality of network devices in a computer network.
[0235] In this embodiment, the computer terminal described above can execute the program code for the following steps in the method for processing image and text data of an application: acquiring the original image and original text to be processed, wherein the original image and original text are used to describe at least one identical object; extracting image context features from the original image and extracting text context features from the original text; concatenating the image context features and text context features to obtain an original feature vector; performing feature encoding on the original feature vector to obtain target text corresponding to the original image and / or target image corresponding to the original text, wherein the target text and / or target image include the same object.
[0236] Optionally, Figure 16 This is a structural block diagram of a computer terminal according to an embodiment of the present invention. Figure 16 As shown, the computer terminal A may include one or more (only one is shown in the figure) processors 1602, memory 1604, and transmission devices 1606.
[0237] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the image and text data processing method and apparatus in this embodiment of the invention. The processor executes various functional applications and predictions by running the software programs and modules stored in the memory, thereby realizing the aforementioned image and text data processing method. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to terminal A via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0238] The processor can invoke information and application programs stored in memory via a transmission device to perform the following steps: acquiring an original image and original text to be processed, wherein the original image and original text are used to describe at least one identical object; extracting image context features from the original image and extracting text context features from the original text; concatenating the image context features and text context features to obtain an original feature vector; and performing feature encoding on the original feature vector to obtain target text corresponding to the original image and / or target image corresponding to the original text, wherein the target text and / or target image include the same object.
[0239] Optionally, the processor may also execute program code for the following steps: mapping the original feature vector to a target feature vector, wherein the target feature vector includes an image feature vector and a text feature vector, wherein the similarity between the object represented by the features in the image feature vector and the object represented by the text context features is greater than a first similarity threshold, and the similarity between the object represented by the features in the text feature vector and the object represented by the image context features is greater than a second similarity threshold; generating a target image corresponding to the original text from the image feature vector, and / or generating a target text corresponding to the original image from the text feature vector.
[0240] Optionally, the processor may also execute program code that performs the following steps: compares the similarity of each feature among multiple features with the features in the original feature vector other than each feature to obtain the features in the image feature vector or the text feature vector corresponding to each feature; and generates a target feature vector based on the features in the image feature vector or the text feature vector corresponding to each feature.
[0241] Optionally, the processor may also execute program code for the following steps: generating a target image corresponding to the original text from the image feature vector in the target feature vector, including: generating a target image corresponding to the original text from the image feature vector in the target feature vector based on an image generation model, wherein the parameters of the image generation model are adjusted by a loss function between the image feature vector and the original image feature vector of the original image, wherein the image generation model is a machine learning model; and / or generating a target text corresponding to the original image from the text feature vector in the target feature vector based on a text generation model, wherein the parameters of the text generation model are adjusted by a loss function between the text feature vector and the original text feature vector of the original text, wherein the text generation model is a machine learning model.
[0242] Optionally, the processor may also execute program code for the following steps: extracting global image features from the original image based on an image feature extraction model, wherein the parameters of the image processing model are adjusted by a loss function between the image feature vector and the original image feature vector of the original image; retrieving the output text with the highest similarity to the original image from a text database based on the global image features; and / or extracting global text features from the original text based on a text feature extraction model, wherein the parameters of the text feature extraction model are adjusted by a loss function between the text feature vector and the original text feature vector of the original text; and retrieving the output image with the highest similarity to the original text from an image database based on the global text features.
[0243] As an optional example, the processor can invoke information and applications stored in memory via a transmission device to perform the following steps: acquiring original image samples and original text samples to be processed, wherein the original image samples and original text samples are used to describe at least one identical object; extracting image context feature samples from the original image samples and extracting text context feature samples from the original text; concatenating the image context feature samples and text context feature samples to obtain original feature vector samples; outputting the original feature vector samples, wherein the text used to describe the original image samples, and / or the image used to describe the original text samples, obtained by feature encoding of the original feature vector samples, is used as training samples to train a text-image processing model; wherein the text-image processing model is used to convert an input image into target text used to describe the input image, the target text including the identical object, and / or the text-image processing model is used to convert the input text into a target image used to describe the input text, the target image including the identical object.
[0244] Optionally, the processor may also execute program code that performs the following steps: acquiring the image to be retrieved; converting the image to be retrieved into the corresponding target text based on the image-text processing model, and extracting global image features from the image to be retrieved based on the target text corresponding to the image to be retrieved; and retrieving the output text with the highest similarity to the image to be retrieved from the text database based on the global image features.
[0245] Optionally, the processor may also execute program code that performs the following steps: acquiring the text to be retrieved; converting the text to be retrieved into a corresponding target image based on the image processing model, and extracting global text features from the text to be retrieved based on the target image corresponding to the text to be retrieved; and retrieving the output image with the highest similarity to the text to be retrieved from the image database based on the global text features.
[0246] As an alternative example, the processor may invoke information and applications stored in memory via a transmission device to perform the following steps: displaying the original image and original text to be processed on the presentation screen of a virtual reality (VR) device or an augmented reality (AR) device, wherein the original image and original text are used to describe at least one identical object; the VR device or AR device extracts image context features from the original image and extracts text context features from the original text; after concatenating the image context features and text context features to obtain an original feature vector, the VR device or AR device is driven to render and display the target text corresponding to the original image obtained by feature encoding the original feature vector, wherein the target text includes the same object; and / or, the VR device or AR device is driven to render and display the target image corresponding to the original text obtained by feature encoding the original feature vector, wherein the target image includes the same object.
[0247] As an optional example, the processor can invoke information and an application stored in memory via a transmission device to perform the following steps: acquiring an original image and original text to be processed by invoking a first interface, wherein the original image and original text are used to describe at least one identical object; extracting image context features from the original image and extracting text context features from the original text, wherein the image context features include image features in different image regions of the original image, and the text context features include text features at different text positions in the original text; concatenating the image context features and the text context features to obtain an original feature vector, wherein the features in the original feature vector are image context features or text context features; performing feature encoding on the original feature vector to obtain target text corresponding to the original image, wherein the target text includes the same object; and / or performing feature encoding on the original feature vector to obtain a target image corresponding to the original text, wherein the target image includes the same object; and outputting the target text and / or target image by invoking a second interface, wherein the second interface includes a second parameter, the value of which is the target text and / or target image.
[0248] This invention obtains the image context features of the original image and the text context features of the original text, and concatenates the image context features and text context features for unified feature encoding processing. This allows for implicit alignment between the original image and the original text, reducing the impact of data noise and enabling the generation of text from images or images from text. This achieves the technical effect of improving the processing efficiency of image and text data and solves the technical problem of low processing efficiency of image and text data.
[0249] Those skilled in the art will understand that Figure 16 The structure shown is for illustrative purposes only. Computer terminal A can also be a smartphone (such as an Android phone, iOS phone, etc.), tablet computer, mobile internet device (MID), PAD and other terminal devices. Figure 16 This does not limit the structure of the aforementioned computer terminal A. For example, computer terminal A may also include components that are more complex than those described above. Figure 16 Showing more or fewer components (such as network interfaces, display devices, etc.), or having the same Figure 16 The different configurations shown.
[0250] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0251] Example 5
[0252] Embodiments of the present invention also provide a computer-readable storage medium. Optionally, in this embodiment, the computer-readable storage medium can be used to store the program code executed by the image and text data processing method provided in Embodiment 1.
[0253] Optionally, in this embodiment, the computer-readable storage medium may be located in any computer terminal in a group of computer terminals in a computer network, or in any mobile terminal in a group of mobile terminals.
[0254] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: acquiring an original image and original text to be processed, wherein the original image and original text are used to describe at least one identical object; extracting image context features from the original image and extracting text context features from the original text; concatenating the image context features and text context features to obtain an original feature vector; performing feature encoding on the original feature vector to obtain target text corresponding to the original image, and / or a target image corresponding to the original text, wherein the target text and / or target image include the same object.
[0255] Optionally, the aforementioned computer-readable storage medium may also execute program code that performs the following steps: mapping the original feature vector to a target feature vector, wherein the target feature vector includes an image feature vector and a text feature vector, the similarity between the object represented by the features in the image feature vector and the object represented by the text context features is greater than a first similarity threshold, and the similarity between the object represented by the features in the text feature vector and the object represented by the image context features is greater than a second similarity threshold; generating a target image corresponding to the original text from the image feature vector, and / or generating a target text corresponding to the original image from the text feature vector.
[0256] Optionally, the computer-readable storage medium may also execute program code that performs the following steps: comparing the similarity of each feature among multiple features with the features in the original feature vector other than each feature to obtain the features in the image feature vector or the text feature vector corresponding to each feature; and generating a target feature vector based on the features in the image feature vector or the text feature vector corresponding to each feature.
[0257] Optionally, the aforementioned computer-readable storage medium may also execute program code that performs the following steps: generating a target image corresponding to the original text from the image feature vector in the target feature vector, including: generating a target image corresponding to the original text from the image feature vector in the target feature vector based on an image generation model, wherein the parameters of the image generation model are adjusted by a loss function between the image feature vector and the original image feature vector of the original image, wherein the image generation model is a machine learning model; and / or generating a target text corresponding to the original image from the text feature vector in the target feature vector based on a text generation model, wherein the parameters of the text generation model are adjusted by a loss function between the text feature vector and the original text feature vector of the original text, wherein the text generation model is a machine learning model.
[0258] Optionally, the aforementioned computer-readable storage medium may also execute program code that performs the following steps: extracting global image features from the original image based on an image feature extraction model, wherein the parameters of the image processing model are adjusted by a loss function between the image feature vector and the original image feature vector of the original image; retrieving the output text with the highest similarity to the original image from a text database based on the global image features; and / or extracting global text features from the original text based on a text feature extraction model, wherein the parameters of the text feature extraction model are adjusted by a loss function between the text feature vector and the original text feature vector of the original text; and retrieving the output image with the highest similarity to the original text from an image database based on the global text features.
[0259] As an optional example, a computer-readable storage medium is configured to store program code for performing the following steps: acquiring original image samples and original text samples to be processed, wherein the original image samples and original text samples are used to describe at least one identical object; extracting image context feature samples from the original image samples and extracting text context feature samples from the original text; concatenating the image context feature samples and text context feature samples to obtain original feature vector samples; outputting original feature vector samples, wherein the text used to describe the original image samples, obtained by feature encoding the original feature vector samples, and / or the image used to describe the original text samples, are used as training samples to train a text-image processing model; wherein the text-image processing model is used to convert an input image into target text used to describe the input image, the target text including the identical object, and / or the text-image processing model is used to convert the input text into a target image used to describe the input text, the target image including the identical object.
[0260] Optionally, the aforementioned computer-readable storage medium may also execute program code that performs the following steps: acquiring an image to be retrieved; converting the image to be retrieved into corresponding target text based on a text processing model, and extracting global image features from the image to be retrieved based on the target text corresponding to the image to be retrieved; and retrieving the output text with the highest similarity to the image to be retrieved from a text database based on the global image features.
[0261] Optionally, the aforementioned computer-readable storage medium may also execute program code that performs the following steps: acquiring the text to be retrieved; converting the text to be retrieved into a corresponding target image based on a text-image processing model, and extracting global text features from the text to be retrieved based on the target image corresponding to the text to be retrieved; and retrieving the output image with the highest similarity to the text to be retrieved from the image database based on the global text features.
[0262] As an optional example, a computer-readable storage medium is configured to store program code for performing the following steps: displaying an original image and original text to be processed on a rendering screen of a virtual reality (VR) device or an augmented reality (AR) device, wherein the original image and original text are used to describe at least one identical object; the VR device or AR device extracts image context features from the original image and extracts text context features from the original text; after concatenating the image context features and text context features to obtain an original feature vector, driving the VR device or AR device to render and display target text corresponding to the original image obtained by feature encoding the original feature vector, the target text including the same object; and / or driving the VR device or AR device to render and display a target image corresponding to the original text obtained by feature encoding the original feature vector, wherein the target image includes the same object.
[0263] As an optional example, a computer-readable storage medium is configured to store program code for performing the following steps: acquiring a raw image and raw text to be processed by calling a first interface, wherein the raw image and raw text are used to describe at least one identical object; extracting image context features from the raw image and extracting text context features from the raw text, wherein the image context features include image features in different image regions of the raw image, and the text context features include text features at different text positions in the raw text; concatenating the image context features and the text context features to obtain a raw feature vector, wherein the features in the raw feature vector are image context features or text context features; performing feature encoding on the raw feature vector to obtain target text corresponding to the raw image, wherein the target text includes the same object; and / or performing feature encoding on the raw feature vector to obtain a target image corresponding to the raw text, wherein the target image includes the same object; and outputting the target text and / or the target image by calling a second interface, wherein the second interface includes a second parameter, the value of which is the target text and / or the target image.
[0264] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0265] In the above embodiments of the present invention, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0266] In the several embodiments provided by this invention, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection of units or modules may be electrical or other forms.
[0267] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0268] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0269] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0270] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A method for processing graphic and textual data, characterized in that, include: Obtain the original image and original text to be processed, wherein the original image and the original text are used to describe at least one identical object; Image context features are extracted from the original image, and text context features are extracted from the original text; The image context features and the text context features are concatenated to obtain the original feature vector; The original feature vector is feature encoded to obtain target text corresponding to the original image, and / or target image corresponding to the original text, wherein the target text and / or the target image include the same object; The process of encoding the original feature vector to obtain target text corresponding to the original image and / or target image corresponding to the original text includes: mapping the original feature vector to a target feature vector, wherein the target feature vector includes an image feature vector and a text feature vector, wherein the similarity between the object represented by the features in the image feature vector and the object represented by the text context features is greater than a first similarity threshold, and the similarity between the object represented by the features in the text feature vector and the object represented by the image context features is greater than a second similarity threshold; generating the target image corresponding to the original text from the image feature vector, and / or generating the target text corresponding to the original image from the text feature vector.
2. The method according to claim 1, characterized in that, The original feature vector includes multiple features. Mapping the original feature vector to the target feature vector includes: Each of the plurality of features is compared with the features in the original feature vector other than each of the features to obtain the feature in the image feature vector or the feature in the text feature vector corresponding to each feature; The target feature vector is generated based on the features in the image feature vector or the features in the text feature vector corresponding to each feature.
3. The method according to claim 1, characterized in that, Generating a target image corresponding to the original text from the image feature vector in the target feature vector includes: generating the target image corresponding to the original text from the image feature vector in the target feature vector based on an image generation model, wherein the parameters of the image generation model are adjusted by a loss function between the image feature vector and the original image feature vector of the original image, and wherein the image generation model is a machine learning model; and / or The text generation model generates the target text corresponding to the original image from the text feature vector in the target feature vector. The parameters of the text generation model are adjusted by the loss function between the text feature vector and the original text feature vector of the original text. The text generation model is a machine learning model.
4. The method according to claim 1, characterized in that, The method further includes: Global image features are extracted from the original image based on an image feature extraction model, wherein the parameters of the image feature extraction model are adjusted by a loss function between the image feature vector and the original image feature vector; based on the global image features, the output text with the highest similarity to the original image is retrieved from a text database; and / or Global text features are extracted from the original text based on a text feature extraction model, wherein the parameters of the text feature extraction model are adjusted by a loss function between the text feature vector and the original text feature vector; and the output image with the highest similarity to the original text is retrieved from the image database based on the global text features.
5. A method for processing graphic and textual data, characterized in that, include: Obtain raw image samples and raw text samples to be processed, wherein the raw image samples and raw text samples are used to describe at least one identical object; Image context feature samples are extracted from the original image samples, and text context feature samples are extracted from the original text; The image context feature samples and the text context feature samples are concatenated to obtain the original feature vector samples; The original feature vector sample is output, wherein the text used to describe the original image sample is obtained by feature encoding the original feature vector sample, and / or the image used to describe the original text sample is used as training samples to train the image and text processing model. The image processing model is used to convert an input image into target text describing the input image, the target text including the same object, and / or the image processing model is used to convert the input text into a target image describing the input text, the target image including the same object; The process of obtaining text describing the original image sample and / or an image describing the original text sample by feature encoding the original feature vector sample includes: mapping the original feature vector sample to a target feature vector sample, wherein the target feature vector sample includes an image feature vector sample and a text feature vector sample, wherein the similarity between the object represented by the features in the image feature vector sample and the object represented by the text context feature sample is greater than a first similarity threshold, and the similarity between the object represented by the features in the text feature vector sample and the object represented by the image context feature sample is greater than a second similarity threshold; generating text describing the original image sample from the text feature vector sample, and / or generating an image describing the original text sample from the image feature vector sample.
6. The method according to claim 5, characterized in that, The method further includes: Obtain the image to be retrieved; Based on the image-text processing model, the image to be retrieved is converted into the corresponding target text, and global image features are extracted from the image to be retrieved based on the target text corresponding to the image to be retrieved. Based on the global image features, the text with the highest similarity to the image to be retrieved is retrieved from the text database.
7. The method according to claim 6, characterized in that, The method further includes: Get the text to be searched; Based on the image processing model, the text to be retrieved is converted into the corresponding target image, and global text features are extracted from the text to be retrieved based on the target image corresponding to the text to be retrieved. Based on the global text features, the image with the highest similarity to the text to be retrieved is retrieved from the image database.
8. A method for processing graphic and textual data, characterized in that, include: Displaying the original image and original text to be processed on the presentation screen of a virtual reality (VR) device or an augmented reality (AR) device, wherein the original image and the original text are used to describe at least one identical object; The VR or AR device extracts image context features from the original image and text context features from the original text; After concatenating the image context features and the text context features to obtain the original feature vector, the VR device or the AR device is driven to render and display the target text corresponding to the original image obtained by feature encoding the original feature vector, wherein the target text includes the same object; and / or, the VR device or the AR device is driven to render and display the target image corresponding to the original text obtained by feature encoding the original feature vector, wherein the target image includes the same object; The method further includes: mapping the original feature vector to a target feature vector, wherein the target feature vector includes an image feature vector and a text feature vector, wherein the similarity between the object represented by the features in the image feature vector and the object represented by the text context features is greater than a first similarity threshold, and the similarity between the object represented by the features in the text feature vector and the object represented by the image context features is greater than a second similarity threshold; generating the target text corresponding to the original image from the text feature vector, and / or generating the target image corresponding to the original text from the image feature vector.
9. A method for processing graphic and textual data, characterized in that, include: The original image and original text to be processed are obtained by calling the first interface, wherein the original image and the original text are used to describe at least one identical object; Image context features are extracted from the original image, and text context features are extracted from the original text. The image context features include image features in different image regions of the original image, and the text context features include text features at different text positions in the original text. The image context features and the text context features are concatenated to obtain an original feature vector, wherein the features in the original feature vector are either the image context features or the text context features; The original feature vector is subjected to feature encoding to obtain target text corresponding to the original image, wherein the target text includes the same object; and / or, the original feature vector is subjected to feature encoding to obtain target image corresponding to the original text, wherein the target image includes the same object; The target text and / or the target image are output by calling the second interface; The process of encoding the original feature vector to obtain target text corresponding to the original image, and / or encoding the original feature vector to obtain target image corresponding to the original text, includes: mapping the original feature vector to a target feature vector, wherein the target feature vector includes an image feature vector and a text feature vector, wherein the similarity between the object represented by the features in the image feature vector and the object represented by the text context features is greater than a first similarity threshold, and the similarity between the object represented by the features in the text feature vector and the object represented by the image context features is greater than a second similarity threshold; generating the target text corresponding to the original image from the text feature vector, and / or generating the target image corresponding to the original text from the image feature vector.
10. A processor, characterized in that, The processor is used to run a program, wherein the program executes the method according to any one of claims 1 to 9 when it runs.