Image data processing method and device, storage medium and processor

By acquiring and processing multimodal information of product objects and using a multimodal network model to generate more accurate product information, the problem of low accuracy in description content in existing technologies is solved, and the accuracy of product information description and publishing efficiency are improved.

CN114168777BActive Publication Date: 2026-07-24ALIBABA DAMO (HANGZHOU) TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ALIBABA DAMO (HANGZHOU) TECH CO LTD
Filing Date
2020-09-10
Publication Date
2026-07-24

Smart Images

  • Figure CN114168777B_ABST
    Figure CN114168777B_ABST
Patent Text Reader

Abstract

The application discloses a kind of processing method, device, storage medium and processor of image data.Therein, the method includes: obtaining the product data of product object, wherein product data includes at least one of the following: product picture information, video information and text information;Analysis product data, generate the multi-modal information of product object, wherein multi-modal information includes: the feature sequence of different modal information;Multi-modal network model is used to process multi-modal information, generate product information for describing product object.The present application solves the technical problem of low precision of product information description content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing, and more specifically, to a method, apparatus, storage medium, and processor for processing image data. Background Technology

[0002] Currently, during the process of publishing product information, the product data of the currently uploaded product object is usually used to automatically fill in the complete product information of the product object before publishing.

[0003] Since the product information to be published is generally composed of product features from different modalities, but the method for generating this product information only models information from a single modality, this results in a technical problem of low accuracy in the description of the generated product information.

[0004] There is currently no effective solution to the above problems. Summary of the Invention

[0005] This invention provides a method, apparatus, storage medium, and processor for processing image data, to at least address the technical problem of low accuracy in the description of product information.

[0006] According to one aspect of the present invention, a method for processing image data is provided. The method may include: acquiring product data of a product object, wherein the product data includes at least one of the following: image information, video information, and text information of the product; analyzing the product data to generate multimodal information of the product object, wherein the multimodal information includes: feature sequences of different modal information; and processing the multimodal information using a multimodal network model to generate product information describing the product object.

[0007] According to one aspect of the present invention, a method for processing image data is also provided. The method may include: entering product data of a product object into an input page on an operation interface, wherein the product data includes at least one of the following: image information, video information, and text information of the product; sensing a text generation instruction within the operation interface, analyzing the product data, and generating multimodal information of the product object, wherein the multimodal information includes: feature sequences of different modal information; and displaying product information describing the product object on the operation interface, wherein the product information is generated by processing the multimodal information using a multimodal network model.

[0008] According to one aspect of the present invention, a method for processing image data is also provided. The method may include: displaying product data of a product object on an interactive interface, wherein the product data includes at least one of the following: image information, video information, and text information of the product; sensing a text generation instruction within the interactive interface; responding to the text generation instruction, analyzing the product data, and generating multimodal information of the product object, wherein the multimodal information includes: feature sequences of different modal information; outputting a selection page on the interactive interface, the selection page providing at least one text option, wherein different text options are used to characterize different processing models used for different modal information; displaying product information describing the product object on the interactive interface, wherein, based on the selected text option, a multimodal network model is used to process the multimodal information to generate product information.

[0009] According to one aspect of the present invention, a method for processing image data is also provided. The method may include: a front-end client uploading product data of a product object, wherein the product data includes at least one of the following: image information, video information, and text information of the product; the front-end client transmitting the product data of the product object to a back-end server; the front-end client receiving multimodal information generated by analyzing the product data returned by the back-end server, wherein the multimodal information includes: feature sequences of different modal information of the product object; and the front-end client processing the multimodal information using a multimodal network model to generate product information describing the product object.

[0010] According to one aspect of the present invention, an image data processing apparatus is also provided. The apparatus may include: an acquisition unit for acquiring product data of a product object, wherein the product data includes at least one of the following: image information, video information, and text information of the product; a first processing unit for analyzing the product data and generating multimodal information of the product object, wherein the multimodal information includes: feature sequences of different modal information; and a second processing unit for processing the multimodal information using a multimodal network model to generate product information describing the product object.

[0011] According to one aspect of the present invention, an image data processing apparatus is also provided. The apparatus may include: an input unit for inputting product data of a product object into an input page on an operation interface, wherein the product data includes at least one of the following: image information, video information, and text information of the product; a third processing unit for sensing a text generation instruction within the operation interface, analyzing the product data, and generating multimodal information of the product object, wherein the multimodal information includes: feature sequences of different modal information; and a first display unit for displaying product information describing the product object on the operation interface, wherein the product information is generated by processing multimodal information using a multimodal network model.

[0012] According to one aspect of the present invention, an image data processing apparatus is also provided. The apparatus may include: a second display unit for displaying product data of a product object on an interactive interface, wherein the product data includes at least one of the following: image information, video information, and text information of the product; a sensing unit for sensing a text generation instruction within the interactive interface; a fourth processing unit for responding to the text generation instruction, analyzing the product data, and generating multimodal information of the product object, wherein the multimodal information includes: feature sequences of different modal information; an output unit for outputting a selection page on the interactive interface, the selection page providing at least one text option, wherein different text options are used to characterize different processing models used for different modal information; and a third display unit for displaying product information describing the product object on the interactive interface, wherein multimodal information is processed using a multimodal network model based on the selected text option to generate product information.

[0013] According to one aspect of the present invention, an image data processing apparatus is also provided. The apparatus may include: an uploading unit for enabling a front-end client to upload product data of a product object, wherein the product data includes at least one of the following: image information, video information, and text information of the product; a transmission unit for enabling the front-end client to transmit the product data of the product object to a back-end server; a receiving unit for enabling the front-end client to receive multimodal information generated by analyzing the product data returned by the back-end server, wherein the multimodal information includes: feature sequences of different modal information of the product object; and a fifth processing unit for enabling the front-end client to process the multimodal information using a multimodal network model to generate product information describing the product object.

[0014] According to one aspect of the present invention, a computer-readable storage medium is also provided. The computer-readable storage medium includes a stored program, wherein, when the program is executed by a processor, it controls the device where the computer-readable storage medium is located to perform the image data processing method of the present invention.

[0015] According to one aspect of the present invention, a processor is also provided. The processor is configured to run a program, wherein the program, when running, executes the image data processing method of the embodiments of the present invention.

[0016] According to one aspect of the present invention, an image data processing system is also provided. The system may include: a processor; and a memory connected to the processor, configured to provide the processor with instructions to perform the following processing steps: acquiring product data of a product object, wherein the product data includes at least one of the following: image information, video information, and text information of the product; analyzing the product data to generate multimodal information of the product object, wherein the multimodal information includes: feature sequences of different modal information; and processing the multimodal information using a multimodal network model to generate product information describing the product object.

[0017] In this embodiment of the invention, product data of a product object is acquired, wherein the product data includes at least one of the following: product image information, video information, and text information; the product data is analyzed to generate multimodal information of the product object, wherein the multimodal information includes: feature sequences of different modal information; the multimodal information is processed using a multimodal network model to generate product information describing the product object. In other words, this application, by acquiring multimodal information of a product object and comprehensively processing the multimodal information based on a multimodal network model, generates more accurate product information describing the product object, solving the technical problem of low accuracy in the description of product information and achieving the technical effect of improving the accuracy of the description of product information. Attached Figure Description

[0018] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings:

[0019] Figure 1 This is a hardware structure block diagram of a computer terminal (or mobile device) for implementing an image data processing method according to an embodiment of the present invention;

[0020] Figure 2 This is a flowchart of an image data processing method according to an embodiment of the present invention;

[0021] Figure 3 This is a flowchart of another image data processing method according to an embodiment of the present invention;

[0022] Figure 4 This is a flowchart of another image data processing method according to an embodiment of the present invention;

[0023] Figure 5 This is a flowchart of another image data processing method according to an embodiment of the present invention;

[0024] Figure 6 This is a schematic diagram of a method for processing commodity image data according to an embodiment of the present invention;

[0025] Figure 7 This is a schematic diagram illustrating the processing of the aforementioned product image, product attribute keywords, and product category keywords using a transformer network model according to an embodiment of the present invention.

[0026] Figure 8A This is a schematic diagram of the interactive interface of an image data processing method according to an embodiment of the present invention;

[0027] Figure 8B This is a schematic diagram of a scene for processing image data according to an embodiment of the present invention;

[0028] Figure 9 This is a schematic diagram of an image data processing apparatus according to an embodiment of the present invention;

[0029] Figure 10 This is an illustration of another image data processing apparatus according to an embodiment of the present invention;

[0030] Figure 11 This is a schematic diagram of another image data processing apparatus according to an embodiment of the present invention;

[0031] Figure 12 This is a schematic diagram of another image data processing apparatus according to an embodiment of the present invention; and

[0032] Figure 13 This is a structural block diagram of a computer terminal according to an embodiment of the present invention. Detailed Implementation

[0033] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0034] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0035] First, some nouns or terms that appear in the description of the embodiments of this application shall be interpreted as follows:

[0036] Convolutional Neural Networks (CNNs) are a type of feedforward neural network where artificial neurons respond to surrounding units and can perform large-scale image processing. They include convolutional layers and pooling layers.

[0037] Long Short-Term Memory (LSTM) networks are a type of temporal recurrent neural network suitable for processing and predicting important events with relatively long intervals and delays in time series.

[0038] The multimodal transformer network model is an end-to-end model that can be viewed as an encoder-decoder structure. It can fully learn the multimodal information of the input using automatic learning methods to generate accurate product information.

[0039] Self-attention is a type of attention mechanism and an important component of transformers. Its purpose is to focus on specific details rather than to analyze the whole picture. The core is how to determine the part to focus on based on the goal, and then analyze it further after finding the details.

[0040] Cross-entropy loss is a loss function commonly used in classification problems.

[0041] Example 1

[0042] According to an embodiment of the present invention, an embodiment of a method for processing image data is also provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0043] The method embodiment provided in Embodiment 1 of this application can be executed on a mobile terminal, computer terminal, or similar computing device. Figure 1 This is a hardware structure block diagram of a computer terminal (or mobile device) for implementing an image data processing method according to an embodiment of the present invention. Figure 1 As shown, the computer terminal 10 (or mobile device 10) may include one or more processors (shown as 102a, 102b, ..., 102n in the figure) (the processor may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, computer terminal 10 may also include... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.

[0044] It should be noted that the aforementioned one or more processors and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits can be implemented wholly or partially as software, hardware, firmware, or any other combination. Furthermore, the data processing circuits can be a single, independent processing module, or wholly or partially integrated into any other element within the computer terminal 10 (or mobile device). As involved in the embodiments of this application, the data processing circuit serves as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).

[0045] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the image data processing method in this embodiment of the invention. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory 104, thereby implementing the image data processing method of the aforementioned application. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor, and these remote memories can be connected to the computer terminal 10 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0046] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the computer terminal 10. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module, used for wireless communication with the Internet.

[0047] The display may be, for example, a touchscreen liquid crystal display (LCD) that allows the user to interact with the user interface of the computer terminal 10 (or mobile device).

[0048] It should be noted here that, in some optional embodiments, the above... Figure 1 The computer device (or mobile device) shown may include hardware elements (including circuitry), software elements (including computer code stored on a computer-readable medium), or a combination of both hardware and software elements. It should be noted that... Figure 1 This is only one instance of a specific particular instance and is intended to illustrate the types of components that may exist in the aforementioned computer device (or mobile device).

[0049] exist Figure 1 Under the operating environment shown, this application provides the following: Figure 2 The image data processing method shown is illustrated. It should be noted that the image reconstruction method in this embodiment can be derived from... Figure 1 The mobile terminal in the illustrated embodiment is executed.

[0050] Figure 2 This is a flowchart of an image data processing method according to an embodiment of the present invention. Figure 2As shown, the method may include the following steps:

[0051] Step S202: Obtain product data for the product object.

[0052] In the technical solution provided by step S202 of the present invention, the product data includes at least one of the following: product image information, video information, and text information.

[0053] In this embodiment, the product object can be a commodity object, such as a new product to be published by a seller. The product data of the aforementioned product object is obtained. This product data can be used to describe the product object from multiple different perspectives, including image information and text information. The image information can include pictures and videos, which are visual information and can include details such as color and texture within the product object. The text information can be used to abstractly describe the high-level semantic information of the commodity. The image information, video information, and text information have strong complementary characteristics.

[0054] Step S204: Analyze product data and generate multimodal information of product objects.

[0055] In the technical solution provided by step S204 of the present invention, after obtaining the product data of the product object, the product data is analyzed to generate multimodal information of the product object, wherein the multimodal information includes: feature sequences of different modal information.

[0056] In this embodiment, analyzing product data can involve detecting the product data and generating keywords for the product object. These keywords can be used to characterize the product object's features. Then, the product data and the product object's keywords are combined to generate multimodal information about the product object. This multimodal information, also known as multimodal data, can include feature sequences of different modalities. The different modalities can be modal information of different modes, and the feature sequences can include image feature sequences and text feature sequences. Optionally, the multimodal information includes image information and text information about the product object.

[0057] Step S206: Use a multimodal network model to process multimodal information and generate product information to describe the product object.

[0058] In the technical solution provided in step S206 of the present invention, after analyzing product data and generating multimodal information of product objects, a multimodal network model is used to process the multimodal information to generate product information for describing product objects.

[0059] In this embodiment, the multimodal network model can be an end-to-end model, which can be viewed as an encoder-decoder structure. It can utilize automatic learning methods to fully learn the input multimodal information to generate accurate product information. This product information can be textual descriptions used to describe product objects, such as commodity information, which may include, but is not limited to, product titles, product selling points, etc. Optionally, the multimodal network model in this embodiment can be a multimodal transformer network model, used to fully learn the correlations between different modal information, thereby generating more accurate product information.

[0060] This embodiment utilizes a multimodal network model to comprehensively leverage multimodal information, resulting in more accurate descriptions of the generated product information. Optionally, this embodiment automatically fills in the generated product information into the required information template when publishing the product object, thereby reducing the time sellers spend manually filling in product information and improving the efficiency of product publishing.

[0061] In the context of intelligent product publishing, automatically generating product information by combining multimodal information about the product is crucial for improving the efficiency of sellers publishing product objects. However, in related technologies, neither unimodal nor multimodal text description generation algorithms can fully utilize the complementary relationships between different modalities, resulting in low accuracy in the generated product information descriptions.

[0062] However, this application, through steps S202 to S206 described above, obtains product data of the product object, wherein the product data includes at least one of the following: product image information, video information, and text information; analyzes the product data to generate multimodal information of the product object, wherein the multimodal information includes: feature sequences of different modal information; and processes the multimodal information using a multimodal network model to generate product information describing the product object. In other words, this application, by obtaining multimodal information of the product object and comprehensively processing the multimodal information based on a multimodal network model, generates more accurate product information describing the product object, thus solving the technical problem of low accuracy in the description of product information and achieving the technical effect of improving the accuracy of the description of product information.

[0063] The method described in this embodiment will be further described below.

[0064] As an optional implementation, in step S206, the multimodal network model generates product information by learning the correlation between different modal information during the process of processing multimodal information.

[0065] In this embodiment, the multimodal information includes feature sequences between different modalities, and these different modalities possess complementary information. In processing this multimodal information, the multimodal network model of this embodiment can learn the relationships between different modalities, fully utilizing the complementary information to generate product information, thereby effectively improving the accuracy of product description.

[0066] As an optional implementation, step S204 involves analyzing product data and generating multimodal information of product objects, including: performing attribute detection and category prediction on product data to generate attribute keywords and category keywords for product objects; and preprocessing different modal information based on product data, attribute keywords, and category keywords to generate multimodal information of product objects.

[0067] In this embodiment, when analyzing product data and generating multimodal information of product objects, the product attribute detection module can perform category detection on the product data after attribute detection to generate attribute keywords for the product objects. This embodiment can also use a category prediction module to predict the category of the product data and generate category keywords for the product objects. After generating the attribute keywords and category keywords for the product objects, this embodiment can preprocess different modal information based on the product data, the attribute keywords, and the category keywords. Optionally, this embodiment can preprocess different modal information based on product images, the attribute keywords, and the category keywords for the product objects to generate multimodal information of the product objects. The method for preprocessing different modal information based on product data, the attribute keywords, and the category keywords of the product objects in this embodiment will be further described below.

[0068] As an optional implementation, based on product data, product object attribute keywords, and category keywords, preprocessing of different modal information is performed, including: the encoder of the multimodal network model uses a convolutional neural network model to extract image features from product images and videos, generating image feature sequences; the encoder of the multimodal network model extracts text structured coding features from product object attribute keywords and category keywords, generating text feature sequences; and the image feature sequences and text feature sequences are concatenated to generate preprocessed results.

[0069] In this embodiment, the encoder of the multimodal network model may include a convolutional neural network model, such as a ResNet-50 convolutional neural network. When preprocessing different modal information based on product data, product object attribute keywords, and category keywords, this embodiment may involve the encoder using the convolutional neural network model to extract features from product images and videos, generating an image feature sequence from the extracted features. This could involve extracting feature maps from images and videos and combining them into an image feature sequence. Optionally, the attribute keywords and category keywords in this embodiment may include text structured encoding features (word embeddings). The encoder can extract these text structured encoding features and combine them into a text feature sequence. In this embodiment, the feature sequences of different modal information include the aforementioned image feature sequence and text feature sequence.

[0070] After generating the above image feature sequence and text feature sequence, the image feature sequence and text feature sequence can be concatenated to generate a preprocessed result, which can then be input into the encoder.

[0071] As an optional implementation, step S206 involves using a multimodal network model to process multimodal information and generate product information to describe the product object. This includes: using an encoder of the multimodal network model to encode the preprocessing result and generate a graphic feature sequence, wherein the graphic feature sequence is a feature sequence containing multimodal temporal attention information of images and text; and using a decoder of the multimodal network model to generate product information based on the graphic feature sequence.

[0072] In this embodiment, the multimodal network model may include an encoder and a decoder, where the encoder can be referred to as an encoder submodule and the decoder as a decoder submodule. This embodiment can preprocess different modal information based on product data, product object attribute keywords, and category keywords to generate multimodal information of the product object. Then, the encoder further encodes the preprocessed results to obtain a text-image feature sequence. This text-image feature sequence contains feature sequences of multimodal temporal attention information for both images and text, and can also be referred to as a multimodal temporal feature sequence.

[0073] After the encoder using a multimodal network model encodes the preprocessing results to generate a graphic feature sequence, this embodiment can use the decoder of the multimodal network model to decode the graphic feature sequence to generate product information of the product object. The decoder can be an LSTM.

[0074] As an optional implementation, an encoder using a multimodal network model encodes the preprocessing results to generate a text-image feature sequence. This includes: the encoder of the multimodal network model models the correlation between different modal information through a self-attention mechanism and generates attention weights, wherein the correlation between different modal information is the correlation between image features and text features; based on the modeling results and attention weights, a text-image feature sequence is generated, wherein the text-image feature sequence is a feature sequence containing multimodal temporal attention information of image information and text information.

[0075] In this embodiment, when the encoder using a multimodal network model encodes the preprocessing results to generate a sequence of image and text features, the encoder of the multimodal network model can use a self-attention mechanism to model the relationship between features corresponding to different modal information. For example, the relationship between image features and text features can be modeled by the self-attention mechanism, so that the relationship between different modal information can be the relationship between image features and text features. The modeling result is obtained, and attention weights are generated. Then, based on the above modeling result and attention weights, a sequence of image and text features containing image information and text information is generated.

[0076] As an optional implementation, the decoder of the multimodal network model generates product information based on the image and text feature sequence, including: extracting the currently stored descriptive text sequence; and performing cross-entropy loss processing based on the descriptive text sequence and the image and text feature sequence to predict the product information.

[0077] In this embodiment, when the decoder of the multimodal network model generates product information based on the image and text feature sequence, the decoder's input includes two parts: a descriptive text sequence and an image and text feature sequence. The decoder of this embodiment can extract the currently stored descriptive text sequence, which can be historical information of the currently generated descriptive text sequence. Then, it performs cross-entropy loss processing based on the descriptive text sequence and the image and text feature sequence. The image and text feature sequence can be the image and text information of the product object. Cross-entropy loss processing is performed on the descriptive text sequence and the image and text feature sequence using the cross-entropy loss function to predict the product information of the product object. This includes predicting the next word in each description. Finally, by iteratively executing the above steps, a complete text description statement is obtained, and this complete description statement is determined as the product information of the product object.

[0078] As an optional implementation, before the decoder of the multimodal network model performs cross-entropy loss processing based on the descriptive text sequence and the image feature sequence to predict product information, the method further includes: calculating the attention weight between the image feature sequence and the descriptive text sequence based on the self-attention mechanism model in the decoder of the multimodal network model.

[0079] In this embodiment, before the decoder of the multimodal network model performs cross-entropy loss processing based on the descriptive text sequence and the image and text feature sequence to predict the product information, the attention weight between the image and text feature sequence and the descriptive text sequence is also calculated through the self-attention mechanism in the decoder of the multimodal network model. Then, the attention weight, the current historical information of the descriptive text sequence, and the image and text information of the product object are combined to predict the next word of the description through the cross-entropy loss function, so as to generate the product information of the product object.

[0080] As an optional implementation, after generating product information to describe the product object in step S206, the method further includes: generating multiple types of product materials based on the product information; and publishing multiple product materials.

[0081] In this embodiment, after generating product information to describe the product object, such as the product title and selling points of the product object, various types of product materials can be generated. These product materials are the materials needed when publishing the product object, and can be image product materials, video product materials, text product materials, etc. Each type of product material can include the aforementioned product information, thereby publishing multiple product materials.

[0082] As an optional implementation, after generating the product materials to be published, the method further includes: uploading the product materials to be published and extracting multiple product contents to be verified from the product materials to be published; determining whether at least one product content to be verified meets the entry criteria; if it does, the product materials are successfully entered into the publishing template; otherwise, the product content that fails verification is preprocessed, and if the preprocessed product content meets the entry criteria, the product materials are entered into the publishing template.

[0083] In this embodiment, after generating the product materials to be published, the product materials can be uploaded to the product publishing platform. Multiple product contents to be verified are extracted from these materials. These multiple product contents are the content that needs to be entered into the publishing template. At least one product content to be verified can be determined from the multiple product contents. Then, it is determined whether the at least one product content meets the entry standard, which is the standard used to determine whether at least one product content conforms to the specification. If it is determined that the at least one product content meets the entry standard, the product material is successfully verified and can be successfully entered into the publishing template. If it is determined that the at least one product content does not meet the standard, the at least one product content fails verification. The failed product content is then preprocessed, for example, modified or adjusted. It is then determined whether the preprocessed product content meets the entry standard. If it is determined that the preprocessed product content meets the entry standard, the product material can be entered into the publishing template, and the product object can be published to the product publishing platform through the publishing template.

[0084] This invention also provides another method for processing image data from the perspective of human-computer interaction.

[0085] Figure 3 This is a flowchart of another image data processing method according to an embodiment of the present invention. Figure 3 As shown, the method may include the following steps:

[0086] Step S302: Enter the product data of the product object on the input page of the operation interface.

[0087] In the technical solution provided by step S302 of the present invention, the product data includes at least one of the following: product image information, video information, and text information.

[0088] In this embodiment, the operation interface displays an input page, which is used to input product data of a product object. The product object can be a commodity object, such as a new product to be released by a seller. The product data can be used to describe the product object from various perspectives, including product image information, video information, and text information. Among them, image information and video information belong to visual information and can include detailed information such as color and texture inside the product object. The text information can be used to abstractly describe the high-level semantic information of the commodity. The image information, video information, and text information have strong complementary characteristics.

[0089] Step S304: Sensing the copywriting generation command within the operation interface, analyzing product data, and generating multimodal information of the product object.

[0090] In the technical solution provided by step S304 of the present invention, after the product data of the product object is entered into the input page on the operation interface, the copywriting generation instruction is sensed in the operation interface, the product data is analyzed, and multimodal information of the product object is generated. The multimodal information includes: feature sequences of different modal information.

[0091] In this embodiment, a text generation instruction can be received and sensed within the user interface. This instruction analyzes product data and generates multimodal information about the product object, which can be accessed by the user through touch on the interface. Optionally, when sensing the text generation instruction and analyzing the product data, this embodiment may detect the product data and generate keywords for the product object. These keywords can be used to characterize the product object's features. Then, based on the product data and the product object's keywords, a combination process is performed to generate multimodal information about the product object. This multimodal information may include feature sequences of different modalities, including image feature sequences and text feature sequences. Optionally, this multimodal information includes image information and text information about the product object.

[0092] Step S306: Display product information describing the product object on the operation interface.

[0093] In the technical solution provided by step S306 of the present invention, after generating the multimodal information of the product object, product information describing the product object can be displayed on the operation interface. This product information is generated by processing multimodal information using a multimodal network model.

[0094] In this embodiment, a multimodal network model can be used to fully learn the input multimodal information using an automatic learning method, thereby generating accurate product information. This product information can be text used to describe the product object, such as commodity information, which may include, but is not limited to, the product title, product selling points, etc., and then the product information describing the product object can be displayed on the operation interface.

[0095] This embodiment utilizes a multimodal network model to comprehensively leverage multimodal information, resulting in higher accuracy and more reasonable descriptions of product information displayed on the user interface. Optionally, this embodiment can automatically fill in product information into the required information template when publishing a product object via the user interface, thereby reducing the time sellers spend manually filling in product information and improving the efficiency of product publishing.

[0096] As an optional implementation, after product information describing the product object is displayed on the operation interface in step S306, the method further includes: popping up guidance information on the operation interface, wherein the guidance information includes defect information of the product information; displaying creative materials generated based on the guidance information on the operation interface, wherein the creative materials are the basic information constituting the product materials; generating multiple types of product materials based on the creative materials; and publishing multiple product materials.

[0097] In this embodiment, after displaying product information describing the product object on the operation interface, guidance information can also pop up on the operation interface. This guidance information may include defect information of the product information, which is used to indicate problems in generating product materials and can be used to guide the generation of creative materials. The creative materials are the basic information constituting the product materials. This embodiment can generate creative materials based on the above guidance information; for example, by identifying and filling in gaps based on the guidance information, creative materials can be generated and then displayed on the operation interface.

[0098] After the creative materials generated based on the guidance information are displayed on the operation interface, various types of product materials can be generated based on the creative materials. These product materials are the materials needed when publishing product objects, and can be image product materials, video product materials, text product materials, etc. Each type of product material can include the above product information, thereby publishing multiple product materials.

[0099] This invention also provides another method for processing image data from the perspective of human-computer interaction.

[0100] Figure 4 This is a flowchart of another image data processing method according to an embodiment of the present invention. Figure 4 As shown, the method may include the following steps:

[0101] Step S402: Display the product data of the product object on the interactive interface.

[0102] In the technical solution provided by step S402 of the present invention, the product data includes at least one of the following: product image information, video information, and text information.

[0103] In this embodiment, product data of a product object is acquired, and then the acquired product data is displayed on the interactive interface. The product object can be a commodity object, such as a new product to be released by a seller. The product data can be used to describe the product object from multiple different perspectives, including product image information, video information, and text information. Among them, the image information and video information belong to visual information and can include detailed information such as color and texture inside the product object. The text information can be used to abstractly describe the high-level semantic information of the commodity. The image information, video information, and text information have strong complementary characteristics.

[0104] Step S404: The text generation command is sensed within the interactive interface.

[0105] In the technical solution provided by step S404 of the present invention, after displaying the product data of the product object on the interactive interface, a text generation command is sensed within the interactive interface.

[0106] In this embodiment, a copywriting generation instruction can be received and sensed within the operation interface. This instruction is used to analyze product data and generate multimodal information about the product object, which can be obtained by the user through touch on the interactive interface.

[0107] Step S406: Respond to the copywriting generation instruction, analyze product data, and generate multimodal information of the product object.

[0108] In the technical solution provided by step S406 of the present invention, after sensing the text generation instruction in the interactive interface, the system responds to the text generation instruction, analyzes product data, and generates multimodal information of the product object. The multimodal information includes feature sequences of different modal information, which may include image feature sequences and text feature sequences.

[0109] In this embodiment, after sensing the text generation instruction, the system can respond to the instruction by analyzing product data. This may involve detecting the product data and generating keywords for the product object. These keywords can be used to characterize the product object's features. Then, based on the product data and the product object's keywords, the system performs combination processing to generate multimodal information about the product object. This multimodal information may include the relationships between different modalities. Optionally, this multimodal information includes image information and text information about the product object.

[0110] Step S408: Output a selection page on the interactive interface. The selection page provides at least one text option.

[0111] In the technical solution provided by step S408 of the present invention, after generating the multimodal information of the product object, a selection page is output on the interactive interface. The selection page provides at least one text option, wherein different text options are used to characterize different processing models for different modal information.

[0112] In this embodiment, a selection page can be output and displayed on the interactive interface. At least one text option is displayed at different positions on the selection page for the user to select. The different text options can be used to characterize the processing model used when processing modal information of different modalities. The processing model may include a multimodal network model.

[0113] Step S410: Display product information describing the product object on the interactive interface.

[0114] In the technical solution provided by step S410 of the present invention, after providing at least one text option on the selection page, product information describing the product object is displayed on the interactive interface. Based on the selected text option, a multimodal network model is used to process multimodal information and generate product information.

[0115] In this embodiment, based on the selected text option, it can be determined that the processing model used when processing multimodal information is a multimodal network model, which can be an end-to-end model, and can be regarded as an encoder-decoder structure. It can use automatic learning methods to fully learn the input multimodal information to generate accurate product information. The product information can be text, used to describe the product object, and then the above product information is displayed on the interactive interface.

[0116] This embodiment utilizes a multimodal network model to comprehensively leverage multimodal information, resulting in more accurate descriptions of product information displayed on the interactive interface. Optionally, this embodiment automatically fills in the generated product information into the required information template when publishing a product object, thereby reducing the time sellers spend manually filling in product information and improving the efficiency of product publishing.

[0117] This invention also provides another method for processing image data from the front-end client side.

[0118] Figure 5 This is a flowchart of another image data processing method according to an embodiment of the present invention. Figure 5 As shown, the method may include the following steps:

[0119] Step S502: The front-end client uploads the product data of the product object.

[0120] In the technical solution provided by step S502 of the present invention, the product data includes at least one of the following: product image information, video information, and text information.

[0121] In this embodiment, the front-end client can be a merchant publishing end, which can receive upload operation instructions on the operation interface, respond to the upload operation instructions, and start uploading product data of the product object. The product object can be a commodity object, and the product data can be used to describe the product object from multiple different perspectives. It can include product image information, video information, and text information. The image information and video information belong to visual information and can include detailed information such as color and texture inside the product object. The text information can be used to abstractly describe the high-level semantic information of the commodity. The image information, video information, and text information have strong complementary characteristics.

[0122] In step S504, the front-end client transmits the product data of the product object to the back-end server.

[0123] In the technical solution provided by step S504 of the present invention, after the front-end client uploads the product data of the product object, the front-end client can transmit the product data of the product object to the back-end server.

[0124] In this embodiment, a communication connection is established between the front-end client and the back-end server, which can transmit product data of the product object to the back-end server so that the back-end server can process the product data.

[0125] In step S506, the front-end client receives multimodal information generated by analyzing product data returned by the back-end server.

[0126] In the technical solution provided by step S506 of the present invention, after the front-end client transmits the product data of the product object to the back-end server, the front-end client receives the multimodal information generated by analyzing the product data returned by the back-end server. The multimodal information includes the feature sequence of different modal information of the product object.

[0127] In this embodiment, after the backend server receives the product data of the product object, it can analyze the product data. Optionally, the backend server in this embodiment can detect the product data, generate keywords for the product object, which can be used to characterize the characteristics of the product object, and then combine the product data and the keywords of the product object to generate multimodal information of the product object, which can include the correlation between different modal information.

[0128] After the backend server generates multimodal information, the frontend client receives the multimodal information generated by the backend server through analysis of product data.

[0129] In step S508, the front-end client uses a multimodal network model to process multimodal information and generate product information to describe the product object.

[0130] In the technical solution provided by step S508 of the present invention, after the front-end client receives the multimodal information generated by the back-end server through analysis of product data, the front-end client uses a multimodal network model to process the multimodal information and generate product information to describe the product object.

[0131] In this embodiment, the front-end client can use a multimodal network model to fully learn the input multimodal information through automatic learning methods to generate accurate product information. This product information can be text used to describe the product object, such as commodity information, which may include, but is not limited to, the product title, product selling points, and other information of the product object.

[0132] This embodiment utilizes a multimodal network model to comprehensively leverage multimodal information, resulting in higher accuracy in the product information descriptions generated on the front-end client. Optionally, this embodiment automatically fills in the generated product information into the information template required when publishing a product object, thereby reducing the time sellers spend manually filling in product information and improving the efficiency of product publishing.

[0133] In related technologies, neither unimodal nor multimodal description generation algorithms can fully utilize the complementary relationships between different modal information, resulting in low accuracy of the generated descriptions. This embodiment, however, provides a multimodal-based method for automatically filling in product information. It simultaneously utilizes the multimodal information of the product object as input and, through a self-attention-based multimodal network model, fully learns the relationships between different modal information, thereby generating more accurate product information. This solves the technical problem of low accuracy in product information descriptions and achieves the technical effect of improving the accuracy of product information descriptions.

[0134] Example 2

[0135] The preferred embodiments of the above-described method of this example will be further described below, specifically using commodities as an example.

[0136] In scenarios involving intelligent product publishing, sellers typically need to manually input a large amount of information when publishing new products, including product titles and selling points. Currently, there is a lack of solutions for automatically filling in product information on the publishing platform. This results in sellers spending a significant amount of time and effort on information entry, impacting the efficiency of product publishing. Furthermore, for new sellers, creating accurate and attractive product titles and selling points is also very difficult, thus affecting product exposure and sales.

[0137] In related technologies, CNNs can be used as encoders to model the visual information of images, and then LSTMs can be used as decoders to generate textual descriptions of the images. However, the problem with this approach is that it only models the visual information of the image, neglecting the supplementary role of high-level semantic information in the text, resulting in low accuracy of the generated descriptions.

[0138] In another related technique, a sequence-to-sequence encoder-decoder model can be constructed using long short-term memory networks. The encoder takes keywords from the input text, and the decoder outputs a complete text description. However, this approach has a problem: it relies solely on unimodal textual information to generate the description, lacking the visual information from images, resulting in insufficient detail about the product.

[0139] Another related technique involves constructing a multimodal encoder that uses a convolutional neural network to extract image features and simultaneously extract text structured coding features (word embeddings). After concatenating and fusing the multimodal features, the result is input into a decoder based on a Long Short-Term Memory (LSTM) network, ultimately outputting text descriptions. However, this approach has a problem: it merely performs simple concatenation and fusion of features from different modalities without fully learning the relationships between information from different modalities, resulting in low accuracy in the generated text descriptions.

[0140] As can be seen from the above, in the relevant technologies, the text description method based on single-modal data only models the information of a single modality, such as image or text, in the encoder part, and lacks comprehensive utilization of multimodal information, which results in low accuracy of the generated text description content or insufficient description of product details.

[0141] On the other hand, while the text description generation methods based on multimodal data utilize information from different modalities such as images and text, they merely splice and fuse features from different modalities in a relatively direct manner, ignoring the correlation between information from different modalities, thus resulting in low accuracy of the generated text description content.

[0142] In the context of intelligent product publishing, a key challenge is how to leverage existing, mature product attribute detection and category prediction results, combined with original product images to form a multimodal information input, and then use machine learning methods to automatically generate complete product titles and selling point descriptions from seller-uploaded product images. Different modalities of data can describe products from multiple perspectives. For example, textual information (such as product attributes and categories) can abstractly describe the high-level semantic information of a product, while visual information from images (product images) includes detailed information such as color and texture. Therefore, data from multiple modalities, including images and text, have strong complementary characteristics. Fully utilizing the complementary information from multiple modalities can effectively improve the accuracy of product descriptions. However, current text description generation methods often only model information from a single modality or fail to fully integrate and utilize the complementary characteristics of different modalities, such as images and text. This results in inaccurate descriptions of the product's core selling points or insufficient descriptions of product details.

[0143] This embodiment proposes an end-to-end model-based method for automatically filling in product titles and selling points based on multimodal input. It can simultaneously process multimodal input information, including product images, product attribute text, and product category information. This embodiment can model the correlation between image spatial characteristics and text temporal characteristics through spatiotemporal joint learning of the end-to-end model, generating more accurate and reasonable product titles and selling point descriptions. This reduces the time sellers spend manually filling in product information, thereby improving the efficiency of product listing.

[0144] Figure 6 This is a schematic diagram of a method for processing product image data according to an embodiment of the present invention. Figure 6 As shown, after uploading a product image, the seller can first use the product attribute detection module to detect the product image and obtain the product's attribute keywords. Then, the seller can use the category prediction module to detect the product image and obtain the product's category keywords. Finally, the product image, product attribute keywords, and product category keywords are input into a multimodal transformer network model. The multimodal transformer network model processes the product image, product attribute keywords, and product category keywords to obtain the text description of the product.

[0145] Figure 7 This is a schematic diagram illustrating the processing of the aforementioned product image, product attribute keywords, and product category keywords using a transformer network model according to an embodiment of the present invention. Figure 7As shown, the transformer network model in this embodiment includes an encoder and a decoder. For the input product image, the feature map of the product image can be extracted first using a convolutional neural network ResNet-50, and the extracted product image feature maps can be combined into an image feature sequence. Then, text structured coding features are extracted from the attribute keywords and category keywords of the product, and the extracted text structured coding features are combined into a text feature sequence. Then, the above image feature sequence and text feature sequence are concatenated, and the concatenated result is input into the encoder network of the transformer model. The relationship between image and text features is modeled through a self-attention mechanism, and attention weights are obtained. Then, based on the modeling results and attention weights, a multimodal temporal attention information image and text feature sequence containing image and text is generated.

[0146] In the decoder submodule, the input is divided into two parts: one is the image-text feature sequence (multimodal temporal feature sequence) obtained from the encoder, and the other is the descriptive text sequence currently generated by the decoder. Similarly, a self-attention mechanism is used to calculate the attention weight between the image-text feature sequence and the descriptive text sequence. Then, combining the historical information of the current descriptive text sequence and the image-text information of the product, the next word in the description is predicted using the cross-entropy loss function. Finally, by repeatedly executing the above steps, the complete text description of the product is obtained.

[0147] Figure 8A This is a schematic diagram of the interactive interface of an image data processing method according to an embodiment of the present invention. Figure 8A As shown, users can enter product data for a product object on the input page of the operation interface. The product data includes at least one of the following: product image information (B), video information (P), and text information (I). By clicking the "Generate Product Information" button, the product data is analyzed, and multimodal information of the product object is generated. This multimodal information includes feature sequences of different modalities. Finally, product information describing the product object is displayed on the operation interface. This product information is generated by processing multimodal information using a multimodal transformer network model. This embodiment, by acquiring multimodal information of the product object and processing it based on a multimodal transformer network model, generates more accurate product information describing the product object, solving the technical problem of low accuracy in product information description and achieving the technical effect of improving the accuracy of product information description.

[0148] Figure 8B This is a schematic diagram of a scene illustrating an image data processing method according to an embodiment of the present invention. Figure 8BAs shown, the computing device acquires product data of a product object, wherein the product data includes at least one of the following: product image information, video information, and text information, and can display the above product data on an interactive interface. Then, it senses a copywriting generation command within the interactive interface, responds to the copywriting generation command, analyzes the product data, generates multimodal information of the product object, and can output a selection page on the interactive interface. This selection page provides at least one copywriting option, wherein different copywriting options are used to represent different processing models for different modal information. The multimodal information includes: feature sequences of different modal information. The multimodal information is input into a multimodal transformer network model, which processes the multimodal information to generate product information describing the product object, and then displays the product information describing the product object on the interactive interface.

[0149] In the context of intelligent product publishing, automatically generating product titles, selling points, and other descriptions by combining multimodal information is crucial for improving the efficiency of product publishing for sellers. However, in related technologies, neither unimodal nor multimodal text description generation algorithms can fully utilize the complementary relationships between different modal data, resulting in low accuracy of the generated text descriptions. This embodiment implements a method for automatically filling in product titles and selling points based on multimodal information input. It can simultaneously utilize multiple modal information inputs of the product and, through a transformer network model based on a self-attention mechanism, fully learn the relationships between different modal information, thereby generating more accurate product titles and selling point descriptions. This solves the technical problem of low accuracy in product information descriptions and ultimately improves the accuracy of product information descriptions.

[0150] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, because according to the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.

[0151] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0152] Example 3

[0153] According to an embodiment of the present invention, an image data processing apparatus for implementing the above-described image data processing method is also provided. It should be noted that the image data processing apparatus of this embodiment can be used to execute the present invention. Figure 2 The image data processing method of the embodiment shown.

[0154] Figure 9 This is a schematic diagram of an image data processing apparatus according to an embodiment of the present invention. Figure 9 As shown, the image data processing device 90 may include: an acquisition unit 91, a first processing unit 92, and a second processing unit 93.

[0155] The acquisition unit 91 is used to acquire product data of the product object, wherein the product data includes at least one of the following: product image information, video information, and text information.

[0156] The first processing unit 92 is used to analyze product data and generate multimodal information of the product object, wherein the multimodal information includes feature sequences of different modal information.

[0157] The second processing unit 93 is used to process multimodal information using a multimodal network model to generate product information for describing product objects.

[0158] It should be noted that the aforementioned acquisition unit 91, first processing unit 92, and second processing unit 93 correspond to steps S202 to S206 in Embodiment 1. The three units and their corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should also be noted that the aforementioned units, as part of the device, can operate in the computer terminal 10 provided in Embodiment 1.

[0159] According to an embodiment of the present invention, another image data processing apparatus for implementing the above-described image data processing method is also provided. It should be noted that the image data processing apparatus of this embodiment can be used to execute the present invention. Figure 3 The image data processing method of the embodiment shown.

[0160] Figure 10 This is a schematic diagram of another image data processing apparatus according to an embodiment of the present invention. Figure 10 As shown, the image data processing device 100 may include: an input unit 101, a third processing unit 102, and a first display unit 103.

[0161] The input unit 101 is used to input product data of a product object on the input page of the operation interface, wherein the product data includes at least one of the following: product image information, video information, and text information.

[0162] The third processing unit 102 is used to sense the copywriting generation instruction within the operation interface, analyze product data, and generate multimodal information of the product object, wherein the multimodal information includes: feature sequences of different modal information.

[0163] The first display unit 103 is used to display product information describing the product object on the operation interface, wherein the product information is generated by processing multimodal information using a multimodal network model.

[0164] It should be noted that the above-mentioned input unit 101, third processing unit 102, and first display unit 103 correspond to steps S302 to S306 in Embodiment 1. The three units and their corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above-mentioned units, as part of the device, can run in the computer terminal 10 provided in Embodiment 1.

[0165] According to an embodiment of the present invention, another image data processing apparatus for implementing the above-described image data processing method is also provided. It should be noted that the image data processing apparatus of this embodiment can be used to execute the present invention. Figure 4 The image data processing method of the embodiment shown.

[0166] Figure 11 This is a schematic diagram of another image data processing apparatus according to an embodiment of the present invention. Figure 11 As shown, the image data processing device 110 may include: a second display unit 111, a sensing unit 112, a fourth processing unit 113, an output unit 114, and a third display unit 115.

[0167] The second display unit 111 is used to display product data of a product object on an interactive interface, wherein the product data includes at least one of the following: product image information, video information, and text information.

[0168] The sensing unit 112 is used to sense the text generation command within the interactive interface.

[0169] The fourth processing unit 113 is used to respond to the copywriting generation instruction, analyze product data, and generate multimodal information of the product object, wherein the multimodal information includes: feature sequences of different modal information.

[0170] The output unit 114 is used to output a selection page on the interactive interface. The selection page provides at least one text option, wherein different text options are used to represent different processing models for modal information of different modalities.

[0171] The third display unit 115 is used to display product information describing the product object on the interactive interface, wherein multimodal information is processed using a multimodal network model based on the selected text option to generate product information.

[0172] It should be noted that the second display unit 111, sensing unit 112, fourth processing unit 113, output unit 114, and third display unit 115 mentioned above correspond to steps S402 to S410 in Embodiment 1. The five units and their corresponding steps implement the same examples and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above units, as part of the device, can operate in the computer terminal 10 provided in Embodiment 1.

[0173] According to an embodiment of the present invention, another image data processing apparatus for implementing the above-described image data processing method is also provided. It should be noted that the image data processing apparatus of this embodiment can be used to execute the present invention. Figure 5 The image data processing method of the embodiment shown.

[0174] Figure 12 This is a schematic diagram of another image data processing apparatus according to an embodiment of the present invention. Figure 12 As shown, the image data processing device 120 may include: an uploading unit 121, a transmission unit 122, a receiving unit 123, and a fifth processing unit 124.

[0175] Upload unit 121 is used to enable the front-end client to upload product data of the product object, wherein the product data includes at least one of the following: product image information, video information, and text information.

[0176] The transmission unit 122 is used to enable the front-end client to transmit product data of the product object to the back-end server.

[0177] The receiving unit 123 is used to enable the front-end client to receive multimodal information generated by the back-end server for analyzing product data. The multimodal information includes feature sequences of different modal information of the product object.

[0178] The fifth processing unit 124 is used to enable the front-end client to process multimodal information using a multimodal network model and generate product information to describe the product object.

[0179] It should be noted that the above-mentioned uploading unit 121, transmitting unit 122, receiving unit 123, and fifth processing unit 124 correspond to steps S502 to S508 in Embodiment 1. The five units and their corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above-mentioned units, as part of the device, can run in the computer terminal 10 provided in Embodiment 1.

[0180] In the image data processing apparatus of this embodiment, by acquiring multimodal information of the product object and comprehensively processing the multimodal information based on a multimodal network model, more accurate product information for describing the product object is generated, thus solving the technical problem of low accuracy of product information description and achieving the technical effect of improving the accuracy of product information description.

[0181] Example 4

[0182] Embodiments of the present invention can provide an image data processing system, which may include a computer terminal, which may be any computer terminal device from a group of computer terminals. Optionally, in this embodiment, the computer terminal may also be replaced by a mobile terminal or other terminal device.

[0183] Optionally, in this embodiment, the computer terminal may be located in at least one of a plurality of network devices in a computer network.

[0184] In this embodiment, the computer terminal described above can execute the program code for the following steps in the image data processing method of the application: obtaining product data of the product object, wherein the product data includes at least one of the following: product image information, video information, and text information; analyzing the product data to generate multimodal information of the product object, wherein the multimodal information includes: feature sequences of different modal information; and processing the multimodal information using a multimodal network model to generate product information for describing the product object.

[0185] Optionally, Figure 13 This is a structural block diagram of a computer terminal according to an embodiment of the present invention. Figure 13As shown, the computer terminal A may include one or more (only one is shown in the figure) processors 1302, memory 1304, and transmission devices 1306.

[0186] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the image data processing method and apparatus in this embodiment of the invention. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby realizing the aforementioned image data processing method. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to the mobile terminal A via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0187] The processor can invoke information and application programs stored in the memory via a transmission device to perform the following steps: acquiring product data of a product object, wherein the product data includes at least one of the following: product image information, video information, and text information; analyzing the product data to generate multimodal information of the product object, wherein the multimodal information includes: feature sequences of different modal information; and processing the multimodal information using a multimodal network model to generate product information describing the product object.

[0188] Optionally, the processor may also execute program code that performs the following steps: In the process of processing multimodal information, the multimodal network model generates product information by learning the correlation between different modal information.

[0189] Optionally, the processor may also execute program code that performs the following steps: performs attribute detection and category prediction on product data to generate attribute keywords and category keywords for product objects; and preprocesses different modal information based on product data, attribute keywords and category keywords for product objects to generate multimodal information for product objects.

[0190] Optionally, the processor may also execute program code for the following steps: the encoder of the multimodal network model uses a convolutional neural network model to extract features from product images and videos, generating image feature sequences; the encoder of the multimodal network model extracts text structured coding features from attribute keywords and category keywords of product objects, generating text feature sequences; and the image feature sequences and text feature sequences are concatenated to generate preprocessed results.

[0191] Optionally, the processor may also execute program code for the following steps: encoding the preprocessing result using an encoder of a multimodal network model to generate a text-image feature sequence, wherein the text-image feature sequence is a feature sequence containing multimodal temporal attention information of images and text; and generating product information based on the text-image feature sequence using a decoder of the multimodal network model.

[0192] Optionally, the processor may also execute program code for the following steps: the encoder of the multimodal network model models the correlation between different modal information through a self-attention mechanism and generates attention weights, wherein the correlation between different modal information is the correlation between image features and text features; based on the modeling results and attention weights, a sequence of image and text features is generated, wherein the sequence of image and text features is a feature sequence containing multimodal temporal attention information of image information and text information.

[0193] Optionally, the processor may also execute program code that performs the following steps: extracts the currently stored descriptive text sequence; the decoder of the multimodal network model performs cross-entropy loss processing based on the descriptive text sequence and the image feature sequence to predict product information.

[0194] Optionally, the processor may also execute program code that performs the following steps: before the decoder of the multimodal network model performs cross-entropy loss processing based on the descriptive text sequence and the image feature sequence to predict product information, it calculates the attention weight between the image feature sequence and the descriptive text sequence based on the self-attention mechanism model in the decoder of the multimodal network model.

[0195] Optionally, the processor may also execute program code that performs the following steps: after generating product information to describe the product object, generating multiple types of product materials based on the product information; and publishing multiple product materials.

[0196] Optionally, the processor may also execute program code that performs the following steps: after generating the product material to be published, upload the product material to be published and extract multiple product contents to be verified from the product material to be published; determine whether at least one product content to be verified meets the entry criteria; if it does, successfully enter the product material into the publishing template; otherwise, preprocess the product content that fails verification, and enter the product material into the publishing template if the preprocessed product content meets the entry criteria.

[0197] As an alternative example, the processor can invoke information and applications stored in memory via a transmission device to perform the following steps: inputting product data of a product object into an input page on the operation interface, wherein the product data includes at least one of the following: product image information, video information, and text information; sensing a copywriting generation instruction within the operation interface, analyzing the product data, and generating multimodal information of the product object, wherein the multimodal information includes: feature sequences of different modal information; and displaying product information describing the product object on the operation interface, wherein the product information is generated by processing multimodal information using a multimodal network model.

[0198] Optionally, the processor may also execute program code that performs the following steps: after displaying product information describing the product object on the operation interface, popping up guidance information on the operation interface, wherein the guidance information includes defect information of the product information; displaying creative materials generated based on the guidance information on the operation interface, wherein the creative materials are the basic information constituting the product materials; generating multiple types of product materials based on the creative materials; and publishing multiple product materials.

[0199] As an alternative example, the processor can invoke information and applications stored in memory via a transmission device to perform the following steps: displaying product data of a product object on an interactive interface, wherein the product data includes at least one of the following: product image information, video information, and text information; sensing a copywriting generation instruction within the interactive interface; responding to the copywriting generation instruction, analyzing the product data, and generating multimodal information of the product object, wherein the multimodal information includes: feature sequences of different modal information; outputting a selection page on the interactive interface, the selection page providing at least one copywriting option, wherein different copywriting options are used to characterize different processing models for different modal information; displaying product information describing the product object on the interactive interface, wherein, based on the selected copywriting option, a multimodal network model is used to process the multimodal information to generate product information.

[0200] As an alternative example, the processor can invoke information and applications stored in memory via a transmission device to perform the following steps: a front-end client uploads product data of a product object, wherein the product data includes at least one of the following: image information, video information, and text information of the product; the front-end client transmits the product data of the product object to a back-end server; the front-end client receives multimodal information generated by analyzing the product data returned by the back-end server, wherein the multimodal information includes: feature sequences of different modal information of the product object; the front-end client processes the multimodal information using a multimodal network model to generate product information describing the product object.

[0201] This invention provides a scheme for processing image data. By acquiring product data of a product object, wherein the product data includes at least one of the following: image information, video information, and text information of the product; analyzing the product data to generate multimodal information of the product object, wherein the multimodal information includes: feature sequences of different modal information; and processing the multimodal information using a multimodal network model to generate product information describing the product object. This application, by acquiring multimodal information of a product object and comprehensively processing the multimodal information based on a multimodal network model, generates more accurate product information describing the product object, solving the technical problem of low accuracy in product information description and achieving the technical effect of improving the accuracy of product information description.

[0202] Those skilled in the art will understand that Figure 13 The structure shown is for illustrative purposes only. Computer terminal A can also be a smartphone (such as an Android phone, iOS phone, etc.), tablet computer, mobile internet device (MID), PAD and other terminal devices. Figure 13 This does not limit the structure of the aforementioned computer terminal A. For example, computer terminal A may also include components that are more complex than those described above. Figure 13 The more or fewer components shown (such as network interfaces, display devices, etc.), or having the same Figure 13 The different configurations shown.

[0203] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0204] Example 5

[0205] Embodiments of the present invention also provide a computer-readable storage medium. Optionally, in this embodiment, the computer-readable storage medium can be used to store the program code executed by the image data processing method provided in Embodiment 1.

[0206] Optionally, in this embodiment, the computer-readable storage medium may be located in any computer terminal in a group of computer terminals in a computer network, or in any mobile terminal in a group of mobile terminals.

[0207] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: obtaining product data of a product object, wherein the product data includes at least one of the following: image information, video information, and text information of the product; analyzing the product data to generate multimodal information of the product object, wherein the multimodal information includes: feature sequences of different modal information; and processing the multimodal information using a multimodal network model to generate product information describing the product object.

[0208] Optionally, the computer-readable storage medium is also configured to store program code for performing the following steps: In the process of processing multimodal information, the multimodal network model generates product information by learning the correlation between different modal information.

[0209] Optionally, the computer-readable storage medium is further configured to store program code for performing the following steps: performing attribute detection and category prediction on product data to generate attribute keywords and category keywords for product objects; preprocessing different modal information based on product data, attribute keywords, and category keywords for product objects to generate multimodal information for product objects. Optionally, the computer-readable storage medium is further configured to store program code for performing the following steps: the encoder of the multimodal network model uses a convolutional neural network model to extract image features from product images and videos to generate image feature sequences; the encoder of the multimodal network model extracts text structured coding features from the attribute keywords and category keywords of product objects to generate text feature sequences; and concatenates the image feature sequences and text feature sequences to generate a preprocessed result.

[0210] Optionally, the computer-readable storage medium is further configured to store program code for performing the following steps: preprocessing different modal information using an encoder employing a multimodal network model; encoding the preprocessing results using an encoder employing a multimodal network model to generate a text-image feature sequence, wherein the text-image feature sequence is a feature sequence containing multimodal temporal attention information of images and text; and generating product information based on the text-image feature sequence using a decoder employing a multimodal network model.

[0211] Optionally, the computer-readable storage medium is further configured to store program code for performing the following steps: the encoder of the multimodal network model models the correlation between different modal information through a self-attention mechanism and generates attention weights, wherein the correlation between different modal information is the correlation between image features and text features; based on the modeling results and attention weights, a sequence of image and text features is generated, wherein the sequence of image and text features is a feature sequence containing multimodal temporal attention information of image information and text information.

[0212] Optionally, the computer-readable storage medium is also configured to store program code for performing the following steps: extracting the currently pre-stored descriptive text sequence; and the decoder of the multimodal network model performing cross-entropy loss processing based on the descriptive text sequence and the image feature sequence to predict product information.

[0213] Optionally, the computer-readable storage medium is also configured to store program code for performing the following steps: before the decoder of the multimodal network model performs cross-entropy loss processing based on the descriptive text sequence and the image feature sequence to predict product information, the attention weight between the image feature sequence and the descriptive text sequence is calculated based on the self-attention mechanism model in the decoder of the multimodal network model.

[0214] Optionally, the computer-readable storage medium is also configured to store program code for performing the following steps: after generating product information to describe the product object, generating multiple types of product materials based on the product information; and publishing multiple product materials.

[0215] Optionally, the computer-readable storage medium is further configured to store program code for performing the following steps: after generating product materials to be published, uploading the product materials to be published and extracting multiple product contents to be verified from the product materials to be published; determining whether at least one product content to be verified meets the entry criteria; if it does, successfully entering the product materials into the publishing template; otherwise, preprocessing the product content that fails verification, and entering the product materials into the publishing template if the preprocessed product content meets the entry criteria.

[0216] As an optional example, the storage medium is configured to store program code for performing the following steps: entering product data of a product object on an input page of the operation interface, wherein the product data includes at least one of the following: product image information, video information, and text information; sensing a copywriting generation instruction within the operation interface, analyzing the product data, and generating multimodal information of the product object, wherein the multimodal information includes: feature sequences of different modal information; and displaying product information describing the product object on the operation interface, wherein the product information is generated by processing the multimodal information using a multimodal network model.

[0217] Optionally, the computer-readable storage medium is further configured to store program code for performing the following steps: after displaying product information describing the product object on the operation interface, popping up guidance information on the operation interface, wherein the guidance information includes defect information of the product information; displaying creative materials generated based on the guidance information on the operation interface, wherein the creative materials are the basic information constituting the product materials; generating multiple types of product materials based on the creative materials; and publishing multiple product materials.

[0218] As an optional example, a computer-readable storage medium is configured to store program code for performing the following steps: displaying product data of a product object on an interactive interface, wherein the product data includes at least one of the following: image information, video information, and text information of the product; sensing a copywriting generation instruction within the interactive interface; responding to the copywriting generation instruction, analyzing the product data, and generating multimodal information of the product object, wherein the multimodal information includes: feature sequences of different modal information; outputting a selection page on the interactive interface, the selection page providing at least one copywriting option, wherein different copywriting options are used to characterize different processing models for different modal information; displaying product information describing the product object on the interactive interface, wherein, based on the selected copywriting option, a multimodal network model is used to process the multimodal information to generate product information.

[0219] As an optional example, a computer-readable storage medium is configured to store program code for performing the following steps: a front-end client uploads product data of a product object, wherein the product data includes at least one of the following: image information, video information, and text information of the product; the front-end client transmits the product data of the product object to a back-end server; the front-end client receives multimodal information generated by analyzing the product data returned by the back-end server, wherein the multimodal information includes: feature sequences of different modal information of the product object; the front-end client processes the multimodal information using a multimodal network model to generate product information describing the product object.

[0220] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0221] In the above embodiments of the present invention, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0222] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.

[0223] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0224] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0225] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0226] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A method for processing image data, characterized in that, include: Obtain product data of the product object, wherein the product data includes at least one of the following: product image information, video information, and text information; Analyze the product data to generate multimodal information of the product object, wherein the multimodal information includes: feature sequences of different modal information; The multimodal information is processed using a multimodal network model to generate product information describing the product object; In the process of processing the multimodal information, the multimodal network model learns the correlation between different modal information; and uses the correlation and complementary information between the multimodal information to generate the product information, wherein the product information is used to describe the product object; The multimodal network model includes an encoder and a decoder. The encoder models the relationships between different modal information using a self-attention mechanism to obtain a modeling result and attention weights. Based on the modeling result and the attention weights, a graphic feature sequence containing image and text information is generated. The decoder calculates the attention weight between the graphic feature sequence and the descriptive text sequence using a self-attention mechanism. Based on the attention weights, the historical information of the descriptive text sequence, and the graphic information corresponding to the product, the product information of the product object is obtained. The step of analyzing the product data to generate multimodal information of the product object includes: detecting the product data to obtain keywords of the product object, wherein the keywords are used to characterize the features of the product object; and combining the product data and the keywords of the product object to obtain the multimodal information of the product object.

2. The method according to claim 1, characterized in that, Analyze the product data to generate multimodal information about the product object, including: Perform attribute detection and category prediction on the product data to generate attribute keywords and category keywords for the product object; Based on the product data, the attribute keywords and category keywords of the product object, the different modal information is preprocessed to generate multimodal information of the product object.

3. The method according to claim 2, characterized in that, Based on the product data, the attribute keywords and category keywords of the product object, preprocessing of different modal information is performed, including: The encoder of the multimodal network model uses a convolutional neural network model to extract features from the images and videos of the product, generating an image feature sequence. The encoder of the multimodal network model extracts the text structured encoding features from the attribute keywords and category keywords of the product object, and generates a text feature sequence. The image feature sequence and the text feature sequence are concatenated to generate a preprocessed result.

4. The method according to claim 2, characterized in that, The multimodal information is processed using a multimodal network model to generate product information describing the product object, including: The preprocessing result is encoded using the encoder of the multimodal network model to generate an image-text feature sequence, wherein the image-text feature sequence is a feature sequence containing multimodal temporal attention information of images and text; The decoder of the multimodal network model generates the product information based on the image and text feature sequence.

5. The method according to claim 4, characterized in that, The encoder of the aforementioned multimodal network model encodes the preprocessed results to generate a sequence of image and text features, including: The encoder of the multimodal network model models the correlation between different modal information through a self-attention mechanism and generates attention weights, wherein the correlation between different modal information is the correlation between image features and text features; Based on the modeling results and attention weights, the image and text feature sequence is generated, wherein the image and text feature sequence is a feature sequence containing multimodal temporal attention information that includes image information and text information.

6. The method according to any one of claims 4 to 5, characterized in that, The decoder of the multimodal network model generates the product information based on the image and text feature sequence, including: Extract the currently stored description text sequence; The decoder of the multimodal network model performs cross-entropy loss processing based on the descriptive text sequence and the image feature sequence to predict the product information.

7. The method according to claim 6, characterized in that, Before the decoder of the multimodal network model performs cross-entropy loss processing based on the descriptive text sequence and the image-text feature sequence to predict the product information, the method further includes: The attention weights between the image feature sequence and the descriptive text sequence are calculated based on the self-attention mechanism model in the decoder of the multimodal network model.

8. The method according to claim 1, characterized in that, After generating product information to describe the product object, the method further includes: Based on the product information, various types of product materials are generated; Multiple product materials mentioned above were released.

9. The method according to claim 8, characterized in that, After generating the product materials to be published, the method further includes: Upload the product materials to be published, and extract multiple product contents to be verified from the product materials to be published; Determine whether the content of at least one product to be verified meets the entry criteria; If the conditions are met, the product materials have been successfully entered into the publishing template; Otherwise, the product content that fails verification is preprocessed, and if the preprocessed product content meets the entry criteria, the product material is entered into the publishing template.

10. The method according to any one of claims 1 to 5 or 7-9, characterized in that, The multimodal network model is a multimodal transformer network model.

11. A method for processing image data, characterized in that, include: Enter product data of the product object on the input page of the operation interface, wherein the product data includes at least one of the following: product image information, video information and text information; The system senses a text generation command within the operation interface, analyzes the product data, and generates multimodal information about the product object, wherein the multimodal information includes feature sequences of different modalities. The user interface displays product information describing the product object, wherein the product information is generated by processing the multimodal information using a multimodal network model. In the process of processing the multimodal information, the multimodal network model learns the correlation between different modal information; and uses the correlation and complementary information between the multimodal information to generate the product information, wherein the product information is used to describe the product object; The multimodal network model includes an encoder and a decoder. The encoder models the relationships between different modal information using a self-attention mechanism to obtain a modeling result and attention weights. Based on the modeling result and the attention weights, a graphic feature sequence containing image and text information is generated. The decoder calculates the attention weight between the graphic feature sequence and the descriptive text sequence using a self-attention mechanism. Based on the attention weights, the historical information of the descriptive text sequence, and the graphic information corresponding to the product, the product information of the product object is obtained. The process of analyzing the product data and generating multimodal information of the product object includes: detecting the product data to obtain keywords of the product object, wherein the keywords are used to characterize the features of the product object; and combining the product data and the keywords of the product object to obtain the multimodal information of the product object.

12. The method according to claim 11, characterized in that, After displaying product information describing the product object on the user interface, the method further includes: A guidance message pops up on the operation interface, wherein the guidance message includes information on defects in the product information; The operation interface displays creative materials generated based on the guidance information, wherein the creative materials are the basic information constituting the product materials; Based on the aforementioned creative materials, various types of product materials can be generated; Multiple product materials mentioned above were released.

13. A method for processing image data, characterized in that, include: The product data of the product object is displayed on the interactive interface, wherein the product data includes at least one of the following: product image information, video information, and text information; The text generation command is sensed within the interactive interface; In response to the copy generation instruction, the product data is analyzed to generate multimodal information of the product object, wherein the multimodal information includes: feature sequences of different modal information; A selection page is output on the interactive interface. The selection page provides at least one text option, wherein different text options are used to characterize different processing models for modal information of different modalities. The interactive interface displays product information describing the product object, wherein the multimodal information is processed using a multimodal network model based on the selected text option to generate the product information; In the process of processing the multimodal information, the multimodal network model learns the correlation between different modal information; and uses the correlation and complementary information between the multimodal information to generate the product information, wherein the product information is used to describe the product object; The multimodal network model includes an encoder and a decoder. The encoder models the relationships between different modal information using a self-attention mechanism to obtain a modeling result and attention weights. Based on the modeling result and the attention weights, a graphic feature sequence containing image and text information is generated. The decoder calculates the attention weight between the graphic feature sequence and the descriptive text sequence using a self-attention mechanism. Based on the attention weights, the historical information of the descriptive text sequence, and the graphic information corresponding to the product, the product information of the product object is obtained. The step of analyzing the product data to generate multimodal information of the product object includes: detecting the product data to obtain keywords of the product object, wherein the keywords are used to characterize the features of the product object; and combining the product data and the keywords of the product object to obtain the multimodal information of the product object.

14. A method for processing image data, characterized in that, include: The front-end client uploads product data of a product object, wherein the product data includes at least one of the following: product image information, video information, and text information; The front-end client transmits the product data of the product object to the back-end server; The front-end client receives multimodal information generated by analyzing the product data returned by the back-end server, wherein the multimodal information includes: feature sequences of different modal information of the product object; The front-end client uses a multimodal network model to process the multimodal information and generate product information to describe the product object; In the process of processing the multimodal information, the multimodal network model learns the correlation between different modal information; and uses the correlation and complementary information between the multimodal information to generate the product information, wherein the product information is used to describe the product object; The multimodal network model includes an encoder and a decoder. The encoder models the relationships between different modal information using a self-attention mechanism to obtain a modeling result and attention weights. Based on the modeling result and the attention weights, a graphic feature sequence containing image and text information is generated. The decoder calculates the attention weight between the graphic feature sequence and the descriptive text sequence using a self-attention mechanism. Based on the attention weights, the historical information of the descriptive text sequence, and the graphic information corresponding to the product, the product information of the product object is obtained. The backend server is used to detect the product data to obtain the keywords of the product object, wherein the keywords are used to characterize the features of the product object; and to combine the product data and the keywords of the product object to obtain the multimodal information of the product object.

15. An image data processing apparatus, characterized in that, include: The acquisition unit is used to acquire product data of a product object, wherein the product data includes at least one of the following: product image information, video information, and text information; The first processing unit is used to analyze the product data and generate multimodal information of the product object, wherein the multimodal information includes: feature sequences of different modal information; The second processing unit is used to process the multimodal information using a multimodal network model to generate product information describing the product object; In the process of processing the multimodal information, the multimodal network model learns the correlation between different modal information; and uses the correlation and complementary information between the multimodal information to generate the product information, wherein the product information is used to describe the product object; The multimodal network model includes an encoder and a decoder. The encoder models the relationships between different modal information using a self-attention mechanism to obtain a modeling result and attention weights. Based on the modeling result and the attention weights, a graphic feature sequence containing image and text information is generated. The decoder calculates the attention weight between the graphic feature sequence and the descriptive text sequence using a self-attention mechanism. Based on the attention weights, the historical information of the descriptive text sequence, and the graphic information corresponding to the product, the product information of the product object is obtained. The first processing unit is further configured to analyze the product data and generate multimodal information of the product object through the following steps: detecting the product data to obtain keywords of the product object, wherein the keywords are used to characterize the characteristics of the product object; and combining the product data and the keywords of the product object to obtain the multimodal information of the product object.

16. An image data processing apparatus, characterized in that, include: The data entry unit is used to enter product data of a product object into the data entry page on the operation interface, wherein the product data includes at least one of the following: product image information, video information, and text information; The third processing unit is used to sense the copywriting generation instruction within the operation interface, analyze the product data, and generate multimodal information of the product object, wherein the multimodal information includes: feature sequences of different modal information; The first display unit is used to display product information describing the product object on the operation interface, wherein the product information is generated by processing the multimodal information using a multimodal network model; In the process of processing the multimodal information, the multimodal network model learns the correlation between different modal information; and uses the correlation and complementary information between the multimodal information to generate the product information, wherein the product information is used to describe the product object; The multimodal network model includes an encoder and a decoder. The encoder models the relationships between different modal information using a self-attention mechanism to obtain a modeling result and attention weights. Based on the modeling result and the attention weights, a graphic feature sequence containing image and text information is generated. The decoder calculates the attention weight between the graphic feature sequence and the descriptive text sequence using a self-attention mechanism. Based on the attention weights, the historical information of the descriptive text sequence, and the graphic information corresponding to the product, the product information of the product object is obtained. The third processing unit is further configured to analyze the product data and generate multimodal information of the product object through the following steps: detecting the product data to obtain keywords of the product object, wherein the keywords are used to characterize the characteristics of the product object; and combining the product data and the keywords of the product object to obtain the multimodal information of the product object.

17. An image data processing apparatus, characterized in that, include: The second display unit is used to display product data of a product object on an interactive interface, wherein the product data includes at least one of the following: product image information, video information, and text information; A sensing unit is used to sense text generation instructions within the interactive interface; The fourth processing unit is used to respond to the copy generation instruction, analyze the product data, and generate multimodal information of the product object, wherein the multimodal information includes: feature sequences of different modal information; An output unit is used to output a selection page on the interactive interface, the selection page providing at least one text option, wherein different text options are used to characterize different processing models for modal information of different modalities; The third display unit is used to display product information describing the product object on the interactive interface, wherein the multimodal information is processed using a multimodal network model based on the selected text option to generate the product information; In the process of processing the multimodal information, the multimodal network model learns the correlation between different modal information; and uses the correlation and complementary information between the multimodal information to generate the product information, wherein the product information is used to describe the product object; The multimodal network model includes an encoder and a decoder. The encoder models the relationships between different modal information using a self-attention mechanism to obtain a modeling result and attention weights. Based on the modeling result and the attention weights, a graphic feature sequence containing image and text information is generated. The decoder calculates the attention weight between the graphic feature sequence and the descriptive text sequence using a self-attention mechanism. Based on the attention weights, the historical information of the descriptive text sequence, and the graphic information corresponding to the product, the product information of the product object is obtained. The fourth processing unit is further configured to analyze the product data and generate multimodal information of the product object through the following steps: detecting the product data to obtain keywords of the product object, wherein the keywords are used to characterize the features of the product object; and combining the product data and the keywords of the product object to obtain the multimodal information of the product object.

18. An image data processing apparatus, characterized in that, include: An upload unit is used to enable the front-end client to upload product data of a product object, wherein the product data includes at least one of the following: product image information, video information, and text information; A transmission unit is used to enable the front-end client to transmit the product data of the product object to the back-end server. A receiving unit is configured to enable the front-end client to receive multimodal information generated by the back-end server through analysis of the product data, wherein the multimodal information includes: feature sequences of different modal information of the product object; The fifth processing unit is used to enable the front-end client to process the multimodal information using a multimodal network model, and generate product information to describe the product object; In the process of processing the multimodal information, the multimodal network model learns the correlation between different modal information; and uses the correlation and complementary information between the multimodal information to generate the product information, wherein the product information is used to describe the product object; The multimodal network model includes an encoder and a decoder. The encoder models the relationships between different modal information using a self-attention mechanism to obtain a modeling result and attention weights. Based on the modeling result and the attention weights, a graphic feature sequence containing image and text information is generated. The decoder calculates the attention weight between the graphic feature sequence and the descriptive text sequence using a self-attention mechanism. Based on the attention weights, the historical information of the descriptive text sequence, and the graphic information corresponding to the product, the product information of the product object is obtained. The backend server is used to detect the product data to obtain the keywords of the product object, wherein the keywords are used to characterize the features of the product object; and to combine the product data and the keywords of the product object to obtain the multimodal information of the product object.

19. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein when the program is run by a processor, it controls the device in which the computer-readable storage medium resides to perform the method according to any one of claims 1 to 14.

20. A processor, characterized in that, The processor is used to run a program, wherein the program, when running, performs the method according to any one of claims 1 to 14.

21. An image data processing system, characterized in that, include: processor; A memory, connected to the processor, is used to provide the processor with instructions to perform the following processing steps: acquiring product data of a product object, wherein the product data includes at least one of the following: image information, video information, and text information of the product; analyzing the product data to generate multimodal information of the product object, wherein the multimodal information includes: feature sequences of different modal information; processing the multimodal information using a multimodal network model to generate product information describing the product object; wherein the multimodal network model learns the correlation relationships between different modal information during the processing of the multimodal information; generating the product information using the correlation relationships and complementary information between the multimodal information, wherein the product information is used to describe the product object; The multimodal network model includes an encoder and a decoder. The encoder models the relationships between different modal information using a self-attention mechanism to obtain a modeling result and attention weights. Based on the modeling result and the attention weights, a graphic feature sequence containing image and text information is generated. The decoder calculates the attention weight between the graphic feature sequence and the descriptive text sequence using a self-attention mechanism. Based on the attention weights, the historical information of the descriptive text sequence, and the graphic information corresponding to the product, the product information of the product object is obtained. The step of analyzing the product data to generate multimodal information of the product object includes: detecting the product data to obtain keywords of the product object, wherein the keywords are used to characterize the features of the product object; and combining the product data and the keywords of the product object to obtain the multimodal information of the product object.