Text generation method and apparatus

By acquiring image and text data of e-commerce products, identifying visual attributes, and combining them with text data to generate descriptive text, the problem of high cost and poor accuracy of manually generated summaries in e-commerce has been solved, achieving efficient and accurate text generation.

CN115496550BActive Publication Date: 2025-12-09ALIBABA (CHINA) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211048016.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-30
Publication Date
2025-12-09
Estimated Expiration
2042-08-30

AI Technical Summary

Technical Problem

In the e-commerce sector, due to the large number of products, manually generating text summaries requires a lot of manpower and high costs, and is subject to uncertainty, resulting in poor accuracy of the generated text summaries.

Method used

By acquiring the image and text data of the target object, identifying visual attribute information, and combining it with text data to determine the object attribute set, the target description text is generated end-to-end.

Benefits of technology

It improves the accuracy and efficiency of text generation, making the generated descriptive text more coherent and accurate, and reducing labor costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115496550B_ABST
    Figure CN115496550B_ABST
Patent Text Reader

Abstract

Embodiments of the present specification provide a text generation method and device, wherein the text generation method comprises: obtaining image-text data of a target object, wherein the image-text data comprises image data and text data; identifying visual attribute information of the target object based on the image data, wherein the visual attribute information represents explicit features of the target object; determining an object attribute set of the target object according to the text data and the visual attribute information; and generating a target description text of the target object based on the object attribute set. By obtaining the multi-modal image-text data of the target object and determining the visual attribute information of the target object, the explicit features of the target object are considered, so that the object attributes of the target object are more comprehensive. Furthermore, by determining the object attribute set of the target object according to the text data and the visual attribute information, the text data and the visual attribute information of the target object are comprehensively considered, so that the generated target description text is more coherent, and the accuracy of the target description text is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present specification relate to the technical field of computer technology, and particularly relate to a text generation method. One or more embodiments of the present specification also relate to a text generation apparatus, a computing device, a computer-readable storage medium, and a computer program. BACKGROUND

[0002] With the development of computer technology, the generation of text summaries has gradually become a hot topic in the field of natural language processing. Taking the e-commerce scenario as an example, in the e-commerce scenario, the description of each commodity is usually composed of a variety of data. In order to better describe the characteristics of the commodity and attract users to purchase, a text summary corresponding to the commodity needs to be generated to enable users to quickly and accurately understand the information of the commodity.

[0003] At present, a host usually fully understands the commodity information and summarizes the prominent features of the commodity. However, since there are a large number of commodities in the e-commerce field, it is necessary to spend a large amount of manpower and pay high costs to obtain a text summary of the commodity by manual summarization. In addition, manual summarization will inevitably introduce a large number of uncertain factors, resulting in poor accuracy of the generated text summary. Therefore, there is an urgent need for an accurate text generation scheme. SUMMARY

[0004] In view of this, the embodiments of the present specification provide a text generation method. One or more embodiments of the present specification also relate to a text generation apparatus, a computing device, a computer-readable storage medium, and a computer program to solve the technical defects in the prior art.

[0005] According to a first aspect of the embodiments of the present specification, a text generation method is provided, comprising:

[0006] obtaining image data and text data of a target object, wherein the image data and the text data of the target object comprise image data and text data;

[0007] identifying visual attribute information of the target object based on the image data, wherein the visual attribute information represents the explicit features of the target object;

[0008] determining an object attribute set of the target object according to the text data and the visual attribute information;

[0009] generating a target description text of the target object based on the object attribute set.

[0010] According to a second aspect of the embodiments of the present specification, a text generation apparatus is provided, comprising:

[0011] an obtaining module configured to obtain image data and text data of a target object, wherein the image data and the text data of the target object comprise image data and text data;

[0012] An identification module is configured to identify visual attribute information of the target object based on the image data, wherein the visual attribute information represents explicit features of the target object.

[0013] A determination module is configured to determine an object attribute set of the target object according to the text data and the visual attribute information.

[0014] A generation module is configured to generate a target description text of the target object based on the object attribute set.

[0015] According to a third aspect of an embodiment of the present specification, a computing device is provided, comprising:

[0016] a memory and a processor;

[0017] The memory is configured to store computer executable instructions, and the processor is configured to execute the computer executable instructions, and the computer executable instructions, when executed by the processor, implement the steps of the text generation method.

[0018] According to a fourth aspect of an embodiment of the present specification, a computer readable storage medium is provided, which stores computer executable instructions, and the instructions, when executed by a processor, implement the steps of the text generation method.

[0019] According to a fifth aspect of an embodiment of the present specification, a computer program is provided, wherein when the computer program is executed in a computer, the computer program causes the computer to execute the steps of the text generation method.

[0020] The text generation method provided by one embodiment of the present specification acquires image-text data of a target object, wherein the image-text data comprises image data and text data; identifies visual attribute information of the target object based on the image data, wherein the visual attribute information represents explicit features of the target object; determines an object attribute set of the target object according to the text data and the visual attribute information; and generates a target description text of the target object based on the object attribute set. By acquiring multi-modal image-text data of the target object, determining visual attribute information of the target object, and considering explicit features of the target object, the object attributes of the target object are more comprehensive. Furthermore, by determining an object attribute set of the target object according to the text data and the visual attribute information, the text data and the visual attribute information of the target object are comprehensively considered, so that the generated target description text is more coherent, and the accuracy of the target description text is further improved. BRIEF DESCRIPTION OF DRAWINGS

[0021] Figure 1 is a framework diagram of a text generation system provided by one embodiment of the present specification;

[0022] Figure 2 is a framework diagram of another text generation system provided by one embodiment of the present specification;

[0023] Figure 3 is a flowchart of a text generation method provided by an embodiment of the present specification;

[0024] Figure 4 is a training flowchart of a text processing model in a text generation method provided by an embodiment of the present specification;

[0025] Figure 5 is a training flowchart of an image classification model in a text generation method provided by an embodiment of the present specification;

[0026] Figure 6 is a processing flowchart of a text generation method provided by an embodiment of the present specification;

[0027] Figure 7 is a schematic diagram of a target commodity detail page in a text generation method provided by an embodiment of the present specification;

[0028] Figure 8 is a display interface schematic diagram of a client in a text generation method provided by an embodiment of the present specification;

[0029] Figure 9 is a structural schematic diagram of a text generation device provided by an embodiment of the present specification;

[0030] Figure 10 is a structural block diagram of a computing device provided by an embodiment of the present specification. DETAILED DESCRIPTION

[0031] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present specification. However, the present specification can be practiced without the specific details, other than in the examples, and it is understood that the present specification will encompass numerous variations beyond those described in the detailed description. It will further be understood that the present specification includes all tweaks and combinations of the described examples.

[0032] The terminology used in one or more embodiments of the present specification is for the purpose of describing particular embodiments only and is not intended to be limiting of one or more embodiments of the present specification. As used in one or more embodiments of the present specification and the accompanying claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in one or more embodiments of the present specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0033] It should be understood that, although the terms first, second, etc. can be employed in describing various information in one or more embodiments of the present specification, the information should not be limited to such terms. These terms are only used to differentiate one piece of information from another piece of information. For example, without departing from the scope of one or more embodiments of the present specification, first can also be referred to as second, and similarly, second can also be referred to as first. Depending on the context, the word "if' as used herein can be interpreted as "when" or "upon" or "in response to determining".

[0034] First, the noun terms related to one or more embodiments of the present specification are explained.

[0035] Modality: refers to the form in which data exists, such as natural language, picture, etc.

[0036] Product summary: based on the information of the product, such as the description, appearance, etc. of the product, a short text summary with the significant information of the product is generated.

[0037] Natural language generation: a function of making a computer have the same expression and writing as a human being. That is, according to some key information and its expression form inside the machine, a high-quality natural language text is automatically generated through a planning process.

[0038] BART (Bidirectional and Auto-Regressive Transformers): a model with both context information and auto-regressive characteristics, which inputs natural language and generates natural language.

[0039] Automatic speech recognition (ASR, Automatic Speech Recognition): a technology that converts human language into corresponding text.

[0040] Part-of-speech tagging: a technology that can mark the part of speech of each word in a sentence.

[0041] Mutual information: the degree of dependence between two random variables.

[0042] In the present specification, a text generation method is provided, and the present specification also relates to a text generation device, a computing device, and a computer-readable storage medium, which are described in detail one by one in the following embodiments.

[0043] With the development of computer technology, the generation of text summary has gradually become a hot topic in the field of natural language processing. For example, in the e-commerce scenario, the description of each commodity is usually composed of a variety of data, such as the title of the commodity, the detailed text description and the image, etc. In order to better describe the characteristics of the commodity and attract users to purchase, it is necessary to generate a text summary corresponding to the commodity to enable users to quickly and accurately understand the information of the commodity.

[0044] At present, the host usually fully understands the commodity information and summarizes the prominent features of the commodity. However, since there are a large number of commodities in the e-commerce field, it takes a lot of manpower and high cost to obtain the text summary of the commodity by manual arrangement, and a large number of uncertain factors will inevitably be introduced by manual arrangement. Most of the text summaries are simply spliced, resulting in poor accuracy of the generated text summary and high modification cost. Therefore, an accurate text generation scheme is urgently needed.

[0045] In order to improve the efficiency and accuracy of text generation, the present scheme provides a scheme for generating description text based on multi-modal data. Given the multi-modal image-text data of a target object, an end-to-end automatic generation can accurately summarize the characteristics of the target object and highlight the advantages of the target object.

[0046] In specific implementation, the text generation method provided by the embodiment of the present specification acquires image-text data of a target object, wherein the image-text data includes image data and text data; identifies visual attribute information of the target object based on the image data, wherein the visual attribute information represents the explicit features of the target object; determines an object attribute set of the target object according to the text data and the visual attribute information; and generates a target description text of the target object based on the object attribute set. By acquiring multi-modal image-text data of the target object and determining visual attribute information of the target object, the explicit features of the target object are considered, so that the object attributes of the target object are more comprehensive. Furthermore, by determining an object attribute set of the target object according to the text data and the visual attribute information, the text data and the visual attribute information of the target object are comprehensively considered, so that the generated target description text is more coherent, and the accuracy of the target description text is further improved.

[0047] Referring to Figure 1 , Figure 1 A framework diagram of a text generation system is shown, wherein the text generation system includes a server and a client:

[0048] Client: sends image-text data of a target object to the server, wherein the image-text data includes image data and text data;

[0049] Server-side: Acquire image and text data of the target object; Based on the image data, identify the visual attribute information of the target object, where the visual attribute information represents the explicit features of the target object; Determine the object attribute set of the target object based on the text data and visual attribute information; Based on the object attribute set, generate the target description text of the target object and send the target description text to the client so that the client can display the target description text.

[0050] Client: Receives and displays the target description text sent by the server, so that users can learn about the target object based on the target description text.

[0051] It is worth noting that the text generation method provided in the embodiments of this specification is generally executed by the server. However, in other embodiments of this specification, the client may also have similar functionality to the server, thereby executing the text generation method provided in the embodiments of this specification. In other embodiments, the text generation method provided in the embodiments of this specification may also be executed jointly by the client and the server.

[0052] The scheme described in this specification involves acquiring image and text data of a target object, where the image and text data includes image data and text data; identifying visual attribute information of the target object based on the image data, whereby the visual attribute information represents the explicit features of the target object; determining the object attribute set of the target object based on the text data and visual attribute information; and generating target description text of the target object based on the object attribute set. By acquiring multimodal image and text data of the target object and determining the visual attribute information of the target object, the explicit features of the target object are considered, making the object attributes of the target object more comprehensive. Furthermore, by determining the object attribute set of the target object based on the text data and visual attribute information, the combined text data and visual attribute information of the target object are used, making the generated target description text more coherent and further improving the accuracy of the target description text.

[0053] The solutions provided in one or more embodiments of this specification can be applied to text generation scenarios, such as e-commerce live streaming, conference scenarios, education scenarios, etc. The specific choice should be made according to the actual situation, and the embodiments of this specification do not limit this in any way.

[0054] See Figure 2 , Figure 2 This specification illustrates a framework diagram of another text generation system provided in one embodiment. The system may include a server 100 and multiple clients 200. The multiple clients 200 can establish communication connections through the server 100. In a text generation scenario, the server 100 provides text generation services between the multiple clients 200, which can act as either senders or receivers, achieving real-time communication through the server 100.

[0055] The user can interact with the server 100 through the client 200 to receive data sent by other clients 200, or send data to other clients 200, etc. In the text generation scenario, the user can publish a data stream to the server 100 through the client 200, and the server 100 pushes the data stream to the clients subscribing to the data stream. The data stream can be, for example, image-text data. In the e-commerce live broadcast scenario, the user can collect image-text data of a target product in real time through the client, and send the image-text data to the server. The server can generate a corresponding product description text according to the image-text data sent by the client, and push the product description text to all live broadcast rooms including the product, so that the host introduces the target product according to the product description text. In the conference scenario, the user can collect image-text data in real time through the client and send it to the server. The server can process the image-text data sent by the client, generate an abstract text, and push the abstract text to the clients of other participants, etc.

[0056] The client 200 and the server 100 establish a connection through a network. The network provides a medium for the communication link between the client and the server. The network can include various connection types, such as wired, wireless communication links, or optical fiber cables, etc. The data transmitted by the client 200 can need to be processed by encoding, transcoding, compression, etc. before being published to the server 100.

[0057] The client 200 can be a browser, an APP (Application), or a web application such as an H5 (HyperText Markup Language 5) application, or a light application (also known as a small program, a lightweight application), or a cloud application, etc. The client 200 can be developed based on the software development kit (SDK) provided by the server for the corresponding service, such as the real-time communication (RTC) SDK. The client 200 can be deployed in an electronic device, and needs to rely on the device or some App in the device to run, etc. The electronic device can have a display screen and support information browsing, etc., such as a personal mobile terminal such as a mobile phone, a tablet computer, a personal computer, etc. Various other types of applications can also be configured in the electronic device, such as human-computer dialogue applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.

[0058] The service end 100 can include a server providing various services, for example, a server providing a communication service for a plurality of clients, for example, a server for background training supporting a model used on a client, for example, a server processing data sent by a client, and the like.

[0059] It should be noted that the service end 100 can be implemented as a distributed server cluster composed of multiple servers, or as a single server. The server can also be a server of a distributed system, or a server combined with a blockchain. The server can also be a cloud server of a cloud service, a cloud database, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content distribution networks (CDN, Content Delivery Network), and big data and artificial intelligence platforms, and the like basic cloud computing services, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.

[0060] Referring to Figure 3 , Figure 3 A flowchart of a text generation method provided by one embodiment of the present specification is shown, which specifically includes the following steps:

[0061] Step 302: Obtain the image-text data of the target object, wherein the image-text data includes image data and text data.

[0062] In one or more embodiments of the present specification, with the development of computer technology, the description form of the target object is also becoming more and more rich, such as the description of the commodity including the title, the detailed text description, and the commodity display image, etc. In order to accurately generate the description text of the target object, the multi-modal data of the target object can be obtained, the multi-modal data can include image data and text data, and further, the target description text of the target object is generated according to the multi-modal image-text data.

[0063] Specifically, the target object refers to an object that needs to generate a target description text, which can also be understood as an object waiting to generate a target description text, including but not limited to commodities, people, landscapes, historical sites, and the like. The image-text data of the target object refers to image data and text data including information related to the target object. The image data can be a design drawing, a photo, a design drawing, and the like of the target object, and the text data can be the name, the structured attribute, the detail information, the process information, and the like of the target object.

[0064] In actual application, there are various ways to obtain the image-text data of the target object, which is specifically selected according to actual conditions, and the present specification does not make any limitation thereto.

[0065] In an optional implementation of the present specification, the image-text data of the target object can be acquired upon receiving the text generation instruction. In one possible manner, the image-text data covering the information of the target object is carried in the text generation instruction. In another possible manner, the unique identifier of the target object is included in the text generation instruction, and the target object can be determined according to the unique identifier, and the image-text data of the target object is further acquired.

[0066] Exemplarily, taking the target object as a target commodity, since there is a large amount of commodity detail information in the detail page of the commodity, and there is context semantic coherence between the entire detail pages, the information covering the target commodity can be completed, therefore, after receiving the text generation instruction, the image-text data of the target object can be acquired from the detail page of the target commodity according to the unique identifier of the target object in the text generation instruction.

[0067] In another optional implementation of the present specification, since the image-text data of the target object is usually variable, the image-text data of the target object can be monitored, and the image-text data of the target object is acquired in real time in the case of image-text data change, and the target description text of the target object is generated, so that the user can immediately query the target description text when the target description text is needed. That is, the step of acquiring the image-text data of the target object can include the following steps:

[0068] monitoring the image-text data of the target object;

[0069] acquiring the image-text data of the target object in the case of image-text data update.

[0070] In the embodiment of the present specification, the update of the image-text data includes addition, deletion, replacement, change and the like, and in the embodiment of the present specification, as long as the image-text data of the target object changes, it is considered that the image-text data of the target object is updated.

[0071] Further, since the generation process of the target description text will take a certain time, in the embodiment of the present specification, the target description text of the target object can also be generated in an offline timing manner. The offline timing manner is to update the target description text of the target object at a specified time.

[0072] It should be noted that before updating the target description text at a specified time, it can be detected whether the image-text data of the target object is changed, that is, when the timing task is started, the image-text data of the current target object is compared with the image-text data of the target object at the last update. If the image-text data is changed, the timing task is triggered, the image-text data of the target object is acquired, and the target description text is generated based on the image-text data; if the image-text data is not changed, the description text of the target object is not updated.

[0073] By applying the scheme of the embodiments of the present specification, the target description text of the target object is automatically generated by monitoring the graphic-text data of the target object and obtaining the graphic-text data of the target object in the case of graphic-text data update, thereby saving the time of the user to obtain the target description text and improving the user experience.

[0074] In step 304, the visual attribute information of the target object is identified based on the image data, wherein the visual attribute information represents the explicit features of the target object.

[0075] In one or more embodiments of the present specification, after obtaining the graphic-text data of the target object, the visual attribute information of the target object can be further identified based on the image data included in the graphic-text data. By generating the visual attribute information, the image data is converted into text data, the multi-modal data of the target object is unified, and the modal heterogeneity between the multiple modalities is reduced.

[0076] Specifically, the visual attribute information represents the explicit features of the target object. The explicit features refer to the features of the target object, which can be the noun features of the target object such as color, shape, etc., and can also be the adjective features such as beautiful, elegant, generous, etc. The specific selection is based on the actual situation, and the embodiments of the present specification do not make any limitation thereto.

[0077] In practical applications, there are various ways to identify the visual attribute information of the target object based on the image data, and the specific selection is based on the actual situation, and the embodiments of the present specification do not make any limitation thereto.

[0078] In an optional implementation of the present specification, since the image data can include the text data of the target object, the text data in the image data can be obtained by using optical character recognition (OCR). The visual attribute information in the image data can also be obtained by using an image color recognition tool.

[0079] In another optional implementation of the present specification, the pre-trained picture classification model can be used to identify the visual attribute information of the target object. That is, the step of identifying the visual attribute information of the target object based on the image data can include the following steps:

[0080] The image data is input into the pre-trained picture classification model, and the visual attribute information of the target object is obtained through the classification and identification of the picture classification model.

[0081] Specifically, the pre-trained picture classification model is a model generated by training a preset classification model. The preset classification model refers to a model capable of achieving classification, such as a Swin Transformer model, a residual neural network (ResNet), or a vision transformer (Vit). The actual situation is selected, and the embodiments of the present specification do not make any limitation thereto.

[0082] Taking the vision transformer as an example, the image data is input into the vision transformer. Unlike the traditional convolutional neural network input picture, the image data is divided into patches here, such as 9 patches. The size of each patch can be specified, such as 16x16, etc. Then each patch is input into an embedding layer. After passing through the layer, a series of vectors (tokens) can be obtained. The 9 patches will all obtain their corresponding vectors. Then a vector for classification is added before all the vectors. The dimension of the classification vector is consistent with the other 9 vectors. In addition, position information also needs to be added. Then all the vectors are input into a Transformer encoder. The Transformer encoder is repeatedly stacked L times. Then the output of the token for classification is input into a multilayer perceptron (MLP) head. Finally, the result of the final classification is obtained.

[0083] By applying the scheme of the embodiments of the present specification, the image data is input into the pre-trained picture classification model, and the visual attribute information of the target object is obtained through the classification and recognition of the picture classification model, thereby improving the efficiency and accuracy of obtaining the visual attribute information of the target object, and further making the subsequently generated target description text more accurate.

[0084] It is worth noting that after obtaining the visual attribute information of the target object, the visual attribute information of the target object can be compared with the text data, and the text data of the target object is modified according to the comparison result.

[0085] For example, the text data of the target object is "red clothes make women look younger", and the visual attribute information of the target object is "pink makes white look younger". Comparing the text data and the visual attribute information, "red" in the text data of the target object is replaced with "pink", and the modified text data is "pink clothes make women look younger".

[0086] Step 306: determining the object attribute set of the target object according to the text data and the visual attribute information.

[0087] In one or more embodiments of the present specification, after obtaining the image-text data of the target object and identifying the visual attribute information of the target object based on the image data, the object attribute set of the target object can be further determined according to the text data and the visual attribute information. By integrating the text data and the visual attribute information, the object attribute of the target object is enriched, and the generated target description text is more coherent and accurate.

[0088] Specifically, the object attribute set refers to a set composed of object attribute information of a plurality of target objects, and the object attribute information includes text data and visual attribute information of the target object. The object attribute information can be understood as text information that completely describes the object attribute of the target object.

[0089] In actual applications, the text data and the visual attribute information can be merged and spliced to determine the object attribute set of the target object. For example, the text data of the target object is "orange cat sofa pillow", and the visual attribute information is "orange high-grade". By splicing the text data and the visual attribute information of the target object, it can be determined that the content included in the object attribute set of the target object is "orange cat sofa pillow orange high-grade".

[0090] Further, in order to reduce the data processing amount and improve the text generation efficiency, when splicing the text data and the visual attribute information, the union of the text data and the visual attribute information can also be taken. Referring to the above example, the determined object attribute set is "orange cat sofa pillow high-grade".

[0091] In an optional implementation of the present specification, taking the target object as a target commodity for example, the step of determining the object attribute set of the target object according to the text data and the visual attribute information can include the following steps:

[0092] Determining a commodity attribute set of the target commodity according to the text data and the visual attribute information, wherein the text data includes at least one of a title, a brief introduction, and product parameters of the target commodity.

[0093] Specifically, the title of the commodity usually includes the brand name of the commodity, and the brief introduction of the commodity usually includes the origin and function of the commodity, and the product parameters of the commodity usually include the size, material, and article number of the commodity. The specific selection is based on the actual situation, and the embodiments of the present specification do not make any limitation thereto.

[0094] For example, taking the target commodity as a throw pillow, the title of the target commodity is "big bear throw pillow, large size back pillow, birthday gift", the brief introduction of the target commodity is "cute and innocent panda-shaped throw pillow, soft to the touch, and a good companion for brushing the phone and reading", and the product parameters of the target commodity are "article number: 00001, material: other, size: 70cm*90cm".

[0095] According to the scheme of the embodiment of the present specification, the product attribute set of the target product is determined according to the text data and the visual attribute information, wherein the text data includes at least one of the title, the introduction, and the product parameter of the target product, the object attribute of the target product is enriched, and the generated product description text is more coherent and accurate.

[0096] Step 308: generating the target description text of the target object based on the object attribute set.

[0097] In one or more embodiments of the present specification, after obtaining the image-text data of the target object, identifying the visual attribute information of the target object based on the image data, and determining the object attribute set of the target object according to the text data and the visual attribute information, the target description text of the target object can be further generated based on the object attribute set.

[0098] Specifically, the target description text refers to a text that can briefly and accurately describe the target object. In the embodiment of the present specification, the description text can also be understood as an abstract text, a script, a summary, a content summary, and an abstract script.

[0099] It should be noted that, taking the target object as the target product, the target description text of the target product is the product description text, and the step of generating the target description text of the target object based on the object attribute set can include the following steps:

[0100] Generating the target description text of the target product based on the product attribute set.

[0101] According to the scheme of the embodiment of the present specification, the image-text data of the target object is obtained, wherein the image-text data includes image data and text data; the visual attribute information of the target object is identified based on the image data, wherein the visual attribute information represents the explicit features of the target object; the object attribute set of the target object is determined according to the text data and the visual attribute information; and the target description text of the target object is generated based on the object attribute set. By obtaining the multi-modal image-text data of the target object, determining the visual attribute information of the target object, considering the explicit features of the target object, the object attribute of the target object is more comprehensive, and according to the text data and the visual attribute information, the object attribute set of the target object is determined, the text data and the visual attribute information of the target object are comprehensively considered, the generated target description text is more coherent, and the accuracy of the target description text is further improved.

[0102] In actual application, there are various ways to generate the target description text of the target object based on the object attribute set, which is selected according to actual conditions, and the present specification does not make any limitation in this regard.

[0103] In an optional implementation of the present specification, the text content in the object attribute set can be subjected to word segmentation processing, and each word obtained through word segmentation can be processed by using a pre-set description text generation template to generate a target description text of the target object. The word segmentation processing can be performed by using a word segmentation tool, or the word segmentation result can be obtained by using a pre-set word list for matching, and the specific selection can be made according to actual conditions, and the present specification does not make any limitation in this regard.

[0104] For example, the text content in the object attribute set is “orange cat sofa pillow high-grade feeling”, and the text content is subjected to word segmentation to obtain a word segmentation result of “orange, cat, sofa pillow, high-grade feeling”. The pre-set description text generation template is “XX is XX shape, and gives a person XX feeling”, and the word segmentation result is filled into the description text generation target to obtain a target description text of “sofa pillow is orange cat shape, and gives a person high-grade feeling”.

[0105] In another optional implementation of the present specification, a pre-trained text processing model can be used to generate the target description text, that is, the step of generating the target description text of the target object based on the object attribute set can include the following steps:

[0106] The object attribute set is input into the pre-trained text processing model to generate the target description text of the target object through the text processing model.

[0107] Specifically, the pre-trained text processing model is a model generated by training a pre-set processing model. The pre-set processing model refers to a model capable of realizing text processing, such as a Transformer model (BART, Bidirectional and Auto-Regressive Transformers) having context information and self-recurrent characteristics, a text-to-text transfer conversion model (T5, Text-to-Text Transfer Transformer), a pre-training model (GPT, Generative Pre-Training), and the like, and the specific selection can be made according to actual conditions, and the present specification does not make any limitation in this regard.

[0108] For example, the BART model is an encoder-decoder (Encoder-Decoder) structure, the input of the Encoder end is a sequence added with noise, the input of the Decoder end is a sequence added with a start symbol (right-shifted), and the target of the Decoder end is the original sequence.

[0109] By applying the scheme of the embodiments of the present specification, the object attribute set is input into the pre-trained text processing model, and the target description text of the target object is generated through the text processing model, thereby improving the efficiency of obtaining the target description text and the accuracy of the generated target description text.

[0110] It is worth noting that after the target description text of the target object is generated, the target description text can be directly displayed on the client. The target description text can also be stored in a preset database, and the target description text is called from the preset database when the current client is associated with the target object, that is, after the step of generating the target description text of the target object based on the object attribute set, the following steps can be further included:

[0111] In the case where the object currently displayed on the client is the target object, the target description text is called from the preset database, wherein the preset database is used to store the generated target description text.

[0112] The target description text is displayed on the client, or the target description text is audio converted to generate and play audio data corresponding to the target description text.

[0113] Specifically, if the object currently displayed on the client is the target object, it indicates that the target description text of the target object needs to be obtained. At this time, the target description text can be searched in the preset database to determine whether the pre-generated target description text exists in the preset database. If it exists, the target description text is directly called from the preset database, and the target description text is displayed on the client. If there is no target description text in the preset database, the target description text can be generated in real time by using the text generation method provided in the embodiments of the present specification, and the generated target description text is displayed on the client.

[0114] Further, since the target description text is displayed on the client, the user can introduce the target object according to the target description text. In order to reduce the workload of the user, the text-audio conversion tool can also be used to audio convert the target description text to generate audio data corresponding to the target description text, and after the audio data is generated, the audio data is actively played to realize the introduction of the target object.

[0115] By applying the scheme of the embodiments of the present specification, in the case where the object currently displayed on the client is the target object, the target description text is called from the preset database, thereby saving the time of the user to obtain the target description text and improving the user experience; the target description text is displayed on the client, without the user needing to carefully understand the target object, and the user can directly introduce the target object according to the target description text; the audio data corresponding to the target description text is generated and played, without the user needing to introduce, thereby saving a large amount of human cost.

[0116] The following describes Figure 1The training manner of the text processing model in the illustrated embodiment is described in detail.

[0117] In one or more embodiments of the present specification, the training manner of the text processing model can include the following steps:

[0118] Obtaining a first sample set, wherein the first sample set includes a plurality of sample objects, and each sample object carries sample text data and sample description text;

[0119] Identifying each sample description text to determine the sample visual attribute information of each sample object;

[0120] Data augmentation is performed on each sample text data to determine the augmented text data of each sample object;

[0121] Based on the sample visual attribute information, sample text data and augmented text data of the plurality of sample objects, a preset processing model is trained to obtain a text processing model.

[0122] Specifically, the sample object is used to train the text processing model, and the sample object includes but is not limited to goods, people, scenery, historical sites, etc. The sample text data carried by the sample object is the text data describing the sample object, such as the name, unique attribute, detail information, process information, etc. of the sample object. The sample description text is the corresponding description text of the sample object, which can also be understood as a sample abstract text, a sample script, a sample summary, a sample content abstract, and a sample abstract script. Generally, the first sample set can be obtained by manually inputting a large amount of sample text data and sample description text to form the first sample set; or reading a large amount of sample text data and sample description text from other data acquisition devices or databases to form the first sample set. The specific selection is based on the actual situation, and the present specification does not make any limitation.

[0123] In practical applications, the manner of identifying each sample description text to determine the sample visual attribute information of each sample object can be to perform word segmentation processing on each sample description text, match each word segmentation result with a pre-set visual attribute word table, and obtain the sample visual attribute information of each sample object; or directly performing part-of-speech tagging on the sample description text, retaining the obtained nouns and adjectives to determine the sample visual attribute information.

[0124] In the present embodiment, considering that the same semantic can correspond to multiple words, such as the words expressing good-looking, including beautiful, beautiful, color value, etc., the sample text data of the sample object can be augmented to expand the sample text data of the sample object, so that the sample text data is more diversified, and a certain noise is added to the sample data, further improving the generalization ability of the trained model.

[0125] Exemplarily, the sample text data of the sample object is "This dress is really good-looking", the "good-looking" in the sample text data is replaced with a synonym of good-looking, the data augmentation of the sample text data is realized, and the augmented text data is obtained as "This dress is really beautiful", "This dress is really beautiful", "This dress is really beautiful", etc. The augmented text data can be one or multiple, which is selected according to actual conditions, and the embodiments of the present specification do not make any limitation thereto.

[0126] The scheme of the embodiments of the present specification is applied to obtain a first sample set, wherein the first sample set includes a plurality of sample objects, each sample object carries sample text data and sample description text, each sample description text is identified, sample visual attribute information of each sample object is determined, each sample text data is data augmented, augmented text data of each sample object is determined, a preset processing model is trained based on the sample visual attribute information, sample text data and augmented text data of the plurality of sample objects, a text processing model is obtained, the explicit features of the sample objects are considered, the object attributes of the sample objects are more comprehensive, the sample text data of the sample objects is expanded, the sample text data is more diversified, the trained model has stronger generalization ability, and the accuracy of the trained model is improved.

[0127] Exemplarily, taking a sample object as a sample commodity as an example, the sample text data and the sample description text can be obtained from the live broadcast room and the commodity detail page of the sample commodity, and the first sample set is further constructed, that is, the above-mentioned step of obtaining the first sample set can include the following steps:

[0128] Live data of each sample commodity is extracted from the live broadcast room of the plurality of sample commodities, wherein the live data includes video data and voice data;

[0129] The live data is identified and converted to generate the sample description text of each sample commodity;

[0130] The sample text data of each sample commodity is extracted from the detail page of the plurality of sample commodities;

[0131] The first sample set is constructed according to the sample text data and the sample description text of the plurality of sample commodities.

[0132] Specifically, since the commodity detail page contains a large amount of commodity detail information, and the context semantic coherence exists between the entire detail pages, the graphic and text data of the commodity can be completely covered. Therefore, sample text data of each sample commodity can be extracted from the detail page of the sample commodity, and the manner of extracting the sample text data includes but is not limited to OCR technology. Moreover, live data of the sample commodity can also be collected from the live room of the sample commodity, and the live data includes video data and voice data. The live data is converted and recognized by using ASR technology to generate sample description text of each sample commodity. After obtaining the sample text data and the sample description text, a first sample set can be constructed, wherein the sample description text can be understood as a sample label carried by the sample object, and the sample label represents a result that a preset processing model actually wants to output.

[0133] By applying the scheme of the embodiments of the present specification, live data of each sample commodity is extracted from the live room of the sample commodity, wherein the live data includes video data and voice data. The live data is converted and recognized to generate sample description text of each sample commodity. Sample text data of each sample commodity is extracted from the detail page of the sample commodity. The first sample set is constructed according to the sample text data and the sample description text of the plurality of sample commodities, which enriches the first sample set, makes the context semantic coherence of the sample text data in the sample set, and further improves the accuracy of the trained model.

[0134] Further, after obtaining the sample visual attribute information, the sample text data and the augmented text data of the plurality of sample objects, the sample text data and the augmented text data can be processed based on the sample visual attribute information to determine the initial training sample and the augmented training sample of each sample object. That is, the step of training the preset processing model based on the sample visual attribute information, the sample text data and the augmented text data of the plurality of sample objects to obtain the text processing model can include the following steps:

[0135] merge the sample text data and the sample visual attribute information of each sample object to determine the initial training sample of each sample object;

[0136] merge the augmented text data and the sample visual attribute information of each sample object to determine the augmented training sample of each sample object;

[0137] train the preset processing model based on the initial training sample, the augmented training sample and the sample description text of the plurality of sample objects to obtain the text processing model.

[0138] Specifically, the sample text data and the sample visual attribute information of each sample object are combined to determine initial training samples of the sample objects, and the augmented text data and the sample visual attribute information of each sample object are combined to determine augmented training samples of the sample objects. The manner of combining the augmented text data and the sample visual attribute information of each sample object to determine the augmented training samples of the sample objects can be text splicing, and the text data after deduplication can also be spliced.

[0139] By applying the scheme of the embodiments of the present specification, the sample text data and the sample visual attribute information of each sample object are combined to determine initial training samples of the sample objects, the augmented text data and the sample visual attribute information of each sample object are combined to determine augmented training samples of the sample objects, the initial training samples and the augmented training samples of the plurality of sample objects and the sample description text are used to train a preset processing model to obtain a text processing model. By comprehensively combining the text data and the sample visual attribute information, the object attributes of the sample objects are enriched, and the generalization of the trained model is improved.

[0140] Further, after obtaining the initial training samples and the augmented training samples of the sample objects, the preset processing model can be trained based on the initial training samples and the augmented training samples, that is, the step of using the initial training samples and the augmented training samples of the plurality of sample objects and the sample description text to train the preset processing model to obtain the text processing model can include the following steps:

[0141] extracting first initial training samples and first augmented training samples of a first sample object, wherein the first sample object is any sample object in a first sample set;

[0142] inputting the first initial training samples into the preset processing model to generate a first predicted description text, and inputting the first augmented training samples into the preset processing model to generate a second predicted description text;

[0143] calculating a first loss value according to the first predicted description text and the first sample description text;

[0144] calculating a second loss value according to the second predicted description text and the first sample description text;

[0145] calculating a third loss value according to the first predicted description text and the second predicted description text;

[0146] adjusting model parameters of the preset processing model based on the first loss value, the second loss value, and the third loss value, and returning to the step of extracting the first initial training samples and the first augmented training samples of the first sample object;

[0147] in a case where a first training stop condition is reached, obtaining a text processing model that is trained.

[0148] Specifically, the first sample description text refers to a result that the preset processing model actually wants to output, that is, the first sample description text is a real result. The first predicted description text generated by inputting the first initial training sample into the preset processing model and the second predicted description text generated by inputting the first augmented training sample into the preset processing model are prediction results generated by the preset processing model. When the difference between the prediction result and the real result is small enough, that is, the first loss value and the second loss value are small enough, it indicates that the prediction result is close enough to the real result.

[0149] In particular, since the first augmented training sample is the first initial training sample with added noise, in order to make the prediction results of the preset processing model on the first initial training sample and the first augmented training sample close, and improve the anti-noise capability of the preset processing model, the third loss value can be calculated according to the first predicted description text and the second predicted description text. Finally, after obtaining the first loss value, the second loss value and the third loss value, the model parameters of the preset processing model can be adjusted based on the first loss value, the second loss value and the third loss value, and the step of extracting the first initial training sample and the first augmented training sample of the first sample object is returned. If the first training stop condition is reached, a trained text processing model is obtained.

[0150] It should be noted that the first loss value and the second loss value can be calculated by using a cross-entropy loss function, and the third loss value can be calculated by using a relative entropy loss function (KLD, Kullback-Leibler Divergence). The first training stop condition includes but is not limited to the first preset threshold and the first preset number of iterations, and is selected according to actual conditions. The embodiments of the present specification do not make any limitation on this.

[0151] By using the scheme of the embodiments of the present specification, the efficiency and accuracy of calculating the first loss value and the second loss value are improved by using the cross-entropy loss function, and the efficiency and accuracy of calculating the third loss value are improved by using the relative entropy loss function, which further makes the trained text processing model more accurate.

[0152] In an optional implementation of the present specification, in order to learn better text features, the mutual information maximization loss function can also be used to constrain the encoder in the preset processing model using the initial training sample of each sample object and the sample description text, that is, the preset processing model includes an encoder. Before the above steps of inputting the first initial training sample into the preset processing model to generate the first predicted description text, and inputting the first augmented training sample into the preset processing model to generate the second predicted description text, the following steps can also be included:

[0153] The first initial training sample is input into the encoder to generate a first feature vector;

[0154] inputting the first sample description text into the encoder to generate a second feature vector;

[0155] calculating an encoding loss value according to the first feature vector and the second feature vector;

[0156] adjusting parameters of the encoder based on the encoding loss value, and returning to perform the step of inputting the first initial training sample into the encoder to generate the first feature vector;

[0157] determining the trained encoder in a case where a second training stop condition is reached.

[0158] Specifically, the encoding loss value can be calculated by using the following formula (1):

[0159]

[0160] wherein B is a size of a batch in a training process (loss of B data needs to be calculated each time the parameters are updated), Z i = avg(Z i ), avg represents an average pooling operation, Z i represents a feature vector obtained after the i th initial training sample is input into the encoder, and z y = avg(Z y ), z y represents a feature vector obtained after the i th sample description text is input into the encoder.

[0161] It should be noted that the second training stop condition includes but is not limited to a second preset threshold and a second preset number of iterations, and is specifically selected according to actual conditions, and the embodiments of the present specification do not make any limitation thereto.

[0162] According to the scheme of the embodiments of the present specification, the first initial training sample is input into the encoder to generate the first feature vector, the first sample description text is input into the encoder to generate the second feature vector, the encoding loss value is calculated according to the first feature vector and the second feature vector, the parameters of the encoder are adjusted based on the encoding loss value, and the step of inputting the first initial training sample into the encoder to generate the first feature vector is returned to be performed, and the trained encoder is determined in a case where the second training stop condition is reached. The mutual information maximization loss function is used to constrain the encoder in the preset processing model, so that the preset processing model can learn better text features, and the trained text processing model is more accurate.

[0163] Referring to Figure 4 , Figure 4A training flowchart of a text generation method provided by one embodiment of the present specification is shown, which specifically includes:

[0164] A plurality of sample objects are obtained, each carrying sample text data and sample description text; each sample description text is identified to determine the sample visual attribute information of each sample object; data augmentation is performed on each sample text data to determine the augmented text data of each sample object; the sample text data and the sample visual attribute information of each sample object are merged, and the merged result is input into the encoder and the decoder of the preset processing model to generate a first predicted description text; the augmented text data and the sample visual attribute information of each sample object are merged, and the merged result is input into the encoder and the decoder of the preset processing model to generate a second predicted description text; a first loss value is calculated according to the first predicted description text and the sample description text; a second loss value is calculated according to the second predicted description text and the sample description text; a third loss value is calculated according to the first predicted description text and the second predicted description text; the model parameters of the preset processing model are adjusted based on the first loss value, the second loss value, and the third loss value, and in the case where a first training stop condition is reached, a trained text processing model is obtained.

[0165] The preset processing model includes an encoder and a decoder, the sample text data and the sample visual attribute information of each sample object after merging are input into the encoder to generate a first feature vector; the sample description text of each sample object is input into the encoder to generate a second feature vector; an encoding loss value is calculated according to the first feature vector and the second feature vector; the parameters of the encoder are adjusted based on the encoding loss value, and in the case where a second training stop condition is reached, a trained encoder is determined.

[0166] The training method of the picture classification model in the embodiment shown in the following detailed description. Figure 1 The training method of the picture classification model in the embodiment shown in the following detailed description.

[0167] In one or more embodiments of the present specification, the training method of the picture classification model can include the following steps:

[0168] A second sample set is obtained, wherein the second sample set includes a plurality of sample objects, each carrying sample image data and sample description text;

[0169] Each sample description text is identified to determine the sample visual attribute information of each sample object;

[0170] The sample image data and the sample visual attribute information of the plurality of sample objects are used to train a preset classification model to obtain a picture classification model.

[0171] Specifically, the specific manner of obtaining the second sample set, identifying each sample description text, and determining the sample visual attribute information of each sample object can refer to the above-mentioned manner of training the text processing model, and the embodiments of the present specification will not be described again. Determining the sample visual attribute information of each sample object takes into account the explicit features of the sample object, making the object attributes of the sample object more comprehensive and improving the accuracy of the trained model.

[0172] Further, the step of training the preset classification model by using the sample image data and the sample visual attribute information of the plurality of sample objects to obtain the picture classification model can include the following steps:

[0173] extracting second sample image data and second sample visual attribute information of a second sample object, wherein the second sample object is any sample object in the second sample set;

[0174] inputting the second sample image data into the preset classification model to obtain predicted visual attribute information of the second sample object;

[0175] calculating a classification loss value of the preset classification model according to the second sample visual attribute information and the predicted visual attribute information of the second sample object;

[0176] adjusting model parameters of the preset classification model according to the classification loss value, and returning to the step of extracting the second sample image data and the second sample visual attribute information of the second sample object;

[0177] in the case where the third training stop condition is reached, obtaining the picture classification model that is completed training.

[0178] It should be noted that the classification loss value can be calculated based on the predicted visual attribute information of the second sample object and the second sample visual attribute information, the second sample visual attribute information represents the result that the preset classification model actually wants to output, and the predicted visual attribute information output by inputting the second sample image data into the preset classification model is the prediction result of the preset classification model. When the difference between the prediction result and the real result is small enough, that is, the classification loss value is small enough, it means that the prediction result is close enough to the real result, at this time the preset classification model is trained, and the picture classification model that is completed training is obtained.

[0179] In the embodiments of the present specification, the difference between the prediction result and the real result of the preset classification model can be intuitively shown by calculating the classification loss value, and the preset classification model can be trained based on the difference in the subsequent, and the parameters of the preset classification model can be adjusted, which can effectively improve the training rate and the training effect of the preset classification model.

[0180] It should be noted that the third training stopping condition includes but is not limited to the third preset threshold and the third preset number of iterations, and is selected according to actual conditions, and the embodiments of the present specification do not make any limitation in this regard.

[0181] In a possible implementation, whether to stop the training can be determined only based on a relationship between the classification loss value and the third preset threshold. Specifically, if the classification loss value is greater than the third preset threshold, it indicates that the difference between the second sample visual attribute information and the predicted visual attribute information of the second sample object is large, and the classification recognition capability of the preset classification model is poor. At this time, the model parameters of the preset classification model can be adjusted, and the step of extracting the second sample image data and the second sample visual attribute information of the second sample object is returned to continue training the preset classification model until the classification loss value is less than or equal to the third preset threshold, indicating that the difference between the second sample visual attribute information and the predicted visual attribute information of the second sample object is small, and the training is stopped to obtain the trained picture classification model.

[0182] The third preset threshold is a critical value of the classification loss value. In the case where the classification loss value is greater than the third preset threshold, it indicates that there is still a certain deviation between the prediction result of the preset classification model and the true result, and the model parameters of the preset classification model still need to be adjusted and the preset classification model still needs to be trained. In the case where the classification loss value is less than or equal to the third preset threshold, it indicates that the closeness between the prediction result of the preset classification model and the true result is sufficient, and the training can be stopped.

[0183] In another possible implementation, in addition to comparing the relationship between the classification loss value and the third preset threshold, the number of iterations can also be combined to determine whether the current preset classification model is trained. Specifically, if the classification loss value is less than or equal to the third preset threshold, it indicates that the difference between the second sample visual attribute information and the predicted visual attribute information of the second sample object is small, and the training is stopped to obtain the trained picture classification model. That is, when the classification loss value is less than or equal to the third preset threshold, the training can be stopped to obtain the trained picture classification model without combining the number of iterations. If the classification loss value is greater than the third preset threshold, it is determined whether the number of iterations at this moment reaches the third preset number of iterations. If the number of iterations at this moment does not reach the third number of iterations, the model parameters of the preset classification model are adjusted, and the step of extracting the second sample image data and the second sample visual attribute information of the second sample object is returned to continue training the preset classification model until the third preset number of iterations is reached. In the case where the training is stopped, the trained picture classification model is obtained.

[0184] The values ​​of the third preset threshold and the third preset number of iterations are selected according to the actual situation, and this specification does not impose any limitations on them in the embodiments. When the number of iterations reaches the third preset number of iterations, it indicates that the training times of the preset classification model have been sufficient, and the prediction results of the preset classification model are close enough to the actual results, so training can be stopped.

[0185] In practical applications, there are many functions for calculating classification loss, such as cross-entropy loss function, L1 norm loss function, maximum loss function, mean squared error loss function, log loss function, etc. The specific function to be selected depends on the actual situation, and the embodiments in this specification do not impose any limitations on this.

[0186] The scheme implemented in this specification can determine the specific training status of the preset classification model based on the classification loss value, and adjust the model parameters of the preset classification model in reverse according to the classification loss value if the training is unsuccessful, so as to improve the classification and recognition ability of the model. The training rate is high and the training effect is good.

[0187] See Figure 5 , Figure 5 This specification illustrates a flowchart of the training process for an image classification model in a text generation method according to an embodiment of the present specification, specifically including:

[0188] Multiple sample objects are acquired, each carrying sample image data and sample description text; each sample description text is identified to determine the sample visual attribute information of each sample object; the sample image data of each sample object is input into a preset classification model to obtain predicted visual attribute information; based on the sample visual attribute information and the predicted visual attribute information, the classification loss value of the preset classification model is calculated; based on the classification loss value, the parameters of the preset classification model are tuned, and when the third training stopping condition is met, the image classification model that has completed training is obtained.

[0189] The following is in conjunction with the appendix Figure 6 Taking the text generation method provided in this specification as an example in the application of e-commerce live streaming, the text generation method will be further explained. Figure 6 The flowchart of a text generation method according to an embodiment of this specification is shown, which specifically includes the following steps:

[0190] Step 602: Obtain the details page data of the target product, wherein the details page data includes image data and text data, and the text data includes at least one of the target product's title, description, and product parameters.

[0191] See Figure 7 , Figure 7 This diagram illustrates a target product details page in a text generation method provided in one embodiment of this specification.

[0192] The image data of the target commodity details page includes the image data of the coffee cup, such as the two coffee cups in the figure, and also includes the title of the target commodity: coffee cup large capacity with spoon; the introduction of the target commodity: high glaze, safe and reliable, warm tone, and brings different experience to life; the product parameters of the target commodity: rich in style, 500ml.

[0193] Step 604: input the image data into the pre-trained picture classification model, and obtain the visual attribute information of the target commodity through the classification and identification of the picture classification model, wherein the visual attribute information represents the explicit features of the target commodity.

[0194] Specifically, the image data is input into the pre-trained picture classification model, and the visual attribute information of the target commodity is obtained through the classification and identification of the picture classification model, which is "white, warm tone brown, striped, non-striped, soft color, simple and elegant".

[0195] Step 606: merge the text data and the visual attribute information to determine the commodity attribute set of the target commodity.

[0196] Specifically, the text data and the visual attribute information are merged to determine the commodity attribute set of the target commodity as "coffee cup large capacity with spoon, high glaze, safe and reliable, warm tone, and brings different experience to life, rich in style, 500ml, white, warm tone brown, striped, non-striped, soft color, simple and elegant".

[0197] Step 608: input the commodity attribute set into the pre-trained text processing model, and generate the target description text of the target commodity through the text processing model.

[0198] Specifically, referring to Figure 8 , Figure 8 The display interface of the client in a text generation method provided by one embodiment of the present specification is shown. The target description text included in the client display interface is "this is a large capacity coffee cup with spoon, which has a capacity of 500ml. This coffee cup is rich in style, has white, warm tone brown, striped and non-striped. The color is soft and simple. The coffee cup is made of high glaze, safe and reliable, and brings different life experience to you."

[0199] Step 610: display the target description text on the client side to make the virtual anchor introduce the target commodity according to the target description text.

[0200] According to the scheme of the embodiment of the present specification, the detail page data of the target commodity is obtained, the image data in the detail page data is input into the pre-trained picture classification model, the visual attribute information of the target commodity is obtained through the classification and recognition of the picture classification model, the text data in the detail page data and the visual attribute information are merged, the commodity attribute set of the target commodity is determined, the commodity attribute set is input into the pre-trained text processing model, the target description text of the target commodity is generated through the text processing model, and the target description text is displayed on the client, so that the virtual host introduces the target commodity according to the target description text. The multi-modal data is combined with the algorithm, applied to the virtual host script construction process, used to guide the content construction conforming to the live scene characteristics, and supports the input of multi-source text data and image data, supports long text generation, so as to realize the automatic generation of commodity abstract.

[0201] Corresponding to the method embodiments described above, the present specification also provides text generation device embodiments, Figure 9 The structure of a text generation device provided by one embodiment of the present specification is shown. As shown in the figure, Figure 9 The device includes:

[0202] The acquisition module 902 is configured to acquire image data and text data of a target object.

[0203] The identification module 904 is configured to identify visual attribute information of the target object based on the image data, wherein the visual attribute information represents the explicit features of the target object.

[0204] The determination module 906 is configured to determine an object attribute set of the target object according to the text data and the visual attribute information.

[0205] The generation module 908 is configured to generate a target description text of the target object based on the object attribute set.

[0206] Optionally, the acquisition module 902 is further configured to monitor the image data and text data of the target object, and acquire the image data and text data of the target object in the case of updating the image data and text data.

[0207] Optionally, the device further includes a calling module configured to call the target description text from a preset database in the case that the object currently displayed on the client is the target object, wherein the preset database is used to store the generated target description text; display the target description text on the client; or perform audio conversion on the target description text, generate and play audio data corresponding to the target description text.

[0208] Optionally, the target object includes a target commodity; the determination module 906 is further configured to determine a commodity attribute set of the target commodity according to the text data and the visual attribute information, wherein the text data includes at least one of a title, a brief introduction, and product parameters of the target commodity;

[0209] The generation module 908 is further configured to generate a target description text of the target commodity based on the commodity attribute set.

[0210] Optionally, the generation module 908 is further configured to input the object attribute set into a pre-trained text processing model, and generate the target description text of the target object through the text processing model;

[0211] The apparatus further includes a text processing model training module configured to obtain a first sample set, wherein the first sample set includes a plurality of sample objects, each sample object carrying sample text data and a sample description text; identify each sample description text to determine sample visual attribute information of each sample object; perform data augmentation on each sample text data to determine augmented text data of each sample object; train a preset processing model based on the sample visual attribute information, the sample text data, and the augmented text data of the plurality of sample objects to obtain the text processing model.

[0212] Optionally, the sample object includes a sample commodity; the text processing model training module is further configured to extract live data of each sample commodity from a live broadcast room of the plurality of sample commodities, wherein the live data includes video data and voice data; perform recognition conversion on the live data to generate the sample description text of each sample commodity; extract sample text data of each sample commodity from a detail page of the plurality of sample commodities; and construct the first sample set according to the sample text data and the sample description text of the plurality of sample commodities.

[0213] Optionally, the text processing model training module is further configured to merge the sample text data and the sample visual attribute information of each sample object to determine initial training samples of each sample object; merge the augmented text data and the sample visual attribute information of each sample object to determine augmented training samples of each sample object; and train the preset processing model using the initial training samples, the augmented training samples, and the sample description text of the plurality of sample objects to obtain the text processing model.

[0214] Optionally, the text processing model training module is further configured to extract a first initial training sample and a first augmented training sample of a first sample object, wherein the first sample object is any sample object in the first sample set; input the first initial training sample into the preset processing model to generate a first predicted description text, and input the first augmented training sample into the preset processing model to generate a second predicted description text; calculate a first loss value according to the first predicted description text and the first sample description text; calculate a second loss value according to the second predicted description text and the first sample description text; calculate a third loss value according to the first predicted description text and the second predicted description text; adjust the model parameters of the preset processing model based on the first loss value, the second loss value, and the third loss value, and return to execute the step of extracting the first initial training sample and the first augmented training sample of the first sample object; and obtain the trained text processing model in the case where a first training stop condition is reached.

[0215] Optionally, the preset processing model comprises an encoder; the apparatus further comprises an encoder training module configured to input the first initial training sample into the encoder to generate a first feature vector; input the first sample description text into the encoder to generate a second feature vector; calculate an encoding loss value according to the first feature vector and the second feature vector; adjust the parameters of the encoder based on the encoding loss value, and return to execute the step of inputting the first initial training sample into the encoder to generate the first feature vector; and determine the trained encoder in the case where a second training stop condition is reached.

[0216] Optionally, the recognition module 904 is further configured to input the image data into a pre-trained picture classification model, and obtain the visual attribute information of the target object through classification and recognition of the picture classification model.

[0217] The apparatus further comprises a picture classification model training module configured to obtain a second sample set, wherein the second sample set comprises a plurality of sample objects, and each sample object carries sample image data and sample description text; recognize each sample description text to determine sample visual attribute information of each sample object; and train a preset classification model using the sample image data and the sample visual attribute information of the plurality of sample objects to obtain a picture classification model.

[0218] Optionally, the picture classification model training module is further configured to extract second sample image data and second sample visual attribute information of a second sample object, wherein the second sample object is any sample object in the second sample set; input the second sample image data into the preset classification model to obtain predicted visual attribute information of the second sample object; calculate a classification loss value of the preset classification model according to the second sample visual attribute information and the predicted visual attribute information of the second sample object; adjust the model parameters of the preset classification model according to the classification loss value, and return to the step of extracting the second sample image data and the second sample visual attribute information of the second sample object; and obtain the picture classification model trained in the case where a third training stop condition is reached.

[0219] By applying the scheme of the embodiment of the present specification, the image-text data of the target object is obtained, wherein the image-text data includes image data and text data; the visual attribute information of the target object is identified based on the image data, wherein the visual attribute information represents the explicit features of the target object; the object attribute set of the target object is determined according to the text data and the visual attribute information; and the target description text of the target object is generated based on the object attribute set. By obtaining the multi-modal image-text data of the target object, determining the visual attribute information of the target object, and considering the explicit features of the target object, the object attributes of the target object are more comprehensive. Moreover, by determining the object attribute set of the target object according to the text data and the visual attribute information, the text data and the visual attribute information of the target object are comprehensively considered, so that the generated target description text is more coherent, and the accuracy of the target description text is further improved.

[0220] The above is a schematic scheme of the text generation device of the embodiment. It should be noted that the technical scheme of the text generation device belongs to the same concept as the technical scheme of the text generation method described above, and the details of the technical scheme of the text generation device that are not described in detail can be referred to the description of the technical scheme of the text generation method.

[0221] Figure 10 A structural block diagram of a computing device is shown. The components of the computing device 1000 include, but are not limited to, a memory 1010 and a processor 1020. The processor 1020 is connected to the memory 1010 through a bus 1030, and a database 1050 is used to save data.

[0222] The computing device 1000 also includes an access device 1040 that enables the computing device 1000 to communicate via one or more networks 1060. Examples of such networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or combinations of such networks, such as the Internet. The access device 1040 can include one or more of any type of network interface (for example, a network interface card (NIC)) such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Global System for Mobile (GSM) interface, a Code Division Multiple Access (CDMA) interface, a Bluetooth interface, a Near Field Communication (NFC) interface, a Universal Serial Bus (USB) interface, a Wi-Fi® interface, a Wi-MAX interface, an Ethernet interface, a token ring interface, a serial bus interface, or the like.

[0223] In one embodiment of the present specification, the above-mentioned components of the computing device 1000 and other components not shown in the Figure 10 may be connected to each other, for example, through a bus. It should be understood that Figure 10 The computing device structure diagram shown is merely for the purpose of example, and is not a limitation on the scope of the present specification. Those skilled in the art can add or replace other components as needed.

[0224] The computing device 1000 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (for example, a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, and the like), a mobile phone (for example, a smartphone), a wearable computing device (for example, a smart watch, smart glasses, and the like), or other types of mobile devices, or a stationary computing device such as a desktop computer or a PC. The computing device 1000 can also be a mobile or stationary server.

[0225] The processor 1020 is configured to execute computer-executable instructions, which, when executed by the processor, implement the steps of the text generation method described above.

[0226] The above is a schematic scheme of the computing device of the embodiment. It should be noted that the technical scheme of the computing device and the technical scheme of the text generation method described above belong to the same concept, and the details of the technical scheme of the computing device that are not described in detail can be referred to the description of the technical scheme of the text generation method.

[0227] An embodiment of the present specification also provides a computer readable storage medium storing computer executable instructions, and the computer executable instructions implement the steps of the text generation method when executed by a processor.

[0228] The above is a schematic scheme of the computer readable storage medium of the embodiment. It should be noted that the technical scheme of the storage medium and the technical scheme of the text generation method described above belong to the same concept, and the details of the technical scheme of the storage medium that are not described in detail can be referred to the description of the technical scheme of the text generation method.

[0229] An embodiment of the present specification also provides a computer program, and the computer program causes a computer to execute the steps of the text generation method when the computer program is executed in the computer.

[0230] The above is a schematic scheme of the computer program of the embodiment. It should be noted that the technical scheme of the computer program and the technical scheme of the text generation method described above belong to the same concept, and the details of the technical scheme of the computer program that are not described in detail can be referred to the description of the technical scheme of the text generation method.

[0231] The specific embodiments of the present specification are described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order and still accomplish desirable results. Additionally, the processes depicted in the figures do not necessarily require the particular order shown, or sequential order to achieve desirable results. In certain implementations, multitasking and parallel processing can be advantageous.

[0232] The computer instructions include computer program code, which can be in the form of source code, object code, executable code, or some intermediate form. The computer readable medium can include any entity or apparatus capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.

[0233] It should be noted that, for the aforementioned method embodiments, the sequences of the described actions are not necessarily required to implement the present application, and certain actions can be performed in other sequences, or even at the same time, in accordance with the present application. Furthermore, certain actions can not be required to implement the present application. Additionally, the described embodiments are not necessarily the only possible implementation of the present application.

[0234] In the above embodiments, the description of each embodiment focuses on different aspects, and the parts not described in detail in a certain embodiment can be referred to the relevant description of other embodiments.

[0235] The above disclosed preferred embodiments of the present application are only used to help explain the present application. Alternative embodiments do not describe all the details and limit the present application to the specific embodiments described. Obviously, according to the content of the embodiments of the present application, many modifications and changes can be made. The present application selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of the present application, so that those skilled in the art can well understand and use the present application. The present application is limited by the claims and their full scope and equivalents.

Claims

1. A text generation method, comprising: Acquire the image and text data of the target object, wherein the image and text data includes image data and text data; Based on the image data, visual attribute information of the target object is identified, wherein the visual attribute information represents the explicit features of the target object; Based on the text data and the visual attribute information, determine the object attribute set of the target object; Based on the object attribute set, target description text of the target object is generated. The target description text is generated by inputting the object attribute set into a text processing model. The text processing model is trained using initial training samples, augmented training samples, and sample description text of the sample object. The initial training samples are obtained by merging sample text data and sample visual attribute information of the sample object. The augmented training samples are obtained by merging augmented text data and sample visual attribute information of the sample object. The augmented text data is obtained by data augmentation of the sample text data. The sample visual attribute information is obtained by recognizing the sample description text. The sample text data and sample description text are carried by the sample object.

2. The method according to claim 1, wherein the step of acquiring the image and text data of the target object includes: Monitor the graphic and textual data of the target object; If the image and text data is updated, the image and text data of the target object are obtained.

3. The method according to claim 1 or 2, after the step of generating the target description text of the target object based on the object attribute set, further comprising: When the object currently displayed on the client is the target object, the target description text is retrieved from a preset database, wherein the preset database is used to store the generated target description text; The target description text is displayed on the client; or, the target description text is converted into audio to generate and play the audio data corresponding to the target description text.

4. The method according to claim 1, wherein the target object includes a target product; the step of determining the object attribute set of the target object based on the text data and the visual attribute information includes: Based on the text data and the visual attribute information, the product attribute set of the target product is determined, wherein the text data includes at least one of the target product's title, description, and product parameters; The step of generating the target description text of the target object based on the object attribute set includes: Based on the product attribute set, generate the target description text for the target product.

5. The method according to claim 1, wherein the step of generating the target description text of the target object based on the object attribute set comprises: The object attribute set is input into the pre-trained text processing model, and the text processing model generates the target description text of the target object. The training method for the text processing model includes: Obtain a first sample set, wherein the first sample set includes multiple sample objects, each sample object carrying sample text data and sample description text; Identify the description text of each sample and determine the visual attribute information of each sample object; Data augmentation is performed on each sample text data to determine the augmented text data for each sample object; Based on the visual attribute information, text data, and augmented text data of the multiple sample objects, a preset processing model is trained to obtain the text processing model.

6. The method according to claim 5, wherein the sample object includes sample goods; the step of obtaining the first sample set includes: Live streaming data for each sample product is extracted from the live streaming rooms of multiple sample products, wherein the live streaming data includes video data and audio data; The live stream data is identified and transformed to generate sample description text for each sample product; Extract sample text data for each sample product from its details page; The first sample set is constructed based on the sample text data and sample description text of the multiple sample products.

7. The method according to claim 5, wherein the step of training a preset processing model based on the sample visual attribute information, sample text data, and augmented text data of the plurality of sample objects to obtain the text processing model includes: Merge the sample text data and sample visual attribute information of each sample object to determine the initial training samples for each sample object; Merge the augmented text data and visual attribute information of each sample object to determine the augmented training samples for each sample object; Using the initial training samples, augmented training samples, and sample description text of the multiple sample objects, a preset processing model is trained to obtain the text processing model.

8. The method according to claim 7, wherein the step of training a preset processing model using the initial training samples, augmented training samples, and sample description text of the plurality of sample objects to obtain the text processing model includes: Extract the first initial training sample and the first augmented training sample of the first sample object, wherein the first sample object is any sample object in the first sample set; The first initial training sample is input into the preset processing model to generate the first predicted description text, and the first augmented training sample is input into the preset processing model to generate the second predicted description text. Calculate the first loss value based on the first predicted description text and the first sample description text; Calculate the second loss value based on the second predicted description text and the first sample description text; Calculate a third loss value based on the first predicted description text and the second predicted description text; Based on the first loss value, the second loss value, and the third loss value, the model parameters of the preset processing model are adjusted, and the process returns to the step of extracting the first initial training sample and the first augmented training sample of the first sample object. Once the first training stopping condition is met, the text processing model that has completed training is obtained.

9. The method according to claim 8, wherein the preset processing model includes an encoder; prior to the steps of inputting the first initial training sample into the preset processing model to generate a first predicted descriptive text, and inputting the first augmented training sample into the preset processing model to generate a second predicted descriptive text, the method further includes: The first initial training sample is input into the encoder to generate a first feature vector; The first sample description text is input into the encoder to generate a second feature vector; Calculate the encoding loss value based on the first feature vector and the second feature vector; Based on the encoding loss value, the parameters of the encoder are adjusted, and the process returns to the step of inputting the first initial training sample into the encoder to generate the first feature vector; If the second training stop condition is met, the encoder that has completed training is determined.

10. The method according to claim 1, wherein the step of identifying the visual attribute information of the target object based on the image data comprises: The image data is input into a pre-trained image classification model, and the visual attribute information of the target object is obtained through classification and recognition by the image classification model. The training method for the image classification model includes: Obtain a second sample set, wherein the second sample set includes multiple sample objects, each sample object carrying sample image data and sample description text; Identify the description text of each sample and determine the visual attribute information of each sample object; Using the sample image data and visual attribute information of the multiple sample objects, a preset classification model is trained to obtain the image classification model.

11. The method according to claim 10, wherein the step of training a preset classification model using the sample image data and sample visual attribute information of the plurality of sample objects to obtain the image classification model includes: Extract the second sample image data and second sample visual attribute information of the second sample object, wherein the second sample object is any sample object in the second sample set; The second sample image data is input into a preset classification model to obtain the predicted visual attribute information of the second sample object; The classification loss value of the preset classification model is calculated based on the visual attribute information of the second sample and the predicted visual attribute information of the second sample object. Based on the classification loss value, adjust the model parameters of the preset classification model, and return to the step of extracting the second sample image data and the second sample visual attribute information of the second sample object; Once the third training stopping condition is met, the image classification model that has completed training is obtained.

12. A text generation apparatus, comprising: The acquisition module is configured to acquire the image and text data of the target object, wherein the image and text data includes image data and text data; The recognition module is configured to recognize visual attribute information of the target object based on the image data, wherein the visual attribute information represents the explicit features of the target object; The determination module is configured to determine the object attribute set of the target object based on the text data and the visual attribute information; The generation module is configured to generate target description text for the target object based on the object attribute set. The target description text is generated by inputting the object attribute set into a text processing model. The text processing model is trained using initial training samples, augmented training samples, and sample description text of the sample object. The initial training samples are obtained by merging sample text data and sample visual attribute information of the sample object. The augmented training samples are obtained by merging augmented text data and sample visual attribute information of the sample object. The augmented text data is obtained by data augmentation of the sample text data. The sample visual attribute information is obtained by recognizing the sample description text. The sample text data and sample description text are carried by the sample object.

13. A computing device, comprising: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, which, when executed by the processor, implement the steps of the text generation method according to any one of claims 1 to 11.

14. A computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the steps of the text generation method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Text generation method and device and storage medium

    CN112270163A

  • Multi-modal pre-training model training method and device, equipment and storage medium

    CN114005012A

  • Image data processing method and device, storage medium and processor

    CN114168777A