Data processing method, device and storage medium

By interactively adjusting the feature vectors of object images and text, the problem of low accuracy in object subject selection and central word extraction in the prior art is solved, and more efficient object understanding ability is achieved.

CN114896979BActive Publication Date: 2025-05-13ALIBABA (CHINA) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210579780.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-26
Publication Date
2025-05-13
Estimated Expiration
2042-05-26

AI Technical Summary

Technical Problem

The prior art is low in accuracy when selecting object subjects and extracting central words in object images and text titles, resulting in insufficient understanding of objects.

Method used

By obtaining the initial feature vectors of the target image and text, and performing interactive adjustments, the adjusted image feature vectors and text feature vectors are obtained, and coordinate regression and named entity recognition are performed to improve the accuracy of extracting object positions and central words.

Benefits of technology

It improves the accuracy of object position regression and central word recognition, making object comprehension capabilities stronger, and is suitable for scenarios such as e-commerce platforms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114896979B_ABST
    Figure CN114896979B_ABST
Patent Text Reader

Abstract

The embodiments of the present application provide a data processing method, device and storage medium. The data processing includes: obtaining an initial image feature vector of a target image and an initial text feature vector of a target text; adjusting the initial image feature vector based on the initial text feature vector to obtain an adjusted image feature vector, and adjusting the initial text feature vector based on the initial image feature vector to obtain an adjusted text feature vector; performing coordinate regression based on the adjusted image semantic vector in the adjusted image feature vector to obtain the position information of the target object in the target image; performing named entity recognition based on each adjusted word vector in the adjusted text feature vector to extract the central word in the target text. The embodiments of the present application can improve the accuracy of target object position detection and central word extraction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of computer technology, and in particular, to a data processing method, device, and storage medium. Background Art

[0002] For specific objects, on the one hand, they can be described by text, and on the other hand, they can also be described by images. However, in addition to the target object, the object image may also include other objects; in addition to the core words describing the object, the file text usually also contains a large number of redundant words. For example: in an e-commerce platform, the display of goods can be presented through product images and product text titles, but usually the target product image also contains other products, and the target product text title also contains other non-core words. If the target product is shorts, its corresponding text title may be "Splicing Workwear Shorts". "Shorts" in the above text title is the core word, while "Splicing" and "Workwear" are non-core words.

[0003] In order to better understand the object, it is usually necessary to select the object body based on the object image, that is, mark the location of the object in the object image; extract the central word from the entire object text title. In the e-commerce scenario, product understanding is the basic capability of the e-commerce platform, which has an important impact on the subsequent product display and product search.

[0004] In the related art, two independent schemes are usually adopted to perform the object subject selection and the core word extraction tasks respectively. In this way, both the accuracy of object subject selection and the accuracy of core word extraction are low. In other words, the above method has poor ability to understand objects. Summary of the invention

[0005] In view of this, embodiments of the present application provide a data processing method, device, and storage medium to at least partially solve the above problems.

[0006] According to a first aspect of an embodiment of the present application, a data processing method is provided, including:

[0007] Acquire an initial image feature vector of a target image and an initial text feature vector of a target text; wherein the initial image feature vector includes: an initial image semantic vector; the initial text feature vector includes: an initial word vector of each word in the target text;

[0008] Adjusting the initial image feature vector based on the initial text feature vector to obtain an adjusted image feature vector, and adjusting the initial text feature vector based on the initial image feature vector to obtain an adjusted text feature vector;

[0009] Coordinate regression is performed based on the adjusted image semantic vector in the adjusted image feature vector to obtain the position information of the target object in the target image; named entity recognition is performed based on each adjusted word vector in the adjusted text feature vector to extract the central word in the target text.

[0010] According to a second aspect of an embodiment of the present application, a data processing method is provided, which is applied to a server device, including:

[0011] Receiving a target image of a commodity and a target text for describing attribute information of the commodity sent by a client device;

[0012] Acquire an initial image feature vector of the target image and an initial text feature vector of the target text; wherein the initial image feature vector includes: an initial image semantic vector; the initial text feature vector includes: an initial word vector of each word in the target text;

[0013] Adjusting the initial image feature vector based on the initial text feature vector to obtain an adjusted image feature vector, and adjusting the initial text feature vector based on the initial image feature vector to obtain an adjusted text feature vector;

[0014] Performing coordinate regression based on the adjusted image semantic vector in the adjusted image feature vector to obtain commodity location information in the target image; performing named entity recognition based on each adjusted word vector in the adjusted text feature vector to extract the central word in the target text;

[0015] The target image, the product location information, the core word and the product information of the product are stored accordingly.

[0016] According to a third aspect of an embodiment of the present application, there is provided a data processing device, including:

[0017] An initial feature vector acquisition module is used to acquire an initial image feature vector of a target image and an initial text feature vector of a target text; wherein the initial image feature vector includes: an initial image semantic vector; the initial text feature vector includes: an initial word vector of each word in the target text;

[0018] A first adjustment module, configured to adjust the initial image feature vector based on the initial text feature vector to obtain an adjusted image feature vector, and to adjust the initial text feature vector based on the initial image feature vector to obtain an adjusted text feature vector;

[0019] The result obtaining module is used to perform coordinate regression according to the adjusted image semantic vector in the adjusted image feature vector to obtain the position information of the target object in the target image; perform named entity recognition according to each adjusted word vector in the adjusted text feature vector to extract the central word in the target text.

[0020] According to the fourth aspect of the embodiments of the present application, there is provided an electronic device, comprising: a processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other through the communication bus; the memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform an operation corresponding to the data processing method described in the first aspect or the second aspect.

[0021] According to a fourth aspect of an embodiment of the present application, a computer storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the data processing method as described in the first aspect or the second aspect is implemented.

[0022] According to the data processing method, device and storage medium provided by the embodiment of the present application, after obtaining the initial image feature vector at the image modality level and the initial text feature vector at the text modality level, the initial image feature vector and the initial text feature vector are interactively adjusted across modalities. Specifically: for the image feature vector at the image modality level, the adjustment process refers to the key information contained in the initial text feature vector at the text modality level, so that the adjusted image feature vector focuses on the information related to the above key information, while ignoring or filtering out the information with a low correlation with the above key information; correspondingly, for the text feature vector at the text modality level, the adjustment process also refers to the key information contained in the initial image feature vector at the image modality level, so that the adjusted text feature vector focuses on the information related to the above key information, while ignoring or filtering out the information with a low correlation with the above key information. Furthermore, the position regression of the target object is performed based on the adjusted image feature vector, and the accuracy of the regression result is higher; the named entity recognition is performed based on the adjusted text feature vector, and the accuracy of the recognition result is also higher. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the embodiments of the present application. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.

[0024] Figure 1 A flowchart of a data processing method according to Embodiment 1 of the present application;

[0025] Figure 2 for Figure 1 A schematic diagram of an example scenario in the illustrated embodiment;

[0026] Figure 3 is a flowchart of a data processing method according to Embodiment 2 of the present application;

[0027] Figure 4 for Figure 3 A schematic diagram of an example scenario in the illustrated embodiment;

[0028] Figure 5 is a flowchart of a data processing method according to Embodiment 3 of the present application;

[0029] Figure 6 is a structural block diagram of a data processing device according to Embodiment 4 of the present application;

[0030] Figure 7 This is a schematic diagram of the structure of an electronic device according to Embodiment 5 of the present application. DETAILED DESCRIPTION

[0031] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments in the embodiments of the present application should fall within the scope of protection of the embodiments of the present application.

[0032] The specific implementation of the embodiment of the present application is further explained below in conjunction with the accompanying drawings of the embodiment of the present application.

[0033] Embodiment 1

[0034] Reference Figure 1 , Figure 1 The following is a flowchart of a data processing method according to Embodiment 1 of the present application. Specifically, the data processing method provided in this embodiment includes the following steps:

[0035] Step 102, obtaining an initial image feature vector of the target image and an initial text feature vector of the target text; wherein the initial image feature vector includes: an initial image semantic vector; and the initial text feature vector includes: an initial word vector of each word in the target text.

[0036] The target image is an image containing a target object. The image semantic vector corresponding to the target image is a vector representing the image information contained in the entire target image. Initially, the image semantic vector can be a preset initial image semantic vector. The target text is a text used to describe the target file. The text can contain multiple words, and each word corresponds to an initial word vector.

[0037] Furthermore, in the embodiment of the present application, the above-mentioned initial image feature vector and initial text feature vector can be obtained by the following method:

[0038] Acquire a target image including a target object, a preset initial image semantic vector, and a target text for describing the target object;

[0039] The target image is divided into blocks to obtain multiple image blocks; each image block is linearly transformed to obtain an initial block vector of each image block; the target text is segmented to obtain multiple word units; each word unit is word embedded to obtain an initial word vector of each word unit.

[0040] Step 104: adjusting the initial image feature vector based on the initial text feature vector to obtain an adjusted image feature vector, and adjusting the initial text feature vector based on the initial image feature vector to obtain an adjusted text feature vector.

[0041] Specifically, for the initial image feature vector, that is, for the initial image semantic vector, the adjustment weight value of the initial image semantic vector can be determined based on the similarity between it and the initial word vector of each word element included in the initial text feature vector, and then based on the obtained adjustment weight value of the initial image semantic vector, the initial image semantic vector is adjusted to obtain the adjusted image semantic vector, that is, the adjusted image feature vector is obtained.

[0042] For the initial text feature vector, for each initial word vector therein, the adjustment weight value of the initial word vector can be determined based on the similarity between the initial word vector and the initial image feature vector, and the initial word vector can be adjusted based on the adjustment weight value of the initial word vector to obtain an adjusted word vector; wherein the adjusted text feature vector includes each adjusted word vector.

[0043] In the above process, when adjusting each type of initial vector (initial image semantic vector or initial word vector), the product of the corresponding adjustment weight value and the initial vector can be used as the adjusted vector (adjusted image semantic vector or adjusted word vector) corresponding to the initial vector.

[0044] Furthermore, the initial image feature vector also includes: initial block vectors of each image block in the target image. Corresponding words, based on the similarity between the initial image feature vector and each initial word vector, determine the adjustment weight value of the initial image feature vector, and adjust the initial image feature vector based on the adjustment weight value to obtain the adjusted image feature vector, which may include:

[0045] The similarities between the initial image semantic vector and each initial word vector are calculated respectively and summed up to serve as the adjustment weight value of the initial image semantic vector; then, based on the adjustment weight value of the initial image semantic vector, the initial image semantic vector is adjusted to obtain an adjusted image semantic vector;

[0046] For each initial block vector, the similarities between the initial block vector and each initial word vector are calculated and summed up to serve as the adjustment weight value of the initial block vector; then based on the adjustment weight value of the initial block vector, the initial block vector is adjusted to obtain an adjusted block vector.

[0047] For each initial word vector, based on the similarity between the initial word vector and the initial image feature vector, the adjustment weight value of the initial word vector is determined, which may correspondingly include:

[0048] For each initial word vector, the similarity between the initial word vector and each initial block vector is calculated as the first similarity; the similarity between the initial word vector and the initial image semantic vector is calculated as the second similarity; the sum of the first similarity and the second similarity is calculated as the adjusted weight value of the initial word vector.

[0049] In the embodiment of the present application, there is no limitation on the specific method for calculating the similarity between two vectors, and a suitable method can be selected according to the actual situation. For example, for the convenience of calculation, the dot product between the two vectors can be directly used as the similarity between the two vectors, or the distance between the two vectors (such as Euclidean distance, Manhattan distance, Chebyshev distance, etc.) can be used as the similarity between the two vectors to improve the accuracy of calculation, etc.

[0050] Step 106, performing coordinate regression based on the adjusted image semantic vector in the adjusted image feature vector to obtain the position information of the target object in the target image; performing named entity recognition based on each adjusted word vector in the adjusted text feature vector to extract the central word in the target text.

[0051] When performing coordinate regression, any existing coordinate regression method may be used, such as a traditional coordinate regression algorithm, or a regression network for implementing coordinate regression, and the like.

[0052] Named Entity Recoginition (NER) aims to identify entities in a string of text and mark the type it refers to. In the embodiment of the present application, it is mainly used to mark each word segment of the target text, so as to determine the entity of the central word therein. In the embodiment of the present application, the specific entity recognition method is not limited, and can be referred to the relevant entity recognition technology, which will not be repeated here.

[0053] See also Figure 2 , Figure 2 This is a schematic diagram of a scene corresponding to the first embodiment of the present application. Figure 2 The schematic diagram shown in the figure illustrates an embodiment of the present application by taking a specific scenario as an example:

[0054] Get a target image containing the target object "shorts" and a target text describing "shorts": "stitched workwear shorts"; get the initial image semantic vector (initial image semantic vector) I 1 , and the target image is divided into blocks to obtain multiple image blocks, and each image block is linearly transformed to obtain the initial block vector I of each image block 2 ,……,I N ; Perform word segmentation on the target text to obtain multiple word units, perform word embedding operations on each word unit, and obtain the initial word vector T of each word unit 1 ,……,T N , wherein N is a natural number greater than 1 (in the embodiment of the present application, the number of image blocks and the number of word units obtained by word segmentation processing are not limited, and the number of image blocks may be greater than the number of word units, or may be less than the number of word units); based on the initial text feature vector, the initial image feature vector is adjusted to obtain an adjusted image feature vector, and based on the initial image feature vector, the initial text feature vector is adjusted to obtain an adjusted text feature vector, specifically: for I 1 (i.e. [REG]), which can be calculated separately from T 1 ,……,T N The similarity (I 1 ·T 1 ,……,I 1 ·T N ) and sum them, and then take the sum of similarities as I 1 The adjustment weight value is used to adjust the initial image semantic vector to the adjusted image semantic vector; 2 ,……,I N Any item in can also be calculated separately with respect to T 1 ,……,T NThe similarities of T are summed up, and the sum of the similarities is used as the adjustment weight value, and each initial block vector is adjusted to an adjusted block vector; for T 1 ,……,T N Any item in can also be calculated separately with I 1 ,……,I N The similarities are summed up, and the sum of the similarities is used as the adjustment weight value, and then each initial word vector is adjusted to an adjusted word vector; coordinate regression is performed based on the adjusted image semantic vector [REG] to obtain the position information (x, y, w, h) of the target object in the target image, where x and y represent the coordinates of the upper left corner of the detection box corresponding to the target object, and w and h represent the width and height of the above detection box, respectively; named entity recognition is performed according to each adjusted word vector in the adjusted text feature vector to obtain the label of each word segmentation, where the labels of the word segmentation "splicing" and "work clothes" are both "O", indicating that they are other words except the central word, that is, they are not the central word; the labels of the word segmentation "short" and "pants" are "B" and "I" respectively, indicating that "short" is the starting position of the central word and "pants" is the middle position of the central word, that is, the central word finally extracted is "shorts".

[0055] According to the data processing method provided by the embodiment of the present application, after obtaining the initial image feature vector at the image modality level and the initial text feature vector at the text modality level, the initial image feature vector and the initial text feature vector are interactively adjusted across modalities. Specifically: for the image feature vector at the image modality level, the adjustment process refers to the key information contained in the initial text feature vector at the text modality level, so that the adjusted image feature vector focuses on representing information related to the above key information, while ignoring or filtering out information with a low degree of correlation with the above key information; correspondingly, for the text feature vector at the text modality level, the adjustment process also refers to the key information contained in the initial image feature vector at the image modality level, so that the adjusted text feature vector focuses on representing information related to the above key information, while ignoring or filtering out information with a low degree of correlation with the above key information. Furthermore, the position regression of the target object is performed based on the adjusted image feature vector, and the accuracy of the regression result is higher; the named entity recognition is performed based on the adjusted text feature vector, and the accuracy of the recognition result is also higher.

[0056] The data processing method of this embodiment can be executed by any appropriate electronic device with data processing capability, including but not limited to: a server, a PC, etc.

[0057] Embodiment 2

[0058] Reference Figure 3 , Figure 3The following is a flowchart of a data processing method according to Embodiment 2 of the present application. Specifically, the data processing method provided in this embodiment includes the following steps:

[0059] Step 302, obtaining an initial image feature vector of the target image and an initial text feature vector of the target text; wherein the initial image feature vector includes: an initial image semantic vector and an initial block vector of each image block in the target image; and the initial text feature vector includes: an initial word vector of each word in the target text.

[0060] Furthermore, in the embodiment of the present application, the above-mentioned initial image feature vector and initial text feature vector can be obtained by the following method:

[0061] Acquire a target image including a target object, a preset initial image semantic vector, and a target text for describing the target object;

[0062] The target image is divided into blocks to obtain multiple image blocks; each image block is linearly transformed to obtain an initial block vector of each image block; the target text is segmented to obtain multiple word units; each word unit is word embedded to obtain an initial word vector of each word unit.

[0063] Step 304: Adjust the initial image semantic vector based on each initial block vector to obtain a transition image semantic vector.

[0064] Specifically, based on the similarity between the initial image semantic vector and each initial block vector, a transition weight value of the initial image feature vector may be determined, and based on the transition weight value, the initial image feature vector may be adjusted to obtain a transition image feature vector.

[0065] Furthermore, the similarities between the initial image semantic vector and each initial block vector can be calculated and summed up to serve as a transition weight value of the initial image feature vector, and the initial image feature vector can be adjusted based on the transition weight value to obtain a transition image feature vector.

[0066] In the embodiment of the present application, there is no limitation on the method for calculating the similarity between two vectors, and it can be selected according to actual conditions.

[0067] Step 306: For each initial block vector, adjust the initial block vector based on the remaining initial block vectors and the initial image semantic vector to obtain a transition block vector.

[0068] Step 304 and step 306 may be performed by a self-attention module for image processing, for example: the module may be any neural network model based on an attention mechanism (such as a transformer-based neural network model), and so on.

[0069] Specifically, for each initial block vector, a transition weight value of the initial block vector can be determined based on the similarity between the initial block vector and other initial block vectors, and the similarity between the initial block vector and the initial image semantic vector, and based on the transition weight value of the initial block vector, the initial block vector is adjusted to obtain a transition block vector.

[0070] Furthermore, the inter-block similarities between the initial block vector and the remaining initial block vectors, as well as the similarities between the initial block vector and the initial image semantic vector can be calculated respectively, and then the sum of the above inter-block similarities and the similarities between the initial block vector and the initial image semantic vector is used as the transition weight value of the initial block vector, and based on the transition weight value of the initial block vector, the initial block vector is adjusted to obtain the transition block vector.

[0071] In the embodiment of the present application, there is no limitation on the method for calculating the similarity between two vectors, and it can be selected according to actual conditions.

[0072] Step 308: For each initial word vector, adjust the initial word vector based on the remaining initial word vectors to obtain a transition word vector.

[0073] This step can be performed by a self-attention module for text processing, for example: the module can be any neural network model based on the attention mechanism (such as a transformer-based neural network model), etc.

[0074] Specifically, for each initial word vector, a transition weight value of the initial word vector can be determined based on the similarity between the initial word vector and the remaining initial word vectors, and based on the transition weight value of the initial word vector, the initial word vector can be adjusted to obtain a transition word vector.

[0075] Furthermore, the similarities between the initial word vector and the remaining initial word vectors can be calculated respectively, and then the calculated similarities are summed up as the transition weight value of the initial word vector, and based on the transition weight value of the initial word vector, the initial word vector is adjusted to obtain a transition word vector.

[0076] In the embodiment of the present application, there is no limitation on the method for calculating the similarity between two vectors, and it can be selected according to actual conditions.

[0077] Step 310: adjusting the transition image feature vector based on the transition text feature vector to obtain an adjusted image feature vector, and adjusting the transition text feature vector based on the transition image feature vector to obtain an adjusted text feature vector.

[0078] The transition image feature vector includes a transition image semantic vector and a transition block vector; and the transition text feature vector includes a transition word vector corresponding to each initial word vector.

[0079] Specifically, for the transition image feature vector, that is, for the transition image semantic vector or each transition block vector, the adjustment weight value of the transition image semantic vector or each transition block vector can be determined based on the similarity between it and the transition word vector of each word element included in the transition text feature vector, and then based on the obtained adjustment weight value, the transition image semantic vector or each transition block vector is adjusted to obtain the adjusted image semantic vector or adjusted block vector, that is, the adjusted image feature vector is obtained.

[0080] For the transition text feature vector, for each transition word vector therein, the adjustment weight value of the transition word vector can be determined based on the similarity between the transition word vector and the transition image feature vector, and the transition word vector can be adjusted based on the adjustment weight value of the transition word vector to obtain an adjusted word vector; wherein the adjusted text feature vector includes each adjusted word vector.

[0081] In the above process, when adjusting each type of transition vector, the product of the corresponding adjustment weight value and the transition vector can be used as the adjusted vector corresponding to the transition vector.

[0082] Further, based on the similarity between the transition image feature vector and each transition word vector, determining an adjustment weight value of the transition image feature vector, and adjusting the transition image feature vector based on the adjustment weight value to obtain an adjusted image feature vector, may include:

[0083] The similarities between the transition image semantic vector and each transition word vector are calculated and summed up respectively, which is used as the adjustment weight value of the transition image semantic vector; then, based on the adjustment weight value of the transition image semantic vector, the transition image semantic vector is adjusted to obtain the adjusted image semantic vector;

[0084] For each transition block vector, the similarities between the transition block vector and each transition word vector are calculated and summed up to serve as the adjustment weight value of the transition block vector; then based on the adjustment weight value of the transition block vector, the transition block vector is adjusted to obtain an adjusted block vector.

[0085] For each transition word vector, based on the similarity between the transition word vector and the transition image feature vector, an adjustment weight value of the transition word vector is determined, which may correspondingly include:

[0086] For each transition word vector, the similarity between the transition word vector and each transition block vector is calculated as the third similarity; the similarity between the transition word vector and the transition image semantic vector is calculated as the fourth similarity; the sum of the third similarity and the fourth similarity is calculated as the adjusted weight value of the transition word vector.

[0087] In the embodiment of the present application, there is no limitation on the specific method for calculating the similarity between two vectors, and a suitable method can be selected according to the actual situation. For example, for the convenience of calculation, the dot product between the two vectors can be directly used as the similarity between the two vectors, or the distance between the two vectors (such as Euclidean distance, Manhattan distance, Chebyshev distance, etc.) can be used as the similarity between the two vectors to improve the accuracy of calculation, etc.

[0088] Step 312: fusing the adjusted image feature vector and the adjusted text feature vector to obtain a fused image feature vector and a fused text feature vector.

[0089] In the embodiments of the present application, there is no limitation on the specific fusion strategy. For example, the fusion process can be performed by directly adding the adjusted image feature vector and the adjusted text feature vector, or the adjusted text feature vector can be processed again based on the adjusted image feature vector, or the adjusted image feature vector can be processed based on the adjusted text feature vector, and so on.

[0090] In this step, it can be performed by an attention module for image-text fusion processing. For example, the module can be any neural network model based on the attention mechanism (such as a transformer-based neural network model), etc.

[0091] Step 314 , performing coordinate regression based on the fused image semantic vector in the fused image feature vector to obtain the position information of the target object in the target image.

[0092] Step 316, performing named entity recognition based on each fused word vector in the fused text feature vector, and extracting the central word in the target text.

[0093] See also Figure 4 , Figure 4 This is a schematic diagram of a scene corresponding to the second embodiment of the present application. Figure 4 The schematic diagram shown in the figure illustrates an embodiment of the present application by taking a specific scenario as an example:

[0094] Figure 4 is Figure 2 Before and after obtaining the adjusted text feature vector, a vector adjustment operation and a vector fusion operation are added respectively.

[0095] The added vector adjustment operation is as follows: Figure 2 After obtaining the initial image semantic vector and the initial block vector, the initial vector is adjusted for the first time based on the self-attention mechanism within the image (image self-attention module), thereby obtaining the transition image semantic vector I 1’ , transition block vector I2’ ,……,I N’ ;exist Figure 2 After obtaining the initial word vector, the initial word vector is adjusted for the first time based on the self-attention mechanism within the text (text self-attention module), thereby obtaining the transition word vector T 1’ , T 2’ ,……,T N’ , and then use the attention mechanism to adjust the transition image feature vector based on the transition text feature vector to obtain the adjusted image feature vector, and adjust the transition text feature vector based on the transition image feature vector to obtain the adjusted text feature vector. Specifically: for I 1’ (i.e. [REG]), which can be calculated separately from T 1’ ,……,T N’ The similarity (I 1’ ·T 1’ ,……,I 1’ ·T N’ ) and sum them, and then take the sum of similarities as I 1’ The adjustment weight value is used to adjust the transition image semantic vector to the adjusted image semantic vector; 2’ ,……,I N’ Any item in can also be calculated separately with respect to T 1’ ,……,T N’ The similarity of T is calculated and summed, and then the sum of the similarities is used as the adjustment weight value, and then each transition block vector is adjusted to an adjusted block vector; for T 1’ ,……,T N’ Any item in can also be calculated separately with I 1’ ,……,I N’ The similarities are summed up, and the sum of the similarities is used as the adjustment weight value, and each transition word vector is adjusted to the adjusted word vector.

[0096] Among them, the added vector fusion operation is specifically as follows: after obtaining the adjusted image feature vector and the adjusted text feature vector, the adjusted image feature vector and the adjusted text feature vector can be input into the image-text fusion module, and the adjusted image feature vector and the adjusted text feature vector are fused to obtain the fused image feature vector and the fused text feature vector.

[0097] Afterwards, coordinate regression is performed based on the fused image semantic vector ([REG]) in the fused image feature vector to obtain the location information of the target object "shorts" in the target image; named entity recognition is performed based on the fused word vector to extract the central word "shorts" of the target text.

[0098] According to the data processing method provided by the embodiment of the present application, after obtaining the initial image feature vector at the image modality level and the initial text feature vector at the text modality level, the initial image feature vector and the initial text feature vector are interactively adjusted across modalities. Specifically: for the image feature vector at the image modality level, the adjustment process refers to the key information contained in the initial text feature vector at the text modality level, so that the adjusted image feature vector focuses on representing information related to the above key information, while ignoring or filtering out information with a low degree of correlation with the above key information; correspondingly, for the text feature vector at the text modality level, the adjustment process also refers to the key information contained in the initial image feature vector at the image modality level, so that the adjusted text feature vector focuses on representing information related to the above key information, while ignoring or filtering out information with a low degree of correlation with the above key information. Furthermore, the position regression of the target object is performed based on the adjusted image feature vector, and the accuracy of the regression result is higher; the named entity recognition is performed based on the adjusted text feature vector, and the accuracy of the recognition result is also higher.

[0099] At the same time, after obtaining the initial image feature vector and the initial text feature vector, on the one hand, the initial image feature vector is first adjusted based on the self-attention mechanism inside the image to obtain the transition image feature vector, and the initial text feature vector is adjusted based on the self-attention mechanism inside the text to obtain the transition text feature vector; then the transition image feature vector at the image modality level and the transition text feature vector at the text modality level are adjusted to obtain the adjusted image feature vector and the adjusted text feature vector; on the other hand, after obtaining the adjusted image feature vector and the adjusted text feature vector, coordinate regression is not directly performed based on the adjusted image feature vector, and named entity recognition is not performed based on the adjusted text feature vector, but the adjusted image feature vector and the adjusted text feature vector are first fused again, and then coordinate regression is performed based on the fused image feature vector, and named entity recognition is performed based on the fused text feature vector. Therefore, the embodiment of the present application can further improve the accuracy of coordinate regression and named entity recognition.

[0100] The data processing method of this embodiment can be executed by any appropriate electronic device with data processing capability, including but not limited to: a server, a PC, etc.

[0101] Embodiment 3

[0102] Reference Figure 5 , Figure 5It is a flow chart of the steps of a data processing method according to the third embodiment of the present application. The application scenario of this embodiment may be: a merchant user sends a target image containing his product and a target text used to describe his product to a server device of an e-commerce platform through an e-commerce platform client device; the server device processes the above target image and target text to determine the product location information in the target image, and extracts the central word in the target text; then, for each product, the target image, product location information, central word and product information of the product (such as: product purchase link, product details link, etc.) are stored in the server device accordingly, so as to provide the buyer with an image search or question search function. Specifically: when the buyer sends an image to be matched or a text to be matched to the server device through a client device, the server device can return the product information that the buyer wants to purchase to the buyer by matching the above image to be matched with a pre-stored target image, or the above text to be matched with a pre-stored central word.

[0103] Specifically, the data processing method provided in this embodiment includes the following steps:

[0104] Step 502: receiving a target image of a commodity and a target text for describing attribute information of the commodity sent by a client device.

[0105] Step 504: Obtain an initial image feature vector of the target image and an initial text feature vector of the target text.

[0106] The initial image feature vector includes: an initial image semantic vector; the initial text feature vector includes: an initial word vector of each word in the target text.

[0107] Step 506: Adjust the initial image feature vector based on the initial text feature vector to obtain an adjusted image feature vector, and adjust the initial text feature vector based on the initial image feature vector to obtain an adjusted text feature vector.

[0108] Step 508, performing coordinate regression based on the adjusted image semantic vector in the adjusted image feature vector to obtain the commodity location information in the target image; performing named entity recognition based on each adjusted word vector in the adjusted text feature vector to extract the central word in the target text.

[0109] In the embodiments of the present application, the specific implementation of steps 504 to 508 can refer to the description of the corresponding steps in embodiment 1 or embodiment 2, and will not be repeated here.

[0110] Step 510 , correspondingly storing the target image, the product location information, the core word, and the product information of the product.

[0111] Specifically, the product information may include: an image of the product, a name of the product, or an information link of the product, such as a purchase link, a detail link, and the like.

[0112] Optionally, in some of the embodiments, the method further comprises:

[0113] Receiving an image to be matched sent by a client device;

[0114] For each stored target image of each product, extract a product image from each target image according to the corresponding product position information;

[0115] Determine, from each product image, a matching product image that matches the image to be matched;

[0116] Returns product information corresponding to the target image to which the matching product image belongs to the client device.

[0117] Specifically, after the server device has stored the target image, product location information, central word and product information of each product, the purchaser can send an image (image to be matched) to the server device and obtain the product information of the product that the purchaser may want to purchase contained in the image to be matched from the server device, thereby realizing the "image search" function.

[0118] Optionally, in some of the embodiments, the method further comprises:

[0119] Receiving text to be matched sent by a client device;

[0120] Determine a matching core word that matches the text to be matched from the stored core words;

[0121] The product information corresponding to the matching core word is returned to the client device.

[0122] Specifically, after the server device has stored the target image, product location information, central word and product information of each product, the buyer can also send text (text to be matched) to the server device to obtain the product information of the product the buyer wants to purchase corresponding to the text to be matched from the server device, thereby realizing the "text search" function.

[0123] According to the data processing method provided by the embodiment of the present application, after obtaining the initial image feature vector at the image modality level and the initial text feature vector at the text modality level, the initial image feature vector and the initial text feature vector are interactively adjusted across modalities. Specifically: for the image feature vector at the image modality level, the adjustment process refers to the key information contained in the initial text feature vector at the text modality level, so that the adjusted image feature vector focuses on representing information related to the above key information, while ignoring or filtering out information with a low degree of correlation with the above key information; correspondingly, for the text feature vector at the text modality level, the adjustment process also refers to the key information contained in the initial image feature vector at the image modality level, so that the adjusted text feature vector focuses on representing information related to the above key information, while ignoring or filtering out information with a low degree of correlation with the above key information. Furthermore, based on the above adjusted image feature vector, the product position regression is performed, and the regression result has a higher accuracy; based on the above adjusted text feature vector, the named entity recognition is performed, and the recognition result has a higher accuracy.

[0124] Embodiment 4

[0125] Reference Figure 6 , Figure 6 1 is a structural block diagram of a data processing device according to Embodiment 4 of the present application. The data processing device provided in the embodiment of the present application includes:

[0126] The initial feature vector acquisition module 602 is used to acquire the initial image feature vector of the target image and the initial text feature vector of the target text; wherein the initial image feature vector includes: the initial image semantic vector; the initial text feature vector includes: the initial word vector of each word in the target text;

[0127] A first adjustment module 604 is used to adjust the initial image feature vector based on the initial text feature vector to obtain an adjusted image feature vector, and to adjust the initial text feature vector based on the initial image feature vector to obtain an adjusted text feature vector;

[0128] The result obtaining module 606 is used to perform coordinate regression according to the adjusted image semantic vector in the adjusted image feature vector to obtain the position information of the target object in the target image; perform named entity recognition according to each adjusted word vector in the adjusted text feature vector to extract the central word in the target text.

[0129] Optionally, in some embodiments, the first adjustment module 604 is specifically configured to:

[0130] Based on the similarity between the initial image feature vector and each initial word vector, an adjustment weight value of the initial image feature vector is determined, and the initial image feature vector is adjusted based on the adjustment weight value to obtain an adjusted image feature vector;

[0131] For each initial word vector, an adjustment weight value of the initial word vector is determined based on the similarity between the initial word vector and the initial image feature vector, and the initial word vector is adjusted based on the adjustment weight value of the initial word vector to obtain an adjusted word vector; wherein the adjusted text feature vector includes each adjusted word vector.

[0132] Optionally, in some of the embodiments, the initial image feature vector further includes: an initial block vector of each image block in the target image;

[0133] The first adjustment module 604, when executing the steps of determining the adjustment weight value of the initial image feature vector based on the similarity between the initial image feature vector and each initial word vector, and adjusting the initial image feature vector based on the adjustment weight value to obtain the adjusted image feature vector, is specifically used for:

[0134] The similarities between the initial image semantic vector and each initial word vector are calculated respectively and summed up to serve as an adjustment weight value of the initial image semantic vector; based on the adjustment weight value of the initial image semantic vector, the initial image semantic vector is adjusted to obtain an adjusted image semantic vector;

[0135] For each initial block vector, the similarities between the initial block vector and each initial word vector are calculated and summed up to serve as an adjustment weight value of the initial block vector; based on the adjustment weight value of the initial block vector, the initial block vector is adjusted to obtain an adjusted block vector.

[0136] Optionally, in some of the embodiments, the initial image feature vector further includes: an initial block vector of each image block in the target image; and the data processing device further includes:

[0137] The second adjustment module is used for, after obtaining the initial image feature vector of the target image and the initial text feature vector of the target text, adjusting the initial image semantic vector based on each initial block vector to obtain a transition image semantic vector; for each initial block vector, adjusting the initial block vector based on the remaining initial block vectors and the initial image semantic vector to obtain a transition block vector; for each initial word vector, adjusting the initial word vector based on the remaining initial word vectors to obtain a transition word vector;

[0138] The first adjustment module 604 is specifically used to: adjust the transition image feature vector based on the transition text feature vector to obtain an adjusted image feature vector, and adjust the transition text feature vector based on the transition image feature vector to obtain an adjusted text feature vector; wherein the transition image feature vector includes a transition image semantic vector and a transition block vector; and the transition text feature vector includes transition word vectors corresponding to each initial word vector.

[0139] Optionally, in some of the embodiments, the result obtaining module 606 is specifically used to:

[0140] The adjusted image feature vector and the adjusted text feature vector are fused to obtain a fused image feature vector and a fused text feature vector;

[0141] Coordinate regression is performed based on the fused image semantic vector in the fused image feature vector to obtain the location information of the target object in the target image; named entity recognition is performed based on each fused word vector in the fused text feature vector to extract the central word in the target text.

[0142] The data processing device of the embodiment of the present application is used to implement the corresponding data processing method in the above-mentioned method embodiment 1 or embodiment 2, and has the beneficial effects of the corresponding method embodiment, which will not be repeated here. In addition, the functional implementation of each module in the data processing device of the embodiment of the present application can refer to the description of the corresponding part in the above-mentioned method embodiment 1 or embodiment 2, which will not be repeated here.

[0143] Embodiment 5

[0144] Reference Figure 7 , shows a schematic diagram of the structure of an electronic device according to Embodiment 5 of the present application. The specific embodiments of the present application do not limit the specific implementation of the electronic device.

[0145] like Figure 7 As shown, the electronic device may include: a processor (processor) 702 , a communication interface (Communications Interface) 704 , a memory (memory) 706 , and a communication bus 708 .

[0146] in:

[0147] The processor 702 , the communication interface 704 , and the memory 706 communicate with each other via a communication bus 708 .

[0148] The communication interface 704 is used to communicate with other electronic devices or servers.

[0149] The processor 702 is used to execute the program 710, and specifically can execute the relevant steps in the above data processing method embodiment.

[0150] Specifically, the program 710 may include program codes, which include computer operation instructions.

[0151] The processor 702 may be a CPU, or an application specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present application. The one or more processors included in the smart device may be processors of the same type, such as one or more CPUs; or processors of different types, such as one or more CPUs and one or more ASICs.

[0152] The memory 706 is used to store the program 710. The memory 706 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.

[0153] Program 710 can be specifically used to enable processor 702 to perform the following operations: obtain an initial image feature vector of a target image and an initial text feature vector of a target text; wherein the initial image feature vector includes: an initial image semantic vector; the initial text feature vector includes: an initial word vector of each word in the target text; adjust the initial image feature vector based on the initial text feature vector to obtain an adjusted image feature vector, and adjust the initial text feature vector based on the initial image feature vector to obtain an adjusted text feature vector; perform coordinate regression based on the adjusted image semantic vector in the adjusted image feature vector to obtain the position information of the target object in the target image; perform named entity recognition based on each adjusted word vector in the adjusted text feature vector to extract the central word in the target text.

[0154] or,

[0155] Program 710 can be specifically used to enable processor 702 to perform the following operations: receive a target image containing a product and a target text for describing the attribute information of the product sent by a client device; obtain an initial image feature vector of the target image and an initial text feature vector of the target text; wherein the initial image feature vector includes: an initial image semantic vector; the initial text feature vector includes: an initial word vector of each word in the target text; adjust the initial image feature vector based on the initial text feature vector to obtain an adjusted image feature vector, and adjust the initial text feature vector based on the initial image feature vector to obtain an adjusted text feature vector; perform coordinate regression based on the adjusted image semantic vector in the adjusted image feature vector to obtain the product location information in the target image; perform named entity recognition based on each adjusted word vector in the adjusted text feature vector to extract the central word in the target text; and store the target image, the product location information, the central word, and the product information of the product accordingly.

[0156] The specific implementation of each step in program 710 can refer to the corresponding description of the corresponding steps and units in the above data processing method embodiment, which will not be repeated here. Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working process of the above-described devices and modules can refer to the corresponding process description in the above method embodiment, which will not be repeated here.

[0157] Through the electronic device of this embodiment, after obtaining the initial image feature vector at the image modality level and the initial text feature vector at the text modality level, the initial image feature vector and the initial text feature vector are adjusted interactively across modalities. Specifically, for the image feature vector at the image modality level, the adjustment process refers to the key information contained in the initial text feature vector at the text modality level, so that the adjusted image feature vector focuses on representing information related to the above key information, while ignoring or filtering out information with a low degree of correlation with the above key information; correspondingly, for the text feature vector at the text modality level, the adjustment process also refers to the key information contained in the initial image feature vector at the image modality level, so that the adjusted text feature vector focuses on representing information related to the above key information, while ignoring or filtering out information with a low degree of correlation with the above key information. Furthermore, the position regression of the target object is performed based on the adjusted image feature vector, and the regression result has a higher accuracy; the named entity recognition is performed based on the adjusted text feature vector, and the recognition result has a higher accuracy.

[0158] An embodiment of the present application also provides a computer program product, including computer instructions, which instruct a computing device to execute operations corresponding to any data processing method in the above-mentioned multiple method embodiments.

[0159] It should be pointed out that, according to the needs of implementation, the various components / steps described in the embodiments of the present application can be split into more components / steps, or two or more components / steps or partial operations of components / steps can be combined into new components / steps to achieve the purpose of the embodiments of the present application.

[0160] The above-mentioned method according to the embodiment of the present application can be implemented in hardware, firmware, or implemented as software or computer code that can be stored in a recording medium (such as a CD ROM, RAM, floppy disk, hard disk or magneto-optical disk), or implemented as a computer code originally stored in a remote recording medium or a non-temporary machine-readable medium downloaded through a network and to be stored in a local recording medium, so that the method described herein can be stored in such software processing on a recording medium using a general-purpose computer, a special-purpose processor or programmable or special-purpose hardware (such as an ASIC or FPGA). It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component (e.g., RAM, ROM, flash memory, etc.) that can store or receive software or computer code, and when the software or computer code is accessed and executed by a computer, a processor or hardware, the data processing method described herein is implemented. In addition, when a general-purpose computer accesses the code for implementing the data processing method shown here, the execution of the code converts the general-purpose computer into a special-purpose computer for executing the data processing method shown here.

[0161] Those of ordinary skill in the art will appreciate that the units and method steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the embodiments of the present application.

[0162] The above implementation methods are only used to illustrate the embodiments of the present application, and are not limitations on the embodiments of the present application. Ordinary technicians in the relevant technical field can make various changes and modifications without departing from the spirit and scope of the embodiments of the present application. Therefore, all equivalent technical solutions also belong to the scope of the embodiments of the present application. The scope of patent protection of the embodiments of the present application should be limited by the claims.

Claims

1. A data processing method, comprising: Acquire an initial image feature vector of a target image and an initial text feature vector of a target text; wherein the initial image feature vector includes: an initial image semantic vector; the initial text feature vector includes: an initial word vector of each word in the target text; Adjusting the initial image feature vector based on the initial text feature vector to obtain an adjusted image feature vector, and adjusting the initial text feature vector based on the initial image feature vector to obtain an adjusted text feature vector; Performing coordinate regression based on the adjusted image semantic vector in the adjusted image feature vector to obtain the position information of the target object in the target image; performing named entity recognition based on each adjusted word vector in the adjusted text feature vector to extract the central word in the target text; Among them, the method also includes: respectively calculating the similarity between the initial image semantic vector and each of the multiple initial word vectors to obtain multiple similarities; determining the sum of the multiple similarities as the adjustment weight value of the initial image semantic vector; based on the adjustment weight value, adjusting the initial image semantic vector to obtain the adjusted image semantic vector.

2. The method according to claim 1, wherein: The adjusting the initial image feature vector based on the initial text feature vector to obtain an adjusted image feature vector, and adjusting the initial text feature vector based on the initial image feature vector to obtain an adjusted text feature vector, comprises: Based on the similarity between the initial image feature vector and each initial word vector, determining an adjustment weight value of the initial image feature vector, and adjusting the initial image feature vector based on the adjustment weight value to obtain an adjusted image feature vector; For each initial word vector, an adjustment weight value of the initial word vector is determined based on the similarity between the initial word vector and the initial image feature vector, and the initial word vector is adjusted based on the adjustment weight value of the initial word vector to obtain an adjusted word vector; wherein the adjusted text feature vector includes each adjusted word vector.

3. The method according to claim 2, wherein: The initial image feature vector also includes: an initial block vector of each image block in the target image; determining an adjustment weight value of the initial image feature vector based on the similarity between the initial image feature vector and each initial word vector, and adjusting the initial image feature vector based on the adjustment weight value to obtain an adjusted image feature vector, including: For each initial block vector, respectively calculate the similarity between the initial block vector and each initial word vector and sum them up to serve as the adjustment weight value of the initial block vector; based on the adjustment weight value of the initial block vector, adjust the initial block vector to obtain an adjusted block vector; The step of determining, for each initial word vector, an adjustment weight value of the initial word vector based on the similarity between the initial word vector and the initial image feature vector comprises: For each initial word vector, the similarity between the initial word vector and each initial block vector is calculated as the first similarity; the similarity between the initial word vector and the initial image semantic vector is calculated as the second similarity; the sum of the first similarity and the second similarity is calculated as the adjusted weight value of the initial word vector.

4. The method according to claim 1, wherein: The initial image feature vector also includes: an initial block vector of each image block in the target image; after obtaining the initial image feature vector of the target image and the initial text feature vector of the target text, the method further includes: The initial image semantic vector is adjusted based on each initial block vector to obtain a transition image semantic vector; for each initial block vector, the initial block vector is adjusted based on the remaining initial block vectors and the initial image semantic vector to obtain a transition block vector; For each initial word vector, adjust the initial word vector based on the other initial word vectors to obtain a transition word vector; The adjusting the initial image feature vector based on the initial text feature vector to obtain an adjusted image feature vector, and adjusting the initial text feature vector based on the initial image feature vector to obtain an adjusted text feature vector, comprises: Adjusting the transition image feature vector based on the transition text feature vector to obtain an adjusted image feature vector, and adjusting the transition text feature vector based on the transition image feature vector to obtain an adjusted text feature vector; Among them, the transition image feature vector includes the transition image semantic vector and the transition block vector; the transition text feature vector includes the transition word vector corresponding to each initial word vector.

5. The method according to claim 1 or 4, wherein: performing coordinate regression according to the adjusted image semantic vector in the adjusted image feature vector to obtain position information of the target object in the target image; Performing named entity recognition according to each adjusted word vector in the adjusted text feature vector to extract the central word in the target text includes: Performing a fusion process on the adjusted image feature vector and the adjusted text feature vector to obtain a fused image feature vector and a fused text feature vector; Coordinate regression is performed based on the fused image semantic vector in the fused image feature vector to obtain the position information of the target object in the target image; named entity recognition is performed based on each fused word vector in the fused text feature vector to extract the central word in the target text.

6. The method according to claim 3 or 4, wherein: The step of obtaining an initial image feature vector of a target image and an initial text feature vector of a target text includes: Acquire a target image including a target object, a preset initial image semantic vector, and a target text for describing the target object; The target image is processed into blocks to obtain a plurality of image blocks; each image block is linearly transformed to obtain an initial block vector of each image block; The target text is segmented to obtain a plurality of word units; each word unit is embedded in a word to obtain an initial word vector of each word unit.

7. A data processing method, applied to a server device, comprising: Receiving a target image of a commodity and a target text for describing attribute information of the commodity sent by a client device; Acquire an initial image feature vector of the target image and an initial text feature vector of the target text; wherein the initial image feature vector includes: an initial image semantic vector; the initial text feature vector includes: an initial word vector of each word in the target text; Adjusting the initial image feature vector based on the initial text feature vector to obtain an adjusted image feature vector, and adjusting the initial text feature vector based on the initial image feature vector to obtain an adjusted text feature vector; Performing coordinate regression based on the adjusted image semantic vector in the adjusted image feature vector to obtain commodity location information in the target image; performing named entity recognition based on each adjusted word vector in the adjusted text feature vector to extract the central word in the target text; Correspondingly storing the target image, the commodity location information, the core word, and the commodity information of the commodity; Among them, the method also includes: respectively calculating the similarity between the initial image semantic vector and each of the multiple initial word vectors to obtain multiple similarities; determining the sum of the multiple similarities as the adjustment weight value of the initial image semantic vector; based on the adjustment weight value, adjusting the initial image semantic vector to obtain the adjusted image semantic vector.

8. The method according to claim 7, wherein: The method further comprises: Receiving an image to be matched sent by a client device; For each stored target image of each product, extract a product image from each target image according to the corresponding product position information; Determine, from each product image, a matching product image that matches the image to be matched; Returning product information corresponding to the target image to which the matching product image belongs to the client device.

9. The method according to claim 7, wherein: The method further comprises: Receiving text to be matched sent by a client device; Determining a matching core word that matches the text to be matched from the stored core words; The commodity information corresponding to the matching core word is returned to the client device.

10. An electronic device comprising: A processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other via the communication bus; The memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform an operation corresponding to the data processing method according to any one of claims 1 to 9.

11. A computer storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the data processing method according to any one of claims 1 to 9 is implemented.

12. A computer program product, comprising computer instructions, wherein the computer instructions instruct a computing device to execute operations corresponding to the data processing method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Target detection method and device

    CN113837257A

  • Article information processing method and device, server and storage medium

    CN113935401A

  • Conversation method and device

    CN114357968A