Data Processing Method, Apparatus, Electronic Device, and Storage Medium

By converting product review data into word vector matrix and inputting joint deep learning models, the rapid and automatic labeling of product reviews is achieved, solving the problem that the large amount of product review data makes it difficult for consumers to quickly find comments of interest, and improving consumers' desire to purchase and the conversion rate of e-commerce platforms.

CN111598596BActive Publication Date: 2025-05-27BEIJING JINGDONG SHANGKE INFORMATION TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201910129643.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-02-21
Publication Date
2025-05-27
Estimated Expiration
2039-02-21

AI Technical Summary

Technical Problem

In the prior art, the amount of product review data is too large, making it difficult for consumers to quickly obtain comment information they are interested in, which eliminates consumers' patience and may lead to consumer loss.

Method used

The technology of converting product review data into word vector matrix and inputting it into the trained joint deep learning model to predict the target label of product reviews, thereby achieving automatic labeling and helping consumers quickly find the comments of interest.

Benefits of technology

It improves data processing efficiency, realizes the rapid and automatic labeling of product reviews, reduces the number of reviews viewed by consumers, increases consumers' desire for purchasing products, and promotes the purchase conversion rate and consumer stickiness of e-commerce platforms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111598596B_ABST
    Figure CN111598596B_ABST
Patent Text Reader

Abstract

Embodiments of the present invention provide a data processing method, apparatus, electronic device, and storage medium, which relate to the field of computer technology. The data processing method includes: obtaining current product review data; obtaining a current word vector matrix according to the current product review data; inputting the current word vector matrix into a trained joint deep learning model to predict one or more target labels of the current product review data. The technical solution of the embodiments of the present invention can automatically label product reviews by using the trained joint deep learning model, which is beneficial for consumers to browse product reviews under corresponding labels according to their personal consumption decision points, and reduces the number of product reviews browsed by consumers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technologies, and in particular, to a data processing method, a data processing device, an electronic device, and a computer-readable storage medium. Background Art

[0002] With the development of Internet e-commerce technologies, consumers' habits of purchasing goods have gradually changed from traditional offline models to online models. When consumers decide whether to purchase a product from a certain e-commerce website, in addition to factors such as the e-commerce platform, product brand, and product details, product reviews are also an aspect that consumers focus on. Consumers can obtain the key points of consumption decision-making needs that they are most concerned about and eager to solve from product reviews, such as information about the appearance, performance, price, logistics, and usage experience of the product.

[0003] In the process of implementing the present invention, the inventors found that there are at least the following problems in the prior art: For some large e-commerce platforms, the reviews of a popular product may reach hundreds of thousands or even millions. It is not only time-consuming and laborious for consumers to obtain the key points of consumption decision-making needs from a large amount of text information, but may even wear out consumers' patience, thereby possibly leading to the loss of consumers.

[0004] Therefore, a new data processing method, device, electronic device, and storage medium are needed.

[0005] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present invention, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0006] The purpose of the embodiments of the present invention is to provide a data processing method, a data processing device, an electronic device, and a computer-readable storage medium, so as to at least to some extent overcome the technical problem that consumers cannot quickly obtain the product reviews they are interested in due to the large amount of product review data in the related art.

[0007] According to a first aspect of the embodiments of the present invention, a data processing method is provided, including: obtaining current product review data; obtaining a current word vector matrix according to the current product review data; inputting the current word vector matrix into a trained joint deep learning model to predict one or more target labels of the current product review data.

[0008] In some exemplary embodiments of the present invention, the joint deep learning model includes a convolutional neural network model, and the convolutional neural network model includes a convolutional layer, a pooling layer, and a fully connected layer; wherein, inputting the current word vector matrix into the trained joint deep learning model to predict one or more target labels of the current product review data includes: inputting the current word vector matrix into the convolutional layer to output a current local feature vector sequence; inputting the current local feature vector sequence into the pooling layer to output a sentence feature vector of a first preset dimension; inputting the sentence feature vector into the fully connected layer to output a semantic vector of a second preset dimension; wherein, the second preset dimension is smaller than the first preset dimension.

[0009] In some exemplary embodiments of the present invention, the convolutional layer includes a plurality of convolutional kernels; wherein, inputting the current word vector matrix into the convolutional layer to output a current local feature vector sequence includes: performing a convolution operation on the current word vector matrix with each convolutional kernel respectively to obtain the context feature corresponding to each convolutional kernel; fusing the context features corresponding to each convolutional kernel to obtain the current local feature vector sequence.

[0010] In some exemplary embodiments of the present invention, the joint deep learning model further includes a recurrent neural network model; wherein, inputting the current word vector matrix into the trained joint deep learning model to predict one or more target labels of the current product review data further includes: inputting the semantic vector into the recurrent neural network model to output the probabilities of each label of the product corresponding to the current product review data; sorting the probabilities of each label and selecting the top k labels with the largest probabilities as the target labels of the current product review data; wherein, k is a positive integer greater than or equal to 1.

[0011] In some exemplary embodiments of the present invention, the method further includes: obtaining the probability distribution of the top k largest probabilities; if the probability distribution meets a preset condition, determining the top k labels with the largest probabilities as the target labels of the current product review data.

[0012] In some exemplary embodiments of the present invention, the method further includes: if the probability distribution does not meet the preset condition, selecting the top m largest probabilities from the top k largest probabilities; determining the labels corresponding to the top m largest probabilities as the target labels of the current product review data; wherein, m is a positive integer less than or equal to k and greater than or equal to 1.

[0013] In some exemplary embodiments of the present invention, obtaining a current word vector matrix according to the current product review data includes: preprocessing the current product review data to obtain a current review word sequence; inputting the current review word sequence into a trained word vector model, and outputting the current word vector matrix.

[0014] In some exemplary embodiments of the present invention, the method further includes: obtaining a training data set, where the training data set includes historical product review data labeled with its labels; training the word vector model according to the historical product review data, and outputting a historical word vector matrix; inputting the historical word vector matrix into the joint deep learning model, and training it according to the labeled labels.

[0015] According to a second aspect of an embodiment of the present invention, there is provided a data processing apparatus, including: a review data acquisition module configured to acquire current product review data; a vector matrix acquisition module configured to obtain a current word vector matrix according to the current product review data; a target label prediction module configured to input the current word vector matrix into a trained joint deep learning model to predict one or more target labels of the current product review data.

[0016] According to a third aspect of an embodiment of the present invention, there is provided an electronic device, including: a processor; and a memory, where a computer-readable instruction is stored on the memory, and when the computer-readable instruction is executed by the processor, it implements the data processing method as described in any one of the above.

[0017] According to a fourth aspect of an embodiment of the present invention, there is provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the data processing method as described in any one of the above.

[0018] One embodiment of the above invention has the following advantages or beneficial effects: on the one hand, because the technical means of converting the current product review data into the current word vector matrix is adopted, the processing efficiency is improved for subsequent operations of the joint deep learning model; on the other hand, because the technical means of inputting the current word vector matrix into the trained joint deep learning model and using the trained joint deep learning model to predict and output the target label of the current product review data is also adopted, the technical effect of automatically tagging the current product review data can be achieved. Consumers can browse the product reviews under the corresponding labels according to their personal consumption decision points, reducing the number of product reviews browsed by consumers, helping consumers quickly obtain the product reviews they are interested in, increasing consumers' desire to purchase products, promoting consumers to place orders on the e-commerce platform, improving the purchase conversion rate of the e-commerce platform, and increasing the consumer stickiness of the e-commerce platform. Therefore, the technical problem in the prior art that consumers cannot quickly obtain the product reviews they are interested in due to a large amount of product review data can be solved.

[0019] Another embodiment of the above invention has the following advantages or beneficial effects: because the technical means of combining a convolutional neural network model and a recurrent neural network model in the joint deep learning model is adopted, the convolutional neural network can well extract the feature information contained in the input current word vector matrix, and at the same time, the recurrent neural network model can effectively consider the feature information contained in a sentence or text as a whole by sequentially modeling the sentence or text. Therefore, applying the joint deep learning model that combines the convolutional neural network model and the recurrent neural network model to automatically tag product reviews can improve the accuracy of the tags.

[0020] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present invention, and are used together with the specification to explain the principles of the present invention. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. In the drawings:

[0022] Figure 1 A schematic flowchart of a data processing method according to some embodiments of the present invention is shown;

[0023] Figure 2 Shows Figure 1 Some embodiment flowcharts of step S130 in

[0024] Figure 3 Shows a schematic flowchart of a data processing method according to other embodiments of the present invention;

[0025] Figure 4 Shows a schematic flowchart of a data processing method according to still other embodiments of the present invention;

[0026] Figure 5 Shows a schematic architecture diagram of a CNN model extracting semantic vectors of text according to some embodiments of the present invention;

[0027] Figure 6 Shows a schematic architecture diagram of an RNN model predicting labels according to some embodiments of the invention;

[0028] Figure 7 Shows a schematic diagram of a product review according to some embodiments of the present invention;

[0029] Figure 8 Shows a schematic block diagram of a data processing apparatus according to some exemplary embodiments of the present invention;

[0030] Figure 9 Shows a schematic structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present invention. Detailed implementation manners

[0031] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this invention will be thorough and complete, and will fully convey the concept of the example embodiments to those skilled in the art. Like reference numerals in the figures denote like or similar parts, and thus their repetitive description will be omitted.

[0032] In addition, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present invention. However, those skilled in the art will realize that the technical solutions of the present invention can be practiced without one or more of the specific details, or can be implemented using other methods, components, devices, steps, etc. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present invention.

[0033] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.

[0034] The flowcharts shown in the accompanying drawings are merely illustrative and not necessarily include all contents and operations / steps, nor are they necessarily executed in the described order. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined. Therefore, the actual execution order may change according to the actual situation.

[0035] In the related art, in order to help consumers quickly obtain the consumer decision-making demand points in product reviews, some e-commerce platforms have done the following work on the order review page to classify product reviews:

[0036] 1) Product rating: According to the product ratings given by consumers, the corresponding product reviews can be divided into three levels: good reviews, medium reviews, and bad reviews. Consumers can obtain valuable information for their own consumption decisions from reviews at different levels.

[0037] 2) Buyer impression: In the buyer impression, consumers can select the tags preset by the operation and maintenance personnel, and at the same time, consumers can also customize tags, enabling other consumers to quickly obtain the consumer decision-making points they are interested in through the buyer impression tags.

[0038] 3) Evaluation and sharing of orders: Consumers can share their shopping experience in the form of pictures for other consumers to refer to.

[0039] However, in the above product rating method, by counting the number of good reviews, medium reviews, and bad reviews, only an intuitive product rating is given. Consumers still need to browse a large amount of text information to obtain their personal consumer decision-making demand points. In the above buyer impression and evaluation and sharing of orders methods, when consumers evaluate the products they purchase, the proportion of the number of selected buyer impression tags or customized tags or shared orders in the entire product review is very small. Consumers still need to browse a large amount of review information from good reviews, medium reviews, and bad reviews.

[0040] Figure 1 The flowchart of the data processing method according to some embodiments of the present invention is shown. The data processing method in the embodiments of the present invention can be executed by any electronic device with computing and processing capabilities, such as a server and / or a terminal device, etc.

[0041] As Figure 1 shown, the data processing method provided in the embodiments of the present invention may include the following steps.

[0042] In step S110, obtain the current product review data.

[0043] In the embodiments of the present invention, the current product review data may come from any e-commerce platform, online food delivery platform, online ticketing platform, online review platform, online taxi-hailing platform, online housekeeping platform, etc. The present invention does not limit this.

[0044] In some embodiments, if the data processing method provided by the embodiments of the present invention is executed by a server, the server may obtain the current product review data from a client, where an application (app) for consumers to submit product reviews or a corresponding website can be installed on the client.

[0045] It should be noted that the product mentioned in the embodiments of the present invention is a broad concept, which may include physical items, such as mobile phones, computers, etc.; or non-physical services, such as housekeeping services, taxi services, etc.

[0046] In step S120, a current word vector matrix is obtained according to the current product review data.

[0047] In an exemplary embodiment, obtaining a current word vector matrix according to the current product review data may include: preprocessing the current product review data to obtain a current review word sequence; inputting the current review word sequence into a trained word vector model, and outputting the current word vector matrix.

[0048] In the embodiments of the present invention, preprocessing the current product review data may include performing operations such as word segmentation, part-of-speech tagging, and removing stop words on the current product review data by using a text preprocessing module. Among them, removing stop words can remove meaningless words such as conjunctions, prepositions, and pronouns according to their parts of speech, and retain meaningful words such as verbs, nouns, and adjectives.

[0049] For example, any suitable word segmentation tool such as open-source stanford word segmentation or jieba word segmentation can be used to complete word segmentation, and these word segmentation tools also support operations such as part-of-speech tagging and removing stop words.

[0050] In the embodiments of the present invention, the word vector model may adopt Word2Vec, but the present invention is not limited thereto, and any suitable word vector model can be used to vectorize the preprocessed current product review data. The characteristic of Word2Vec is to vectorize all words, so that the relationship between words can be quantitatively measured, and the connection between words can be mined.

[0051] Among them, Word2Vec is an NLP (Neuro-Linguistic Programming) tool open-sourced by Google in 2013, including two models: CBOW (Continuous Bag-of-Words) and Skip-gram. Through this word segmentation tool, the text information processing method can be transformed from a traditional high-dimensional sparse vector space into a low-dimensional word vector space. In the embodiments of the present invention, a large amount of product review text data can be mapped to a semantic space through Word2Vec.

[0052] In step S130, the current word vector matrix is input into the trained joint deep learning model to predict one or more target labels of the current product review data.

[0053] In the embodiments of the present invention, the joint deep learning model can combine a Convolutional Neural Network (CNN) model and a Recurrent Neural Network (RNN) model. The CNN model is used to extract the text features of the product review and transform them into high-level semantic vectors, and the RNN model is used to perform multi-label prediction on the product review, so as to assign one or more target labels to the product review. Consumers can browse the product reviews under the corresponding labels according to their desired consumption decision points, reducing the number of product reviews browsed and increasing consumers' desire to purchase products.

[0054] On the one hand, the data processing method provided by the embodiments of the present invention, because it adopts the technical means of converting the current product review data into the current word vector matrix, improves the processing efficiency for the subsequent operation of the joint deep learning model; on the other hand, because it also adopts the technical means of inputting the current word vector matrix into the trained joint deep learning model and using the trained joint deep learning model to predict and output the target labels of the current product review data, the technical effect of automatically tagging the current product review data can be achieved. Consumers can browse the product reviews under the corresponding labels according to their personal consumption decision points, reducing the number of product reviews browsed by consumers, helping consumers quickly obtain the product reviews they are interested in, increasing consumers' desire to purchase products, promoting consumers to place orders on the e-commerce platform, increasing the purchase conversion rate of the e-commerce platform, and increasing the consumer stickiness of the e-commerce platform. Therefore, it can solve the technical problem in the prior art that consumers cannot quickly obtain the product reviews they are interested in due to the large amount of product review data.

[0055] Figure 2 shows Figure 1Schematic flow diagrams of some embodiments of step S130 in []. In the embodiments of the present invention, the joint deep learning model may include a convolutional neural network model, and the convolutional neural network model may include a convolutional layer, a pooling layer, and a fully connected layer.

[0056] As Figure 2 shown, in the embodiments of the present invention, the above step S130 may further include the following steps.

[0057] In step S131, the current word vector matrix is input into the convolutional layer, and a current local feature vector sequence is output.

[0058] In an exemplary embodiment, the convolutional layer may include a plurality of convolutional kernels.

[0059] In an exemplary embodiment, inputting the current word vector matrix into the convolutional layer and outputting the current local feature vector sequence may include: performing a convolution operation on the current word vector matrix with each convolutional kernel respectively to obtain the context features corresponding to each convolutional kernel; fusing the context features corresponding to each convolutional kernel to obtain the current local feature vector sequence.

[0060] In the embodiments of the present invention, the convolutional layer in the CNN model captures the n-gram (n-gram, n is a positive integer greater than or equal to 1) context features of words through a sliding window. It performs a convolution operation on the current word vector matrix output by Word2Vec and a plurality of convolutional kernels to generate an output, that is, the current local feature vector sequence.

[0061] In step S132, the current local feature vector sequence is input into the pooling layer, and a sentence feature vector of a first preset dimension is output.

[0062] In the embodiments of the present invention, after the current local feature vector sequence of local context features is extracted by the convolutional layer, an aggregation operation needs to be performed on these local features to obtain a sentence-level vector feature with a fixed size, regardless of the length of the input word sequence. Therefore, it is necessary to ignore the local features that have no significant impact on the sentence semantics and only retain the feature vectors that have semantics for the sentence in the global feature vector. For this purpose, a pooling operation is used to force the network to retain the most useful local features generated by the convolutional layer, that is, to select the maximum neuron activation value in each pooling area of the feature map.

[0063] In the embodiments of the present invention, the value range of the first preset dimension may be 100-200, but the present invention is not limited thereto and can be set independently according to actual needs.

[0064] In step S133, the sentence feature vector is input into the fully connected layer, and a semantic vector of a second preset dimension is output.

[0065] Among them, the second preset dimension is smaller than the first preset dimension. That is, the fully connected layer is used to reduce the dimension of the sentence feature vector output by the pooling layer, which can reduce the subsequent data processing volume and improve the data processing efficiency.

[0066] In the embodiment of the present invention, between the fully connected layers, vectors generated in multiple vertical directions need to be fused together through a merge layer. After generating the sentence feature vector of the sentence-level vector feature, a non-linear transformation is applied to extract the high-level semantic representation.

[0067] Continue to refer to Figure 2 , the joint deep learning model may further include a recurrent neural network model.

[0068] In step S134, the semantic vector is input into the recurrent neural network model, and the probabilities of each label of the commodity corresponding to the current commodity review data are output.

[0069] Among them, the RNN model is a neural network for processing sequence data, such as time series. Here, it can be used to process commodity review data with multiple label sequences.

[0070] In the embodiment of the present invention, the recurrent neural network model may be an LSTM (Long short term memory) neural network model. Among them, LSTM is one of the most successful variants of RNN, which includes three gates: an input gate, a forget gate, and an output gate. In the embodiment of the present invention, LSTM can be used for label prediction of commodity reviews.

[0071] In the embodiment of the present invention, at least one label can be set in advance for different commodities, or at least one label can be set in advance for different commodity categories. For example, multiple labels such as "logistics", "performance", "appearance", and "screen" are set in advance for the mobile phone category. After that, when a new commodity review of a certain mobile phone is received, when using the RNN model provided by the embodiment of the present invention to predict the target label of the new commodity review, the RNN model can output the probabilities of each label corresponding to the new commodity review.

[0072] In step S135, the probabilities of each label are sorted, and the top k labels with the largest probabilities are selected as the target labels of the current commodity review data.

[0073] Among them, k is a positive integer greater than or equal to 1.

[0074] Just like a movie can have one or more labels such as "love", "action", and "comedy" at the same time, commodity reviews can also be given one or more labels.

[0075] In the embodiments of the present invention, after prediction by the RNN model, the probabilities of each tag of a product review of a certain mobile phone can be output. After sorting the probabilities of each tag in descending order (of course, it can also be sorted in ascending order), the top k tags with the largest probabilities can be selected as its target tags. For example, assuming that the information described in a product review includes "logistics", "price", and "appearance", then this review can be tagged with the three tags "logistics", "price", and "appearance".

[0076] In the data processing method provided by the embodiment of the present invention, because the combined deep learning model combines the technical means of the convolutional neural network model and the recurrent neural network model, the convolutional neural network can well extract the feature information contained in the input current word vector matrix, and at the same time, the recurrent neural network model can effectively consider the feature information contained in the whole sentence or text by sequentially modeling the sentence or text. Therefore, applying the combined deep learning model that combines the convolutional neural network model and the recurrent neural network model to automatically tag product reviews can improve the accuracy of the tags.

[0077] Figure 3 The flowchart of the data processing method according to other embodiments of the present invention is shown.

[0078] As Figure 3 shown, compared with the above embodiments, the difference of the data processing method provided by the embodiment of the present invention is that it may further include the following steps.

[0079] In step S310, obtain the probability distribution of the top k maximum probabilities.

[0080] In the embodiments of the present invention, after prediction by the RNN model, the probabilities of each tag of a product review can be output. After sorting the probabilities of each tag in descending order, the top k maximum probabilities can be selected.

[0081] In step S320, determine whether the probability distribution meets a preset condition; if the probability distribution meets the preset condition, then proceed to step S330; if the probability distribution does not meet the preset condition, then jump to step S340.

[0082] In step S330, if the probability distribution meets the preset condition, then determine the top k tags with the largest probabilities as the target tags of the current product review data.

[0083] For example, assume that the probabilities of the tags "logistics", "appearance", "performance", "screen", and "battery" corresponding to a comment on a mobile phone are 0.3, 0.3, 0.3, 0.05, and 0.05 respectively. Assume further that k = 3. Then the first three probabilities are 0.3, 0.3, and 0.3. From this, it can be known that these first three probabilities are equal. Therefore, the three tags "logistics", "appearance", and "performance" corresponding to these first three probabilities can all be used as the target tags of this comment.

[0084] It should be noted that the above preset conditions do not limit that each of the top k maximum probabilities must be equal. The above is only an example. As long as the difference value between two adjacent probabilities among the top k maximum probabilities is less than the preset threshold, that is, the values of the probabilities among the top k maximum probabilities do not differ too much, it can be considered that the probability distribution meets the preset conditions. The value of the preset threshold can be adjusted independently according to specific requirements, and the present invention does not limit this.

[0085] In step S340, if the probability distribution does not meet the preset conditions, then select the first m maximum probabilities from the top k maximum probabilities.

[0086] For example, assume that the probabilities of the tags "logistics", "appearance", "performance", "screen", and "battery" corresponding to a comment on a mobile phone are 0.8, 0.1, 0.09, 0.06, and 0.05 respectively. Assume further that k = 3. Then the first three probabilities are 0.8, 0.1, and 0.09. From this, it can be known that the difference between the first probability 0.8 and the second probability 0.1 is relatively large. At this time, it can be considered that the probability distribution does not meet the preset conditions. Select the first probability 0.8 from these 3 maximum probabilities, and discard the probabilities 0.1 and 0.09.

[0087] In step S350, determine the tags corresponding to the first m maximum probabilities as the target tags of the current product comment data.

[0088] Wherein, m is a positive integer less than or equal to k and greater than or equal to 1.

[0089] For example, finally determine that "logistics" corresponding to the above first probability 0.8 is the target tag of this mobile phone comment.

[0090] The data processing method provided by the embodiment of the present invention can further improve the accuracy of product comment prediction by processing the probabilities predicted by the RNN model.

[0091] Figure 4 Shows a schematic flowchart of a data processing method according to still other embodiments of the present invention.

[0092] Such as Figure 4As shown, compared with the above embodiments, the data processing method provided by the embodiments of the present invention is different in that it may further include the following steps.

[0093] In step S410, a training data set is obtained, and the training data set includes historical commodity review data with its labels annotated.

[0094] In the embodiments of the present invention, a large amount of historical commodity review data can be collected from e-commerce platforms, etc. first, and then each piece of historical commodity review data can be labeled with its true label by means of manual annotation.

[0095] In step S420, the word vector model is trained according to the historical commodity review data, and a historical word vector matrix is output.

[0096] In the embodiments of the present invention, the word vector model, such as Word2Vec, can be trained according to the historical commodity review data in the training data set. The historical word vector matrix of each piece of historical commodity review data can also be output by using the trained word vector model.

[0097] In step S430, the historical word vector matrix is input into the joint deep learning model, and it is trained according to its annotated label.

[0098] In the embodiments of the present invention, taking the word vector model as Word2Vec and the joint deep learning model as the CNN-RNN model as an example, the training process may include two steps: First, the Word2Vec model is trained by using all the annotated historical commodity review data as unlabeled data. Then the output of the Word2Vec model is fed into the second training step, that is, the supervised training of the CNN-RNN model. For example, a softmax classifier can be used for the upper layer of the RNN for label prediction, and then the cross-entropy loss is propagated back from the RNN to the CNN to update the weights of the CNN-RNN model.

[0099] In the embodiments of the present invention, the Adam optimization algorithm can also be used to make the model converge quickly.

[0100] In some other embodiments, for regularization, the L2 norm can be used to constrain all the weights in the CNN and the RNN.

[0101] Therefore, for each piece of historical commodity review, different lengths of label sequences will be predicted. Ideally, the label sequence of each input historical commodity review completely matches the subset of labels belonging to the historical commodity review (that is, the set of true labels pre-annotated for the historical commodity review).

[0102] Figure 5The figure shows a schematic architecture diagram of a CNN model extracting semantic vectors from text according to some embodiments of the present invention.

[0103] As Figure 5 shown, first, the labeled product review dataset is used as the training dataset to train the Word2Vec model and the CNN-RNN model. After the model training is completed, the current product review is input into the text preprocessing module, and then the output information of the text preprocessing module is input into the trained Word2Vec model. Here, it is assumed that the output information of the text preprocessing module includes n words (n is a positive integer greater than or equal to 1). Then the Word2Vec model will output x1, x2, x3 until xn word vectors. The dimensions of each word vector can be fixed and the same. Combining all the word vectors together, a word vector matrix is obtained. Then the word vector matrix output by the Word2Vec model is input into the convolutional layer of the trained CNN model, and the word vector matrix is respectively convolved with multiple convolutional kernels (1-gram, … 3-gram, more up to n-gram) in the convolutional layer. Then the outputs of each convolutional operation are concatenated to obtain a local feature vector sequence. Then the local feature vector sequence is input into the pooling layer of the CNN model to obtain a sentence feature vector of a fixed size. Then the sentence feature vector is input into the fully connected layer of the CNN model to output the semantic vector X.

[0104] Figure 6 The figure shows a schematic architecture diagram of an RNN model predicting labels according to some embodiments of the invention.

[0105] As Figure 6 shown, here it is assumed that the labels pre-set for the product corresponding to the current product review are successively label_logistics (tag_logistics), label_price (tag_price), label_appearance (tag_appearance), label_performance (tag_performance), label_... (tag_...), and it is assumed that the RNN model is LSTM. Among them, the bottom layer Xt (t is a positive integer greater than or equal to 0) is the input at the t-th moment, that is, the high-level semantic vector of the product review; the middle layer Ht is the hidden state at the t-th moment, which is responsible for the memory function of the entire neural network. Ht (t is a positive integer greater than or equal to 0) is jointly determined by the hidden state of the previous moment and the input of the current moment of this layer; the upper layer Yt (t is a positive integer greater than or equal to 1) is the output at the t-th moment, that is, the softmax layer calculates the probability of each label through a linear transformation.

[0106] In the embodiments of the present invention, the following calculation formula can be adopted:

[0107] Xi = X, i = 0, 1, 2,..., t

[0108] Hi = f(U (i) Xi + W (i) H(i - 1) + b (i) ), i = 1, 2, ..., t

[0109] Yi = g(V i Hi), i = 1, 2, ..., t

[0110] Among them, in the above formula, f is the activation function, U(i) is the weight matrix of the input Xi, W(i) is the weight matrix with the previous value H(i - 1) as the input of this time, and b(i) is the bias term; g is the activation function, and V(i) is the weight matrix of the output layer Yi.

[0111] Specifically, when inputting a product review, it first goes through text preprocessing to output the preprocessed word sequence, uses the word vector model to vectorize the words and feeds them into the CNN model, and sequentially passes through the convolutional layer, pooling layer, and fully connected layer to output the semantic vector; after feeding the output semantic vector as the initial state of label prediction into the LSTM, the entire network can predict the relevant label sequence of the input product review based on the features extracted by the CNN. Among them, the RNN model starts with <START>( <start>)Start label sequence prediction. First, use the top softmax layer to calculate the probability of each label through linear transformation. Then, predict one or more target labels with the highest probability; finally, the label prediction is based on <end>(<End>) marks the end.

[0112] Figure 7 A schematic diagram of product reviews according to some embodiments of the present invention is shown.

[0113] As Figure 7 shown, it is a multi-label classification sample diagram of product reviews. Here, it is assumed that a review of a certain mobile phone is "I have tried the mobile phone for two weeks, and it's generally good. First, I give full marks to XX Mall. It was said to be shipped after one week, but it was shipped in two days and arrived in one day. It's amazing. Let's talk about the mobile phone. The advantages are high cost performance. It's currently the cheapest 845. The camera is also good. The disadvantages are that the screen is just average, the battery drains a bit fast. After looking at the screen for a long time, I feel a bit dizzy. The automatic screen brightness is sometimes very dim. I don't know what's going on. Everything else is fine. Generally, the mobile phone is good. I was about to give up after grabbing it for a few minutes, but it prompted me to try changing a finger, and then I changed a finger and immediately grabbed it. Hahaha". The description information of this product review includes "logistics", "price", "camera", "battery", "screen". Therefore, this review can be labeled with the tags "logistics", "price", "camera", "battery", "screen".

[0114] The data processing method provided by the embodiments of the present invention uses a CNN-RNN model to predict one or more tags for product reviews. Through three-stage operations of training a word vector model by Word2Vec, extracting text features of product reviews by a CNN model, and performing multi-label prediction on product reviews by an RNN model, one or more tags are marked for each product review to help consumers quickly obtain consumption decision demand points from a large amount of text information.

[0115] In addition, in the embodiments of the present invention, a data processing device is also provided. Referring to Figure 8 shown, the data processing device 800 may include: a review data acquisition module 810, a vector matrix acquisition module 820, and a target tag prediction module 830.

[0116] Among them, the review data acquisition module 810 may be configured to acquire current product review data. The vector matrix acquisition module 820 may be configured to obtain a current word vector matrix according to the current product review data. The target tag prediction module 830 may be configured to input the current word vector matrix into a trained joint deep learning model to predict one or more target tags of the current product review data.

[0117] In an exemplary embodiment, the joint deep learning model may include a convolutional neural network model, and the convolutional neural network model may include a convolutional layer, a pooling layer, and a fully connected layer. Among them, the target label prediction module 830 may include: a local feature extraction unit, configured to input the current word vector matrix into the convolutional layer and output a sequence of current local feature vectors; a sentence feature acquisition unit, configured to input the sequence of current local feature vectors into the pooling layer and output a sentence feature vector of a first preset dimension; a semantic vector generation unit, configured to input the sentence feature vector into the fully connected layer and output a semantic vector of a second preset dimension; wherein, the second preset dimension is smaller than the first preset dimension.

[0118] In an exemplary embodiment, the convolutional layer may include a plurality of convolutional kernels. Among them, the local feature extraction unit may include: a context feature extraction subunit, configured to perform a convolution operation on the current word vector matrix with each convolutional kernel respectively to obtain context features corresponding to each convolutional kernel; a context feature fusion subunit, configured to fuse the context features corresponding to each convolutional kernel to obtain the sequence of current local feature vectors.

[0119] In an exemplary embodiment, the joint deep learning model may further include a recurrent neural network model. Among them, the target label prediction module 830 may further include: a label probability prediction unit, configured to input the semantic vector into the recurrent neural network model and output the probabilities of each label of the commodity corresponding to the current commodity review data; a target label selection unit, configured to sort the probabilities of each label and select the top k labels with the largest probabilities as the target labels of the current commodity review data; wherein, k is a positive integer greater than or equal to 1.

[0120] In an exemplary embodiment, the data processing device 800 may further include: a probability distribution acquisition module, configured to acquire the probability distribution of the top k maximum probabilities; a first label determination module, configured to determine the top k labels with the largest probabilities as the target labels of the current commodity review data if the probability distribution meets a preset condition.

[0121] In an exemplary embodiment, the data processing device 800 may further include: a probability selection module, configured to select the top m maximum probabilities from the top k maximum probabilities if the probability distribution does not meet the preset condition; a second label determination module, configured to determine the labels corresponding to the top m maximum probabilities as the target labels of the current commodity review data; wherein, m is a positive integer less than or equal to k and greater than or equal to 1.

[0122] In an exemplary embodiment, the vector matrix obtaining module 820 may include: a word sequence obtaining unit configured to preprocess the current product review data to obtain a current review word sequence; and a vector matrix obtaining unit configured to input the current review word sequence into a trained word vector model and output the current word vector matrix.

[0123] In an exemplary embodiment, the data processing device 800 may further include: a training data obtaining module configured to obtain a training data set, where the training data set includes historical product review data labeled with its tags; a first model training module configured to train the word vector model according to the historical product review data and output a historical word vector matrix; and a second model training module configured to input the historical word vector matrix into the joint deep learning model and train it according to its labeled tags.

[0124] Since each functional module of the data processing device 800 in the exemplary embodiment of the present invention corresponds to the steps in the exemplary embodiment of the above data processing method, it will not be described in detail herein.

[0125] In an exemplary embodiment of the present invention, an electronic device capable of implementing the above method is further provided.

[0126] Next, refer to Figure 9 , which shows a schematic structural diagram of a computer system 900 of an electronic device suitable for implementing the embodiments of the present invention. Figure 9 The shown computer system 900 of the electronic device is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present invention.

[0127] As Figure 9 shown, the computer system 900 includes a central processing unit (CPU) 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage section 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for system operation are also stored. The CPU 901, ROM 902, and RAM 903 are connected to each other through a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0128] The following components are connected to the I / O interface 905: an input section 906 including a keyboard, a mouse, etc.; an output section 907 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. as well as a speaker, etc.; a storage section 908 including a hard disk, etc.; and a communication section 909 including a network interface card such as a LAN card, a modem, etc. The communication section 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to the I / O interface 905 as required. A removable medium 911 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is mounted on the drive 910 as required so that a computer program read therefrom is installed into the storage section 908 as required.

[0129] Specifically, according to an embodiment of the present invention, the processes described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 909, and / or installed from the removable medium 911. When the computer program is executed by a central processing unit (CPU) 901, the above-described functions defined in the system of the present application are executed.

[0130] It should be noted that the computer-readable medium shown in the present invention can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, a computer-readable storage medium can be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present invention, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination of the above.

[0131] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0132] The modules and / or units and / or subunits involved in the embodiments of the present invention can be implemented in software or in hardware. The described modules and / or units and / or subunits can also be provided in a processor. Wherein, the names of these modules and / or units and / or subunits do not, in some cases, constitute a limitation on the modules and / or units and / or subunits themselves.

[0133] As another aspect, the present application also provides a computer-readable medium, which may be included in the electronic device described in the above embodiments; or may exist separately without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the one or more programs are executed by an electronic device, the electronic device is caused to implement the data processing method as described in the above embodiments.

[0134] For example, the electronic device can implement as Figure 1 shown in: Step S110, obtaining current product review data; Step S120, obtaining a current word vector matrix according to the current product review data; Step S130, inputting the current word vector matrix into the trained joint deep learning model to predict one or more target labels of the current product review data.

[0135] It should be noted that although several modules and / or units and / or subunits of a device or apparatus for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present invention, the features and functions of the two or more modules and / or units and / or subunits described above can be embodied in one module and / or unit and / or subunit. Conversely, the features and functions of one module and / or unit and / or subunit described above can be further divided and embodied by multiple modules and / or units and / or subunits.

[0136] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present invention can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present invention.

[0137] Other embodiments of the present invention will be readily apparent to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include known or customary techniques in the art not disclosed herein. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the invention are pointed out by the following claims.

[0138] It should be understood that the present invention is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.< / end> < / start>

Claims

1. A data processing method, characterized in that, comprising: Obtaining current product review data; Obtaining a current word vector matrix according to the current product review data; Inputting the current word vector matrix into a trained joint deep learning model to predict one or more target labels of the current product review data, the joint deep learning model including a convolutional neural network model and a recurrent neural network model, the convolutional neural network model including a convolutional layer, a pooling layer and a fully connected layer, the convolutional layer including a plurality of convolutional kernels, which includes: Performing a convolution operation on the current word vector matrix with each convolutional kernel respectively to obtain context features corresponding to each convolutional kernel; Fusing the context features corresponding to each convolutional kernel to obtain a current local feature vector sequence; Inputting the current local feature vector sequence into the pooling layer to output a sentence feature vector of a first preset dimension; Inputting the sentence feature vector into the fully connected layer to output a semantic vector of a second preset dimension; wherein, the second preset dimension is less than the first preset dimension; Inputting the semantic vector into the recurrent neural network model to output the probabilities of each label of the product corresponding to the current product review data; Sorting the probabilities of each label and selecting the top k labels with the largest probabilities as the target labels of the current product review data; wherein, k is a positive integer greater than or equal to 1.

2. The data processing method according to claim 1, characterized in that, further comprising: Obtaining the probability distribution of the top k largest probabilities; If the probability distribution meets a preset condition, determining the top k labels with the largest probabilities as the target labels of the current product review data.

3. The data processing method according to claim 2, characterized in that, further comprising: If the probability distribution does not meet the preset condition, selecting the top m largest probabilities from the top k largest probabilities; Determining the labels corresponding to the top m largest probabilities as the target labels of the current product review data; wherein, m is a positive integer less than or equal to k and greater than or equal to 1.

4. The data processing method according to claim 1, characterized in that, Obtaining a current word vector matrix according to the current product review data, including: Preprocessing the current product review data to obtain a current review word sequence; Inputting the current review word sequence into a trained word vector model to output the current word vector matrix.

5. The data processing method according to claim 4, characterized in that, further comprising: Obtaining a training data set, the training data set including historical product review data labeled with its labels; Training the word vector model according to the historical product review data and outputting a historical word vector matrix; Inputting the historical word vector matrix into the joint deep learning model and training it according to its labeled labels.

6. A data processing device, characterized in that, comprising: A review data acquisition module configured to acquire current product review data; A vector matrix acquisition module configured to obtain a current word vector matrix according to the current product review data; The target label prediction module is configured to input the current word vector matrix into the trained joint deep learning model to predict one or more target labels of the current product review data. The joint deep learning model includes a convolutional neural network model and a recurrent neural network model. The convolutional neural network model includes a convolutional layer, a pooling layer, and a fully connected layer. The convolutional layer includes multiple convolutional kernels, and it includes: performing a convolution operation on the current word vector matrix with each convolutional kernel to obtain the context features corresponding to each convolutional kernel; fusing the context features corresponding to each convolutional kernel to obtain the current local feature vector sequence; inputting the current local feature vector sequence into the pooling layer to output a sentence feature vector of a first preset dimension; inputting the sentence feature vector into the fully connected layer to output a semantic vector of a second preset dimension; wherein, the second preset dimension is smaller than the first preset dimension; inputting the semantic vector into the recurrent neural network model to output the probabilities of each label of the product corresponding to the current product review data; sorting the probabilities of each label and selecting the top k labels with the largest probabilities as the target labels of the current product review data; wherein, k is a positive integer greater than or equal to 1.

7. An electronic device, characterized in that it includes: a processor; and a memory, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by the processor, the data processing method according to any one of claims 1 to 5 is implemented.

8. A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the data processing method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Commodity comment data labeling system and method based on hierarchical AP clustering

    CN107633007A

  • Emotion analysis method based on context word vectors and depth learning

    CN108427670A