Label extraction method and device

Through the combination of emotional classification model and comprehensive labels, the special comments in hotel reviews are screened, and the problem of low success rate of special comment mining in the existing technology is solved, achieving more accurate and efficient review extraction.

CN120429437APending Publication Date: 2025-08-05HUAWEI TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410157431.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-02-01
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

It is difficult for existing technology to effectively discover high-quality special comments in hotel reviews, resulting in a low success rate of special comments.

Method used

By introducing an emotion classification model, classifying user comments, combining emotional scores and comprehensive labels, special comments are selected.

Benefits of technology

It improves the success rate of mining special comments, ensures that the extracted comments are more in line with users' preferences, and reduces the impact of meaninglessness and repeated comments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120429437A_ABST
    Figure CN120429437A_ABST
Patent Text Reader

Abstract

The invention provides a label extraction method and device in the field of artificial intelligence, which are used for analyzing user comments and determining characteristic comments in the user comments in combination with emotion scores of the user comments and comprehensive labels. The method comprises the following steps: classifying user comments according to a sentiment classification model to obtain sentiment categories of the user comments; according to sentiment words in the user comments and related vocabularies of the sentiment words, calculating to obtain sentiment scores of the user comments; according to the emotion category, analyzing the user comments to obtain a comprehensive label of the user comments, the comprehensive label being used for summarizing feature information of the user comments; and obtaining characteristic comments in the user comments according to the emotion scores and the comprehensive labels.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and in particular to a label extraction method and device. Background Art

[0002] Hotel reviews are a valuable asset for online travel agencies (OTAs) and travel management companies (TMCs). Filtering valid user reviews from these reviews and deriving insights into users' diverse experiences with hotels is crucial for enriching a hotel's image, supporting hotel searches, guiding user orders, and helping other users quickly understand hotel information. Therefore, identifying high-quality reviews, summarizing them into comprehensive hotel review tags, and particularly identifying distinctive reviews from these reviews, has become a key technical challenge for platforms.

[0003] At present, relevant solutions have been proposed for extracting labels from reviews. By constructing a label system, if the keywords in the label system appear in the review, the review is judged to be a valid review. Alternatively, by marking a certain number of featured reviews, the number of stop words, the number of featured words, the number of dimension words and the length of the review in the featured reviews are taken as quadruple, and the similarity between the quadruple of unlabeled reviews and the quadruple of featured reviews is calculated to determine whether the review is a featured review.

[0004] Existing solutions can only determine whether a review is a featured review by whether there are keywords or the number of specific words, which results in some featured reviews being unable to be extracted, and the success rate of mining featured reviews is low. Summary of the Invention

[0005] This application provides a label extraction method and device, which are applied in the field of artificial intelligence, for analyzing user reviews, combining the sentiment scores and comprehensive labels of user reviews, and determining the distinctive reviews in the user reviews.

[0006] In view of this, in the first aspect, the present application provides a label extraction method. First, user comments are obtained, and the user comments include the user's evaluation or opinion on a certain thing; after obtaining the user comments, the user comments can be classified according to the sentiment classification model to obtain the sentiment category of the user comments, and the user comments include one or more first review sentences; based on the sentiment words in the user comments and the associated words associated with the sentiment words, the sentiment score of the user comments can be calculated, and the sentiment score is used to express the user's liking or affirmation of the review object; based on the sentiment category of the user comments, the user comments are analyzed to obtain a comprehensive label of the user comments, and the comprehensive label is used to summarize the characteristic information of the user comments; based on the aforementioned sentiment score and comprehensive label, characteristic comments can be determined from the user comments.

[0007] In this embodiment, the sentiment scores of user reviews can be incorporated into the user reviews, combined with the comprehensive tags of the user reviews, to extract distinctive reviews from the user reviews. Screening distinctive reviews based on sentiment scores takes into account user experience, making the extracted distinctive reviews more in line with user preferences. Furthermore, distinctive reviews can be extracted from the comprehensive tags, avoiding situations where distinctive reviews cannot be extracted, thereby improving the success rate of distinctive review mining.

[0008] In a possible implementation, the aforementioned classification of user reviews according to the sentiment classification model to obtain the sentiment categories of the user reviews may include: dividing the user reviews to obtain one or more first review sentences, where the user reviews are obtained after being reviewed by the content review model; using the one or more first review sentences as input to the sentiment classification model to obtain sentiment categories, where the sentiment categories include one or more of positive, negative, or neutral.

[0009] In an embodiment of the present application, user reviews can be screened through a content review model to filter out meaningless and repetitive reviews and reviews involving illegal and banned words, thereby reducing the workload of analyzing user reviews and improving the efficiency of mining distinctive reviews.

[0010] In a possible implementation, the aforementioned calculation of the sentiment score of the user review based on the sentiment words and associated words in the user review may include: determining the sentiment words and associated words by querying a sentiment vocabulary, the sentiment vocabulary including multiple sentiment words and multiple associated words; and calculating the sentiment score of the first review sentence based on the sentiment words and associated words.

[0011] In an embodiment of the present application, the sentiment score of a user review sentence can be calculated. The sentiment score represents the user's attitude towards the review object, and can provide a screening basis for the subsequent extraction of featured reviews, helping to screen out positive user reviews.

[0012] In a possible implementation, the aforementioned analysis of user comments according to sentiment categories to obtain comprehensive labels for user comments may include: preprocessing the first review sentence to obtain a second review sentence; if the second review sentence conforms to the target syntactic structure, then analyzing the second review sentence using dependency syntactic analysis according to the sentiment category to obtain a comprehensive label, where the target syntactic structure includes one or more of a subject-predicate structure, a subject-predicate relationship, or a verb-object relationship; if the second review sentence does not conform to the target syntactic structure, then analyzing the second review sentence using a sentence matching model to obtain a comprehensive label; if the second review sentence does not conform to the target syntactic structure and does not conform to the target sentence pattern, then analyzing the second review sentence based on a keyword library to obtain a comprehensive label.

[0013] In the embodiment of the present application, a variety of methods can be used to extract comprehensive tags from user reviews to improve the success rate of comprehensive tag extraction, thereby improving the success rate of subsequent extraction of special tags from user reviews based on comprehensive tags.

[0014] In a possible implementation, the method may further include: segmenting the second review sentence to obtain a plurality of words.

[0015] In a possible implementation, the aforementioned analysis of the second review sentence using dependency syntactic analysis based on the sentiment category to obtain a comprehensive label may include: determining the evaluation vocabulary of the second review sentence based on the vocabulary; calculating the first similarity between the evaluation vocabulary and the opinion subscript based on the sentiment category of the second review sentence, the opinion subscript including the evaluation of the first review object, the opinion subscript including the positive opinion subscript and the negative opinion subscript, the positive opinion subscript being a positive evaluation of the review object, and the negative opinion subscript being a negative evaluation of the review object; obtaining the second review object of the second review sentence; and obtaining a comprehensive label based on the first similarity and the second review object.

[0016] In the embodiment of the present application, the similarity between words is taken into consideration, and similarity processing is performed on the evaluation words in the user comments. The comprehensive labels in the user comments can be extracted based on the calculated similarity between the evaluation words and the opinion subscripts, thereby avoiding the extraction of only words consistent with the opinion subscripts and improving the success rate of comprehensive label extraction.

[0017] In a possible implementation, the aforementioned calculation of the first similarity between the evaluation vocabulary and the opinion subscript based on the sentiment category of the second review sentence may include: when the sentiment category is positive, calculating the first similarity between the evaluation vocabulary and the positive opinion subscript.

[0018] In a possible implementation, the aforementioned acquisition of the second comment object of the second comment sentence may include: acquiring the part of speech of the vocabulary; determining the second comment object of the second comment sentence based on the part of speech and the entity knowledge base, the entity knowledge base including the first comment object, the first keyword and the opinion subscript.

[0019] In a possible implementation, the aforementioned determination of the second review object of the second review sentence based on part of speech and the entity knowledge base may include: determining the review subject of the second review sentence based on part of speech; and determining the second review object based on the similarity between the review subject and the first keyword in the entity knowledge base.

[0020] In a possible implementation, the aforementioned obtaining of a comprehensive label based on the first similarity and the second comment object may include: if the first similarity is higher than or equal to a first preset value, then using the second comment object and the positive opinion subscript as a comprehensive label; if the first similarity is lower than the first preset value, then using the second comment object and the evaluation vocabulary as a comprehensive label.

[0021] In a possible implementation, the aforementioned use of a sentence matching model to analyze the second review sentence to obtain a comprehensive label may include: vectorizing the second review sentence to obtain a processed third review sentence; calculating a second similarity between the third review sentence and a model file, where the model file is a vectorized file of the target sentence; and obtaining a comprehensive label based on the second similarity.

[0022] In a possible implementation, the aforementioned obtaining of the comprehensive label according to the second similarity may include: if the second similarity is higher than or equal to a second preset value, using the corresponding target sentence as the comprehensive label.

[0023] In a possible implementation, the aforementioned analysis of the second review sentence based on the keyword library to obtain a comprehensive label may include: calculating a third similarity between the vocabulary and the second keyword in the keyword library; if the third similarity is higher than or equal to a third preset value, and the sentiment category of the second review sentence is positive, then using the second keyword as the comprehensive label.

[0024] In a possible implementation, the aforementioned obtaining of featured comments in user reviews based on sentiment scores and comprehensive labels may include: determining the priority of the second review object based on the comprehensive label; and determining the featured comments in user reviews based on the priority and sentiment scores of the second review object.

[0025] In a possible implementation, the aforementioned determination of featured comments in user reviews based on the priorities of the second review objects and sentiment scores may include: if the priorities of the second review objects are the same, comparing the sentiment scores, and determining the featured comments in user reviews based on the sorting of the sentiment scores; if the priorities of the second review objects are different, determining the featured comments in user reviews based on the sorting of the priorities of the second review objects.

[0026] In a second aspect, the present application provides a label extraction device, comprising:

[0027] The analysis module is used to classify user comments according to the sentiment classification model to obtain the sentiment category of the user comments;

[0028] A calculation module is used to calculate the sentiment score of the user review based on the sentiment words and related words in the user review;

[0029] The above analysis module is also used to analyze user reviews according to sentiment categories to obtain comprehensive labels for user reviews. The comprehensive labels are used to summarize the characteristic information of user reviews.

[0030] The processing module is used to obtain the characteristic comments in the user comments based on the sentiment score and comprehensive labels.

[0031] In one possible implementation, the above-mentioned analysis module is specifically used to divide user comments to obtain one or more first review sentences, and the user comments are obtained after being reviewed by a content review model; the one or more first review sentences are used as input to a sentiment classification model to obtain sentiment categories, and the sentiment categories include one or more of positive, negative or neutral.

[0032] In a possible implementation, the calculation module is specifically used to determine the sentiment words and related words by querying the sentiment word library, where the sentiment word library includes multiple sentiment words and multiple related words; and calculate the sentiment score of the first review sentence based on the sentiment words and related words.

[0033] In one possible implementation, the above-mentioned analysis module is specifically used to preprocess the first comment sentence to obtain a second comment sentence; if the second comment sentence conforms to the target syntactic structure, the second comment sentence is analyzed using dependency syntactic analysis according to the sentiment category to obtain a comprehensive label, and the target syntactic structure includes one or more of the subject-predicate structure, the attributive-predicate relationship, or the verb-object relationship; if the second comment sentence does not conform to the target syntactic structure, the second comment sentence is analyzed using a sentence matching model to obtain a comprehensive label; if the second comment sentence does not conform to the target syntactic structure and does not conform to the target sentence pattern, the second comment sentence is analyzed based on the keyword library to obtain a comprehensive label.

[0034] In a possible implementation, the device further includes: a word segmentation module, configured to segment the second review sentence to obtain a plurality of words.

[0035] In a possible implementation, the above-mentioned analysis module is specifically used to determine the evaluation vocabulary of the second review sentence based on the vocabulary; calculate the first similarity between the evaluation vocabulary and the opinion subscript based on the sentiment category of the second review sentence, the opinion subscript includes the evaluation of the first review object, the opinion subscript includes a positive opinion subscript and a negative opinion subscript, the positive opinion subscript is a positive evaluation of the review object, and the negative opinion subscript is a negative evaluation of the review object; obtain the second review object of the second review sentence; and obtain a comprehensive label based on the first similarity and the second review object.

[0036] In a possible implementation, the analysis module is specifically configured to calculate a first similarity between the evaluation vocabulary and the positive opinion subscript if the sentiment category is positive.

[0037] In a possible implementation, the above-mentioned analysis module is specifically used to obtain the part of speech of the vocabulary; determine the second comment object of the second comment sentence based on the part of speech and the entity knowledge base, and the entity knowledge base includes the first comment object, the first keyword and the opinion subscript.

[0038] In a possible implementation, the analysis module is specifically configured to determine the review subject of the second review sentence based on the part of speech; and determine the second review object based on the similarity between the review subject and the first keyword in the entity knowledge base.

[0039] In a possible implementation, the above-mentioned analysis module is specifically used to use the second comment object and the positive opinion subscript as a comprehensive label if the first similarity is higher than or equal to a first preset value; if the first similarity is lower than the first preset value, use the second comment object and the evaluation vocabulary as a comprehensive label.

[0040] In one possible implementation, the above-mentioned analysis module is specifically used to vectorize the second comment sentence to obtain a processed third comment sentence; calculate the second similarity between the third comment sentence and the model file, where the model file is a vectorized file of the target sentence; and obtain a comprehensive label based on the second similarity.

[0041] In a possible implementation, the analysis module is specifically configured to use the corresponding target sentence as a comprehensive label if the second similarity is higher than or equal to a second preset value.

[0042] In a possible implementation, the above-mentioned analysis module is specifically used to calculate the third similarity between the vocabulary and the second keyword in the keyword library; if the third similarity is higher than or equal to a third preset value, and the sentiment category of the second review sentence is positive, the second keyword is used as a comprehensive label.

[0043] In a possible implementation, the processing module is specifically configured to determine the priority of the second review object based on the comprehensive tag; and determine the featured review in the user review based on the priority and sentiment score of the second review object.

[0044] In one possible implementation, the above-mentioned processing module is specifically used to compare the sentiment scores if the priorities of the second review objects are the same, and determine the distinctive reviews in the user reviews based on the sorting of the sentiment scores; if the priorities of the second review objects are different, determine the distinctive reviews in the user reviews based on the sorting of the priorities of the second review objects.

[0045] In a third aspect, the present application provides a label extraction device, which includes: a processor, a memory, an input / output device, and a bus; the memory stores computer instructions; when the processor executes the computer instructions in the memory, the memory stores computer instructions; when the processor executes the computer instructions in the memory, it is used to implement any one of the implementation methods of the first aspect.

[0046] In a fourth aspect, an embodiment of the present application provides a chip system comprising a processor and an input / output port, wherein the processor is used to implement the processing functions involved in the method described in the first aspect above, and the input / output port is used to implement the transceiver functions involved in the method described in the first aspect above.

[0047] In one possible design, the chip system also includes a memory, which is used to store program instructions and data for implementing the functions involved in the method described in the first aspect above.

[0048] The chip system may be composed of chips, or may include chips and other discrete devices.

[0049] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium having computer instructions stored therein; when the computer instructions are executed on a computer, the computer executes the method described in the possible implementation of the first aspect.

[0050] In a sixth aspect, an embodiment of the present application provides a computer program product. The computer program product includes a computer program or instructions, which, when executed on a computer, causes the computer to execute the method described in the possible implementation of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 The complete process of extracting featured reviews provided for this application;

[0052] Figure 2 A schematic diagram of a system framework provided for this application;

[0053] Figure 3 A flowchart of a tag extraction method provided in this application;

[0054] Figure 4 Comparison chart of model evaluation indicators provided for this application;

[0055] Figure 5 Comparison chart of model training time provided for this application;

[0056] Figure 6 A flowchart of another tag extraction method provided in this application;

[0057] Figure 7 The secondary comprehensive label distribution map obtained by the label extraction method provided in this application;

[0058] Figure 8 A schematic diagram of the structure of a label extraction device provided in this application;

[0059] Figure 9 A schematic diagram of the structure of another tag extraction device provided in this application;

[0060] Figure 10 A schematic diagram of the structure of a chip provided in this application. DETAILED DESCRIPTION

[0061] The following will describe the technical solutions in the embodiments of this application in conjunction with the drawings in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of this application.

[0062] The method provided in this application can be applied to artificial intelligence (AI) scenarios. AI is a theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a branch of computer science that attempts to understand the essence of intelligence and produce a new type of intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making. Research in the field of artificial intelligence includes robotics, natural language processing, computer vision, decision-making and reasoning, human-computer interaction, recommendation and search, basic AI theory, etc.

[0063] First, we describe the overall workflow of AI systems. The following section elaborates on the aforementioned AI framework from two perspectives: the "intelligent information chain" (horizontal axis) and the "IT value chain" (vertical axis). The "intelligent information chain" reflects the entire process from data acquisition to processing. For example, this could be the general process of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, and intelligent execution and output. Throughout this process, data undergoes a condensed journey from "data-information-knowledge-wisdom." The "IT value chain," spanning the underlying infrastructure of human intelligence, information (provided and processed by technology), and the system's industrial ecosystem, reflects the value that AI brings to the information technology industry.

[0064] (1) Infrastructure

[0065] Infrastructure provides computing power for AI systems, enabling communication with the outside world and supporting this through a foundational platform. External communication occurs through sensors; computing power is provided by intelligent chips (CPUs, NPUs, GPUs, ASICs, FPGAs, and other hardware accelerators). The foundational platform includes a distributed computing framework and network-related platform guarantees and support, including cloud storage and computing, and interconnected networks. For example, sensors communicate with the outside world to acquire data, which is then fed into the intelligent chips within the distributed computing system provided by the foundational platform for computation.

[0066] (2) Data

[0067] Data above the infrastructure layer represents data sources for AI. This data includes graphics, images, voice, and text, as well as IoT data from traditional devices. This includes business data from existing systems and sensor data such as force, displacement, liquid level, temperature, and humidity.

[0068] (3) Data processing

[0069] Data processing generally includes data training, machine learning, deep learning, search, reasoning, decision-making, etc.

[0070] Among them, machine learning and deep learning can symbolize and formalize data for intelligent information modeling, extraction, preprocessing, and training.

[0071] Reasoning refers to the process of simulating human intelligent reasoning in computers or intelligent systems, using formalized information to perform machine thinking and solve problems based on reasoning control strategies. Typical functions are search and matching.

[0072] Decision-making refers to the process of making decisions after intelligent information is reasoned, and usually provides functions such as classification, sorting, and prediction.

[0073] (4) General ability

[0074] After the data has undergone the data processing mentioned above, some general capabilities can be further formed based on the results of the data processing, such as algorithms or a general system, for example, translation, text analysis, computer vision processing, speech recognition, image recognition, etc.

[0075] (5) Smart products and industry applications

[0076] Smart products and industry applications refer to the products and applications of artificial intelligence systems in various fields. They are the encapsulation of the overall artificial intelligence solution, which productizes intelligent information decision-making and realizes practical application. Its application areas mainly include: smart terminals, smart transportation, smart medical care, autonomous driving, smart cities, etc.

[0077] The embodiments of the present application involve related applications of neural networks. In order to better understand the solutions of the embodiments of the present application, the following first introduces the relevant terms and concepts of the neural networks that may be involved in the embodiments of the present application, as well as the relevant terms and concepts that may be involved in the embodiments of the present application.

[0078] (1) Neural Network

[0079] A neural network can be composed of neural units. A neural unit can refer to an operation unit with xs and intercept 1 as input. The output of the operation unit can be shown as formula (1-1):

[0080]

[0081] Where s = 1, 2, ... n, n is a natural number greater than 1, Ws is the weight of xs, and b is the bias of the neural unit. f is the activation function of the neural unit, which is used to introduce nonlinear characteristics into the neural network to convert the input signal of the neural unit into the output signal. The output signal of the activation function can be used as the input of the next convolutional layer, and the activation function can be a sigmoid function. A neural network is a network formed by connecting multiple single neural units mentioned above, that is, the output of one neural unit can be the input of another neural unit. The input of each neural unit can be connected to the local receptive field of the previous layer to extract the features of the local receptive field. The local receptive field can be an area composed of several neural units.

[0082] (2) Convolutional Neural Network

[0083] A convolutional neural network (CNN) is a deep neural network with a convolutional architecture. It consists of a feature extractor consisting of convolutional layers and subsampling layers, which can be viewed as a filter. A convolutional layer is a layer of neurons that performs convolution processing on the input signal. In a convolutional layer of a CNN, a neuron can only connect to a subset of neurons in adjacent layers. A convolutional layer typically contains several feature planes, each of which is composed of a rectangular arrangement of neurons. Neurons in the same feature plane share weights, which are referred to as convolution kernels. Shared weights can be understood as position-independent feature extraction. Convolution kernels can be formalized as matrices of random size, and during CNN training, the kernels can be learned to acquire reasonable weights. Furthermore, shared weights have the direct benefit of reducing the number of connections between layers of the CNN, thereby reducing the risk of overfitting.

[0084] (3) Long Short-Term Memory Network

[0085] Long short-term memory (LSTM) networks are a type of recurrent neural network. The LSTM network structure consists of one or more units with both forgetting and remembering capabilities. Key components include the forget gate, input gate, and output gate. The forget gate determines which information is forgotten from the memory unit. It uses a sigmoid activation function, which outputs a value between 0 and 1, indicating the proportion of retained information. The input gate determines which new information is stored in the memory unit. It consists of two parts: a sigmoid activation function to determine the update, and a tanh activation function to generate candidate values. The output gate determines the next hidden state (i.e., the output at the next time step).

[0086] (4) FastText Model

[0087] The FastText model is a supervised model that takes a sequence of words (a paragraph of text or a sentence) as input and outputs the probability that the sequence belongs to different categories. It has a wide range of applications in natural language processing, including text classification and sentiment analysis. It offers the advantages of high speed and accuracy.

[0088] (5) Chi-square test

[0089] The chi-square test is a commonly used hypothesis method based on the chi-square distribution. It is mainly used to compare the correlation analysis of two or more sample rates (constituent ratios) and two categorical variables. Its fundamental idea is to compare the degree of fit between the theoretical frequency and the actual frequency or the goodness of fit problem. The basic principle is to first assume that H0 is true, that is, there is no difference between the observed frequency and the expected frequency, based on the previously calculated value, first assume that H0 is true, based on the previously calculated chi-square test, 2 The value indicates the degree of deviation between the observed value and the theoretical value. 2 The distribution and degrees of freedom can determine the probability P of the current statistic and more extreme cases if the H0 hypothesis is true. If the current statistic is greater than the P value, it means that the observed value deviates too much from the theoretical value, and the null hypothesis H0 should be rejected, indicating that there is a significant difference between the observed and expected frequencies; otherwise, the null hypothesis H0 cannot be rejected, and it cannot be concluded that the actual situation represented by the sample is different from the theoretical hypothesis.

[0090]

[0091] Among them, A is the observed value, E is the theoretical value, k is the number of observed values, n is the frequency, p is the theoretical frequency, and n*p is the theoretical frequency (theoretical value).

[0092] (6) Word frequency-inverse text frequency

[0093] Term frequency-inverse document frequency (TF-IDF) is a statistical method used to assess the importance of a word to a document set or a document in a corpus. It is widely used in information retrieval and text mining and is a common weight calculation method. TF stands for term frequency, which is the frequency with which a word appears in the current document, while IDF stands for inverse document frequency, which is the inverse of the frequency with which a word appears in a document. TF-IDF is actually the product of TF and IDF. If a word appears frequently in a document set, then the TF value is high. If a word appears rarely or not at all in a document, then the IDF value is high. Therefore, the larger the TF-IDF value, the more important the word is to the document. Using the TF-IDF algorithm, we can extract the features of important words in a document set, thereby performing operations such as classification and clustering.

[0094] The following describes the complete process of extracting featured reviews in this embodiment of the application.

[0095] See also Figure 1 , Figure 1 It shows the complete process of extracting featured reviews and the hardware involved in the process. Users can review hotels through terminal devices, where the terminal devices can be mobile phones, PCs or tablets, and the specific details are not limited here. User review information will be transmitted to the inventory server in real time. Before extracting featured reviews from user reviews, the user reviews in the inventory server can be reviewed for content, and reviews containing pornographic, advertising, insulting or spamming content can be removed. Featured tags are extracted from the reviewed user reviews and stored in the inventory server. The extracted featured tags can be used to display the hotel details page to help users quickly understand hotel-related information. They can also be used as a list page for hotel selection to provide a basis for users to choose when booking a hotel.

[0096] It should be understood that the method provided in this application can be applied not only to extracting featured comments from hotel reviews, but also to other evaluations of services, products or facilities. For example, it can be applied to extracting featured comments from movie reviews, and it can also be applied to reviews of travel services, and it can also be applied to reviews of shopping products. The specific details are not limited here.

[0097] The following introduces the system framework provided by the embodiments of the present application.

[0098] See also Figure 2, an embodiment of the present application provides a system framework 200. As shown in the system framework 200, the content review module 201 is used to review the acquired user reviews to obtain the reviewed user reviews. Based on the reviewed user reviews, the sentiment polarity module 202 is used to divide the reviewed user reviews into short sentences to obtain a plurality of review short sentences, and can divide the multiple review short sentences into sentiment categories and calculate sentiment scores. Based on the sentiment categories of the multiple review short sentences, the comprehensive label extraction module 203 can extract comprehensive labels for the multiple review short sentences to obtain comprehensive labels for the multiple review short sentences. The comprehensive label construction and maintenance module 204 can extract label information of multiple dimensions from hotel reviews, analyze the dimensions of label extraction from multiple angles, and construct a comprehensive label library, which includes primary labels, secondary labels, and tertiary labels, as shown in Table 1. If the comprehensive labels subsequently extracted by the comprehensive label extraction module 203 do not exist in the three-level label system, they can be automatically added to the three-level labels of the comprehensive label library to update the comprehensive label library. Based on the obtained comprehensive labels and sentiment scores of the review sentences, the characteristic label extraction module 205 can determine characteristic reviews from the user reviews.

[0099] It should be noted that the user reviews obtained by the content audit module 201 can be obtained from the aforementioned inventory server, or can be obtained from other storage devices for screening user reviews, and the specific details are not limited here. Figure 2 It is only a schematic diagram of a system framework provided in an embodiment of the present application. The positional relationship between the devices, components, modules, etc. shown in the figure does not constitute any limitation. For example, the content review module 201 can be connected to the sentiment polarity module 202, or it can be connected to the comprehensive label extraction module 203.

[0100] Table 1

[0101]

[0102]

[0103] The following is an introduction to the method flow provided by this application in conjunction with the aforementioned system framework.

[0104] See Figure 3 , a flow chart of a label extraction method provided in this application is described as follows.

[0105] 301. Obtain user reviews;

[0106] The user reviews may be obtained by reading from an inventory server or a database, which is not limited here. The user reviews may include the user's evaluation of a service, product, or facility, which is not limited here.

[0107] 302. Divide the user reviews to obtain one or more first review sentences;

[0108] After obtaining user reviews, the user reviews may be divided into short sentences to obtain one or more first review short sentences.

[0109] Optionally, a content review model can be used to review the obtained user reviews and eliminate reviews containing banned words or spam reviews. Banned words include pornographic words, abusive words, and advertising words. Spam reviews include repeated reviews or reviews that are irrelevant to the review object. After obtaining the reviewed user reviews, text review can also be implemented through convolutional neural networks (CNN), long short-term memory networks (LSTM), or gated recurrent units (GRU), etc. to obtain the reviewed reviews. The specific details are not limited here.

[0110] In the embodiment of the present application, user reviews are screened through a content review model, which can remove reviews containing illegal or banned words and repeated and meaningless reviews, reduce the number of reviews for feature review extraction, and thus improve the efficiency of feature review mining.

[0111] Optionally, the reviewed user review can be divided according to delimiters to obtain one or more first review sentences. A delimiter is a symbol used to separate characters in the user review. A delimiter can be a comma, a period, a semicolon, an exclamation point, or a question mark, the specifics of which are not limited here. If a user review does not contain a delimiter, the review can be directly used as the first review sentence. If a user review contains one or more delimiters, multiple first review sentences can be obtained.

[0112] 303. Analyze the first review sentence according to the sentiment classification model to obtain the sentiment category of the first review sentence;

[0113] After obtaining one or more review sentences, analysis can be performed based on the sentiment classification model to obtain the sentiment category of the first review sentence. The sentiment category may include one or more of positive, negative, or neutral, which is not specifically limited here.

[0114] Optionally, the obtained first review sentence may be segmented to obtain one or more first words; and the obtained one or more first words may be input into a sentiment classification model to obtain a sentiment category of the first review sentence.

[0115] The sentiment classification model can be an improved FastText model, that is, a classification model that is improved on the FastText model. The FastText model uses a continuous bag of words (CBOW) structure, which includes an input layer, a hidden layer, and an output layer. Based on the FastText model, the input layer of the model is improved to obtain the improved FastText model.

[0116] For example, a chi-square test and a fusion of term frequency and inverse document frequency (TF-IDF) are used to extract keywords from text. The chi-square test calculates the relevance of each feature (word) to the classification result. A greater relevance helps the classifier perform classification. However, the chi-square test ignores the importance of word frequency. Therefore, the TF-IDF algorithm is combined. The TF-IDF algorithm evaluates the value of vocabulary for text classification by calculating the frequency of each word in the text and whether the word has the ability to distinguish between texts.

[0117] By combining the chi-square test with the TF-IDF algorithm to improve the model input layer, the improved FastText model can focus more on discriminative words during training, thereby improving the accuracy of model classification. In addition, when the limited number of words in the text leads to insufficient training data for the model, the TF-IDF algorithm can extract features from the text and represent the text as a word frequency vector, thereby alleviating the sparsity problem when processing text. It can also filter the noise in the text, reducing the impact of noise on model training and improving the generalization ability of the model.

[0118] The improved FastText model is compared with the original FastText model. The comparison results are as follows: Figure 4 、 Figure 5 As shown. Figure 4 It can be seen that the recall rate and precision of the improved FastText model are higher than those of the original FastText model. Figure 5 As can be seen, the training time of the improved FastText model is shorter than that of the original FastText model. This shows that the improved FastText model shortens the training time by reducing redundant vocabulary and introduces sentiment features into the features constructed based on the original words, which increases the model's feature expression capability, minimizes false positives, and thus improves the model's accuracy.

[0119] 304. Calculate the first comment sentence to obtain a sentiment score of the first comment sentence;

[0120] After obtaining one or more first review sentences, the sentiment score of the first review sentence can also be calculated to provide a screening basis for the subsequent extraction of featured reviews. The sentiment score represents the user's review attitude. The higher the sentiment score, the higher the user's degree of affirmation or liking for the review object.

[0121] Optionally, the sentiment word and associated words associated with the sentiment word in the first review sentence can be determined by querying the sentiment word library. The sentiment word represents the user's attitude or evaluation of the review object. For example, if the first review sentence is "I like this dress", the sentiment word of this sentence is "like". Another example is "The restaurant's dishes are relatively monotonous", and the sentiment word of this sentence is "monotonous". For example, if the first review sentence is "The hotel waiter is very responsible", the sentiment word "very" in this sentence is an associated word associated with the sentiment word. Subsequently, the sentiment score of the first review sentence can be calculated based on the obtained sentiment words and associated words.

[0122] The emotional polarity of the sentiment word can be used to determine the positive or negative score of the sentiment score of the word, thereby calculating the sentiment score of the review sentence. Sentiment words include positive and negative words. The weight of positive words is positive, while the weight of negative words is negative. For example, in the sentence "I like this hotel," where the sentiment word is "like," its lexical weight can be 1, so the sentiment score of the sentence is 1. For another example, in the sentence "I hate the hotel's swimming pool," the sentiment word identified as "hate" can be -2, so the sentiment score of the sentence is -2. The lexical weight can be 1, 2, or other values, which are not limited here and can be set according to actual conditions.

[0123] Associated words may include adverbs of degree or negative adverbs, where different adverbs of degree have different sentiment scores. For example, the sentiment score of "very" or "extremely" may be 2 points, the sentiment score of "extraordinarily" or "especially" may be 1.5 points, the sentiment score of "more" or "more and more" may be 1.25 points, and the sentiment score of "some" or "a little" may be 0.25 points. It can also be set according to actual conditions and is not limited here.

[0124] For short sentences, the sentiment score can also be calculated based on the number of negative adverbs preceding the sentiment word. If the number of negative adverbs is odd, the sentiment score for the sentence can be -1*score, which is the score obtained based on one or more of the sentiment words or related words. If the number of negative adverbs is even, the sentiment score for the sentence can be 1*score.

[0125] Taking a short user review as an example, the calculation of the sentiment score is as follows. For example, for the review sentence "The transportation to the amusement park is very inconvenient", the words obtained after word segmentation of this review sentence are "amusement park", "transportation", "very", "not", "convenient". According to the related word "very" and the negative adverb "not", the calculated sentiment score of this review sentence can be -1.5 points.

[0126] Among them, there is no necessary sequence between step 303 and step 304. Step 303 can be executed first, or step 304 can be executed first. Specifically, no limitation is made here.

[0127] 305. Analyze the first review sentence according to the sentiment category to obtain the comprehensive label of the first review sentence;

[0128] After obtaining the sentiment category, three methods can be used to extract the comprehensive label of the first review sentence, so as to obtain the comprehensive label of the first review sentence.

[0129] Optionally, the first review sentence can be preprocessed to obtain a second review sentence. The preprocessing can specifically include removing invalid symbols and emoticons, and can also convert traditional Chinese characters in the review into simplified Chinese characters. Specifically, no limitation is made here. For example, "真是開心啊--" can be converted into "真是开心啊" after preprocessing.

[0130] After obtaining the second review sentence, if the second review sentence conforms to the target syntactic structure, the dependency syntactic analysis method can be used to analyze the second review sentence first to obtain the comprehensive label. By calculating the similarity between the evaluation words in the second review sentence and the view subscripts in the entity knowledge base, and according to this similarity and the review object of the second review sentence, the comprehensive label is determined. Among them, the target syntactic structure includes one or more of the subject-predicate structure, the attributive-middle relationship or the verb-object relationship.

[0131] If the second review sentence does not conform to the target syntactic structure, a sentence pattern matching model can be used to analyze the second review sentence to obtain the comprehensive label. By calculating the similarity between the second review sentence and the target sentence pattern in the sentence pattern matching model, when the similarity is higher than the preset value, the target sentence pattern corresponding to the second review sentence is used as the comprehensive label. Among them, the target sentence pattern can include nearby landmarks, multiple stays, want to stay again, feel at home or the hotel is luxurious, etc., and can also be other target sentence patterns. Specifically, no limitation is made here.

[0132] If the second review sentence does not conform to the target syntactic structure and also does not conform to the target sentence pattern, then the second review sentence can be analyzed according to the keyword library to obtain the comprehensive label. By calculating the similarity between the words in the second review sentence and the keywords in the keyword library, the words in the second review sentence with a similarity higher than the preset value are used as the comprehensive label, and the sentiment category of this second review sentence is positive.

[0133] There is no necessary order between step 304 and step 305. Step 304 can be performed first, or step 305 can be performed first. The specific order is not limited here.

[0134] In the implementation manner of the present application, comprehensive label extraction can be performed on the second review sentence in three ways, and text similarity processing can be performed on the review object and review attitude in the review sentence, avoiding the extraction of only the review object and review attitude that appear in the label maintenance system, so that words with high similarity with those in the label maintenance system can also be extracted, thereby improving the success rate of comprehensive label mining.

[0135] 306. Obtain featured comments from user reviews based on sentiment scores and comprehensive tags.

[0136] After obtaining the sentiment score and comprehensive label, the featured review in the user review can be determined based on the priority and sentiment score of the second review object of the second review sentence. The featured review is a positive and high-quality review that best highlights the characteristics of the review object.

[0137] Optionally, you can first compare the priorities of the second comment objects of the second comment sentences. If the priorities of the second comment objects are different, you can sort the second comment objects according to their priorities and select the second comment sentences corresponding to the second comment objects with the highest priority as the featured comments. You can also select the second comment objects with the top three priorities based on the number of second comment objects and use their corresponding second comment sentences as the featured comments. The specific details are not limited here.

[0138] If the second review objects have the same priority, the sentiment scores of the second review sentences can be compared, and the featured review can be determined based on the sentiment score sorting. The sentiment scores can be sorted from high to low or from low to high. If the sorting method is high to low, the second review sentence with the highest sentiment score can be selected as the featured review, or the top three second review sentences can be selected as the featured review, or the top five second review sentences can be selected as the featured review. The specific decision can be made based on the number of candidate reviews and the actual situation, and is not limited here.

[0139] Among them, since the second review sentence is obtained by preprocessing the first review sentence and does not involve relevant processing of sentiment words and conjunctions, the sentiment score of the second review sentence is consistent with the sentiment score of the first review sentence.

[0140] In the embodiment of the present application, in the process of extracting comprehensive tags, the similarity of vocabulary is taken into consideration, text similarity processing is performed on user reviews, and user reviews are divided into sentiment categories, thereby avoiding the extraction of only the review subjects and sentiment categories maintained in the tag system, thereby improving the success rate of mining distinctive reviews, and introducing sentiment scores to judge user reviews, which can more objectively screen out distinctive reviews.

[0141] See Figure 6 , a flow chart of another tag extraction method provided in this application is described as follows.

[0142] 601. Obtain user reviews;

[0143] Among them, the user review can be the user's evaluation of the hotel, the user's evaluation of the travel service, or the user's evaluation of the purchased goods. The specific details are not limited here. The embodiment of this application only takes the user's evaluation of the hotel as an example for introduction.

[0144] 602. Divide the user reviews to obtain one or more first review sentences;

[0145] 603. Analyze the first review sentence according to the sentiment classification model to obtain the sentiment category of the first review sentence;

[0146] 604. Calculate the first comment sentence to obtain a sentiment score of the first comment sentence;

[0147] In the embodiment of the present application, steps 601 to 604 are the same as those described above. Figure 3 Steps 301 to 304 in the illustrated embodiment are similar and will not be described in detail here.

[0148] 605. Preprocess the first review sentence to obtain a second review sentence;

[0149] After obtaining the first review sentence, the first review sentence may be subjected to text processing. For example, invalid symbols in the first review sentence may be removed, and traditional Chinese characters in the first review sentence may be converted into simplified Chinese characters. The specific details are not limited here.

[0150] 606. Determine whether the second comment sentence conforms to the target syntactic structure;

[0151] After obtaining the second review sentence, it can be determined whether the second review sentence conforms to the target syntactic structure. The target syntactic structure can include one or more of a subject-verb structure (SBV), an attribute structure (ATT), or a verb-object structure (VOB). The subject-verb structure can extract a noun plus an adjective combination, which can be processed to extract label information such as "large room", "good breakfast", and "low cost performance"; the attribute relationship can extract an adjective plus noun combination, such as "beautiful scenery" and "warm service"; and the verb-object structure can extract a verb plus noun combination, such as "providing pick-up service" and "providing shuttle bus" and other label information. If the second review sentence conforms to the target syntactic structure, step 607 is executed; otherwise, step 608 is executed.

[0152] 607. Analyze the second review sentence using dependency parsing based on the sentiment category to obtain a comprehensive label for the second review sentence.

[0153] When the second comment sentence conforms to the aforementioned target syntactic structure, the second comment sentence can be analyzed using dependency syntactic analysis.

[0154] Optionally, after obtaining the second review sentence, the second review sentence can be segmented to obtain multiple second words; based on the multiple second words and the parts of speech of the multiple second words, the evaluation words in the second review sentence can be determined; based on the sentiment category of the second review sentence, the first similarity between the evaluation words and the opinion subscript in the entity knowledge base can be calculated; based on the second review object obtained for the second review sentence and the calculated first similarity, a comprehensive label for the second review sentence can be obtained.

[0155] The entity knowledge base includes multiple first keywords, first review objects corresponding to the first keywords, and opinion subscripts, as shown in Table 2. The opinion subscript is the evaluation of the first review object, and the opinion subscript can include positive opinion subscripts and negative opinion subscripts. The positive opinion subscript is a positive evaluation of the first review object, and the negative opinion subscript is a negative evaluation of the first review object, as shown in Table 3.

[0156] Optionally, the second review object of the second review sentence can be determined based on the part of speech of the second vocabulary and the entity knowledge base. Specifically, the review subject of the second review sentence can be determined based on the part of speech of the second vocabulary; and the similarity between the review subject and the first keyword in the entity knowledge base can be calculated to determine the second review object.

[0157] Among them, if the similarity between the comment subject and the first keyword is higher than or equal to the preset value, the comment object corresponding to the first keyword can be used as the second comment object; if the similarity between the comment subject and the first keyword is lower than the preset value, the comment subject can be used as the second comment object, and based on the frequency of appearance of the comment subject, when the frequency is more than the preset value, the comment subject can be added to the comment object list in the entity knowledge base, thereby updating the entity knowledge base.

[0158] Optionally, when the sentiment category of the second review sentence is positive, a first similarity between the review word and the positive opinion subscript can be calculated. If the first similarity is greater than or equal to a first preset value, the second review object and the positive opinion subscript can be used as a comprehensive label; otherwise, the second review object and the review word can be used as a comprehensive label.

[0159] Table 2

[0160]

[0161] Table 3

[0162]

[0163] 608. Analyze the second review sentence using a sentence pattern matching model to determine whether the second review sentence conforms to the target sentence pattern;

[0164] If the second review sentence does not conform to the target syntactic structure, a sentence matching model can be used to analyze the second review sentence and determine whether it conforms to the target sentence structure based on the analysis results. The sentence matching model can be a doc2vec model or other text classification model, the specific structure of which is not limited here. The target sentence structure can include one or more of the following: nearby landmarks, multiple stays, would like to stay again, feel like home, or luxurious hotel.

[0165] Among them, the second review sentence is analyzed, and a sentence matching model can be used to calculate the second similarity between the second review sentence and the target sentence. If the second similarity is higher than or equal to a second preset value, that is, the second review sentence conforms to the target sentence, then step 609 is executed; otherwise, step 610 is executed.

[0166] 609. Use the target sentence pattern as the comprehensive label for the second review sentence;

[0167] When the second similarity is greater than or equal to a second preset value, the target sentence pattern corresponding to the second review sentence can be directly used as the comprehensive label of the second review sentence. For example, if the second review sentence is "I'll come again next time" and the similarity between the review sentence and the target sentence pattern "I want to stay again" is greater than a preset value, "I want to stay again" can be used as the comprehensive label of the review sentence.

[0168] 610. Analyze the second review sentence based on the keyword library to obtain a comprehensive label for the second review sentence;

[0169] When the second similarity is lower than the second preset value, the second review sentence can be analyzed according to the keyword library, where the keyword library includes multiple second keywords, and the second keywords include one or more of parent-child, family, business trip, or couple.

[0170] Specifically, the third similarity between the second vocabulary obtained and the second keyword in the keyword library is calculated. If the third similarity is higher than or equal to a third preset value, and the sentiment category of the second review sentence is positive, the second keyword can be used as a comprehensive label.

[0171] 611. Based on the sentiment score and comprehensive tags, obtain the featured comments from the user reviews.

[0172] After obtaining the sentiment score and comprehensive label, the comprehensive label can also be classified to determine the priority of the second review object, and the featured review in the user review can be determined based on the priority and sentiment score of the second review object.

[0173] Among them, taking into account the user's concerns and the nature of the product, in the embodiment of this method, for hotel user reviews, the priority of the second review object can be sorted in the following order: nearby landmarks>service provision>surroundings, shopping, catering>taxi, travel>hotel facilities>landscape>breakfast, food taste, meals>environment, hygiene>multiple stays, want to stay again, feel at home, luxurious hotel, suitable for the crowd>others.

[0174] Optionally, before determining the featured review based on the priority of the second review object and the sentiment score of the second review sentence, the sentiment score of the user review can be calculated based on the sentiment scores of multiple second review sentences, and the user reviews ranked in the top ten by sentiment score can be selected as candidate reviews for the featured review, and the user reviews ranked in the top twenty by sentiment score can also be selected as candidate reviews for the featured review. The specific details are not limited here, and then the featured review is selected from the candidate reviews.

[0175] After selecting the candidate reviews, the priorities of the second review objects in the second review sentences of each candidate review can be compared. If the priorities of the second review objects are different, the second review sentence corresponding to the second review object with the highest priority can be selected as the candidate featured review. If the priorities of the second review objects are the same, the sentiment scores of the second review sentences are compared. Based on the sentiment score ranking, the second review sentence with the highest sentiment score can be selected as the candidate featured review. In this way, a review sentence from each candidate review is determined as a candidate featured review.

[0176] After obtaining multiple candidate featured reviews, the multiple candidate featured reviews can be screened and candidate featured reviews with a character length of less than or equal to 15 characters can be selected as featured reviews, thereby obtaining featured reviews in user reviews. The featured reviews can be used to display on the hotel details page to help users quickly understand the hotel's featured information, and can also be used to display on the hotel selection list page to help users make hotel reservation choices.

[0177] For example, the user review with the highest sentiment score for a hotel is as follows: "The stay experience was very, very good! The front desk lady was very nice and the service was excellent! In particular, Miss Wang Xue at the front desk was very considerate and enthusiastic, which made me very happy! She helped with things like extending check-out and ordering meals. The stay experience was excellent and made me feel right at home. Thank you again, Miss Wang Xue!" After analyzing this user review, the analysis results will be expressed in the form of a four-tuple "sentiment score, comment phrase, comment object, and comprehensive label". The analysis results of this user review can be "4. Very, very good stay experience, cost-effectiveness, and cost-effectiveness"; "1. Excellent stay experience, cost-effectiveness, and cost-effectiveness"; "1. Makes you feel like you're at home, suitable for everyone, and makes you feel at home."

[0178] Among them, the review object is the first priority, and the emotional score is the second priority. First, the priorities of the review objects are compared. The priority of being suitable for the crowd is higher than the cost-effectiveness, so it can obtain the characteristic label "making people feel at home".

[0179] For ease of understanding, the following exemplifies the effect of a tag extraction method provided in this application by taking the analysis of actual hotel reviews as an example.

[0180] Among them, comprehensive labels are extracted from hotel reviews. The comprehensive labels are concentrated on secondary labels such as stay experience, surrounding evaluation, environment atmosphere and catering evaluation. The distribution of secondary labels in the comprehensive labels is as follows: Figure 7 The third-level labels under some of the second-level labels are shown in Table 4, and some of the extracted featured reviews are shown in Table 5. 100 hotel reviews were randomly selected and manually labeled with sentiment categories as comparison data. These 100 reviews were analyzed using the sentiment classification model, and the macro-F1 value was 0.95. The closer the macro-F1 value is to 1, the better the model performance is, indicating that the sentiment classification model used in the embodiment of this application has a high accuracy rate.

[0181] Table 4

[0182]

[0183]

[0184] Table 5

[0185]

[0186]

[0187] The above describes the method flow provided by the present application. Based on the above method flow, the following describes the device provided by the present application.

[0188] See Figure 8 , the present application provides a structural diagram of a tag extraction device, including:

[0189] Analysis module 801, used to classify user comments according to the sentiment classification model to obtain the sentiment category of the user comments;

[0190] A calculation module 802 is used to calculate the sentiment score of the user review based on the sentiment words and the associated words with the sentiment words in the user review;

[0191] The analysis module 801 is further configured to analyze user comments based on sentiment categories to obtain comprehensive labels for the user comments. The comprehensive labels are used to summarize characteristic information of the user comments.

[0192] The processing module 803 is used to obtain the featured comments in the user comments according to the sentiment score and the comprehensive label.

[0193] In one possible implementation, the above-mentioned analysis module 801 is specifically used to divide user comments to obtain one or more first review sentences, where the user comments are obtained after being reviewed by a content review model; and one or more first review sentences are used as input to a sentiment classification model to obtain sentiment categories, where the sentiment categories include one or more of positive, negative or neutral.

[0194] In a possible implementation, the calculation module 802 is specifically configured to determine sentiment words and associated words by querying a sentiment word library, where the sentiment word library includes multiple sentiment words and multiple associated words; and calculate the sentiment score of the first review sentence based on the sentiment words and associated words.

[0195] In one possible implementation, the above-mentioned analysis module 801 is specifically used to preprocess the first comment sentence to obtain a second comment sentence; if the second comment sentence conforms to the target syntactic structure, the second comment sentence is analyzed using dependency syntactic analysis according to the sentiment category to obtain a comprehensive label, and the target syntactic structure includes one or more of the subject-predicate structure, the attributive-predicate relationship, or the verb-object relationship; if the second comment sentence does not conform to the target syntactic structure, the second comment sentence is analyzed using a sentence matching model to obtain a comprehensive label; if the second comment sentence does not conform to the target syntactic structure and does not conform to the target sentence pattern, the second comment sentence is analyzed according to the keyword library to obtain a comprehensive label.

[0196] In a possible implementation, the device further includes: a word segmentation module 804, configured to segment the second review sentence to obtain a plurality of words.

[0197] In a possible implementation, the above-mentioned analysis module 801 is specifically used to determine the evaluation vocabulary of the second review sentence based on the vocabulary; calculate the first similarity between the evaluation vocabulary and the opinion subscript according to the sentiment category of the second review sentence, the opinion subscript includes the evaluation of the first review object, the opinion subscript includes a positive opinion subscript and a negative opinion subscript, the positive opinion subscript is a positive evaluation of the review object, and the negative opinion subscript is a negative evaluation of the review object; obtain the second review object of the second review sentence; and obtain a comprehensive label based on the first similarity and the second review object.

[0198] In a possible implementation, the analysis module 801 is specifically configured to calculate a first similarity between the evaluation vocabulary and the positive opinion subscript if the sentiment category is positive.

[0199] In a possible implementation, the above-mentioned analysis module 801 is specifically used to obtain the part of speech of the vocabulary; determine the second comment object of the second comment sentence based on the part of speech and the entity knowledge base, and the entity knowledge base includes the first comment object, the first keyword and the opinion subscript.

[0200] In a possible implementation, the analysis module 801 is specifically configured to determine the review subject of the second review sentence based on the part of speech; and determine the second review object based on the similarity between the review subject and the first keyword in the entity knowledge base.

[0201] In a possible implementation, the above-mentioned analysis module 801 is specifically used to use the second comment object and the positive opinion subscript as a comprehensive label if the first similarity is higher than or equal to the first preset value; if the first similarity is lower than the first preset value, use the second comment object and the evaluation vocabulary as a comprehensive label.

[0202] In one possible implementation, the analysis module 801 is specifically configured to vectorize the second comment sentence to obtain a processed third comment sentence; calculate a second similarity between the third comment sentence and a model file, where the model file is a vectorized file of the target sentence; and obtain a comprehensive label based on the second similarity.

[0203] In a possible implementation, the analysis module 801 is specifically configured to use the corresponding target sentence as a comprehensive label if the second similarity is higher than or equal to a second preset value.

[0204] In a possible implementation, the above-mentioned analysis module 801 is specifically used to calculate the third similarity between the vocabulary and the second keyword in the keyword library; if the third similarity is higher than or equal to a third preset value, and the sentiment category of the second review sentence is positive, the second keyword is used as a comprehensive label.

[0205] In a possible implementation, the processing module 803 is specifically configured to determine the priority of the second review object based on the comprehensive tag; and determine a featured review in the user review based on the priority and sentiment score of the second review object.

[0206] In one possible implementation, the processing module 803 is specifically used to compare sentiment scores if the priorities of the second review objects are the same, and determine the featured reviews in the user reviews based on the sorting of sentiment scores; if the priorities of the second review objects are different, determine the featured reviews in the user reviews based on the sorting of the priorities of the second review objects.

[0207] See Figure 9 , this application provides a structural diagram of another tag extraction device, including:

[0208] The tag extraction device may include a processor 901 and a memory 902. The processor 901 and the memory 902 are interconnected via a circuit. The memory 902 stores program instructions and data.

[0209] The memory 902 stores the aforementioned Figure 3 and Figure 6 The program instructions and data corresponding to the steps in .

[0210] Processor 901 is used to execute the above Figure 3 and Figure 6 The method steps performed by the label extraction device shown in any embodiment.

[0211] Optionally, the tag extraction device may further include a transceiver 903 for receiving or sending data.

[0212] The present application also provides a computer-readable storage medium in which a program is stored. When the program is run on a computer, the computer executes the above-mentioned Figure 3 and Figure 6 The illustrated embodiments describe steps in a method.

[0213] Optionally, the aforementioned Figure 9 The label extraction device shown in is a chip.

[0214] The embodiment of the present application also provides a label extraction device, which can also be called a digital processing chip or chip. The chip includes a processing unit and a communication interface. The processing unit obtains program instructions through the communication interface. The program instructions are executed by the processing unit. The processing unit is used to execute the aforementioned Figure 3 and Figure 6 The method steps performed by the label extraction device shown in any embodiment.

[0215] The present application also provides a digital processing chip. This digital processing chip integrates circuitry and one or more interfaces for implementing the functions of the processor 901 described above. If the digital processing chip incorporates memory, it can perform the method steps of any one or more of the aforementioned embodiments. If the digital processing chip does not incorporate memory, it can connect to an external memory via a communication interface. The digital processing chip implements the actions performed by the label extraction device in the aforementioned embodiments based on program code stored in the external memory.

[0216] The present application also provides a computer program product which, when executed on a computer, enables the computer to execute the aforementioned Figure 3 or Figure 6 The illustrated embodiments describe method steps.

[0217] The label extraction device provided in the embodiment of the present application can be a chip, which includes: a processing unit and a communication unit. The processing unit can be, for example, a processor, and the communication unit can be, for example, an input / output interface, a pin or a circuit. The processing unit can execute the computer execution instructions stored in the storage unit to enable the chip in the server to execute the above Figure 3 or Figure 6 The method described in the embodiment shown. Optionally, the storage unit is a storage unit within the chip, such as a register, a cache, etc. The storage unit may also be a storage unit located outside the chip within the wireless access device, such as a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM), etc.

[0218] Specifically, the aforementioned processing unit or processor may be a central processing unit (CPU), a neural-network processing unit (NPU), a graphics processing unit (GPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.

[0219] For example, see Figure 10 , Figure 10 This is a schematic diagram of the structure of a chip provided in an embodiment of the present application. The chip can be represented as a neural network processor NPU 1000. NPU 1000 is mounted on the host CPU (HostCPU) as a coprocessor and is assigned tasks by the Host CPU. The core of the NPU is arithmetic circuit 1003, which is controlled by controller 1004 to extract matrix data from memory and perform multiplication operations.

[0220] In some implementations, arithmetic circuit 1003 includes multiple processing engines (PEs). In some implementations, arithmetic circuit 1003 is a two-dimensional systolic array. Arithmetic circuit 1003 can also be a one-dimensional systolic array or other electronic circuitry capable of performing mathematical operations such as multiplication and addition. In some implementations, arithmetic circuit 1003 is a general-purpose matrix processor.

[0221] For example, assume there are input matrix A, weight matrix B, and output matrix C. The arithmetic circuit retrieves the corresponding data of matrix B from weight memory 1002 and caches it on each PE in the arithmetic circuit. The arithmetic circuit retrieves the data of matrix A from input memory 1001 and performs a matrix operation on matrix B. The partial or final matrix result is stored in accumulator 1008.

[0222] Unified memory 1006 is used to store input and output data. Weight data is directly transferred to weight memory 1002 through direct memory access controller (DMAC) 1005. Input data is also transferred to unified memory 1006 through DMAC.

[0223] The bus interface unit (BIU) 1010 is used for interaction between the AXI bus, the DMAC, and the instruction fetch buffer (IFB) 1009 .

[0224] The bus interface unit 1010 (BIU) is used for the instruction fetch memory 1009 to obtain instructions from the external memory, and is also used for the storage unit access controller 1005 to obtain the original data of the input matrix A or the weight matrix B from the external memory.

[0225] DMAC is mainly used to transfer input data in the external memory DDR to the unified memory 1006 or transfer weight data to the weight memory 1002 or transfer input data to the input memory 1001.

[0226] The vector calculation unit 1007 includes multiple operation processing units. When necessary, it further processes the output of the operation circuit, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc. It is mainly used for non-convolutional / fully connected layer network calculations in neural networks, such as batch normalization, pixel-level summation, and upsampling of feature planes.

[0227] In some implementations, the vector calculation unit 1007 can store the processed output vector to the unified memory 1006. For example, the vector calculation unit 1007 can apply a linear function and / or a nonlinear function to the output of the operation circuit 1003, such as linear interpolation of the feature plane extracted by the convolution layer, or accumulate a vector of values to generate an activation value. In some implementations, the vector calculation unit 1007 generates a normalized value, a pixel-level summed value, or both. In some implementations, the processed output vector can be used as an activation input to the operation circuit 1003, for example, for use in a subsequent layer in a neural network.

[0228] An instruction fetch buffer 1009 connected to the controller 1004 is used to store instructions used by the controller 1004;

[0229] Unified memory 1006, input memory 1001, weight memory 1002, and instruction fetch memory 1009 are all on-chip memories. External memories are private to the NPU hardware architecture.

[0230] The processor mentioned in any of the above can be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits for controlling the above Figure 3 、 Figure 4 or Figure 5 Method of procedure.

[0231] It should also be noted that the device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the device embodiments provided in this application, the connection relationship between the modules indicates that there is a communication connection between them, which can be specifically implemented as one or more communication buses or signal lines.

[0232] Through the description of the above implementation methods, technical personnel in the relevant field can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0233] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0234] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0235] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0236] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, read-only memory), random access memory (RAM, random access memory), disk or optical disk, and other media that can store program code.

[0237] The terms "first," "second," "third," "fourth," and the like (if any) in the specification and claims of this application and in the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or sequential sequence. It should be understood that the terms used in this manner are interchangeable where appropriate so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" and "having," and any variations thereof, are intended to cover non-exclusive inclusions, e.g., a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0238] Finally, it should be noted that the above is only a specific implementation method of the present application, but the protection scope of the present application is not limited to this. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed in this application, which should be covered by the protection scope of the present application.

Claims

1. A tag extraction method, characterized in that: include: Classify user comments according to the sentiment classification model to obtain the sentiment category of the user comments; Calculate the sentiment score of the user review based on the sentiment words in the user review and the associated words of the sentiment words; Analyzing the user comments according to the sentiment category to obtain a comprehensive label of the user comments, wherein the comprehensive label is used to summarize characteristic information of the user comments; According to the sentiment score and the comprehensive label, characteristic comments in the user comments are obtained.

2. The method according to claim 1, characterized in that The user comments are classified according to the sentiment classification model to obtain the sentiment categories of the user comments, including: Dividing the user reviews to obtain one or more first review sentences, wherein the user reviews are obtained after being reviewed by a content review model; The one or more first review sentences are used as inputs of the sentiment classification model to obtain the sentiment category, where the sentiment category includes one or more of positive, negative, or neutral.

3. The method according to claim 2, characterized in that The step of calculating the sentiment score of the user review based on the sentiment words in the user review and the associated words with the sentiment words includes: Determine the emotional words and the associated words by querying an emotional word library, wherein the emotional word library includes a plurality of emotional words and a plurality of associated words; Calculate the sentiment score of the first review sentence based on the sentiment words and the associated words.

4. The method according to any one of claims 2 or 3, characterized in that Analyzing the user comments according to the emotion category to obtain a comprehensive label of the user comments includes: Preprocessing the first review sentence to obtain a second review sentence; If the second review sentence conforms to the target syntactic structure, then according to the sentiment category, the second review sentence is analyzed using dependency syntactic analysis to obtain the comprehensive label, wherein the target syntactic structure includes one or more of a subject-predicate structure, an attributive-predicate relationship, or a verb-object relationship; If the second review sentence does not conform to the target syntactic structure, analyzing the second review sentence using a sentence matching model to obtain the comprehensive label; If the second review sentence does not conform to the target syntactic structure and does not conform to the target sentence pattern, the second review sentence is analyzed according to the keyword library to obtain the comprehensive label.

5. The method according to claim 4, characterized in that The method further comprises: The second review sentence is segmented to obtain a plurality of words.

6. The method according to claim 5, characterized in that The step of analyzing the second review sentence using dependency parsing according to the sentiment category to obtain the comprehensive label includes: Determining, based on the vocabulary, an evaluation vocabulary for the second review sentence; Calculating, based on the sentiment category of the second review sentence, a first similarity between the review word and an opinion subscript, wherein the opinion subscript includes an evaluation of the first review object, and the opinion subscript includes a positive opinion subscript and a negative opinion subscript, wherein the positive opinion subscript represents a positive evaluation of the review object, and the negative opinion subscript represents a negative evaluation of the review object; Obtain the second review object of the second review sentence; The comprehensive label is obtained according to the first similarity and the second comment object.

7. The method according to claim 6, characterized in that Calculating the first similarity between the evaluation vocabulary and the opinion subscript according to the sentiment category of the second review sentence includes: When the sentiment category of the second review sentence is positive, a first similarity between the evaluation vocabulary and the positive opinion subscript is calculated.

8. The method according to any one of claims 6 or 7, characterized in that The step of obtaining the second review object of the second review sentence includes: Obtaining the part of speech of the vocabulary; A second review object of the second review sentence is determined according to the part of speech and an entity knowledge base, wherein the entity knowledge base includes the first review object, the first keyword, and the viewpoint subscript.

9. The method according to claim 8, characterized in that The determining, based on the part of speech and the entity knowledge base, a second review object of the second review sentence includes: determining the review subject of the second review sentence according to the part of speech; The second comment object is determined according to the similarity between the comment subject and the first keyword in the entity knowledge base.

10. The method according to any one of claims 6 to 9, characterized in that Obtaining the comprehensive label according to the first similarity and the second comment object includes: If the first similarity is higher than or equal to a first preset value, the second comment object and the positive opinion subscript are used as a comprehensive label; If the first similarity is lower than the first preset value, the second comment object and the evaluation vocabulary are used as a comprehensive label.

11. The method according to any one of claims 4 to 10, characterized in that The sentence matching model is used to analyze the second review sentence to obtain the comprehensive label, including: Vectorize the second review sentence to obtain a processed third review sentence; Calculating a second similarity between the third review sentence and a model file, where the model file is a vectorized file of the target sentence; The comprehensive label is obtained according to the second similarity.

12. The method according to claim 11, characterized in that Obtaining the comprehensive label according to the second similarity includes: If the second similarity is higher than or equal to a second preset value, the corresponding target sentence pattern is used as a comprehensive label.

13. The method according to claim 5, characterized in that The second review sentence is analyzed based on the keyword library to obtain the comprehensive label, including: calculating a third similarity between the vocabulary and a second keyword in the keyword library; If the third similarity is higher than or equal to a third preset value, and the sentiment category of the second review sentence is positive, the second keyword is used as the comprehensive label.

14. The method according to any one of claims 1 to 13, characterized in that Obtaining the featured comments from the user comments based on the sentiment score and the comprehensive label includes: Determining the priority of the second comment object according to the comprehensive tag; Determine a featured review among the user reviews according to the priority of the second review object and the sentiment score.

15. The method according to claim 14, characterized in that Determining the featured comments in the user comments based on the second comment object and the priority of the sentiment score includes: If the priorities of the second review objects are the same, comparing the sentiment scores, and determining the featured reviews among the user reviews based on the ranking of the sentiment scores; If the priorities of the second review objects are different, the featured reviews in the user reviews are determined according to the priority ranking of the second review objects.

16. A label extraction device, characterized in that: include: An analysis module is used to classify user comments according to a sentiment classification model to obtain the sentiment category of the user comments; A calculation module, configured to calculate the sentiment score of the user review based on the sentiment words in the user review and the associated words with the sentiment words; The analysis module is further configured to analyze the user comments according to the sentiment category to obtain a comprehensive label for the user comments, wherein the comprehensive label is used to summarize characteristic information of the user comments; A processing module is used to obtain characteristic comments from the user comments based on the sentiment score and the comprehensive label.

17. A label extraction device, characterized in that: include: a processor and a memory, the processor being coupled to the memory; The memory is used to store programs; The processor is configured to execute the program in the memory so as to perform the method according to any one of claims 1 to 15.

18. A computer-readable storage medium comprising instructions, which, when executed on a computer, enable the computer to perform the method according to any one of claims 1 to 15.

19. A computer program product comprising instructions which, when run on a computer, cause the computer to perform the method according to any one of claims 1 to 15.

Citation Information

Patent Citations

  • Hotel feature comment extraction method

    CN107122471A

  • Comment analysis method based on word vectors and syntactic features and visual interactive interface

    CN110175325A

  • Movie comment viewpoint emotion tendency analysis method

    CN110825876A