Method, System, and Storage Medium for Identifying User Evaluation Triples Based on Large Language Models

Through an automated process based on large models, user comment data is collected from the Internet, data screening, group segmentation and sentiment analysis are carried out, and evaluation dimension framework is constructed in combination with phrase clustering models, which solves the problems of low efficiency and low accuracy of user evaluation triple recognition in the existing technology, and realizes efficient and accurate evaluation triple recognition and flexible dimension framework design.

CN119397015BActive Publication Date: 2025-05-27GUANGZHOU DATASTORY INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411308982.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-19
Publication Date
2025-05-27
Estimated Expiration
2044-09-19

AI Technical Summary

Technical Problem

In the prior art, the user evaluation triple recognition method has problems such as incomplete manual recognition and low efficiency, as well as high cost, poor flexibility and low accuracy for intelligent model recognition.

Method used

The user evaluation triple recognition method based on the big model is adopted. By collecting user comment data from the Internet, pre-trained large language model is used for data screening, group segmentation, sentiment analysis and dimension extraction, combined with the phrase clustering model for preliminary and secondary clustering, the final evaluation dimension framework is constructed to realize fully automated evaluation triple recognition.

Benefits of technology

It significantly improves development efficiency, reduces labor costs, improves the flexibility and adaptability of identification accuracy and dimension framework, and can quickly adjust the evaluation dimension to adapt to the needs of different fields and application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119397015B_ABST
    Figure CN119397015B_ABST
Patent Text Reader

Abstract

The present invention provides a method, system, and storage medium for identifying user evaluation triples based on large models. The method includes: using a first large language model to perform quality screening on the collected comment data, using a second large language model to perform sense group segmentation on the screened high-quality comments, performing sentiment analysis and evaluation word extraction on each sense group segment, and summarizing all evaluation words in combination with semantic information to obtain evaluation dimensions; using a phrase clustering model to perform preliminary clustering on the evaluation dimensions, cleaning and naming normalization on the preliminary clustering results; performing secondary clustering and normalization on the redefined evaluation dimensions; reclassifying the outlier dimensions removed during the two clustering processes and incorporating them into the evaluation dimensions after secondary clustering and normalization; finally, jointly saving the evaluation dimensions, their corresponding core evaluation words, and sentiment tendencies as comment triples. The present invention can improve the efficiency and accuracy of identifying evaluation triples while reducing the dependence on manual intervention.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and more specifically, to a method, system, and storage medium for identifying user evaluation triples based on a large model. Background Art

[0002] In the current information age, user reviews have become an important data source for evaluating the quality of products and services. These reviews contain rich user feedback information, which is of great value for enterprises to conduct market analysis, improve product quality, and enhance customer satisfaction. In user review analysis, evaluation triples, as an effective representation, usually include three main elements: evaluation dimension, evaluation word, and sentiment tendency. The evaluation dimension is used to define the theme or aspect of the review, the evaluation word specifically describes the evaluation situation of this dimension, and the sentiment tendency indicates the emotional attitude of the review.

[0003] Currently, the identification of evaluation triples mainly relies on the following two schemes:

[0004] 1) Manual dimension code table: This scheme relies on a manually compiled list of evaluation dimensions and evaluation words. By manually collecting and organizing evaluation vocabulary and using rules such as regular expressions for matching, corresponding dimension labels are assigned to the review data. The dimension framework of this scheme can be adjusted and updated at any time according to needs, with high flexibility. However, it requires a large amount of manpower to collect evaluation words and write matching logic, resulting in low efficiency. Moreover, as the number of evaluation dimensions increases, the maintenance and update costs will also increase significantly. In addition, due to the limitations of manual coding, the evaluation dimensions may not be comprehensive enough, leading to low accuracy and coverage of the identification results.

[0005] 2) Intelligent triple model: This scheme usually relies on deep learning technologies, especially pre-trained models such as BERT. By annotating a large amount of training data for model training, the automatic identification of evaluation triples is achieved. Although the trained model can automatically identify evaluation triples, reducing manual intervention and providing higher identification accuracy than manual methods, it requires a large amount of labeled data for model training, resulting in high data annotation costs and long development cycles. In addition, once the evaluation dimension framework is determined, it is difficult to modify and expand it, with relatively low flexibility.

[0006] In addition, the expressions of user reviews are very rich, and new content emerges every once in a while. The existing manual dimension code table is very prone to mis-matching and recalling irrelevant data. Although the intelligent triple model can improve the identification accuracy, it has insufficient generalization ability for new user expressions. Summary of the Invention

[0007] To overcome the deficiencies of incomplete and inefficient manual recognition, as well as high cost, poor flexibility, and low accuracy in intelligent model recognition in the above-mentioned prior art, the present invention provides a method, system, and storage medium for identifying user evaluation triples based on a large model. Based on a large language model, it can achieve fully automated identification of evaluation triples, reduce dependence on manual labor, improve recognition efficiency and accuracy, and at the same time enhance the flexibility and adaptability of dimension framework modification.

[0008] To solve the above technical problems, the technical solution of the present invention is as follows:

[0009] A method for identifying user evaluation triples based on a large model, comprising the following steps:

[0010] S1: Collect user review data from multiple platforms on the Internet and construct a review data set;

[0011] S2: Set quality screening conditions for user review data, and use a pre-trained first large language model to screen the review data set to obtain a high-quality review data set;

[0012] S3: Use a pre-trained second large language model to perform sense group segmentation on the high-quality review data set, and respectively segment each high-quality user review data into several sense group segments;

[0013] S4: Use a pre-trained second large language model to perform sentiment analysis on each sense group segment to obtain the sentiment tendency of each sense group segment; and extract the core evaluation words from each sense group segment, and summarize the extracted core evaluation words in combination with the semantic information of the corresponding sense group segment to refine specific evaluation dimensions;

[0014] S5: Use a pre-trained phrase clustering model to perform preliminary clustering on all the refined evaluation dimensions according to similarity to obtain a preliminary clustering result;

[0015] S6: Use a pre-trained second large language model to clean the preliminary clustering result, remove outlier dimensions that do not meet the preset conditions, and redefine the evaluation dimensions for the cleaned preliminary clustering result to normalize the naming of all evaluation dimensions to obtain a redefined preliminary clustering result;

[0016] S7: Repeat steps S5 - S6 to perform secondary clustering and naming normalization on the redefined preliminary clustering result to obtain a redefined secondary clustering result;

[0017] S8: Reclassify the outlier dimensions removed during the two clustering processes and incorporate them into the redefined secondary clustering result to construct a final evaluation dimension framework;

[0018] Match the core evaluation words and sentiment tendencies of each semantic group segment with the final evaluation dimension framework to obtain the evaluation triple of each semantic group segment, and complete the identification of the user evaluation triple.

[0019] Preferably, in the step S2, the set quality screening conditions for user comment data include any one of comment length, grammatical integrity, comment paragraph density, and object richness included.

[0020] Preferably, in the step S3, the semantic group segment is specifically a segment expressing a specific evaluation dimension in high-quality user comment data, and each semantic group segment includes at least one evaluation word;

[0021] Use the semantic structure analysis function of the second large language model to perform segmentation of semantic group segments.

[0022] Preferably, in the step S4, the sentiment tendency includes any one of positive, neutral, and negative;

[0023] The core evaluation words include: evaluation attribute words and evaluation description words;

[0024] The evaluation attribute words are used to reflect the main evaluation dimensions of the comments; the evaluation description words are used to describe the characteristics or states of the evaluation dimensions;

[0025] Use the sentiment classification function of the second large language model to perform sentiment analysis; use the vocabulary extraction function of the second large language model to extract core evaluation words; use the semantic understanding function of the second large language model to obtain the semantic information of each semantic group segment and perform refinement of evaluation dimensions.

[0026] Preferably, in the step S6, the outlier dimensions that do not meet the preset conditions include: evaluation dimensions that cannot be matched with all core evaluation words, and evaluation dimensions that do not meet the preset definition specifications.

[0027] Preferably, the number of parameters of the first large language model is less than that of the second large language model.

[0028] Preferably, the first large language model is specifically the mini version after quantization of the SocialGPT large language model;

[0029] The second large language model is specifically the SocialGPT large language model;

[0030] The phrase clustering model is specifically the BGE-M3 semantic vector model.

[0031] Preferably, the number of parameters of the first large language model is specifically 0.2B; the number of parameters of the second large language model is specifically 32B.

[0032] The present invention also provides a user evaluation triple recognition system based on a large model, which applies the above-mentioned user evaluation triple recognition method based on a large model, and includes:

[0033] A comment collection unit: used to collect user comment data from multiple platforms on the Internet and construct a comment data set;

[0034] A comment screening unit: used to set quality screening conditions for user comment data, and use a pre-trained first large language model to screen the comment data set to obtain a high-quality comment data set;

[0035] A comment segmentation unit: used to use a pre-trained second large language model to perform sense segmentation on the high-quality comment data set, and respectively segment each high-quality user comment data into several sense segments;

[0036] An emotion analysis and dimension extraction unit: used to use a pre-trained second large language model to perform emotion analysis on each sense segment to obtain the emotion tendency of each sense segment; and extract core evaluation words from each sense segment, and combine the semantic information of the corresponding sense segment to summarize the extracted core evaluation words, and refine specific evaluation dimensions;

[0037] A clustering unit: used to use a pre-trained phrase clustering model to perform preliminary clustering on all the refined evaluation dimensions according to similarity to obtain a preliminary clustering result;

[0038] A naming normalization unit: used to use a pre-trained second large language model to clean the preliminary clustering result, eliminate outlier dimensions that do not meet the preset conditions, and redefine the evaluation dimensions for the cleaned preliminary clustering result to normalize the naming of all evaluation dimensions to obtain a redefined preliminary clustering result;

[0039] A secondary normalization unit: used to repeat the execution of the clustering unit and the naming normalization unit to perform secondary clustering and naming normalization on the redefined preliminary clustering result to obtain a redefined secondary clustering result;

[0040] An evaluation triple recognition unit: used to reclassify the outlier dimensions eliminated in the two clustering processes and incorporate them into the redefined secondary clustering result to construct a final evaluation dimension framework;

[0041] Match the core evaluation words and emotion tendencies of each sense segment with the final evaluation dimension framework to obtain the evaluation triples of each sense segment, and complete the recognition of user evaluation triples.

[0042] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps in the above method are implemented.

[0043] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:

[0044] The present invention provides a method, system and storage medium for identifying user evaluation triples based on a large model. The solution of the present invention adopts a three-stage process design, namely data screening, dimension generation and dimension normalization. After collecting the comment data, first use the first large language model to screen the data quality; then use the second large language model to perform sense group segmentation on the screened data, perform sentiment analysis on each sense group segment, extract the core evaluation words from each sense group segment, and combine the semantic information of the sense group segment to summarize the extracted evaluation words, and refine the specific evaluation dimensions; then use the phrase clustering model to perform preliminary clustering on the evaluation dimensions, and clean the preliminary clustering results; after cleaning, redefine the evaluation dimensions, and then perform secondary clustering and normalization on the redefined evaluation dimensions; reclassify the outlier dimensions removed during the two clustering processes and incorporate them into the evaluation dimensions after secondary clustering and normalization; finally, jointly save the evaluation dimensions, their corresponding core evaluation words and sentiment tendencies as comment triples;

[0045] The present invention has the following advantages:

[0046] 1) Significantly improve development efficiency: The present invention can completely automate the identification of evaluation triples, including dimension generation, evaluation word extraction and sentiment analysis, reducing manual intervention in traditional methods, making it more efficient to extract valuable evaluation information from a large number of user comments; through the fully automated evaluation triple identification solution, the present invention replaces the time-consuming manual operations in traditional methods with automatic processing, greatly shortening the time from comment data to the generation of evaluation triples; specifically, the present invention can shorten the cycle of triple identification from several months in traditional methods to several days, greatly improving the processing speed;

[0047] 2) Reduce labor costs: The automated process of the present invention significantly reduces the dependence on manual intervention and reduces labor costs; through the automatic processing and dimension design of the large language model, enterprises do not need a large amount of human resources for data annotation or code table collection, thus saving labor and economic costs;

[0048] 3) Improve recognition accuracy: In the data processing flow of the present invention, two large language models are combined. One is used for data screening and preprocessing, with the advantage of high computational efficiency; the other is responsible for in-depth semantic analysis and dimension generation, ensuring high-precision recognition results. This combination improves the processing efficiency while maintaining a high level of model performance. By leveraging the powerful semantic understanding and analysis capabilities of the large language model, the present invention can accurately extract evaluation dimensions and evaluation words, enhancing the accuracy and reliability of the recognition results, thereby making the comment analysis more precise.

[0049] 4) The dimension normalization process of the present invention includes two rounds of processing: preliminary clustering and secondary clustering, ensuring the accuracy and consistency of the evaluation dimensions. Preliminary clustering is used to merge similar dimensions, while secondary clustering further optimizes the dimension framework, removing redundant dimensions and ensuring a clear definition of the final dimensions. This dual processing strategy improves the accuracy of dimension normalization and the reliability of the system.

[0050] 5) The present invention also introduces a reclassification mechanism for outlier dimensions, analyzing and reclassifying the outlier dimensions generated during the two rounds of clustering to ensure the matching of each dimension with the evaluation words and incorporating them into the optimized dimension system. This mechanism solves the problem of dimension inconsistency and improves the comprehensiveness and accuracy of the dimension framework.

[0051] 6) Enhance dimension flexibility: The dynamic dimension generation mechanism of the present invention enables the evaluation dimensions to be quickly adjusted according to changes in actual data and application scenarios, overcoming the problem of poor flexibility and inability to update in a timely manner in traditional triple models. At the same time, due to the adaptability and flexibility of the large language model, the present invention can process data in different fields, including but not limited to different industries and different framework requirements. This enables the technical solution of the present invention to achieve excellent results in a wide range of application scenarios. Description of the Drawings

[0052] Figure 1 It is a flowchart of a method for identifying user evaluation triples based on a large model provided in Embodiment 1.

[0053] Figure 2 It is a framework diagram of a method for identifying user evaluation triples based on a large model provided in Embodiment 2.

[0054] Figure 3 It is a structural diagram of a system for identifying user evaluation triples based on a large model provided in Embodiment 3. Detailed Embodiments

[0055] The drawings are only for illustrative purposes and should not be construed as a limitation of this patent.

[0056] To better illustrate this embodiment, some components in the drawings are omitted, enlarged or reduced, which does not represent the size of the actual product;

[0057] For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.

[0058] The technical solution of the present invention will be further described below with reference to the drawings and embodiments.

[0059] Embodiment 1

[0060] As Figure 1 shown, this embodiment provides a method for identifying user evaluation triples based on a large model, including the following steps:

[0061] S1: Collect user comment data from multiple platforms on the Internet and construct a comment data set;

[0062] S2: Set the quality screening conditions for user comment data, and use a pre-trained first large language model to screen the comment data set to obtain a high-quality comment data set;

[0063] S3: Use a pre-trained second large language model to perform sense segmentation on the high-quality comment data set, and segment each high-quality user comment data into several sense segments;

[0064] S4: Use a pre-trained second large language model to perform sentiment analysis on each sense segment to obtain the sentiment tendency of each sense segment; and extract the core evaluation words from each sense segment, and summarize the extracted core evaluation words in combination with the semantic information of the corresponding sense segment to refine specific evaluation dimensions;

[0065] S5: Use a pre-trained phrase clustering model to perform preliminary clustering on all the refined evaluation dimensions according to similarity to obtain a preliminary clustering result;

[0066] S6: Use a pre-trained second large language model to clean the preliminary clustering result, eliminate the outlier dimensions that do not meet the preset conditions, and redefine the evaluation dimensions of the cleaned preliminary clustering result to normalize the naming of all evaluation dimensions to obtain a redefined preliminary clustering result;

[0067] S7: Repeat steps S5 - S6 to perform secondary clustering and naming normalization on the redefined preliminary clustering result to obtain a redefined secondary clustering result;

[0068] S8: Reclassify the outlier dimensions eliminated in the two clustering processes and incorporate them into the redefined secondary clustering result to construct a final evaluation dimension framework;

[0069] Match the core evaluation words and sentiment tendencies of each semantic group segment with the final evaluation dimension framework to obtain the evaluation triple of each semantic group segment, and complete the identification of the user evaluation triple.

[0070] In the specific implementation process, the main purpose of this method is to achieve fully automated evaluation triple identification by combining two large language models, reduce the dependence on manual labor, improve the identification effect, and provide flexible dimension framework design capabilities;

[0071] This method includes three main stages: data screening, dimension generation, and dimension normalization; the processes of each stage will be described in detail below;

[0072] 1) Data collection and screening:

[0073] Data collection: First, use web crawler technology to legally crawl user comment data from multiple platforms on the Internet to construct a comment data set. In this embodiment, the user comment data are all comment data of a specific industry or brand;

[0074] Data quality assessment: In the data screening stage, it is first necessary to assess the quality of the comment data to ensure the accuracy of subsequent processing and reduce subsequent resource consumption; use an LLM model with a small number of parameters (such as a fine-tuned model with 0.2B parameters), and set the evaluation criteria for data quality according to conditions such as the density and object richness of the comments; the LLM model with a small number of parameters serves as a data quality screening tool in this step, which can effectively filter out high-quality comment data and avoid the negative impact of low-quality data on the final effect;

[0075] 2) Dimension generation:

[0076] Segment semantic group segments: Segment the screened comment data; a semantic group is an independent meaningful segment in a comment, and here it is required that each semantic group focuses on one evaluation dimension; use an LLM model with a large number of parameters (such as a fine-tuned model with 32B parameters) to analyze the semantic structure of the comment and segment the comment into semantic group segments with clear semantics. This process can accurately identify and divide the main evaluation cores in the comment;

[0077] Judge comment sentiment: Conduct sentiment analysis on each semantic group segment to judge its sentiment tendency as positive, negative, or neutral; sentiment analysis is one of the key steps in evaluating triples. Through the sentiment classification ability of the large language model, an accurate sentiment label can be assigned to each semantic group segment;

[0078] Extract evaluation words: Extract the core evaluation words from each semantic group segment; evaluation words are divided into attribute words and descriptive words, where attribute words reflect the main evaluation dimensions of the review (such as "comfort"), and descriptive words describe the characteristics or states of the evaluation dimensions (such as "very good"); through the vocabulary extraction function of a large-scale language model (LLM), relevant evaluation words can be accurately extracted from the semantic group segments and classified;

[0079] Refine evaluation dimensions: Combine the semantic information of the semantic group segments to summarize the extracted evaluation words and refine specific evaluation dimensions; the dimension refinement process utilizes the semantic understanding ability of the large language model to map the evaluation words into a structured dimension framework to form a hierarchical evaluation dimension;

[0080] 3) Dimension normalization:

[0081] Evaluation dimension clustering: Conduct preliminary clustering on the refined evaluation dimensions, use a phrase clustering model to merge the dimensions, and set a loose threshold during the preliminary clustering process to initially classify similar dimensions into one category and associate the evaluation words with these clustering results; this process aims to reduce redundant dimensions and improve the aggregation effect of the dimensions;

[0082] Cleaning of clustering results: After preliminary clustering, it is necessary to clean the clustering results and remove outlier dimensions that do not meet the specifications; this includes those dimensions that are clearly not matched with the evaluation words or do not meet the definition specifications; the cleaning process helps to improve the accuracy and consistency of the dimension framework;

[0083] Redefinition of evaluation dimensions: Based on the cleaned clustering results, redefine a unified dimension label for each class; for example, merge "wearing experience - comfort - softness", "product fabric - fabric experience - softness", and dimensions with the same semantics in this class into a unified "wearing experience - comfort - softness" dimension; the redefinition process ensures the consistency and accuracy of the dimension labels;

[0084] Secondary clustering and normalization of evaluation dimensions: Conduct secondary clustering and normalization on the redefined dimensions to ensure that the final dimensions are more precise; set a more strict and precise threshold for the second round of clustering in this step, perform secondary cleaning and secondary redefinition after secondary clustering to further optimize the accuracy and consistency of the dimensions;

[0085] Reclassification of outlier dimensions: Reclassify the outlier dimensions excluded during the two rounds of cleaning processes and incorporate them into the evaluation dimensions after secondary normalization; for dimensions that cannot be classified, exclude them to ensure that the final obtained dimension framework has high integrity and consistency;

[0086] This method can reduce manual intervention. Through automated model training and data processing workflows, it reduces human participation and improves efficiency. At the same time, this method can also improve the recognition effect: by leveraging the powerful semantic understanding ability of large language models, it enhances the precision and recall rate of evaluating triple recognition. Additionally, this method can enhance dimensional flexibility, enabling the automatic generation and dynamic adjustment of dimensional frameworks to meet the requirements of different fields and application scenarios.

[0087] Example 2

[0088] This embodiment provides a method for identifying user evaluation triples based on a large model, including the following steps:

[0089] S1: Collect user comment data from multiple platforms on the Internet and construct a comment data set;

[0090] S2: Set quality screening conditions for user comment data, and use a pre-trained first large language model to screen the comment data set to obtain a high-quality comment data set;

[0091] S3: Use a pre-trained second large language model to perform sense segmentation on the high-quality comment data set, and separately segment each high-quality user comment data into several sense segments;

[0092] S4: Use a pre-trained second large language model to perform sentiment analysis on each sense segment to obtain the sentiment tendency of each sense segment; and extract core evaluation words from each sense segment, and summarize the extracted core evaluation words in combination with the semantic information of the corresponding sense segment to refine specific evaluation dimensions;

[0093] S5: Use a pre-trained phrase clustering model to perform preliminary clustering on all the refined evaluation dimensions according to similarity to obtain a preliminary clustering result;

[0094] S6: Use a pre-trained second large language model to clean the preliminary clustering result, remove outlier dimensions that do not meet the preset conditions, and redefine the evaluation dimensions for the cleaned preliminary clustering result to normalize the naming of all evaluation dimensions to obtain a redefined preliminary clustering result;

[0095] S7: Repeat steps S5 - S6 to perform secondary clustering and naming normalization on the redefined preliminary clustering result to obtain a redefined secondary clustering result;

[0096] S8: Reclassify the outlier dimensions removed during the two clustering processes and incorporate them into the redefined secondary clustering result to construct a final evaluation dimension framework;

[0097] Match the core evaluation words and sentiment tendencies of each semantic group segment with the final evaluation dimension framework to obtain the evaluation triple of each semantic group segment, and complete the identification of the user evaluation triple;

[0098] In step S2, the set quality screening conditions for user comment data include any one of comment length, grammatical integrity, comment paragraph density, and object richness included;

[0099] In step S3, the semantic group segment is specifically a segment expressing a specific evaluation dimension in high-quality user comment data, and each semantic group segment includes at least one evaluation word;

[0100] Use the semantic structure analysis function of the second large language model to perform segmentation of semantic group segments;

[0101] In step S4, the sentiment tendency includes any one of positive, neutral, and negative;

[0102] The core evaluation words include evaluation attribute words and evaluation description words;

[0103] The evaluation attribute words are used to reflect the main evaluation dimensions of the comment; the evaluation description words are used to describe the characteristics or states of the evaluation dimensions;

[0104] Use the sentiment classification function of the second large language model to perform sentiment analysis; use the vocabulary extraction function of the second large language model to extract core evaluation words; use the semantic understanding function of the second large language model to obtain the semantic information of each semantic group segment and perform refinement of evaluation dimensions;

[0105] In step S6, the outlier dimensions that do not meet the preset conditions include evaluation dimensions that cannot be matched with all core evaluation words and evaluation dimensions that do not meet the preset definition specifications;

[0106] The number of parameters of the first large language model is less than that of the second large language model; in this embodiment, the first large language model is specifically the mini version after quantization of the SocialGPT large language model, and the number of parameters is specifically 0.2B;

[0107] The second large language model is specifically the SocialGPT large language model, and the number of parameters is specifically 32B;

[0108] The phrase clustering model is specifically the BGE-M3 semantic vector model.

[0109] In the specific implementation process, as Figure 2 shown, this method includes three main stages: data screening, dimension generation, and dimension normalization; the processes of each stage will be described in detail below;

[0110] 1) Data collection and screening:

[0111] Data collection: First, user review data is legally crawled from multiple social media and e-commerce platforms on the Internet through web crawler technology to construct a review dataset. In this embodiment, the user review data are all review data about clothing and footwear products;

[0112] Data quality assessment: In the data screening stage, it is first necessary to conduct quality assessment and screening on the review data to ensure the accuracy of subsequent processing and reduce subsequent resource consumption; use an LLM model with a small number of parameters (the model used in this embodiment is the mini version quantized from the self-developed SocialGPT, with 0.2B parameters), and set the evaluation criteria for data quality according to conditions such as the density and object richness of the reviews; the LLM model with a small number of parameters serves as a data quality screening tool in this step, which can effectively filter out high-quality review data and avoid the negative impact of low-quality data on the final effect; as shown in Table 1, Table 1 is an example of the original review data and the data after screening, where " / " indicates discarding the corresponding review data;

[0113] Table 1 Example of original review data and screened data

[0114] Example of original data Example of filtered data This pair of shoes is very comfortable and the price is also reasonable. This pair of shoes is very comfortable and the price is also reasonable. Good / The shoes are stylishly designed, but the soles are very hard and not comfortable enough. The shoes are stylishly designed, but the soles are very hard and not comfortable enough. Recommended by someone / Not worth the price. Not worth the price. The pants have a good fit and are very comfortable to wear. The pants have a good fit and are very comfortable to wear.

[0115] As can be seen from Table 1, through data screening, some low-quality reviews, such as ambiguous or information-insufficient reviews, are removed, ensuring relatively high-quality data for subsequent processing;

[0116] 2) Dimension generation:

[0117] Segment semantic groups: Segment the screened review data into semantic groups; a semantic group is a segment with independent meaning in a review, and here it is required that each semantic group focuses on one evaluation dimension; use an LLM model with a large number of parameters (in this embodiment, it is the self-developed SocialGPT model, which has been fine-tuned on specific label recognition tasks and has 32B parameters) to analyze the semantic structure of the review and segment the review into semantic groups with clear semantics. This process can accurately identify and divide the main evaluation cores in the review;

[0118] Judge review sentiment: Conduct sentiment analysis on each semantic group segment to judge its sentiment tendency as positive, negative, or neutral; sentiment analysis is one of the key steps in evaluating triple recognition. Through the sentiment classification ability of the large language model, an accurate sentiment label can be assigned to each semantic group segment;

[0119] Extract evaluation words: Extract the core evaluation words from each semantic group segment; evaluation words are divided into attribute words and descriptive words. Among them, attribute words reflect the main evaluation dimensions of the comment (such as "comfort", "style design"), and descriptive words describe the characteristics or states of the evaluation dimensions (such as "very good", "very", "fashionable"); through the vocabulary extraction function of the large parameter LLM model, relevant evaluation words can be accurately extracted from the semantic group segment and classified;

[0120] Refine evaluation dimensions: Combine the semantic information of the semantic group segment to summarize the extracted evaluation words and refine specific evaluation dimensions; the dimension refinement process utilizes the semantic understanding ability of the large language model to map the evaluation words into a structured dimension framework to form a hierarchical evaluation dimension. Table 2 shows a data example after the dimension generation stage;

[0121] Table 2 Data example after the dimension generation stage

[0122] Semantic group segment Emotional tendency Core evaluation word Evaluation dimension The comfort level is very good Positive Comfort#Very good Product experience - Usage experience - Comfort level The price is also very reasonable Positive Price#Very reasonable Price / promotion - Price - Price satisfaction The shoes are stylishly designed Positive Design#Stylish Product design - Appearance design - Style design The soles are very hard and not comfortable enough Negative Comfort#The soles are very hard Usage experience - Comfort level - Not comfortable enough Not worth the price Negative Price#Not worth it Price - Value The pants have a good fit Positive Fit#Good Fit design Very comfortable to wear Positive Comfortable#To wear Wearing experience - Wearing comfort - Comfort level

[0123] As can be seen from Table 2, in the dimension generation stage, through the in-depth semantic analysis and semantic group segmentation of the large parameter LLM model, the effective extraction of different dimension information in the comment is realized. At the same time, the classification of evaluation words and the refinement of dimensions enable the accurate identification and summary of the core information of each comment;

[0124] 3) Dimension normalization:

[0125] Evaluation dimension clustering: Conduct preliminary clustering on the refined evaluation dimensions. Use a phrase clustering model (in this embodiment, the BGE-M3 semantic vector model fine-tuned with comment phrases is used) to merge the dimensions, and set a loose threshold during the preliminary clustering process to ensure a large range of clustering. This step is used to initially classify similar dimensions into one category and associate the evaluation words with these clustering results; this process aims to reduce redundant dimensions and improve the aggregation effect of dimensions. Table 3 shows an example of the preliminary clustering results;

[0126] Table 3 Example of preliminary clustering results

[0127]

[0128]

[0129] Clustering result cleaning: After preliminary clustering, it is necessary to clean the clustering results, analyze each clustering result, find out the outlier dimensions that do not meet the specifications or are significantly mismatched with the evaluation words and remove them; the cleaning process helps to improve the accuracy and consistency of the dimension framework. Table 4 shows an example of the preliminary clustering results after cleaning;

[0130] Example of the preliminary clustering results after cleaning in Table 4

[0131]

[0132] Redefinition of evaluation dimensions: Based on the clustering results after cleaning, redefine unified dimension labels for each class; for example, merge "Wearing experience - Comfort - Softness", "Product fabric - Fabric experience - Softness", and dimensions with the same semantics in this class into a unified "Wearing experience - Comfort - Softness" dimension; the redefinition process ensures the consistency and accuracy of dimension labels;

[0133] Secondary clustering and normalization of evaluation dimensions: Perform secondary clustering and normalization on the redefined dimensions to ensure that the final dimensions are more precise; set more strict and precise thresholds for the second round of clustering in this step, and perform secondary cleaning and secondary redefinition after secondary clustering to further optimize the precision and consistency of dimensions; Table 5 shows an example of the evaluation dimensions after secondary clustering and normalization;

[0134] Table 5 Example of the evaluation dimensions after secondary clustering and normalization

[0135]

[0136] Re - classification of outlier dimensions: Re - classify the outlier dimensions excluded during the two - round cleaning process and incorporate them into the evaluation dimensions after secondary normalization; exclude dimensions that cannot be classified to ensure that all evaluation words can be accurately matched to the final dimension system, and at the same time ensure that the final obtained dimension framework has a high degree of integrity and consistency; Table 6 shows an example of the re - classification results of outlier dimensions;

[0137] Cluster ID Outlier dimension Evaluation word Normalized dimension name 1 2 Price - Value Price#Not worth it Price / promotion - Price - Price satisfaction 3 Fit design Fit#Good Product design - Appearance design - Style design

[0138] Finally, obtain the evaluation triples for each semantic group segment. Table 7 shows an example of the evaluation triples for the semantic group segments finally obtained;

[0139] Table 7 Example of the evaluation triples for the semantic group segments finally obtained

[0140] Semantic group segment Emotional tendency Evaluation word Final evaluation dimension The comfort level is very good Positive Comfort#Very good Product experience - Wearing experience - Comfort level The price is also very reasonable Positive Price#Very reasonable Price / promotion - Price - Price satisfaction The shoes are stylishly designed Positive Design#Stylish Product design - Appearance design - Style design The soles are very hard and not comfortable enough Negative Comfort#The soles are very hard Product experience - Wearing experience - Comfort level Not worth the price Negative Price#Not worth it Price / promotion - Price - Price satisfaction The pants have a good fit Positive Fit#Good Product design - Appearance design - Style design Very comfortable to wear Positive Comfortable#To wear Product experience - Wearing experience - Comfort level

[0141] As can be seen from Table 7, the dimension normalization stage optimizes the classification of evaluation dimensions through preliminary and secondary clustering, making the dimension framework more accurate and consistent; at the same time, the cleaning and redefinition steps ensure the unity of each dimension label and the effective matching of evaluation words;

[0142] It is worth mentioning that the first large language model, the second large language model, and the phrase clustering model used in this method do not depend on a specific model. Various currently commonly used open-source and closed-source large models can achieve the expected goals and effects, and this embodiment does not make specific restrictions here either;

[0143] This method can reduce manual intervention. Through an automated model training and data processing process, it reduces human participation and improves efficiency. At the same time, this method can also improve the recognition effect: Utilizing the powerful semantic understanding ability of the large language model, it improves the accuracy and recall rate of the evaluation triple recognition. In addition, this method can also enhance dimensional flexibility, realize the automatic generation and dynamic adjustment of the dimension framework to meet the needs of different fields and application scenarios.

[0144] Embodiment 3

[0145] As Figure 3 shown, this embodiment provides a user evaluation triple recognition system based on a large model, applying the method for recognizing user evaluation triples based on a large model described in Embodiment 1 or 2, including:

[0146] A comment collection unit 301: used to collect user comment data from multiple platforms on the Internet and construct a comment data set;

[0147] A comment screening unit 302: used to set quality screening conditions for user comment data, and use a pre-trained first large language model to screen the comment data set to obtain a high-quality comment data set;

[0148] A comment segmentation unit 303: used to perform sense group segmentation on the high-quality comment data set using a pre-trained second large language model, and respectively segment each high-quality user comment data into several sense group segments;

[0149] An emotion analysis and dimension extraction unit 304: used to perform emotion analysis on each sense group segment using a pre-trained second large language model to obtain the emotion tendency of each sense group segment; and extract core evaluation words from each sense group segment, and summarize the extracted core evaluation words in combination with the semantic information of the corresponding sense group segment to refine specific evaluation dimensions;

[0150] A clustering unit 305: used to perform preliminary clustering on all the refined evaluation dimensions according to similarity using a pre-trained phrase clustering model to obtain a preliminary clustering result;

[0151] Named normalization unit 306: It is used to clean the preliminary clustering result by using a pre-trained second large language model, eliminate the outlier dimensions that do not meet the preset conditions, redefine the evaluation dimensions for the cleaned preliminary clustering result, normalize the names of all evaluation dimensions, and obtain the redefined preliminary clustering result;

[0152] Secondary normalization unit 307: It is used to repeatedly execute the clustering unit 305 and the named normalization unit 306, perform secondary clustering and named normalization on the redefined preliminary clustering result, and obtain the redefined secondary clustering result;

[0153] Evaluation triple recognition unit 308: It is used to reclassify the outlier dimensions eliminated in the two clustering processes and incorporate them into the redefined secondary clustering result to construct the final evaluation dimension framework;

[0154] Match the core evaluation words and sentiment tendencies of each semantic group segment with the final evaluation dimension framework to obtain the evaluation triple of each semantic group segment, and complete the recognition of the user evaluation triple.

[0155] In the specific implementation process, first, the comment collection unit 301 collects data from multiple platforms on the Internet;

[0156] After collecting the comment data, the comment screening unit 302 first uses the first large language model to screen the data quality;

[0157] Then, the comment segmentation unit 303 uses the second large language model to segment the screened data into semantic groups. The sentiment analysis and dimension extraction unit 304 performs sentiment analysis on each semantic group segment, extracts the core evaluation words from each semantic group segment, and summarizes the extracted evaluation words in combination with the semantic information of the semantic group segment to refine specific evaluation dimensions;

[0158] After that, the clustering unit 305 performs preliminary clustering on the evaluation dimensions using a phrase clustering model; the named normalization unit 306 cleans the preliminary clustering result; after cleaning, the evaluation dimensions are redefined;

[0159] After that, the secondary normalization unit 307 performs secondary clustering and normalization on the redefined evaluation dimensions;

[0160] Finally, the evaluation triple recognition unit 308 reclassifies the outlier dimensions eliminated in the two clustering processes and incorporates them into the evaluation dimensions after secondary clustering and normalization; saves the final evaluation dimensions and their corresponding core evaluation words and sentiment tendencies as comment triples;

[0161] This system can reduce manual intervention. Through automated model training and data processing processes, it reduces human participation and improves efficiency. At the same time, this system can also improve the recognition effect: by leveraging the powerful semantic understanding ability of the large language model, it enhances the accuracy and recall rate of the evaluation triple recognition. In addition, this system can also enhance dimensional flexibility, enabling the automatic generation and dynamic adjustment of the dimension framework to meet the requirements of different fields and application scenarios.

[0162] The same or similar reference numerals correspond to the same or similar components;

[0163] The terms used to describe the positional relationship in the drawings are for illustrative purposes only and should not be construed as a limitation of this patent;

[0164] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the claims of the present invention.

Claims

1. A method for identifying user evaluation triples based on a large model, characterized in that: The following steps are involved: S1: Collect user comment data from multiple platforms on the Internet and build a comment dataset; S2: setting a quality screening condition for user comment data, and using the pre-trained first language model to screen the comment data set to obtain a high-quality comment data set; S3: using the pre-trained second largest language model to segment the high-quality comment data set into semantic groups, and segmenting each high-quality user comment data into a plurality of semantic group segments; S4: using the pre-trained second language model to perform sentiment analysis on each of the phrase segments to obtain the sentiment tendency of each phrase segment; And extract the core evaluation words from each phrase segment, and summarize the extracted core evaluation words in combination with the semantic information of the corresponding phrase segment to extract specific evaluation dimensions; S5: Use the pre-trained phrase clustering model to perform preliminary clustering on all the extracted evaluation dimensions according to similarity to obtain preliminary clustering results; S6: using the pre-trained second language model to clean the preliminary clustering results, remove outlier dimensions that do not meet the preset conditions, redefine the evaluation dimensions of the cleaned preliminary clustering results, normalize the names of all evaluation dimensions, and obtain the redefined preliminary clustering results; S7: repeating steps S5 to S6, performing secondary clustering and naming normalization on the redefined preliminary clustering results, and obtaining the redefined secondary clustering results; S8: Reclassify the outlier dimensions removed in the two clustering processes and incorporate them into the redefined secondary clustering results to construct the final evaluation dimension framework; The core evaluation words and sentiment tendency of each idea group segment are matched with the final evaluation dimension framework to obtain the evaluation triplet of each idea group segment, thereby completing the identification of the user evaluation triplet.

2. According to the large model-based user evaluation triple identification method of claim 1, it is characterized in that: In the step S2, the quality screening conditions of the user comment data are set to include: any one of the comment length, grammatical completeness, comment paragraph density and the richness of the objects included.

3. The method for identifying user evaluation triples based on a large model according to claim 1, characterized in that: In step S3, the idea group segment is specifically a segment expressing a specific evaluation dimension in the high-quality user review data, and each idea group segment includes at least one evaluation word; The semantic structure analysis function of the second largest language model is used to segment the meaning groups.

4. The method for identifying user evaluation triples based on a large model according to claim 1, characterized in that: In step S4, the sentiment tendency includes: any one of positive, neutral and negative; The core evaluation words include: evaluation attribute words and evaluation description words; The evaluation attribute words are used to reflect the main evaluation dimensions of the review; the evaluation description words are used to describe the characteristics or status of the evaluation dimensions; The sentiment classification function of the second largest language model is used to perform sentiment analysis; the vocabulary extraction function of the second largest language model is used to extract core evaluation words; the semantic understanding function of the second largest language model is used to obtain the semantic information of each of the meaning groups, and to refine the evaluation dimensions.

5. The method for identifying user evaluation triples based on a large model according to claim 1, characterized in that: In step S6, the outlier dimensions that do not meet the preset conditions include: evaluation dimensions that cannot match all core evaluation words, and evaluation dimensions that do not meet the preset definition specifications.

6. A method for identifying user evaluation triples based on a large model according to any one of claims 1 to 5, characterized in that: The first largest language model has a smaller number of parameters than the second largest language model.

7. The method for identifying user evaluation triples based on a large model according to claim 6, characterized in that: The first language model is specifically a quantized mini version of the SocialGPT large language model; The second largest language model is specifically the SocialGPT large language model; The phrase clustering model is specifically a BGE-M3 semantic vector model.

8. The method for identifying user evaluation triples based on a large model according to claim 7, characterized in that: The parameter amount of the first largest language model is specifically 0.2B; the parameter amount of the second largest language model is specifically 32B.

9. A user evaluation triplet recognition system based on a large model, applying a user evaluation triplet recognition method based on a large model as claimed in any one of claims 1 to 8, characterized in that: include: Comment collection unit: used to collect user comment data from multiple platforms on the Internet and build a comment dataset; A comment screening unit: used to set the quality screening conditions of the user comment data, and use the pre-trained first language model to screen the comment data set to obtain a high-quality comment data set; A comment segmentation unit is used to segment the high-quality comment data set into semantic groups using the pre-trained second language model, and to segment each high-quality user comment data into a plurality of semantic group segments; Sentiment analysis and dimension extraction unit: used for performing sentiment analysis on each of the phrase segments using a pre-trained second language model to obtain the sentiment tendency of each phrase segment; And extract the core evaluation words from each phrase segment, and summarize the extracted core evaluation words in combination with the semantic information of the corresponding phrase segment to extract specific evaluation dimensions; Clustering unit: used to use the pre-trained phrase clustering model to perform preliminary clustering on all extracted evaluation dimensions according to similarity and obtain preliminary clustering results; A naming normalization unit is used to clean the preliminary clustering results using a pre-trained second language model, remove outlier dimensions that do not meet preset conditions, redefine the evaluation dimensions of the cleaned preliminary clustering results, normalize the names of all evaluation dimensions, and obtain the redefined preliminary clustering results; Secondary normalization unit: used to repeatedly execute the clustering unit and the naming normalization unit, perform secondary clustering and naming normalization on the redefined preliminary clustering results, and obtain the redefined secondary clustering results; Evaluation triplet identification unit: used to reclassify the outlier dimensions removed in the two clustering processes and incorporate them into the redefined secondary clustering results to construct the final evaluation dimension framework; The core evaluation words and sentiment tendency of each idea group segment are matched with the final evaluation dimension framework to obtain the evaluation triplet of each idea group segment, thereby completing the identification of the user evaluation triplet.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps in the method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Integrated evaluation method for E-commerce service quality

    CN108446813A

  • Product key user demand mining method driven by small sample comment data

    CN115713349A