Data Processing Method and Device Based on User's New Media APP Interaction Operations
By real-time acquisition and semantic analysis of user comment content in the new media APP, combined with the preset dynamic theme library to identify semantic clusters and determine the demand label, the problems of low user participation and data deviation in the existing technology are solved, and efficient and accurate user data collection and analysis are achieved.
Patent Information
- Application Number
- CN202510338874.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-03-21
AI Technical Summary
In the prior art, relying on users to actively submit information or obtain data by a single sensor leads to low user participation and limited coverage, and the data is easily affected by subjective factors, making it difficult to fully reflect the user's real needs and behavior patterns.
By obtaining the user's comment content in the new media APP in real time, combining the preset dynamic theme library for context semantic analysis, identifying semantic clusters, and determining the demand label based on the preset quantitative parameters, storing the ternary mapping relationship between the product logo, topic and demand label.
It realizes efficient and accurate collection and processing of user interaction data, reduces misjudgments caused by the diversity of language expression, adapts to changes in user concerns, and fully reflects the user's real needs and behavior patterns.
Smart Images

Figure CN119862890B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and particularly to a data processing method and device based on user interaction operations on new media APPs. Background Art
[0002] In today's digital age, new media APPs (such as social media and e-commerce) have become important platforms for people to obtain information, express opinions, and interact. Users' interaction behaviors in APPs, such as comments, likes, and shares, contain rich information, which is of great value for product optimization, market research, and personalized recommendations. Therefore, how to efficiently and accurately collect and process users' interaction data in new media APPs has become the focus of many enterprises and research institutions.
[0003] In related technologies, data collection technologies mainly rely on users to actively submit information or obtain data through a single sensor. For example, some APPs collect user feedback through questionnaires, and users need to manually fill in the questionnaire content and submit it. The methods that rely on users to actively submit, such as questionnaires, have obvious passivity, low user participation, limited coverage, and are easily affected by users' subjective factors, resulting in biased data collection and difficulty in comprehensively reflecting users' real needs and behavior patterns. Summary of the Invention
[0004] Embodiments of this application provide a data processing method and device based on user interaction operations on new media APPs. To provide a basic understanding of some aspects of the disclosed embodiments, a simple summary is given below. This summary part is not a general review, nor is it intended to identify key / important constituent elements or delineate the protection scope of these embodiments. Its sole purpose is to present some concepts in a simple form as a prelude to the subsequent detailed description.
[0005] In a first aspect, embodiments of this application provide a data processing method based on user interaction operations on new media APPs, the method including:
[0006] In response to a user's interaction instruction for a comment box of a target object, real-time obtain the content input by the user for the comment box; wherein, the target object is a work browsed by the user in the new media APP, and the work is used to describe a target product;
[0007] Perform context semantic analysis on the content input into the comment box in real time to identify whether there is a semantic cluster related to the content input into the comment box in real time among the semantic clusters corresponding to each theme in a preset dynamic theme library; wherein, the preset dynamic theme library is periodically updated according to multiple original comments existing in the work within a preset period;
[0008] In the case where there is a semantic cluster related to the content in the real-time input comment box, obtain the target theme to which the semantic cluster related to the content in the real-time input comment box belongs, capture a regional screenshot of the comment box, and extract all user comment texts in the regional screenshot;
[0009] Determine the demand label of the user for the target product according to the preset quantization parameter and the user comment text;
[0010] In the database, store the ternary mapping relationship among the product identifier of the target product, the target theme, and the demand label.
[0011] In a second aspect, an embodiment of the present application provides a data processing device based on user interaction operations on a new media APP. The device includes:
[0012] A response acquisition module, configured to respond to an interaction instruction of the user for a comment box of a target object, and acquire in real time the content input by the user for the comment box; wherein, the target object is a work browsed by the user in the new media APP, and the work is used to describe the target product;
[0013] A semantic cluster determination module, configured to perform context semantic analysis on the content input in the real-time comment box to identify whether there is a semantic cluster related to the content input in the real-time comment box in the semantic clusters corresponding to each theme in a preset dynamic theme library; wherein, the preset dynamic theme library is periodically updated according to multiple original comments existing in the work within a preset period;
[0014] A text extraction module, configured to, in the case where there is a semantic cluster related to the content input in the real-time comment box, obtain the target theme to which the semantic cluster related to the content input in the real-time comment box belongs, capture a regional screenshot of the comment box, and extract all user comment texts in the regional screenshot;
[0015] A demand label determination module, configured to determine the demand label of the user for the target product according to the preset quantization parameter and the user comment text;
[0016] A ternary mapping relationship storage module, configured to store, in the database, the ternary mapping relationship among the product identifier of the target product, the target theme, and the demand label.
[0017] The technical solution provided by the embodiment of the present application may include the following beneficial effects:
[0018] In an embodiment of the present application, on the one hand, by obtaining the user input content in real time and performing context semantic analysis in combination with each theme in the preset dynamic theme library, the system can real-time identify semantic clusters in the user comments, rather than just keywords. This semantic-based analysis method can effectively reduce misjudgments caused by the diversity of language expressions, thereby ensuring high accuracy of the collected data. At the same time, the system regularly updates the semantic clusters through the preset dynamic theme library, which can adapt to the changes in user focus and further improve the accuracy and timeliness of semantic recognition. On the other hand, the system determines the demand tags according to the preset quantization parameters and the user comment text, and stores the ternary mapping relationship between the product identifier of the target product, the target theme and the demand tags. This structured data storage method is not only convenient for subsequent analysis, but also can comprehensively reflect the real needs and behavior patterns of users through multi-dimensional data fusion.
[0019] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0021] Figure 1 is a schematic flowchart of a data processing method based on user interaction operations on a new media APP according to an embodiment of the present application;
[0022] Figure 2 is a schematic diagram of a comment box page according to an embodiment of the present application;
[0023] Figure 3 is a schematic diagram of prompt word information according to an embodiment of the present application;
[0024] Figure 4 is a schematic diagram of a demand tag visualization interface according to an embodiment of the present application;
[0025] Figure 5 is another schematic diagram of a demand tag visualization interface according to an embodiment of the present application;
[0026] Figure 6 is a schematic flowchart of a model training method for a semantic cluster analysis model according to an embodiment of the present application;
[0027] Figure 7 is a schematic structural diagram of a data processing device based on user interaction operations on a new media APP according to an embodiment of the present application;
[0028] Figure 8 is a schematic structural diagram of an electronic device according to an embodiment of the present application. Detailed implementation manners
[0029] The following description and the accompanying drawings fully disclose specific implementation manners of the present application, enabling those skilled in the art to practice them.
[0030] It should be clear that the described embodiments are only a part of the embodiments of the present application, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts belong to the scope protected by the present application.
[0031] When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0032] In the description of the present application, it should be understood that terms such as "first", "second", etc. are only used for descriptive purposes and cannot be construed as indicating or implying relative importance. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances. In addition, in the description of the present application, unless otherwise specified, "a plurality of" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.
[0033] Currently, data acquisition technologies mainly rely on users to actively submit information or obtain data through a single sensor. For example, some APPs collect user feedback through questionnaire surveys, and users need to manually fill in the questionnaire content and submit it.
[0034] The inventor realizes that methods relying on users to actively submit, such as questionnaire surveys, have obvious passivity, low user participation, limited coverage, and are easily affected by user subjective factors, resulting in deviations in the collected data and making it difficult to comprehensively reflect the true needs and behavior patterns of users.
[0035] To solve the above problems, the present application provides a data processing method and apparatus based on user interaction operations with new media APPs to address the problems existing in the above-related technical problems. In an embodiment of the present application, on the one hand, by obtaining user input content in real time and performing context semantic analysis in combination with each theme in a preset dynamic theme library, the system can real-time identify semantic clusters in user comments, rather than just keywords. This semantic-based analysis method can effectively reduce misjudgments caused by the diversity of language expressions, thereby ensuring high accuracy of the collected data. At the same time, the system regularly updates semantic clusters through the preset dynamic theme library, which can adapt to changes in user focus points and further improve the accuracy and timeliness of semantic recognition. On the other hand, the system determines demand tags according to preset quantization parameters and user comment texts, and stores the ternary mapping relationship between the product identifier of the target product, the target theme, and the demand tags. This structured data storage method is not only convenient for subsequent analysis but also can comprehensively reflect the real needs and behavior patterns of users through multi-dimensional data fusion. The following will be described in detail with exemplary embodiments.
[0036] The following will be combined with the attached Figure 1 - attached Figure 6 , to introduce in detail the data processing method based on user interaction operations with new media APPs provided by the embodiments of the present application. This method can be implemented depending on a computer program and can run on a data processing apparatus based on the von Neumann architecture for user interaction operations with new media APPs. This computer program can be integrated into an application or run as an independent tool-type application.
[0037] Please refer to Figure 1 , which is a schematic flowchart of a data processing method based on user interaction operations with new media APPs provided by an embodiment of the present application. As Figure 1 shown, the method of the embodiment of the present application may include the following steps:
[0038] S101, in response to an interaction instruction of a user for a comment box of a target object, obtain in real time the content input by the user for the comment box; wherein, the target object is a work browsed by the user in the new media APP, and the work is used to describe the target product;
[0039] Wherein, the user refers to an individual who browses content and performs interactions in the new media APP. The target object is the specific content browsed by the user in the new media APP and is a work. The comment box is an input box in the APP interface for the user to input comments. The interaction instruction is an operation performed by the user in the comment box, such as clicking, inputting text, etc. The content is the specific text information input by the user in the comment box. The new media APP is, for example, a social media, e-commerce, or other application programs, and the user browses and interacts through these platforms. The work is the specific content of the target object and may be a product introduction, article, video, etc. The target product is the product described in the work.
[0040] In some embodiments of the present application, when a user browses a work describing a target product in a new media APP, if the user triggers the comment box to input a comment, the system will respond to the user's operation (interaction instruction) in real time and obtain the content input by the user for the comment box.
[0041] For example, user Xiaoming is browsing the introduction page of a smart watch in a certain e-commerce APP. There are product pictures, function introductions, and a comment box for users to post comments on the page. Xiaoming triggers the comment box, and a comment box page pops up. For example Figure 2 as shown, Xiaoming inputs in the comment box: "The appearance design of this smart watch is very fashionable, but the battery life is too short", and the system will immediately respond to Xiaoming's operation and obtain the comment content he input in real time for subsequent analysis.
[0042] S102, perform context semantic analysis on the content input into the comment box in real time to identify whether there is a semantic cluster related to the content input into the comment box in real time in the semantic clusters corresponding to each theme in the preset dynamic theme library; wherein, the preset dynamic theme library is obtained by periodically updating based on multiple original comments existing in the works within a preset period;
[0043] Among them, context semantic analysis is used to combine the character with the previously input characters for semantic analysis every time a character is input into the comment box. Since the processing speed of the algorithm is much faster than the speed of the user inputting characters, every time the user inputs a character, it can be combined with the previously input characters to form new content for real-time analysis, thereby improving the real-time performance of the analysis. The preset dynamic theme library is a database that is periodically updated based on multiple original comments collected within a preset period, and is used to store and manage the themes related to the product and the semantic clusters corresponding to the themes. A semantic cluster is a group of related words and phrases around a certain theme or concept, and these words are semantically similar or related. The original comments are the historical views expressed by different users in the new media APP for a certain product or content. The preset period is the time interval for periodically updating the dynamic theme library.
[0044] In some embodiments of the present application, after the client obtains the content input by the user for the comment box, for example, at the current moment, the last character "le" of the sentence "The appearance design of this smart watch is very fashionable, but the battery life is too short" input by Xiaoming, at this time, the client can splice the last character "le" with the previously input "The appearance design of this smart watch is very fashionable, but the battery life is too short" to obtain the content input into the comment box in real time, that is, "The appearance design of this smart watch is very fashionable, but the battery life is too short le".
[0045] For example, the system performs context semantic analysis on the comment input by Xiaoming. The system searches in the preset dynamic theme library to check if there is a semantic cluster related to the sentence "The appearance design of this smart watch is very fashionable, but the battery life is too short". Suppose the semantic clusters corresponding to each theme in the dynamic theme library are as follows (appearance design: fashionable, design sense, high appearance value, fashionable appearance; battery life: short battery life, durable battery). At this time, through searching, the system can identify that the sentence "The appearance design of this smart watch is very fashionable, but the battery life is too short" exists in two semantic clusters in the dynamic theme library respectively.
[0046] In the embodiment of the present application, the specific process of identifying whether there is a semantic cluster corresponding to each theme in the preset dynamic theme library that is related to the content in the real-time input comment box includes: traversing and obtaining the semantic clusters corresponding to each theme from the preset dynamic theme library; extracting the semantic feature vector of the content in the real-time input comment box as the center point vector; taking the embedding vectors of the semantic clusters corresponding to each theme as multiple edge vectors of the center point vector, and establishing the relationship between the center point vector and each edge vector to obtain multiple objects to be analyzed; for each object to be analyzed, calculating the semantic similarity between the center point vector and each edge vector to obtain multiple semantic similarities of each object to be analyzed; calculating the similarity mean value of the multiple semantic similarities of each object to be analyzed to obtain the target semantic similarity of each object to be analyzed; when there is an object to be analyzed with a target semantic similarity greater than the preset threshold, it is determined that there is a semantic cluster corresponding to each theme in the preset dynamic theme library that is related to the content in the real-time input comment box; or, when there is no object to be analyzed with a target semantic similarity greater than the preset threshold, it is determined that there is no semantic cluster corresponding to each theme in the preset dynamic theme library that is related to the content in the real-time input comment box.
[0047] For example, the dynamic theme library contains the following themes and corresponding semantic clusters. Theme: Product appearance, semantic cluster embedding vector: [0.8, 0.1, 0.2] (representing the embedding vectors of words such as "fashionable", "design sense", "high appearance value"). Theme: Product performance, semantic cluster embedding vector: [0.1, 0.7, 0.3] (representing the embedding vectors of words such as "short battery life", "durable battery"). Theme: Price, semantic cluster embedding vector: [0.2, 0.2, 0.6] (representing the embedding vectors of words such as "high cost performance", "reasonable price"). Xiaoming's comment content is: "The appearance design of this smart watch is very fashionable, but the battery life is too short." The system extracts the semantic feature vector of this comment through natural language processing (NLP) technology. Suppose the obtained center point vector is: [0.7, 0.4, 0.5]. Taking the embedding vectors of the semantic clusters corresponding to each theme in the dynamic theme library as edge vectors and establishing a relationship with the center point vector, multiple objects to be analyzed are obtained:
[0048] Object to be analyzed 1: Center point vector: [0.7, 0.4, 0.5], Edge vector (product appearance): [0.8, 0.1, 0.2]. Object to be analyzed 2: Center point vector: [0.7, 0.4, 0.5], Edge vector (product performance): [0.1, 0.7, 0.3]. Object to be analyzed 3: Center point vector: [0.7, 0.4, 0.5], Edge vector (price): [0.2, 0.2, 0.6].
[0049] At this time, for each object to be analyzed, calculate the semantic similarity between the center point vector and the edge vector (for example, using cosine similarity). The calculation results are as follows: Object to be analyzed 1 (product appearance): Similarity = 0.85, Object to be analyzed 2 (product performance): Similarity = 0.78, Object to be analyzed 3 (price): Similarity = 0.35. Assume that the preset threshold is 0.6. According to the target semantic similarity: Product appearance (0.85) > 0.6, there is a relevant semantic cluster. Product performance (0.78) > 0.6, there is a relevant semantic cluster. Price (0.35) < 0.6, there is no relevant semantic cluster.
[0050] At this time, the system can determine that: Xiaoming's comment is relevant to the semantic cluster in the "product appearance" theme (such as "fashionable"). Xiaoming's comment is relevant to the semantic cluster in the "product performance" theme (such as "short battery life"). Xiaoming's comment is not relevant to the semantic cluster in the "price" theme.
[0051] In the embodiment of the present application, the preset dynamic theme library can be continuously updated. The specific update process is as follows: Obtain multiple original comments existing in the works within a preset period; From the multiple original comments, analyze the original comments that are not relevant to each theme in the preset dynamic theme library as the comments to be analyzed; Input each comment to be analyzed into a pre-trained semantic cluster analysis model, and output the semantic cluster corresponding to each comment to be analyzed; Among them, the pre-trained semantic cluster analysis model is a data model obtained by machine learning based on the historical user comment data stored within a preset time period; Perform theme classification on the semantic cluster corresponding to each comment to be analyzed to obtain the theme corresponding to each comment to be analyzed; From the themes corresponding to each comment to be analyzed, count the number of themes of the same theme; In the case where the number of themes of the same theme is greater than or equal to the preset number threshold, store the binary mapping relationship between the same theme and all semantic clusters corresponding to the same theme in the preset dynamic theme library. Among them, by regularly updating the semantic cluster through the preset dynamic theme library, it is possible to adapt to the changes in user focus and further improve the accuracy and timeliness of semantic recognition.
[0052] For example, the system obtains user comments on a smart watch in the past week. The comments are as follows:
[0053] 1. "The exterior design of this smartwatch is very fashionable.";
[0054] 2. "The battery life is too short. I hope it can be improved.";
[0055] 3. "The price is a bit high and the cost performance is not good.";
[0056] 4. "The waterproof function of the watch is very good.";
[0057] 5. "The operation interface is very complex and not very user-friendly."
[0058] Assume that the existing themes in the dynamic theme library are "Product Appearance" and "Product Performance". The system analyzes these comments and finds that the 3rd and 5th comments are not relevant to the existing themes. Therefore, they are marked as comments to be analyzed. The comments to be analyzed are input into the semantic cluster analysis model, and the model outputs the following content:
[0059] The semantic clusters corresponding to the 3rd comment: Price, Cost performance. The semantic clusters corresponding to the 5th comment: Operation interface, User experience.
[0060] The system classifies these semantic clusters into the corresponding themes: The 3rd comment: The theme is "Price". The 5th comment: The theme is "User experience".
[0061] Count the number of comments for each theme: Theme "Price": 1 comment, Theme "User experience": 1 comment. Assume that the preset quantity threshold is 2, and currently, the number of comments for no theme reaches the threshold. Assume that after a period of time, the system continues to collect comments and finds that the number of comments for the "Price" theme reaches 3 (exceeding the threshold). Then, the "Price" theme and its corresponding semantic clusters are stored in the dynamic theme library. Among them, the updated binary mapping relation table in the dynamic theme library is shown in Table 1 for example.
[0062] Table 1
[0063]
[0064] In the embodiments of the present application, the specific process of generating a pre-trained semantic cluster analysis model includes: collecting historical user comment data stored within a preset time period from a new media APP, where the historical user comment data covers multiple works; performing data preprocessing on the historical user comment data to obtain a vocabulary sequence corresponding to each piece of historical user comment data, and the data preprocessing includes text cleaning, word segmentation, stop word removal, and part-of-speech tagging; capturing semantic information for the vocabulary sequence corresponding to each piece of historical user comment data to convert the vocabulary sequence corresponding to each piece of historical user comment data into a high-dimensional vector, and obtaining multiple vocabulary features for each sample comment data; constructing a semantic cluster for each sample comment data according to the multiple vocabulary features of each sample comment data; using the semantic cluster of each sample comment data as a label to label the vocabulary sequence corresponding to each piece of historical user comment data to obtain model training samples; creating a semantic cluster analysis model; inputting the model training samples into the semantic cluster analysis model for machine learning and outputting a model loss value; when the model loss value reaches the minimum, generating a pre-trained semantic cluster analysis model.
[0065] Among them, text cleaning includes removing noises in the comments, such as HTML tags, special characters, emojis, etc. Word segmentation includes splitting the comment text into separate vocabulary sequences. For example, splitting "This mobile phone is very useful" into "This / mobile phone / very / useful". Stop word removal includes removing common words that contribute less to semantics, such as "of", "is", "and", etc. Part-of-speech tagging includes tagging the part of speech (noun, verb, adjective, etc.) for each word. For example, the original comment: "The battery life of this mobile phone is really good!", after preprocessing: ["This", "mobile phone", "battery", "life", "really", "good"].
[0066] Among them, semantic information capture converts the vocabulary sequence into a high-dimensional vector through algorithms (such as Word2Vec, BERT, etc.), and these vectors can capture the semantic information of the vocabulary. The high-dimensional vector is that each word or comment is represented as a point in a high-dimensional space, and similar words or comments are closer in the space. For example, the vocabulary sequence: ["mobile phone", "battery", "life"], after conversion into a high-dimensional vector is [[0.1, 0.2, 0.3], [0.4, 0.5, 0.6], [0.7, 0.8, 0.9]].
[0067] Among them, using the semantic cluster of each comment as a label for supervised learning. The labeled data is used to train the semantic cluster analysis model. For example, the comment is ["mobile phone", "battery", "life"], and the semantic cluster label is ["battery", "life"], and the final training sample of this comment is ([0.1, 0.2, 0.3], ["battery", "life"]).
[0068] During the training process, the model loss value gradually decreases from 1.0 to 0.1. At this time, the model training is completed, and a semantic cluster analysis model is generated.
[0069] In the embodiment of the present application, the specific process of constructing the semantic cluster of each sample comment data according to multiple lexical features of each sample comment data is as follows: Initialize multiple lexical features of each sample comment data to mark the multiple lexical features as unvisited; Traverse the first unvisited lexical feature among the multiple lexical features of each sample comment data, and use the traversed lexical feature as the to-be-analyzed lexical feature, and mark the status of the to-be-analyzed lexical feature as visited; Query all unvisited first lexical features within the domain of the to-be-analyzed lexical feature; Calculate the similarity between the to-be-analyzed lexical feature and each first lexical feature; Group the first lexical features with a similarity greater than or equal to the preset similarity threshold and the to-be-analyzed lexical feature into one cluster to obtain the clustering clusters of each sample comment data; Continue to execute the step of traversing the first unvisited lexical feature among the multiple lexical features of each sample comment data until all the multiple lexical features are marked as visited, and obtain multiple clustering clusters of each sample comment data; Reverse-transform the lexical features in each clustering cluster of each sample comment data to obtain the lexical sequence of each clustering cluster of each sample comment data; Input the lexical sequence of each clustering cluster of each sample comment data into a preset large language model to analyze the representative vocabulary of each clustering cluster of each sample comment data, and obtain the semantic cluster of each sample comment data.
[0070] Among them, the calculation formula of the similarity is:
[0071] ;
[0072] Among them, is the similarity, is the th first lexical feature, is the to-be-analyzed lexical feature, is the number of features of the first lexical feature.
[0073] Among them, the model training samples include the semantic cluster of each sample comment data and each sample comment data; the semantic cluster analysis model includes an acquisition module, a probability fitting module, a probability prediction module, a cross-entropy loss value calculation module, a model loss value calculation module, and a loss value output module.
[0074] In the embodiments of the present application, the specific process of inputting model training samples into a semantic cluster analysis model for machine learning and outputting a model loss value is as follows: The acquisition module acquires the semantic clusters of each sample comment data; the probability fitting module fits the first probability distribution of the semantic clusters of each sample comment data; the probability prediction module predicts the second probability distribution of each sample comment data; the cross-entropy loss value calculation module calculates the cross-entropy loss value of each sample comment data according to the first probability distribution and the second probability distribution; the model loss value calculation module calculates the cross-entropy loss value of the model training samples based on the cross-entropy loss values of each sample comment data as the model loss value; the loss value output module outputs the model loss value.
[0075] Among them, the loss function of the model loss value is:
[0076]
[0077] Among them, is the model loss value, is the number of sample comment data in the model training samples, is the number of semantic clusters, is the th true label of the th semantic cluster of the th sample. If the th sample belongs to the th semantic cluster, then = 1, otherwise = 0, is the probability that the model predicts the th sample belongs to the
[0078] S103. In the case where there is a semantic cluster related to the content in the real-time input comment box, obtain the target theme to which the semantic cluster related to the content in the real-time input comment box belongs, intercept a regional screenshot of the comment box, and extract all user comment texts in the regional screenshot;
[0079] In the embodiments of the present application, the system first identifies whether the comment content input by the user is related to a preset semantic cluster through semantic analysis. If there is a correlation, determine the theme to which the semantic cluster related to the user input content belongs. For example, if the user comment is related to the semantic cluster of "battery life", then the comment belongs to the "product performance" theme. The system intercepts the screen area where the comment box is located and generates a screenshot. The purpose of this step is to capture the visual information of the comment box and its related products. The system extracts the text content of all user comments from the screenshot through OCR (Optical Character Recognition) technology. This step can obtain the comments input by the user and further replies made by other users to this comment. The relevant comments of other users can assist in comprehensive analysis.
[0080] For example, the system performs semantic analysis on Xiaoming's comment and identifies the following semantic clusters: "Design Patent" (related to the "Product Appearance" theme), "Battery Life" (related to the "Product Performance" theme). The system determines the themes related to Xiaoming's comment: for "Design Patent", the target theme is "Product Appearance". For "Battery Life", the target theme is "Product Performance". The system automatically captures the screen area where the comment box is located and generates a screenshot. This screenshot may include the comment box, Xiaoming's comment content, and the comments of other users. The system extracts the text content of all user comments from the screenshot through OCR technology. Suppose the screenshot contains the following comments: Xiaoming's comment: "The design of this smartwatch is very fashionable, but the battery life is too short", Other user's comment: "Indeed, the battery life is a problem", "The design is great".
[0081] S104, determine the demand tags of the user for the target product according to the preset quantization parameters and the user's comment text;
[0082] Among them, the preset quantization parameters include an emotion threshold and an attention threshold.
[0083] In the embodiment of the present application, the specific process of determining the demand tags of the user for the target product according to the preset quantization parameters and the user's comment text includes: performing sentiment tendency analysis on the user's comment text through a sentiment analysis tool to obtain the user's sentiment polarity score; determining the user's sentiment tendency description word according to the emotion threshold and the user's sentiment polarity score; matching the user's comment text with a preset keyword template through the TF-IDF algorithm to determine multiple keywords in the user's comment text; counting the length of the user's comment text as the attention of the user to the target product; determining the user's attention tendency description word according to the attention threshold and the attention; constructing a prompt word for the target product according to the sentiment tendency description word, the attention tendency description word, and the multiple keywords, where the prompt word is used to guide a preset large language model to generate the demand tags of the user for the target product; inputting the prompt word into the preset large language model and outputting the demand tags of the user for the target product; among them, the preset large language model is ChatGPT.
[0084] For example, use a sentiment analysis tool to analyze Xiaoming's comment and the comments of other users who replied to this comment, and obtain sentiment polarity scores. Suppose the scores output by the sentiment analysis tool are: positive sentiment: 0.7, negative sentiment: 0.6, neutral sentiment: 0.1. The sentiment polarity scores indicate that the comment contains both positive sentiment ("The design is very fashionable") and negative sentiment ("The battery life is too short"). According to the preset sentiment threshold (assuming the positive sentiment threshold is 0.8 and the negative sentiment threshold is 0.6), determine the sentiment tendency descriptor. The negative sentiment score of 0.6 reaches the threshold, so the sentiment tendency descriptor is "negative". Use the TF-IDF algorithm to extract the keywords in the comment. Suppose the preset keyword template includes "appearance design", "battery life", "price", etc., and the extraction result is: keywords: ["appearance design", "battery life"]. Count the length of Xiaoming's comment (counted by the number of characters or words). Suppose the comment length is 35 characters. According to the preset attention threshold (assuming the high attention threshold is 30 characters), determine the attention tendency descriptor. The comment length of 35 characters exceeds the threshold, so the attention tendency descriptor is "high attention". At this time, construct a prompt for the target product to guide the large language model to generate demand tags. The prompt is as follows Figure 3 as shown. Input the above information into a preset large language model (such as ChatGPT), and the model outputs the demand tags of the user for the target product as follows:
[0085] "The demand tags of the user for the smart watch: satisfied with the appearance design, but hope to improve the battery life".
[0086] S105. In the database, store the triple mapping relationship between the product identifier of the target product, the target theme, and the demand tags.
[0087] Among them, the product identifier of the target product is that each product has a unique identifier in the database (such as product ID) to distinguish different products. The target theme is the theme involved in the user's comment, such as "product appearance", "product performance", "price", etc. The demand tag is a summary description of the user's needs, such as "satisfied with the appearance design", "battery life to be optimized", etc. The triple mapping relationship stores the relationship between the product identifier, the theme, and the demand tag in the form of a triple for quick query and analysis.
[0088] Among them, the triple mapping relationship is as shown in Table 2.
[0089] Table 2
[0090]
[0091] In some embodiments of the present application, after storing the ternary mapping relationship between the product identifier of the target product, the target theme, and the demand label, in some visualization display scenarios, a data query request input by the client is received. The data query request carries the product identifier to be analyzed and the theme to be analyzed; from the ternary mapping relationship, all demand labels corresponding to the product identifier to be analyzed and the theme to be analyzed are obtained; all the same demand labels are aggregated to obtain non-repeated demand labels; the non-repeated demand labels are imported into a visualization component to visually display the non-repeated demand labels. Among them, the product identifier to be analyzed and the theme to be analyzed are selected and submitted in the demand label visualization interface, and the demand label visualization interface is, for example Figure 4 as shown, and the display result of the visual display of the demand label is, for example Figure 5 as shown.
[0092] In the embodiments of the present application, on the one hand, by real-time obtaining the user input content and performing context semantic analysis in combination with each theme in the preset dynamic theme library, the system can real-time identify the semantic clusters in the user comments, rather than just keywords. This semantic-based analysis method can effectively reduce misjudgments caused by the diversity of language expressions, thereby ensuring high accuracy of the collected data. At the same time, the system regularly updates the semantic clusters through the preset dynamic theme library, which can adapt to the changes in the user's focus of attention and further improve the accuracy and timeliness of semantic recognition. On the other hand, the system determines the demand labels according to the preset quantization parameters and the user comment text, and stores the ternary mapping relationship between the product identifier of the target product, the target theme, and the demand label. This structured data storage method is not only convenient for subsequent analysis, but also can comprehensively reflect the true needs and behavior patterns of users through multi-dimensional data fusion.
[0093] Please refer to Figure 6 , which is a schematic flowchart of a model training method for a semantic cluster analysis model provided by an embodiment of the present application. As Figure 6 shown, the method of the embodiment of the present application may include the following steps:
[0094] S201, collect historical user comment data stored within a preset time period from the new media APP, and the historical user comment data covers multiple works;
[0095] S202, perform data preprocessing on the historical user comment data to obtain a vocabulary sequence corresponding to each historical user comment data. The data preprocessing includes text cleaning, word segmentation, stop word removal, and part-of-speech tagging;
[0096] S203, capture semantic information for the vocabulary sequence corresponding to each historical user comment data to convert the vocabulary sequence corresponding to each historical user comment data into a high-dimensional vector, and obtain multiple vocabulary features of each sample comment data;
[0097] S204. Construct a semantic cluster for each sample comment data according to multiple lexical features of each sample comment data;
[0098] S205. Use the semantic cluster of each sample comment data as a label to label the lexical sequence corresponding to each historical user comment data, and obtain model training samples;
[0099] S206. Create a semantic cluster analysis model;
[0100] S207. Input the model training samples into the semantic cluster analysis model for machine learning, and output the model loss value;
[0101] S208. When the model loss value reaches the minimum, generate a pre-trained semantic cluster analysis model.
[0102] In the embodiment of the present application, on the one hand, by adopting real-time acquisition of user input content and combining with each theme in the preset dynamic theme library for context semantic analysis, the system can real-time identify the semantic clusters in the user comments, rather than just keywords. This semantic-based analysis method can effectively reduce misjudgments caused by the diversity of language expressions, thereby ensuring high accuracy of the collected data. At the same time, the system regularly updates the semantic clusters through the preset dynamic theme library, which can adapt to the changes in user focus and further improve the accuracy and timeliness of semantic recognition. On the other hand, the system determines the demand labels according to the preset quantization parameters and the user comment text, and stores the ternary mapping relationship between the product identifier of the target product, the target theme and the demand labels. This structured data storage method is not only convenient for subsequent analysis, but also can comprehensively reflect the real needs and behavior patterns of users through multi-dimensional data fusion.
[0103] The following is an embodiment of the device of the present application, which can be used to execute the embodiment of the method of the present application. For the details not disclosed in the embodiment of the device of the present application, please refer to the embodiment of the method of the present application.
[0104] Please refer to Figure 7 , which shows a schematic structural diagram of a data processing device based on user interaction operations on a new media APP provided by an exemplary embodiment of the present application. The data processing device based on user interaction operations on a new media APP can be implemented as all or part of an electronic device through software, hardware or a combination of both. The device 1 includes a response acquisition module 10, a semantic cluster determination module 20, a text extraction module 30, a demand label determination module 40, and a ternary mapping relationship storage module 50.
[0105] The response acquisition module 10 is configured to, in response to an interaction instruction of a user for a comment box of a target object, real-time acquire the content input by the user for the comment box; wherein, the target object is a work browsed by the user in the new media APP, and the work is used to describe the target product;
[0106] The semantic cluster determination module 20 is used to perform context semantic analysis on the content in the real-time input comment box to identify whether there is a semantic cluster related to the content in the real-time input comment box in the semantic clusters corresponding to each theme in the preset dynamic theme library; wherein, the preset dynamic theme library is obtained by periodically updating multiple original comments existing in the works within a preset period.
[0107] The text extraction module 30 is used to, when there is a semantic cluster related to the content in the real-time input comment box, obtain the target theme to which the semantic cluster related to the content in the real-time input comment box belongs, capture a regional screenshot of the comment box, and extract all user comment texts in the regional screenshot.
[0108] The requirement label determination module 40 is used to determine the requirement labels of the user for the target product according to the preset quantization parameters and the user comment texts.
[0109] The triple mapping relationship storage module 50 is used to store the triple mapping relationship among the product identifier of the target product, the target theme, and the requirement labels in the database.
[0110] It should be noted that when the data processing device based on the user's new media APP interaction operation provided in the above embodiment executes the data processing method based on the user's new media APP interaction operation, only the above division of each functional module is used for illustration. In actual application, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the data processing device based on the user's new media APP interaction operation provided in the above embodiment and the data processing method embodiment based on the user's new media APP interaction operation belong to the same concept, and the implementation process thereof is detailed in the method embodiment, which will not be repeated here.
[0111] The serial numbers of the above embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments.
[0112] In an embodiment of the present application, on the one hand, by obtaining the user input content in real time and performing context semantic analysis in combination with each theme in the preset dynamic theme library, the system can identify semantic clusters in the user comments in real time, rather than just keywords. This semantic-based analysis method can effectively reduce misjudgments caused by the diversity of language expressions, thereby ensuring high accuracy of the collected data. At the same time, the system updates the semantic clusters regularly through the preset dynamic theme library, which can adapt to the changes in the user's focus of attention and further improve the accuracy and timeliness of semantic recognition. On the other hand, the system determines demand tags according to the preset quantization parameters and the user comment text, and stores the ternary mapping relationship between the product identifier of the target product, the target theme, and the demand tags. This structured data storage method is not only convenient for subsequent analysis, but also can comprehensively reflect the user's real needs and behavior patterns through multi-dimensional data fusion.
[0113] The present application also provides a computer-readable medium, on which program instructions are stored. When the program instructions are executed by a processor, the data processing method based on the user's new media APP interaction operation provided by each of the above method embodiments is implemented.
[0114] The present application also provides a computer program product containing instructions. When it runs on a computer, it causes the computer to execute the data processing method based on the user's new media APP interaction operation provided by each of the above method embodiments.
[0115] Please refer to Figure 8 , which is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 8 shown, the electronic device 1000 may include: at least one processor 1001, at least one network interface 1004, a user interface 1003, a memory 1005, and at least one communication bus 1002.
[0116] Among them, the communication bus 1002 is used to realize the connection and communication between these components.
[0117] Among them, the user interface 1003 may include a display screen (Display) and a camera (Camera). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface.
[0118] Among them, the network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface).
[0119] Among them, the processor 1001 may include one or more processing cores. The processor 1001 connects various parts within the entire electronic device 1000 through various interfaces and lines. By running or executing instructions, programs, code sets, or instruction sets stored in the memory 1005, and by calling the data stored in the memory 1005, it performs various functions of the electronic device 1000 and processes data. Optionally, the processor 1001 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 1001 may integrate one or a combination of several of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, application programs, etc.; the GPU is responsible for rendering and drawing the content to be displayed on the display screen; the modem is used to process wireless communication. It can be understood that the above-mentioned modem may not be integrated into the processor 1001 and may be implemented separately by a single chip.
[0120] Among them, the memory 1005 may include random access memory (RAM) and may also include read-only memory. Optionally, the memory 1005 includes a non-transitory computer-readable storage medium. The memory 1005 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 1005 may include a program storage area and a data storage area. Among them, the program storage area can store instructions for implementing the operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area can store the data involved in the above-mentioned various method embodiments. Optionally, the memory 1005 may also be at least one storage system located far from the aforementioned processor 1001. As Figure 8 shown, the memory 1005, as a computer storage medium, may include an operating system, a network communication module, a user interface module, and a data processing application program for user interaction operations based on a new media APP.
[0121] In Figure 8In the electronic device 1000 shown, the user interface 1003 is mainly used to provide an interface for the user to input and obtain the data input by the user. The processor 1001 can be used to call the data processing application program stored in the memory 1005 based on the user's new media APP interaction operation, and specifically perform the following operations:
[0122] In response to the user's interaction instruction for the comment box of the target object, real-time obtain the content input by the user for the comment box; wherein, the target object is the work browsed by the user in the new media APP, and the work is used to describe the target product;
[0123] Perform context semantic analysis on the content input into the comment box in real time to identify whether there is a semantic cluster related to the content input into the comment box in real time in the semantic clusters corresponding to each theme in the preset dynamic theme library; wherein, the preset dynamic theme library is obtained by regularly updating based on multiple original comments existing in the works within a preset period;
[0124] In the case where there is a semantic cluster related to the content input into the comment box in real time, obtain the target theme to which the semantic cluster related to the content input into the comment box in real time belongs, intercept the area screenshot of the comment box, and extract all the user comment texts in the area screenshot;
[0125] Determine the demand label of the user for the target product according to the preset quantization parameter and the user comment text;
[0126] In the database, store the ternary mapping relationship between the product identifier of the target product, the target theme and the demand label.
[0127] In one embodiment, the processor 1001 also performs the following operations:
[0128] Obtain multiple original comments existing in the works within a preset period;
[0129] Analyze the original comments that are not related to each theme in the preset dynamic theme library from the multiple original comments as the comments to be analyzed;
[0130] Input each comment to be analyzed into the pre-trained semantic cluster analysis model, and output the semantic cluster corresponding to each comment to be analyzed; wherein, the pre-trained semantic cluster analysis model is a data model obtained by machine learning based on the historical user comment data stored within a preset time period;
[0131] Perform theme classification on the semantic cluster corresponding to each comment to be analyzed to obtain the theme corresponding to each comment to be analyzed;
[0132] From the themes corresponding to each comment to be analyzed, count the number of themes of the same theme;
[0133] When the number of topics of the same theme is greater than or equal to a preset number threshold, store the binary mapping relationship between the same theme and all semantic clusters corresponding to the same theme in the preset dynamic theme library.
[0134] In one embodiment, when the processor 1001 executes to generate the pre-trained semantic cluster analysis model, it specifically performs the following operations:
[0135] Collect historical user comment data stored within a preset time period from the new media APP. The historical user comment data covers multiple works;
[0136] Perform data preprocessing on the historical user comment data to obtain the vocabulary sequence corresponding to each historical user comment data. The data preprocessing includes text cleaning, word segmentation, stop word removal, and part-of-speech tagging;
[0137] Capture semantic information for the vocabulary sequence corresponding to each historical user comment data to convert the vocabulary sequence corresponding to each historical user comment data into a high-dimensional vector, and obtain multiple vocabulary features of each sample comment data;
[0138] Construct the semantic cluster of each sample comment data according to the multiple vocabulary features of each sample comment data;
[0139] Use the semantic cluster of each sample comment data as a label to label the vocabulary sequence corresponding to each historical user comment data, and obtain the model training samples;
[0140] Create a semantic cluster analysis model;
[0141] Input the model training samples into the semantic cluster analysis model for machine learning, and output the model loss value;
[0142] When the model loss value reaches the minimum, generate the pre-trained semantic cluster analysis model.
[0143] In one embodiment, when the processor 1001 executes to construct the semantic cluster of each sample comment data according to the multiple vocabulary features of each sample comment data, it specifically performs the following operations:
[0144] Initialize the multiple vocabulary features of each sample comment data to mark the multiple vocabulary features as unvisited;
[0145] Traverse the first unvisited vocabulary feature among the multiple vocabulary features of each sample comment data, and use the traversed vocabulary feature as the vocabulary feature to be analyzed, and mark the status of the vocabulary feature to be analyzed as visited;
[0146] Query all unvisited first vocabulary features within the domain of the vocabulary feature to be analyzed;
[0147] Calculate the similarity between the lexical features to be analyzed and each first lexical feature;
[0148] Group the first lexical features with a similarity greater than or equal to the preset similarity threshold and the lexical features to be analyzed into a cluster to obtain the clustering clusters of each sample comment data;
[0149] Continue to execute the step of traversing the first unvisited lexical feature among the multiple lexical features of each sample comment data until all the multiple lexical features are marked as visited, to obtain multiple clustering clusters of each sample comment data;
[0150] Reverse-transform the lexical features in each clustering cluster of each sample comment data to obtain the lexical sequence of each clustering cluster of each sample comment data;
[0151] Input the lexical sequence of each clustering cluster of each sample comment data into a preset large language model to analyze the representative words of each clustering cluster of each sample comment data, so as to obtain the semantic clusters of each sample comment data.
[0152] In one embodiment, when the processor 1001 executes inputting the model training samples into the semantic cluster analysis model for machine learning and outputting the model loss value, it specifically performs the following operations:
[0153] The acquisition module acquires the semantic clusters of each sample comment data;
[0154] The probability fitting module fits the first probability distribution of the semantic clusters of each sample comment data;
[0155] The probability prediction module predicts the second probability distribution of each sample comment data;
[0156] The cross-entropy loss value calculation module calculates the cross-entropy loss value of each sample comment data according to the first probability distribution and the second distribution probability;
[0157] The model loss value calculation module calculates the cross-entropy loss value of the model training samples according to the cross-entropy loss value of each sample comment data as the model loss value;
[0158] The loss value output module outputs the model loss value.
[0159] In one embodiment, when the processor 1001 executes identifying whether there is a semantic cluster related to the content in the real-time input comment box in the semantic clusters corresponding to each theme in the preset dynamic theme library, it specifically performs the following operations:
[0160] Traverse and obtain the semantic clusters corresponding to each theme from the preset dynamic theme library;
[0161] Extract the semantic feature vector of the content in the real-time input comment box as the center point vector;
[0162] Take the embedding vectors of the semantic clusters corresponding to each theme as multiple marginal vectors of the center point vector, and establish the relationship between the center point vector and each marginal vector to obtain multiple objects to be analyzed;
[0163] For each object to be analyzed, calculate the semantic similarity between the center point vector and each marginal vector to obtain multiple semantic similarities of each object to be analyzed;
[0164] Calculate the similarity mean of the multiple semantic similarities of each object to be analyzed to obtain the target semantic similarity of each object to be analyzed;
[0165] When there is an object to be analyzed whose target semantic similarity is greater than the preset threshold, it is determined that there is a semantic cluster related to the content in the real-time input comment box in the semantic clusters corresponding to each theme in the preset dynamic theme library; or,
[0166] When there is no object to be analyzed whose target semantic similarity is greater than the preset threshold, it is determined that there is no semantic cluster related to the content in the real-time input comment box in the semantic clusters corresponding to each theme in the preset dynamic theme library.
[0167] In one embodiment, when the processor 1001 executes to determine the demand label of the user for the target product according to the preset quantization parameter and the user's comment text, the following operations are specifically executed:
[0168] Perform sentiment tendency analysis on the user's comment text through a sentiment analysis tool to obtain the user's sentiment polarity score;
[0169] Determine the user's sentiment tendency description word according to the sentiment threshold and the user's sentiment polarity score;
[0170] Match the user's comment text with the preset keyword template through the TF-IDF algorithm to determine multiple keywords in the user's comment text;
[0171] Count the length of the user's comment text as the user's attention to the target product;
[0172] Determine the user's attention tendency description word according to the attention threshold and the attention;
[0173] Construct a prompt word for the target product according to the sentiment tendency description word, the attention tendency description word and multiple keywords, and the prompt word is used to guide the preset large language model to generate the demand label of the user for the target product;
[0174] Input the prompt word into the preset large language model to output the demand label of the user for the target product; where,
[0175] The preset large language model is ChatGPT.
[0176] In one embodiment, the processor 1001 further performs the following operations:
[0177] Receive a data query request for client input, where the data query request carries the product identifier to be analyzed and the subject to be analyzed;
[0178] Obtain all requirement tags that meet the product identifier to be analyzed and the corresponding subject to be analyzed from the ternary mapping relationship;
[0179] Aggregate all the same requirement tags to obtain non-duplicate requirement tags;
[0180] Import the non-duplicate requirement tags into the visualization component to visually display the non-duplicate requirement tags.
[0181] In the embodiments of the present application, on the one hand, by adopting real-time acquisition of user input content and performing context semantic analysis in combination with each theme in the preset dynamic theme library, the system can real-time identify semantic clusters in user comments, rather than just keywords. This semantic-based analysis method can effectively reduce misjudgments caused by the diversity of language expressions, thereby ensuring high accuracy of the collected data. At the same time, the system regularly updates the semantic clusters through the preset dynamic theme library, which can adapt to changes in user focus and further improve the accuracy and timeliness of semantic recognition. On the other hand, the system determines requirement tags according to preset quantization parameters and user comment texts, and stores the ternary mapping relationship between the product identifier of the target product, the target theme, and the requirement tags. This structured data storage method is not only convenient for subsequent analysis, but also can comprehensively reflect the real needs and behavior patterns of users through multi-dimensional data fusion.
[0182] Those of ordinary skill in the art can understand that all or part of the processes of implementing the above method embodiments can be completed by instructing relevant hardware through a computer program. The program for data processing based on user interaction operations on the new media APP can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above method embodiments. Among them, the storage medium of the program for data processing based on user interaction operations on the new media APP can be a magnetic disk, an optical disk, a read-only memory, or a random access memory, etc.
[0183] The above-disclosed are only the preferred embodiments of the present application. Of course, the scope of rights of the present application cannot be limited thereby. Therefore, equivalent changes made according to the claims of the present application still fall within the scope covered by the present application.
Claims
1. A data processing method based on user new media APP interactive operation, characterized in that: The method comprises: In response to an interactive instruction of a user on a comment box of a target object, obtaining content input by the user on the comment box in real time; wherein the target object is a work browsed by the user in a new media APP, and the work is used to describe a target product; Performing contextual semantic analysis on the content input into the comment box in real time to identify whether there is a semantic cluster related to the content input into the comment box in real time in the semantic clusters corresponding to each theme in the preset dynamic theme library, the semantic cluster being output by a pre-trained semantic cluster analysis model; wherein the preset dynamic theme library is obtained by regularly updating a plurality of original comments existing in the work within a preset period; In the case where there is a semantic cluster related to the content input into the comment box in real time, obtaining a target topic to which the semantic cluster related to the content input into the comment box in real time belongs, taking a regional screenshot of the comment box, and extracting all user comment texts in the regional screenshot; Determining the user's demand label for the target product according to the preset quantitative parameters and the user comment text; In a database, a ternary mapping relationship between the product identification of the target product, the target subject and the demand tag is stored; and a pre-trained semantic cluster analysis model is generated according to the following steps, including: Collecting historical user comment data stored in a preset time period from the new media APP, where the historical user comment data covers multiple works; Performing data preprocessing on the historical user comment data to obtain a vocabulary sequence corresponding to each piece of historical user comment data, wherein the data preprocessing includes text cleaning, word segmentation, stop word removal, and part-of-speech tagging; Capturing semantic information of the vocabulary sequence corresponding to each piece of historical user comment data, so as to convert the vocabulary sequence corresponding to each piece of historical user comment data into a high-dimensional vector, and obtaining a plurality of vocabulary features of each piece of historical user comment data; Constructing a semantic cluster of each sample comment data according to multiple lexical features of each sample comment data; Using the semantic cluster of each sample comment data as a label, annotating the vocabulary sequence corresponding to each historical user comment data, and obtaining a model training sample; Create semantic cluster analysis model; Inputting the model training samples into the semantic cluster analysis model for machine learning, and outputting a model loss value; When the model loss value reaches a minimum, generating a pre-trained semantic cluster analysis model; The step of constructing a semantic cluster of each piece of sample comment data according to the plurality of lexical features of each piece of sample comment data comprises: Initializing a plurality of lexical features of each sample comment data to mark the plurality of lexical features as unvisited; Traversing the first unvisited lexical feature among the multiple lexical features of each sample comment data, taking the traversed lexical feature as the lexical feature to be analyzed, and marking the state of the lexical feature to be analyzed as visited; Querying all unvisited first vocabulary features in the field of the vocabulary feature to be analyzed; Calculating the similarity between the to-be-analyzed vocabulary feature and each first vocabulary feature; Classifying the first vocabulary feature whose similarity is greater than or equal to a preset similarity threshold and the vocabulary feature to be analyzed into one cluster, to obtain a clustering cluster for each sample comment data; Continue to perform the step of traversing the first unvisited lexical feature among the multiple lexical features of each sample comment data until all the multiple lexical features are marked as visited, thereby obtaining multiple clusters of each sample comment data; Reversely transforming the vocabulary features in each cluster of each sample comment data to obtain the vocabulary sequence of each cluster of each sample comment data; The vocabulary sequence of each cluster of each sample comment data is input into a preset large language model to analyze the representative vocabulary of each cluster of each sample comment data to obtain the semantic cluster of each sample comment data.
2. The method according to claim 1, characterized in that The method further comprises: Obtain multiple original reviews of the work within a preset period; Analyzing the original comments that are not related to the topics in the preset dynamic topic library from the multiple original comments as the comments to be analyzed; Input each comment to be analyzed into a pre-trained semantic cluster analysis model, and output the semantic cluster corresponding to each comment to be analyzed; wherein the pre-trained semantic cluster analysis model is a data model obtained by machine learning based on historical user comment data stored in a preset time period; Performing topic classification on the semantic cluster corresponding to each comment to be analyzed to obtain the topic corresponding to each comment to be analyzed; From the topics corresponding to each of the comments to be analyzed, count the number of topics with the same topic; When the number of topics of the same topic is greater than or equal to a preset number threshold, a binary mapping relationship between the same topic and all semantic clusters corresponding to the same topic is stored in the preset dynamic topic library.
3. The method according to claim 1, characterized in that The model training sample includes the semantic cluster of each sample comment data and each sample comment data; the semantic cluster analysis model includes an acquisition module, a probability fitting module, a probability prediction module, a cross entropy loss value calculation module, a model loss value calculation module and a loss value output module; The step of inputting the model training sample into the semantic cluster analysis model for machine learning and outputting the model loss value comprises: The acquisition module acquires the semantic cluster of each sample comment data; The probability fitting module fits a first probability distribution of a semantic cluster of each sample comment data; The probability prediction module predicts a second probability distribution of each sample comment data; The cross entropy loss value calculation module calculates the cross entropy loss value of each sample comment data according to the first probability distribution and the second probability distribution; The model loss value calculation module calculates the cross entropy loss value of the model training sample according to the cross entropy loss value of each sample comment data as the model loss value; The loss value output module outputs the model loss value.
4. The method according to claim 1, characterized in that The step of identifying whether there is a semantic cluster related to the content input into the comment box in real time in the semantic clusters corresponding to each theme in the preset dynamic theme library, includes: From the preset dynamic theme library, traverse and obtain the semantic clusters corresponding to each theme; Extracting a semantic feature vector of the content input into the comment box in real time as a center point vector; Using the embedding vectors of the semantic clusters corresponding to the topics as multiple edge vectors of the center point vector, and establishing a relationship between the center point vector and each edge vector to obtain multiple objects to be analyzed; For each object to be analyzed, calculating the semantic similarity between the center point vector and each edge vector to obtain multiple semantic similarities of each object to be analyzed; Calculating a similarity mean of multiple semantic similarities of each object to be analyzed to obtain a target semantic similarity of each object to be analyzed; When there is an object to be analyzed whose target semantic similarity is greater than a preset threshold, it is determined that there is a semantic cluster related to the content input into the comment box in real time among the semantic clusters corresponding to each topic in the preset dynamic topic library; or, when there is no object to be analyzed whose target semantic similarity is greater than the preset threshold, it is determined that there is no semantic cluster related to the content input into the comment box in real time among the semantic clusters corresponding to each topic in the preset dynamic topic library.
5. The method according to claim 1, characterized in that The preset quantitative parameters include an emotion threshold and an attention threshold; The step of determining the user's demand label for the target product according to the preset quantitative parameter and the user comment text includes: Performing sentiment analysis on the user's comment text to obtain the user's sentiment polarity score; Determining a descriptive word for the user's emotional tendency according to the emotional threshold and the user's emotional polarity score; Matching the user comment text with a preset keyword template by using a TF-IDF algorithm to determine a plurality of keywords in the user comment text; Counting the length of the user's comment text as the user's attention to the target product; Determining a descriptive word for the user's attention tendency according to the attention threshold and the attention degree; Constructing a prompt word for the target product according to the sentiment tendency description word, the attention tendency description word and the multiple keywords, wherein the prompt word is used to guide a preset large language model to generate a demand label of the user for the target product; The prompt word is input into a preset large language model, and a demand label of the user for the target product is output.
6. The method according to claim 1, characterized in that After storing the ternary mapping relationship between the product identification of the target product, the target subject and the demand tag, the method further includes: Receiving a data query request input by a client, wherein the data query request carries an identification of a product to be analyzed and a subject to be analyzed; From the ternary mapping relationship, obtain all the required tags that satisfy the product identifier to be analyzed and the subject to be analyzed; Aggregate all the same requirement tags to obtain non-repeated requirement tags; Import non-repeated requirement tags into the visualization component to visualize the non-repeated requirement tags.
7. A data processing device based on user new media APP interactive operation implemented by the method described in any one of claims 1 to 6, characterized in that: The device comprises: A response acquisition module, for responding to an interactive instruction of a user on a comment box of a target object, and acquiring in real time the content input by the user on the comment box; wherein the target object is a work browsed by the user in a new media APP, and the work is used to describe a target product; A semantic cluster determination module is used to perform contextual semantic analysis on the content input into the comment box in real time, so as to identify whether there is a semantic cluster related to the content input into the comment box in real time in the semantic clusters corresponding to each theme in the preset dynamic theme library; wherein the preset dynamic theme library is obtained by regularly updating a plurality of original comments existing in the work within a preset period; A text extraction module, for obtaining, when there is a semantic cluster related to the content input into the comment box in real time, a target topic to which the semantic cluster related to the content input into the comment box in real time belongs, taking a regional screenshot of the comment box, and extracting all user comment texts in the regional screenshot; A demand label determination module, used to determine the user's demand label for the target product according to preset quantitative parameters and the user comment text; The ternary mapping relationship storage module is used to store the ternary mapping relationship between the product identification of the target product, the target subject and the demand tag in a database.
Citation Information
Patent Citations
Emotion analysis method and device for commodity comment big data and storage medium
CN118350377A
New media content value evaluation method and device and computer equipment
CN119128464A