Text analysis method, device, medium, and program product based on feedback adjustment
Patent Information
- Application Number
- CN202611055474.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-15
- Publication Date
- 2026-09-22
AI Technical Summary
[0005]本申请实施例提供一种基于反馈调整的文本分析方法、设备、介质及程序产品,用以解决现有技术中分析结果无法根据业务反馈持续优化的技术问题
[0013]在本实施例中,通过获取人工标注数据,并根据该数据对文本分析过程中使用的至少一个模型进行训练或微调,以及调整聚类操作对应的参数,并基于更新后的模型和调整后的参数重新执行分析,直至满足预设的迭代终止条件,形成了人工反馈与算法优化的闭环迭代机制。由此,解决了现有文本分析方法中人工标注结果无法用于优化算法模型和聚类参数、系统难以根据业务反馈持续改进分析效果的技术问题,实现了分析模型的逐步优化和分析结果的精准调整,减少了重复人工干预,使系统能够适应业务需求的变化。
Smart Images

Figure CN122796201A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of natural language processing technology, and in particular to a text analysis method, device, medium, and program product based on feedback adjustment. Background Technology
[0002] With the digital transformation of the financial services industry, customer satisfaction surveys have become an important means for banks to optimize service quality. Banks collect customer feedback through open-ended questionnaires, which contain a large amount of unstructured text data. How to efficiently and accurately automate the analysis of this text has become a pressing technical problem for the banking industry.
[0003] Currently, the analysis of questionnaire text data typically relies on traditional machine learning methods. These methods generally include the following steps: First, classification rules are constructed by manually defining keywords or regular expressions; or, feature engineering is used to extract statistical information such as term frequency-inverse document frequency (TF-IDF) features and part-of-speech tagging from the text. Then, the extracted features are input into classification models such as support vector machines (SVM) or random forests for training. Finally, the trained model is used to classify new questionnaire texts by topic.
[0004] However, the aforementioned traditional machine learning methods suffer from the problem that the analysis results cannot be continuously optimized based on business feedback. Summary of the Invention
[0005] This application provides a text analysis method, device, medium, and program product based on feedback adjustment to solve the technical problem in the prior art that the analysis results cannot be continuously optimized based on business feedback.
[0006] In a first aspect, embodiments of this application provide a text analysis method based on feedback adjustment, the method comprising:
[0007] Obtain the preprocessed text data;
[0008] Perform clustering operations on the preprocessed text data to generate at least one text cluster class;
[0009] Perform semantic analysis on the text contained in the text cluster class to generate a category topic summary corresponding to the text cluster class;
[0010] Obtain manually labeled data for text clusters and category topic summaries, and train or fine-tune at least one model used in the text analysis process based on the manually labeled data, as well as adjust the parameters corresponding to the clustering operation;
[0011] Based on the trained or fine-tuned model and the adjusted parameters, the clustering and semantic analysis operations are re-executed until the preset iteration termination condition is met.
[0012] Output the text cluster class and corresponding category topic summary generated after the iteration terminates.
[0013] In this embodiment, manually labeled data is acquired, and at least one model used in the text analysis process is trained or fine-tuned based on this data. The parameters corresponding to the clustering operation are also adjusted. The analysis is then re-executed based on the updated model and adjusted parameters until a preset iteration termination condition is met, forming a closed-loop iterative mechanism of human feedback and algorithm optimization. This solves the technical problems in existing text analysis methods where manually labeled results cannot be used to optimize algorithm models and clustering parameters, and the system struggles to continuously improve analysis results based on business feedback. It achieves gradual optimization of the analysis model and precise adjustment of analysis results, reduces repetitive manual intervention, and enables the system to adapt to changes in business needs.
[0014] In one possible implementation, clustering is performed on the preprocessed text data to generate at least one text cluster class, including:
[0015] The preprocessed text data is input into the text embedding model to obtain the text vector corresponding to each text in the text data.
[0016] Density-based clustering algorithms are used to cluster text vectors to generate at least one text cluster class.
[0017] Text vectors that are not classified into any text cluster are marked as noise points.
[0018] In this implementation, text data is input into a text embedding model and transformed into text vectors, enabling unstructured text to be processed by clustering algorithms. A density-based clustering algorithm is used to cluster the text vectors, requiring no pre-defined number of categories and capable of recognizing clusters of arbitrary shapes, making it suitable for text data with complex semantic distributions. By marking text vectors that do not belong to any text cluster as noise points, low-quality data is automatically filtered, avoiding interference with the clustering results. This achieves the technical effects of automated vector transformation, flexible clustering, and noise filtering of text data.
[0019] In one possible implementation, the preprocessed text data is input into a text embedding model to obtain text vectors corresponding to each text in the text data, including:
[0020] The preprocessed text data is input into multiple text embedding models to obtain the text vectors output by each of the multiple text embedding models.
[0021] The text vectors output by multiple text embedding models are weighted and concatenated to obtain a fused text vector, which serves as the text vector corresponding to each text in the text data.
[0022] In this implementation, text data is input into multiple text embedding models, combining the semantic understanding advantages of different models. The text vectors output by each model are weighted and concatenated to generate a fused vector with richer semantic representation. This fused vector is then used as the text vector for each text, providing higher-quality input data for subsequent clustering. This improves the precision of text semantic representation and the accuracy of clustering results.
[0023] In one possible implementation, a density-based clustering algorithm performs clustering operations on text vectors, including:
[0024] Obtain the density distribution information of the text vector;
[0025] Based on the density distribution information, determine the parameters of the density-based clustering algorithm;
[0026] Clustering operations are performed based on parameters.
[0027] In this implementation, the density distribution information of text vectors is obtained, and the parameters of the clustering algorithm are determined based on this distribution information, allowing the parameter settings to adapt to the local density characteristics of the data. Clustering operations are then performed based on the determined parameters, avoiding the clustering bias caused by traditional fixed parameter settings—that is, too many clusters when the parameters are too small, and too few clusters when the parameters are too large. This solves the technical problem of clustering parameters relying on human experience and being unable to adapt to changes in data distribution, achieving adaptive adjustment of clustering parameters and optimization of clustering results.
[0028] In one possible implementation, the parameters of the density-based clustering algorithm are determined based on the density distribution information, including:
[0029] Calculate the variance of inter-cluster distances based on density distribution information;
[0030] Based on the variance of inter-cluster distances, determine the neighborhood radius parameter and the minimum number of points parameter for the clustering algorithm.
[0031] In this implementation, the neighborhood radius and minimum number of points are determined based on the variance of the inter-cluster distance. This allows the clustering algorithm's parameter settings to be adjusted according to the dispersion of text clusters, avoiding the problem of fixed parameters failing to match the characteristics of different text distributions. This improves the clustering algorithm's adaptability to changes in text data distribution and the accuracy of the clustering results.
[0032] In one possible implementation, after performing semantic analysis on the text contained in the text cluster class to generate a category topic summary corresponding to the text cluster class, the following steps are also included:
[0033] Construct a cluster relationship graph with text clusters as nodes and semantic similarity between text clusters as edges;
[0034] The cluster relationship graph is input into the graph neural network model to learn the node embedding vectors of each text cluster.
[0035] Based on node embedding vectors, the topic evolution path between text clusters is analyzed.
[0036] This implementation constructs a cluster relationship graph and uses a graph neural network to learn node embeddings, analyzing the evolutionary path of topics. Isolated topics are transformed into logically related evolutionary chains, revealing causal relationships and development trends between issues. This addresses the technical problem of traditional methods that only output independent topics and lack in-depth analysis, providing a deeper basis for business decisions.
[0037] In one possible implementation, the method further includes:
[0038] The original text data is preprocessed to obtain preprocessed text data. The preprocessing includes: marking low-quality text in the original text data, and using the marked low-quality text as noise point references for clustering operations.
[0039] In this implementation, the original text data is preprocessed to identify low-quality text, which is then used as noise points for clustering. This allows the clustering algorithm to exclude low-quality data during runtime, preventing interference with the clustering results. This reduces the impact of noise on clustering analysis and improves the accuracy and reliability of the clustering results.
[0040] In a second aspect, embodiments of this application provide an electronic device, including: a processor and a memory communicatively connected to the processor;
[0041] The memory stores the instructions that the computer executes;
[0042] The processor executes computer-executable instructions stored in memory to implement any of the methods of the first aspect.
[0043] Thirdly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the method of any one of the first aspects.
[0044] Fourthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the method of any one of the first aspects. Attached Figure Description
[0045] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0046] Figure 1 This is a flowchart illustrating a text analysis method based on feedback adjustment, provided as an embodiment of this application.
[0047] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0048] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the embodiments of this application.
[0049] It should be noted that the text analysis method, device, storage medium, and program product based on feedback adjustment provided in this application can be used in the field of natural language processing technology, or in any field other than natural language processing technology. This application does not limit the application field of the text analysis method, device, storage medium, and program product based on feedback adjustment.
[0050] Specific application scenarios for this application include customer satisfaction surveys and analysis in the financial services industry. In daily bank operations, open-ended questionnaires are typically used to collect customer feedback on various services, including but not limited to evaluations of wealth management product functions, suggestions on counter service processes, complaints about the user experience of mobile banking apps, and praise or criticism of account managers' service attitude. This feedback exists in unstructured text format, allowing customers to freely fill in specific content without being restricted by options. For example, a customer might describe in detail procedural problems encountered when applying for a loan, or specifically explain why the return on a certain wealth management product did not meet expectations.
[0051] This application embodiment automates the processing of these massive questionnaire texts to achieve rapid classification, core theme extraction, and structured analysis of customer feedback. Ultimately, it generates analysis reports that can be directly used by product, operations, or customer service departments, helping banks accurately identify service pain points, optimize product functions, and improve customer experience.
[0052] For the above application scenarios, existing technologies have developed different implementation methods for processing questionnaire text data as technology has advanced.
[0053] Initially, analysis relied primarily on manual methods. Business personnel or data analysts needed to read through each customer questionnaire, manually categorizing the feedback into different themes based on their understanding of the business and experience, and summarizing key information to form an analysis report. For example, when receiving feedback such as "The loan approval process is too complicated, and it took half a month to disburse the loan," business personnel would categorize it under "Loan Business" or "Service Efficiency" and record the key issue of "customer complaining about excessively long loan approval times" in the report. While this method ensured the interpretability of the analysis results, it was inefficient.
[0054] To improve processing efficiency, rule-based text classification methods emerged. Technical staff pre-analyze common feedback expressions relevant to the business, building a keyword library and regular expression rules. For example, for the service attitude category, keywords such as "poor attitude," "impatient," "indifferent," and "unresponsive" are defined; for the system performance category, keywords such as "crashes," "lag," "login failure," and "slow response" are defined. When processing questionnaire text, the system checks for these keywords through string matching. If a match is found, the text is automatically categorized into the corresponding preset category. When a customer fills in "the customer service attitude is very poor, unresponsive," the system categorizes it as a service attitude problem because it matches the keywords "poor attitude" and "unresponsive."
[0055] With the development of machine learning technology, traditional machine learning models have been introduced into text classification tasks. This method first preprocesses the questionnaire text, including stop word removal and word segmentation, dividing continuous text into independent word units. Then, feature engineering is used to extract numerical features from the text, with TF-IDF being the most commonly used. This method counts the frequency of each word in a single text (term frequency) while considering the prevalence of the word in the entire corpus (inverse document frequency), ultimately transforming each text into a fixed-dimensional numerical vector. For example, a text containing "loan," "approval," and "too slow" would be mapped to a high-dimensional sparse vector, where the dimensions corresponding to words like "loan," "approval," and "too slow" have higher weights. Next, these feature vectors and their corresponding category labels are input into classification models such as SVM or Random Forest for training, allowing the model to learn the feature distribution patterns of different text categories. SVM separates feature vectors of different categories by finding the optimal hyperplane, while Random Forest classifies by constructing multiple decision trees and combining voting results. Finally, the trained model is used to perform the same feature extraction process on newly collected questionnaire texts, inputting the resulting feature vectors into the model, which then outputs the topic category to which the text belongs.
[0056] In recent years, deep learning models have further improved the automation level of text processing. This method first performs word segmentation and serialization on the questionnaire text, transforming each text into a sequence composed of word indices, such as mapping service attitude difference to a numerical sequence like [1024, 356, 78]. Then, it constructs neural network structures such as Long Short-Term Memory (LSTM) or Convolutional Neural Network (CNN), and maps each word index to a low-dimensional dense word vector through embedding layers, making semantically similar words closer together in the vector space.
[0057] LSTM captures long-distance dependencies in text through gating mechanisms, enabling it to understand the semantic shifts in complex expressions such as "Although the customer service attitude is very good, the processing speed is too slow." CNN, on the other hand, extracts local features from text through convolutional kernels to identify key phrase patterns. Then, the network is trained using a large amount of labeled training data, and the network weights are continuously adjusted through backpropagation, allowing the model to gradually learn the mapping relationship between text content and category labels.
[0058] After training, the questionnaire text to be classified is input into the model. Through forward propagation calculations via embedding layers, LSTM or CNN layers, and fully connected layers, the model outputs the probability distribution of the text belonging to each category. The category with the highest probability is taken as the classification result. For example, when the input is "Mobile banking login always crashes," the model might output a probability of 0.85 for the system performance category, 0.10 for the product function category, and 0.05 for the service attitude category, ultimately classifying the text as a system performance problem.
[0059] However, the above implementation method has the following technical problems:
[0060] The aforementioned manual analysis methods face bottlenecks when processing massive amounts of questionnaire data. Business personnel need to read and manually categorize customer feedback item by item. With tens of thousands of questionnaires, this approach is not only time-consuming and labor-intensive, but also makes it difficult to complete the analysis of the entire dataset within a limited timeframe. This results in a large amount of customer feedback not being promptly transformed into a basis for service improvement. While rule-based text classification methods improve efficiency, their classification effectiveness depends entirely on the completeness of the keyword database. However, new colloquial expressions constantly emerge in customer feedback, and the rule database updates often lag behind actual needs, causing a large amount of semantically relevant feedback to be missed because the keywords are not matched.
[0061] At the semantic understanding level, traditional machine learning models (such as SVM and Random Forest) rely on statistical features like TF-IDF for text classification. These features only reflect the frequency of word occurrence and cannot understand the semantic connections between different expressions such as "interest rates are too high" and "returns are too low," leading to semantically similar texts being scattered into different categories. While small-scale deep learning models (such as LSTM and CNN) can capture some semantic information through word vectors, their ability to process financial terminology (such as liquidity risk and provision coverage ratio) and long texts remains limited. They also require a large amount of high-quality labeled data for training, lack the ability to identify rare but important customer feedback, and the fragmentation of classification results remains a prominent issue.
[0062] Furthermore, none of the aforementioned methods established an effective filtering mechanism for low-quality responses when processing questionnaire text. A large number of perfunctory responses such as "just fill in," "none," and "okay" were mixed in with valid feedback and included in the analysis process. This low-quality text not only increased unnecessary computational burden but also acted as noise, interfering with the accuracy of clustering and classification results, making it difficult for the final statistical reports to truly reflect customer opinions.
[0063] More critically, existing technical solutions all employ a one-way processing flow: text input is followed by model analysis and output. Once generated, this result cannot be corrected or optimized based on the actual judgment of business personnel. When business personnel discover classification errors or inaccurate topic summarization, these corrective suggestions cannot be fed back to the algorithm model, potentially leading to repeated occurrences of the same errors. When business needs shift from service attitude analysis to product function analysis, the system cannot adapt to the changing scenario, requiring the re-collection of training data and adjustment of the model or rules, lacking a closed-loop mechanism for continuous optimization of analysis effectiveness during use.
[0064] The feedback-based text analysis method provided in this application aims to solve the aforementioned technical problems of the prior art. By combining clustering and semantic analysis operations, unstructured text is transformed into interpretable structured topics. It acquires manually labeled data and trains or fine-tunes at least one model used in the text analysis process based on this data, adjusting the parameters corresponding to the clustering operation. The analysis is then re-executed based on the updated model and adjusted parameters, forming an iterative mechanism of manual feedback and algorithm optimization until a preset iteration termination condition is met, at which point the final analysis result is output. This solves the problems in the prior art where analysis results cannot be continuously optimized based on business feedback, and the system struggles to adapt to changes in business needs.
[0065] Optionally, the clustering operation converts text into vectors through a text embedding model, performs clustering using a density-based clustering algorithm, and marks text vectors that do not belong to any cluster as noise points to achieve automatic filtering of low-quality data.
[0066] To further improve the clustering effect, the parameters of the clustering algorithm can be determined based on the density distribution information of the text vectors, and the neighborhood radius and minimum number of points can be adjusted by calculating the variance of the inter-cluster distance, so that the clustering effect can match the distribution characteristics of different text data.
[0067] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.
[0068] First, combine Figure 1 This application introduces a text analysis method based on feedback adjustment provided in its embodiments. Figure 1 A flowchart illustrating a feedback-based text analysis method provided in this application embodiment is shown below. Figure 1 As shown, the method includes:
[0069] S101. Obtain the preprocessed text data.
[0070] Specifically, the text data obtained in this step is preprocessed questionnaire text data. Preprocessing operations include basic cleaning and effective text filtering. Basic cleaning involves cleaning the original questionnaire text, removing invalid characters, duplicate content, and blank text.
[0071] After basic text cleaning, effective text screening is performed. Specifically, the cleaned text is input into a pre-trained effective text screening model. This model is a binary classification model; its input is a single text, and its output is the classification result of whether the text is valid or invalid. The model is constructed as follows: historical questionnaire texts labeled as valid or invalid are obtained as training samples; the text of each sample is converted into a text embedding vector; the embedding vector is used as the feature, and the labeled category is used as the label to train the classifier.
[0072] As a specific implementation, the text embedding vectors can be transformed using the same BGE-M3 model (a general multilingual text embedding model, where BGE stands for BAAI general text embedding and M3 stands for multilingual, multi-granular, and multi-functional) as the subsequent clustering steps. This maps each text to a 768-dimensional embedding vector. The classifier can be logistic regression or a multilayer perceptron, with parameters optimized using the cross-entropy loss function. After training, the model can automatically determine whether any input text is valid feedback. Text deemed valid by the model is retained, while invalid text is marked as low-quality text and used as noise points for subsequent clustering operations. For example, for two customer feedback entries, "Poor service attitude, hope for improvement" and "None," the former is considered valid text, while the latter is considered invalid. The preprocessed text data serves as the basic input for subsequent analysis.
[0073] S102. Perform clustering operations on the preprocessed text data to generate at least one text cluster class.
[0074] First, the preprocessed text data is input into a text embedding model to obtain text vectors corresponding to each text. As a specific implementation, the text embedding model can be, for example, the BGE-M3 model. This model maps each input text to a fixed-dimensional numerical vector; for example, the text "poor service attitude" is mapped to a 768-dimensional vector representation of [0.32, -0.45, 0.78, ...]. Texts with similar semantics are closer together in the vector space, while texts with different semantics are farther apart.
[0075] Then, a density-based clustering algorithm is used to cluster the text vectors. As a specific implementation, the clustering algorithm can be, for example, the density-based spatial clustering of applications with noise (DBSCAN) algorithm. By setting neighborhood radius and minimum number of points, density-connected vectors are grouped into the same cluster. For example, text vectors with similar semantics such as "poor service attitude," "impatient employees," and "indifferent customer service" can be clustered into the same cluster, generating text clusters representing the service attitude theme.
[0076] Finally, text vectors that do not belong to any text cluster class are marked as noise points. These noise points correspond to low-quality text or isolated feedback unrelated to the current topic, such as "The weather is nice today" or garbled content.
[0077] S103. Perform semantic analysis on the text contained in the text cluster class to generate a category topic summary corresponding to the text cluster class.
[0078] This step organizes the text contained in each text cluster and inputs it into the large language model (LLM). Leveraging the LLM's semantic understanding and text generation capabilities, a concise and readable category topic summary is generated for each cluster. It should be noted that the large language model can be a pre-trained base model or a model fine-tuned in S104.
[0079] As a specific implementation method, LLM can adopt open-source models that are privately deployed within an enterprise, such as Qwen, Kimi, and Minimax.
[0080] Before inputting text into the LLM, the text within each cluster needs to be organized. Specifically, for each text cluster, a structured prompt word template is constructed, and the text within the cluster is filled into the prompt word template before being input into the LLM. If a cluster contains too many samples, causing the total input to exceed the LLM context window limit, the text can first be ranked by cluster score. For example, the similarity between each text and the text at the cluster center can be calculated, and the texts can be sorted from high to low similarity. The top-ranked texts (e.g., the top K texts) can then be concatenated to control the input length. The input text can be concatenated by joining the selected texts in descending order of similarity or in their original order, with each text separated by a delimiter. For example, "poor service attitude | impatient employees | indifferent customer service" can be used as input.
[0081] After receiving input text, LLM generates a summary description of the core theme of the text cluster based on pre-trained knowledge and contextual understanding capabilities. For example, for a cluster containing text such as "poor service attitude," "impatient employees," and "indifferent customer service," LLM generates a category theme summary of "customer feedback on service attitude issues"; for a cluster containing text such as "app crashes," "login lag," and "unable to receive verification codes," the generated category theme summary is "app system performance issues"; and for a cluster containing text such as "interest rates are too low," "returns are not as expected," and "financial products are not profitable," the generated category theme summary is "product return issues."
[0082] This operation transforms vector clusters that were originally only recognized by machines into structured topics that can be directly understood, facilitating subsequent review and business decisions.
[0083] S104. Obtain manually labeled data for text clusters and category topic summaries, and train or fine-tune at least one model used in the text analysis process based on the manually labeled data, and adjust the parameters corresponding to the clustering operation.
[0084] Specifically, this step involves acquiring manually annotated questionnaire data and cleaning the data. The cleaned, manually annotated data is then used for feedback training of a text analysis workflow based on clustering and LLM.
[0085] In step S101, the effective text filtering model is iteratively trained based on the aforementioned manually labeled data; in step S102, the clustering operation's algorithm parameters are adjusted based on feedback from the aforementioned manually labeled data; in step S103, the LLM is fine-tuned based on the aforementioned manually labeled data, taking into account both business benefits and training costs, using one of the following methods: supervised fine-tuning (SFT), low-rank adaptation (Lora), or instruction tuning.
[0086] The acquired and cleaned manually labeled data, along with the model training or fine-tuning based on this data, are used to adjust the parameters of the clustering operation and the LLM in subsequent steps.
[0087] S105. Based on the trained or fine-tuned model and the adjusted parameters, re-execute the clustering operation and semantic analysis operation until the preset iteration termination condition is met.
[0088] Specifically, iterative optimization is performed based on the manually labeled data obtained in S104 and the training or fine-tuning results of the model, as well as the correction instructions generated based on the data.
[0089] First, the system analyzes the feedback information contained in the correction instructions. For example, when the correction instruction is to merge two clusters, it indicates that the neighborhood radius parameter of the current clustering algorithm may be set too small, causing semantically similar texts to be scattered into different clusters; when the correction instruction is to split a cluster, it indicates that the radius parameter may be set too large, causing texts on different topics to be incorrectly merged. Based on this feedback information, the system adjusts the parameters of the clustering algorithm, such as increasing or decreasing the neighborhood radius and adjusting the minimum number of points.
[0090] Then, based on the model trained or fine-tuned in S104 (including the updated effective text filtering model and / or the fine-tuned large language model), and the adjusted clustering parameters, the clustering operation in step S102 and the semantic analysis operation in step S103 are re-executed to generate new text clusters and category topic summaries. This process is repeated until a preset iteration termination condition is met. The iteration termination condition may be, for example, that the number of correction instructions reflected in the manually labeled data is lower than a preset threshold, indicating that the current result basically meets the business requirements; or that the clustering results tend to stabilize after several consecutive iterations, and the cluster division no longer changes significantly; or that the loss function value of the model training or fine-tuning converges to a preset range.
[0091] S106. Output the text cluster class and corresponding category topic summary generated after the iteration terminates.
[0092] Once the iteration termination condition is met, the final generated text clusters and their corresponding category topic summaries are output. These outputs are presented in a structured format, such as tables or tree diagrams, showing each topic category and its typical customer feedback, for subsequent applications by the business department, such as service quality analysis and product optimization decisions.
[0093] The feedback-adjusted text analysis method provided in this application acquires manually labeled data and trains or fine-tunes at least one model used in the text analysis process based on this data, as well as adjusting the parameters corresponding to the clustering operation. The analysis is then re-executed based on the updated model and adjusted parameters, forming a closed-loop iterative mechanism. Specifically, after obtaining manually labeled questionnaire data, the system iteratively trains the effective text selection model and fine-tunes the large language model based on this data. Simultaneously, the labeled data is transformed into the basis for adjusting clustering parameters. Clustering and semantic analysis operations are re-executed based on the updated model and adjusted parameters, iterating repeatedly until a preset termination condition is met. This process allows the effective text selection capability, clustering accuracy, and topic summary generation quality to continuously improve with the accumulation of manually labeled data. It solves the technical problem in existing technologies where analysis results cannot be continuously optimized once generated, achieving adaptive improvement of text analysis results, reducing repetitive manual intervention, and improving the efficiency and accuracy of questionnaire text analysis.
[0094] In one possible implementation, the preprocessed text data is input into a text embedding model to obtain the text vector corresponding to each text.
[0095] As a specific implementation, the text embedding model can employ the BGE-M3 model, which maps each piece of input text to a 768-dimensional numerical vector. For example, the text "poor service attitude" can be mapped to a 768-dimensional vector representation of [0.32, -0.45, 0.78, ...]. Texts with similar semantics are closer together in the vector space, while texts with different semantics are farther apart.
[0096] As an alternative implementation, a multi-model fusion strategy can also be adopted. Specifically, the preprocessed text data is input into multiple different text embedding models, such as the BGE-M3 model and the Qwen3-Embedding model, to obtain the text vectors output by each model. Then, the text vectors output by each model are weighted and fused, such as by weighted concatenation or weighted averaging, to obtain the fused text vector as the final text vector corresponding to the text.
[0097] By employing the multi-model fusion strategy described above, the advantages of different text embedding models can be combined. For example, the BGE-M3 model has efficient multilingual processing capabilities, and the Qwen3-Embedding model performs well in certain semantic understanding tasks. Through weighted fusion, text vectors with richer semantic representations and stronger generalization capabilities can be generated, providing higher-quality input data for subsequent clustering operations.
[0098] In one possible implementation, the process of clustering text vectors using density-based clustering algorithms can be further improved by using an adaptive parameter adjustment mechanism.
[0099] In practice, the first step is to obtain the density distribution information of the text vectors. This density distribution information reflects the distribution characteristics of the text data in the vector space, including the Euclidean distance distribution between vectors and local density variations. For example, the distance between each pair of text vectors can be calculated, and statistical measures such as the mean and variance of the distances can be obtained. Alternatively, the probability density distribution of the vectors can be fitted using methods such as kernel density estimation.
[0100] Then, based on the acquired density distribution information, the parameters of the density-based clustering algorithm are determined. Taking the DBSCAN algorithm as an example, its core parameters include the neighborhood radius and the minimum number of points. The neighborhood radius can be determined based on the statistical characteristics of the vector distance distribution, for example, by using a certain quantile of the distance distribution (such as the median or the upper quartile) as the initial neighborhood radius; the minimum number of points can be set according to the data scale and business requirements, for example, set to one-thousandth of the total data or set between 3 and 5 based on experience.
[0101] Finally, clustering is performed based on the determined parameters. In this way, the parameters of the clustering algorithm can adaptively match the distribution characteristics of the current text data, avoiding clustering bias caused by fixed parameter settings. That is, too small parameters will produce too many fragmented clusters, while too large parameters will cause texts on different topics to be incorrectly merged.
[0102] The aforementioned adaptive parameter adjustment mechanism enables the clustering algorithm to automatically adjust parameters according to the distribution characteristics of different batches of data, thereby improving the adaptability of the clustering effect to data changes.
[0103] In one possible implementation, the process of determining the parameters of the clustering algorithm based on density distribution information can be further optimized by calculating the variance of the inter-cluster distance.
[0104] In practice, the first step is to calculate the inter-cluster distance variance based on the acquired density distribution information. The inter-cluster distance variance quantifies the dispersion of different text clusters, reflecting the dispersion or concentration of text topics in the vector space. For example, preliminary clustering of text vectors can be performed to obtain several temporary clusters. Then, the distance between the centroids of each cluster is calculated, and the variance of these distances is statistically analyzed. A larger inter-cluster distance variance indicates a more dispersed distribution of topic clusters; a smaller variance indicates a more concentrated distribution of topic clusters.
[0105] Then, based on the calculated variance of inter-cluster distances, the neighborhood radius parameter and minimum number of points parameter for the clustering algorithm are determined. The specific parameter adjustment strategy is as follows:
[0106] When the variance of inter-cluster distance is large, it indicates that the text topics are relatively dispersed. In this case, the neighborhood radius should be appropriately increased to merge texts with similar meanings but slightly different distances into the same cluster. At the same time, the minimum point parameter should be appropriately reduced to retain valuable sparse clusters. When the variance of inter-cluster distance is small, it indicates that the text topics are relatively concentrated. In this case, the neighborhood radius should be appropriately reduced to distinguish the boundaries of different topics and avoid the erroneous merging of texts with different topics. At the same time, the minimum point parameter should be appropriately increased to filter low-quality texts and outliers.
[0107] For example, when analyzing feedback related to service attitude, the text distribution may be relatively concentrated with small variance in inter-cluster distance. In this case, using a smaller neighborhood radius and a higher minimum point count can accurately distinguish subtle differences such as "poor attitude," "impatience," and "indifference." When the analysis shifts to product feature feedback, the types of issues involved in the text are more diverse (such as benefit issues, operational issues, and functional defects), and the distribution is relatively dispersed with larger variance in inter-cluster distance. In this case, expanding the neighborhood radius and reducing the minimum point count can merge semantically similar scattered texts into meaningful topic clusters.
[0108] Through the parameter adjustment mechanism based on inter-cluster distance variance, the parameter settings of the clustering algorithm can adapt to the text distribution characteristics under different business scenarios, solving the problem that traditional fixed parameter settings are difficult to adapt to data changes, and improving the accuracy and reliability of clustering results.
[0109] In one possible implementation, after generating the category topic summary in step S103, the relationships between text clusters can be further explored to reveal the evolution path of the topics.
[0110] In practice, the generated text clusters are first used as nodes, and the semantic similarity between text clusters is used as edges to construct a cluster relationship graph. Semantic similarity can be obtained by calculating the cosine similarity of the central vectors of each cluster. For example, for a cluster representing "poor service attitude" and a cluster representing "insufficient employee training", if the cosine similarity of the central vectors of the two clusters is high, then an edge is established between the two nodes, and the weight of the edge is the similarity value.
[0111] In this way, all semantically related clusters are connected to form a graph structure that reflects the topic relationships. This cluster relationship graph serves as the input to the graph neural network model and contains two parts of information: a set of nodes (each text cluster) and a set of edges (semantic relationships between clusters).
[0112] Then, the constructed cluster relationship graph is input into the graph neural network model, which learns and outputs node embedding vectors for each text cluster. As a specific implementation, the graph neural network model can, for example, employ a graph sampling and aggregation (GraphSAGE) model, which updates the vector representation of each target node by aggregating the neighboring node information. After multiple layers of iterative learning, the model outputs a fixed-dimensional node embedding vector for each text cluster. These output vectors contain not only the semantic information of the cluster's own text but also the structural position information of the cluster in the cluster relationship graph and the association information between the cluster and other adjacent clusters.
[0113] Finally, based on the node embedding vectors output by the model, the topic evolution paths between text clusters are analyzed. The position and direction of the node embedding vectors in the vector space reflect the deep connections between clusters. Through algorithms such as clustering, dimensionality reduction, or path mining, causal topic evolution chains can be discovered. For example, when the node embedding vectors of the "poor service attitude" cluster and the "insufficient employee training" cluster are spatially adjacent, and the "insufficient employee training" cluster is associated with the "service process optimization needs" cluster, the evolution path of "poor service attitude → insufficient employee training → service process optimization needs" can be revealed. In this way, the root cause of the problem can be identified. In practical application scenarios, optimization can be prioritized starting with employee training, rather than simply addressing the superficial service attitude problem.
[0114] By using the above methods, isolated topics are transformed into logically related evolutionary chains, uncovering causal relationships and development trends between problems, and providing a deeper basis for business decisions.
[0115] The electronic device provided in this application embodiment can execute the method provided in the above method embodiment. Its implementation principle and technical effect are similar, and will not be described in detail here.
[0116] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the methods in any of the above method embodiments.
[0117] All or part of the steps in the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a readable memory. When the program is executed, it performs the steps of the above method embodiments; and the aforementioned memory (storage medium) includes: read-only memory (ROM), RAM, flash memory, hard disk, solid-state drive, magnetic tape, floppy disk, optical disk, and any combination thereof.
[0118] This application describes embodiments with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processing unit of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processing unit of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0119] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0120] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0121] Obviously, those skilled in the art can make various modifications and variations to the embodiments of this application without departing from the spirit and scope of this application. Therefore, if these modifications and variations to the embodiments of this application fall within the scope of the claims of this application and their equivalents, this application also intends to include these modifications and variations.
[0122] In this application, the term "comprising" and its variations can refer to non-limiting inclusion; the term "or" and its variations can refer to "and / or". The terms "first", "second", etc., in this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. In this application, "multiple" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.
[0123] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to this application.
[0124] It should be further noted that although the steps in the flowchart are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowchart may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0125] Furthermore, unless otherwise specified, the functional units / modules in the various embodiments of this application can be integrated into one unit / module, or each unit / module can exist physically separately, or two or more units / modules can be integrated together. The integrated units / modules described above can be implemented in hardware or as software program modules.
[0126] In the above embodiments, the descriptions of each embodiment have their own emphasis. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments. The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combination of these technical features does not contradict each other, it should be considered within the scope of this specification.
[0127] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the following claims.
[0128] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
Claims
1. A text analysis method based on feedback adjustment, characterized in that, The method includes: Obtain the preprocessed text data; Clustering operations are performed on the preprocessed text data to generate at least one text cluster class; Perform semantic analysis on the text contained in the text cluster class to generate a category topic summary corresponding to the text cluster class; Obtain manually labeled data for the text clusters and the topic summaries of the categories, and train or fine-tune at least one model used in the text analysis process based on the manually labeled data, and adjust the parameters corresponding to the clustering operation; Based on the trained or fine-tuned model and the adjusted parameters, the clustering operation and the semantic analysis operation are re-executed until the preset iteration termination condition is met. Output the text cluster class and corresponding category topic summary generated after the iteration terminates.
2. The method according to claim 1, characterized in that, The clustering operation on the preprocessed text data to generate at least one text cluster class includes: The preprocessed text data is input into a text embedding model to obtain the text vector corresponding to each text in the text data. A density-based clustering algorithm is used to cluster the text vectors to generate at least one text cluster class; Text vectors that are not classified into any text cluster are marked as noise points.
3. The method according to claim 2, characterized in that, The step of inputting the preprocessed text data into a text embedding model to obtain the text vector corresponding to each text in the text data includes: The preprocessed text data is input into multiple text embedding models to obtain the text vectors output by each of the multiple text embedding models. The text vectors output by the multiple text embedding models are weighted and concatenated to obtain a fused text vector, which serves as the text vector corresponding to each text in the text data.
4. The method according to claim 2, characterized in that, The density-based clustering algorithm performs clustering operations on the text vectors, including: Obtain the density distribution information of the text vector; Based on the density distribution information, the parameters of the density-based clustering algorithm are determined; The clustering operation is performed based on the parameters.
5. The method according to claim 4, characterized in that, Determining the parameters of the density-based clustering algorithm based on the density distribution information includes: Calculate the inter-cluster distance variance based on the density distribution information; Based on the inter-cluster distance variance, the neighborhood radius parameter and minimum number of points parameter of the clustering algorithm are determined.
6. The method according to any one of claims 1-5, characterized in that, After performing semantic analysis on the text contained in the text cluster class to generate a category topic summary corresponding to the text cluster class, the method further includes: A cluster relationship graph is constructed using the text clusters as nodes and the semantic similarity between the text clusters as edges. The cluster relationship graph is input into a graph neural network model to learn the node embedding vectors of each text cluster. Based on the node embedding vectors, the topic evolution paths among the text clusters are analyzed.
7. The method according to any one of claims 1-5, characterized in that, The method further includes: The original text data is preprocessed to obtain the preprocessed text data; wherein, the preprocessing includes: marking low-quality text in the original text data, and using the marked low-quality text as noise point references for the clustering operation.
8. An electronic device, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the method as described in any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1 to 7.
10. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method of any one of claims 1 to 7.