Text data processing method and device, computer equipment and medium
By classifying and clustering the text data of macro-category assets, determining the hot topic vector cluster and generating abstracts, the problem of difficulty in identifying hot topics in the existing technology is solved, and the effect of accurately identifying and improving the quality of the abstract is achieved.
Patent Information
- Application Number
- CN202411988059.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-06
AI Technical Summary
In the prior art, it is difficult to identify hot topics of text data of macro-category assets, making manual evaluation more difficult.
By obtaining the list of opinions of pending text data, classifying asset classes, and obtaining the set of opinions; then clustering the list of opinions vectors, determining the cluster of hot topic vectors, and generating a summary based on the target opinions.
Accurate identification of hot topics in text data of macro-category assets has been achieved, and the quality and credibility of view summary have been improved.
Smart Images

Figure CN119938906A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a text data processing method, device, computer equipment and medium. Background Art
[0002] The analysis of hot topics in macro asset classes aims to help investors understand market trends, formulate investment strategies, avoid risks, and capture market sentiment and changes through forecasts of economic conditions and trends of various assets (such as stocks, bonds, crude oil, gold, etc.), which is of great value to investors.
[0003] In the related art, hot topics of macro asset class text data are usually identified through manual evaluation. However, with the rapid development of the Internet and social media, text data about macro asset classes spreads rapidly on these platforms, and the amount of information increases dramatically, making it increasingly difficult to manually evaluate hot topics.
[0004] Therefore, it is urgent to propose a new text data processing method. Summary of the invention
[0005] The present application provides a text data processing method, apparatus, computer equipment and medium, which solve the technical problem of difficulty in identifying hot topics in macro-asset text data in related technologies, and achieve the technical effect of accurately identifying hot topics in macro-asset text data.
[0006] In order to achieve the above objectives, the main technical solutions adopted in this application include:
[0007] In a first aspect, an embodiment of the present application provides a text data processing method, the method comprising:
[0008] Obtain a list of viewpoints corresponding to a plurality of text data to be processed;
[0009] Classifying the opinion list of each to-be-processed text data according to the asset category to obtain an opinion set for each asset category; wherein the opinion set corresponds to an opinion vector list;
[0010] Performing clustering processing on the opinion vector list to obtain a plurality of opinion vector clusters, and determining a hot topic vector cluster in the plurality of opinion vector clusters according to the number of vectors in the opinion vector clusters; wherein the hot topic vector cluster corresponds to a target opinion in the opinion set;
[0011] A summary is generated based on the target viewpoint to obtain a core topic name and a viewpoint summary of the target viewpoint.
[0012] Optionally, obtaining a list of viewpoints corresponding to a plurality of text data to be processed includes:
[0013] Acquire the plurality of text data to be processed and a first prompt word template;
[0014] For each text data to be processed, the large model prompt engineering technology is used to extract attribute-level opinions through the first prompt word template to obtain a list of opinions for each text data to be processed.
[0015] Optionally, the prompt words in the first prompt word template are customized based on asset categories and analysis angles.
[0016] Optionally, a list of viewpoint vectors corresponding to the viewpoint set is obtained in the following manner:
[0017] The opinions in the opinion set are vectorized and represented by a vectorized embedding model to obtain a vector list corresponding to the opinion set.
[0018] Optionally, clustering the opinion vector list to obtain a plurality of opinion vector clusters includes:
[0019] The opinion vector list is clustered by using an agglomerative hierarchical clustering algorithm to obtain the plurality of opinion vector clusters.
[0020] Optionally, determining a hot topic vector cluster from among the plurality of opinion vector clusters according to the number of vectors in the opinion vector cluster includes:
[0021] According to the number of vectors of the opinion vector clusters, determining k opinion vector clusters having the largest number of vectors among the multiple opinion vector clusters;
[0022] The k opinion vector clusters are determined as the hot topic vector clusters.
[0023] Optionally, the generating a summary based on the target viewpoint to obtain the core topic name and viewpoint summary of the target viewpoint includes:
[0024] By using the large model prompt engineering technology, the target viewpoint is summarized through the second prompt word template to obtain the core topic name and viewpoint summary of the target viewpoint.
[0025] In a second aspect, an embodiment of the present application provides a text data processing device, the device comprising:
[0026] A viewpoint list acquisition module is used to acquire a viewpoint list corresponding to a plurality of text data to be processed;
[0027] A classification module, used to classify the opinion list of each to-be-processed text data according to the asset category, and obtain an opinion set for each asset category; wherein the opinion set corresponds to an opinion vector list;
[0028] A clustering module, configured to perform clustering processing on the opinion vector list to obtain a plurality of opinion vector clusters, and determine a hot topic vector cluster from the plurality of opinion vector clusters according to the number of vectors in the opinion vector clusters; wherein the hot topic vector cluster corresponds to a target opinion in the opinion set;
[0029] The summary generation module is used to generate a summary based on the target viewpoint to obtain the core topic name and viewpoint summary of the target viewpoint.
[0030] In a third aspect, an embodiment of the present application provides a computer device, including:
[0031] A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the method described in any of the above embodiments by executing the computer instructions.
[0032] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which computer instructions are stored, and the computer instructions are used to enable a computer to execute the method described in any of the above embodiments.
[0033] In the embodiment of the present application, first, a list of viewpoints corresponding to a plurality of text data to be processed is obtained; secondly, the list of viewpoints of each text data to be processed is classified according to the asset category to obtain a set of viewpoints for each asset category, and the viewpoint set corresponds to a list of viewpoint vectors; then, the viewpoint vector list is clustered to obtain a plurality of viewpoint vector clusters, and a hot topic vector cluster is determined in the plurality of viewpoint vector clusters according to the number of vectors in the viewpoint vector clusters, and the hot topic vector cluster corresponds to a target viewpoint in the viewpoint set; finally, a summary is generated based on the target viewpoint to obtain the core topic name and viewpoint summary of the target viewpoint. By determining the hot topics by the number of vectors in the viewpoint vector clusters obtained by clustering, an accurate assessment of the popularity of the hot topics is achieved, thereby improving the quality and credibility of the viewpoint summary. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to more clearly illustrate the specific implementation methods of the present application or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0035] Figure 1a A flowchart of a text data processing method provided in an embodiment of this specification;
[0036] Figure 1bA schematic diagram of a text data processing method provided in an embodiment of this specification;
[0037] Figure 2 A flowchart of a text data processing method provided in an embodiment of this specification;
[0038] Figure 3 A schematic diagram of a text data processing device provided in an embodiment of this specification;
[0039] Figure 4 A schematic diagram of the structure of a computer device provided in an embodiment of this specification. DETAILED DESCRIPTION
[0040] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present application.
[0041] The analysis of macro-asset classes usually requires extracting key hot topic information from massive text data and summarizing it effectively. This information is crucial for investors because refined hot information can help them quickly understand market trends, identify potential risks and opportunities, and make more informed investment decisions.
[0042] In related technologies, some directly use vector models to vectorize the split text data and then analyze it. This method is prone to losing a lot of information because for vector models of the same dimension, the longer the input text data, the lower its compression rate, resulting in less sufficient information reflected in the original text data. Some also lack quantitative evaluation methods for the popularity of hot topics in text data, resulting in the generated hot topic summary may not accurately reflect the actual popularity distribution in the text data, thus affecting the quality and credibility of the summary.
[0043] Based on this, the present application provides a method for processing text data: first, obtain a list of opinions corresponding to multiple text data to be processed; second, classify the list of opinions for each text data to be processed according to the asset category, and obtain a set of opinions for each asset category, and the set of opinions corresponds to a list of opinion vectors; then, cluster the list of opinion vectors to obtain multiple opinion vector clusters, and determine the hot topic vector clusters in the multiple opinion vector clusters according to the number of vectors in the opinion vector clusters, and the hot topic vector clusters correspond to the target opinion in the opinion set; finally, generate a summary based on the target opinion to obtain the core topic name and opinion summary of the target opinion. Compared with the method of vectorizing the split text data in the related art, this method obtains the list of opinions corresponding to the text data, which can better reflect the information in the original text data. By determining the hot topics by the number of vectors in the opinion vector clusters obtained by clustering, an accurate assessment of the popularity of the hot topics is achieved, thereby improving the quality and credibility of the opinion summary.
[0044] According to an embodiment of the present application, an embodiment of a text data processing method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0045] See also Figure 1a and Figure 1b In this embodiment, a text data processing method is provided, the steps of which include:
[0046] S110: Obtain a viewpoint list corresponding to a plurality of text data to be processed.
[0047] The text data to be processed may be an article analyzing macro-asset categories published by market institutions and professionals in the financial field through social media platforms (such as WeChat public accounts, Weibo, financial websites, etc.).
[0048] The opinion list may be a list of opinions on macro-class assets in each text data to be processed, wherein an opinion may be an insight, analysis or view on a certain asset class.
[0049] S120 , classifying the opinion list of each to-be-processed text data according to the asset category to obtain an opinion set for each asset category.
[0050] The opinion set corresponds to a list of opinion vectors. Asset categories can be investable targets such as stocks, crude oil, gold, bonds, and commodities.
[0051] The opinion set can be obtained by aggregating the opinion lists in multiple text data to be processed according to the asset categories involved in the opinions. For example, all opinions about crude oil in multiple text data to be processed are aggregated into the crude oil opinion set. The advantage of this is that by focusing on analyzing the opinion set, we can gain an in-depth understanding of the market participants' views, expectations and potential risks on crude oil, thereby providing more targeted support for related investment decisions. The opinions in the opinion set are vectorized to obtain a opinion vector list corresponding to the opinion set.
[0052] S130 , clustering the opinion vector list to obtain a plurality of opinion vector clusters, and determining a hot topic vector cluster among the plurality of opinion vector clusters according to the number of vectors in the opinion vector clusters.
[0053] Among them, the hot topic vector cluster corresponds to a target viewpoint in the viewpoint set. The target viewpoint is the viewpoint corresponding to the viewpoint vector in the hot topic vector cluster. Since the viewpoint vector and the viewpoint are in a one-to-one correspondence, after obtaining the hot topic vector cluster, the target viewpoint corresponding to each hot topic vector cluster can be obtained through this correspondence.
[0054] S140, generating a summary based on the target viewpoint to obtain the core topic name and viewpoint summary of the target viewpoint.
[0055] The core topic name can be the most representative and highly summarized topic label extracted from the target viewpoint after analysis and extraction, which can accurately reflect the core content discussed in the target viewpoint, usually a concise word or phrase. The viewpoint summary can be a refinement and summary of the target viewpoint, which can retain the core information of the original text while removing redundant parts, and present the main content in a concise and easy-to-understand way.
[0056] Specifically, the text data processing method first obtains multiple text data to be processed, and then uses a large language model (such as DeepSeek-V2.5-1210 or Qwen2.5-72B) to extract the opinions in each text data to be processed using prompt word engineering technology to obtain a list of opinions in each text data to be processed, and each list of opinions in the text data to be processed may involve different asset categories. Compared with the related art that uses Map-reduce to segment text data according to the longest context length processed by the large language model, which is prone to abnormal sentences and semantic loss, this method does not need to segment text data, can ensure the coherence and integrity of opinions, and avoid semantic errors caused by segmentation.
[0057] Furthermore, since the opinion list of each text data to be processed may involve different asset categories, in order to analyze each asset category, the opinion list of each text data to be processed can be classified according to the asset category to obtain an opinion set for each asset category, such as a crude oil opinion set, a Hong Kong stock opinion set, a gold opinion set, etc.
[0058] Furthermore, the opinions in the opinion set are vectorized through a vector embedding model (such as Tao-8k) to obtain a corresponding opinion vector list. In order to identify hot topics in each asset category, the opinion vector list corresponding to each asset category can be clustered to obtain multiple opinion vector clusters. Since each opinion vector cluster corresponds to a topic, the popularity of the topic can be measured by the number of vectors in the opinion vector cluster. By determining the top N opinion vector clusters with the largest number of vectors, the hot topic vector cluster can be identified. Since opinion vectors correspond to opinions one by one, according to this correspondence, the hot topic vector cluster can be mapped to the corresponding target opinion in natural language.
[0059] Finally, a large language model (such as DeepSeek-V2.5-1210 or Qwen2.5-72B) is used again to generate a summary of the target opinion corresponding to each hot topic vector cluster using prompt word engineering technology to obtain the core topic name and opinion summary of the target opinion.
[0060] In the above embodiment, first, a list of viewpoints corresponding to a plurality of text data to be processed is obtained; secondly, the list of viewpoints of each text data to be processed is classified according to the asset category to obtain a set of viewpoints for each asset category, and the viewpoint set corresponds to a list of viewpoint vectors; then, the viewpoint vector list is clustered to obtain a plurality of viewpoint vector clusters, and a hot topic vector cluster is determined in the plurality of viewpoint vector clusters according to the number of vectors in the viewpoint vector cluster, and the hot topic vector cluster corresponds to a target viewpoint in the viewpoint set; finally, a summary is generated based on the target viewpoint to obtain the core topic name and viewpoint summary of the target viewpoint. Compared with the method of vectorizing the split text data in the related art, this method obtains a list of viewpoints corresponding to the text data, which better reflects the information in the original text data. The hot topics are determined by the number of viewpoint vectors in the viewpoint vector cluster obtained by clustering, and an accurate assessment of the popularity of the hot topics is achieved, thereby improving the quality and credibility of the viewpoint summary.
[0061] See also Figure 2 In some embodiments, obtaining a list of viewpoints corresponding to a plurality of text data to be processed includes:
[0062] S210, obtaining a plurality of text data to be processed and a first prompt word template;
[0063] S220 , for each text data to be processed, using the large model prompt engineering technology, extracting attribute-level opinions through the first prompt word template, and obtaining a list of opinions for each text data to be processed.
[0064] The first prompt word template is used to extract a viewpoint list from the text data to be processed. Exemplarily, the content of the first prompt word template is as follows:
[0065] "###role play
[0066] You are an excellent macroeconomic analyst, especially good at reading and analyzing macroeconomic market analysis reports.
[0067] ###Quest Requirements
[0068] The following market analysis report involves the description and analysis of the current status of macro asset classes. Please use the format of the table below to summarize the relevant information and opinions of each asset class from the following perspectives.
[0069] {Enter analysis angle here}
[0070] ###Notes
[0071] {Enter notes here}
[0072] Output format
[0073] |Asset Classes|Summary of Views|
[0074] {Enter the asset class of interest here}
[0075] ###Article Content
[0076] {Enter the text data to be processed here}".
[0077] Attribute-level perspectives can be perspectives related to asset attributes, which can be risk attributes (such as price volatility), income attributes (such as yield, appreciation potential), macroeconomic-related attributes (such as inflation), etc. In macro asset analysis, the overall-level perspective focuses on the overall or comprehensive evaluation of assets or markets, while the attribute-level perspective focuses on analyzing individual attributes from a more detailed perspective. The attribute-level perspective helps investors more accurately understand the impact of various factors on asset performance by deeply studying the key factors of assets. In addition, the attribute-level perspective can reveal potential risks and opportunities, enabling investors to make more scientific and rational decisions in a complex market environment.
[0078] Specifically, first, the relevant information required for analysis can be input into the first prompt word template, including the asset category of interest and the analysis angle, where the analysis angle can be the asset attribute, so as to extract the attribute-level viewpoint of the corresponding asset category; then the text data to be processed (usually an article) is input into the first prompt word template to generate the first prompt word corresponding to the text data to be processed; finally, the first prompt word is input into the large language model, and the large language model outputs the viewpoint list of the text data to be processed. Multiple text data to be processed are processed in sequence to obtain the viewpoint list of each text data to be processed.
[0079] In some embodiments, the prompt words in the first prompt word template are customized based on asset categories and analysis angles.
[0080] Specifically, the asset categories and analysis angles can be customized according to the needs or research objectives of a specific field. By designing the first prompt word, the analysis process of the large language model can be accurately guided, thereby achieving in-depth analysis of the needs or research objectives of a specific field. This customized analysis method not only focuses on core issues more specifically, but also improves the relevance and practicality of the analysis results, better meeting specific research needs.
[0081] In some embodiments, a list of opinion vectors corresponding to the opinion set is obtained in the following manner: the opinions in the opinion set are vectorized by using a vectorized embedding model to obtain a list of vectors corresponding to the opinion set.
[0082] Opinion analysis usually requires identifying hot topics in each asset class, which can be achieved by converting the opinions in the opinion set into vectors and then using clustering for analysis. Therefore, it is necessary to vectorize the opinions in the opinion set and obtain a list of opinion vectors corresponding to the opinion set.
[0083] Specifically, the opinions in the opinion set can be input into a vector embedding model (such as Tao-8k) in sequence. The vector embedding model converts the opinions into vector representations in a high-dimensional space to obtain a list of opinion vectors.
[0084] In some embodiments, clustering the opinion vector list to obtain a plurality of opinion vector clusters includes: clustering the opinion vector list using an agglomerative hierarchical clustering algorithm to obtain a plurality of opinion vector clusters.
[0085] Among them, the agglomerative hierarchical clustering algorithm starts from each opinion vector, regards each opinion vector as a separate cluster, and then gradually merges the most similar clusters until the clustering categories reach a given number and stops clustering. It can also stop clustering when the similarity of the most similar vector pairs is less than a preset threshold.
[0086] Specifically, first, obtain a list of opinion vectors as an input data set for the agglomerative hierarchical clustering algorithm. Then, calculate the similarity between all opinion vectors, merge the most similar opinion vector pairs, and the similarity metric can be cosine similarity or Euclidean distance. The merged vector can be used as a new clustering category. In order to represent this new clustering category, the mean vector of the two vectors can be used to represent the clustering category, which is used to continue clustering until the similarity of the most similar vector pair is less than a preset threshold or the clustering category reaches a given number of categories, stop clustering, and return multiple opinion vector clusters. It should be noted that the clustering operation needs to pay attention to the impact of noise on the clustering results. For example, after the clustering is completed, the point whose distance from the center point of the nearest opinion vector cluster is greater than a certain threshold can be regarded as noise and removed from the opinion vector cluster. In other embodiments, noise points are identified and marked by an external noise detection method, such as DBSCAN (density-based clustering algorithm). DBSCAN determines whether a data point belongs to a cluster by density, and marks points without sufficient neighborhood density as noise. DBSCAN can be used to perform preliminary processing on the opinion vector list to identify noise, and then the remaining points are input into the agglomerative hierarchical clustering algorithm for clustering.
[0087] It should be noted that in the embodiment of the present application, the hot topics are determined in combination with the number of vectors in the opinion vector cluster. It can be seen that the accuracy of the clustering results of the opinion vector list is very important for accurately evaluating the hot topics. Therefore, in the embodiment of the present application, it is necessary to break the clustering performance bottleneck caused by the preset number of clusters in the related technology and accurately cluster the opinion vector list. First, the number of clusters k is pre-specified in combination with the number of historical hot topics. Secondly, the opinion vector list is hard divided through a partitioning clustering algorithm to obtain k initial opinion vector clusters.
[0088] Furthermore, considering that the clustering results obtained by the partitioning clustering algorithm are determined based on historical situations, and there may be differences between historical situations and current actual situations, it is necessary to verify the number of hard partition clusters k. Specifically, the opinion vector cluster to be verified is determined in the k initial opinion vector clusters according to the number of opinion vectors in the initial opinion vector cluster. For any opinion vector cluster to be verified, a hierarchical clustering algorithm (such as an agglomerative hierarchical clustering algorithm) is used to soft-partition the opinion vector cluster to be verified, and multiple first opinion vector clusters are obtained.
[0089] Further, the initial opinion vector cluster whose number of opinion vectors exceeds the preset number threshold is determined as the opinion vector cluster to be verified. The initial opinion vector cluster whose number of opinion vectors does not exceed the preset number threshold is determined as the second opinion vector cluster; and the final clustering result is constructed based on the multiple second opinion vector clusters and the multiple first opinion vector clusters, that is, the multiple opinion vector clusters obtained by clustering the opinion vector list.
[0090] In the above embodiment, considering that the amount of information increases sharply, the text data to be processed belongs to a large-scale data set. Correspondingly, the scale of the opinion vector list is also large. In order to improve the data processing efficiency, a partition type clustering method (such as K-means) is used to obtain the initial opinion vector cluster. On the one hand, compared with the opinion vector list before clustering, the initial opinion vector cluster belongs to a small data set or a small-scale data set. On the other hand, the accuracy of the cluster number k needs to be improved. Therefore, in order to improve the accuracy of data processing, a hierarchical clustering algorithm is used to soft-partition the opinion vector cluster to be verified, so as to find the internal hierarchical structure of the opinion vector cluster to be verified, thereby completing the verification of the opinion vector cluster to be verified, and obtaining a more accurate number of opinion vector clusters, thereby ensuring that the number of vectors in the opinion vector cluster is credible, and providing a good data basis for determining hot topics based on the number of vectors in the opinion vector cluster.
[0091] In some embodiments, based on the number of vectors in the opinion vector clusters, a hot topic vector cluster is determined among multiple opinion vector clusters, including: based on the number of vectors in the opinion vector clusters, determining k opinion vector clusters with the largest number of vectors among multiple opinion vector clusters; and determining the k opinion vector clusters as hot topic vector clusters.
[0092] Specifically, through clustering, multiple opinion vector clusters are obtained, each opinion vector cluster contains several opinion vectors, and these opinion vectors represent the opinions extracted from the text data to be processed. In order to identify the most representative hot topics, the number of opinion vectors in each opinion vector cluster is first counted. Then, all opinion vector clusters are sorted according to the number of opinion vectors, and the top k clusters with the largest number are selected as hot topic vector clusters. Exemplarily, there are 5 opinion vector clusters, containing 10, 15, 20, 5 and 30 opinion vectors respectively. According to the sorting of the number of opinion vectors, the top k = 3 opinion vector clusters are opinion vector clusters containing 30, 20 and 15 opinion vectors respectively. These three opinion vector clusters are determined to be hot topic vector clusters. They represent the most important topics, and the subsequent opinion summary will be further performed based on the target opinions corresponding to these three hot topic vector clusters.
[0093] In the above embodiment, the hot topics are determined by the number of opinion vectors in the opinion vector clusters obtained by clustering, which achieves accurate evaluation of the hot topics and makes the hot topic analysis more accurate.
[0094] In some embodiments, a summary is generated based on the target viewpoint to obtain the core topic name and viewpoint summary of the target viewpoint, including:
[0095] By using the large model prompt engineering technology, the target viewpoint is summarized through the second prompt word template to obtain the core topic name and viewpoint summary of the target viewpoint.
[0096] The second prompt word template is used to extract the core topic name from the target viewpoint and summarize the viewpoint. For example, the content of the second prompt word template is as follows:
[0097] "###role play
[0098] You are an excellent macroeconomic analyst, especially good at extracting core themes and summarizing opinions from a large number of opinions.
[0099] ###Quest Requirements
[0100] Below are the core views of market analysts on {asset class}. Please extract the themes of these views and summarize them.
[0101] ###Notes
[0102] {Enter notes here}
[0103] Output format
[0104] {"Opinion topic":"","Opinion summary":""}
[0105] ###Market View
[0106] {Enter target viewpoints line by line here}".
[0107] Specifically, first obtain the target viewpoint corresponding to a hot topic of a certain asset category, and enter the asset category and precautions in the second prompt word template; then enter the target viewpoint into the second prompt word template to generate the second prompt word corresponding to the hot topic; finally, input the second prompt word into the large language model, and the large language model outputs the core topic name and viewpoint summary of the target viewpoint. Further, the target viewpoints corresponding to multiple hot topics of a certain asset category are processed in sequence to obtain all the core topic names and corresponding viewpoint summaries of the asset category. Further, the core topic names and viewpoint summaries of multiple asset categories extracted from multiple text data to be processed are summarized to form a complete viewpoint summary of the research field.
[0108] See also Figure 3 In this embodiment, a text data processing device 600 is also provided. The text data processing device 600 includes:
[0109] The viewpoint list acquisition module 610 is used to acquire a viewpoint list corresponding to a plurality of text data to be processed;
[0110] A classification module 620 is used to classify the opinion list of each to-be-processed text data according to the asset category to obtain an opinion set for each asset category; wherein the opinion set corresponds to an opinion vector list;
[0111] The clustering module 630 is used to perform clustering processing on the opinion vector list to obtain multiple opinion vector clusters, and determine a hot topic vector cluster in the multiple opinion vector clusters according to the number of vectors in the opinion vector clusters; wherein the hot topic vector cluster corresponds to a target opinion in the opinion set;
[0112] The summary generation module 640 is used to generate a summary based on the target viewpoint to obtain the core topic name and viewpoint summary of the target viewpoint.
[0113] In some implementations, the viewpoint list acquisition module 610 further includes:
[0114] A data and module acquisition unit, used for acquiring a plurality of text data to be processed and a first prompt word template;
[0115] The viewpoint extraction unit is used to extract attribute-level viewpoints for each text data to be processed by using the large model prompt engineering technology and the first prompt word template to obtain a viewpoint list for each text data to be processed.
[0116] In some embodiments, the viewpoint list acquisition module 610 is further used to set prompt words, and the prompt words in the first prompt word template are customized based on asset categories and analysis angles.
[0117] In some implementations, the classification module 620 further includes:
[0118] The vectorization unit is used to vectorize the opinions in the opinion set through a vectorized embedding model to obtain a vector list corresponding to the opinion set.
[0119] In some implementations, the clustering module 630 further includes:
[0120] The clustering unit is used to cluster the opinion vector list through an agglomerative hierarchical clustering algorithm to obtain multiple opinion vector clusters.
[0121] In some implementations, the clustering module 630 further includes:
[0122] a vector cluster determination unit, configured to determine k opinion vector clusters having the largest number of vectors among the plurality of opinion vector clusters according to the number of vectors of the opinion vector clusters;
[0123] The hotspot vector cluster determination unit is used to determine k opinion vector clusters as hotspot topic vector clusters.
[0124] In some implementations, the summary generation module 640 further includes:
[0125] The summary generation unit is used to generate a summary of the target viewpoint by using the large model prompt engineering technology through the second prompt word template to obtain the core topic name and viewpoint summary of the target viewpoint.
[0126] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.
[0127] The text data processing device in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.
[0128] See also Figure 4 , Figure 4 is a schematic diagram of the structure of a computer device provided in an embodiment of the present application, such as Figure 4 As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components are connected to each other using different buses for communication, and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the computer device, including instructions stored in or on the memory to display the graphical information of the GUI on an external input / output device (such as, a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 4 A processor 10 is taken as an example.
[0129] The processor 10 may be a central processing unit, a network processor or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.
[0130] The memory 20 stores instructions executable by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiment.
[0131] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely arranged relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0132] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 20 may also include a combination of the above types of memory.
[0133] The computer device also includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30 and the output device 40 may be connected via a bus or other means. Figure 4 The example of connecting through bus is taken in the following.
[0134] The input device 30 can receive input digital or character information, and generate key signal input related to the user settings and function control of the computer device, such as a touch screen, a keypad, a mouse, a track pad, a touch pad, an indicator bar, one or more mouse buttons, a trackball, a joystick, etc. The output device 40 may include a display device, an auxiliary lighting device (e.g., an LED) and a tactile feedback device (e.g., a vibration motor), etc. The above-mentioned display device includes but is not limited to a liquid crystal display, a light emitting diode, a display and a plasma display. In some optional embodiments, the display device can be a touch screen.
[0135] The embodiment of the present application also provides a computer-readable storage medium. The above method according to the embodiment of the present application can be implemented in hardware, firmware, or can be implemented as a computer code that can be recorded in a storage medium, or can be implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and will be stored in a local storage medium through a network download, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state hard disk, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor, or hardware, the method shown in the above embodiment is implemented.
[0136] The embodiment of the present application provides a computer program product, which includes computer instructions, which are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method of any embodiment of the present application.
[0137] Although the embodiments of the present application are described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present application, and such modifications and variations are all within the scope defined by the appended claims.
[0138] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0139] For the convenience of description, the above device is described in various units according to their functions. Of course, when implementing the present application, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0140] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.
[0141] The present application is described with reference to the flowchart and / or block diagram of the method, device (system) and computer program product according to the embodiment of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, and the combination of the process and / or box in the flowchart and / or block diagram can be realized by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for realizing the function specified in one process or multiple processes in the flowchart and / or one box or multiple boxes in the block diagram.
[0142] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0143] These computer program instructions may also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0144] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0145] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0146] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.
[0147] Although the embodiments of the present application have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present application, and such modifications and variations are all within the scope defined by the appended claims.
Claims
1. A text data processing method, characterized in that: The method comprises: Obtain a list of viewpoints corresponding to a plurality of text data to be processed; Classifying the opinion list of each to-be-processed text data according to the asset category to obtain an opinion set for each asset category; wherein the opinion set corresponds to an opinion vector list; Performing clustering processing on the opinion vector list to obtain a plurality of opinion vector clusters, and determining a hot topic vector cluster in the plurality of opinion vector clusters according to the number of vectors in the opinion vector clusters; wherein the hot topic vector cluster corresponds to a target opinion in the opinion set; A summary is generated based on the target viewpoint to obtain a core topic name and a viewpoint summary of the target viewpoint.
2. The method according to claim 1, characterized in that The step of obtaining a list of viewpoints corresponding to a plurality of text data to be processed includes: Acquire the plurality of text data to be processed and a first prompt word template; For each text data to be processed, the large model prompt engineering technology is used to extract attribute-level opinions through the first prompt word template to obtain a list of opinions for each text data to be processed.
3. The method according to claim 2, characterized in that The prompt words in the first prompt word template are customized based on asset categories and analysis angles.
4. The method according to claim 1, characterized in that Get the opinion vector list corresponding to the opinion set in the following way: The opinions in the opinion set are vectorized and represented by a vectorized embedding model to obtain a vector list corresponding to the opinion set.
5. The method according to claim 1, characterized in that The clustering process is performed on the opinion vector list to obtain a plurality of opinion vector clusters, including: The opinion vector list is clustered by using an agglomerative hierarchical clustering algorithm to obtain the plurality of opinion vector clusters.
6. The method according to claim 1, characterized in that The step of determining a hot topic vector cluster from among the plurality of opinion vector clusters according to the number of vectors of the opinion vector clusters comprises: According to the number of vectors of the opinion vector clusters, determining k opinion vector clusters having the largest number of vectors among the multiple opinion vector clusters; The k opinion vector clusters are determined as the hot topic vector clusters.
7. The method according to claim 6, characterized in that The generating of a summary based on the target viewpoint to obtain the core topic name and viewpoint summary of the target viewpoint includes: By using the large model prompt engineering technology, the target viewpoint is summarized through the second prompt word template to obtain the core topic name and viewpoint summary of the target viewpoint.
8. A text data processing device, characterized in that: The device comprises: A viewpoint list acquisition module is used to acquire a viewpoint list corresponding to a plurality of text data to be processed; A classification module, used to classify the opinion list of each to-be-processed text data according to the asset category, and obtain an opinion set for each asset category; wherein the opinion set corresponds to an opinion vector list; A clustering module, configured to perform clustering processing on the opinion vector list to obtain a plurality of opinion vector clusters, and determine a hot topic vector cluster in the plurality of opinion vector clusters according to the number of vectors in the opinion vector clusters; wherein the hot topic vector cluster corresponds to a target opinion in the opinion set; The summary generation module is used to generate a summary based on the target viewpoint to obtain the core topic name and viewpoint summary of the target viewpoint.
9. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the method according to any one of claims 1 to 7 by executing the computer instructions.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the method according to any one of claims 1 to 7.
Citation Information
Cited By
File access cold and hot degree calculation method and system based on large model and medium
CN120670335A