Industrial risk identification method based on hotspot clustering algorithm and industrial large model agent
Patent Information
- Application Number
- CN202510294547.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-13
AI Technical Summary
The existing technology cannot accurately identify hot topics involving industrial risks from massive industrial text data and generate industrial development suggestions.
Using a method based on hot-spot clustering algorithm and industrial big model agent, we identify hot topics related to industry risks through semantic vectorization, connectivity graph clustering, and industry-related judgments through industry-related methods, and generate industry suggestions based on big models.
It has achieved accurate identification of industrial risks and generated industrial development suggestions, and improved the ability to discover and identify useful information in massive data.
Smart Images

Figure CN120218612A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of industrial risk identification, and particularly relates to an industrial risk identification method based on a hot topic clustering algorithm and an industrial large model agent. Background Art
[0002] A large language model (LLM) is a neural network model based on deep learning, usually having billions or even hundreds of billions of parameters. These models are obtained through pre-training using self-supervised learning or semi-supervised learning on large-scale text data. Since they have learned the grammar, semantics, and context information of the text from massive data, they can perform excellently in natural language processing (NLP) tasks and are widely used in multiple fields such as text generation, machine translation, question answering systems, and sentiment analysis.
[0003] A large amount of hidden information is hidden in the massive text data on the Internet. How to timely and accurately discover the required information from the massive data is a problem that researchers are committed to solving. Through topic recognition and clustering algorithms, variable and messy text data (such as news data and forum data) can be aggregated together to form hot topics, thereby discovering useful information hidden in the data. Summary of the Invention
[0004] In order to solve the technical problem that the prior art cannot accurately screen out hot topics related to industrial risks from massive industrial text data and summarize industrial development suggestions, the present invention proposes an industrial risk identification method based on a hot topic clustering algorithm and an industrial large model agent, which realizes the accurate identification of industrial risks and gives industrial development suggestions.
[0005] The specific solution is as follows: An industrial risk identification method based on a hot topic clustering algorithm and an industrial large model agent, S1, constructing an industrial data set: constructing an industrial data set through an industrial database and scraping online data; the industrial data set reflects the characteristics of a specified industry and contains data with negative views; S2, hot topic clustering; performing semantic vectorization processing on the data in the industrial data set described in S1, and performing similarity calculation to evaluate data correlation; performing clustering analysis on the data through a connected graph, and optimizing the clustering result through splitting subgraphs; performing reverse sorting on the clusters in the clustering result based on the number of members in the class to obtain industrial hot topic texts; S3, screening industrial risk hot topics: constructing an industrial large model agent based on an industrial knowledge base and a large model and configuring tool components, inputting the industrial hot topic texts into the industrial large model agent, and running the industrial large model agent to perform sentiment analysis and industry-related judgment to obtain hot topics related to industrial risks; S4, Industrial Recommendation Generation: Based on large models, generate industrial recommendations for industrial risk hotspots.
[0006] Preferably, the method for semantic vectorization of data in step S2 is as follows: S21, Model Construction: Construct a neural network model with a GPT architecture. The neural network model includes at least one Transformer decoder layer and is configured with a position-weighted average pooling module; S22, Parameter Optimization: Use the bias tensor contrast fine-tuning method to optimize the parameters of the neural network model. Introduce a learnable bias tensor into the loss function for similarity comparison calculation; S23, Generate Semantic Vectors: Traverse all text elements Ti in the text collection T to be analyzed. Through the neural network model, perform semantic encoding on each Ti to generate the corresponding semantic vector Vi; S24, Construct a Semantic Vector Set: Aggregate the semantic vectors Vi corresponding to all text elements Ti to construct a structured semantic vector set V = {Vi|Ti ∈ T}.
[0007] Preferably, the method for similarity calculation in step S2 is as follows: Traverse any two vector combinations in the semantic vector set V, and calculate the cosine value between any two vectors through the vector cosine calculation formula. According to the incoming similarity threshold and the magnitude of the cosine value between vectors, determine whether there is a relevant relationship between any two vectors. If the cosine value exceeds the threshold, it is determined that there is a correlation.
[0008] Preferably, the method for using correlation to achieve data clustering analysis through a connected graph in step S2 is as follows: S2a, Construct a Topological Graph: Use the text data in the industrial dataset as vertices to construct a vertex set. Use the correlation between text data as the basis for whether there is an edge between vertices to construct an adjacency matrix. Use the cosine value between text data as the weight of the edge to construct a weighted undirected graph G of the dataset T based on text similarity relationships; S2b, Connected Graph Clustering: The vertices that are connected to each other in the weighted undirected graph G have semantic relevance; According to the connectivity between the vertices of the weighted undirected graph G, divide the graph G into multiple connected graphs. Each connected graph can be regarded as a cluster, and the vertices Ti included in the connected graph are the class member objects included in this cluster; S2c, Split Optimization: If the scale of the class exceeds the set range, remove the relationships with lower weights and re-split the subgraphs according to the connectivity of the graph to obtain new clusters.
[0009] Preferably, in step 2, there is semantic correlation among the data of each cluster in the clustering result, and the number of class members, that is, the amount of data contained in the cluster, reflects the popularity of the corresponding topic; the clusters are sorted in reverse order according to the number of members of the class, and the clusters before the threshold are set as the hot topics included in the industrial dataset and output as industrial hot texts.
[0010] Preferably, in step 3, the method for constructing the industrial large model agent is as follows: S31. Construct an industrial knowledge base: The data constituting the industrial knowledge base mainly includes, but is not limited to, industrial-related policies and regulations, industrial knowledge graphs, and industrial historical news data; vectorized feature extraction and storage are performed on the data in the industrial knowledge base. S32. Selection and loading of the large language model base: Select a large language model as the large model base of the industrial large model agent, and perform supervised fine-tuning on the large model base through labeled data analysis examples. S33. Configure the tool components to build a processing flow: Configure the tool components of the agent, and the tool components include, but are not limited to, input nodes, large model calls, vector library retrieval, external links, logical processing, special result output, and data processing. S34. Publish the function interface: The industrial large model agent provides services as a function interface.
[0011] Preferably, in step 3, the method for running the industrial large model agent is as follows: S3a. First, input the preprocessed industrial hot text as a parameter into the industrial large model agent. S3b. The industrial large model agent performs industrial knowledge base retrieval and network retrieval based on the input industrial hot text, and uses the retrieved data as background knowledge. S3c. Sentiment analysis: Based on the praise and criticism analysis ability of the large model, construct praise and criticism analysis prompt words through prompt engineering, fill the background knowledge and industrial hot text into the large model prompt words, and obtain and mark whether the industrial hot text contains a derogatory sentiment. S3d. Industrial judgment: Based on the logical judgment ability of the large model, construct industrial-related judgment prompt words, fill the background knowledge and industrial hot text into the large model prompt words, and obtain and mark whether the industrial hot text is related to the industry. S3e. According to the marking results of steps S3c and S3d, filter out irrelevant hot topic data, and screen the filtered data that contains both derogatory sentiment and is related to the industry as the hot topics related to industrial risks and output them.
[0012] Preferably, in step 4, the method for generating industrial suggestions based on the large model is as follows: S41. Aggregate industrial risk hot topics: Select representative data from each topic obtained in step 3 and splice them; S42. Construct prompt words: Construct prompt words for generating industrial suggestions through prompt engineering, and fill the risk hot topic data into the prompt words; S43. Invoke the large model to generate industrial suggestions for industrial risk hot topics and output them.
[0013] An industrial risk identification framework based on a hot topic clustering algorithm and an industrial large model agent, including: Application layer: Support interface calls and embedded calls; Deploy as a service through middleware, support external applications to call through the REST protocol; Provide support encapsulated as a method library, support external applications to perform embedded calls by introducing the method library and calling the interface; Industrial large model agent: Include: Invoking external industrial knowledge bases and large language model services, prompt words constructed through prompt engineering, and tool components for enhancing the capabilities of large models; Industrial risk hot topic identification service unit, including: A hot topic clustering module for mining hot topics from industrial data, a semantic vectorization module for converting text data into semantic vectors, a data classification module for preprocessing industrial data and filtering interfering data; Invoking the large language model service, a sentiment analysis module for performing sentiment analysis on hot topic data, invoking the large language model service, an industrial relevance identification module for judging the industrial relevance of hot topic data, invoking the large language model service, and an industrial suggestion generation module for generating industrial development suggestions for industrial negative hot topic data.
[0014] Beneficial effects: The present invention proposes an industrial risk identification method based on a hot topic clustering algorithm and an industrial large model agent, which combines the clustering algorithm and the large model to discover useful information under a large amount of data and accurately identify industrial risks. First, through the hot topic clustering algorithm, text semantic vectorization is realized through the position weighted average pooling and contrast fine-tuning method of the bias tensor, and hot topics are found from a large amount of industrial text data through connected graph clustering and splitting optimization. Secondly, based on the semantic understanding ability of the large model, an industrial large model agent is constructed, and the hot topics are judged through the industrial large model agent, including sentiment analysis and industrial judgment, so as to screen out the hot topics involving industrial risks. Compared with using the large model and the clustering algorithm alone, the synergistic effect of the large model and the clustering algorithm makes the identification of hot topics of industrial risks more accurate. At the same time, input the clustered hot topics into the industrial large model. Compared with directly inputting the text into the large model, the processing time of the large model is saved by optimizing the process, thereby improving the efficiency. Thirdly, apply the inductive summary ability of the large model to summarize industrial development suggestions based on the summarized industrial risk hot topics. Description of the Drawings
[0015] Figure 1 Flowchart of an industrial risk identification method based on a hot - spot clustering algorithm and an industrial large - model agent in the embodiment
[0016] Figure 2 Flowchart of hot - spot clustering in the embodiment
[0017] Figure 3 Example graph of undirected graph G in the embodiment
[0018] Figure 4 Flowchart of industrial risk judgment in the embodiment
[0019] Figure 5 Flowchart of industrial recommendation generation in the embodiment
[0020] Figure 6 Program framework diagram of an industrial risk identification based on a hot - spot clustering algorithm and an industrial large - model agent in the embodiment Specific implementation manners
[0021] The present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners
[0022] As Figure 1 shown, an industrial risk identification method based on a hot - spot clustering algorithm and an industrial large - model agent S1. Construct an industrial data set: Construct an industrial data set through an industrial database and scraping online data; the industrial data set reflects the characteristics of a specified industry and contains data with negative views S2. Hot - spot clustering: Perform semantic vectorization processing on the data in the industrial data set described in S1, and execute similarity calculation to evaluate data correlation; perform clustering analysis on the data through a connected graph, and optimize the clustering result through splitting sub - graphs; reverse - order sort the clusters in the clustering result based on the number of members in the class to obtain industrial hot - spot texts S3. Screening of industrial risk hot - spots: Construct an industrial large - model agent based on an industrial knowledge base and a large - model and configure tool components, input the industrial hot - spot texts into the industrial large - model agent, run the industrial large - model agent for sentiment analysis and industry - related judgment to obtain hot - topics related to industrial risks S4. Industrial recommendation generation: Generate industrial recommendations based on a large - model for industrial risk hot - spots
[0023] 1. Construct an industrial data set To identify the hot topics in a specified industry, a certain amount of industry-related text data needs to be obtained first. The data types include, but are not limited to, industry news, industry public opinion data, industry-related regulations, etc. Industry data generally comes from the data in existing industry databases or data crawled from the Internet.
[0024] When all the obtained industry data is aggregated together, an industry data set to be analyzed can be constructed. The industry data set generally needs to have the following characteristics: 1) The data included in the industry data set needs to have a certain relevance to the specified industry.
[0025] 2) The industry data set needs to have a certain quantity scale to support the subsequent clustering steps to form industry hotspots.
[0026] 3) The data in the industry data set needs to have good data quality. The length of each piece of data should not be too short, and interference information should be avoided as much as possible.
[0027] 4) The industry data set needs to contain a certain number of data with negative views on the specified industry, otherwise industry risk topics cannot be analyzed.
[0028] 2. Industry data hot topic clustering The main process of industry data hot topic clustering is as Figure 2 shown: 1) Data semantic vectorization The main purpose of semantic vectorization is to extract the semantic features contained in the data text and encode them into vector representations in a low-dimensional space through a model, which facilitates the subsequent similarity calculation and relevance judgment.
[0029] The vectorization model used in this method is the GPT structure, which uses the decoder layer of the Transformer structure to achieve the vectorization of text semantics through the position-weighted average pooling and the contrast fine-tuning method of the bias tensor.
[0030] We denote the text set to be analyzed as T, and Ti is a member of T, that is, Ti ∈ T. This method traverses all the data in T and extracts the semantic vector Vi of the text based on Ti through the vectorization method, thus constructing a semantic vector set V.
[0031] 2) Data similarity calculation and relevance judgment Since texts with similar semantics also have similar vectors, the semantic similarity between two texts can be judged by calculating the distance between the semantic vectors of the two texts, and whether the two texts are relevant can be judged according to the similarity threshold.
[0032] This method traverses any two-vector combinations in the semantic vector set V and calculates the cosine value between any two vectors through the vector cosine calculation formula. According to the input similarity threshold and the magnitude of the cosine value between vectors, it is determined whether there is a correlation between any two vectors. If the cosine value exceeds the threshold, it is judged that there is a correlation.
[0033] 3) Construct a topological graph As Figure 3 shown, this method constructs a vertex set with text data as vertices, constructs an adjacency matrix based on the correlation between data as the basis for whether there is an edge between vertices, and constructs a weighted undirected graph G of the dataset T based on text similarity relationships with the cosine value between data as the edge weights.
[0034] 4) Connected graph clustering In graph theory, it is defined that in an undirected graph, if there is a path connecting vertex i to vertex j, then i and j are said to be connected. If any two points in the graph are connected, then the graph is called a connected graph.
[0035] This method believes that the vertices that are connected to each other in the undirected graph G have semantic correlation relationships. According to the connectivity between the vertices of the undirected graph G, the graph G is divided into multiple connected graphs, and each connected graph can be regarded as a cluster. The vertices Ti included in the connected graph are the class member objects included in the cluster.
[0036] 5) Split and optimize This method analyzes the results of connected graph clustering. If the scale of a class exceeds a certain range, then the class may need to be split and optimized to prevent data that is actually semantically irrelevant from aggregating into one class due to individual abnormal data. For those clusters that need to be optimized, remove the relationships with lower weights and split out subgraphs again according to the connectivity of the graph to obtain more compact clusters.
[0037] 6) Result collation Through the above steps, we cluster the industrial data into multiple clusters. The data within each cluster has semantic correlation. This method believes that these clusters represent the topics included in the industrial dataset, and the number of class members, that is, the amount of data included in the cluster, reflects the popularity of the corresponding topic. We sort the clusters in descending order according to the number of class members, and the clusters ranked in the front are the hot topics included in the industrial dataset.
[0038] 3. Identification of hot topics of industrial risks based on large language models Since large language models learn and store a large amount of knowledge during the training phase, including common sense, professional knowledge, cultural knowledge, etc., large language models possess certain logical reasoning abilities and can perform reasoning and judgment based on the given information. Through the above steps, we have obtained the hot topics of industrial data. This method realizes the function of screening out topics related to industrial risks from the hot topics based on the semantic understanding and reasoning abilities of large models.
[0039] 1) Data preprocessing For the hot topics clustered from the hotspots, we need to perform data preprocessing to convert the clustering results into a data format convenient for large models to analyze. Since each cluster contains multiple industrial data texts as cluster member objects, and the prompts that large models can handle have length limitations, it is necessary to sample and select several data texts in each hot topic and splice them together for convenient subsequent calls.
[0040] 2) Construct an industrial large model agent In order to fully activate the semantic understanding and logical reasoning abilities of large language models, this method needs to first construct an industrial large model agent based on the large language model. The main work process of constructing an industrial large model agent is as follows: a) Construct an industrial knowledge base Optionally, in order to improve the judgment accuracy of large language models, we need to construct an external industrial knowledge base to expand the industrial knowledge available for large models. The data constituting the industrial knowledge base mainly includes, but is not limited to: industrial-related policies and regulations, industrial knowledge graphs, industrial historical news data, etc.
[0041] After this method obtains industrial-related data, it performs vectorized feature extraction and storage on these data to facilitate subsequent improvement of the accuracy, relevance, and timeliness of analysis results through the technology of Retrieval-Augmented Generation (RAG).
[0042] b) Selection and loading of the large language model base The core of the industrial large model agent is the large language model. We need to select a suitable open-source or non-open-source large language model as the large model base of the industrial large model agent. Under permitted conditions, we can perform supervised finetuning (SFT) on the large model base by annotating data analysis examples to make the model adapt to the analysis task and improve the accuracy and stability of the analysis.
[0043] c) Construct a processing flow through tool components The agent adds additional tool components on the basis of the large language model, expanding the capabilities of the large model to adapt to more complex processing flows. Common tool components generally include: input nodes, large model calls, vector library retrieval, external links, logical processing, special result output, data processing and other tool components. Based on these components and the capabilities of the large model, this method plans the processing flow of the agent and constructs an intelligent agent for judging industrial risk hotspots.
[0044] d) Publish as a functional interface After the intelligent agent is constructed, it needs to be published externally as a callable service. In the subsequent actual implementation process, the industrial large model intelligent agent provides services externally as a functional interface.
[0045] 3) Industrial risk identification operation process of the industrial large model intelligent agent Figure 4 The specific operation process of the industrial large model intelligent agent is shown as follows: a) First, input the preprocessed industrial hotspot text as a parameter into the intelligent agent.
[0046] b) Based on the input industrial hotspot text, the intelligent agent conducts knowledge base retrieval and network retrieval, and uses the retrieved data as background knowledge.
[0047] c) Based on the large model's praise and criticism analysis ability, construct praise and criticism analysis prompts (Prompts) through Prompt Engineering. Fill the background knowledge and industrial hotspot text into the large model prompts to obtain and mark whether the industrial hotspot text contains negative emotions.
[0048] d) Based on the large model's logical judgment ability, construct industrial-related judgment prompts. Fill the background knowledge and industrial hotspot text into the large model prompts to obtain and mark whether the industrial hotspot text is related to the industry.
[0049] e) According to the marking results of c and d, filter out irrelevant hotspot topic data and screen out hotspot topics related to industrial risks.
[0050] 4. Generation of industrial suggestions based on the large model Through the above steps, we have analyzed hotspot topics related to industrial risks from the industrial dataset. This method, based on the inductive summary ability of the large language model, summarizes all the discovered industrial risk hotspot topics and generates suggestions for the subsequent development of the industry. As Figure 5 shown, the main process is as follows: 1) Summarize industrial risk hotspot topics, select several pieces of data in each topic as topic representatives, and splice all the data as the data basis for the large model to generate suggestions.
[0051] 2) Construct prompts for generating industry suggestions through Prompt Engineering and fill the risk hot spot data into the prompts.
[0052] 3) Invoke the large model to generate industry suggestions for industry risk hot spots and output them.
[0053] As Figure 6 shown, an industry risk identification framework based on a hot spot clustering algorithm and an industry large model agent includes: Application layer: Supports interface calls and embedded calls; Deployed as a service through middleware, supports external applications to call through the REST protocol; Provides support encapsulated as a method library, supports external applications to make embedded calls by introducing the method library and calling the interface. Industry large model agent: Includes: Invoking external industry knowledge bases and large language model services, prompts constructed through prompt engineering, and tool components for enhancing the capabilities of large models. Industry risk hot spot identification service unit, including: A hot spot clustering module for mining hot topics from industry data, a semantic vectorization module for converting text data into semantic vectors, a data classification module for preprocessing industry data and filtering interference data; Invoking the large language model service, a sentiment analysis module for performing sentiment analysis on hot topic data, invoking the large language model service, an industry relevance identification module for judging the industry relevance of hot topic data, invoking the large language model service, and an industry suggestion generation module for generating industry development suggestions for industry negative hot spot data.
[0054] 1. This method relies on the capabilities of large language models, and the underlying layer is an industry large model agent based on large models. It includes external industry knowledge bases (industry knowledge vector bases), large language model services, prompts constructed through prompt engineering, and tool components for enhancing the capabilities of large models.
[0055] 2. In the actual implementation process, this method provides services to the application side as an independent service, mainly including the following functional modules: 1) The hot spot clustering module is responsible for mining hot topics from industry data 2) The semantic vectorization module is responsible for converting text data into semantic vectors.
[0056] 3) The data classification module is mainly responsible for preprocessing industry data and filtering interference data.
[0057] 4) The sentiment analysis module is responsible for performing sentiment analysis on hot topic data based on the praise and criticism analysis capabilities of the large model.
[0058] 5) The industry relevance identification module is responsible for judging the industry relevance of hot topic data based on the inference ability of the large model.
[0059] 6) The industry suggestion generation module is responsible for generating industry development suggestions for negative hot industry data based on the inductive and summarizing ability of the large model.
[0060] 3. The method can be called by external applications in various ways as a functional service: 1) This method supports being deployed as a service through middleware and supports external applications to call through the REST protocol.
[0061] 2) This method supports being encapsulated as a method library and supports external applications to make embedded calls by introducing the method library and calling the interface.
[0062] Example 2: A program constructed based on the method of the present invention is applied to Industry A for illustration: 1. Through network collection, news data related to Industry A is obtained. By means of keyword retrieval, time period filtering, data preprocessing, etc., a news data set related to Industry A is constructed, with the data volume being 39,809 pieces, and the data fields including: text, title, time, unique value, etc.
[0063] 2. Through the hot topic clustering algorithm, hot topics are discovered in the data set. By setting parameter thresholds, the top 200 topics with the highest popularity are found as shown in Table 1, where the topic name is the inductive hot topic theme information, and the topic popularity is the news data included under the topic:
[0064] 3. Using the semantic understanding ability and logical judgment ability of the large model, sentiment judgment and industry relevance screening are carried out on the hot topics, and several hot topics related to industry risks are screened out from 200 hot topics as shown in Table 2:
[0065] 4. Using the analysis ability of the large model, the industry risk hot topics obtained in the previous steps are used as public opinion information and input into the large model. Through the industry risk overview obtained by the large model analysis, the analysis results output by the large model agent are as follows: {The risks currently faced by Industry A mainly include the following points: 1. Technical route and market acceptance risks 2. Infrastructure and cost risks 3. Policy risks 4. Market competition risks In summary, the A industry faces significant risks in aspects such as technology route selection, infrastructure construction, policy dependence, and market competition, which require the joint efforts of enterprises, the government, and society to address. Based on the large model's risk analysis of the A industry, it is possible to further use the large model to present challenges and problems for the A industry, as well as corresponding suggestions, respectively for the government, consumers, and enterprises under the existing risk situation. The industry suggestions generated by this large model are as follows: I. Challenges and Suggestions for the Government 1. Challenges Lagging infrastructure construction, insufficient policy stability, and supply chain security risks 2. Suggestions Strengthen infrastructure planning, improve the policy support system, and ensure supply chain security.
[0066] II. Challenges and Suggestions for Consumers III. Challenges and Suggestions for Enterprises It should be noted that the above-described specific implementation manners can enable those skilled in the art to more comprehensively understand the present invention, but do not limit the present invention in any way. Therefore, although this specification has described the present invention in detail with reference to the accompanying drawings and embodiments, those skilled in the art should understand that the present invention can still be modified or equivalently replaced. In short, all technical solutions and their improvements that do not depart from the spirit and scope of the present invention should be covered by the protection scope of the patent of the present invention.
Claims
1. An industrial risk identification method based on hotspot clustering algorithm and industrial large model intelligent agent, characterized in that: S1, constructing an industry data set: constructing an industry data set by crawling an industry database and online data; the industry data set includes data with specified industry characteristics and containing negative opinions; S2, hotspot clustering; semantic vectorization processing is performed on the industry data set described in S1, and similarity calculation is performed to obtain semantic correlation relationships, and a connectivity graph is further constructed to obtain clustering results, and the clustering results are optimized by splitting subgraphs; based on the number of members of the class in the connectivity graph, the clusters in the clustering results are sorted in reverse order to obtain industry hotspot texts; S3, screening of industry risk hot spots: fine-tune the big model, configure the industry knowledge base and big model tool components, and obtain the industry big model intelligent body; input hot spot clustering into the industry big model intelligent body to obtain the industry hot spot text, run the industry big model intelligent body to perform sentiment analysis and industry-related judgment in turn, and screen out hot topics related to industry risks; S4, industry suggestion generation: summarize industry risk hot topics, construct prompt words, call the big model to generate industry suggestions for industry risk hotspots and output them.
2. The industrial risk identification method based on hotspot clustering algorithm and industrial large model intelligent agent according to claim 1 is characterized in that: The method for data semantic vectorization in step S2 is: S21, constructing a model: constructing a neural network model with a GPT architecture, wherein the neural network model includes at least one Transformer decoder layer and is configured with a position weighted average pooling module; S22, parameter optimization: optimizing the parameters of the neural network model by using a bias tensor comparison fine-tuning method, and introducing a learnable bias tensor into the loss function to perform similarity comparison calculation; S23, generating a semantic vector: traversing all text elements Ti in the text set T to be analyzed, performing semantic encoding on each Ti through the neural network model, and generating a corresponding semantic vector Vi; S24, construct a semantic vector set: aggregate the semantic vectors Vi corresponding to all text elements Ti, and construct a structured semantic vector set V = {Vi|Ti∈T}.
3. The industrial risk identification method based on hotspot clustering algorithm and industrial large model intelligent agent according to claim 1 is characterized in that: The similarity calculation method in step S2 is: traverse any two vector combinations in the semantic vector set V, and calculate the cosine value between any two vectors using the vector cosine calculation formula. According to the input similarity threshold and the cosine value between the vectors, it is determined whether there is a correlation between any two vectors. If the cosine value exceeds the threshold, it is determined that there is a correlation.
4. The industrial risk identification method based on hotspot clustering algorithm and industrial large model intelligent agent according to claim 1 is characterized in that: The method for realizing data cluster analysis by using correlation through connectivity graph in step S2 is: S2a, constructing a topological graph: using the text data in the industry dataset as vertices to construct a vertex set, using the correlation between the text data as the basis for whether there is an edge between the vertices to construct an adjacency matrix, using the cosine value between the text data as the weight of the edge, and constructing a weighted undirected graph G based on the text similarity relationship of the dataset T; S2b, connected graph clustering: the mutually connected vertices in the weighted undirected graph G have semantically related relationships; according to the connectivity between the vertices of the weighted undirected graph G, the graph G is divided into multiple connected graphs, each of which can be regarded as a class cluster, and the vertices Ti contained in the connected graph are the class member objects contained in the class cluster; S2c, split optimization: If the size of a class exceeds the set range, remove the relationships with lower weights and re-split the subgraph according to the connectivity of the graph to obtain a new cluster.
5. The industrial risk identification method based on hotspot clustering algorithm and industrial large model intelligent agent according to claim 1 is characterized in that: In step 2, the data of each cluster in the clustering results have semantic correlation, and the number of class members, that is, the amount of data contained in the cluster, reflects the popularity of the corresponding topic; the clusters are sorted in reverse order according to the number of class members, and the clusters before the threshold are set as hot topics contained in the industry data set, and output as industry hot text.
6. The industrial risk identification method based on hotspot clustering algorithm and industrial large model intelligent agent according to claim 1 is characterized in that: In step 3, the method of constructing the industrial large model intelligent agent is: S31, constructing an industry knowledge base: the data constituting the industry knowledge base mainly includes but is not limited to: industry-related policies and regulations, industry knowledge graphs, and industry historical news data; vectorized feature extraction and storage of data in the industry knowledge base; S32, large language model base selection and loading: select a large language model as the large model base of the industrial large model agent, and perform supervised fine-tuning on the large model base by annotating data analysis samples; S33, configure tool components to build a processing flow: configure tool components of the agent, the tool components include but are not limited to: input nodes, large model calls, vector library retrieval, external links, logic processing, special result output, data processing; S34, release functional interface: The industrial large model intelligent body provides services to the outside world as a functional interface.
7. The industrial risk identification method based on hotspot clustering algorithm and industrial large model intelligent agent according to claim 1 is characterized in that: In step 3, the method for running the industrial large model agent is: S3a, firstly, the pre-processed industry hot text is input as a parameter into the industry big model intelligent agent; S3b, the industry big model agent performs industry knowledge base retrieval and network retrieval based on the input industry hot text, and uses the retrieved data as background knowledge; S3c, sentiment analysis: Based on the big model's ability to analyze praise and criticism, the prompt engineering is used to construct praise and criticism analysis prompt words, and background knowledge and industry hotspot texts are filled into the big model prompt words to obtain and mark whether the industry hotspot texts contain derogatory sentiments; S3d, industry judgment: Based on the logical judgment ability of the big model, construct industry-related judgment prompt words, fill the background knowledge and industry hot text into the big model prompt words, obtain and mark whether the industry hot text is related to the industry; S3e, according to the marking results of steps S3c and S3d, filter out irrelevant hot topic data, select and output the filtered data that contains both derogatory sentiment and is related to the industry as hot topics related to industrial risks.
8. The industrial risk identification method based on hotspot clustering algorithm and industrial large model intelligent agent according to claim 1 is characterized in that: In step 4, the method by which the big model generates industry recommendations for hot topics is as follows: S41, summarize hot topics of industry risks: select representative data from each topic obtained in step 3 and splice them; S42, constructing prompt words: constructing prompt words for generating industry suggestions through prompt engineering, and filling risk hotspot data into the prompt words; S43, call the big model to generate and output industry recommendations for industry risk hotspots.
Citation Information
Patent Citations
Industrial chain multi-collaborative intelligent decision-making method based on knowledge graph
CN117172725A
Smart park industry cluster data contrastive analysis system and method
CN117251492A
Enterprise operation supervision multi-agent cooperation method and system based on large language model
CN119107051A
Cited By
Industrial development multi-dimensional collaborative evaluation method based on artificial intelligence large model
CN120822854A
An industry development multi-dimensional collaborative evaluation method based on an artificial intelligence large model
CN120822854B