Specific information studying and judging method and system based on large language model

By adopting a combination method of large language model and sliding window technology in the process of data collection, information recognition and specific information analysis, the problems of low data collection efficiency, incomplete information recognition and in real-time analysis and judgment of specific information in the prior art are solved, and efficient and accurate data management and three-dimensional character attribute description are achieved.

CN120045763APending Publication Date: 2025-05-27NAT COMP NETWORK & INFORMATION SECURITY MANAGEMENT CENT
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411949680.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The prior art has problems such as inefficiency, low data quality, lack of automated classification and classification capabilities and difficulty in adapting to data scale growth in data collection, information identification, specific information analysis and system scalability.

Method used

Using a method based on a large language model, data is obtained through a combination of API access, reverse analysis APP and network crawlers, and cleaned and standardized. The basic attribute information in the information data is extracted using a large language model, a relational knowledge base is constructed, and a three-dimensional attribute description is obtained through the fusion of online and offline features. Information analysis and analysis are carried out based on sliding window technology to realize real-time prompts.

Benefits of technology

It improves data acquisition efficiency and quality, realizes efficient management of multi-source heterogeneous data, the system supports real-time data transmission and efficient prompts, improves the accuracy of information extraction and management, and builds a three-dimensional character attribute description system, supporting dynamic updates and real-time processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045763A_ABST
    Figure CN120045763A_ABST
Patent Text Reader

Abstract

The invention provides a specific information research and judgment method and system based on a large language model, and the method comprises the steps: obtaining information data through the combination of API access, a reverse analysis APP and a web crawler, carrying out the cleaning and standardization processing of the information data, and storing the information data into a distributed database for unified management; the method comprises the following steps: preprocessing information data based on a large language model, extracting basic attribute information of a to-be-analyzed object in the information data by adopting a mode of combining pre-training and fine tuning, and constructing a relationship knowledge base based on the basic attribute information; obtaining online features of the to-be-analyzed object through the online dimension, obtaining offline features of the to-be-analyzed object through the offline dimension, and carrying out feature fusion on the online features and the offline features to obtain three-dimensional attribute description; and based on a sliding window technology, carrying out information research and judgment analysis on the text determined by the relation knowledge base and the three-dimensional attribute description, and carrying out real-time prompting on abnormal information according to a research and judgment analysis result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of big data analysis and artificial intelligence applications, and particularly relates to a method and system for judging specific information based on a large language model. Background Art

[0002] Currently, the systematic collection, collation, judgment, and analysis of certain individuals and their specific information are carried out. At present, this kind of work is mainly carried out manually, and it is difficult to meet the actual needs in terms of efficiency and scale.

[0003] The existing related technologies mainly have the following problems:

[0004] 1. In terms of data collection: The data collection of existing technologies for social media and news platforms is usually scattered and independent, lacking a unified data access and management mechanism. The anti-crawler strategies of different platforms are different, and it is difficult to collect data from multiple platforms through a single technology. In addition, due to the inconsistent data formats of different information sources and uneven data quality, there is a lack of effective data cleaning and standardization processing solutions, resulting in low efficiency in data acquisition, merging, and integration.

[0005] 2. In terms of information recognition: Existing technologies are difficult to effectively identify and associate the account information of the same natural person on different platforms, lacking the ability to analyze and mine the social relationship network of certain individuals on different platforms. This results in the inability to accurately construct the attribute description of the individual, making the information about the individual incomplete and inaccurate.

[0006] 3. In terms of judging specific information: Existing technologies mainly rely on manual methods to judge specific information, which is not only inefficient but also easily affected by personal subjective factors, and cannot guarantee the consistency and accuracy of the judgment results. At the same time, due to the lack of automated classification and grading capabilities for specific information, real-time processing and prompting of specific information cannot be achieved.

[0007] 4. In terms of system scalability: Most existing systems are rigidly designed, difficult to adapt to the growing data scale and new data sources, and lack a perfect data security protection mechanism. This results in insufficient scalability and maintainability of the system, making it difficult to meet the growing business needs.

[0008] In view of the above problems, there is an urgent need to design a technical solution that can automatically collect data, intelligently analyze and judge, and perform real-time processing and prompting. The present invention is a solution proposed for these technical problems. Summary of the Invention

[0009] The present invention provides a method and system for judging specific information based on a large language model to solve the problems raised in the background art.

[0010] A method for judging specific information based on a large language model includes:

[0011] Step 1: Obtain information data through a combination of API access, reverse analysis of APPs, and web crawlers, clean and standardize the information data, and then store it in a distributed database for unified management;

[0012] Step 2: After preprocessing the information data based on a large language model, extract the basic attribute information of the objects to be analyzed in the information data by combining pre-training and fine-tuning, and construct a relationship knowledge base based on the basic attribute information;

[0013] Step 3: Obtain the online features of the objects to be analyzed through the online dimension, obtain the offline features of the objects to be analyzed through the offline dimension, and fuse the online features and offline features to obtain a three-dimensional attribute description;

[0014] Step 4: Based on the sliding window technology, conduct information research and judgment analysis on the text determined by the relationship knowledge base and the three-dimensional attribute description, and give real-time prompts for abnormal information according to the research and judgment analysis results.

[0015] Preferably, in Step 1, obtaining information data through a combination of API access, reverse analysis of APPs, and web crawlers includes:

[0016] Deploy a self-developed distributed crawler program on a special-purpose computer through anti-crawling restriction technology, set reasonable concurrency and download delay, and collect data;

[0017] Design a custom RESTful-style API interface, and through an authentication mechanism and security control, transmit the collected data to obtain information data.

[0018] Preferably, transmitting the collected data also includes:

[0019] Use the HMAC-SHA256 algorithm for request signature authentication during the data transmission of the collected data.

[0020] Preferably, in Step 1, after cleaning and standardizing the information data and storing it in a distributed database for unified management, it includes:

[0021] Based on a self-developed algorithm, sequentially perform data cleaning, multi-dimensional duplicate data detection, outlier processing based on statistical methods, and missing value filling with multiple strategies on the information data;

[0022] Store and uniformly manage the information data in multiple shard nodes and replica sets in the distributed database.

[0023] Preferably, in the second step, after preprocessing the information data based on the large language model, the basic attribute information of the object to be analyzed in the information data is extracted by combining pre-training and fine-tuning, and a relationship knowledge base is constructed based on the basic attribute information, including:

[0024] Deploy a large language model pre-trained with customization, and perform standardized input processing using a class normalization method. The calculation formula is as follows:

[0025]

[0026] where x' represents the output vector after preprocessing the information data based on the large language model, n represents the vector dimension, and x i represents the i-th input vector of the information data;

[0027] Optimize the large language model according to the positional encoding technique. The calculation formula is as follows;

[0028] RoPE(x,m,θ)=[x 1 cos(mθ)-x 2 sin(mθ),x 1 sin(mθ)+x 2 cos(mθ)]

[0029] where m represents the position index, θ represents the base frequency parameter, x 1 and x 2 are adjacent dimensions of the input vector; RoPE(x,m,θ) represents introducing a rotation transformation related to the position;

[0030] Design a specific prompt template to guide the automated extraction of information data from the optimized large language model to obtain the extracted text;

[0031] And identify the key entities in the extracted text based on entity recognition technology. Based on the entity semantics of the key entities, associate the information data with different expressions referring to the same entity through entity linking to form a unified knowledge representation of the basic attribute information of the object to be analyzed, and construct a relationship knowledge base according to the knowledge representation.

[0032] Preferably, in the third step, obtain the online features of the object to be analyzed through the online dimension, obtain the offline features of the object to be analyzed through the offline dimension, and perform feature fusion on the online features and offline features to obtain a three-dimensional attribute description, including:

[0033] Construct the online features of the object to be analyzed in the online dimension based on the online quantization index;

[0034] Construct a social network structure based on the offline indicators, and calculate the offline centrality of the object to be analyzed according to the following formula;

[0035]

[0036] Among them, Centrality(v) represents the offline centrality of the object to be analyzed, v represents the target node, s represents the node position of any first offline quantization index in the social network structure, t represents the node position of any second offline quantization index in the social network structure, and shortest p aths(s,t|v) represents the number of shortest paths passing through the target node, and shortest p aths(s,t) represents the number of all shortest paths in the social network structure;

[0037] Offline features constructed based on offline indicators and offline centrality;

[0038] Based on the following formula, the online features and offline features are weighted and fused to obtain a comprehensive score;

[0039]

[0040] Among them, Score represents the three-dimensional attribute description, g represents the total number of features of the online features and offline features, w j represents the weight of the j-th feature, and f j represents the standardized score of the j-th feature.

[0041] Based on the online features, offline features, offline centrality and comprehensive score, a three-dimensional attribute description of the object to be analyzed is constructed.

[0042] Preferably, in the fourth step, based on the sliding window technology, information research and judgment analysis is carried out on the text determined by the relationship knowledge base and the three-dimensional attribute description, and real-time prompts for abnormal information are given according to the research and judgment analysis results, including:

[0043] Carry out topic classification on the text determined by the relationship knowledge base and the three-dimensional attribute description;

[0044] Based on the sliding window technology, set the time window size and sliding step length, and continuously process the change in the discussion volume of each topic within a specific time period to realize the update of the text;

[0045] Based on the improved LDA topic model, text clustering analysis is carried out, and its topic-word probability calculation formula is:

[0046]

[0047] Among them, P(w|d) represents the topic-word probability, w represents the word, d represents the text, z represents the topic, P(w|z) represents the probability of the word under the topic, and P(z|d) represents the probability of the topic in the text;

[0048] Determine the mean and standard deviation of historical data of the text based on the theme-word probability;

[0049] Perform anomaly detection based on a statistical threshold, and the calculation formula of the threshold is as follows:

[0050] Threshold = μ + k * σ

[0051] Where Threshold represents the threshold, μ represents the mean of historical data, σ represents the standard deviation of historical data, and k represents the adjustment parameter;

[0052] Based on the four-level impact rating standard with different settings of the adjustment parameter, determine the rating threshold corresponding to each level. When the real-time threshold of the latest text exceeds the rating threshold, trigger a real-time prompt.

[0053] Preferably, the manifestation of triggering the real-time prompt is a keyword cloud map and a trend chart.

[0054] Preferably, the values of the parameters are different under different impact rating standards, and are specifically set in advance according to the actual situation.

[0055] A specific information research and judgment system based on a large language model, comprising:

[0056] A multi-channel data acquisition and management module, which is used to obtain information data based on a combination of API access, reverse analysis of APPs, and web crawlers, and after cleaning and normalizing the information data, store it in a distributed database for unified management;

[0057] An information extraction and knowledge base module, which is used to preprocess the information data based on a large language model, extract the basic attribute information of the object to be analyzed in the information data by combining pre-training and fine-tuning, and construct a relationship knowledge base based on the basic attribute information;

[0058] A feature processing module, which is used to obtain the online features of the object to be analyzed through the online dimension, obtain the offline features of the object to be analyzed through the offline dimension, and perform feature fusion on the online features and the offline features to obtain a three-dimensional attribute description;

[0059] An information research and judgment module, which is used to perform information research and judgment analysis on the text determined by the relationship knowledge base and the three-dimensional attribute description based on the sliding window technology, and give a real-time prompt for abnormal information according to the research and judgment analysis results.

[0060] Compared with the prior art, the present invention has achieved the following beneficial effects:

[0061] Improved the efficiency and quality of data collection: Through a unified data collection framework and standardized data processing procedures, efficient collection and management of multi-source heterogeneous data were achieved. The system supports the access and collection of no less than 10 information types, and the data collection communication supports real-time data transmission of more than 10 seconds, improving the accuracy of information extraction and management. Based on the information extraction method of the large language model, the automatic extraction accuracy of essential fields exceeds 80%, and the field accuracy in the delivered knowledge base exceeds 95%. A person information database has been successfully constructed, containing no less than 8 types of attribute information and capable of dynamic updates; A three-dimensional person attribute description system has been realized. Through the fusion analysis of multi-dimensional features online and offline, the description of person and organization attributes has been constructed. The scientific quantification of person influence is realized using the super IP theory, and the description system can be dynamically updated to maintain the timeliness and accuracy of information; Support for real-time processing and efficient prompting. The system can achieve a data processing speed of 300 records per second in a single thread, and the query response for person information remains at the second level; An influence judgment system containing no less than 5 categories and 4 levels has been established to achieve precise identification and timely prompting of specific information.

[0062] Other features and advantages of the present invention will be described in the following specification, and some of them will be obvious from the specification or understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the structures specifically pointed out in this application document.

[0063] The technical solution of the present invention will be further described in detail below through the accompanying drawings and embodiments. Description of the Drawings

[0064] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation to the present invention. In the drawings:

[0065] Figure 1 It is a flowchart of a method for judging specific information based on a large language model in an embodiment of the present invention;

[0066] Figure 2 It is a structural diagram of a system for judging specific information based on a large language model in an embodiment of the present invention;

[0067] Figure 3 It is a display diagram of the hierarchical architecture design in an embodiment of the present invention. Detailed Embodiments

[0068] The following describes the preferred embodiments of the present invention with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention and are not used to limit the present invention.

[0069] Embodiment 1:

[0070] A method for judging specific information based on large language models, as Figure 1 shown, includes:

[0071] Step 1: Obtain information data through a combination of API access, reverse analysis of APPs, and web crawlers, clean and standardize the information data, and store it in a distributed database for unified management;

[0072] Step 2: After preprocessing the information data based on a large language model, adopt a combination of pre-training and fine-tuning to extract the basic attribute information of the object to be analyzed in the information data, and construct a relationship knowledge base based on the basic attribute information;

[0073] Step 3: Obtain the online characteristics of the object to be analyzed through the online dimension, obtain the offline characteristics of the object to be analyzed through the offline dimension, and perform feature fusion on the online characteristics and offline characteristics to obtain a three-dimensional attribute description;

[0074] Step 4: Based on the sliding window technology, conduct information judgment and analysis on the text determined by the relationship knowledge base and the three-dimensional attribute description, and give real-time prompts for abnormal information according to the judgment and analysis results.

[0075] In this embodiment, the information data is, for example, social media and news website data.

[0076] In this embodiment, the relationship knowledge base includes a knowledge base with multiple attributes such as foreign names, foreign abbreviations, Chinese names, date of birth, etc.

[0077] In this embodiment, the online dimension includes characteristics such as personnel behavior, content generation, and dissemination impact; the offline dimension includes characteristics such as professional background, institutional position, and social relationships.

[0078] In this embodiment, based on the sliding window technology, specific information is sorted and judged, relevant remarks are classified according to the theme, and are divided into levels one to four according to the degree of influence. The LDA topic model and semantic analysis method are used to identify abnormal information patterns, and real-time prompts are given for information exceeding the set threshold.

[0079] The beneficial effects of the above design solution are as follows: It improves the efficiency and quality of data collection. Through a unified data collection framework and a standardized data processing process, the efficient collection and management of multi-source heterogeneous data are realized. The system supports the access and collection of no less than 10 types of information, and the data collection communication supports real-time data transmission for more than 10 seconds, enhancing the accuracy of information extraction and management. Based on the information extraction method of the large language model, the automatic extraction accuracy of essential fields exceeds 80%, and the field accuracy in the delivered knowledge base exceeds 95%. A person information database has been successfully constructed, including no less than 8 types of attribute information, and can be dynamically updated. A three-dimensional person attribute description system has been realized. Through the fusion analysis of multi-dimensional features online and offline, the descriptions of person and institution attributes have been constructed, and the scientific quantification of person influence has been realized using the super IP theory. The description system can be dynamically updated to maintain the timeliness and accuracy of information. It supports real-time processing and efficient prompting. The system can achieve a data processing speed of 300 records per second in a single thread, and the query response for person information remains at the second level. An influence judgment system including no less than 5 categories and 4 levels has been established to achieve the accurate identification and timely prompting of specific information.

[0080] Embodiment 2:

[0081] Based on Embodiment 1, an embodiment of the present invention provides a method for judging specific information based on a large language model. In Step 1, information data is obtained through a combination of API access, reverse analysis of APP, and web crawler, including:

[0082] Through anti-crawling restriction technology, a self-developed distributed crawler program is deployed on a special-purpose computer, and reasonable concurrency and download delay are set to collect data.

[0083] By designing a custom RESTful-style API interface, through an authentication mechanism and security control, the collected data is transmitted to obtain information data.

[0084] The working principle and beneficial effects of the above design solution are as follows: Through technologies such as a dynamic proxy pool and request header rotation, the anti-crawling restriction is broken through. A self-developed distributed crawler program is deployed on a special-purpose computer such as a bastion host. By setting reasonable concurrency and download delay, the stability of the crawler operation is ensured, and the efficient collection of data is realized. A custom RESTful-style API interface is designed to achieve standardized docking with third-party data sources, including a complete authentication mechanism and security control. Self-developed algorithms are used for data cleaning and data standardization processing, including multi-dimensional duplicate data detection. The processed data is stored using a distributed database cluster, with multiple shard nodes and replicas configured, and the sharding strategy and index design are optimized to improve data storage and access efficiency.

[0085] Embodiment 3:

[0086] Based on Example 2, the embodiment of the present invention provides a specific information analysis method based on a large language model, and transmitting the collected data also includes:

[0087] The collected data is transmitted using the HMAC-SHA256 algorithm for request signature authentication.

[0088] The beneficial effect of the above design scheme is: using the HMAC-SHA256 algorithm to perform request signature authentication to ensure the security of data transmission.

[0089] Embodiment 4:

[0090] Based on Example 1, the embodiment of the present invention provides a specific information analysis method based on a large language model. In the step 1, the information data is cleaned and normalized and then stored in a distributed database for unified management, including:

[0091] Based on self-developed algorithms, the information data is cleaned, repeated data is detected in multiple dimensions, outlier processing is done based on statistical methods, and missing value filling is done using multiple strategies;

[0092] In a distributed database, multiple shard nodes and replica sets are used to store and uniformly manage information data.

[0093] The beneficial effects of the above design scheme are: using self-developed algorithms for data cleaning and data standardization, including multi-dimensional duplicate data detection, outlier processing based on statistical methods, missing value filling with multiple strategies, etc., to improve data quality.

[0094] Embodiment 5:

[0095] Based on Example 1, the embodiment of the present invention provides a method for analyzing specific information based on a large language model. In step 2, after preprocessing the information data based on the large language model, a combination of pre-training and fine-tuning is used to extract basic attribute information of the object to be analyzed in the information data, and a relationship knowledge base is constructed based on the basic attribute information, including:

[0096] Deploy a custom pre-trained large language model and use the class normalization method to perform standardized input processing. The calculation formula is as follows:

[0097]

[0098] Among them, x' represents the output vector after preprocessing the information data based on the large language model, n represents the vector dimension, and x i Represents the i-th input vector of information data;

[0099] Optimize the large language model according to the position encoding technology, and the calculation formula is as follows;

[0100] RoPE(x,m,θ) = [x 1 cos(mθ) - x 2 sin(mθ), x 1 sin(mθ) + x 2 cos(mθ)]

[0101] Among them, m represents the position index, θ represents the base frequency parameter, x 1 and x 2 are adjacent dimensions of the input vector; RoPE(x,m,θ) represents the introduction of a position-related rotation transformation;

[0102] Design a specific prompt template to guide the automated extraction of information data from the optimized large language model to obtain the extracted text;

[0103] And based on entity recognition technology, identify the key entities in the extracted text. Based on the entity semantics of the key entities, associate the different expression information data referring to the same entity through entity linking to form a knowledge representation of the basic attribute information of the unified object to be analyzed, and construct a relationship knowledge base according to the knowledge representation.

[0104] The working principle and beneficial effects of the above design scheme are as follows: Deploy an open-source large language model with custom pre-training, use the RMSNorm-like normalization method for standardized input processing, and introduce the RoPE technology to process position information at the same time. RoPE can better capture the position information in the sequence through the introduction of a position-related rotation transformation, improving the model's processing ability for long texts; the performance and stability of the model are improved through the combination of these optimization technologies; based on the strategy of combining pre-training and fine-tuning, design a specific prompt template to guide the model for information extraction, extract multi-dimensional information including personal basic information, institutional attributes, social relationships, etc., and achieve automated information extraction; use entity recognition technology to identify the key entities in the text, including people, geographical locations, etc., and associate the different expression information referring to the same entity through entity linking to form a unified knowledge representation, construct a knowledge base containing 11 attributes, realize the systematic management of personal and institutional information, support flexible information retrieval and relationship query, and form a structured knowledge system.

[0105] Example 6:

[0106] Based on Embodiment 1, an embodiment of the present invention provides a method for judging specific information based on a large language model. In Step 3, online features of the object to be analyzed are obtained through the online dimension, and offline features of the object to be analyzed are obtained through the offline dimension. The online features and offline features are fused to obtain a three-dimensional attribute description, including:

[0107] Construct online features of the object to be analyzed in the online dimension based on online quantization indicators;

[0108] Construct a social network structure based on offline indicators, and calculate the offline centrality of the object to be analyzed according to the following formula;

[0109]

[0110] where Centrality(v) represents the offline centrality of the object to be analyzed, v represents the target node, s represents the node position of any first offline quantization indicator in the social network structure, t represents the node position of any second offline quantization indicator in the social network structure, shortest p aths(s,t|v) represents the number of shortest paths passing through the target node, shortest p aths(s,t) represents the number of all shortest paths in the social network structure;

[0111] Offline features constructed based on offline indicators and offline centrality;

[0112] Perform weighted fusion on the online features and offline features according to the following formula to obtain a comprehensive score;

[0113]

[0114] where Score represents the three-dimensional attribute description, g represents the total number of features of the online features and offline features, w j represents the weight of the jth feature, f j represents the standardized score of the jth feature.

[0115] Construct a three-dimensional attribute description of the object to be analyzed based on the online features, offline features, offline centrality, and comprehensive score.

[0116] The working principle and beneficial effects of the above design solution are as follows: Construct the behavioral characteristics of a person in the online dimension, including the quantification of indicators such as content generation ability, interaction frequency, and influence. By analyzing the activity trajectory of the person on the social media platform, construct the online part of the person attribute description; Integrate information such as the professional background, institutional position, and social relationships of the person in the offline dimension, calculate the centrality of the person using social network analysis technology, design a scientific feature weight system, and perform weighted fusion of the features in the online and offline dimensions to construct a complete three-dimensional description. Through this weighted fusion method, achieve the accurate evaluation and quantitative expression of the person's influence; Support the dynamic update of the attribute description system, and adjust the description content in real time according to the changes in the person's behavior and attributes to maintain the timeliness and accuracy of the attribute description.

[0117] Embodiment 7:

[0118] Based on Embodiment 1, the embodiment of the present invention provides a method for judging specific information based on a large language model. In step 4, based on the sliding window technology, perform information judgment and analysis on the text determined by the relationship knowledge base and the three-dimensional attribute description, and give real-time prompts for abnormal information according to the judgment and analysis results, including:

[0119] Perform topic classification on the text determined by the relationship knowledge base and the three-dimensional attribute description;

[0120] Based on the sliding window technology, set the time window size and sliding step length, and continuously process the change in the discussion volume of each topic within a specific time period to achieve the update of the text;

[0121] Perform text clustering analysis based on the improved LDA topic model. The topic-word probability calculation formula is:

[0122]

[0123] Among them, P(w|d) represents the topic-word probability, w represents the word, d represents the text, z represents the topic, P(w|z) represents the probability of the word under the topic, and P(z|d) represents the probability of the topic in the text;

[0124] Determine the historical data mean and historical data standard deviation of the text based on the topic-word probability;

[0125] Perform anomaly detection based on the statistical threshold. The calculation formula of the threshold is as follows:

[0126] Threshold = μ + k * σ

[0127] Among them, Threshold represents the threshold, μ represents the historical data mean, σ represents the historical data standard deviation, and k represents the adjustment parameter;

[0128] Based on the four-level impact rating standard with different settings of adjustment parameters, determine the rating thresholds corresponding to each level. When the real-time threshold of the latest text exceeds the rating threshold, trigger a real-time prompt.

[0129] The working principle and beneficial effects of the above design scheme are as follows: The sliding window technology is adopted to realize the real-time processing of information flow. By setting the time window size and sliding step length, continuously process the change in the discussion volume of each topic within a specific time period; Use the improved LDA topic model for text clustering analysis. Through this improved model, automatically discover and track hot topics, and combine semantic analysis technology to identify abnormal information patterns. Use statistical thresholds for anomaly detection, set a four-level impact rating standard, and form a clear hierarchical system from the highest-level specific domain security threats to general daily information. Take corresponding disposal measures for different levels; Realize the visual display of prompt information, including various display methods such as keyword cloud maps and trend charts, which is convenient for relevant personnel to quickly master the prompt information and take corresponding measures.

[0130] Example 8:

[0131] An embodiment of the present invention provides a specific information research and judgment system based on a large language model, as Figure 2 shown, including:

[0132] A multi-channel data acquisition and management module, which is used to obtain information data based on a combination of API access, reverse analysis of APPs, and web crawlers, and after cleaning and normalizing the information data, store it in a distributed database for unified management;

[0133] An information extraction and knowledge base module, which is used to preprocess information data based on a large language model, and then extract the basic attribute information of the objects to be analyzed in the information data by combining pre-training and fine-tuning, and construct a relationship knowledge base based on the basic attribute information;

[0134] A feature processing module, which is used to obtain the online features of the object to be analyzed through the online dimension, obtain the offline features of the object to be analyzed through the offline dimension, and perform feature fusion on the online features and offline features to obtain a three-dimensional attribute description;

[0135] An information research and judgment module, which is used to perform information research and judgment analysis on the text determined by the relationship knowledge base and the three-dimensional attribute description based on the sliding window technology, and give a real-time prompt for abnormal information according to the research and judgment analysis results.

[0136] In this embodiment, the information data is, for example, social media and news website data.

[0137] In this embodiment, the relationship knowledge base contains a knowledge base with multiple attributes such as foreign names, foreign abbreviations, Chinese names, date of birth, etc.

[0138] In this embodiment, the online dimension includes features such as personnel behavior, content generation, and dissemination impact; the offline dimension includes features such as professional background, institutional position, and social relationships.

[0139] In this embodiment, specific information is sorted out and judged based on the sliding window technology. Relevant remarks are classified according to the theme and divided into four levels from one to four according to the degree of influence. The LDA topic model and semantic analysis method are used to identify abnormal information patterns, and information exceeding the set threshold is prompted in real time.

[0140] The beneficial effects of the above design scheme are as follows: The efficiency and quality of data collection are improved: Through a unified data collection framework and a standardized data processing process, the efficient collection and management of multi-source heterogeneous data are realized. The system supports the access and collection of no less than 10 pieces of information, and the data collection communication supports real-time data transmission of more than 10 seconds, improving the accuracy of information extraction and management. Based on the information extraction method of the large language model, the automatic extraction accuracy of the necessary fields exceeds 80%, and the field accuracy in the delivered knowledge base exceeds 95%. A person information database has been successfully constructed, containing no less than 8 types of attribute information and capable of realizing dynamic updates; a three-dimensional person attribute description system has been realized: Through the fusion analysis of multi-dimensional features online and offline, the descriptions of person and institution attributes are constructed, and the scientific quantification of the influence of a person is realized by using the super IP theory. The description system can be dynamically updated to maintain the timeliness and accuracy of information; Support real-time processing and efficient prompting: The system can achieve a data processing speed of 300 pieces / second in a single thread, and the query response to person information remains at the second level; An influence judgment system containing no less than 5 categories and 4 levels has been established to realize the accurate identification and timely prompting of specific information.

[0141] The system of the present invention adopts a hierarchical architecture design, as Figure 3 shown, mainly including the following levels:

[0142] 1. Data collection layer: Responsible for collecting data through distributed crawlers and API interfaces. The system supports the data collection of news website platforms, and the collected content includes basic information, information on participated activities, public homepage information, relevant public report information, etc.

[0143] 2. Data preprocessing layer: Responsible for cleaning and standardizing the collected raw data. Data cleaning is performed through self-developed algorithms, including multi-dimensional duplicate data detection, outlier processing, missing value filling, etc. At the same time, the data format is standardized to ensure the quality of the data.

[0144] 3. Information Extraction Layer: Information extraction is performed based on the Llama 3 large language model. The input is standardized using a class RMSNorm normalization method, the commonly used ReLU activation function is replaced with SwiGLU, and the RoPE technology is introduced to process position information. By designing specific prompt templates, the automated extraction of entity information and relationship information is achieved.

[0145] 4. Knowledge Base Construction Layer: It includes three main parts: the person information database, the attribute description system, and the label knowledge base. The person information database stores the basic information of persons; the attribute description system defines a unified attribute expression framework; the label knowledge base classifies and marks information.

[0146] 5. Feature Processing Layer: The attribute description system is constructed based on the Super IP theory. The online feature analysis unit analyzes the behavioral characteristics of different persons on social media; the offline feature analysis unit integrates information such as the professional background and institutional positions of persons; finally, a complete three-dimensional attribute description is formed through feature fusion.

[0147] 6. Judgment and Prompt Layer: The sliding window technology is used to achieve real-time processing. Text clustering is performed through an improved LDA topic model, and specific information is classified and prompted at different levels.

[0148] 7. Application Display Layer: It provides a friendly user interface and realizes functions such as data visualization display, interactive query, and analysis.

[0149] This system is deployed using a distributed architecture. On the hardware environment, high-performance servers with Hygon 7390 CPUs, 128G of memory, and 2T of hard disks are configured. A total of 5 application servers, 4 crawler servers, and 3 database servers are deployed in the system. All servers are interconnected through a gigabit switch, and a firewall is configured for security protection. In terms of the software environment, the system runs on the Ubuntu 21.04 operating system, uses MongoDB as the main data storage system, uses Redis for cache management, realizes the message queue through ZeroMQ, and uses Elasticsearch to provide search services. The system uses Flask as the Web development framework, uses Scrapy to realize the data collection function, and is mainly developed using Python 3.9 and above versions.

[0150] In practical applications, system administrators first need to configure data sources through the management interface. The system supports data collection from at least 10 major platforms. Administrators can set the collection frequency and configure the field content to be collected, including basic person information, speech and activity information, public home pages, etc. At the same time, administrators can also set data cleaning rules to ensure the quality of the collected data. During the data collection process, the system will automatically handle abnormal situations such as anti-crawler restrictions and data missing to ensure the stability and integrity of data collection.

[0151] In terms of information repository management, the system provides a complete function for maintaining the person information repository. The administrator can quickly establish the person information repository through batch import, support the editing and maintenance of no less than 8 types of attribute information, including basic attributes, social relationships, influence indicators, etc. The system also supports flexible information retrieval and export functions, facilitating users to conduct data analysis and utilization. For the management of the person attribute description system, the system supports viewing, editing, updating, and exporting, and users can dynamically adjust relevant information according to actual needs.

[0152] In terms of prompt management, the system implements a complete prompt configuration and processing process. The administrator can set prompt rules, adjust prompt thresholds, and configure different levels of prompt grades according to actual needs. When the system detects specific information, it will automatically trigger the prompt mechanism, generate prompt information, and display it to the user through a visual interface. The prompt information includes a detailed analysis report, supports manual review and rule adjustment, ensuring the accuracy and effectiveness of the prompt.

[0153] This system demonstrates good performance indicators during actual operation. In terms of data processing, the system can achieve a processing speed of 300 records per second in single-threaded mode, and the accuracy rate of the necessary fields for information extraction reaches over 80%. The system response time remains at the second level. To ensure the stable operation of the system, the system implements a complete operation and maintenance management system, including system status monitoring, data backup and recovery, log management, performance monitoring, etc. In terms of security management, the system implements strict user permission management, data access control, operation log auditing and other mechanisms to ensure the security of the system and data.

[0154] Through deployment and operation tests, it shows that this system can effectively meet the needs of collecting, analyzing, and judging certain person and organization information, providing reliable technical support for relevant decision-making. The system has good scalability and maintainability, can be functionally extended and optimized according to actual needs, and adapts to the changing business requirements.

[0155] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of this application document and its equivalent technologies, the present invention also intends to include these changes and modifications.

Claims

1. A specific information analysis method based on a large language model, characterized in that: include: Step 1: Acquire information data through a combination of API access, reverse analysis APP and web crawlers, and clean and normalize the information data before storing it in a distributed database for unified management; Step 2: After preprocessing the information data based on the large language model, the basic attribute information of the object to be analyzed in the information data is extracted by combining pre-training and fine-tuning, and a relational knowledge base is constructed based on the basic attribute information; Step 3: Obtain online features of the object to be analyzed through the online dimension, obtain offline features of the object to be analyzed through the offline dimension, and fuse the online features and offline features to obtain a three-dimensional attribute description; Step 4: Based on the sliding window technology, perform information analysis on the text determined by the relationship knowledge base and the three-dimensional attribute description, and provide real-time prompts for abnormal information based on the analysis results.

2. According to claim 1, a specific information analysis method based on a large language model is characterized in that: In the step 1, information data is obtained by combining API access, reverse analysis of APP and web crawlers, including: Through anti-crawling restriction technology, we deploy self-developed distributed crawler programs on special-purpose computers, set reasonable concurrency numbers and download delays, and collect data; By designing a custom RESTful-style API interface, the collected data is transmitted through authentication mechanism and security control to obtain information data.

3. According to claim 2, a specific information analysis method based on a large language model is characterized in that: Transmitting the collected data to the data also includes: The collected data is transmitted using the HMAC-SHA256 algorithm for request signature authentication.

4. According to claim 1, a method for analyzing specific information based on a large language model is characterized in that: In the step 1, the information data is cleaned and normalized and then stored in a distributed database for unified management, including: Based on self-developed algorithms, the information data is cleaned, repeated data is detected in multiple dimensions, outlier processing is done based on statistical methods, and missing value filling is done using multiple strategies; In a distributed database, multiple shard nodes and replica sets are used to store and uniformly manage information data.

5. According to the method of specific information analysis based on a large language model as described in claim 1, it is characterized in that: In the step 2, after preprocessing the information data based on the large language model, the basic attribute information of the object to be analyzed in the information data is extracted by combining pre-training and fine-tuning, and a relational knowledge base is constructed based on the basic attribute information, including: Deploy a custom pre-trained large language model and use the class normalization method to perform standardized input processing. The calculation formula is as follows: Among them, x , represents the output vector after preprocessing the information data based on the large language model, n represents the vector dimension, x i Represents the i-th input vector of information data; The large language model is optimized based on the position encoding technology. The calculation formula is as follows; RoPE(x,m,θ)=[x1cos(mθ)-x2sin(mθ),x1sin(mθ)+x2cos(mθ)] Where m represents the position index, θ represents the base frequency parameter, x1 and x2 are adjacent dimensions of the input vector; RoPE(x,m,θ) represents the rotation transformation by introducing position-related; Design a specific prompt template to guide the automatic extraction of information data from the optimized large language model to obtain the extracted text; Based on entity recognition technology, key entities in the extracted text are identified and extracted. Based on the entity semantics of key entities, different expression information data referring to the same entity are associated through entity linking to form a unified knowledge representation of the basic attribute information of the object to be analyzed, and a relational knowledge base is constructed based on the knowledge representation.

6. The specific information analysis method based on a large language model according to claim 1 is characterized in that: In the step three, the online features of the object to be analyzed are obtained through the online dimension, the offline features of the object to be analyzed are obtained through the offline dimension, and the online features and the offline features are fused to obtain a three-dimensional attribute description, including: Construct online features of the object to be analyzed in the online dimension based on online quantitative indicators; The social network structure is constructed based on offline indicators, and the offline centrality of the object to be analyzed is calculated according to the following formula; Among them, Centrality(v) represents the offline centrality of the object to be analyzed, v represents the target node, s represents the node position of any first offline quantitative indicator in the social network structure, t represents the node position of any second offline quantitative indicator in the social network structure, shortest p aths(s,t|v) represents the number of shortest paths passing through the target node. p aths(s,t) represents the number of all shortest paths in the social network structure; Offline features constructed based on offline indicators and offline centrality; Based on the following formula, online features and offline features are weighted and integrated to obtain a comprehensive score; Among them, Score represents the three-dimensional attribute description, g represents the total number of online features and offline features, and w j represents the weight of the jth feature, f j represents the standardized score of the jth feature; A three-dimensional attribute description of the object to be analyzed is constructed based on the online features, offline features, offline centrality and comprehensive scores.

7. The specific information analysis method based on a large language model according to claim 1 is characterized in that: In the fourth step, based on the sliding window technology, information analysis is performed on the text determined by the relationship knowledge base and the three-dimensional attribute description, and abnormal information is prompted in real time according to the analysis results, including: Topic classification of texts determined by relational knowledge base and stereo attribute description; Based on the sliding window technology, the time window size and sliding step size are set to continuously process the changes in the discussion volume of each topic within a specific time period to update the text; Based on the improved LDA topic model, text clustering analysis is performed, and the topic-word probability calculation formula is: Among them, P(w|d) represents the topic-word probability, w represents the word, d represents the text, z represents the topic, P(w|z) represents the probability of the word under the topic, and P(z|d) represents the probability of the topic in the text; Determine the historical data mean and historical data standard deviation of the text based on the topic-word probability; Anomaly detection is performed based on statistical thresholds. The threshold calculation formula is as follows: Threshold=μ+k*σ Among them, Threshold represents the threshold, μ represents the mean of historical data, σ represents the standard deviation of historical data, and k represents the adjustment parameter; Based on the different settings of adjustment parameters, four-level impact level assessment standards are set to determine the level threshold corresponding to each level. When the real-time threshold of the latest text monitored exceeds the level threshold, a real-time prompt is triggered.

8. The specific information analysis method based on a large language model according to claim 7 is characterized in that: The triggering real-time prompt is expressed in the form of a keyword cloud map and a trend chart.

9. The specific information analysis method based on a large language model according to claim 7 is characterized in that: The values ​​of the adjustment parameters are different under different impact level assessment standards, and are pre-set based on actual conditions.

10. A specific information analysis system based on a large language model, used to implement the steps of the method according to any one of claims 1 to 9, characterized in that: include: Multi-channel data collection and management module, used to obtain information data based on a combination of API access, reverse analysis APP and web crawlers, and to clean and normalize the information data and store it in a distributed database for unified management; The information extraction and knowledge base module is used to pre-process the information data based on the large language model, extract the basic attribute information of the object to be analyzed in the information data by combining pre-training and fine-tuning, and build a relational knowledge base based on the basic attribute information; A feature processing module is used to obtain online features of the object to be analyzed through the online dimension, obtain offline features of the object to be analyzed through the offline dimension, and perform feature fusion on the online features and the offline features to obtain a three-dimensional attribute description; The information analysis module is used to perform information analysis on the text determined by the relationship knowledge base and the three-dimensional attribute description based on the sliding window technology, and to provide real-time prompts for abnormal information based on the analysis results.

Citation Information

Cited By

  • Data classification and grading method and system based on large language model

    CN120950688A