Information identification method and system based on network index

By combining information identification systems with network crawlers, inverted indexes and deep learning models, the problem that information monitoring systems in the prior art are difficult to capture complex public opinion trends, and the rapid and accurate identification and analysis of network information, especially the detection of false information.

CN120508655APending Publication Date: 2025-08-19中科天玑数据科技股份有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510510568.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

Existing information monitoring systems are difficult to capture complex public opinion trends efficiently and accurately, especially in network information monitoring and analysis. Keyword matching and simple text classification methods are not enough to deal with multilingual and multi-dimensional information changes.

Method used

Optimized network crawling technology, inverted index structure and deep learning model are adopted, and combined with natural language processing and time series analysis, an information recognition system is built, and abnormal peaks of network information are identified and monitored through distributed storage and multilingual sentiment analysis.

Benefits of technology

It realizes rapid and accurate identification and analysis of network information, can detect false information and malicious dissemination behavior, and improves the accuracy and efficiency of information monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120508655A_ABST
    Figure CN120508655A_ABST
Patent Text Reader

Abstract

The invention provides an information identification method and system based on a network index. The method comprises the following steps: capturing network text data from a plurality of online resources by using an optimized web crawler; constructing and maintaining an inverted index based on the web text data, wherein the inverted index is used for accelerating retrieval of keywords in the captured web text data; performing natural language data processing on the web text data; performing sentiment analysis and subject classification on the web text data after natural language data processing by using the trained deep learning model to obtain an analysis result; a time sequence model is constructed according to the reverse index and the analysis result, the time sequence model is suitable for data fluctuation characteristics in different time periods, the abnormal peak value of the information is judged according to the time sequence model, and the method can effectively, rapidly and accurately recognize and analyze the information on the network, so that the network quality is improved. The accuracy and efficiency of information identification are improved, and powerful data support is provided for decision makers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of big data technology, and in particular relates to an information identification method and system based on network indexing. Background Art

[0002] With the development of the Internet and social media, there is a lot of information online about various hot news, discussions, comments, etc. How to monitor, analyze and manage the information in a timely manner has become a difficult task we face.

[0003] This information encompasses the social attitudes formed and held by the public, as the subject, regarding the occurrence, development, and changes of intermediary social events within a given social space—the social managers, businesses, individuals, and other organizations, as the subjects, and their social and moral orientations. It is the sum total of the beliefs, attitudes, opinions, and emotions expressed by a broad range of people regarding various social phenomena and issues.

[0004] After an event occurs, the general public will learn the truth of the matter through various channels, followed by a flood of comments. These comments include rational and objective evaluations, exaggerated expressions, and even subjective remarks that incite subversion.

[0005] Previous information monitoring systems often relied on keyword matching and simple text classification, making it difficult to capture complex public opinion trends efficiently and accurately.

[0006] Therefore, it is particularly important to monitor, analyze and guide network information. Summary of the Invention

[0007] This application provides an information identification method based on network indexing. This solution combines advanced web crawler technology, inverted index structure and deep learning model to realize the identification and analysis of information in large-scale network data.

[0008] In a first aspect, an embodiment of the present application provides an information identification method based on network indexing, the method comprising using an optimized web crawler to crawl network text data from multiple online resources, the crawler automatically detecting and updating a target website list according to preset rules; constructing and maintaining an inverted index based on the network text data, the inverted index being used to accelerate the retrieval of keywords in the crawled network text data, and performing distributed storage and query optimization on the inverted index; performing natural language data processing on the network text data, including but not limited to word segmentation, stop word removal, part-of-speech tagging, named entity recognition, and semantic role tagging; training a deep learning model with labeled information data, the trained deep learning model having the ability to distinguish between positive and negative emotions and multiple emotion intensities, performing sentiment analysis and topic classification on the network text data after natural language data processing using the trained deep learning model to obtain analysis results; constructing a time series model based on the inverted index and analysis results, the time series model being applicable to data fluctuation characteristics in different time periods, and judging abnormal peaks of information based on the time series model.

[0009] By adopting the above scheme, the network index-based information identification method of this scheme can combine advanced web crawler technology, inverted index structure and deep learning model to achieve rapid and accurate identification and analysis of information in large-scale online text data on the network.

[0010] In some embodiments of the present invention, before using an optimized web crawler to crawl web text data from multiple online resources, a list of target websites to be crawled and their update frequency are determined, and the crawling is timely and the latest information is obtained according to a preset time range.

[0011] In some embodiments of the present invention, the inverted index is a distributed index structure that adopts a hash partitioning or range partitioning strategy.

[0012] In some embodiments of the present invention, the natural language processing further includes establishing a customized vocabulary for specialized terms in a specific field.

[0013] In some embodiments of the present invention, the natural language data processing also includes: data cleaning, data deduplication, and data sampling; during the data cleaning process, the hyperlinks and content in brackets in the network text data are retained; the data deduplication is to sort the network text data according to the collection time, and remove duplicate data according to the cluster authentication code to which the network text data belongs; the data sampling is to sort the weights of the keywords, filter out keywords with low weights, and sample and delete the keywords with low weights.

[0014] In some embodiments of the present invention, the deep learning model is trained by one or more combinations of a recursive neural network, a long short-term memory network, a gated recurrent unit, and a transformer structure that supports multiple languages.

[0015] In some embodiments of the present invention, an interface is further included for visually displaying the analysis results to a user, wherein the interface supports user-defined query conditions and display parameters, and allows the user to annotate and evaluate the analysis results.

[0016] In some embodiments of the present invention, an anomaly detection module is also included for performing anomaly detection on network text data after natural language data processing, for identifying potential false information or malicious dissemination behavior. The anomaly detection module monitors the network text data in real time based on a machine learning algorithm.

[0017] In some embodiments of the present invention, the anomaly detection module includes but is not limited to multi-dimensional feature extraction, social network propagation path analysis, content consistency check, user behavior analysis, collaborative filtering and group behavior analysis, and adaptive learning mechanism.

[0018] The above scheme adopts the existing information monitoring method that relies too much on keyword matching and text classification. This scheme improves the accuracy of word segmentation by using multiple data processing methods of natural language processing for network text data and establishing a customized vocabulary, combined with an inverted index structure that supports efficient management and query of large-scale data sets; through the analysis of network text data by supporting multi-language deep learning models, it realizes the recognition and analysis of information in multi-language environments, and through the anomaly detection module, it analyzes the information change trend over time and the characteristics of multiple dimensions of text content to realize the diagnosis and monitoring of false information or malicious dissemination behavior on the Internet.

[0019] In second aspect, an embodiment of the present application provides an information identification system based on network index, which includes a computer device, the computer device including a processor and a memory, the memory storing computer instructions, the processor being used to execute the computer instructions stored in the memory, and when the computer instructions are executed by the processor, the system implements the information identification method based on network index.

[0020] Additional advantages, objects, and features of the present invention will be described in part in the following description and will become apparent to those skilled in the art after studying the following or may be learned by practice of the present invention. The objects and other advantages of the present invention may be particularly pointed out and attained in the description and drawings.

[0021] Those skilled in the art will understand that the purposes and advantages that can be achieved by the present invention are not limited to the above specific descriptions, and the above and other purposes that can be achieved by the present invention will be more clearly understood based on the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The accompanying drawings are used to provide further understanding of the present disclosure and constitute a part of the specification. Together with the following detailed description, they are used to explain the present disclosure, but do not constitute a limitation of the present disclosure.

[0023] In the attached figure:

[0024] Figure 1 is a schematic diagram of an implementation of the network index-based information identification method;

[0025] Figure 2 It is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0026] In order to more clearly understand the above-mentioned objectives, features and advantages of the present disclosure, the scheme of the present disclosure will be further described below. It should be noted that the embodiments of the present disclosure and the features therein can be combined with each other in the absence of conflict.

[0027] In the following description, many specific details are set forth to facilitate a full understanding of the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; it is obvious that the embodiments in the specification are only part of the embodiments of the present disclosure, rather than all of the embodiments.

[0028] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprises" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a..." do not exclude the presence of other identical elements in the process, method, article or device that includes the elements.

[0029] Generally speaking, the existing network information supervision system is difficult to capture complex public opinion trends efficiently and accurately.

[0030] Therefore, in order to supervise network information, the present application provides an information identification method and system based on network indexing.

[0031] Figure 1 This is a schematic diagram of an information identification method based on network indexing provided in one embodiment of the present application.

[0032] like Figure 1 As shown, an embodiment of the present application provides an information identification method based on a network index, the method comprising the following steps:

[0033] S1: Use an optimized web crawler to crawl web text data from multiple online resources. The crawler automatically detects and updates the target website list according to preset rules; uses an intelligent algorithm to evaluate the importance of the target website and prioritizes crawling high-weight sites.

[0034] The preset rules are generated based on the robots.txt file of the target website, frequency limits for crawling requests, content-based crawling limits (such as sensitive data such as personal privacy, financial information, medical records, etc.), crawling depth or time limits, and a handling mechanism for crawling errors.

[0035] S2: constructing and maintaining an inverted index based on the network text data, wherein the inverted index is used to accelerate the retrieval of keywords in the captured network text data, and performing distributed storage and query optimization on the inverted index.

[0036] An inverted index is a data structure used to quickly find documents containing specific terms. It enables efficient query processing by mapping terms to lists of documents containing those terms. The key is a term, and the value is a list of documents or document IDs. Using an inverted index can support efficient management and querying of large datasets. Distributed storage and query optimization of the inverted index can reduce storage space and improve query efficiency.

[0037] S3: Processing the web text data using natural language processing, including but not limited to word segmentation, stop word removal, part-of-speech tagging, named entity recognition, and semantic role labeling. Processing the web text data using natural language processing methods such as word segmentation, stemming, stop word elimination, morphological processing, and noise filtering can construct a streamlined, standardized web text corpus, improving data quality and usability. The word segmentation method also includes constructing a factor combination corresponding to the keyword based on the number and position of occurrences of the web text data in articles on different platforms. The keyword factor combination is used to describe the logical relationship between different keywords.

[0038] S4: The deep learning model is trained with the labeled information data. The trained deep learning model has the ability to distinguish between positive and negative emotions and multiple emotion intensities. The trained deep learning model is used to perform sentiment analysis and topic classification on the network text data after natural language data processing to obtain analysis results.

[0039] S5: Constructing a time series model based on the inverted index and the analysis results. The time series model is applicable to data fluctuation characteristics in different time periods, and abnormal peaks of information are determined based on the time series model.

[0040] By adopting the above scheme, the network index-based information identification method of this scheme can combine advanced web crawler technology, inverted index structure and deep learning model, and analyze the changing trend of information over time by using time series model, and then detect abnormal peak conditions of information, so as to realize the rapid and accurate identification and analysis of information in large-scale online text data on the network.

[0041] In some embodiments of the present invention, before step S1, a list of target websites to be crawled and their update frequency are determined, and the crawling is timely and the latest information is obtained according to a preset time range.

[0042] In some embodiments of the present invention, the inverted index is a structure that adopts a hash partitioning or range partitioning strategy; an inverted index that adopts such a structure can ensure load balancing and scalability.

[0043] In some embodiments of the present invention, the natural language processing further includes establishing a customized vocabulary for specialized terms in a specific field, which can improve the accuracy of word segmentation.

[0044] In the specific implementation process, the natural language processing in this solution also includes word segmentation of the network text data, and word frequency statistics and synonym conversion of keywords.

[0045] In some embodiments of the present invention, the natural language data processing also includes: data cleaning, data deduplication, and data sampling; during the data cleaning process, the hyperlinks and content in brackets in the network text data are retained; the data deduplication is to sort the network text data according to the collection time, and remove duplicate data according to the cluster authentication code to which the network text data belongs; the data sampling is to sort the weights of the keywords, filter out keywords with low weights, and sample and delete the keywords with low weights.

[0046] In the specific implementation process, by performing the above-mentioned processing on network text data, the data processing results are made faster and more accurate.

[0047] In some embodiments of the present invention, the deep learning model is trained using one or more combinations of a multi-language recurrent neural network (RNN), a long short-term memory network (LTSM), a gated recurrent unit (GRU), and a transformer neural network. Using a model with multiple combinations involves training the deep learning model by adding another neural network, such as an LTSM, to the training unit of a neural network, such as a transformer, to achieve a faster training process and a more intelligent deep learning model after training.

[0048] In some embodiments of the present invention, an interface is further included for visually displaying the analysis results to a user, wherein the interface supports user-defined query conditions and display parameters, and allows the user to annotate and evaluate the analysis results.

[0049] In some embodiments of the present invention, an anomaly detection module is also included for performing anomaly detection on network text data after natural language data processing, for identifying potential false information or malicious dissemination behavior. The anomaly detection module monitors the network text data in real time based on a machine learning algorithm.

[0050] In some embodiments of the present invention, the anomaly detection module includes but is not limited to multi-dimensional feature extraction, social network propagation path analysis, content consistency check, user behavior analysis, collaborative filtering and group behavior analysis, and adaptive learning mechanism.

[0051] In its implementation, multi-dimensional feature extraction involves extracting features from text content across multiple dimensions, including but not limited to sentiment, vocabulary frequency, and user interaction patterns. Social network communication path analysis involves constructing a communication graph to track the path of information between different nodes or platforms. Content consistency checking involves comparing content from different sources on the same topic to identify highly similar or large-scale misleading content from multiple platforms or resources. User behavior analysis involves creating profiles of user behavior, recording information such as posting habits, activity time periods, and frequently used topics. Collaborative filtering and group behavior analysis involve detecting abnormal behavior in a user and conducting extended analysis of their associated users. The adaptive learning mechanism involves the anomaly detection module continuously learning from new cases to update its own algorithms and rules.

[0052] The above solution can solve the problem of existing information monitoring methods being overly dependent on keyword matching and text classification. This solution improves the accuracy of word segmentation by using multiple data processing methods of natural language processing for network text data and establishing a customized vocabulary, combined with an inverted index structure that supports efficient management and query of large-scale data sets; it analyzes network text data through a deep learning model that supports multiple languages, thereby realizing the recognition and analysis of information in a multilingual environment; it forms a feedback mechanism through user annotation and evaluation of analysis results, thereby realizing continuous improvement of the performance and accuracy of the deep learning model; and it uses an anomaly detection module to analyze the changing trends of information over time and the characteristics of multiple dimensions of text content to realize the diagnosis and monitoring of false information or malicious dissemination behaviors on the network.

[0053] In second aspect, an embodiment of the present application provides an information identification system based on network index, which includes a computer device, the computer device including a processor and a memory, the memory storing computer instructions, the processor being used to execute the computer instructions stored in the memory, and when the computer instructions are executed by the processor, the system implements the steps implemented by the information identification method based on network index.

[0054] In a third aspect, an embodiment of the present application provides a computer-readable storage medium having computer program instructions stored thereon. When the computer program instructions are executed by a processor, the above-mentioned network index-based information identification system is implemented.

[0055] Figure 2 It is a structural diagram of an electronic device provided in one embodiment of the present application.

[0056] like Figure 2 As shown, an embodiment of the present application provides an electronic device, which includes: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, the above-mentioned network index-based information identification system is implemented.

[0057] The electronic device may include a processor 1201 and a memory 1202 storing computer program instructions.

[0058] Specifically, the processor 1201 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.

[0059] The memory 1202 may include a large capacity memory for data or instructions. By way of example and not limitation, the memory 1202 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory 1202 may include removable or non-removable (or fixed) media. Where appropriate, the memory 1202 may be inside or outside the integrated gateway disaster recovery device. In a specific embodiment, the memory 1202 is a non-volatile solid-state memory.

[0060] The memory may include read-only memory (ROM), random access memory (RAM), magnetic disk storage media devices, optical storage media devices, flash memory devices, electrical, optical or other physical / tangible memory storage devices. Thus, generally, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to an aspect of the present disclosure.

[0061] The processor 1201 reads and executes computer program instructions stored in the memory 1202 to implement any one of the methods for determining battery thermal runaway parameters in the above embodiments.

[0062] In one example, the electronic device may further include a communication interface 1203 and a bus 1210. Figure 2 As shown, the processor 1201 , the memory 1202 , and the communication interface 1203 are connected via a bus 1210 and communicate with each other.

[0063] The communication interface 1203 is mainly used to implement communication between various modules, devices, units and / or equipment in the embodiments of the present application.

[0064] Bus 1210 includes hardware, software or both, couples the parts of electronic equipment to each other.For example, but not limitation, bus can include accelerated graphics port (AGP) or other graphics bus, enhanced industry standard architecture (EISA) bus, front side bus (FSB), hypertransport (HT) interconnection, industry standard architecture (ISA) bus, infinite bandwidth interconnection, low pin count (LPC) bus, memory bus, micro channel architecture (MCA) bus, peripheral component interconnection (PCI) bus, PCI-Express (PCI-X) bus, serial advanced technology attachment (SATA) bus, video electronics standard association local (VLB) bus or other suitable bus or two or more of these combinations. In appropriate cases, bus 1210 can include one or more buses. Although the present application embodiment describes and shows specific bus, the application considers any suitable bus or interconnection.

[0065] It should be understood that the present application is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, a detailed description of known methods is omitted here. In the above embodiments, several specific steps are described and illustrated as examples. However, the method process of the present application is not limited to the specific steps described and illustrated. Those skilled in the art can make various changes, modifications, and additions, or change the order of the steps after understanding the spirit of the present application.

[0066] The functional blocks shown in the above-described block diagram can be implemented as hardware, software, firmware or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of the present application are programs or code segments that are used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link by a data signal carried in a carrier wave. "Machine-readable medium" can include any medium that can store or transmit information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROMs, flash memories, erasable ROMs (EROMs), floppy disks, CD-ROMs, optical disks, hard disks, optical fiber media, radio frequency (RF) links, etc. The code segment can be downloaded via a computer network such as the Internet, an intranet, etc.

[0067] It should also be noted that the exemplary embodiments mentioned in this application describe some methods or systems based on a series of steps or devices. However, this application is not limited to the order of the above steps. In other words, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.

[0068] Aspects of the present disclosure have been described above with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present disclosure. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer or other programmable data processing device to produce a machine so that these instructions executed by the processor of the computer or other programmable data processing device enable the implementation of the function / action specified in one or more boxes of the flowchart and / or block diagram. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor or a field programmable logic circuit. It is also understood that each box in the block diagram and / or flowchart and the combination of the boxes in the block diagram and / or flowchart can also be implemented by dedicated hardware that performs the specified function or action, or can be implemented by a combination of dedicated hardware and computer instructions.

[0069] The above description is only a specific embodiment of the present application. Those skilled in the art will clearly understand that for the convenience and brevity of description, the specific working processes of the systems, modules and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. It should be understood that the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical scope disclosed in the present application, and these modifications or replacements should be included in the scope of protection of the present application.

Claims

1. A method for identifying information based on network indexing, characterized in that: The method comprises the following steps: S1: Using an optimized web crawler to crawl web text data from multiple online resources, the crawler automatically detects and updates a list of target websites according to preset rules; S2: constructing and maintaining an inverted index based on the web text data, wherein the inverted index is used to accelerate the retrieval of keywords in the captured web text data, and performing distributed storage and query optimization on the inverted index; S3: performing natural language data processing on the network text data, including but not limited to word segmentation, stop word removal, part-of-speech tagging, named entity recognition, and semantic role labeling; S4: Using labeled information data to train a deep learning model, wherein the trained deep learning model has the ability to distinguish between positive and negative emotions and multiple emotion intensities, and using the trained deep learning model to perform sentiment analysis and topic classification on network text data processed from natural language data to obtain analysis results; S5: Constructing a time series model based on the inverted index and the analysis results. The time series model is applicable to data fluctuation characteristics in different time periods, and abnormal peaks of information are determined based on the time series model.

2. The information identification method based on network index according to claim 1, characterized in that: The step S1 further includes: before step S1, determining a list of target websites to be crawled and their update frequency, and determining whether the crawling is timely and the latest information is obtained according to a preset time range.

3. The information identification method based on network index according to claim 1, characterized in that: The inverted index is an index structure that adopts a hash partitioning or range partitioning strategy.

4. The information identification method based on network index according to claim 1, characterized in that: The natural language processing also includes establishing a customized vocabulary for specialized terminology in a specific field.

5. The information identification method based on network index according to claim 1 or 2, characterized in that: The natural language data processing also includes: data cleaning, data deduplication, and data sampling; during the data cleaning process, the hyperlinks and content in brackets in the network text data are retained; the data deduplication is to sort the network text data according to the collection time, and remove duplicate data according to the cluster authentication code to which the network text data belongs; the data sampling is to sort the weights of the keywords, filter out keywords with low weights, and sample and delete the keywords with low weights.

6. The information identification method based on network index according to claim 1, characterized in that: The deep learning model is trained through one or more combinations of recursive neural networks, long short-term memory networks, gated recurrent units, and transformer structures that support multiple languages.

7. The information identification method based on network index according to claim 6, characterized in that: It also includes an interface for visually displaying the analysis results to the user, wherein the interface supports user-defined query conditions and display parameters, and allows the user to mark and evaluate the analysis results.

8. The information identification method based on network index according to claim 1, characterized in that: It also includes an anomaly detection module for performing anomaly detection on network text data after natural language data processing. The anomaly detection module is used to identify potential false information or malicious dissemination behavior. The anomaly detection module monitors the network text data in real time based on a machine learning algorithm.

9. The information identification method based on network indexing according to claim 8, characterized in that: The anomaly detection module includes but is not limited to multi-dimensional feature extraction, social network propagation path analysis, content consistency check, user behavior analysis, collaborative filtering and group behavior analysis, and adaptive learning mechanism.

10. An information identification system based on network indexing, characterized in that: The system includes a computer device, which includes a processor and a memory. The memory stores computer instructions, and the processor is used to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the system implements the information identification method based on network index as described in any one of claims 1 to 9.