A network data authenticity analysis method and system
By dynamically dividing data levels according to the propagation path length and adaptive sampling of the propagation level, combined with a long text priority strategy, the problem of high efficiency and low efficiency caused by random processing of network data is solved, and efficient and accurate network data authenticity analysis is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-11
- Publication Date
- 2026-04-07
AI Technical Summary
In existing technologies, random processing of crawled web data results in a large amount of unprocessed data, which affects analysis efficiency.
We dynamically divide data into levels based on the length of the propagation path, focusing on high-impact data. We determine the set of target authenticity levels through adaptive sampling of propagation levels and a long text priority strategy, and combine semantic tags to conduct authenticity analysis of network data.
While ensuring the accuracy of the analysis, we aim to reduce the amount of data processing and improve the relevance and efficiency of the analysis.
Smart Images

Figure CN120950755B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of data analysis technology, and in particular relates to a method and system for analyzing the authenticity of network data. Background Technology
[0002] With the rapid development of social media and mobile internet, online data is experiencing explosive growth, generating billions of pieces of public content daily. False information spreads virally through multi-level marketing mechanisms, amplifying its impact exponentially; for example, a highly disseminated rumor can reach over a million users within three hours. Traditional manual verification methods have become completely ineffective in terms of data scale and response speed. Current technologies typically employ random processing of crawled online data, often resulting in a large amount of unprocessed data, thus impacting analysis efficiency. Summary of the Invention
[0003] This invention provides a method and system for analyzing the authenticity of network data, which addresses the technical problem that randomly processing crawled network data often results in a large amount of unprocessed network data, thus affecting analysis efficiency.
[0004] In a first aspect, the present invention provides a method for analyzing the authenticity of network data, comprising:
[0005] Crawl at least one piece of network data, and behavioral data corresponding to the at least one piece of network data;
[0006] Based on the behavioral data, the at least one network data is first divided using a preset first-level division rule to obtain at least one network data set, wherein a network data set contains at least one network data at the same level.
[0007] Based on a preset second-level partitioning rule, each network data in a certain network data set is partitioned in the second way to obtain at least one network data subset corresponding to the certain network data set, and semantic tags corresponding to the at least one network data subset. The level corresponding to the certain network data set is the level with the highest priority among all levels.
[0008] Based on the aforementioned level, at least one target network data is selected from the at least one subset of network data, and the authenticity of the at least one target network data is analyzed to obtain a set of authenticity levels corresponding to the at least one subset of network data. The set of authenticity levels contains the authenticity levels corresponding to at least one target network data in the same subset of network data.
[0009] Determine the degree deviation of each degree of authenticity in a certain degree of authenticity set corresponding to a certain subset of network data, and update the certain degree of authenticity set based on each degree deviation using a preset set update strategy to obtain a certain target degree of authenticity set.
[0010] The authenticity of each piece of network data in a certain subset of network data is determined based on the authenticity of each target in the set of authenticity of a certain target, and the authenticity of each piece of network data in other subsets of network data is determined based on a certain semantic label of the certain subset of network data.
[0011] Secondly, the present invention provides a network data authenticity analysis system, comprising:
[0012] The crawling module is configured to crawl at least one piece of network data and behavioral data corresponding to the at least one piece of network data;
[0013] The first partitioning module is configured to perform a first partitioning on the at least one network data according to the behavioral data and a preset first-level partitioning rule to obtain at least one network data set, wherein a network data set contains at least one network data at the same level.
[0014] The second partitioning module is configured to perform a second partitioning on each network data in a certain network data set based on a preset second-level partitioning rule, to obtain at least one network data subset corresponding to the certain network data set, and semantic tags corresponding to the at least one network data subset, wherein the certain level corresponding to the certain network data set is the level with the highest priority among all levels.
[0015] The analysis module is configured to select at least one target network data in the at least one network data subset based on a certain level, and perform authenticity analysis on the at least one target network data to obtain a set of authenticity levels corresponding to the at least one network data subset, wherein a set of authenticity levels contains the authenticity levels corresponding to at least one target network data in the same network data subset.
[0016] The update module is configured to determine the degree deviation of each degree of authenticity in a certain degree of authenticity set corresponding to a certain subset of network data, and update the certain degree of authenticity set based on each degree deviation using a preset set update strategy to obtain a certain target degree of authenticity set.
[0017] The determination module is configured to determine the authenticity level of each network data in a certain network data subset based on the authenticity level of each target in the set of authenticity levels of a certain target, and to determine the authenticity level of each network data in other network data subsets based on a certain semantic label of the certain network data subset.
[0018] Thirdly, an electronic device is provided, comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the steps of the network data authenticity analysis method of any embodiment of the present invention.
[0019] Fourthly, the present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein when the program instructions are executed by a processor, the processor performs the steps of the network data authenticity analysis method of any embodiment of the present invention.
[0020] The network data authenticity analysis method and system of this application dynamically divides data levels by propagation path length, focuses on high-impact data, improves the targeting of analysis, and determines the target authenticity degree set by adaptive sampling of propagation level and long text priority strategy. Based on the authenticity degree of each target in the target authenticity degree set, the authenticity degree of each network data in the network data subset is determined. This can reduce the amount of data processing as much as possible while ensuring the accuracy of network data analysis. Attached Figure Description
[0021] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 A flowchart illustrating a method for analyzing the authenticity of network data provided in an embodiment of the present invention;
[0023] Figure 2 This is a structural block diagram of a network data authenticity analysis system provided in an embodiment of the present invention;
[0024] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0025] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0026] Please see Figure 1 The diagram shows a flowchart of a network data authenticity analysis method according to this application.
[0027] like Figure 1 As shown, the method for analyzing the authenticity of network data specifically includes the following steps:
[0028] Step S101: Crawl at least one piece of network data and behavioral data corresponding to the at least one piece of network data.
[0029] In this step, the behavioral data includes the path length of at least one propagation path corresponding to the network data.
[0030] Specifically, to crawl network data and its corresponding behavioral data, a distributed crawler cluster is first deployed to connect to the target data source API interface and web page. Dedicated crawling modules are configured for social media platforms such as Weibo and Twitter. During the initialization phase, the crawling scope is defined to include text data (news articles, user comments, post content) and its associated behavioral data (path length of the propagation path). The crawler system adopts a dual-channel acquisition architecture: the main data channel uses the Scrapy framework to build a targeted crawler, configuring XPath / CSS selectors to accurately extract the core elements (text content) of the network data, while the behavioral data channel uses the platform's open API (such as the Weibo repost_timeline interface) or reverse engineering to parse private interfaces to capture the propagation path associated with the network data in real time and record the user ID of each forwarding node.
[0031] Step S102: Based on the behavioral data, the at least one network data is first divided using a preset first-level division rule to obtain at least one network data set, wherein a network data set contains at least one network data at the same level.
[0032] In this step, at least one standard path length range is set, wherein one standard path length range corresponds to one propagation degree; the path length of each propagation path corresponding to a certain network data is obtained, and the standard path length range to which each path length belongs is found to obtain the propagation degree corresponding to each path length; the propagation degrees are added together to obtain the propagation degree corresponding to the certain network data, and at least one network data within the same propagation degree range is divided into the same network data set to obtain at least one network data set.
[0033] Specifically, several standard path length ranges are pre-defined, and a propagation level value is assigned to each range. For example: path length between 1 and 5: propagation level value is 1 (indicating primary propagation); path length between 6 and 20: propagation level value is 2 (indicating intermediate propagation); path length between 21 and 100: propagation level value is 3 (indicating explosive propagation).
[0034] Extracting the propagation path length of a single network data entry:
[0035] For each piece of network data, obtain all its propagation paths from the associated behavioral data. The length of each propagation path is defined as the number of forwarding nodes in that path (for example, if user A→user B→user C, the path length is 3). If a piece of network data has multiple propagation paths (for example, it is forwarded multiple times to form multiple propagation chains), then record the length of all propagation paths.
[0036] Convert path length to propagation extent:
[0037] For each propagation path, its standard path length range is determined based on its path length, and the corresponding propagation degree value is obtained. For example, a propagation path with a length of 8 belongs to the range of 6-20, and its corresponding propagation degree value is 2.
[0038] In this embodiment, the degree of propagation of network data represents the influence of the network data. The greater the degree of propagation, the more priority should be given to analyzing the authenticity of the network data, and the more accurate its authenticity should be. Therefore, different priority sets of network data are set by the first-level division rule, and different selection rules are executed for network data sets with different priorities (such as selecting at least one target network data in the at least one network data subset based on the first-level division rule in step S104). This facilitates the effective allocation of the amount of data to be processed, thereby reducing the amount of data to be processed while ensuring the accuracy of subsequent authenticity analysis.
[0039] Step S103: Based on the preset second-level partitioning rules, perform a second partitioning on each network data in a certain network data set to obtain at least one network data subset corresponding to the certain network data set, and semantic tags corresponding to the at least one network data subset, wherein the level corresponding to the certain network data set is the level with the highest priority among all levels.
[0040] In this step, semantic tags include semantic sub-tags and numerical sub-tags for network data. Specifically, semantic sub-tags represent the type of text content, such as "mobile phone" or "car." Numerical sub-tags represent the numerical values appearing in the text, such as "100," "above 500," or "500-1000."
[0041] It should be noted that when there are no numerical values in the network data, the numerical sub-label is set to "none".
[0042] Specifically, at least one first keyword of a first network data and at least one second keyword of a second network data in a certain network dataset are obtained to obtain a first keyword set and a second keyword set, wherein the first network data and the second network data are any two network data in the certain network dataset; each first keyword in the first keyword set is simultaneously input into a preset semantic recognition model, and the semantic recognition model outputs a first semantic label corresponding to the first keyword set; each second keyword in the second keyword set is simultaneously input into the preset semantic recognition model, and the semantic recognition model outputs a second semantic label corresponding to the second keyword set; it is determined whether the first semantic label and the second semantic label are the same; if they are the same, the first network data and the second network data are divided into the same network data subset; otherwise, the first network data and the second network data are divided into different network data subsets.
[0043] In one specific embodiment, the extracted keywords are input into a preset semantic recognition model, which outputs semantic sub-labels and numerical sub-labels of the network data.
[0044] A semantic recognition model can be a multi-label classification model that, after training, can classify text (to obtain semantic sub-labels) and recognize numerical values (to obtain numerical sub-labels).
[0045] For example:
[0046] Online data 1: "Apple releases a new iPhone priced at $699."
[0047] Semantic sub-tag: technology;
[0048] Numeric sub-tag: 699;
[0049] Online data 2: The Lakers won the NBA Finals 4-2.
[0050] Semantic sub-tag: sports;
[0051] Numerical sub-tag: 4:2;
[0052] Division rules:
[0053] Network data with the same semantic sub-labels and the same numerical sub-labels are grouped into the same subset.
[0054] For example:
[0055] Subset 1: All data with semantic tags (technology, 699)
[0056] Subset 2: All data with the semantic label (sports, 4:2)
[0057] Handling multi-tag scenarios:
[0058] If a piece of data has multiple semantic sub-labels or numerical sub-labels, we adopt the following strategy:
[0059] Select the semantic sub-label with the highest confidence and the numerical sub-label.
[0060] If there is no numeric sub-tag, the numeric sub-tag is recorded as "none".
[0061] Suppose that the high-propagation set contains the following three data points:
[0062] Data 1: "Apple releases new phone, priced at $699." (Keywords: Apple, phone, phone, price, $699) → Semantic tags (Apple phone, $699);
[0063] Data 2: "Apple releases new phone, priced at $899." (Keywords: Apple, phone, phone, price, $899) → Semantic tags (Apple phone, $899);
[0064] Data 4: "Tesla Model 3 is expected to be sold at a price increase in the first quarter due to increased costs." (Keywords: Tesla, Model 3, price increase) → Semantic tags (Tesla cars, none)
[0065] Data 4: "Tesla Model 3 is expected to see a significant price reduction." (Keywords: Tesla, Model 3, price reduction) → Semantic tags (Tesla cars, none);
[0066] Partition results:
[0067] Network data subset A (Apple phone, 699): contains data 1;
[0068] Network data subset B (Apple phone, 899): contains data 2;
[0069] Network data subset B (Tesla cars): contains data 3 and data 4.
[0070] Step S104: Based on the certain level, select at least one target network data in each of the at least one network data subset, and perform authenticity analysis on the at least one target network data to obtain a set of authenticity levels corresponding to the at least one network data subset. The set of authenticity levels contains the authenticity levels corresponding to at least one target network data in the same network data subset.
[0071] In this step, a preset selection quantity is determined based on a certain level, where different levels correspond to different selection quantities; the network data in a certain network data subset are sorted based on data length to obtain a network data subsequence, and a certain selection quantity of network data is selected from the network data subsequence in descending order to obtain at least one target network data corresponding to a certain network data subset, where a certain network data subset is any network data subset among at least one network data subset, and the data length is the number of characters in the text of the network data; the at least one target network data corresponding to a certain network data subset is input into a preset realism classification model, and the realism classification model outputs the realism degree corresponding to each target network data; each realism degree is divided into the same realism degree set to obtain a realism degree set corresponding to a certain network data subset.
[0072] In this embodiment, different selection quantities are set according to different levels. The higher the priority, the more selection quantities are set. The sampling quantity is adaptively configured according to the propagation level, so that computing resources are focused on high-impact data. Secondly, since long texts contain richer verifiable entities and logical chains, long text samples are selected in descending order of text character count. Finally, the target network data is input into the authenticity classification model to generate the authenticity level.
[0073] It should be noted that the realism classification model can be obtained through iterative training of the neural network, so it will not be elaborated on here.
[0074] Step S105: Determine the degree deviation of each degree of authenticity in a certain degree of authenticity set corresponding to a certain subset of network data, and update the certain degree of authenticity set based on each degree deviation using a preset set update strategy to obtain a certain target degree of authenticity set.
[0075] In this step, the absolute value of the difference between a certain degree of realism and other degrees of realism in a certain set of realism is obtained, and the absolute values of each difference are added together to obtain a certain degree deviation of the certain degree of realism; it is determined whether the certain degree deviation is greater than a preset deviation threshold; if it is not greater than the preset deviation threshold, the certain degree of realism corresponding to the certain degree deviation is not removed from the certain set of realism; if it is greater than the preset deviation threshold, the certain degree of realism corresponding to the certain degree deviation is removed from the certain set of realism to obtain an updated set of a certain target degree of realism.
[0076] Step S106: Determine the authenticity level of each network data in a certain network data subset based on the authenticity level of each target in the set of authenticity levels of a certain target, and determine the authenticity level of each network data in other network data subsets based on a certain semantic label of the certain network data subset.
[0077] In this step, the average value of a set of target authenticity levels is calculated to obtain an average authenticity level, which is then directly used as the authenticity level of each piece of network data in a subset of network data.
[0078] Furthermore, it is determined whether other semantic labels of other network data subsets are consistent with a certain semantic label, wherein any network data subset within the other network data subsets of the other network data subsets; if consistent, then a certain average authenticity level is directly used as the authenticity level of each network data in the other network data subsets; if inconsistent, then based on other levels corresponding to other network data subsets, at least one target other network data is selected from each of the other network data subsets, and authenticity analysis is performed on at least one target other network data to obtain a set of other authenticity levels corresponding to at least one other network data subset; the authenticity level of each network data in the other network data subsets is determined according to each other authenticity level in the set of other authenticity levels.
[0079] After obtaining the degree of authenticity of each network data in other network data subsets, the degree deviation of each degree of authenticity is calculated, and the other degree of authenticity sets are updated according to the degree deviation. The implementation principle is the same as the implementation steps of step S105, so it will not be described in detail here.
[0080] Specifically, other network data subsets are other network data subsets, and other semantic labels are semantic labels for other network data subsets.
[0081] In summary, the method of this application dynamically divides data levels by propagation path length, focuses on high-impact data, and improves the targeting of analysis. Furthermore, by using adaptive sampling of propagation levels and combining a long text priority strategy to determine the set of target authenticity, and then determining the authenticity of each network data in the network data subset based on the authenticity of each target in the target authenticity set, it can reduce the amount of data processing as much as possible while ensuring the accuracy of network data analysis.
[0082] Please see Figure 2 The diagram shows a structural block diagram of a network data authenticity analysis system according to this application.
[0083] like Figure 2 As shown, the network data authenticity analysis system 200 includes a crawling module 210, a first partitioning module 220, a second partitioning module 230, an analysis module 240, an update module 250, and a determination module 260.
[0084] The crawling module 210 is configured to crawl at least one network data and corresponding behavioral data. The first partitioning module 220 is configured to perform a first partitioning of the at least one network data according to the behavioral data using a preset first-level partitioning rule, resulting in at least one network data set, wherein each network data set contains at least one network data at the same level. The second partitioning module 230 is configured to perform a second partitioning of each network data in a network data set based on a preset second-level partitioning rule, resulting in at least one network data subset corresponding to the network data set and semantic tags corresponding to the at least one network data subset, wherein the level corresponding to the network data set is the level with the highest priority among all levels. The analysis module 240 is configured to select from the at least one network data subset based on the level. At least one target network data is collected, and authenticity analysis is performed on the at least one target network data to obtain a set of authenticity levels corresponding to the at least one network data subset. Each set of authenticity levels contains authenticity levels corresponding to at least one target network data within the same network data subset. An update module 250 is configured to determine the degree deviation of each authenticity level in the set of authenticity levels corresponding to the network data subset, and update the set of authenticity levels based on each degree deviation using a preset set update strategy to obtain a target authenticity level set. A determination module 260 is configured to determine the authenticity level of each network data in a network data subset based on the target authenticity levels in the target authenticity level set, and to determine the authenticity level of each network data in other network data subsets based on a semantic tag of the network data subset.
[0085] It should be understood that Figure 2 The modules and references described in the document Figure 1 The steps described in the text correspond to those in the method described above. Therefore, the operations, features, and corresponding technical effects described above also apply to the method described in the text. Figure 2 The various modules in the document will not be described in detail here.
[0086] In other embodiments, the present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein when the program instructions are executed by a processor, the processor performs the network data authenticity analysis method in any of the above method embodiments.
[0087] In one embodiment, the computer-readable storage medium of the present invention stores computer-executable instructions, which are configured as follows:
[0088] Crawl at least one piece of network data, and behavioral data corresponding to the at least one piece of network data;
[0089] Based on the behavioral data, the at least one network data is first divided using a preset first-level division rule to obtain at least one network data set, wherein a network data set contains at least one network data at the same level.
[0090] Based on a preset second-level partitioning rule, each network data in a certain network data set is partitioned in the second way to obtain at least one network data subset corresponding to the certain network data set, and semantic tags corresponding to the at least one network data subset. The level corresponding to the certain network data set is the level with the highest priority among all levels.
[0091] Based on the aforementioned level, at least one target network data is selected from the at least one subset of network data, and the authenticity of the at least one target network data is analyzed to obtain a set of authenticity levels corresponding to the at least one subset of network data. The set of authenticity levels contains the authenticity levels corresponding to at least one target network data in the same subset of network data.
[0092] Determine the degree deviation of each degree of authenticity in a certain degree of authenticity set corresponding to a certain subset of network data, and update the certain degree of authenticity set based on each degree deviation using a preset set update strategy to obtain a certain target degree of authenticity set.
[0093] The authenticity of each piece of network data in a certain subset of network data is determined based on the authenticity of each target in the set of authenticity of a certain target, and the authenticity of each piece of network data in other subsets of network data is determined based on a certain semantic label of the certain subset of network data.
[0094] Computer-readable storage media may include a stored program area and a stored data area, wherein the stored program area may store an operating system and an application program required for at least one function; the stored data area may store data created based on the use of the network data authenticity analysis system, etc. Furthermore, the computer-readable storage medium may include high-speed random access memory, and may also include memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, the computer-readable storage medium may optionally include memory remotely disposed relative to a processor, which can be connected to the network data authenticity analysis system via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0095] Figure 3 This is a schematic diagram of the structure of the electronic device provided in the embodiment of the present invention, such as... Figure 3 As shown, the device includes a processor 310 and a memory 320. The electronic device may also include an input device 330 and an output device 340. The processor 310, memory 320, input device 330, and output device 340 can be connected via a bus or other means. Figure 3 Taking a bus connection as an example, the memory 320 is the computer-readable storage medium described above. The processor 310 executes various server functions and data processing by running non-volatile software programs, instructions, and modules stored in the memory 320, thereby implementing the network data authenticity analysis method described in the above embodiment. The input device 330 can receive input digital or character information and generate key signal inputs related to user settings and function control of the network data authenticity analysis system. The output device 340 may include a display screen or other display device.
[0096] The aforementioned electronic device can execute the method provided in the embodiments of the present invention, and has the corresponding functional modules and beneficial effects for executing the method. Technical details not described in detail in this embodiment can be found in the method provided in the embodiments of the present invention.
[0097] In one implementation, the above-described electronic device is used in a network data authenticity analysis system for a client, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to:
[0098] Crawl at least one piece of network data, and behavioral data corresponding to the at least one piece of network data;
[0099] Based on the behavioral data, the at least one network data is first divided using a preset first-level division rule to obtain at least one network data set, wherein a network data set contains at least one network data at the same level.
[0100] Based on a preset second-level partitioning rule, each network data in a certain network data set is partitioned in the second way to obtain at least one network data subset corresponding to the certain network data set, and semantic tags corresponding to the at least one network data subset. The level corresponding to the certain network data set is the level with the highest priority among all levels.
[0101] Based on the aforementioned level, at least one target network data is selected from the at least one subset of network data, and the authenticity of the at least one target network data is analyzed to obtain a set of authenticity levels corresponding to the at least one subset of network data. The set of authenticity levels contains the authenticity levels corresponding to at least one target network data in the same subset of network data.
[0102] Determine the degree deviation of each degree of authenticity in a certain degree of authenticity set corresponding to a certain subset of network data, and update the certain degree of authenticity set based on each degree deviation using a preset set update strategy to obtain a certain target degree of authenticity set.
[0103] The authenticity of each piece of network data in a certain subset of network data is determined based on the authenticity of each target in the set of authenticity of a certain target, and the authenticity of each piece of network data in other subsets of network data is determined based on a certain semantic label of the certain subset of network data.
[0104] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of various embodiments or some parts of embodiments.
[0105] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for analyzing the authenticity of network data, characterized in that, include: Crawl at least one piece of network data, and behavioral data corresponding to the at least one piece of network data; Based on the behavioral data, the at least one network data is first divided using a preset first-level division rule to obtain at least one network data set, wherein a network data set contains at least one network data at the same level, and the behavioral data includes the path length of at least one propagation path corresponding to the network data. The step of performing a first division on the at least one network data set based on the behavioral data and using a preset first-level division rule to obtain at least one network data set includes: Set at least one standard path length range, where each standard path length range corresponds to a propagation degree; Obtain the path length of each propagation path corresponding to a certain network data, and find the standard path length range to which each path length belongs, to obtain the propagation degree corresponding to each path length. The various propagation levels are added together to obtain the propagation level corresponding to a certain network data, and at least one network data within the same propagation level range is divided into the same network data set to obtain at least one network data set; Based on a preset second-level partitioning rule, each network data in a certain network data set is partitioned in the second way to obtain at least one network data subset corresponding to the certain network data set, and semantic tags corresponding to the at least one network data subset. The level corresponding to the certain network data set is the level with the highest priority among all levels. Based on the aforementioned level, at least one target network data is selected from the at least one subset of network data, and the authenticity of the at least one target network data is analyzed to obtain a set of authenticity levels corresponding to the at least one subset of network data. The set of authenticity levels contains the authenticity levels corresponding to at least one target network data in the same subset of network data. Determine the degree deviation of each degree of authenticity in a set of authenticity levels corresponding to a certain subset of network data, and update the set of authenticity levels based on each degree deviation using a preset set update strategy to obtain a certain target set of authenticity levels. The authenticity of each piece of network data in a certain subset of network data is determined based on the authenticity of each target in the set of authenticity of a certain target, and the authenticity of each piece of network data in other subsets of network data is determined based on a certain semantic label of the certain subset of network data.
2. The method for analyzing the authenticity of network data according to claim 1, characterized in that, The semantic tags include semantic sub-tags and numerical sub-tags for network data. The second partitioning of each network data in a certain network dataset based on a preset second-level partitioning rule yields at least one network data subset corresponding to the certain network dataset, and the semantic tags corresponding to the at least one network data subset include: Obtain at least one first keyword of a first network data and at least one second keyword of a second network data in a certain network data set to obtain a first keyword set and a second keyword set, wherein the first network data and the second network data are any two network data in the certain network data set; Each first keyword in the first keyword set is simultaneously input into a preset semantic recognition model, and the semantic recognition model outputs a first semantic label corresponding to the first keyword set; Each second keyword in the second keyword set is simultaneously input into a preset semantic recognition model, and the semantic recognition model outputs a second semantic label corresponding to the second keyword set; Determine whether the first semantic tag and the second semantic tag are the same; If they are the same, the first network data and the second network data are assigned to the same network data subset; otherwise, the first network data and the second network data are assigned to different network data subsets.
3. The method for analyzing the authenticity of network data according to claim 1, characterized in that, Based on the aforementioned level, at least one target network data is selected from each of the at least one subset of network data, and authenticity analysis is performed on the at least one target network data to obtain a set of authenticity levels corresponding to the at least one subset of network data, including: A preset selection quantity is determined based on a certain level, wherein different levels correspond to different selection quantities; The network data in a certain subset of network data is sorted according to the data length to obtain a network data subsequence. Then, a certain number of network data are selected from the network data subsequence in descending order to obtain at least one target network data corresponding to the certain subset of network data. The certain subset of network data is any one of the at least one subset of network data. The data length is the number of characters in the text of the network data. At least one target network data corresponding to a certain subset of network data is input into a preset authenticity classification model, and the authenticity classification model outputs the degree of authenticity corresponding to each target network data. Each degree of authenticity is assigned to the same set of authenticity levels to obtain a set of authenticity levels corresponding to a certain subset of network data.
4. The method for analyzing the authenticity of network data according to claim 1, characterized in that, The step of determining the degree deviation of each degree of authenticity in a certain degree of authenticity set corresponding to a certain subset of network data, and updating the certain degree of authenticity set based on each degree deviation using a preset set update strategy to obtain a certain target degree of authenticity set includes: Obtain the absolute value of the difference between a certain degree of authenticity and other degrees of authenticity in a certain set of degrees of authenticity, and add the absolute values of each difference to obtain a certain degree deviation of the certain degree of authenticity; Determine whether the deviation of a certain degree is greater than a preset deviation threshold; If the deviation is not greater than a preset deviation threshold, then the degree of authenticity corresponding to the deviation will not be removed from the set of degrees of authenticity. If the deviation exceeds a preset threshold, the degree of authenticity corresponding to the deviation will be removed from the set of degrees of authenticity, resulting in an updated set of degrees of authenticity for the target.
5. The method for analyzing the authenticity of network data according to claim 1, characterized in that, The steps of determining the authenticity level of each piece of network data in a certain subset of network data based on the authenticity levels of each target in the set of authenticity levels of a certain target, and determining the authenticity level of each piece of network data in other subsets of network data based on a semantic label of the certain subset of network data, include: The average value of the set of target authenticity is calculated to obtain an average authenticity level, and the average authenticity level is directly used as the authenticity level of each network data in a certain network data subset. Determine whether other semantic tags of other subsets of network data are consistent with the given semantic tag, wherein the other subsets of network data are any subsets of network data in other subsets of network data; If they are consistent, then the average degree of authenticity shall be directly used as the degree of authenticity of each network data in the other subset of network data; If there is a discrepancy, then based on other levels corresponding to the other network data subsets, at least one target other network data is selected from the other network data subsets respectively, and the authenticity analysis is performed on the at least one target other network data to obtain other authenticity degree sets corresponding to the at least one other network data subsets; The degree of authenticity of each network data in the other subset of network data is determined based on the other degrees of authenticity in the other set of degrees of authenticity.
6. A network data authenticity analysis system, characterized in that, include: The crawling module is configured to crawl at least one piece of network data and behavioral data corresponding to the at least one piece of network data; The first partitioning module is configured to perform a first partitioning of the at least one network data according to the behavioral data and a preset first-level partitioning rule to obtain at least one network data set, wherein a network data set contains at least one network data at the same level, and the behavioral data includes the path length of at least one propagation path corresponding to the network data. The step of performing a first division on the at least one network data set based on the behavioral data and using a preset first-level division rule to obtain at least one network data set includes: Set at least one standard path length range, where each standard path length range corresponds to a propagation degree; Obtain the path length of each propagation path corresponding to a certain network data, and find the standard path length range to which each path length belongs, to obtain the propagation degree corresponding to each path length. The various propagation levels are added together to obtain the propagation level corresponding to a certain network data, and at least one network data within the same propagation level range is divided into the same network data set to obtain at least one network data set; The second partitioning module is configured to perform a second partitioning on each network data in a certain network data set based on a preset second-level partitioning rule, to obtain at least one network data subset corresponding to the certain network data set, and semantic tags corresponding to the at least one network data subset, wherein the certain level corresponding to the certain network data set is the level with the highest priority among all levels. The analysis module is configured to select at least one target network data in the at least one network data subset based on a certain level, and perform authenticity analysis on the at least one target network data to obtain a set of authenticity levels corresponding to the at least one network data subset, wherein a set of authenticity levels contains the authenticity levels corresponding to at least one target network data in the same network data subset. The update module is configured to determine the degree deviation of each degree of authenticity in a set of authenticity levels corresponding to a certain subset of network data, and update the set of authenticity levels based on each degree deviation using a preset set update strategy to obtain a certain target set of authenticity levels. The determination module is configured to determine the authenticity level of each network data in a certain network data subset based on the authenticity level of each target in the set of authenticity levels of a certain target, and to determine the authenticity level of each network data in other network data subsets based on a certain semantic label of the certain network data subset.
7. An electronic device, characterized in that, include: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method described in any one of claims 1 to 5.
Citation Information
Patent Citations
Computer network data acquisition, analysis and management method and device and storage medium
CN117194754A
Network node, core network and communication method
EP4443322A1