Information tracking method and apparatus

CN118861395BActive Publication Date: 2026-09-29CHINA MOBILE GROUP DESIGN INST +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410909508.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-08
Publication Date
2026-09-29
Estimated Expiration
2044-07-08

AI Technical Summary

Technical Problem

[0004]本申请实施例提供一种信息追踪方法及装置,用以解决当前的信息追踪方法只能筛选出与目标信息相似性较高的信息,而无法筛选出与目标信息关联性较高的信息,导致信息筛选的准确性较低,从而导致信息追踪的准确性较低的技术问题

Benefits of technology

[0016]本申请提供的信息追踪方法及装置,获取与当前目标信息相关的多条当前爬虫信息,将任一条当前爬虫信息按照不同的数据颗粒度进行切片,得到该条当前爬虫信息对应的多条当前切片信息,将当前目标信息和该条当前爬虫信息对应的多条当前切片信息进行向量化,得到当前目标信息向量和该条当前爬虫信息对应的多条当前切片信息向量,计算该条当前爬虫信息对应的多条当前切片信息向量与当前目标信息向量之间的欧氏距离平均值,若欧氏距离平均值小于平均值阈值,则对该条当前爬虫信息进行追踪。本申请中,先获取与当前目标信息相关的多条当前爬虫信息,这些爬虫信息与当前目标信息的关联性有高有低,对每条当前爬虫信息按照不同的数据颗粒度进行切片,可以得到针对每条当前爬虫信息的具备不同数据颗粒度的多条当前切片信息,由于数据颗粒度不同会导致信息语义的不同,不同颗粒度对应的当前切片信息与当前目标信息的语义关联性也会不同,则将每条当前爬虫信息划分为多条与当前目标信息关联性不同的当前切片信息,可以更加精准地捕获到与当前目标信息关联性较高的当前爬虫信息,再将这些当前切片信息和当前目标信息向量化后计算欧式距离平均值,通过与阈值的比较筛选出多条当前爬虫信息中与当前目标信息在不同关联性下综合相似性较高的当前爬虫信息,从而能够综合关联性和相似性的影响,筛选出与当前目标信息同时具备高关联性和高相似性的当前爬虫信息进行追踪,提高信息筛选的准确性,进而提高信息追踪的准确性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118861395B_ABST
    Figure CN118861395B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of information tracking, and provides an information tracking method and device. The method comprises the following steps: acquiring a plurality of current crawler information related to current target information; slicing any current crawler information according to different data granularities to obtain a plurality of current slice information; vectorizing the current target information and the plurality of current slice information to obtain a current target information vector and a plurality of current slice information vectors; calculating the average value of the Euclidean distance between the plurality of current slice information vectors and the current target information vector; and if the average value of the Euclidean distance is less than an average value threshold, tracking the any current crawler information. The information tracking method and device provided by the application can comprehensively consider the influence of correlation and similarity, filter out the current crawler information with high correlation and high similarity at the same time to perform tracking, improve the accuracy of information filtering, and further improve the accuracy of information tracking.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of information tracking technology, specifically to an information tracking method and apparatus. Background Technology

[0002] With the rapid development of the internet in recent years, it has become an important channel for people to obtain information. Continuous tracking of target information via the internet has also become an urgent need. Continuous tracking of target information involves continuously tracking related information, including information that is highly relevant to the target information and information that is highly similar to it.

[0003] Current information tracking methods calculate the similarity between target information and internet information using conventional similarity algorithms, and then filter out internet information with high similarity for tracking. This method can only filter out information with high similarity to the target information, but cannot filter out information with high relevance to the target information, resulting in low accuracy of information filtering and thus low accuracy of information tracking. Summary of the Invention

[0004] This application provides an information tracking method and apparatus to solve the technical problem that current information tracking methods can only filter out information that is highly similar to the target information, but cannot filter out information that is highly related to the target information, resulting in low accuracy of information filtering and thus low accuracy of information tracking.

[0005] In a first aspect, embodiments of this application provide an information tracking method, including: Retrieve multiple pieces of current crawler information related to the current target information; Slice any current crawler information according to different data granularities to obtain multiple current slice information corresponding to the current crawler information; Vectorize the current target information and the multiple current slice information corresponding to any current crawler information to obtain the current target information vector and the multiple current slice information vector corresponding to any current crawler information; Calculate the average Euclidean distance between the multiple current slice information vectors corresponding to any current crawler information and the current target information vector; If the average Euclidean distance is less than the average threshold, then any current crawler information will be tracked.

[0006] In one embodiment, the step of slicing any current crawler information according to different data granularities to obtain multiple current slice information corresponding to the any current crawler information includes: The characters in any given current crawler information are sliced ​​at equal intervals according to different spacings to obtain multiple current slice information corresponding to any given current crawler information.

[0007] In one embodiment, the step of vectorizing the current target information and the multiple current slice information corresponding to any current crawler information to obtain the current target information vector and the multiple current slice information vector corresponding to any current crawler information includes: Input the current target information and multiple current slice information corresponding to any current crawler information into the first embedding layer of the information type recognition model to obtain the character vector of each character in the current target information and the character vector of each character in the multiple current slice information corresponding to any current crawler information. Input the character vectors of each character in the current target information and the character vectors of each character in multiple current slice information corresponding to any current crawler information into the convolutional layer of the information type recognition model to obtain the current target information vector and the multiple current slice information vectors corresponding to any current crawler information. The information type recognition model is trained on the basis of the BERT-CNN model using historical information and its type labels. The historical information includes historical target information and historical crawler information.

[0008] In one embodiment, the information type identification model is determined based on the following method: Slice any historical crawler information related to any historical target information into slices according to different data granularities to obtain multiple historical slice information corresponding to the historical crawler information. Multiple historical target information and multiple historical crawler information related to the multiple historical target information are input into the embedding layer of the BERT-CNN model to obtain the character vectors of each character in the multiple historical target information and the character vectors of each character in the multiple historical crawler information. The character vectors of each character in the multiple historical target information and the character vectors of each character in the multiple historical slice information corresponding to the multiple historical crawler information are input into the convolutional layer of the BERT-CNN model to obtain the vectors of multiple historical target information and the vectors of multiple historical slice information corresponding to the multiple historical crawler information. The multiple historical target information vectors, the multiple historical crawler information vectors corresponding to the multiple historical slice information, and the type labels corresponding to each historical information are input into the nonlinear classifier of the BERT-CNN model to obtain the type prediction probability of the multiple historical target information and the type prediction probability of the multiple historical crawler information corresponding to the multiple historical slice information. If the cross-entropy loss value between the predicted probability of each information type and the true probability of the type is greater than or equal to the loss value threshold, then after adjusting the parameter weights of the BERT-CNN model, the process returns to the step of inputting multiple historical target information and multiple historical crawler information related to the multiple historical target information into the embedding layer of the BERT-CNN model until the cross-entropy loss value is less than the loss value threshold, and the BERT-CNN model at this time is determined as the information type recognition model.

[0009] In one embodiment, obtaining multiple pieces of current crawler information related to the current target information includes: Obtain multiple high-frequency information items related to the current target information in the network; The multiple high-popularity information entries are input into the second embedding layer of the information type recognition model to obtain the character vector of each character in the multiple high-popularity information entries; The character vectors of each character in the multiple high-popularity information are input into the convolutional layer of the information type recognition model to obtain the multiple high-popularity information vectors corresponding to the multiple high-popularity information. The multiple high-popularity information vectors are input into the nonlinear classifier of the information type recognition model to obtain the type prediction probability of the multiple high-popularity information vectors. Based on the predicted probabilities of the types of the multiple high-profile information items, the source of the crawler information is determined; Obtain multiple current crawler information entries related to the current target information from the crawler information source.

[0010] In one embodiment, determining the source of crawler information based on the probability prediction of the types of the multiple high-popularity information items includes: Sort the multiple type prediction probabilities of any high-popularity information in descending order, and obtain the first target information source of the type to which the multiple type prediction probabilities at the top of the sort are located. If there is a first matching information source in the first target information source that overlaps with the preset information source, then the first matching information source is determined as the crawler information source. If there is no first matching information source that overlaps with the preset information source in the first target information source, then after waiting for a preset time, return to the step of obtaining multiple high-popularity information related to the current target information in the network, until the second target information source is obtained; If there is a second matching information source in the second target information source that overlaps with the preset information source, then the second matching information source is determined as the crawler information source; If there is no second matching information source in the second target information source that overlaps with the preset information source, then the second target information source is matched with the first target information source; If there is a third matching information source in the second target information source that overlaps with the first target information source, then the third matching information source is determined as the crawler information source.

[0011] In one embodiment, tracking any of the current crawler information includes: Set a listening interface for the information source of any of the current crawler information; Periodically retrieve multiple monitoring messages related to the current target information from the monitoring interface; The multiple monitoring messages are treated as multiple current crawler messages. The process of slicing any current crawler message according to different data granularities is returned to obtain multiple current slice messages corresponding to any current crawler message, until the Euclidean distance between the multiple slice message vectors corresponding to any monitoring message and the current target message vector is obtained. Based on the slice information vector corresponding to the minimum Euclidean distance, a summary information and a jump link are generated, and the summary information and the jump link are pushed to the user.

[0012] Secondly, embodiments of this application provide an information tracking device, comprising: The information acquisition module is used to: acquire multiple pieces of current crawler information related to the current target information; The slicing module is used to slice any current crawler information according to different data granularities to obtain multiple current slice information corresponding to the current crawler information. The vectorization module is used to: vectorize the current target information and multiple current slice information corresponding to any current crawler information to obtain the current target information vector and the multiple current slice information vector corresponding to any current crawler information; The distance calculation module is used to: calculate the average Euclidean distance between multiple current slice information vectors corresponding to any current crawler information and the current target information vector; The information tracking module is used to: track any current crawler information if the average Euclidean distance is less than the average threshold.

[0013] Thirdly, embodiments of this application provide an electronic device, including a processor and a memory storing a computer program, wherein the processor executes the program to implement the steps of the information tracking method described in the first aspect.

[0014] Fourthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the steps of the information tracking method described in the first aspect.

[0015] Fifthly, embodiments of this application provide a non-transitory computer-readable storage medium, including a computer program, which, when executed by a processor, implements the steps of the information tracking method described in the first aspect.

[0016] The information tracking method and apparatus provided in this application acquire multiple pieces of current crawler information related to the current target information, slice any one piece of current crawler information according to different data granularities to obtain multiple pieces of current slice information corresponding to the current crawler information, vectorize the current target information and the multiple pieces of current slice information corresponding to the current crawler information to obtain the current target information vector and the multiple pieces of current slice information vector corresponding to the current crawler information, calculate the average Euclidean distance between the multiple pieces of current slice information vector corresponding to the current crawler information and the current target information vector, and if the average Euclidean distance is less than the average threshold, then the current crawler information is tracked. In this application, multiple current crawler information related to the current target information are first obtained. The correlation between these crawler information and the current target information varies. Each current crawler information is sliced ​​according to different data granularities, resulting in multiple current slice information with different data granularities for each current crawler information. Since different data granularities lead to different information semantics, the semantic correlation between the current slice information corresponding to different granularities and the current target information will also be different. Therefore, dividing each current crawler information into multiple current slice information with different correlations with the current target information can more accurately capture current crawler information with high correlations with the current target information. Then, after vectorizing these current slice information and the current target information, the average Euclidean distance is calculated. By comparing with a threshold, current crawler information with high comprehensive similarity to the current target information under different correlations is selected. This can comprehensively consider the influence of correlation and similarity, and select current crawler information with both high correlation and high similarity to the current target information for tracking, thereby improving the accuracy of information filtering and thus improving the accuracy of information tracking. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is one of the flowcharts illustrating the information tracking method provided in the embodiments of this application; Figure 2 This is a second schematic flowchart of the information tracking method provided in the embodiments of this application; Figure 3 This is the third flowchart illustrating the information tracking method provided in the embodiments of this application; Figure 4 This is the fourth flowchart illustrating the information tracking method provided in the embodiments of this application; Figure 5 This is the fifth flowchart illustrating the information tracking method provided in the embodiments of this application; Figure 6 This is a schematic diagram of the structure of the information tracking device provided in the embodiments of this application; Figure 7 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0020] Figure 1 This is one of the flowcharts illustrating the information tracking method provided in this application. (Refer to...) Figure 1 This application provides an information tracking method, which may include: 101. Obtain multiple pieces of current crawler information related to the current target information; 102. Slice any current crawler information according to different data granularities to obtain multiple current slice information corresponding to the current crawler information; 103. Vectorize the current target information and the multiple current slice information corresponding to the current crawler information to obtain the current target information vector and the multiple current slice information vector corresponding to the current crawler information; 104. Calculate the average Euclidean distance between the current slice information vectors corresponding to the current crawler information and the current target information vector; 105. If the average Euclidean distance is less than the average threshold, then the current crawler information will be tracked.

[0021] In step 102, the characters in the current crawler information are sliced ​​at equal intervals according to different spacings to obtain multiple current slice information corresponding to the current crawler information.

[0022] Assume the current crawler information is It is sliced ​​according to different data granularities to obtain , , and Then we have: ; ; ; ; in, Indicates current crawler information The first in One character, The current slice information is obtained by slicing the slices at equal intervals with a spacing of 1 character. The current slice information is obtained by slicing the image at equal intervals of two characters. The current slice information is obtained by slicing the slices at equal intervals of 3 characters. The current slice information is obtained by slicing the data at equal intervals of 4 characters. Therefore, the current crawler information is... There are 4 current slice information entries. Since the granularity of 1, 2, 3, and 4 is sufficient to cover most of the granular features in the current crawler information over a large area without causing excessive computational growth, this slicing method is adopted in this embodiment.

[0023] In addition, compared with traditional word segmentation methods, the slicing method in this embodiment segments adjacent characters by gradually increasing the character spacing, so that the slice clusters in the slice information can obtain higher semantic accuracy by continuously increasing the number of adjacent characters, and avoids the large computational overhead brought by word segmentation algorithms.

[0024] It should be noted that the slicing method of the current crawler information can be set according to actual needs, and is not limited here. For example, based on the above slicing, it can be sliced ​​according to other character spacing to obtain the corresponding crawler information. More information about the current slice.

[0025] Furthermore, the current crawler information can be pre-processed. Noise reduction processing is performed, followed by segmentation to further improve the current crawler information. The accuracy of the noise reduction process is ensured. The noise reduction method can be selected based on actual needs and is not limited here. In this embodiment, regular expressions and stop word lists can be used to filter the current crawler information to eliminate noise.

[0026] The information tracking method provided in this embodiment obtains multiple current crawler information related to the current target information, slices any current crawler information according to different data granularities to obtain multiple current slice information corresponding to the current crawler information, vectorizes the current target information and the multiple current slice information corresponding to the current crawler information to obtain the current target information vector and the multiple current slice information vector corresponding to the current crawler information, calculates the average Euclidean distance between the multiple current slice information vector corresponding to the current crawler information and the current target information vector, and if the average Euclidean distance is less than the average threshold, then the current crawler information is tracked. In this embodiment, multiple current crawler information related to the current target information are first acquired. The correlation between these crawler information and the current target information varies. Each current crawler information is sliced ​​according to different data granularities, resulting in multiple current slice information with different data granularities for each current crawler information. Since different data granularities lead to different information semantics, the semantic correlation between the current slice information corresponding to different granularities and the current target information will also be different. Therefore, dividing each current crawler information into multiple current slice information with different correlations with the current target information can more accurately capture current crawler information with high correlations with the current target information. Then, after vectorizing these current slice information and the current target information, the average Euclidean distance is calculated. By comparing with a threshold, current crawler information with high comprehensive similarity to the current target information under different correlations is selected. This can comprehensively consider the influence of correlation and similarity, and select current crawler information with both high correlation and high similarity to the current target information for tracking, thereby improving the accuracy of information filtering and thus improving the accuracy of information tracking.

[0027] In one embodiment, vectorizing the current target information and multiple current slice information corresponding to any current crawler information to obtain a current target information vector and multiple current slice information vectors corresponding to that current crawler information may include: Input the current target information and the multiple current slice information corresponding to the current crawler information into the first embedding layer of the information type recognition model to obtain the character vector of each character in the current target information and the character vector of each character in the multiple current slice information corresponding to the current crawler information. Input the character vectors of each character in the current target information and the character vectors of each character in the multiple current slice information corresponding to the current crawler information into the convolutional layer of the information type recognition model to obtain the current target information vector and the multiple current slice information vectors corresponding to the current crawler information. The information type recognition model is trained using historical information and its type labels, based on the BERT-CNN (Bidirectional Encoder Representations from Transformers-Convolutional Neural Networks) model. The historical information includes historical target information and historical crawler information.

[0028] For the current slice information For example, after the first embedding layer of its input information type recognition model, it can obtain the character vectors of each character: ; in, express character vectors, Based on the current slice information The generated vocabulary, This means that in Get from A character vector.

[0029] After the above character vector sequence is input into the convolutional layer of the information type recognition model, the vector corresponding to each character is first obtained. Feature sequences : ; in, This represents the activation function. Indicates weight transpose, Indicates the first bias. This indicates the number of convolution kernels in the convolutional layer.

[0030] By concatenating each feature sequence, the information of the current slice can be obtained. The corresponding current slice information vector.

[0031] For the current target information vector and the current slice information , and The acquisition of the corresponding current slice information vector is similar, and will not be elaborated here.

[0032] This embodiment inputs multiple current slices of information corresponding to the current target information and the current crawler information into the information type recognition model trained by the BERT-CNN model. This model can vectorize each piece of information while fully preserving its features, making subsequent calculations more accurate and efficient.

[0033] Figure 2 This is a second schematic flowchart of the information tracking method provided in this application embodiment. (Refer to...) Figure 2 In one embodiment, the information type identification model can be determined based on the following method: 201. Slice any historical crawler information related to any historical target information into slices according to different data granularities to obtain multiple historical slice information corresponding to that historical crawler information. 202. Input multiple historical target information and multiple historical crawler information related to multiple historical target information into the embedding layer of the BERT-CNN model to obtain the character vectors of each character in the multiple historical target information and the character vectors of each character in the multiple historical crawler information corresponding to the multiple historical slice information. 203. Input the character vectors of each character in multiple historical target information and the character vectors of each character in multiple historical slice information corresponding to multiple historical crawler information into the convolutional layer of the BERT-CNN model to obtain the vectors of multiple historical target information and the vectors of multiple historical slice information corresponding to multiple historical crawler information. 204. Input multiple historical target information vectors, multiple historical crawler information vectors corresponding to multiple historical slice information, and type labels corresponding to each historical information into the nonlinear classifier of the BERT-CNN model to obtain the type prediction probability of multiple historical target information and the type prediction probability of multiple historical crawler information corresponding to multiple historical slice information. 205. If the cross-entropy loss value between the predicted probability of the type and the true probability of the type of each piece of information is greater than or equal to the loss value threshold, then adjust the parameter weights of the BERT-CNN model and return to step 202. 206. If the cross-entropy loss value is less than the loss value threshold, the BERT-CNN model at this time is determined as an information type recognition model.

[0034] In step 204, any information can be calculated according to the following formula. Type prediction probability : ; in, Information The vector, Indicates weight transpose, This indicates the second bias.

[0035] This information There are multiple types of predicted probabilities, namely It is a set of prediction probabilities of multiple types. Assuming that the preset information types are A, B, and C, then each piece of information has a prediction probability corresponding to A. The predicted probability corresponding to B and the predicted probability corresponding to C , Then it is , and A set of.

[0036] In step 205, the cross-entropy loss value between the predicted type probability and the true type probability of each piece of information can be calculated according to the following formula. : ; in, This indicates the number of information entries in the embedding layer of the BERT-CNN model during this training. Indicates the number of information types. Indicates the first The message is the first The true probability of each information type Indicates the first The message is the first Predicted probability of each information type Indicates the first The message is the first The predicted probability of each information type.

[0037] In this embodiment, since BERT can learn deep semantic representations and CNN is good at capturing local features, combining the two and training the BERT-CNN model with a large amount of historical target information and historical crawler information can make full use of BERT's semantic understanding ability and CNN's local feature extraction ability, thereby improving the overall classification performance of the model for target information and crawler information. At the same time, taking the convergence of the cross-entropy loss between the predicted probability of information type and the true probability of type as the training objective can minimize the error of the model's output classification results.

[0038] Figure 3 This is the third flowchart illustrating the information tracking method provided in this application's embodiments. (Refer to...) Figure 3 In one embodiment, obtaining multiple pieces of current crawler information related to the current target information may include: 301. Obtain multiple high-frequency information items related to the current target information in the network; 302. Input multiple high-popularity information messages into the second embedding layer of the information type recognition model to obtain the character vectors of each character in the multiple high-popularity information messages; 303. Input the character vectors of each character in multiple high-profile information into the convolutional layer of the information type recognition model to obtain multiple high-profile information vectors corresponding to multiple high-profile information; 304. Input multiple high-popularity information vectors into the nonlinear classifier of the information type recognition model to obtain the type prediction probability of multiple high-popularity information. 305. Based on the probability prediction of multiple high-traffic information types, determine the information source for web crawlers; 306. Obtain multiple pieces of current crawler information related to the current target information from the crawler information source.

[0039] In step 301, the information related to the current target information in the network can be sorted according to its popularity, and the top 10 pieces of information in the sorted list can be selected as high-popularity information.

[0040] In step 302, the second embedding layer and the aforementioned first embedding layer are two independent parts of the embedding layer of the information type recognition model.

[0041] When using web crawling technology to crawl information on the Internet, in order to capture as much information as possible related to the target information, it is necessary to crawl a large range of information, which requires a lot of expensive computing resources.

[0042] This embodiment first obtains a small amount of high-popularity information related to the current target information, and determines the crawler information source based on its type probability prediction. This can filter out a small range of crawler information sources from a large range of network information sources, narrowing the scope of crawled information and thus reducing the consumption of computing resources. At the same time, using high-popularity information as a guide can also avoid information omissions caused by narrowing the scope of crawled information.

[0043] Figure 4 This is the fourth flowchart illustrating the information tracking method provided in this application. (Refer to...) Figure 4 In one embodiment, determining the source of crawler information based on the predicted probability of the types of the multiple high-popularity information items may include: 401. Sort the multiple type prediction probabilities of any high-heat information from largest to smallest, and obtain the first target information source of the type to which the multiple type prediction probabilities at the top of the sort are located. 402. If there is a first matching information source in the first target information source that overlaps with the preset information source, then the first matching information source is determined as the crawler information source; 403. If there is no first matching information source that overlaps with the preset information source in the first target information source, then after waiting for a preset time, return to the step of obtaining multiple high-heat information related to the current target information in the network until the second target information source is obtained. 404. If there is a second matching information source in the second target information source that overlaps with the preset information source, then the second matching information source is determined as the crawler information source; 405. If there is no second matching information source in the second target information source that overlaps with the preset information source, then the second target information source is matched with the first target information source; 406. If there is a third matching information source in the second target information source that overlaps with the first target information source, then the third matching information source is determined as the crawler information source.

[0044] In step 401, the information sources corresponding to the top 5 types of predicted probabilities among the multiple types of highly popular information can be used as the first target information source.

[0045] In step 402, information from various mainstream media on the Internet can be divided into multiple information types, and representative websites from each information type can be selected as preset information sources.

[0046] In step 403, the preset duration can be set according to the actual situation, and there is no limitation here. In this embodiment, the preset duration can be set to 10 minutes.

[0047] In step 406, if there is no third matching information source in the second target information source that overlaps with the first target information source, it is considered that no unified statement has been formed regarding the current target information, and the credibility of the information currently circulating on the Internet is low, so data crawling will not be performed for the time being.

[0048] This embodiment selects representative information sources based on mainstream media information types, matches these representative information sources with representative information sources of high-profile information, and uses the successfully matched information sources as crawler information sources. This ensures that the crawler information sources possess both mainstream appeal and high popularity, making the information crawled from these sources more likely to be highly relevant to the current target information. Furthermore, if the representative information sources of mainstream media information and high-profile information do not match, the two high-profile information sources are matched, and the successfully matched information sources are used as crawler information sources. This ensures that the crawler information sources maintain high popularity over time, and the information crawled from these sources is also highly relevant to the current target information.

[0049] Figure 5 This is the fifth flowchart illustrating the information tracking method provided in this application. (Refer to...) Figure 5 In one embodiment, after tracking any current crawler information, the process may include: 501. Set a listening interface for any current crawler information source; 502. Periodically retrieve multiple monitoring messages related to the current target information from the monitoring interface; 503. Treat multiple monitoring messages as multiple current crawler messages, return the steps of slicing any current crawler message according to different data granularities to obtain multiple current slice messages corresponding to the current crawler message, until the Euclidean distance between the multiple slice message vectors corresponding to the monitoring message and the current target message vector is obtained; 504. Generate summary information and jump links based on the slice information vector corresponding to the minimum Euclidean distance, and push the summary information and jump links to the user.

[0050] Current information tracking methods require manual screening of the acquired internet information, which is time-consuming, labor-intensive, and can lead to a loss of timeliness in information tracking.

[0051] The method in this embodiment does not require manual screening. By periodically monitoring the information sources of the selected crawler information, the latest monitoring information in the information sources can be obtained regularly. For each latest monitoring information, the slice information with high relevance and similarity to the current target information is filtered and fed back to the user for timely review and tracking. This enables continuous tracking of the current target information, improves information tracking efficiency, and ensures the timeliness of information tracking.

[0052] The information tracking device provided in the embodiments of this application is described below. The information tracking device described below and the information tracking method described above can be referred to in correspondence.

[0053] Figure 6 This is a schematic diagram of the information tracking device provided in an embodiment of this application. (Refer to...) Figure 6 This application provides an information tracking device, which may include: Information acquisition module 601 is used to: acquire multiple pieces of current crawler information related to the current target information; The slicing module 602 is used to slice any current crawler information according to different data granularities to obtain multiple current slice information corresponding to the current crawler information. The vectorization module 603 is used to: vectorize the current target information and multiple current slice information corresponding to any current crawler information to obtain the current target information vector and the multiple current slice information vector corresponding to any current crawler information; The distance calculation module 604 is used to: calculate the average Euclidean distance between multiple current slice information vectors corresponding to any current crawler information and the current target information vector; The information tracking module 605 is used to: track any current crawler information if the average Euclidean distance is less than the average threshold.

[0054] The information tracking device provided in this embodiment acquires multiple current crawler information related to the current target information, slices any current crawler information according to different data granularities to obtain multiple current slice information corresponding to the current crawler information, vectorizes the current target information and the multiple current slice information corresponding to the current crawler information to obtain the current target information vector and the multiple current slice information vector corresponding to the current crawler information, calculates the average Euclidean distance between the multiple current slice information vector corresponding to the current crawler information and the current target information vector, and if the average Euclidean distance is less than the average threshold, then the current crawler information is tracked. In this embodiment, multiple current crawler information related to the current target information are first acquired. The correlation between these crawler information and the current target information varies. Each current crawler information is sliced ​​according to different data granularities, resulting in multiple current slice information with different data granularities for each current crawler information. Since different data granularities lead to different information semantics, the semantic correlation between the current slice information corresponding to different granularities and the current target information will also be different. Therefore, dividing each current crawler information into multiple current slice information with different correlations with the current target information can more accurately capture current crawler information with high correlations with the current target information. Then, after vectorizing these current slice information and the current target information, the average Euclidean distance is calculated. By comparing with a threshold, current crawler information with high comprehensive similarity to the current target information under different correlations is selected. This can comprehensively consider the influence of correlation and similarity, and select current crawler information with both high correlation and high similarity to the current target information for tracking, thereby improving the accuracy of information filtering and thus improving the accuracy of information tracking.

[0055] In one embodiment, the slicing module 602 is specifically used for: The characters in any given current crawler information are sliced ​​at equal intervals according to different spacings to obtain multiple current slice information corresponding to any given current crawler information.

[0056] In one embodiment, the vectorization module 603 is specifically used for: Input the current target information and multiple current slice information corresponding to any current crawler information into the first embedding layer of the information type recognition model to obtain the character vector of each character in the current target information and the character vector of each character in the multiple current slice information corresponding to any current crawler information. Input the character vectors of each character in the current target information and the character vectors of each character in multiple current slice information corresponding to any current crawler information into the convolutional layer of the information type recognition model to obtain the current target information vector and the multiple current slice information vectors corresponding to any current crawler information. The information type recognition model is trained on the basis of the BERT-CNN model using historical information and its type labels. The historical information includes historical target information and historical crawler information.

[0057] In one embodiment, the model building module (not shown in the figure) is used for: Slice any historical crawler information related to any historical target information into slices according to different data granularities to obtain multiple historical slice information corresponding to the historical crawler information. Multiple historical target information and multiple historical crawler information related to the multiple historical target information are input into the embedding layer of the BERT-CNN model to obtain the character vectors of each character in the multiple historical target information and the character vectors of each character in the multiple historical crawler information. The character vectors of each character in the multiple historical target information and the character vectors of each character in the multiple historical slice information corresponding to the multiple historical crawler information are input into the convolutional layer of the BERT-CNN model to obtain the vectors of multiple historical target information and the vectors of multiple historical slice information corresponding to the multiple historical crawler information. The multiple historical target information vectors, the multiple historical crawler information vectors corresponding to the multiple historical slice information, and the type labels corresponding to each historical information are input into the nonlinear classifier of the BERT-CNN model to obtain the type probability of the multiple historical target information and the type prediction probability of the multiple historical slice information corresponding to the multiple historical crawler information. If the cross-entropy loss value between the predicted probability of each information type and the true probability of the type is greater than or equal to the loss value threshold, then after adjusting the parameter weights of the BERT-CNN model, the process returns to the step of inputting multiple historical target information and multiple historical crawler information related to the multiple historical target information into the embedding layer of the BERT-CNN model until the cross-entropy loss value is less than the loss value threshold, and the BERT-CNN model at this time is determined as the information type recognition model.

[0058] In one embodiment, the information acquisition module 601 is specifically used for: Obtain multiple high-frequency information items related to the current target information in the network; The multiple high-popularity information entries are input into the second embedding layer of the information type recognition model to obtain the character vector of each character in the multiple high-popularity information entries; The character vectors of each character in the multiple high-popularity information are input into the convolutional layer of the information type recognition model to obtain the multiple high-popularity information vectors corresponding to the multiple high-popularity information. The multiple high-popularity information vectors are input into the nonlinear classifier of the information type recognition model to obtain the type prediction probability of the multiple high-popularity information vectors. Based on the predicted probabilities of the types of the multiple high-profile information items, the source of the crawler information is determined; Obtain multiple current crawler information entries related to the current target information from the crawler information source.

[0059] In one embodiment, the information acquisition module 601 is specifically used for: Sort the multiple type prediction probabilities of any high-popularity information in descending order, and obtain the first target information source of the type to which the multiple type prediction probabilities at the top of the sort are located. If there is a first matching information source in the first target information source that overlaps with the preset information source, then the first matching information source is determined as the crawler information source. If there is no first matching information source that overlaps with the preset information source in the first target information source, then after waiting for a preset time, return to the step of obtaining multiple high-popularity information related to the current target information in the network, until the second target information source is obtained; If there is a second matching information source in the second target information source that overlaps with the preset information source, then the second matching information source is determined as the crawler information source; If there is no second matching information source in the second target information source that overlaps with the preset information source, then the second target information source is matched with the first target information source; If there is a third matching information source in the second target information source that overlaps with the first target information source, then the third matching information source is determined as the crawler information source.

[0060] In one embodiment, the periodic monitoring module (not shown in the figure) is used for: Set a listening interface for the information source of any of the current crawler information; Periodically retrieve multiple monitoring messages related to the current target information from the monitoring interface; The multiple monitoring messages are treated as multiple current crawler messages. The process of slicing any current crawler message according to different data granularities is returned to obtain multiple current slice messages corresponding to any current crawler message, until the Euclidean distance between the multiple slice message vectors corresponding to any monitoring message and the current target message vector is obtained. Based on the slice information vector corresponding to the minimum Euclidean distance, a summary information and a jump link are generated, and the summary information and the jump link are pushed to the user.

[0061] Figure 7 is a schematic diagram of the structure of the electronic device provided in an embodiment of this application, as shown below. Figure 7 As shown, the electronic device may include: a processor 710, a communication interface 720, a memory 730, and a communication bus 740, wherein the processor 710, the communication interface 720, and the memory 730 communicate with each other via the communication bus 740. The processor 710 can call a computer program in the memory 730 to execute the steps of the information tracking method, such as including: Retrieve multiple pieces of current crawler information related to the current target information; Slice any current crawler information according to different data granularities to obtain multiple current slice information corresponding to the current crawler information; Vectorize the current target information and the multiple current slice information corresponding to any current crawler information to obtain the current target information vector and the multiple current slice information vector corresponding to any current crawler information; Calculate the average Euclidean distance between the multiple current slice information vectors corresponding to any current crawler information and the current target information vector; If the average Euclidean distance is less than the average threshold, then any current crawler information will be tracked.

[0062] Furthermore, the logical instructions in the aforementioned memory 730 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0063] On the other hand, this application also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can perform the steps of the information tracking methods provided in the above embodiments, such as: Retrieve multiple pieces of current crawler information related to the current target information; Slice any current crawler information according to different data granularities to obtain multiple current slice information corresponding to the current crawler information; Vectorize the current target information and the multiple current slice information corresponding to any current crawler information to obtain the current target information vector and the multiple current slice information vector corresponding to any current crawler information; Calculate the average Euclidean distance between the multiple current slice information vectors corresponding to any current crawler information and the current target information vector; If the average Euclidean distance is less than the average threshold, then any current crawler information will be tracked.

[0064] On the other hand, embodiments of this application also provide a non-transitory computer-readable storage medium storing a computer program thereon, the computer program being used to cause a processor to execute the steps of the information tracking methods provided in the above embodiments, for example including: Retrieve multiple pieces of current crawler information related to the current target information; Slice any current crawler information according to different data granularities to obtain multiple current slice information corresponding to the current crawler information; Vectorize the current target information and the multiple current slice information corresponding to any current crawler information to obtain the current target information vector and the multiple current slice information vector corresponding to any current crawler information; Calculate the average Euclidean distance between the multiple current slice information vectors corresponding to any current crawler information and the current target information vector; If the average Euclidean distance is less than the average threshold, then any current crawler information will be tracked.

[0065] The non-transitory computer-readable storage medium can be any available medium or data storage device that the processor can access, including but not limited to magnetic memory (e.g., floppy disk, hard disk, magnetic tape, magneto-optical disk (MO)), optical memory (e.g., CD, DVD, BD, HVD), and semiconductor memory (e.g., ROM, EPROM, EEPROM, non-volatile memory (NAND FLASH), solid-state drive (SSD)).

[0066] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0067] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0068] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. An information tracking method, characterized in that, include: Retrieve multiple pieces of current crawler information related to the current target information; Each current crawler record is sliced ​​according to different data granularities to obtain multiple current slice records corresponding to the given current crawler record, including: The characters in any one of the current crawler information are sliced ​​at equal intervals according to different spacings to obtain multiple current slice information corresponding to any one of the current crawler information; Vectorize the current target information and the multiple current slice information corresponding to any current crawler information to obtain the current target information vector and the multiple current slice information vector corresponding to any current crawler information; Calculate the average Euclidean distance between the multiple current slice information vectors corresponding to any current crawler information and the current target information vector; If the average Euclidean distance is less than the average threshold, then any current crawler information will be tracked.

2. The information tracking method according to claim 1, characterized in that, The step of vectorizing the current target information and the multiple current slice information corresponding to any current crawler information to obtain the current target information vector and the multiple current slice information vectors corresponding to any current crawler information includes: Input the current target information and multiple current slice information corresponding to any current crawler information into the first embedding layer of the information type recognition model to obtain the character vector of each character in the current target information and the character vector of each character in the multiple current slice information corresponding to any current crawler information. Input the character vectors of each character in the current target information and the character vectors of each character in multiple current slice information corresponding to any current crawler information into the convolutional layer of the information type recognition model to obtain the current target information vector and the multiple current slice information vectors corresponding to any current crawler information. The information type recognition model is trained on the basis of the BERT-CNN model using historical information and its type labels. The historical information includes historical target information and historical crawler information.

3. The information tracking method according to claim 2, characterized in that, The information type identification model is determined based on the following method: Slice any historical crawler information related to any historical target information into slices according to different data granularities to obtain multiple historical slice information corresponding to the historical crawler information. Multiple historical target information and multiple historical crawler information related to the multiple historical target information are input into the embedding layer of the BERT-CNN model to obtain the character vectors of each character in the multiple historical target information and the character vectors of each character in the multiple historical crawler information. The character vectors of each character in the multiple historical target information and the character vectors of each character in the multiple historical slice information corresponding to the multiple historical crawler information are input into the convolutional layer of the BERT-CNN model to obtain the vectors of multiple historical target information and the vectors of multiple historical slice information corresponding to the multiple historical crawler information. The multiple historical target information vectors, the multiple historical crawler information vectors corresponding to the multiple historical slice information, and the type labels corresponding to each historical information are input into the nonlinear classifier of the BERT-CNN model to obtain the type prediction probability of the multiple historical target information and the type prediction probability of the multiple historical crawler information corresponding to the multiple historical slice information. If the cross-entropy loss value between the predicted probability of each information type and the true probability of the type is greater than or equal to the loss value threshold, then after adjusting the parameter weights of the BERT-CNN model, the process returns to the step of inputting multiple historical target information and multiple historical crawler information related to the multiple historical target information into the embedding layer of the BERT-CNN model until the cross-entropy loss value is less than the loss value threshold, and the BERT-CNN model at this time is determined as the information type recognition model.

4. The information tracking method according to claim 2, characterized in that, The acquisition of multiple current crawler information related to the current target information includes: Obtain multiple high-frequency information items related to the current target information in the network; The multiple high-popularity information entries are input into the second embedding layer of the information type recognition model to obtain the character vector of each character in the multiple high-popularity information entries; The character vectors of each character in the multiple high-popularity information are input into the convolutional layer of the information type recognition model to obtain the multiple high-popularity information vectors corresponding to the multiple high-popularity information. The multiple high-popularity information vectors are input into the nonlinear classifier of the information type recognition model to obtain the type prediction probability of the multiple high-popularity information vectors. Based on the predicted probabilities of the types of the multiple high-profile information items, the source of the crawler information is determined; Obtain multiple current crawler information entries related to the current target information from the crawler information source.

5. The information tracking method according to claim 4, characterized in that, The step of determining the source of crawler information based on the probability prediction of the types of the multiple high-popularity information items includes: Sort the multiple type prediction probabilities of any high-popularity information in descending order, and obtain the first target information source of the type to which the multiple type prediction probabilities at the top of the sort are located. If there is a first matching information source in the first target information source that overlaps with the preset information source, then the first matching information source is determined as the crawler information source. If there is no first matching information source that overlaps with the preset information source in the first target information source, then after waiting for a preset time, return to the step of obtaining multiple high-popularity information related to the current target information in the network, until the second target information source is obtained; If there is a second matching information source in the second target information source that overlaps with the preset information source, then the second matching information source is determined as the crawler information source; If there is no second matching information source in the second target information source that overlaps with the preset information source, then the second target information source is matched with the first target information source; If there is a third matching information source in the second target information source that overlaps with the first target information source, then the third matching information source is determined as the crawler information source.

6. The information tracking method according to claim 1, characterized in that, After tracking any of the current crawler information, the process includes: Set a listening interface for the information source of any of the current crawler information; Periodically retrieve multiple monitoring messages related to the current target information from the monitoring interface; The multiple monitoring messages are treated as multiple current crawler messages. The process of slicing any current crawler message according to different data granularities is returned to obtain multiple current slice messages corresponding to any current crawler message, until the Euclidean distance between the multiple slice message vectors corresponding to any monitoring message and the current target message vector is obtained. Based on the slice information vector corresponding to the minimum Euclidean distance, a summary information and a jump link are generated, and the summary information and the jump link are pushed to the user.

7. An information tracking device, characterized in that, For performing the information tracking method of claim 1, comprising: The information acquisition module is used to: acquire multiple pieces of current crawler information related to the current target information; The slicing module is used to slice any current crawler information according to different data granularities to obtain multiple current slice information corresponding to the current crawler information. The vectorization module is used to: vectorize the current target information and multiple current slice information corresponding to any current crawler information to obtain the current target information vector and the multiple current slice information vector corresponding to any current crawler information; The distance calculation module is used to: calculate the average Euclidean distance between multiple current slice information vectors corresponding to any current crawler information and the current target information vector; The information tracking module is used to: track any current crawler information if the average Euclidean distance is less than the average threshold.

8. An electronic device comprising a processor and a memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the information tracking method according to any one of claims 1 to 6.

9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the information tracking method according to any one of claims 1 to 6.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the information tracking method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Similarity-based pest image retrieval method

    CN112559792A

  • Dialogue representation-based triage method, apparatus and device, and storage medium

    CN113223735A