Cross-domain document online traceability method and device

By time-slicing and parallel processing of streaming incremental Internet copy information, local semantic clusters are established and greedy merged, combined with active dynamic monitoring mechanisms, the problems of poor universality and difficulty in real-time processing in cross-domain copy tracing are solved, and accurate identification of cross-domain copy and real-time efficient processing of massive data are achieved.

CN120124633AActive Publication Date: 2025-06-10INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510123506.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-26
Publication Date
2025-06-10
Estimated Expiration
2045-01-26

AI Technical Summary

Technical Problem

The existing technology is poor in versatility when applied across fields, and it is difficult to meet the real-time processing requirements of massive data, and it is impossible to accurately identify cross-domain copywriting semantic associations.

Method used

By sharding the streaming incremental Internet copy information in time, extracting the earliest batch of uncalculated data from each time slice to form an incremental data set, the incremental data sets of different time slices are processed in parallel, local semantic clusters are established in each incremental data set, and greedily merged with the historical clustering results, an active class dynamic monitoring mechanism is introduced, and active classes whose members change in the clustering results are continuously tracked.

Benefits of technology

It realizes accurate identification and semantic correlation of cross-domain copywriting, supports real-time and efficient processing of massive data, and improves the universality of cross-domain applications, system throughput and response timeliness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120124633A_ABST
    Figure CN120124633A_ABST
Patent Text Reader

Abstract

The invention provides a cross-domain copywriting online traceability method and device, and the method comprises the steps: carrying out the time fragmentation of the obtained streaming incremental Internet copywriting information according to the data storage time, obtaining the data in each time slice, and extracting the uncalculated earliest batch of data from the data in the time slices to form an incremental data set; performing parallel processing on the incremental data sets of different time slices, establishing a local semantic cluster in each incremental data set, and performing greedy merging on the local semantic cluster and the historical clustering result to obtain a clustering result; introducing an active class dynamic monitoring mechanism, and continuously tracking the active class with the member number changing in the clustering result to obtain a tracing result; and storing the clustering result and the traceability result into a distributed search engine. According to the method, cross-domain copywriting semantic association can be accurately identified, and real-time efficient processing of mass data is supported.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of online traceability of copywriting, and particularly to a cross-domain copywriting online traceability method and device. Background Art

[0002] In the Internet era, the dissemination of information has shown an explosive growth, and various types of copywriting content continue to be generated. As a basic requirement, the core of copywriting traceability technology is to identify groups of copywriting with similar content and publishing intentions through semantic analysis technology, and determine their dissemination sources and scales. This technology has important application values in multiple fields, including intellectual property protection, brand reputation monitoring, public opinion analysis, marketing effect evaluation, etc. These application scenarios all rely on accurate and rapid copywriting traceability analysis results as the basis for decision-making support.

[0003] However, traditional keyword matching-based methods rely on domain-specific keyword lists and have poor generality in cross-domain applications. Methods based on pre-trained language models such as BERT have high computational complexity and are difficult to meet the real-time processing requirements of massive data. Therefore, there is an urgent need to propose a copywriting traceability method that can accurately identify cross-domain copywriting semantic associations and support real-time and efficient processing of massive data. Summary of the Invention

[0004] The present invention provides a cross-domain copywriting online traceability method and device to solve the defects of poor generality in cross-domain applications and difficulty in meeting the real-time processing requirements of massive data in the prior art, and to achieve accurate identification of cross-domain copywriting semantic associations and support real-time and efficient processing of massive data. The technical solutions proposed by the present invention are as follows: In a first aspect, the present invention provides a cross-domain copywriting online traceability method, including: Performing time slicing on the obtained streaming incremental Internet copywriting information according to the data storage time to obtain data within each time slice, and extracting the earliest batch of uncalculated data from the data within the time slice to form an incremental data set; Processing the incremental data sets of different time slices in parallel, establishing local semantic clusters within each incremental data set, and performing greedy merging with historical clustering results to obtain a clustering result; Introducing an active class dynamic monitoring mechanism to continuously track the active classes with changes in the number of members in the clustering result to obtain a traceability result; Storing the clustering result and the traceability result in a distributed search engine.

[0005] Optionally, the establishing local semantic clusters within each incremental data set and performing greedy merging with historical clustering results to obtain a clustering result includes: For each incremental data set, perform feature extraction and similarity calculation on the incremental data set, and construct a corresponding semantic graph based on the extracted semantic features and the calculated similarity relationships; wherein, the vertex set of the semantic graph corresponds to all texts in the incremental data set, the parameters of the vertices represent the semantic features of the corresponding texts, and the edge set of the semantic graph is generated according to the similarity relationships; Use the depth-first search algorithm to identify all connected components in the semantic graph, and each connected component forms an online cluster, and extract the centroid of each online cluster; For the centroid of each online cluster, use the fuzzy query mechanism to perform beam search, and adopt a greedy strategy based on large cluster priority for inter-cluster merging to obtain the clustering result.

[0006] Optionally, for the centroid of each online cluster, using the fuzzy query mechanism to perform beam search, and adopting a greedy strategy based on large cluster priority for inter-cluster merging to obtain the clustering result, includes: For the centroid of each online cluster, use the fuzzy query mechanism to perform beam search, retrieve the historical similar texts and their category identifiers similar to the centroid in the historical clustering results, and form a corresponding candidate category set according to the retrieval results; Adopt a greedy strategy based on large cluster priority for inter-cluster merging of online clusters and historical clusters, and select the category with the most intra-class nodes in the candidate category set as the attribution of the current online cluster to obtain the clustering result.

[0007] Optionally, the centroid of the online cluster is determined by the following method: In the online cluster, select the text with the earliest release time as the centroid of the online cluster.

[0008] Optionally, introducing an active class dynamic monitoring mechanism to continuously track the active classes whose member numbers change in the clustering result to obtain the traceability result, includes: For each time, determine the number of members of each cluster according to the clustering result at that time, and compare it with the number of members in the previous time window to determine the set of active classes at that time; For each time, calculate the active class traceability function, and the active class traceability function outputs the traceability results of all active classes in the active class set that meet the activity threshold; wherein, the traceability results include source information, time information, propagation path and propagation scale.

[0009] Optionally, the active class function traceability function is: Wherein, is the active class function traceability function, is the set of active classes, is the th active class in the set of active classes, For the active class The traceability result is time For the active class at time the number of members For the active class at time the number of members is the activity threshold is the time window

[0010] In a second aspect, the present invention further provides a cross - domain copywriting online traceability device, including the following modules: A time - slicing module, configured to perform time - slicing on the obtained streaming incremental Internet copywriting information according to the data storage time to obtain the data within each time slice, and extract the earliest batch of uncalculated data from the data within the time slice to form an incremental data set; A parallel copywriting aggregation module, configured to process the incremental data sets of different time slices in parallel, establish local semantic clusters within each incremental data set, and perform greedy merging with the historical clustering results to obtain a clustering result; A global copywriting traceability module, configured to introduce an active class dynamic monitoring mechanism to continuously track the active classes whose member numbers change in the clustering result to obtain a traceability result; A traceability result storage module, configured to store the clustering result and the traceability result into a distributed search engine.

[0011] In a third aspect, the present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and running on the processor, where when the processor executes the computer program, it implements the cross - domain copywriting online traceability method as described in the first aspect above.

[0012] In a fourth aspect, the present invention further provides a non - transitory computer - readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the cross - domain copywriting online traceability method as described in the first aspect above.

[0013] In a fifth aspect, the present invention further provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the cross - domain copywriting online traceability method as described in the first aspect above.

[0014] Based on the above technical solutions, the beneficial effects of the present invention compared with the prior art are: The cross - domain copywriting online traceability method and device provided by the present invention slice the streaming incremental Internet copywriting information by time, and extract the earliest batch of uncalculated data from each time slice to form an incremental data set. This processing method avoids processing the entire data set at once, thereby reducing the computational complexity. At the same time, since only the incremental data within each time slice is processed, real - time processing of new data can be achieved. This method processes the incremental data sets of different time slices in parallel, making full use of multi - core processors or distributed computing resources. Through parallel processing and the greedy merging algorithm, the data processing speed can be significantly improved, thus meeting the real - time processing requirements of massive data. At the same time, by introducing an active class dynamic monitoring mechanism, the active classes with changes in the number of members in the clustering results can be continuously tracked, so as to quickly obtain the traceability results. Moreover, this method establishes local semantic clusters within each incremental data set. These semantic clusters are divided based on the semantic similarity of the data, rather than relying on domain - specific keyword vocabularies. Therefore, even in data from different domains, as long as the copywriting information has similar themes or content, they will be divided into the same semantic cluster, thereby improving the generality of cross - domain applications. Through the greedy merging algorithm, the new incremental data set is integrated with the existing historical clustering results to gradually construct a more comprehensive and accurate clustering result. This integration method does not rely on specific domain knowledge, but is merged based on the semantic similarity of the data, and is applicable to cross - domain data processing.

[0015] Other features and advantages of the present invention will be described in the following specification, and, in part, will be obvious from the specification, or will be understood by implementing the present invention. The objectives and other advantages of the present invention are realized and obtained by the structures specifically pointed out in the specification, claims, and drawings.

[0016] To make the above - mentioned objectives, features, and advantages of the present invention more obvious and understandable, the following specifically gives preferred embodiments and, in conjunction with the accompanying drawings, makes a detailed description as follows. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0018] Figure 1 It is a schematic flowchart of the cross - domain copywriting online traceability method provided by the present invention.

[0019] Figure 2 It is a schematic overall framework diagram of the cross - domain copywriting online traceability method provided by the present invention.

[0020] Figure 3 It is a schematic diagram of the streaming copywriting clustering process provided by the present invention.

[0021] Figure 4 It is a schematic diagram of the structure of the cross-domain copywriting online traceability device provided by the present invention.

[0022] Figure 5 It is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed implementation manners

[0023] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.

[0024] The following combines Figures 1-4 to describe the cross-domain copywriting online traceability method and device of the present invention. The input is streaming incremental Internet copywriting information, including data storage time, publication time, publication content, and publication account. The output is copywriting groups, copywriting group sizes, and copywriting sources. The copywriting sources include key information such as the first publication time, first user, and platform.

[0025] The present invention proposes corresponding solutions for three core technical problems in the online copywriting traceability system. First, aiming at the problem that traditional keyword matching algorithms are difficult to effectively process cross-domain copywriting with significantly different semantic features, a cross-domain adaptive copywriting semantic clustering algorithm based on similarity connectivity is proposed. Second, to address the dual challenges of efficient clustering of stored data and real-time response to incremental data in a massive copywriting data environment, the clustering algorithm is reconstructed into a streaming computing mode, and combined with a parallel stream clustering algorithm of a dynamic semantic graph, a category information real-time update mechanism, and an efficient indexing mechanism, significantly improving the throughput capacity and response timeliness of the system. Aiming at the problem of diverse copywriting traceability requirements in different application scenarios, a clustering-traceability asynchronous decoupling mechanism is constructed. By separating the calculation of copywriting traceability and streaming clustering, the system can flexibly adapt to the traceability logics of different scenarios such as advertising marketing and public opinion supervision, significantly improving the scalability at the application level. Refer to Figure 1 As shown, the cross-domain copywriting online traceability method includes the following: Step S110: Perform time slicing on the obtained streaming incremental Internet copywriting information according to the data storage time to obtain the data within each time slice, and extract the earliest batch of uncalculated data from the data within the time slice to form an incremental data set.

[0026] Divide the streaming incremental copywriting information obtained from the Internet according to the data storage time, and slice it into several time slices. Each time slice contains data within a certain time period. Then, extract the earliest batch of data that has not been calculated or processed from the data within each time slice. These data constitute the incremental data set. This step ensures the timeliness of the data and the priority of processing.

[0027] Since the spread speed of information in the social network decays exponentially with time, most content copying and adaptation behaviors occur within a short time after the original information is released, making the text similarity higher for texts close in time. Therefore, the present invention introduces a data sharding mechanism based on the storage time window to shard the data, which not only ensures the local aggregation of high-similarity texts but also significantly reduces the overhead of cross-shard calculation.

[0028] The present invention proposes an adaptive parallel scheduling mechanism for time-series data streams. Based on the temporal locality characteristics of the data, this mechanism uses a dynamic time window to shard the data and constructs each time segment as an independent computing task unit. Specifically, the present invention proposes a method for information sharding and data set construction based on time-series consistent hashing. First, divide the streaming incremental copywriting information obtained from the Internet according to the data storage time. The purpose of this step is to slice the data into several time slices, and each time slice contains data within a certain time period. The size of the time slice (denoted as ) can be adjusted according to the dynamic situation of the actual computing resources to ensure the balance of the computing load. Through time sharding, the data can be divided into more manageable and processable units, while ensuring the local aggregation of high-similarity texts and reducing the overhead of cross-shard calculation.

[0029] Next, extract the earliest batch of data that has not been calculated or processed from the data within each time slice. These data constitute the incremental data set. This step is called micro-batch data set construction. Specifically, for the data within the time slice , , extract the earliest batch of uncalculated data in the order of storage time to form the incremental data set . This mechanism ensures that each piece of data will only be processed once, avoiding repeated calculations, while maintaining the temporality of data processing and ensuring the completeness of information traceability as much as possible. The incremental data set can be expressed as: where represents a single data record, which refers to one of the streaming incremental copywriting information obtained from the Internet. represents the start time point of the time slice. Indicates the processing status of the data, which is a boolean function that returns true when has been processed, and false otherwise. Indicates the data record that has not been processed yet. Among them, is the logical NOT symbol. Indicates the data record of the storage time. Indicates the size of the time slice, that is, the time span of the data contained within the time slice.

[0030] The above formula means: From all the data records whose storage times are within the , interval, filter out those data records that have not been processed yet to form the incremental data set within the time slice. . This index preparation stage lays a data foundation for subsequent parallel clustering and global traceability.

[0031] During the process of constructing the incremental data set, a multi-dimensional index structure is also established using the Elasticsearch distributed search engine, including basic attribute fields such as text content (content), publish time (publish_time), and storage time (insert_time). For the data that has completed clustering, its category identifier (cluster_id) and similarity features (similarity_features) are additionally indexed. These indexes lay a data foundation for subsequent parallel clustering and global traceability.

[0032] At the task scheduling level, the present invention designs an adaptive scheduling algorithm based on resource awareness. This algorithm realizes the on-demand allocation of computing resources and the dynamic scheduling of tasks by maintaining a fixed-scale computing resource pool and a dynamic task queue. For each incremental data set , it is constructed as an independent computing task unit and dynamically scheduled according to the current state of the computing resources and the length of the task queue. This mechanism realizes the efficient utilization of computing resources through task life cycle management, while ensuring the integrity and continuity of time series data processing.

[0033] Step S120: Process the incremental data sets of different time slices in parallel, establish local semantic clusters within each incremental data set, and perform greedy merging with the historical clustering results to obtain the clustering results.

[0034] Parallel processing is performed on the incremental data sets of different time slices to improve processing efficiency. Inside each incremental data set, local semantic clusters are established through specific algorithms or techniques. These semantic clusters represent collections of copywriting information in the data set that have similar themes or content. The local semantic clusters established within each time slice are greedily merged with the historical clustering results. Greedy merging is a method of gradually constructing a solution that makes the optimal choice at each step in the hope of obtaining a globally optimal solution. Through greedy merging, new incremental data can be gradually integrated with the existing clustering results to form a more comprehensive and accurate clustering result.

[0035] Utilize multi-core processors or distributed computing resources to perform parallel processing on the incremental data sets of different time slices. The incremental data set of each time slice is responsible for being processed by an independent computing unit (such as a thread, process, or computing node). Manage and coordinate the work of each computing unit through a parallel computing framework (such as Apache Spark, Dask, etc.). Refer to Figure 2 As shown, the incremental data set ((Micro-Chunk)) of each time slice respectively corresponds to a streaming copywriting clustering task, such as Figure 2 the streaming copywriting clustering task 1, the streaming copywriting clustering task 2,..., the streaming copywriting clustering task n in. The number of streaming copywriting clustering tasks corresponds to the number of time slices.

[0036] Specifically, preprocess each piece of copywriting information in the incremental data set, including steps such as word segmentation, stop word removal, and stemming, to extract key information. Use techniques such as TF-IDF, word vectors (such as Word2Vec, GloVe), or context embeddings (such as BERT, GPT, etc.) to convert the preprocessed copywriting information into numerical feature vectors. Select a clustering algorithm (such as K-means, DBSCAN, hierarchical clustering, etc.), and establish local semantic clusters within the incremental data set according to the feature vectors. The choice of clustering algorithm should be weighed according to the characteristics of the data and the clustering requirements. Thus, within each incremental data set, identify and extract collections of copywriting information with similar themes or content to form local semantic clusters.

[0037] Load the existing historical clustering results from the storage system, which may contain clustering information from previous time slices. Design and implement a greedy merging algorithm that selects the optimal merging method at each step in the current state. Specific greedy strategies can include: distance-based merging. Calculate the distances (such as Euclidean distance, cosine similarity, etc.) between the local semantic clusters and each cluster in the historical clustering results, and select the closest cluster for merging. Similarity-based merging, utilize the similarity features of the copywriting information (such as text similarity scores) to select the cluster with the highest similarity for merging. Size-based merging of clusters, consider the size of the clusters (i.e., the number of copywriting information within the clusters), and select clusters with similar sizes for merging to avoid forming overly large clusters. Post-process the merged clustering results, such as removing duplicate clusters, adjusting cluster boundaries, etc., to ensure the accuracy and consistency of the clustering results. Store the updated clustering results in the storage system for subsequent use. At the same time, update the index structure to ensure that the new clustering results can be efficiently retrieved and queried.

[0038] Step S130: Introduce an active class dynamic monitoring mechanism to continuously track the active classes with changing member numbers in the clustering results to obtain the traceability results.

[0039] The present invention proposes an asynchronous global traceability method based on active class awareness, which introduces an active class dynamic monitoring mechanism into the global traceability framework to continuously track the active classes with changing member numbers in the clustering results. These active classes represent the popular or widely concerned copywriting information themes on the current Internet. By monitoring the changes in the active classes, the source and propagation path of cross-domain copywriting can be discovered and tracked in a timely manner, thereby obtaining the traceability results. This mechanism continuously tracks the active classes with significant changes in member numbers in the clustering results, senses the propagation dynamics of copywriting information on the Internet in real time, and realizes a rapid response to popular or widely concerned copywriting information themes.

[0040] The system adopts an asynchronous decoupled architecture to decouple the active class monitoring and traceability analysis. This means that even when the traceability analysis is still in progress, the system can continue to monitor new active classes, thereby improving the overall processing efficiency. All traceability results are uniformly stored in a distributed search engine (such as ElasticSearch). This design not only ensures data consistency and traceability but also facilitates subsequent data retrieval and analysis. The system supports expanding multi-dimensional traceability rules according to the requirements of different application scenarios. For example, new algorithms or metrics can be introduced to optimize processes such as source identification and propagation path construction. Through the loose coupling design concept, each module of the system (such as the parallel copywriting aggregation module, the global copywriting traceability module, etc.) can evolve and be upgraded independently, thereby reducing the overall complexity and maintenance cost of the system.

[0041] Step S140: Store the clustering results and the traceability results in a distributed search engine.

[0042] Store the clustering results and traceability results in a distributed search engine. A distributed search engine is a system that can efficiently process and retrieve large-scale data. It improves retrieval efficiency and scalability by dispersing data storage across multiple nodes. This architecture can not only effectively manage massive data but also provide flexible data retrieval and analysis capabilities while maintaining high performance. In a distributed search engine, users can conveniently search for relevant document information and its traceability results through keywords or other search conditions. By storing the clustering results and traceability results in the distributed search engine in this invention, it can ensure that the clustering results and traceability results can be stored for a long time, facilitating subsequent data analysis and retrieval. Utilizing the efficient retrieval ability of the distributed search engine, users can quickly find relevant document information and its traceability results, thereby improving work efficiency. Through a unified data storage layer, it ensures the consistency between the clustering results and traceability results, avoiding problems of data conflicts or inconsistencies.

[0043] Before storage, perform necessary preprocessing on the clustering results and traceability results, such as data cleaning, format conversion, etc., to ensure data quality and compatibility. Utilize the indexing function provided by the distributed search engine to build indexes for the clustering results and traceability results. These indexes can be based on fields such as keywords, timestamps, propagation paths, etc., to improve retrieval efficiency. Upload the preprocessed clustering results and traceability results to each node of the distributed search engine. During the upload process, the system needs to ensure data integrity and consistency, avoiding data loss or duplication. After the upload is completed, the system conducts data verification to ensure that all clustering results and traceability results have been correctly stored and can be retrieved through indexes. Users can input search conditions such as keywords, time ranges, propagation paths, etc. through the search interface or API provided by the distributed search engine to quickly find relevant document information and its traceability results. Based on the stored clustering results and traceability results, users can conduct further data analysis, such as trend prediction, impact assessment, etc., to deeply understand the propagation dynamics and effects of document information. Utilize visualization tools to display the clustering results and traceability results in the form of charts, maps, etc., to help users more intuitively understand the data and analysis results.

[0044] As the data volume increases and retrieval requirements become more diverse, the distributed search engine needs to be continuously expanded and upgraded. This includes improvements in aspects such as increasing the number of nodes, optimizing the indexing algorithm, and enhancing the retrieval speed. At the same time, the system also needs to maintain compatibility with the clustering algorithm and the traceability analysis module to ensure seamless data docking and efficient processing. The above system refers to the overall technical solution or software architecture for performing clustering analysis, traceability analysis, and related data storage and retrieval functions.

[0045] The cross-domain copywriting online traceability method provided by the present invention slices the streaming incremental Internet copywriting information according to time, and extracts the earliest batch of uncalculated data from each time slice to form an incremental data set. This processing method avoids processing the entire data set at once, thereby reducing the computational complexity. At the same time, since only the incremental data within each time slice is processed, real-time processing of new data can be achieved. The method processes the incremental data sets of different time slices in parallel, making full use of multi-core processors or distributed computing resources. Through parallel processing and the greedy merging algorithm, the data processing speed can be significantly improved, thus meeting the real-time processing requirements of massive data. At the same time, by introducing an active class dynamic monitoring mechanism, the active classes with changes in the number of members in the clustering results can be continuously tracked, so as to quickly obtain the traceability results. Moreover, the method establishes local semantic clusters within each incremental data set, and these semantic clusters are divided based on the semantic similarity of the data, rather than relying on domain-specific keyword vocabularies. Therefore, even in data from different domains, as long as the copywriting information has similar themes or contents, they will be divided into the same semantic cluster, thereby improving the generality of cross-domain applications. Through the greedy merging algorithm, the new incremental data set is integrated with the existing historical clustering results to gradually construct a more comprehensive and accurate clustering result. This integration method does not rely on specific domain knowledge, but merges based on the semantic similarity of the data, and is applicable to cross-domain data processing.

[0046] In an optional embodiment, a streaming semantic clustering algorithm for cross-domain copywriting is proposed. This method captures the potential semantic associations between text copywritings through the Simhash algorithm combined with the Hamming distance, breaking through the limitations of traditional keyword matching methods in cross-domain scenarios. By calculating the semantic similarity between copywritings, a dynamically evolving similarity graph structure is constructed, and an adaptive clustering algorithm based on graph topology is adopted to effectively avoid the dependence of parametric clustering methods (such as k-means) on the prior number of categories. At the streaming processing level, this framework designs an incremental clustering strategy based on semantic consistency. This strategy realizes the dynamic update and maintenance of the clustering results by establishing local semantic clusters within a small batch of data and greedily merging them with the historical clustering results. This progressive clustering scheme not only ensures computational efficiency but also maintains semantic consistency across time windows. Experimental results show that this framework demonstrates significant advantages in semantic discovery and organization of cross-domain copywriting, and is particularly suitable for large-scale text clustering scenarios that require real-time processing and dynamic evolution.

[0047] The establishment of local semantic clusters within each incremental data set and the greedy merging with the historical clustering results to obtain the clustering result described in step S120 above includes: S1201. For each incremental data set, perform feature extraction and similarity calculation on the incremental data set, and construct a corresponding semantic graph based on the extracted semantic features and the calculated similarity relationship; wherein the vertex set of the semantic graph corresponds to all texts in the incremental data set, the parameters of the vertices represent the semantic features of the corresponding texts, and the edge set of the semantic graph is generated according to the similarity relationship.

[0048] For each incremental data set, feature extraction is first performed. The text data in the incremental data set (Micro-Chunk) is converted into a numerical feature vector that can represent its semantic content. Feature extraction methods can use bag-of-words model, TF-IDF, word embedding (such as Word2Vec, BERT), etc. Next, similarity calculation is performed. Based on the extracted semantic features, the similarity between each text in the incremental data set is calculated. Similarity calculation can use methods such as cosine similarity, Euclidean distance, Manhattan distance, etc.

[0049] After obtaining the similarity information, the corresponding semantic graph (SimGraph) is constructed based on the extracted semantic features and the calculated similarity relationship. The vertex set of the semantic graph corresponds to all the texts in the incremental dataset, and the parameters of each vertex represent the semantic features of the corresponding text. The edge set of the semantic graph is generated based on the similarity relationship between the texts: if the similarity between two texts exceeds the preset threshold, an edge is established between them. In this way, the semantic graph intuitively reflects the similarity relationship between the texts in the incremental dataset.

[0050] The above similarity algorithm can adopt Simhash algorithm. Figure 3 For example, taking the dynamic generation of online semantic graph based on Simhash as an example, the SimHash algorithm is used to extract features and calculate similarity of each text Doc1, Doc2, Doc3...Docm obtained from the incremental data set (Micro-Chunk). The SimHash algorithm maps the text into a 64-bit binary fingerprint to obtain the feature vector W of each text 1 , W 2 , W 3 … W m . The Hamming distance is used to measure text similarity. When the Hamming distance (HD) of the SimHash values ​​of two texts is ≤ 1, the two texts are considered similar. Based on the calculated similarity relationship, a semantic graph simG(V,E) is constructed, where the vertex set V corresponds to all texts in the data set, and the edge set E is generated according to the similarity relationship: if and only if the SimHash Hamming distance of two texts i and j is not greater than 1, an edge e(i,j) is established between them.

[0051] S1202. Identify all connected components in the semantic graph using the depth - first search algorithm. Each connected component forms an online cluster, and extract the centroid of each online cluster.

[0052] After obtaining the semantic graph, perform online micro - cluster identification based on graph semantic connectivity. Use the depth - first search algorithm (DFS) to identify all connected components in the semantic graph. Each connected component naturally forms an online cluster (CluOnline) C k , and these online clusters reflect the local semantic structure of the text in the incremental dataset. In each online cluster C k , extract the centroid centroid(C k ) as the representative of the cluster. The centroid can be the text with the earliest publishing time within the cluster. The centroid will be used in the subsequent cluster - merging steps. As Figure 3 shown, for each online cluster Micro CluOnline includes the following information: centroid, insert time, and feature. The centroids corresponding to Micro CluOnline1 and MicroCluOnline2 are Doc a and Doc z respectively, and the corresponding features are w a , w z . The texts of all online clusters are aggregated to obtain an online - cluster text list (CluOnline doc list), which includes doc a, doc b, doc c, doc d, doc e, doc f.

[0053] S1203. For the centroid of each online cluster, perform beam search using the fuzzy query mechanism and adopt a greedy strategy based on large - cluster priority for cluster - merging to obtain the clustering result.

[0054] After extracting the centroid of each online cluster, perform online - cluster - historical - cluster fusion based on fuzzy retrieval. For the centroid of each online cluster, perform beam search using the fuzzy query mechanism. The fuzzy query mechanism allows considering the approximate matching relationship between texts during the merging process, thereby increasing the flexibility and accuracy of the merging. Based on the beam search, adopt a greedy strategy based on large - cluster priority for cluster - merging. This strategy tends to merge larger clusters first because larger clusters usually contain more information and are more stable. During the merging process, continuously calculate the centroid of the newly merged cluster and update the clustering result. The centroid of the old cluster is called Old centroid, and the centroid of the new cluster is called New centroid.

[0055] The above-mentioned step S120 is not only applicable to the clustering process of incremental data sets, but also can be seamlessly docked with historical clustering results. Through the greedy merging strategy, the dynamic update and maintenance of clustering results are realized, which not only ensures the computing efficiency but also maintains the semantic consistency across time windows. In addition, the present invention combines a variety of advanced technologies such as a parallel-stream hybrid processing framework, an adaptive parallel scheduling mechanism for time-series data streams, and an asynchronous global traceability method based on active class awareness, further improving the overall performance and practicality of the system. Experimental results show that this framework exhibits significant advantages in semantic discovery and organization of cross-domain copywriting, and is particularly suitable for large-scale text clustering scenarios that require real-time processing and dynamic evolution. At the macro level, for multiple incremental data sets in different time slices parallel processing is achieved, and each incremental data set can be assigned to an independent computing node for processing. In each incremental data set internally, the clustering process is executed in strict serial steps. This hybrid clustering method improves the system throughput through parallel processing among incremental data sets, and at the same time adopts a strict serial processing mechanism within a single incremental data set, maximizing the stability of clustering results. In addition, the hybrid indexing mechanism based on SimHash and ElasticSearch enables the system to achieve efficient similar text retrieval and category merging while maintaining clustering flexibility.

[0056] The present invention adopts the method of incremental clustering, only processing the newly added data set each time, avoiding recalculating the clustering results of the entire data set, thereby improving the clustering efficiency. This method can flexibly process continuously changing data sets and adapt to the dynamic changes of data by continuously updating the clustering results. Through steps such as feature extraction, similarity calculation, semantic graph construction, and greedy merging, this method can maintain a high clustering quality. In particular, the fuzzy query mechanism and the greedy strategy based on large cluster priority help to maintain the stability and representativeness of clusters during the merging process. The constructed semantic graph and the extracted centroids can be used to support complex query operations such as semantic-based similarity search and text classification.

[0057] In an optional embodiment, for the centroid of each online cluster in the above-mentioned step S1203, beam search is performed using the fuzzy query mechanism, and cluster merging is performed using the greedy strategy based on large cluster priority to obtain the clustering result, including: S12031. For the centroid of each online cluster, beam search is performed using the fuzzy query mechanism to retrieve the historical similar texts and their category identifiers similar to the centroid in the historical clustering results, and a corresponding candidate category set is formed according to the retrieval results.

[0058] For the centroid of each online cluster, use a fuzzy query mechanism (such as fuzzy matching, similarity calculation, etc.) to search for historical similar texts similar to the centroid in the historical clustering results, and obtain a list of similar texts (similar doc list). Fuzzy queries allow a certain degree of error, so that relevant historical data can be found even if there is no exact match. In the results of the fuzzy query, use the beam search algorithm to limit the search space and only retain several historical texts and their category identifiers (Cluster id) that are most similar to the centroid, such as Figure 3 in . The beam search avoids the explosion of the search space by maintaining a candidate set of a fixed size. According to the results of the beam search, form a candidate category set corresponding to each online cluster, and these sets contain the category identifiers related to the historical similar texts.

[0059] S12032. Adopt a greedy strategy based on large clusters first for the inter-cluster merging of online clusters and historical clusters, and select the category with the most intra-class nodes in the candidate category set as the attribution of the current online cluster to obtain the clustering result.

[0060] After obtaining the candidate category set for each online cluster, adopt a greedy strategy for inter-cluster merging. The core idea of the greedy strategy is to make the optimal choice at each step, that is, the choice that seems optimal in the current state, hoping to lead to a globally optimal (or close to globally optimal) result. In the greedy strategy, larger clusters (i.e., clusters containing more nodes) are preferentially considered as historical clusters for merging. This is because larger clusters usually contain more information and are more likely to represent stable categories. For each online cluster, select the category with the most intra-class nodes in its candidate category set as the attribution of the current online cluster. This step ensures that the online cluster is classified into the most similar and stable category. After the above steps, the final clustering result is obtained, where each online cluster is classified into a specific category.

[0061] Specifically, adopt a greedy strategy based on large clusters first for the inter-cluster merging of online clusters and historical clusters, and select the category with the most intra-class nodes in the candidate category set as the attribution of the current online cluster. This process is as follows: S210. First, perform initialization: Create an empty cluster set clusters to store all online clusters. Create a candidate category set categories for each cluster to store the categories to which each cluster belongs.

[0062] S220. For each online cluster, if it has not been assigned to a certain category yet, traverse the category set categories, find the cluster with the largest number of nodes, i.e., the largest-sized cluster, as the historical cluster to be merged, obtain the category of this historical cluster, and use the category of this historical cluster as the merged category. Assign the current online cluster to this category, assign the merged category to all the text data of the online cluster, and update the number of nodes of the corresponding category in the candidate category set categories.

[0063] S230. Repeat step S220, continuously receive new text data, and repeat the above steps to perform the selection of inter-cluster merging and category assignment with historical clusters until the stop condition is met (such as reaching a preset time window, processing all input data, etc.).

[0064] S240. The cluster set clusters will contain all the merged online clusters, and each online cluster has a clear category assignment. Output the clustering result, which can include information such as the text content, number of nodes, and category assignment of each online cluster.

[0065] By performing inter-cluster merging of online clusters and historical clusters through a greedy strategy based on large clusters first, larger clusters can be retained while reducing the number of small clusters, thereby improving the overall quality of clustering. Selecting the category of the historical cluster with the most nodes within the class as the assignment can ensure that each cluster has a clear category label, which helps with subsequent analysis and processing. Moreover, the amount of updated historical information is minimized, which can accelerate the calculation. Compared with other complex clustering algorithms, such as hierarchical clustering and spectral clustering, the computational complexity of this algorithm is relatively low and is suitable for processing large-scale online data. By setting a similarity threshold, unnecessary similarity calculations can be reduced, further improving the efficiency of the algorithm. This algorithm can adapt to different data types and feature representations, and only needs to select a suitable similarity calculation method. The algorithm has good adaptability to the addition of new data and can update the clustering result in real time.

[0066] Through the fuzzy query mechanism and beam search, the present invention can find the online cluster most similar to the historical data, thereby improving the accuracy of clustering. At the same time, the greedy strategy based on large clusters first ensures the stability of the clustering result. The beam search algorithm limits the size of the search space and avoids excessive consumption of computing resources. The greedy strategy reduces the overall computational complexity by selecting the optimal solution at each step. This process can process online data (i.e., data that arrives in real time) and dynamically classify it into appropriate categories, and is suitable for real-time clustering tasks. Due to the adoption of fuzzy query and greedy strategy, this process can be easily extended to large-scale data sets while maintaining good performance and accuracy.

[0067] In an optional embodiment, the centroid of the online cluster is determined in the following manner: In the online clusters, the text with the earliest release time is selected as the centroid of the online cluster. Specifically, in each online cluster, the text with the earliest release time is found and recorded. This earliest released text is used as the centroid of the online cluster. This selection not only ensures the timeliness of tracing the source but also reduces the computational complexity.

[0068] By selecting the text with the earliest release time as the centroid, the present invention can reflect the origin and early development of the online cluster, which helps to understand the formation background and initial content of the cluster. Moreover, selecting the text with the earliest release time as the centroid is simpler and more direct in calculation, reducing the amount of calculation. In an online environment, text data is updated in real time. Selecting the earliest released text as the centroid can quickly determine the representative of each cluster, which is very beneficial for real-time analysis and processing. As the representative of the cluster, the content and characteristics of the centroid are easy to understand and interpret. Selecting the text with the earliest release time as the centroid, its content and background are easier to understand and accept.

[0069] In an optional embodiment, on the basis of cluster analysis, in order to continuously track the active classes whose member numbers change in the clustering results and obtain their source tracing results, the present invention introduces an active class dynamic monitoring mechanism. The core of this mechanism lies in analyzing the clustering results of each time window, determining the set of active classes, and calculating the active class source tracing function to obtain the source tracing results.

[0070] Introducing the active class dynamic monitoring mechanism described in step S130 above and continuously tracking the active classes whose member numbers change in the clustering results to obtain the source tracing results includes: S1301. For each time, determine the number of members of each cluster according to the clustering results at this time, compare it with the number of members in the previous time window, and determine the set of active classes at this time.

[0071] The present invention will be at a specific time , the set of clustering result sets where all member numbers change significantly is defined as the set of active classes , For all active classes at the moment. Each active class represents a theme of copywriting information that is currently widely concerned. Define the source tracing result: . The number of members in each time window is pre-stored in ES. When generating the set of active classes, it can be obtained in real time from ES. Among them,

[0072] where is the source information function; is the time information function; is the propagation path function; is the propagation scale function.

[0073] The system regularly counts the number of members in each cluster and compares the changes in the number of members within adjacent time windows (such as and ). If the change in the number of members of a certain cluster exceeds the preset activity threshold , it is regarded as an active class. After detecting an active class, the centroid and scale of the online cluster are updated in real time.

[0074] For each time window, first determine the number of members in each cluster according to the clustering result at that time. Then, compare the number of members in the current time window with the number of members in the previous time window. If the number of members of a certain cluster has changed significantly (increased or decreased), it is regarded as an active class and added to the set of active classes in the current time window.

[0075] S1302. For each time, calculate the active class traceability function, and the active class traceability function outputs the traceability results of all active classes in the active class set that meet the activity threshold; where the traceability results include source information, time information, propagation path, and propagation scale.

[0076] After determining the active class set, calculate the active class traceability function to obtain the traceability results. The active class traceability function is a complex calculation process that considers multiple factors such as the source information, time information, propagation path, and propagation scale of the active class. The output of the function is a set containing the traceability results of all active classes that meet the activity threshold.

[0077] The active class function traceability function can be defined as: where is the active class function traceability function, is the active class set, is the th active class in the active class set, is the traceability result of the active class , is the time, is the number of members of the active class at time , is the number of members of the active class at time , is the activity threshold, is the time window.

[0078] The traceability result It can include information in multiple aspects, such as source information (i.e., the origin or the initial formation reason of the active class), time information (including the time when the active class appears, the duration, etc.), the propagation path (i.e., the path by which the active class spreads or evolves over time), and the propagation scale (i.e., the number of members or the influence size of the active class in different time windows).

[0079] For each identified as an active class, the system performs the following steps for calculating the traceability results: obtaining source information, recording time information, constructing the propagation path, and evaluating the propagation scale. Obtaining source information means that through functions, the system traces and determines the original source of the copywriting information. Recording time information means using functions to record the key time points when the copywriting information spreads from the source. Constructing the propagation path means through functions to construct the propagation path of the copywriting information from the source to the current active class. Evaluating the propagation scale means using functions to evaluate the influence range and scale of the copywriting information during the propagation process.

[0080] By introducing the active class dynamic monitoring mechanism, the present invention can monitor the change in the number of members in the clustering results in real time and respond quickly to these changes. This helps to timely discover and handle potential hotspots or trends. The active class traceability functions comprehensively consider multiple factors such as source information, time information, propagation path, and propagation scale, thereby improving the accuracy of traceability. This helps to more deeply understand the formation reason and propagation process of the active class. By obtaining the traceability results of the active class, it can provide strong support for decision-making. For example, in the field of marketing, this information can be used to formulate targeted marketing strategies or adjust the product direction. This mechanism can adapt to different data types and feature representations, and only needs to appropriately adjust the active class traceability functions. In addition, it can also be integrated with other algorithms or modules to form a complete processing flow.

[0081] The cross-domain copywriting online traceability method proposed by the present invention involves adaptive parallel scheduling for time-series data streams, adaptive stream clustering based on dynamic semantic graphs, and an asynchronous global traceability framework based on active class perception. Through the adaptive parallel scheduling method for time-series data streams of the present invention, it is possible to dynamically adjust the computing resource allocation strategy according to the data stream characteristics, improving the system processing capacity and resource utilization efficiency. This method adaptively performs task allocation and load balancing by real-time sensing the time-series characteristics and load changes of the data stream, effectively solving the performance bottleneck problem of traditional fixed scheduling schemes in the face of burst traffic.

[0082] The above-mentioned adaptive flow clustering method based on dynamic semantic graphs can, compared with traditional static clustering algorithms, update the semantic feature space in real time and dynamically adjust the clustering results, greatly improving the generality and real-time performance of semantic clustering. By constructing and maintaining a dynamic semantic graph, this method enables the clustering results to adaptively adjust with the evolution of text semantics, overcoming the problems of lag and insufficient accuracy in traditional clustering methods when dealing with dynamic text streams.

[0083] The asynchronous global traceability framework based on active class awareness proposed in the present invention adopts an asynchronous processing strategy with priority given to active classes, greatly improving the response speed and processing efficiency of the system. By monitoring and tracking the evolution of active classes in real time, this framework preferentially allocates resources for traceability analysis of key clusters, effectively solving the performance and real-time problems of traditional methods in the face of large-scale text traceability. At the same time, the asynchronous design ensures the scalability and fault tolerance of the system.

[0084] In summary, the present invention greatly improves the timeliness of online traceability, has stronger domain adaptation ability and transferability, and at the same time improves the scalability and stability of the system. These technological innovations effectively solve the key problems such as efficiency bottlenecks, lack of universality, and resource scheduling in copywriting traceability, providing a comprehensive and reliable solution for large-scale copywriting traceability applications.

[0085] The cross-domain copywriting online traceability device provided by the present invention will be described below. The cross-domain copywriting online traceability device described below can be correspondingly referred to the cross-domain copywriting online traceability method described above.

[0086] The cross-domain copywriting online traceability device provided by the present invention, with reference to Figure 4 shown, includes: A time slicing module 310, configured to perform time slicing on the obtained streaming incremental Internet copywriting information according to the data storage time to obtain the data within each time slice, and extract the earliest uncalculated batch of data from the data within the time slice to form an incremental data set; A parallel copywriting aggregation module 320, configured to process the incremental data sets of different time slices in parallel, establish local semantic clusters within each incremental data set, and perform greedy merging with the historical clustering results to obtain a clustering result; A global copywriting traceability module 330, configured to introduce an active class dynamic monitoring mechanism to continuously track the active classes whose member numbers change in the clustering result to obtain a traceability result; A traceability result storage module 340, configured to store the clustering result and the traceability result in a distributed search engine.

[0087] Figure 5 Illustrates a schematic diagram of the physical structure of an electronic device, as Figure 5As shown in the figure, the electronic device may include: a processor 410, a communications interface 420, a memory 430, and a communication bus 440. Among them, the processor 410, the communications interface 420, and the memory 430 complete their mutual communication through the communication bus 440. The processor 410 may call the logical instructions in the memory 430 to execute the online traceability method for cross-domain copywriting.

[0088] In addition, when the logical instructions in the above-mentioned memory 430 can be implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0089] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the online traceability method for cross-domain copywriting provided by the above-mentioned various methods.

[0090] On yet another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the online traceability method for cross-domain copywriting provided by the above-mentioned various methods.

[0091] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.

[0092] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0093] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A cross-domain copywriting online tracing method, characterized in that: include: The acquired streaming incremental Internet copywriting information is time-sliced ​​according to the data storage time to obtain the data in each time slice, and the earliest batch data that has not been calculated is extracted from the data in the time slice to form an incremental data set; Incremental data sets of different time slices are processed in parallel, local semantic clusters are established in each incremental data set, and clustering results are obtained by greedy merging with historical clustering results; Introduce a dynamic monitoring mechanism for active classes to continuously track active classes whose number of members has changed in the clustering results to obtain traceability results; The clustering result and the tracing result are stored in a distributed search engine.

2. The cross-domain copywriting online tracing method according to claim 1 is characterized in that: The local semantic clusters are established in each incremental data set, and greedily merged with the historical clustering results to obtain the clustering results, including: For each incremental data set, feature extraction and similarity calculation are performed on the incremental data set, and a corresponding semantic graph is constructed based on the extracted semantic features and the calculated similarity relationship; wherein the vertex set of the semantic graph corresponds to all texts in the incremental data set, the parameters of the vertices represent the semantic features of the corresponding texts, and the edge set of the semantic graph is generated according to the similarity relationship; A depth-first search algorithm is used to identify all connected components in the semantic graph. Each connected component constitutes an online cluster, and the centroid of each online cluster is extracted. For the centroid of each online cluster, a beam search is performed using a fuzzy query mechanism, and a greedy strategy based on large cluster priority is adopted to merge clusters to obtain the clustering results.

3. The cross-domain copywriting online tracing method according to claim 2 is characterized in that: For the centroid of each online cluster, a beam search is performed using a fuzzy query mechanism, and a greedy strategy based on large cluster priority is adopted to merge clusters to obtain clustering results, including: For the centroid of each online cluster, a beam search is performed using a fuzzy query mechanism to retrieve historical similar texts and their category identifiers that are similar to the centroid in the historical clustering results, and a corresponding candidate category set is formed based on the retrieval results; A greedy strategy based on large cluster priority is used to merge online clusters and historical clusters, and the category with the most nodes in the candidate category set is selected as the current online cluster to obtain the clustering result.

4. The cross-domain copywriting online tracing method according to claim 2 is characterized in that: The centroid of an online cluster is determined as follows: In the online cluster, the text with the earliest publishing time is selected as the centroid of the online cluster.

5. The cross-domain copywriting online tracing method according to claim 1 is characterized in that: The active class dynamic monitoring mechanism is introduced to continuously track the active classes whose number of members in the clustering results has changed to obtain the tracing results, including: For each time, the number of members in each cluster is determined based on the clustering results at that time, and compared with the number of members in the previous time window to determine the active class set at that time; For each time, an active class tracing function is calculated, and the active class tracing function outputs the tracing results of all active classes in the active class set that meet the activity threshold; wherein the tracing results include source information, time information, propagation path, and propagation scale.

6. The cross-domain copywriting online tracing method according to claim 5 is characterized in that: The active class function tracing function is: in, For active class function tracing function, is the active class set, is the first in the active class set Active classes, For active class The traceability results, For time, For active class In time The number of members, For active class In time The number of members, is the activity threshold, is the time window.

7. A cross-domain copywriting online source tracing device, characterized in that: include: The time slicing module is used to time slice the acquired streaming incremental Internet copywriting information according to the data storage time to obtain the data in each time slice, and extract the earliest batch data that has not been calculated from the data in the time slice to form an incremental data set; The parallel copywriting aggregation module is used to process incremental data sets of different time slices in parallel, establish local semantic clusters in each incremental data set, and greedily merge with historical clustering results to obtain clustering results; The global copywriting tracing module is used to introduce a dynamic monitoring mechanism for active classes, and continuously track the active classes whose number of members has changed in the clustering results to obtain tracing results; The tracing result storage module is used to store the clustering result and the tracing result in a distributed search engine.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, it implements the cross-domain document online tracing method as described in any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the cross-domain document online tracing method as described in any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the cross-domain document online tracing method as described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Quick hot spot detection method and system based on mass news data

    CN108304502A

  • Distributed system fault root cause tracing method based on knowledge graph technology

    CN113377567A

  • Intrusion behavior-oriented tracing data clustering method and device

    CN113612749A

  • Network data tracing method and device, electronic equipment and storage medium

    CN115757912A

  • Text verbal skill clustering method and device, computer equipment and storage medium

    CN116484003A