Vertical website deep crawling method based on intention recognition

By building a dynamic knowledge graph and multi-source crawling strategy, identifying weak signal paths, and simulating user behavior, we solved the adaptability and depth problems of vertical website data crawling, and achieved efficient and accurate data acquisition and risk verification.

CN120780889AActive Publication Date: 2025-10-14ZHEJIANG FULIN TECH CO LTD

Patent Information

Application Number
CN202511301800.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-12
Publication Date
2025-10-14
Estimated Expiration
2045-09-12

AI Technical Summary

Technical Problem

Existing technologies have poor adaptability in vertical website data crawling, high maintenance costs, difficulty in obtaining deep data, limited data extraction accuracy, and are unable to simulate complex user interaction processes, resulting in insufficient crawling depth and breadth.

Method used

By building a dynamic knowledge graph, identifying weak signal association paths, generating data collection intentions, adopting a multi-source crawling strategy, simulating real user behavior, executing deep crawling tasks, and performing target data processing and risk hypothesis verification.

Benefits of technology

It realizes intelligent, efficient and accurate data capture of vertical websites, reduces dependence on website structure, improves the pertinence and automation of data collection, and enhances the ability to collect in-depth information in complex financial fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120780889A_ABST
    Figure CN120780889A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of information retrieval, in particular to a vertical website deep crawling method based on intention recognition, which comprises the following steps: crawling data from a plurality of heterogeneous information sources, and constructing a dynamic knowledge graph; identifying a weak signal association path formed by connecting a plurality of relation edges in the dynamic knowledge graph, and converting the weak signal association path into a to-be-verified risk hypothesis; generating a data acquisition intention, and based on the to-be-verified risk hypothesis, generating a rejection query for searching reverse evidence as a first data acquisition intention; receiving an input query instruction, and performing semantic analysis on the query instruction to generate a second data acquisition intention; a multi-source crawling strategy is generated based on the type of the data collection intention, and the multi-source crawling strategy comprises a plurality of vertical website sources and crawling priorities of the vertical website sources; executing a deep crawling task according to the multi-source crawling strategy to obtain target data; and processing the target data, verifying the to-be-verified risk hypothesis, and judging whether the reverse evidence exists or not and judging the intensity of the reverse evidence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information retrieval technology, and in particular to a vertical website deep crawling method based on intent recognition. Background Art

[0002] Web crawlers, as programs that automatically retrieve webpage information, are core tools for data capture and information retrieval. Vertical crawlers, among others, focus on websites with specific themes or industries, aiming to accurately retrieve structured data related to those specific fields. Currently, the mainstream approach to crawling vertical websites relies on pre-configured crawling rules.

[0003] However, existing technologies suffer from at least the following technical issues: First, they suffer from poor adaptability and high maintenance costs. Crawling methods based on fixed rules are highly dependent on the website's page structure. Once the target website's page layout, tag attributes, or front-end code changes, the preset extraction rules become invalid, resulting in data crawling failures or errors. This requires significant manpower to reanalyze and reconfigure the rules, resulting in a low level of automation. Second, "deep" data is difficult to effectively capture. High-value data on vertical websites typically requires a series of interactive operations, such as searching, page turning, and filtering. Traditional crawlers struggle to simulate this complex, user-intent-based interaction process, resulting in an inability to access and crawl the "deep" data hidden behind these interactions, resulting in a serious lack of depth and breadth in crawling. Finally, the accuracy of data extraction needs to be improved. Existing technologies have limited ability to distinguish between core and non-core content on a page, often capturing a large amount of irrelevant, redundant information. This not only increases the burden of subsequent data cleaning but also reduces the signal-to-noise ratio of the resulting structured data.

[0004] Therefore, how to reduce dependence on specific website structures, automatically understand and simulate users' intentions to obtain core data, and thus achieve intelligent, efficient and accurate capture of "deep" data of vertical websites is a technical problem that needs to be solved urgently in this field.

[0005] To this end, a vertical website deep crawling method based on intent recognition is proposed. Summary of the Invention

[0006] The purpose of the present invention is to provide a vertical website deep crawling method based on intent recognition, including crawling data from multiple heterogeneous sources to construct a dynamic knowledge graph; identifying a weak signal association path composed of multiple relationship edge connections in the dynamic knowledge graph, and converting the weak signal association path into a risk hypothesis to be verified; generating a data collection intention, and based on the risk hypothesis to be verified, generating a refutation query for finding reverse evidence as a first data collection intention; receiving an input query instruction, performing semantic analysis on the query instruction, and generating a second data collection intention; based on the type of data collection intention, generating a multi-source crawling strategy, including multiple vertical website sources and their crawling priorities; executing a deep crawling task according to the multi-source crawling strategy to obtain target data; processing the target data, verifying the risk hypothesis to be verified, and judging whether the reverse evidence exists and its strength.

[0007] To achieve the above object, the present invention provides the following technical solutions: A vertical website deep crawling method based on intent recognition, including: Crawl data from multiple heterogeneous sources to construct a dynamic knowledge graph, which includes entity nodes and time-stamped edges that represent the relationships between entity nodes. Identify weak signal association paths formed by multiple edge connections in the dynamic knowledge graph and convert these weak signal association paths into risk hypotheses to be verified. Generate a data collection intention, where the types of data collection intentions include: generating a refutation query to find counter-evidence based on a risk hypothesis to be verified, and using the refutation query as a first data collection intention; receiving an input query instruction, and performing semantic analysis on the query instruction to generate a second data collection intention of the query type; Based on the type of data collection intention, a multi-source crawling strategy is generated, wherein the multi-source crawling strategy includes multiple vertical website sources to be crawled and their crawling priorities; according to the multi-source crawling strategy, a deep crawling task is executed to obtain target data; According to the type of the data collection intention, the target data is processed, the risk hypothesis to be verified is verified, and the existence and strength of the counter-evidence are determined.

[0008] Preferably, the specific process of obtaining the weak signal correlation path includes: Executing a multi-hop path traversal algorithm on the dynamic knowledge graph to identify entity node connection paths whose lengths meet a preset value; Performing a timing consistency analysis on the connection paths to select path combinations whose timestamps have logical relevance; Calculate the weight of each relationship edge in the path, the weight value is comprehensively evaluated based on the relationship strength, time correlation and rarity; Based on a path importance scoring algorithm that comprehensively considers path length, node importance, and relationship edge weight, weak signal association paths with scores exceeding a preset threshold are selected.

[0009] Preferably, the specific process of obtaining the risk hypothesis to be verified includes: The weak signal association path is input into a pre-trained financial domain language model to extract key entities and relationship descriptions in the path; the path information is converted into a structured causal relationship description, which includes preconditions, reasoning process and expected results; and the causal relationship description is converted into a risk hypothesis to be verified in natural language form.

[0010] Preferably, the specific process of generating the data collection intention includes: Based on the risk hypothesis to be verified, three types of refutation queries are automatically constructed through refutation queries: disproof query, strength query and alternative explanation query. The disproof query is used to find evidence that can negate the premise of the risk hypothesis. The strength query is used to verify the logical strength of the causal relationship in the hypothesis. The alternative explanation query is used to discover explanatory factors that weaken the risk hypothesis. Performing semantic analysis on the query instruction using a preset intent classification model to classify the query instruction into one of a factual query, a list query, and a comparative query; Based on the second data collection intention, a search is first performed in the dynamic knowledge graph. If there is data that meets the query instruction in the dynamic knowledge graph, the query instruction is directly responded to. If the query is incomplete, the missing data part is used as the target of the subsequent crawling task.

[0011] Preferably, according to the type of the data collection intention, multiple vertical website sources are selected from a preset information source library, and the information source library includes general search engines, academic databases, legal information databases, and financial regulatory agency websites; Based on the urgency of data collection and data quality requirements, different crawling priorities are set for the selected multiple vertical website sources, including high priority, medium priority and low priority; Develop exclusive crawling parameter configurations for each vertical website source, including access frequency limit, number of concurrent threads, timeout period, and number of retries.

[0012] Preferably, the specific process of acquiring the target data includes: Perform deep crawling tasks on each vertical website source in turn according to the crawling priority order; For each vertical website source, we first conduct a website structure analysis to identify the page hierarchy and URL pattern where the target data is located; Adopt a crawling strategy that simulates real user behavior, including randomizing access intervals, simulating mouse click trajectories, and constructing request header information that conforms to website characteristics; For dynamically loaded content encountered during the crawling process, the page is fully loaded before data extraction; Establish a crawling progress monitoring mechanism to track the data acquisition status of each website source in real time, and automatically switch to alternative crawling strategies under abnormal circumstances.

[0013] Preferably, the specific process of determining whether counter-evidence exists and its strength includes: Based on the target data obtained by the first data collection intention, the evidence fragments related to the refutation query are identified through text similarity algorithm and keyword matching technology; the evidence fragments are classified and labeled as strong reverse evidence, weak reverse evidence and irrelevant evidence, and the strength value of the evidence is calculated; according to the labeling and strength value of the reverse evidence, the confidence of the risk hypothesis to be verified is scored; when the strength of the strong reverse evidence exceeds a preset threshold, the corresponding risk hypothesis is marked as falsified.

[0014] Compared with the prior art, the present invention has the following beneficial effects: 1. The present invention provides an intent recognition mechanism, specifically dividing data collection tasks into multiple types of collection intents, generating a refutation query as the first data collection intent based on a risk hypothesis to be verified, and generating a second data collection intent based on a query instruction after semantic analysis, thereby matching the optimal crawling strategy according to different data requirements. This improves the targeted nature of data collection and avoids the waste of extensive information crawling. At the same time, driven by different types of intents, it can automatically adjust crawling priority, information source selection, and crawling depth to achieve scenario-oriented data acquisition accuracy and effectively meet the in-depth information collection needs of complex financial verticals.

[0015] 2. The present invention constructs a dynamic knowledge graph, establishes semantic relationship edges between entities that evolve over time, and then mines weak signal association paths based on the graph structure to provide data support for risk discovery. On this basis, a pre-trained language model is used to convert the path into a risk hypothesis to be verified with causal logic, thereby realizing the automatic conversion from graph data to natural language risk expression. Through technical means such as multi-hop traversal, path importance scoring, and causal modeling, potential high-risk signals are extracted from weak and non-explicit data, and further drive the subsequent data collection and verification process to form a complete closed loop. This weak signal-driven risk hypothesis generation and verification method has the ability to actively discover and verify potential risks, breaking through the limitations of existing systems that can only passively crawl and statically analyze data.

[0016] 3. The present invention constructs a multi-source crawling strategy. Driven by the intention of data collection, it assigns different access parameters to each type of website and determines the crawling priority. It also supports simulating real user behaviors such as mouse tracks and page loading order, thereby circumventing the anti-crawling mechanism. At the same time, through the structural analysis and content recognition technology of the crawling process, it can quickly lock the target data page location and realize data extraction by loading the complete page content, avoiding information loss due to asynchronous loading or deep structure. In addition, it also has a built-in crawling progress monitoring and abnormal switching mechanism, which can enable backup strategies when specific websites are blocked or connections fail, thereby improving the stability and continuity of the collection task. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 A flowchart of a vertical website deep crawling method based on intent recognition provided by the present invention; Figure 2 A schematic diagram of the reverse evidence acquisition process according to an embodiment of the present invention; Figure 3 A schematic diagram of the deep crawling structure provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0018] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0019] Example 1: See also Figure 1 , the present invention provides a vertical website deep crawling method based on intent recognition, the technical solution is as follows: Crawl data from multiple heterogeneous sources and build a dynamic knowledge graph. The dynamic knowledge graph includes entity nodes and relationship edges with timestamps. The relationship edges represent the relationships between entity nodes. After crawling, the data is preprocessed and then built into the dynamic knowledge graph. The data preprocessing includes data standardization, normalization, and data conversion. Identifying a weak signal association path formed by a plurality of the relationship edge connections in the dynamic knowledge graph, and converting the weak signal association path into a risk hypothesis to be verified; The specific process of obtaining the weak signal correlation path includes: Executing a multi-hop path traversal algorithm on the dynamic knowledge graph to identify entity node connection paths with a length of a preset value; the preset value is obtained through experience; the preset value is obtained through experience; the multi-hop path traversal algorithm performs a depth-first search from a starting node, records the number of hops, and selects paths that meet the conditions; Performing a timing consistency analysis on the connection paths to select path combinations whose timestamps have logical relevance; Calculate the weight of each relationship edge in the path, the weight value is comprehensively evaluated based on the relationship strength, time correlation and rarity; A path importance scoring algorithm is applied to select weak signal association paths whose scores exceed a preset threshold. The path importance scoring algorithm comprehensively considers path length, node importance, and relationship edge weight; the preset threshold is obtained through experience.

[0020] The specific calculation process of the weight value includes: Obtain the relationship strength parameter of the relationship edge. The relationship strength parameter is quantified based on the importance of the relationship type. Investment holding relationships, guarantee relationships, and related-party transaction relationships correspond to different strength coefficients. Calculate the time correlation parameter of the relationship edge. The time correlation parameter is calculated by decaying according to the time distance between the relationship edge timestamp and the current analysis time point. The closer the time distance, the higher the weight. Evaluate the rarity parameter of the relationship edge. The rarity parameter is determined by counting the frequency of occurrence of this type of relationship in the entire knowledge graph. The lower the frequency, the higher the rarity. The relationship strength parameter, time correlation parameter and rarity parameter are weighted and combined to generate the comprehensive weight value of the relationship edge.

[0021] The specific process of obtaining the path importance score includes: Calculate the path length score, determine the length score based on the number of hops contained in the path, set the optimal path length range, and give a score penalty to paths that exceed or fall short of the range; Obtain node importance scores, and conduct a comprehensive evaluation based on the node's connectivity in the knowledge graph, node type weight, and node entity influence; Summarize the weight values ​​of all relationship edges in the path and calculate the total weight score of the path; The path length score, node importance score and total weight score are weighted and summed according to the preset ratio to obtain the final score of path importance.

[0022] It also includes the incremental update process of the dynamic knowledge graph: Monitor changes in preset data sources and trigger incremental update mechanisms when new data is detected; Perform entity recognition and relationship extraction on the newly added data to generate candidate entity nodes and relationship edges; Perform graph consistency checks to detect conflicts between new content and existing graph structures, including entity duplication checks, relationship logic checks, and temporal rationality checks; Adopting an incremental graph fusion algorithm, the newly verified entities and relationships are integrated into the existing dynamic knowledge graph, while updating the importance scores of the relevant nodes and the relationship edge weights. Establish a graph version management mechanism to record the timestamp, data source and change content of each update, and support historical backtracking and version comparison of graph status.

[0023] In this embodiment, by crawling data from multiple heterogeneous sources and constructing a dynamic knowledge graph containing a time dimension, the timeliness of information association is improved. Through multi-hop path traversal and temporal consistency analysis, potential weak signal association paths are identified from the graph structure, and a comprehensive evaluation is performed in combination with relationship strength, time correlation and rarity to achieve efficient mining of hidden risk clues. A path importance scoring algorithm is introduced to ensure that the identified weak signal paths have high logical credibility and analytical value; a dynamic incremental update mechanism ensures that new data can be integrated in a timely manner, and the stability and rationality of the graph structure are guaranteed through consistency verification. The version management mechanism supports the backtracking and analysis of the knowledge graph evolution process, enhancing the interpretability and maintainability of the system. This provides strong data support and reasoning basis for subsequent risk warning and verification.

[0024] The specific process of obtaining the risk hypothesis to be verified includes: Input the weak signal association path into a pre-trained financial domain language model to extract key entities and relationship descriptions in the path; The path information is converted into a structured causal relationship description through a templated reasoning engine, wherein the causal relationship description includes premise conditions, reasoning process and expected results; Apply risk narrative generation algorithms to convert structured descriptions into risk hypothesis statements in natural language; A unique identifier is assigned to each of the risk assumptions to be verified, and its source path and generation timestamp are recorded.

[0025] In this embodiment, by introducing a pre-trained financial domain language model, deep semantic analysis of key entities and relationships in weak signal association paths can be performed to accurately extract the core elements of potential risk factors. Combined with a templated reasoning engine, unstructured path information is converted into a causal relationship description with a logical chain, enhancing the interpretability of risk assumptions and the clarity of reasoning. The risk narrative generation algorithm automatically generates risk statements in natural language, improving the readability and dissemination efficiency of the output results. By assigning a unique identifier to each risk assumption and recording the source information, traceable management of the entire risk reasoning process is achieved, which helps to improve the systematicity and reliability of risk identification and response.

[0026] Generate a data collection intention, where the types of data collection intentions include: generating a refutation query to find counter-evidence based on a risk hypothesis to be verified, and using the refutation query as a first data collection intention; receiving an input query instruction, and performing semantic analysis on the query instruction to generate a second data collection intention of the query type; The specific process of generating the data collection intention includes: In response to the first data collection intention, based on the risk hypothesis to be verified, a refutation query generation module automatically constructs three types of refutation queries: disproof queries, strength questioning queries, and alternative explanation queries. The disproof queries are used to find evidence that can negate the premise of the risk hypothesis. The strength questioning queries are used to verify the logical strength of the causal relationship in the hypothesis. The alternative explanation queries are used to discover explanatory factors that weaken the risk hypothesis. For the second data collection intention, semantically analyze the query instruction using a preset intent classification model to classify the query instruction into one of a factual query, a list query, or a comparative query; Based on the second data collection intention, a search is first performed in the dynamic knowledge graph. If data that meets the query instruction exists in the dynamic knowledge graph, the query instruction is directly responded to. If it does not exist or does not exist completely, the missing data part is used as the target of the subsequent crawling task.

[0027] In this embodiment, by generating multiple types of data collection intentions and automatically constructing query requests with reverse verification significance based on the risk hypothesis to be verified, the risk identification system's logical rigor and falsifiability are enhanced. Refutation queries help examine the rationality of hypotheses from multiple perspectives, reducing the risk of misjudgment. Furthermore, the intent classification model performs semantic parsing of user query instructions, enabling accurate identification and response to different query types. Combined with the dynamic knowledge graph's prioritized retrieval and automatic identification of missing data, the system can dynamically adjust data crawling strategies to improve data response efficiency and coverage completeness.

[0028] Based on the type of the data collection intention, a multi-source crawling strategy is generated, wherein the multi-source crawling strategy includes multiple vertical website sources to be crawled and their crawling priorities; The specific acquisition process of the multi-source crawling strategy includes: According to the type of the data collection intention, multiple vertical website sources are selected from a preset information source library, wherein the information source library includes general search engines, academic databases, legal information databases, and financial regulatory agency websites; Based on the urgency of data collection and data quality requirements, different crawling priorities are set for the selected multiple vertical website sources, including high priority, medium priority and low priority; Develop exclusive crawling parameter configurations for each vertical website source, including access frequency limit, number of concurrent threads, timeout period and number of retries, including user agent rotation, IP proxy pool and access behavior simulation.

[0029] In this embodiment, by generating a multi-source crawling strategy based on data collection intent type, we achieve targeted crawling of different vertical website sources, improving the relevance of data acquisition. By combining the urgency of data collection intent with data quality requirements, we dynamically assign crawling priorities and achieve rational optimization of resource scheduling. This improves the efficiency, reliability, and adaptability of heterogeneous data collection, providing richer and higher-quality data support for subsequent risk verification.

[0030] According to the multi-source crawling strategy, perform deep crawling tasks to obtain target data; The specific process of obtaining the target data includes: Perform deep crawling tasks on each vertical website source in turn according to the crawling priority order; For each vertical website source, we first conduct a website structure analysis to identify the page hierarchy and URL pattern where the target data is located; Adopt a crawling strategy that simulates real user behavior, including randomizing access intervals, simulating mouse click trajectories, and constructing request header information that conforms to website characteristics; For dynamically loaded content encountered during the crawling process, the JavaScript rendering engine is used to fully load the page before extracting the data; Establish a crawling progress monitoring mechanism to track the data acquisition status of each website source in real time, and automatically switch to alternative crawling strategies under abnormal circumstances.

[0031] In this embodiment, by executing priority-based deep crawling tasks, target data is efficiently and orderly acquired from multiple vertical website sources, ensuring that key data is captured first. Website structure analysis and URL pattern recognition are employed to improve the accuracy of target positioning. A crawling strategy that simulates real user behavior is introduced to increase the success rate and stability of data collection. For dynamically loaded content, JavaScript rendering technology is used to extract full page data, enhancing adaptability to modern web page structures.

[0032] According to the type of the data collection intention, the target data is processed, the risk hypothesis to be verified is verified, and the existence and strength of the counter-evidence are determined.

[0033] The specific process of judging whether there is counter-evidence and its strength includes: Figure 2 : Based on the target data acquired by the first data collection intention, identifying evidence fragments related to the refutation query through a text similarity algorithm and keyword matching technology; Classify and annotate the identified evidence fragments, marking them as strong counter-evidence, weak counter-evidence, or irrelevant evidence; Calculate the strength of the counter-evidence based on a quantitative assessment of the credibility of the source, the relevance of the content, and the timeliness of the evidence; Based on the existence and strength of the counter-evidence, the confidence of the risk hypothesis to be verified is scored. The confidence score uses a numerical range of 0 to 1. When the strength of strong counter-evidence exceeds a preset threshold, the corresponding risk hypothesis is marked as falsified.

[0034] The specific process of obtaining the evidence strength value includes: Evaluate the credibility of the evidence source, grading it based on the authority of the data source, with regulatory agency websites and exchange announcements at the highest level, professional financial websites at a medium level, and social media and forums at a low level; Calculate the content relevance score. Use the text similarity algorithm to calculate the semantic match between the evidence fragment and the refutation query. The higher the match, the higher the relevance score. Determine the timeliness score based on the correspondence between the time when the evidence was generated and the time window of the risk hypothesis. The stronger the timeliness score, the higher the timeliness score. The credibility of the evidence source, the content relevance score, and the timeliness score are normalized and weighted to generate a comprehensive strength value of the evidence.

[0035] The specific process of obtaining the confidence score of the risk hypothesis to be verified includes: Count the distribution of strong counter-evidence, weak counter-evidence, and irrelevant evidence, and calculate the quantitative weight of each type of evidence; The strength values ​​of each type of evidence were weighted averaged to obtain a comprehensive score of the strength of supporting evidence and the strength of opposing evidence; A confidence scoring model was established, with the strength of supporting evidence as a positive factor and the strength of opposing evidence as a negative factor, and a comprehensive calculation was performed based on the weight of the amount of evidence. The calculation results are mapped to a numerical range of 0 to 1, where 0 represents complete unreliability and 1 represents complete credibility. The falsification threshold is set to 0.3. When the threshold is lower than this, the risk hypothesis is marked as falsified.

[0036] When the strength of the negative evidence is in an uncertain range, multiple rounds of iterative verification are performed: Set an uncertainty interval threshold for the confidence score. When the confidence score of the risk hypothesis is within the preset uncertainty interval, initiate a multi-round iterative verification mechanism. Based on the weak links in the first round of verification results, generate targeted supplementary data collection intentions, including evidence deep mining queries, cross-validation queries, and time series expansion queries; Expand the time range and source range of data collection, and conduct targeted in-depth crawling of missing key evidence; Combining the accumulated evidence obtained from multiple rounds of validation, the Bayesian updating method is used to dynamically adjust the confidence score of the risk hypothesis; Set a maximum number of iteration rounds, terminate the iteration when the maximum number of rounds is reached or the confidence score converges stably, and output the final verification conclusion.

[0037] In this embodiment, by combining the text similarity algorithm with the keyword matching technology, the evidence fragments related to the refutation query are efficiently identified from the collected target data, thereby realizing the automated verification of the risk hypothesis. Through the classification and labeling of the evidence fragments and the quantitative evaluation of their strength, the system can comprehensively consider the credibility, relevance and timeliness of the source of the evidence, and improve the scientificity and accuracy of the judgment. Scoring the confidence of the risk hypothesis based on the reverse evidence strength value helps to dynamically update the validity judgment of the hypothesis and avoid subjective human intervention. Further setting the uncertainty interval and introducing multiple rounds of iterative verification and Bayesian update mechanism, the verification process is continuously optimized when key evidence is insufficient, effectively avoiding misjudgment and missed judgment, and enhancing the system's adaptability and reasoning depth to complex risk scenarios. Intelligent, data-driven and evidence-oriented risk identification is achieved, effectively improving the accuracy and response efficiency of risk management.

[0038] The present invention also includes a cross-domain adaptation mechanism to extend the method to different vertical fields: Build a domain ontology library to predefine entity types, relationship types, and professional terms in different vertical fields, including finance, healthcare, law, and manufacturing; Establish a domain feature model library and configure exclusive weak signal identification patterns, risk hypothesis generation templates, and evidence assessment standards for each vertical field; Design a domain-adaptive learning mechanism to automatically adjust the path traversal algorithm parameters, weight calculation formula, and confidence scoring model by analyzing the data characteristics and business rules of the target domain; Achieve cross-domain knowledge transfer, and transfer weak signal patterns and verification strategies that have been proven effective in one field to other similar fields; Establish a domain effect evaluation system to continuously optimize the performance of cross-domain adaptation algorithms by comparing the risk identification accuracy and verification efficiency in different fields.

[0039] The present invention improves the applicability and promotion value of risk identification methods in various vertical industries through a cross-domain adaptation mechanism. By constructing a domain ontology library and a feature model library, the entity relationships and weak signal features unique to each field can be identified, thereby enhancing the professionalism and pertinence of risk modeling. The adaptive learning mechanism dynamically adjusts the core algorithm parameters according to the data characteristics of the target field, ensuring the adaptation effect and stability of the model in different environments. At the same time, the knowledge transfer strategy can achieve efficient reuse of high-value models between similar fields, reducing the cost of system training and deployment. The supporting evaluation system provides a quantitative basis for cross-domain performance optimization, ensuring that the system continues to maintain high accuracy and high verification efficiency in multiple fields.

[0040] This invention provides a vertical website deep crawling method based on intent recognition, constructs a complete closed-loop process from weak signal recognition to risk verification, and improves the intelligence level and timeliness of risk identification and response. Figure 3 By integrating data from multiple heterogeneous sources and constructing a dynamic knowledge graph, the system captures entity relationships with temporal attributes, enhancing the temporal logic and dynamic expressiveness of the information structure. Furthermore, it uses multi-hop path traversal and path importance scoring models to identify potential weak signal association paths, effectively uncovering potential risk clues hidden in complex data networks. Combining financial domain language models with inference engines, the system converts path information into structured causal relationships with clear logical chains and generates risk hypotheses in natural language, making risk representation more intuitive, readable, and traceable. By constructing refutational data collection intentions and designing three targeted query methods, the system significantly enhances the falsifiability and logical integrity of risk hypotheses. Furthermore, through intent classification and semantic parsing, the system achieves precise understanding and automatic response to user queries, improving the efficiency of human-computer interaction. At the data acquisition level, the system dynamically generates a multi-source crawling strategy based on the collection intention, setting differentiated priorities and anti-crawl parameters for different vertical website sources, improving the adaptability and stability of the crawling process. During deep crawling, structural analysis, behavioral simulation, and dynamic loading processing techniques are combined to ensure a high success rate in acquiring target data. Ultimately, the system achieves automatic verification and confidence scoring of risk hypotheses through similarity analysis and evidence strength quantification, promoting risk judgment from experience-based to data-driven, and has strong practical value and expansion potential.

[0041] Example 2: This embodiment applies the above-mentioned vertical website deep crawling method based on intent recognition to the financial risk analysis scenario. The specific technical solution is as follows: We crawl financial data from multiple heterogeneous sources, including listed company announcement websites, financial news portals, regulatory agencies' websites, and industry databases, to construct a dynamic knowledge graph for the financial sector. This dynamic knowledge graph includes enterprise entity nodes, person entity nodes, event entity nodes, and time-stamped relationship edges. The relationship edges represent the relationships between entity nodes, such as investment relationships, guarantee relationships, related-party transaction relationships, and equity change relationships.

[0042] Specifically, it crawls the regular reports of listed companies from the official websites of the Shenzhen Stock Exchange and the Shanghai Stock Exchange to extract structured data such as basic company information, financial data, and related party information; crawls real-time news information from financial portals to identify event information such as corporate mergers and acquisitions, executive changes, and performance warnings; obtains regulatory updates such as penalty announcements and regulatory letters from regulatory websites; and supplements industry analysis reports and macroeconomic data from professional databases such as Wind and Bloomberg.

[0043] During the knowledge graph construction process, for example, the system identified Company G as having a purchasing relationship with supplier Company A (timestamp: 2023-03-15), an investment and holding relationship with subsidiary Company B (timestamp: 2023-01-20), and a part-time position with Company C (timestamp: 2023-02-10). Each relationship edge records specific timestamp information, forming a corporate relationship network with time series characteristics.

[0044] A multi-hop path traversal algorithm is executed on the constructed financial dynamic knowledge graph to identify entity node connection paths with a length of 8 hops. Taking a potential financial fraud risk as an example, the system identifies the following associated paths: Path example: Listed Company A → Related Party Transaction → Supplier B → Equity Investment → Investment Institution C → Actual Control → Related Party D → Fund Flow → Listed Company A; A temporal consistency analysis of the connection path revealed that the timestamps of the relationship edges in the path showed obvious logical correlation: the related-party transaction occurred in January 2023, the equity investment was completed in February 2023, the actual control relationship was established in March 2023, and the capital transactions were concentrated in April 2023, forming a complete time chain.

[0045] The weight value of each relationship edge in the path is calculated, and a comprehensive evaluation is performed based on the relationship strength, time correlation and rarity to obtain the comprehensive weight of the path.

[0046] The path importance scoring algorithm is applied to comprehensively consider the path length, node importance (listed company node importance, related party node importance) and relationship edge weight (the importance score of the weak signal correlation path is calculated. If it exceeds the preset threshold, it will be identified as an important weak signal correlation path by the system.

[0047] The weak signal association path is input into a pre-trained financial language model (based on the BERT architecture, and domain-adaptive training is performed on financial text corpus) to extract the key entities in the path (listed company A, supplier B, investment institution C, related party D) and relationship descriptions (related transactions, equity investment, actual control, and capital transactions).

[0048] Through the templated reasoning engine, the path information is converted into a structured causal relationship description: Prerequisite: Listed company A has a large-scale related-party transaction with supplier B; Reasoning: Supplier B receives investment from investment institution C, which is actually controlled by related party D, which has financial transactions with listed company A. Expected outcome: The risk of financial fraud through profit manipulation through related parties.

[0049] Using a risk narrative generation algorithm, we convert structured descriptions into natural language-based risk hypotheses to be verified: "Listed Company A may use Supplier B through a complex network of related parties to commit financial fraud by inflating revenue. Specifically, they provide financial support to Supplier B through Investment Institution C controlled by Related Party D, which then purchases from Listed Company A, creating a closed-loop financial operation." Assign a unique identifier to the risk hypothesis to be verified, and record its source path and generation timestamp.

[0050] For the first data collection intention, based on the above-mentioned risk hypothesis to be verified, the refutation query generation module automatically constructs three types of refutation queries: Conduct counter-evidence inquiries to find evidence that can negate the risk assumptions, such as: "Does the transaction between listed company A and supplier B have a real business background?"; "Does supplier B have independent production and operation capabilities?"; "Is the source of funds of investment institution C clear and compliant?"

[0051] Strength questioning query verifies the logical strength of the causal relationship in the hypothesis. Examples include: "What is the degree of control that related party D has over investment institution C?"; "Is there a clear correspondence between the timing of fund transactions and related-party transactions?"; "How common is a similar operating model in the industry?"

[0052] Alternative explanation queries reveal other explanatory factors that may weaken the risk hypothesis. Examples include: "Is there a legitimate reason for Listed Company A to need this type of supplier for business expansion?"; "Does Investment Institution C have other investment motives?"; "Do changes in industry policies affect the rationality of related transactions?"

[0053] Regarding the second data collection intention, assuming that the system receives the user's query instruction "Query the related-party transactions of listed company A in the past three years", the query instruction is semantically analyzed through the preset intent classification model and identified as a factual query.

[0054] First, a search was performed in the dynamic knowledge graph, and it was found that only partial related transaction data for 2023 existed in the graph, and complete data for 2021 and 2022 was missing. Therefore, "Detailed data of related transactions of listed company A from 2021 to 2022" was set as the target of the subsequent crawling task.

[0055] Based on the type of data collection intent, select multiple vertical website sources from the preset financial information source library: 1) General search engine for news retrieval and public opinion analysis; 2) Academic database, used for industry research reports and academic papers; 3) Legal information database, used for legal disputes and litigation information; 4) Financial regulatory agency websites, for regulatory penalties and announcement information; 5) Professional financial website for market analysis and investor discussions.

[0056] Set crawling priorities based on the urgency of data collection and data quality requirements: High priority, high authority, and accurate data; Medium priority, strong professionalism, and good analytical depth; Low priority, with wide coverage but variable data quality.

[0057] In order of crawling priority, deep crawling tasks are performed on each vertical website source in turn. The specific examples are as follows: The crawling process for the Shanghai Stock Exchange's official website: First, conduct a website structure analysis to identify the page level where the listed company's announcements are located. The URL mode is a dynamic parameter splicing form.

[0058] Adopt a crawling strategy that simulates real user behavior: Randomize the access interval and randomly generate the access interval between 2-5 seconds; simulate the mouse click trajectory and simulate the user's operation process such as clicking on the company code and date filtering; construct the request header information that conforms to the website characteristics and set the Referer, User-Agent and other fields.

[0059] For dynamically loaded content encountered during the crawling process, wait for the page to be fully loaded before extracting the announcement PDF link and basic information.

[0060] A crawling progress monitoring mechanism was established to track the data acquisition status of each website source in real time. For example, 126 Shanghai Stock Exchange announcements, 483 news items related to Eastmoney, and 15 CSRC penalty announcements were successfully obtained. When a website connection timed out or returned a 403 error, the crawling strategy was automatically switched to a backup strategy, retrying using a different IP proxy and access parameters.

[0061] Based on the target data obtained by the first data collection intention, the text similarity algorithm (cosine similarity) and keyword matching technology are used to identify evidence fragments related to the refutation query: Through counter-information inquiry and identification of relevant evidence, it was found from the industrial and commercial registration information of Supplier B that its registered capital was only 500,000 yuan, it had no self-owned production equipment, and its main business was trade agency, which was inconsistent with the "core supplier" status claimed by Listed Company A.

[0062] The strength of the questioning query was used to identify relevant evidence. From the equity structure of investment institution C, it was found that the shareholding ratio of related party D was only 15%, and there were 3 other shareholders. The degree of control was questionable, and the similarity score of the relevant evidence fragments was obtained.

[0063] Relevant evidence identification was conducted through alternative explanation queries. It was found from industry policy documents that policy adjustments did occur in relevant industries in early 2023, which may affect the rationality of the supply chain layout, and the relevant evidence similarity score was obtained.

[0064] Classify and label the identified evidence fragments: Strong counter-evidence, including industrial and commercial information indicating that Supplier B has insufficient production capacity; Weak reverse evidence, the equity information of the dispersed control rights of investment institution C; Irrelevant evidence, information on industry policy adjustments (due to low correlation with core assumptions) Calculate the strength of the counter-evidence and conduct a quantitative assessment based on the credibility of the evidence source, content relevance, and timeliness. Score the confidence level of the risk hypothesis to be verified based on the existence and strength of the counter-evidence.

[0065] Analyze the relationship between the final confidence score and the preset falsification threshold. If it is lower than the preset threshold, the risk hypothesis will be marked as "falsified", indicating that through deep data crawling and evidence analysis, the financial fraud risk hypothesis lacks sufficient supporting evidence.

[0066] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A vertical website deep crawling method based on intent recognition, characterized in that: include: Crawl data from multiple heterogeneous sources and build a dynamic knowledge graph. The dynamic knowledge graph includes entity nodes and relationship edges with timestamps. The relationship edges represent the relationships between entity nodes. Identifying a weak signal association path formed by a plurality of the relationship edge connections in the dynamic knowledge graph, and converting the weak signal association path into a risk hypothesis to be verified; Generate a data collection intention, where the types of data collection intentions include: generating a refutation query to find counter-evidence based on a risk hypothesis to be verified, and using the refutation query as a first data collection intention; receiving an input query instruction, and performing semantic analysis on the query instruction to generate a second data collection intention of the query type; Based on the type of data collection intention, a multi-source crawling strategy is generated, wherein the multi-source crawling strategy includes multiple vertical website sources to be crawled and their crawling priorities; according to the multi-source crawling strategy, a deep crawling task is executed to obtain target data; According to the type of the data collection intention, the target data is processed, the risk hypothesis to be verified is verified, and the existence and strength of the counter-evidence are determined.

2. A vertical website deep crawling method based on intent recognition according to claim 1, characterized in that: The specific process of obtaining the weak signal correlation path includes: Executing a multi-hop path traversal algorithm on the dynamic knowledge graph to identify entity node connection paths whose lengths meet a preset value; Performing a timing consistency analysis on the connection paths to select path combinations whose timestamps have logical relevance; Calculate the weight of each relationship edge in the path, the weight value is comprehensively evaluated based on the relationship strength, time correlation and rarity; Based on a path importance scoring algorithm that comprehensively considers path length, node importance, and relationship edge weight, weak signal association paths with scores exceeding a preset threshold are selected.

3. The vertical website deep crawling method based on intent recognition according to claim 2 is characterized by: The specific process of obtaining the risk hypothesis to be verified includes: The weak signal association path is input into a pre-trained financial domain language model to extract key entities and relationship descriptions in the path; the path information is converted into a structured causal relationship description, which includes preconditions, reasoning process and expected results; and the causal relationship description is converted into a risk hypothesis to be verified in natural language form.

4. The method for deep crawling vertical websites based on intent recognition according to claim 1 is characterized in that: The specific process of generating the data collection intention includes: Based on the risk hypothesis to be verified, three types of refutation queries are automatically constructed through refutation queries: disproof query, strength query and alternative explanation query. The disproof query is used to find evidence that can negate the premise of the risk hypothesis. The strength query is used to verify the logical strength of the causal relationship in the hypothesis. The alternative explanation query is used to discover explanatory factors that weaken the risk hypothesis. Performing semantic analysis on the query instruction using a preset intent classification model to classify the query instruction into one of a factual query, a list query, and a comparative query; Based on the second data collection intention, a search is first performed in the dynamic knowledge graph. If there is data that meets the query instruction in the dynamic knowledge graph, the query instruction is directly responded to. If the query is incomplete, the missing data part is used as the target of the subsequent crawling task.

5. The method for deep crawling vertical websites based on intent recognition according to claim 4 is characterized in that: According to the type of the data collection intention, multiple vertical website sources are selected from a preset information source library, wherein the information source library includes general search engines, academic databases, legal information databases, and financial regulatory agency websites; Based on the urgency of data collection and data quality requirements, different crawling priorities are set for the selected multiple vertical website sources, including high priority, medium priority and low priority; Develop exclusive crawling parameter configurations for each vertical website source, including access frequency limit, number of concurrent threads, timeout period, and number of retries.

6. The method for deep crawling vertical websites based on intent recognition according to claim 5, characterized in that: The specific process of obtaining the target data includes: Perform deep crawling tasks on each vertical website source in turn according to the crawling priority order; For each vertical website source, we first conduct a website structure analysis to identify the page hierarchy and URL pattern where the target data is located; Adopt a crawling strategy that simulates real user behavior, including randomizing access intervals, simulating mouse click trajectories, and constructing request header information that conforms to website characteristics; For dynamically loaded content encountered during the crawling process, the page is fully loaded before data extraction; Establish a crawling progress monitoring mechanism to track the data acquisition status of each website source in real time, and automatically switch to alternative crawling strategies under abnormal circumstances.

7. The method for deep crawling vertical websites based on intent recognition according to claim 1, characterized in that: The specific process of determining whether there is counter-evidence and its strength includes: Based on the target data obtained by the first data collection intention, the evidence fragments related to the refutation query are identified through text similarity algorithm and keyword matching technology; the evidence fragments are classified and labeled as strong reverse evidence, weak reverse evidence and irrelevant evidence, and the strength value of the evidence is calculated; according to the labeling and strength value of the reverse evidence, the confidence of the risk hypothesis to be verified is scored; when the strength of the strong reverse evidence exceeds a preset threshold, the corresponding risk hypothesis is marked as falsified.

Citation Information

Patent Citations

  • Knowledge graph-based intelligent quantitative research platform system and method thereof

    CN119693155A

  • Water conservancy intelligent question and answer interaction platform and method driven by knowledge graph

    CN120196707A

  • Multi-source knowledge processing and querying method and device, equipment and medium

    CN120197681A

  • Privacy enhanced intelligent search method and system based on multi-round iteration

    CN120596652A

  • Risk profiling and rating of extended relationships using ontological databases

    US20210019674A1

Cited By

  • Agricultural operation multi-source image processing method based on artificial intelligence

    CN121686254A

  • An agricultural operation multi-source image processing method based on artificial intelligence

    CN121686254B

  • Intention recognition method based on multi-source evidence fusion reasoning

    CN121881279A