A data collection and report generation method, medium and device

By generating universal access policies and using pre-trained large models to parse text data, the limitations of network site defense systems on data collection are solved, achieving efficient and reliable data collection and report generation.

CN121117355BActive Publication Date: 2026-06-26BEIJING SHOUFA INTELLIGENT TECHNOLOGY CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING SHOUFA INTELLIGENT TECHNOLOGY CO LTD
Filing Date
2025-09-04
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

Existing data collection and report generation methods are easily identified and restricted by network site defense systems, resulting in wasted collection node resources and task interruptions, low efficiency, and failure to effectively combine access policies for node attribute verification.

Method used

By analyzing access status data based on reference collection nodes, a universal access strategy applicable to all target websites is generated. Suitable collection nodes are selected, and a pre-trained large model is used to parse text data. Consistency verification and feature profile generation are performed in conjunction with information update time. Finally, similarity matching is performed to generate a data report.

Benefits of technology

It improves the efficiency and reliability of data collection, ensures the compliance of collection nodes and the continuity of data, enhances parsing efficiency and accuracy, and enables the efficient generation of structured reports.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121117355B_ABST
    Figure CN121117355B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of data processing, and especially relates to a data collection and report generation method, medium and equipment, a general access strategy is obtained by analyzing a reference collection node, target collection nodes are screened based on the general access strategy and text data is collected, irregular nodes are removed from the source, and corresponding measures are taken in time when irregular behaviors of the nodes are found, so that the compliance of the collection nodes and the continuity of the collected data are guaranteed; a pre-trained large model is used to analyze the text data to generate an attribute information set, so that the analysis efficiency and accuracy are improved; according to target network site information update time, the multi-source attribute information of a data subject is verified, multi-site data conflicts are accurately identified, and the subject information consistency and the integrity of the feature portrait are ensured; the feature description data of the target identification code and the feature portrait are subjected to similarity matching, multi-source network data is converted into a structured report available for a third-party system, and reliable data collection and report generation are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a data acquisition and report generation method, medium, and device. Background Technology

[0002] With the explosive growth of internet information, there is an increasing demand for automated data collection from multiple websites and the generation of structured reports in scenarios such as business intelligence analysis and personal information aggregation.

[0003] Existing data collection and report generation methods typically use a fixed or small number of collection nodes, employing a uniform, preset access frequency and pattern to access target sites to obtain text data, parse the text to extract structured information, and finally summarize the information into a report. However, website defense systems can easily detect such automated behavior, leading to the blocking of collection node IP addresses or strict restrictions on access. Furthermore, the aforementioned methods, when screening collection nodes, simply determine whether a node can log in to the site without combining access policies to verify the compliance of node attributes. As a result, some nodes, although able to log in, are still easily restricted due to attribute conflicts with site rules, causing waste of node resources and leading to interruptions, inefficiencies, or even complete failures in data collection tasks.

[0004] Therefore, how to achieve reliable data collection and report generation has become an urgent problem to be solved. Summary of the Invention

[0005] To address the aforementioned technical problems, the present invention provides a data acquisition and report generation method, which includes the following steps:

[0006] S1. Based on the access status data and access results obtained when each reference acquisition node accesses each target network site, strategy analysis is performed to obtain a general access strategy applicable to all target network sites.

[0007] S2 selects several target collection nodes based on a general access strategy and controls each target collection node to collect text data from each target network site.

[0008] S3 uses a pre-trained large model to parse all text data corresponding to each target website and obtain the attribute information set corresponding to each target website. The attribute information set includes the attribute information of several data subjects.

[0009] S4. Based on the information update time of each target network site, perform consistency verification on the attribute information of each data subject from all target network sites, and generate a feature profile corresponding to each data subject based on the verification results.

[0010] S5, perform similarity matching on the feature description data and feature profile of each received target identifier code, and generate a data report corresponding to each target identifier code based on the matching results.

[0011] The present invention also provides a non-transitory computer-readable storage medium storing at least one instruction or at least one program, wherein the at least one instruction or at least one program is loaded and executed by a processor to implement the above-described data acquisition and report generation method.

[0012] The present invention also provides an electronic device, including a processor and the aforementioned non-transitory computer-readable storage medium.

[0013] This invention has at least the following beneficial effects: By performing policy analysis based on access status data of reference collection nodes, the scattered restriction rules of multiple sites are transformed into a unified and executable general access policy, solving the policy adaptation problem of cross-site collection and improving the collection efficiency and feasibility of subsequent target collection nodes; Target collection nodes are filtered based on the general access policy, and the collection of text data by nodes is controlled, eliminating nodes with illegal address types from the source. Through periodic comparison, illegal behavior of target collection nodes is detected in a timely manner, avoiding the imposition of access restrictions on nodes by sites due to prolonged illegal behavior. When the behavior of a node changes from compliant to illegal, collection is immediately suspended to prevent the risk from escalating, while retaining the collected compliant text data, reducing data loss, and ensuring the quality of the collection process. The system ensures compliance and continuity of data collection; it employs a pre-trained large model to parse the text data of each target network site, generating attribute information sets to improve parsing efficiency and accuracy; based on the update time of the target network site information, it verifies the consistency of multi-source attribute information of the data subject, accurately identifies data conflicts between multiple sites, avoids feature confusion caused by direct integration of conflicting data, and ensures the consistency of subject information and the integrity of feature profiles; it performs similarity matching between the feature description data of the target identifier code and the feature profile to generate corresponding data reports. Thus, through a series of compliant collection, accurate parsing, reliable integration, and efficient matching operations, multi-source network data is transformed into structured reports usable by third-party systems, improving data value and achieving reliable data collection and report generation. Attached Figure Description

[0014] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0015] Figure 1This is a flowchart of a data acquisition and report generation method provided in Embodiment 1 of the present invention. Detailed Implementation

[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0017] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It is understood that, where appropriate, the terms used to distinguish similar objects can be interchanged so that the invention can also be implemented in other embodiments besides the illustrated or described embodiments. Furthermore, the terms "including," "having," and any variations are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or devices.

[0018] Example 1

[0019] This first embodiment provides a data acquisition and report generation method, such as... Figure 1 As shown, this data collection and report generation method includes the following steps:

[0020] S1. Based on the access status data and access results obtained when each reference acquisition node accesses each target network site, strategy analysis is performed to obtain a general access strategy applicable to all target network sites.

[0021] The reference collection nodes are pre-deployed test data collection carriers used to test the access rules of target websites. A predetermined number of reference collection nodes are used to initiate access to all target websites, simulating different access behaviors such as frequency, duration, and time periods. This generates corresponding data on access behaviors and results, providing realistic and reusable samples for subsequent policy analysis and avoiding the risk of node restrictions due to directly using target collection nodes for testing. The predetermined number can be set to 30-50 to ensure that the reference collection nodes cover different network identifiers and different access behaviors as samples.

[0022] The reference collection node possesses core attributes completely identical to the subsequent target collection node, including network identification attributes: static IP / dynamic IP pool, geographical region of the IP, and network access type (e.g., residential network, enterprise network); access permission attributes: it has completed authentication of the target network site (e.g., registering an account, obtaining an API access token), and has the same functional access permissions as the target collection node (e.g., data viewing, list download); hardware and software configuration: the same collection tools (e.g., compliant access client, API calling program), the same request sending protocol (e.g., HTTPS), and the same session management mechanism (e.g., cookie expiration settings).

[0023] The target website is a publicly accessible online information carrier that needs to collect text data. It contains non-confidential or specially authorized text data that can be publicly accessed, and has built-in access restriction rules such as request frequency limits, IP blacklists, and session duration limits to prevent malicious collection or excessive access. These include, but are not limited to, commercial information platforms (such as corporate information query websites), public affairs platforms (such as government information disclosure websites), and industry data platforms (such as industry report publishing websites).

[0024] The access result is determined by the response information of the target website. The access result includes successful access and access restrictions imposed. When the access request does not have prompts such as frequent operation or restricted account and can obtain text data normally, the access result is determined to be successful access. When access restrictions are imposed, the access request returns prompts such as permission denied, unauthorized, service unavailable, IP blocked, excessive request frequency, or account frozen, the access result is determined to be access restrictions imposed.

[0025] The access results generated by the reference collection nodes when accessing the target network sites reflect the access restriction rules of the target network sites, providing a basis for extracting common restriction features and ensuring that the final generated general access policy can be adapted to the rules of all target network sites.

[0026] Access status data is a collection of behavioral characteristics recorded by reference collection nodes during access to target websites, including at least one of the following: login duration, access frequency, access time period, network address type, number of requests, and time interval. Each type of status data corresponds to a dimension of access behavior. By analyzing the correlation between access status data and access results, key behavioral characteristics that lead to node restrictions can be accurately identified. For example, an access frequency exceeding 40 times per hour is likely to be restricted, providing data support for the generation of general access strategies.

[0027] The general access strategy is a set of common access rules applicable to all target websites, derived from the access status data of all reference collection nodes. It includes two types of rules: static adaptation rules and dynamic behavior rules. Static adaptation rules are used to filter collection nodes that meet the requirements, such as the network address must be a residential IP and the geographical area of ​​the node must cover the service area of ​​the target website. Dynamic behavior rules are used to constrain the access behavior of collection nodes, such as the frequency of a single node accessing any website not exceeding 35 times per hour and the duration of a single session not exceeding 2 hours.

[0028] The above method simulates different access behaviors such as access frequency, duration, and time period of the target network site by referring to the collection node. It analyzes the correlation between access status data and access results, accurately locates the key behavioral characteristics that cause the node to be restricted, and extracts a set of common access rules applicable to all target network sites. This transforms the scattered restriction rules of multiple sites into a unified and executable operation standard, solves the problem of cross-site collection strategy adaptation, and improves the collection efficiency and feasibility of subsequent target collection nodes.

[0029] In one specific embodiment, S1 includes the following steps:

[0030] S11 records the access status data generated by each reference acquisition node when accessing each target network site within a preset observation period.

[0031] S12, according to the order of data sampling time, construct a status data sequence of all access status data corresponding to each reference acquisition node.

[0032] S13, associate and label each access status data with the access result of the corresponding access operation to obtain the labeled status data sequence.

[0033] S14. Based on the labeled state data sequence generated by all reference acquisition nodes, an access result prediction model is trained. The access result prediction model takes the state data sequence as input and takes successful access or access restriction as output label.

[0034] S15, analyze the importance of each state feature in the access result prediction model, extract several key state features that lead to the prediction result of being subject to access restrictions, and the threshold range of each key state feature.

[0035] S16. Based on several key status features and the threshold range of each key status feature, generate a general access policy applicable to all target network sites. The general access policy includes network address switching rules, access request frequency limits, daily access time periods, and single access duration suggestions.

[0036] The preset observation period is a pre-set time period for collecting access data from reference collection nodes, usually 7-30 days. The specific duration can be adjusted according to the stability of the access rules of the target website. For example, government websites with stable rules can be set to 30 days, while e-commerce websites with variable rules can be set to 7 days, to ensure coverage of the access rules of the target website at different times, such as the differences in restrictions between weekdays and holidays.

[0037] Data sampling time is synchronized with access behavior; for example, access status data is recorded once for each request initiated. Discrete access status data of the same reference acquisition node and the same target network site within a preset observation period are concatenated into a continuous sequence according to the data sampling time to reflect the pattern of access behavior changes over time, providing temporal feature support for subsequent model learning of the correlation between behavior changes and access results.

[0038] For each state data sequence, a corresponding access result is matched to the access behavior, constructing labeled samples of input behavior sequences and output access results for training the access result prediction model. Specifically, Long Short-Term Memory (LSTM) networks or Gated Recurrent Units (GRUs) are selected as the base models, as they are adept at capturing long-term dependencies in time-series data and can identify key change nodes in the access behavior sequences. All labeled state data sequences are divided into a training set for model parameter learning, a validation set for model hyperparameter tuning, and a test set for model performance evaluation in a 7:2:1 ratio.

[0039] During model training, non-numerical features (such as network address patterns) in the state data sequence are converted into numerical codes (e.g., residential network IP is encoded as 1, data center IP is encoded as 0). The encoded state sequence data is then input into the access result prediction model to be trained. Using the access result label as the training objective, the model prediction error is calculated using the cross-entropy loss function. The model parameters are iteratively adjusted using the Adam optimizer; for example, the learning rate is set to 0.001, and the number of iterations is set to 100 rounds. After each round of training, the model accuracy is evaluated using a validation set. Training stops when the validation set accuracy no longer improves after 5 consecutive rounds. The model performance is evaluated using a test set, requiring an accuracy ≥90% and a recall ≥85% to ensure no high-risk behaviors are missed. If these targets are not met, the model parameters are readjusted, such as increasing the number of hidden layers, adjusting the learning rate, and retraining. This results in a high-precision access result prediction model capable of determining whether access behavior sequences will be restricted, providing model support for subsequent extraction of key restriction features.

[0040] Model interpretation tools, such as Shapley Additive Explanations (SHAP) and Local Interpretable Model-Agnostic Explanations (LIME), are used to analyze the output of the access outcome prediction model. The contribution of each state feature to the prediction result, i.e., the feature importance score, is calculated. For each key state feature, a grid search + probability curve analysis method is used to determine the threshold range. This extracts understandable and implementable key state features and threshold ranges, providing a clear technical basis for generating a general access strategy. Taking access frequency as an example, multiple test points are set within the feature value range (e.g., 0-100 times / hour). The model is input to calculate the probability of being restricted for each test point. An access frequency-restriction probability curve is plotted. The feature value range corresponding to a restriction probability ≥90% is determined as the restriction threshold range. For example, if the access frequency exceeds 38 times / hour and the restriction probability is ≥90%, then the threshold range is >38 times / hour.

[0041] Each key state feature and threshold range is mapped to a specific policy rule to ensure that the subsequent screening and access operations of target collection nodes have clear standards. At the same time, a single policy can be adapted to all target network sites, solving the problem of fragmented cross-site policies.

[0042] Specifically, if network address type is a key characteristic, then if the probability of a data center IP being restricted is ≥90%, the rule is translated into a network address switching rule: only use residential network IPs to access target sites; if the same node continuously accesses the same site for more than 2 hours, it must switch to a residential IP in the same area. If access frequency is a key characteristic, then the lowest security frequency threshold among all target sites is taken as the upper limit for access request frequency. For example, if site A has a security threshold of 40 times / hour, site B 35 times / hour, and site C 38 times / hour, this can be translated into an upper limit for access request frequency: the frequency of a single node accessing any target site shall not exceed 35 times / hour. If access time period is a key characteristic, then if the probability of access being restricted between 20:00 and 23:00 is ≥80%, this can be translated into: access during the daily access period is only initiated between 8:00 and 20:00. If login duration is a key characteristic, then the threshold range can be translated into a single access duration suggestion. For example, a threshold range >300 seconds is translated into: the duration of a single login session shall not exceed 300 seconds, and a new login session shall be attempted after the timeout.

[0043] Those skilled in the art will know that the model structures and training methods of LSTM and GRU in the prior art, as well as the application methods of SHAP and LIME, fall within the protection scope of this invention, and will not be elaborated here.

[0044] As described above, by recording data across all dimensions within a preset observation period, the acquired access status data covers different scenarios, avoiding policy misjudgments due to incomplete data. By constructing a status data sequence and associating and labeling it with access results, the model can learn the temporal variation patterns of access behavior. Through key feature and threshold extraction and policy generation, policy rules are quantified and executable, and a single policy can be adapted to all target sites, improving the feasibility and security of subsequent target collection nodes, thereby enhancing the efficiency and reliability of data collection.

[0045] S2 selects several target collection nodes based on a general access strategy and controls each target collection node to collect text data from each target network site.

[0046] In one specific embodiment, S2 includes the following steps:

[0047] S21, perform a conformity judgment on the network address type and general access policy of each candidate collection node, and obtain the conformity judgment result corresponding to each candidate collection node, wherein the conformity judgment result is conforming or not conforming.

[0048] S22, the candidate acquisition nodes that meet the conformity judgment result are determined as the target acquisition nodes.

[0049] S23, During the data collection process, the access status data and general access strategy of each target data collection node are compared according to a preset comparison period to obtain the deviation judgment result of each target data collection node in each preset comparison period, wherein the deviation judgment result is either deviation or no deviation.

[0050] S24, when the deviation judgment result of the target acquisition node changes from not deviating to deviating, the acquisition behavior of the target acquisition node is paused, and the text data collected by the target acquisition node from each target network site when it is not deviating is obtained.

[0051] The candidate data collection node pool is a pre-defined resource library of potential data collection nodes, containing a sufficient number of nodes with diverse attributes. The types of candidate nodes include dedicated data collection terminals, compliant cloud server instances, and authorized user terminal devices. All nodes have completed basic configurations such as IP allocation, network access authentication, and target website authentication. The number of nodes in the candidate pool can be set to 2-3 times the final target number of data collection nodes. This avoids insufficient candidate nodes, which would prevent meeting the needs of parallel data collection at multiple sites, while ensuring that the selected nodes are compatible with common access strategies through attribute diversity. It also provides backup resources for replacing abnormal nodes later.

[0052] Different target websites often set access restrictions for specific network address types. For example, some websites prohibit high-frequency access from data center IPs. Network address type is the attribute classification of the network address used by candidate collection nodes when accessing target websites, including residential network IPs, enterprise network IPs, and data center IPs. It is the core judgment dimension of static adaptation rules in general access policies. For example, a general access policy may limit access to only residential network IPs. By verifying the conformity of the candidate node's network address type with the general access policy, nodes that are easily restricted due to address type violations can be eliminated from the source, ensuring that nodes entering the next stage have the compliance basis at the address level and reducing the risk of restrictions triggered by address type during subsequent collection.

[0053] The preset comparison period is a pre-set time interval used to periodically compare the access behavior of target collection nodes with the general access policy, with a value ranging from 5 to 30 minutes. If the comparison period is too long, such as 1 hour, it may be possible for a node to have been continuously accessing the system beyond the policy for half an hour before being detected, increasing the risk of being restricted; if the period is too short, such as 1 minute, it will increase the system's computing power consumption. It can be set according to the sensitivity of the dynamic rules in the general access policy. For example, if the threshold for the access frequency rule is ≤38 requests per hour, then the preset comparison period can be set to 10-15 minutes to ensure timely detection of behaviors that exceed the frequency limit.

[0054] The deviation assessment result is determined by comparing the real-time access status data of the target data collection node with the dynamic behavioral rules of the general access policy, such as access frequency and access duration, to determine whether the behavior is compliant. No deviation means that the access behavior complies with the policy, while deviation means that any access behavior, such as exceeding the access frequency limit or violating the access time period, exceeds the policy restrictions.

[0055] The target data collection nodes collect text data from various target websites, including but not limited to web page content (such as company information pages and personnel introduction pages), list text (such as the text content of data tables), and descriptive text (such as experience descriptions and service introductions).

[0056] As described above, by judging the compliance of candidate node network address types with policies, nodes with non-compliant address types are eliminated from the source. Through periodic comparison, violations of target data collection nodes are promptly detected, avoiding the imposition of access restrictions on nodes due to prolonged violations. When a node's behavior changes from compliant to non-compliant, data collection is immediately suspended to prevent the risk from escalating, while retaining the collected compliant text data to reduce data loss and ensure that the overall progress of the data collection task is not seriously affected.

[0057] S3 uses a pre-trained large model to parse all text data corresponding to each target website and obtain the attribute information set corresponding to each target website. The attribute information set includes the attribute information of several data subjects.

[0058] Among them, pre-trained large models are deep learning models that are pre-trained on large-scale general corpora and have powerful natural language understanding, information extraction and semantic analysis capabilities. For example, the BERT, GPT and ERNIE series of models under the Transformer architecture can be adapted to specific information extraction tasks through prompt words, adapt to the differences in text format of different target websites, and greatly improve the accuracy and generalization ability of attribute information extraction.

[0059] The preprocessed text data and prompt word templates are input into a pre-trained large-scale model. The pre-trained model, through learned general semantic knowledge and the task objectives guided by the prompt words, identifies the data subject and extracts the attribute values ​​of corresponding attribute fields from the text data, thus transforming unstructured text into structured attribute information. Those skilled in the art will recognize that existing Transformer architecture models such as BERT, GPT, and ERNIE, and their applications, fall within the scope of this invention, and will not be elaborated upon here.

[0060] The data subject is the object entity in text data that carries core information. The specific type varies depending on the content domain of the target website. For example, for commercial information websites, it includes enterprises and legal persons; for academic websites, it includes researchers and papers; and for government websites, it includes government agencies and matters. It has multi-dimensional attribute information that can be extracted, such as the registered capital and establishment time of enterprises, and the work experience of personnel.

[0061] An attribute information set is a structured collection of attribute information formed after parsing all text data of a target website. Each attribute information includes a mapping relationship between the data subject, attribute field, and attribute value. For example, the data subject is company A, the attribute field is registered capital, and the attribute value is 10 million yuan; the data subject is person B, the attribute field is work experience, and the attribute value is product manager of a company from 2020 to 2023.

[0062] It is understandable that, given the differences in text data formats across different target websites, such as the presence of HTML tags and redundant advertising text, preprocessing operations can be used to unify the data format and remove invalid information. This ensures that the text data input to the pre-trained large model is clean and well-organized, avoiding parsing errors caused by format interference. Those skilled in the art will recognize that the preprocessing operations of unifying the data format and removing invalid information in the prior art fall within the scope of protection of this invention, and will not be elaborated upon here.

[0063] In one specific embodiment, S3 includes the following steps:

[0064] S31, Obtain the mapping relationship library between the domain name of the target website and the information extraction prompt template, wherein the information extraction prompt template specifies the attribute fields and format to be extracted from the text data.

[0065] S32, based on the domain name of the target website from which each text data originates, retrieve the information extraction prompt template corresponding to each text data from the mapping relationship database.

[0066] S33, input each text data and the corresponding information extraction prompt template into the pre-trained large model, and obtain the natural language response output by the pre-trained large model.

[0067] S34 parses and performs preset standardization processing on the natural language response, generating a list of attribute information corresponding to each text data.

[0068] S35, based on the list of attribute information corresponding to all text data of each target website, form an attribute information set corresponding to each target website.

[0069] The mapping database stores the correspondence between the content domains of the target website domains and the exclusive information extraction prompt templates for those websites. The content domains include various fields such as education information (e.g., xxx.education.cn), medical information (e.g., xxx.medical.com), enterprise information (e.g., xxx.enterprise.com), and school information (e.g., xxx.school.edu.cn). The implementer can set these fields according to the actual situation, which facilitates direct association of exclusive templates through domain names, avoids manually matching templates for each piece of text data, and solves the problem of low template selection efficiency when parsing text from multiple websites.

[0070] The information extraction prompt template includes the attribute fields to be extracted and the output format requirements. The standardized domain name is input into the retrieval interface of the mapping relation database. The corresponding template is quickly located through the domain name index, ensuring that all text data from the same site uses a unified template. This avoids attribute field confusion caused by inconsistent templates and lays a unified foundation for subsequent attribute information integration. For example, the domain name "xxx.enterprise.com" corresponds to the enterprise information domain, with the corresponding prompt template ID "TPL-ENT-001," and the template content is "Extract enterprise name, registered capital, and establishment date, output format is JSON"; the domain name "xxx.school.edu.cn" corresponds to the school information domain, with the corresponding prompt template ID "TPL-EDU-002," and the template content is "Extract school name, educational level, and establishment date, date format is YYYY-MM-DD."

[0071] Natural language responses to different text data may have variations in expression. Pre-defined standardization processing addresses these format differences in natural language responses by using predefined structured transformation and format unification rules. These rules include structured transformation: extracting the field-value correspondences from the natural language responses into machine-readable structured data; and format unification: correcting attribute value formats and filling in missing format information according to pre-defined standards. For example, the establishment date of May 2020 is standardized to 2020-05-01, and the registered capital of 10 million is standardized to 10 million yuan.

[0072] Standardization ensures that all attribute values ​​are formatted consistently, providing a unified data format for subsequent integration of attribute information lists and cross-site validation.

[0073] An attribute information list is a collection of list items consisting of "attribute field, attribute value, and extraction time" formed after parsing and standardizing a single piece of text data. For example, "Company Name, Company A, 2025-08-29 10:00"; "Registered Capital, 10 million yuan, 2025-08-29 10:00"; "Establishment Time, 2025-05-01, 2025-08-29 10:00".

[0074] The above-mentioned closed loop of mapping library construction, template retrieval, model parsing, standardization processing, and information integration transforms unstructured text data from multiple target websites into structured, high-quality attribute information sets. This achieves a systematic improvement in multi-site adaptation, efficient parsing, and data value-added, realizing full automation and efficient adaptation of multi-site text parsing, significantly reducing manual costs and time consumption, and providing a high-value data foundation for subsequent analysis.

[0075] S4. Based on the information update time of each target network site, perform consistency verification on the attribute information of each data subject from all target network sites, and generate a feature profile corresponding to each data subject based on the verification results.

[0076] In one specific embodiment, S4 includes the following steps:

[0077] S41, for any data subject, group the attribute information of the current data subject from all target network sites according to the attribute items, and obtain the attribute information group corresponding to each attribute item for the current data subject.

[0078] S42, for any attribute item, perform temporal logic conflict detection on each attribute information in the attribute information group corresponding to the current attribute item, and obtain the conflict detection result corresponding to each attribute information, wherein the conflict detection result is either conflict or no conflict.

[0079] S43. If the conflict detection result of the attribute information is no conflict, then the attribute information is determined to be reliable information.

[0080] S44. If the conflict detection result of the attribute information is a conflict, then among all the attribute information that conflicts with the attribute information, the attribute information with the latest information update time of the corresponding target network site is determined as reliable information.

[0081] S45. Generate a feature profile of the current data subject based on all reliable information in the attribute information group corresponding to all attribute items.

[0082] Specifically, for each data subject, the attribute information collected from all target network sites will be grouped by attribute item name to form an attribute information group. This provides a unified analysis unit for subsequent time-series logic conflict detection, avoids verification omissions caused by scattered attribute information, and ensures that verification covers all source data.

[0083] The attribute items include several sequential attribute items and several non-sequential attribute items.

[0084] Temporal attributes are attribute categories within the data subject's attribute information that may change over time, and these changes exhibit a chronological logic. Examples include work experience, registered capital, and job level. Because the attribute values ​​of temporal attributes change over time, information collected from multiple sites is prone to conflicts such as overlapping or reversed time sequences. For instance, site A might display work from 2020-2023, while site B displays work from 2022-2024. Therefore, it is necessary to filter reliable information through temporal logic verification to avoid the feature profile containing contradictory time-related data.

[0085] Non-time-series attributes are attribute categories within the data subject's attribute information that remain stable over a long period, or whose changes are not directly related to time, and whose attribute values ​​are unique within the same time dimension. Only one true attribute value exists at any given time point, and different attribute values ​​collected from multiple sites will inevitably conflict. Conflicts must be identified through uniqueness checks, and reliable values ​​must be filtered based on the information update time to ensure the accuracy of stable attributes in the feature profile and avoid subsequent data association failures due to incorrect core identifiers.

[0086] For each attribute in the attribute information group, differentiated detection rules are formulated based on the time-series / non-time-series attribute item type to determine whether there are logical conflicts in the attribute values.

[0087] Specifically, for time-series attribute items (such as work experience), time range overlap detection and time sequence detection are adopted. By parsing the time range in the attribute value (such as 2020-2023), it is determined whether there is time overlap (such as 2020-2023 overlapping with 2022-2024) or time reversal (such as recording 2023-2025 first and then recording 2020-2022). If either of these conditions exists, it is determined to be a conflict.

[0088] For non-temporally ordered attribute items, a uniqueness check is used. If there are two or more different attribute values ​​in the attribute information group, all different values ​​are judged as conflicting; if all attribute values ​​are the same, they are judged as not conflicting.

[0089] If the attribute information in an attribute information group is found to have no logical conflicts, it is directly identified as reliable information without additional filtering, simplifying the process while ensuring information credibility. If there are conflicts in the attribute information group, it is assumed that the later the information update time, the closer the data is to the current true state. The attribute information with the latest update time among the conflicting information is selected as reliable information, balancing timeliness and credibility and avoiding subjective bias from human judgment.

[0090] By classifying and integrating reliable information from all attribute items, a unique feature identifier can be assigned to each data subject, forming a structured and complete feature profile. This transforms multi-source, scattered data into unified and reliable features, providing a standardized carrier for subsequent matching. Specifically, a hash algorithm (such as SHA-256) is used to encrypt and calculate the identity information of the data subject, generating a unique feature identifier to ensure that the identifier is unique for the same data subject and not repeated for different data subjects. The feature profile is organized according to a preset format (such as JSON or XML), including fields such as feature identifier, non-time-series attribute set, time-series attribute set, profile generation time, and a list of data source sites, ensuring a clear structure.

[0091] The above-mentioned methods accurately identify data conflicts across multiple sites by grouping attribute information by item and detecting temporal logical conflicts, avoiding feature confusion caused by directly integrating conflicting data. By directly identifying reliable information without conflicts and filtering conflicting information by update time, objective and repeatable reliable information screening standards are established, significantly improving the accuracy of reliable information screening. Through the classification and integration of reliable information and the generation of unique feature identifiers, standardized feature profiles are formed, solving the problem of fragmented multi-source data, providing a high-quality data source for subsequent target identifier matching, and improving the reliability of data collection and report generation.

[0092] In one specific embodiment, S45 includes the following steps:

[0093] S451, quantify and vectorize each piece of reliable information corresponding to the current data subject to obtain the attribute feature vector.

[0094] S452, perform preset processing on all attribute feature vectors corresponding to the current data subject to obtain the feature profile of the current data subject. The preset processing includes splicing, dimensionality reduction and feature fusion.

[0095] Numericalization involves converting non-numerical attribute values ​​such as text and categorical data in reliable information into machine-calcifiable numerical forms. Vectorization encoding converts the numericalized attribute values ​​or the original numerical attribute values ​​into fixed-dimensional vectors, avoiding deviations in subsequent feature processing due to differences in attribute value types and scales, and ensuring the comparability and computational efficiency of attribute features.

[0096] The concatenation process involves joining all attribute feature vectors of the current data subject end-to-end according to the order of attribute items to form a high-dimensional merged vector. This avoids integration omissions caused by scattered storage of attribute features and ensures that subsequent dimensionality reduction and fusion operations can cover all attribute information. The dimensionality reduction process uses dimensionality reduction algorithms (such as Principal Component Analysis (PCA) and t-SNE) to compress the dimension of the concatenated high-dimensional merged vector. This reduces the vector dimension while retaining core feature information, balancing the preservation of core information with the reduction of computational cost, and improving the efficiency of subsequent processing. The feature fusion process uses feature fusion algorithms (such as attention mechanisms and weighted averages) to weight the attribute importance of the dimensionality-reduced vector, strengthening core attribute features and weakening secondary attribute features, thereby improving the accuracy of subsequent feature profiling and target identifier matching.

[0097] Those skilled in the art will recognize that any numerical method, vectorized encoding method, splicing method, dimensionality reduction method, and feature fusion method in the prior art falls within the protection scope of this invention, and will not be elaborated further here.

[0098] S5, perform similarity matching on the feature description data and each feature profile of each received target identifier code, and generate a data report corresponding to each target identifier code based on the matching results. Here, the target identifier code is a unique code used to identify a specific data subject, and the feature description data is the attribute description of the data subject corresponding to the target identifier code.

[0099] In one specific embodiment, S5 includes the following steps:

[0100] S51, receive a query request sent by a third-party system, wherein the query request contains several target identifier codes and feature description data corresponding to each target identifier code.

[0101] S52 converts each feature description data into a corresponding query feature vector.

[0102] S53, calculate the similarity between each query feature vector and each feature profile.

[0103] S54, for any target identifier, the feature image with the highest similarity among the feature images that have a similarity exceeding a preset similarity threshold with the current target identifier is determined as the target image corresponding to the current target identifier.

[0104] S55 generates a data report corresponding to each target identifier code based on the target profile corresponding to each target identifier code.

[0105] The third-party system refers to an external system that needs to be linked to the feature profile generated by this technical solution to obtain complete information about the data subject. Examples include enterprise CRM systems (customer relationship management), government information query systems, and academic talent management systems. The query request is a structured request data sent by the third-party system that contains information about the data subject to be matched.

[0106] The feature description data is transformed into a query dimension vector through numerical and vectorized encoding. The similarity between the query feature vector and each feature profile is calculated, quantifying the degree of correlation between the two and providing a mathematical basis for subsequent matching decisions. Those skilled in the art will recognize that any similarity calculation method in the prior art falls within the protection scope of this invention, such as cosine similarity, and will not be elaborated upon here.

[0107] First, filter out feature images with low relevance by setting a preset similarity threshold, and then select the image with the highest similarity from the remaining images as the target image to ensure that the matching results are both reliable and optimal, and avoid mismatches and missed matches.

[0108] Based on the preset report template, key attribute information is extracted from the target profile and organized into a data report that is easy for third-party systems to read and use, realizing the transformation of feature profiles into business information and providing data value to third-party systems.

[0109] The above-mentioned encoding rules are used to transform the query feature vector to achieve same-dimensional matching between third-party data and feature profiles. Through similarity calculation, preset thresholds and core attribute confidence secondary screening, objective matching standards are established to improve the reliability of matching results, thereby improving the reliability of report generation.

[0110] In one specific embodiment, S55 includes the following steps:

[0111] S551, obtain a predefined report template, wherein the report template includes several attribute items and the presentation order of each attribute item.

[0112] S552, extract the values ​​of each attribute item from the attribute information corresponding to the target image.

[0113] S553, fill the extracted attribute values ​​into the corresponding positions in the report template to generate a data report corresponding to each target identifier code.

[0114] The predefined report templates are pre-designed report frameworks with fixed structures and fields, tailored to different business scenarios. They include several attribute items and the presentation order of each attribute item, providing clear guidance for subsequent attribute extraction and population. Based on the business scenario or specific requirements of the third-party system, the corresponding predefined report template is automatically or manually selected from the template library, ensuring a high degree of match between the report template and the third-party's needs, and avoiding information redundancy or omissions caused by generic templates.

[0115] The extracted and formatted attribute values ​​are used to replace the corresponding placeholders in the report template. The report is then rendered according to the output format specified in the template to generate the final structured data report, ensuring that the report format is correct, the content is complete, and it can be used directly.

[0116] The above approach, through policy analysis based on access status data from reference collection nodes, transforms the dispersed restriction rules of multiple sites into a unified and executable general access policy, resolving the policy adaptation issue across sites and improving the collection efficiency and feasibility of subsequent target collection nodes. Based on the general access policy, target collection nodes are filtered, and the collection of text data by these nodes is controlled. Nodes with non-compliant address types are eliminated at the source. Periodic comparisons promptly detect violations by target collection nodes, preventing access restrictions from being imposed on nodes due to prolonged violations. When a node's behavior changes from compliant to non-compliant, collection is immediately paused to prevent escalation of risk, while retaining already collected compliant text data to minimize data loss, thus ensuring the compliance of the collection nodes. The system ensures the continuity of data collection; it employs a pre-trained large model to parse the text data of each target network site, generating attribute information sets and improving parsing efficiency and accuracy; based on the update time of the target network site information, it verifies the consistency of multi-source attribute information of the data subject, accurately identifies data conflicts between multiple sites, avoids feature confusion caused by direct integration of conflicting data, and ensures the consistency of subject information and the integrity of feature profiles; it performs similarity matching between the feature description data of the target identifier code and the feature profile, generating corresponding data reports. Thus, through a series of compliant collection, accurate parsing, reliable integration, and efficient matching operations, multi-source network data is transformed into structured reports usable by third-party systems, improving data value and achieving reliable data collection and report generation.

[0117] Example 2

[0118] Embodiment 2 of the present invention provides a non-transitory computer-readable storage medium, which can be disposed in an electronic device to store at least one instruction or at least one program related to implementing a method in the method embodiment. The at least one instruction or at least one program is loaded and executed by the processor to implement the data acquisition and report generation method provided in the above embodiment.

[0119] Example 3

[0120] Embodiment 3 of the present invention provides an electronic device, which includes a processor and the non-transitory computer-readable storage medium of Embodiment 2 of the present invention.

[0121] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.

Claims

1. A data acquisition and report generation method, characterized in that, The data acquisition and report generation method includes the following steps: S1, based on the access status data and access results obtained by each reference acquisition node when accessing each target network site, performs policy analysis to obtain a general access policy applicable to all target network sites. S1 includes the following steps: S11, record the access status data generated by each reference acquisition node when accessing each target network site within a preset observation period, wherein the access status data includes at least one of login duration, access frequency, access time period, network address type, number of requests and time interval; S12, according to the order of data sampling time, construct a status data sequence of all access status data corresponding to each reference acquisition node; S13, associate and label each access status data with the access result of the corresponding access operation to obtain the labeled status data sequence, wherein the access result is successful access or access restriction is imposed; S14. Based on the labeled state data sequence generated by all reference acquisition nodes, an access result prediction model is trained, wherein the access result prediction model takes the state data sequence as input and takes successful access or being subject to access restrictions as output labels. S15, Analyze the importance of each state feature in the access result prediction model, and extract several key state features that lead to the prediction result of being subject to access restrictions and the threshold range of each key state feature. S16. Based on the aforementioned key state features and the threshold range of each key state feature, a general access policy applicable to all target network sites is generated, wherein the general access policy includes network address switching rules, access request frequency limits, daily access time periods, and single access duration suggestions. S2, based on the general access strategy, select several target collection nodes, and control each target collection node to collect text data from each target network site; S3, a pre-trained large model is used to parse all text data corresponding to each target network site to obtain the attribute information set corresponding to each target network site, wherein the attribute information set includes the attribute information of several data subjects; S4. Based on the information update time of each target network site, perform consistency verification on the attribute information of each data subject from all target network sites, and generate a feature profile corresponding to each data subject based on the verification results. S5, perform similarity matching on the feature description data and feature profile of each received target identifier code, and generate a data report corresponding to each target identifier code based on the matching results.

2. The data acquisition and report generation method according to claim 1, characterized in that, S2 includes the following steps: S21, perform a conformity judgment on the network address type and general access policy of each candidate collection node, and obtain the conformity judgment result corresponding to each candidate collection node, wherein the conformity judgment result is conforming or not conforming; S22, determine the candidate acquisition nodes that meet the compliance judgment result as the target acquisition nodes; S23, During the acquisition process, the access status data of each target acquisition node and the general access strategy are compared according to a preset comparison period to obtain the deviation judgment result of each target acquisition node in each preset comparison period, wherein the deviation judgment result is either deviation or no deviation; S24, when the deviation judgment result of the target acquisition node changes from not deviating to deviating, the acquisition behavior of the target acquisition node is paused, and the text data collected by the target acquisition node from each target network site when it is not deviating is obtained.

3. The data acquisition and report generation method according to claim 1, characterized in that, S3 includes the following steps: S31, Obtain the mapping relationship library between the domain name of the target website and the information extraction prompt template, wherein the information extraction prompt template specifies the attribute fields and format to be extracted from the text data; S32, based on the domain name of the target website from which each text data originates, retrieve the information extraction prompt template corresponding to each text data from the mapping relationship database; S33, input each text data and the corresponding information extraction prompt template into the pre-trained large model, and obtain the natural language response output by the pre-trained large model; S34, the natural language response is parsed and pre-standardized to generate a list of attribute information corresponding to each text data; S35, based on the list of attribute information corresponding to all text data of each target website, form an attribute information set corresponding to each target website.

4. The data acquisition and report generation method according to claim 1, characterized in that, S4 includes the following steps: S41, For any data subject, group the attribute information of the current data subject from all target network sites according to the attribute items, and obtain the attribute information group corresponding to each attribute item of the current data subject; S42, for any attribute item, perform temporal logic conflict detection on each attribute information in the attribute information group corresponding to the current attribute item, and obtain the conflict detection result corresponding to each attribute information, wherein the conflict detection result is either conflict or no conflict; S43, if the conflict detection result of the attribute information is no conflict, then the attribute information is determined to be reliable information; S44, If the conflict detection result of the attribute information is a conflict, then among all the attribute information that conflicts with the attribute information, the attribute information with the latest information update time of the corresponding target network site is determined as reliable information. S45. Generate a feature profile of the current data subject based on all reliable information in the attribute information group corresponding to all attribute items.

5. The data acquisition and report generation method according to claim 4, characterized in that, S45 includes the following steps: S451, quantify and vectorize each piece of reliable information corresponding to the current data subject to obtain the attribute feature vector; S452, perform preset processing on all attribute feature vectors corresponding to the current data subject to obtain the feature profile of the current data subject, wherein the preset processing includes splicing, dimensionality reduction and feature fusion.

6. The data acquisition and report generation method according to claim 1, characterized in that, S5 includes the following steps: S51, receive a query request sent by a third-party system, wherein the query request includes a number of target identifier codes and feature description data corresponding to each target identifier code; S52, convert each feature description data into a corresponding query feature vector; S53, calculate the similarity between each query feature vector and each feature profile; S54, For any target identifier, the feature image with the highest similarity among the feature images that have a similarity exceeding a preset similarity threshold with the current target identifier is determined as the target image corresponding to the current target identifier. S55 generates a data report corresponding to each target identifier code based on the target profile corresponding to each target identifier code.

7. The data acquisition and report generation method according to claim 6, characterized in that, S55 includes the following steps: S551, Obtain a predefined report template, wherein the report template includes several attribute items and the presentation order of each attribute item; S552, extract the values ​​of each attribute item from the attribute information corresponding to the target image; S553, fill the extracted attribute values ​​into the corresponding positions of the report template to generate a data report corresponding to each target identifier code.

8. A non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores at least one instruction or at least one program segment, characterized in that, The at least one instruction or the at least one program segment is loaded and executed by the processor to implement the data acquisition and report generation method as described in any one of claims 1-7.

9. An electronic device, characterized in that, Includes a processor and the non-transitory computer-readable storage medium as described in claim 8.