A method and system for collecting Internet information using artificial intelligence

CN122570802APending Publication Date: 2026-08-14HUAIAN WANZHIYUAN TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-29
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

传统互联网信息采集以单模态抓取为主,缺乏针对多模态数据的标准化分离与差异化预处理,未实现跨模态特征的统一对齐与协同建模

Benefits of technology

一、本发明通过多模态数据标准化分离与差异化预处理,对不同类型网页数据开展规范清洗、去重与元数据提取,保留数据单元关键属性与位置标识,建立统一的数据映射关系,从源头消除冗余、低质数据干扰,夯实多模态信息采集的数据基础;通过跨模态特征提取与多维度量化对齐,融合语义、时间、空间及模态差异相关维度,实现不同模态数据特征的协同匹配与关联,解决传统技术中多模态数据割裂、关联缺失的问题;基于多模态特征构建因果图并采用结构因果模型完成因果效应计算,精准识别数据单元间的真实关联,有效排除干扰因素影响,让数据关联判定更贴合互联网信息的实际传播与关联逻辑,大幅提升采集结果的精准性,实现从显性数据抓取到深层关联解析的升级,全面拓展采集内容的覆盖维度与深度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122570802A_ABST
    Figure CN122570802A_ABST
Patent Text Reader

Abstract

This invention discloses an artificial intelligence-based method and system for collecting internet information, relating to the fields of internet information collection and artificial intelligence technology. The specific steps of the method are as follows: performing multimodal data separation, obtaining raw data from the target webpage and separating it into four modal independent units, including text, and performing standardized preprocessing; cross-modal feature extraction, generating feature vectors and aligning them; constructing an initial causal graph to determine preliminary associations; using a structural causal model to calculate causal effects and filter node pairs; mining implicit information, fusing explicit and implicit information to generate complete results and storing them, supporting incremental updates. This invention eliminates redundant and low-quality data interference from the source, solves the problem of fragmented multimodal data, significantly improves collection accuracy, and achieves an upgrade from explicit crawling to deep analysis; it can mine implicit association information and fuse it to generate high-value collection results; and it adopts an incremental update mechanism to reduce resource consumption and improve update efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of Internet information collection and artificial intelligence technology, specifically to an Internet information artificial intelligence collection method and system. Background Technology

[0002] With the exponential growth of internet information, web page content has formed a multimodal ecosystem integrating text, images, videos, and structured tables. Information iterates rapidly and its logical connections are complex, continuously increasing the demands for the completeness, depth, and real-time nature of data collection. Traditional internet information collection primarily relies on single-modal crawling, lacking standardized separation and differentiated preprocessing for multimodal data, and failing to achieve unified alignment and collaborative modeling of cross-modal features. Furthermore, existing technologies largely depend on explicit content matching, failing to introduce causal reasoning mechanisms to uncover implicit correlations, and data updates often employ a full re-collection model, making it difficult to meet the precise, in-depth, and efficient collection needs of multimodal internet information.

[0003] Existing internet information collection technologies have many shortcomings: multimodal data processing is fragmented, lacking dedicated cleaning, deduplication, and metadata extraction processes for text, images, videos, and structured tables; it also fails to combine timestamps, page positions, and modal differences to achieve accurate quantification and alignment of cross-modal features, easily leading to the loss of intermodal correlations; correlation analysis is limited to semantic co-occurrence, failing to construct multimodal causal graphs, lacking structural causal models to quantify causal effects, and failing to effectively identify and eliminate confounding variables, resulting in low accuracy in causal relationship judgments; implicit information mining capabilities are weak, lacking a deep traversal mechanism based on causal paths, making it impossible to extract deep implicit correlations such as supply chain risks and public opinion transmission; data updates use full refreshes without event-driven incremental mechanisms, resulting in high resource consumption, poor real-time performance, and difficulty in adapting to frequently iterated pages.

[0004] In summary, existing technologies have significant shortcomings in multimodal collaborative processing, cross-modal feature alignment, causal reasoning, latent information mining, and efficient incremental updates, and cannot meet the current needs for the collection and in-depth analysis of multimodal Internet information. Summary of the Invention

[0005] The purpose of this invention is to overcome the shortcomings of existing technologies and provide an artificial intelligence-based method and system for collecting internet information. This method involves multimodal data separation, processing four modalities of data including text, cross-modal feature extraction, generating feature vectors and aligning them, constructing an initial causal graph to determine associations, using a structural causal model to calculate causal effects and filter node pairs, mining implicit information, fusing explicit and implicit information to generate complete results and supporting incremental updates.

[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: In one aspect, an artificial intelligence method for collecting internet information, the method comprising: S100, Multimodal Data Separation: Obtain the raw data of the target Internet page, separate it into four independent data units of four modalities: text, images, videos, and structured tables, and complete standardized preprocessing; S200, Cross-modal feature extraction: Perform feature extraction on data units of each modality to generate corresponding time series feature vectors and semantic feature vectors, and complete the alignment processing between features of different modalities; S300, Initial Cause-Effect Graph Construction: Construct an initial multimodal cause-effect graph based on multimodal feature vectors to determine the preliminary correlation between data units; S400, Causal Effect Calculation: The structural causal model is used to calculate the causal effect of the initial multimodal causal graph and screen out node pairs with significant causal relationships; S500 Implicit Information Fusion: Based on the filtered causal relationship node pairs, it mines implicit related information that is not directly displayed in the original page, merges explicit collected information with implicit related information, generates complete collection results and stores them, and supports incremental updates of collection results.

[0007] Furthermore, the text modality removes HTML tags, special characters, and duplicate content, segments the text into independent text units at the sentence level, and labels them with publication timestamps; the image modality removes duplicate images using perceptual hashing technology, filters low-quality images with resolutions below 72 dpi, extracts shooting time, device model, GPS location, and image size information from EXIF ​​metadata, fills null values ​​into metadata fields that fail to extract, and marks abnormal states; the video modality extracts keyframes at a frame rate of 1 fps, removes blurry and duplicate frames using structural similarity technology, and extracts video duration, resolution, encoding format, and upload time information; the structured table modality identifies and merges tables with the same structure, removes blank rows and columns, converts table data into key-value pair format, and labels the data update time; all preprocessed data units retain the relative position coordinates information from the original page, are uniformly assigned a 128-bit unique identifier ID, and establish a one-to-one mapping relationship between the ID and the original data source URL, collection time, and collection node number.

[0008] Furthermore, the alignment process between the different modal features utilizes the following calculation formula: in, The feature alignment weights between the i-th arbitrary modal data unit and the j-th arbitrary modal data unit; Cosine similarity between the semantic feature vectors of two data units; and These are the timestamps for two data units; The time decay coefficient, with a value of 86400s, is used to adjust the intensity of the influence of time difference on the association weight, and is determined by the statistics of the half-life of Internet information propagation. This represents the pixel distance between two data units on the page, in pixels (px). The modality difference coefficient is 1 for text-to-text modality, 1.2 for text-to-image modality, 1.5 for text-to-video modality, 1.3 for text-to-structured table modality, 1.4 for image-to-video modality, 1.6 for image-to-structured table modality, and 1.7 for video-to-structured table modality. The feature alignment weight The edge weights used for initial multimodal causal graph assignment ensure that multimodal data units that are temporally close, semantically related, and geographically adjacent receive higher association priority.

[0009] Furthermore, in the cross-modal feature extraction: for text modality, a bidirectional long short-term memory network combined with an attention mechanism is used to extract semantic and temporal features; for image modality, a visual Transformer is used to extract visual semantic and temporal features; for video modality, a 3D convolutional neural network is used to extract spatiotemporal features, which are then concatenated with keyframe visual semantic features for dimensionality reduction; for structured table modality, a fully connected neural network is used to extract numerical and statistical temporal features; all four modalities generate a 768-dimensional unified feature vector; for time series features, a sliding window method is used with a window size of 5 seconds and a step size of 1 second to extract the change features of the corresponding modality data, generating a 256-dimensional time series feature vector, and all time series feature vectors are subjected to L2 normalization.

[0010] Furthermore, the initial multimodal causal graph uses each independent data unit as a graph node. The node attributes include the modality type, timestamp, semantic feature vector, and page position information of the data unit. The edges between nodes are determined comprehensively based on the cross-modal feature alignment results of all pairs of modalities, cross-page co-occurrence, and semantic association. The direction of the edge is determined by the order of the timestamps of the two data units, with the earlier data unit pointing to the later data unit. For data units published at the same time, the direction of the edge is determined based on the semantic dominance relationship. The weight of the edge is divided into three levels: high, medium, and low, based on the association strength, and only the medium-level and above association edges are retained.

[0011] Furthermore, the structural causal model is used to calculate the causal effects of the initial multimodal causal graph, and the formula is as follows: in, The average causal effect of node u on node v; and These are two nodes in the initial multimodal causal graph; X is the intervention variable, representing whether the characteristics of node u are intervened. When X=1, it means that the characteristics of node u exist or have changed. When X=0, it means that the characteristics of node u do not exist or have not changed. The result variable represents the feature value of node v; The set of obfuscated variables, i.e., the set of nodes that simultaneously affect X and Y; The expected value of outcome variable Y is given when intervention variable X takes the value 1 and the set of confounding variables Z is fixed. The expected value of outcome variable Y is given when intervention variable X takes the value of 0 and the set of confounding variables Z is fixed. These are the quantized values ​​corresponding to the initial edge weight levels. The quantized values ​​for the high, medium, and low levels are 0.85, 0.50, and 0.15, respectively. This formula introduces an initial edge weight correction term, which solves the problem of causal effect bias caused by edge weight differences in the initial multimodal causal graph; when If the value is greater than 0.5, the node pair is considered to have a significant causal relationship.

[0012] Furthermore, the identification of confounding variables in the causal effect calculation adopts a graph neural network-based method. The initial multimodal causal graph is input into a two-layer graph attention network. The influence of any node on other nodes is evaluated through a multi-head attention mechanism. When the influence of any node on both node u and node v is greater than 0.5, the node is marked as a confounding variable and added to the confounding variable set Z. The size of the confounding variable set Z is capped at 10. When the number of identified confounding variables exceeds 10, only the top 10 nodes with the highest influence are retained. When calculating the causal effect of any node pair with a significant causal relationship, the corresponding confounding variable set Z is substituted as a conditional variable into the causal effect calculation formula to exclude the influence of confounding variables. The system fine-tunes the graph attention network every 30 days using new labeled data.

[0013] Furthermore, the implicit correlation information mining adopts a depth-first path traversal method. Starting from each node pair with a significant causal relationship, it traverses all causal paths with a length not exceeding 3, and integrates all node information on each path to generate implicit correlation information. Twelve common implicit information type templates are preset, including supply chain risks, event causal relationships, user behavior preferences, product quality issues, market trend changes, public opinion dissemination paths, competitor dynamic correlations, financial risk transmission, user churn causes, content dissemination patterns, security vulnerability correlations, and compliance risk hazards. The integrated implicit correlation information is matched with the implicit information type templates to generate structured implicit correlation information. A unique identifier is generated for each structured implicit correlation information, and the corresponding causal path, causal effect value, and number of data sources involved are recorded.

[0014] Furthermore, the incremental update adopts an event-driven model. When the target page data changes, the system only extracts the data units of the changed part and completes feature extraction, updating the affected nodes and edges in the initial multimodal causal graph. The system maintains a causal effect cache table to store the causal effect values ​​of all generated node pairs. When the features of any node change, only the causal effect values ​​of all node pairs directly connected to that node are regenerated. The incremental update time interval is set to 5 minutes. For high-priority target pages, the time interval can be adjusted to 1 minute. During the incremental update process, historical causal relationship data is retained to track the change trend.

[0015] On the other hand, an internet information artificial intelligence collection system includes: Multimodal data separation module: used to acquire the raw data of the target Internet page, separate it into four independent data units of four modalities: text, images, videos and structured tables, and complete standardized preprocessing; Cross-modal feature extraction module: used to extract features from data units of each modality, generate corresponding time series feature vectors and semantic feature vectors, and complete the alignment processing between features of different modalities; Initial Cause-Effect Graph Construction Module: Used to construct an initial multimodal cause-effect graph based on multimodal feature vectors, and to determine the preliminary correlation between data units; Causal effect calculation module: used to calculate the causal effect of the initial multimodal causal graph using a structural causal model, and to screen out node pairs with significant causal relationships; Implicit Information Fusion Module: Based on the filtered causal relationship node pairs, it mines implicit related information that is not directly displayed in the original page, merges explicit collected information with implicit related information, generates complete collection results and stores them, and supports incremental updates of collection results.

[0016] Compared with existing technologies, this method and system for collecting internet information using artificial intelligence has the following advantages: I. This invention utilizes standardized separation and differentiated preprocessing of multimodal data to perform standardized cleaning, deduplication, and metadata extraction on different types of web page data. It retains key attributes and location identifiers of data units, establishes a unified data mapping relationship, and eliminates redundant and low-quality data interference from the source, thus solidifying the data foundation for multimodal information collection. Through cross-modal feature extraction and multi-dimensional quantification alignment, it integrates semantic, temporal, spatial, and modal difference-related dimensions to achieve collaborative matching and association of features from different modalities, solving the problems of fragmented multimodal data and missing associations in traditional technologies. Based on multimodal features, it constructs a causal graph and uses a structural causal model to calculate causal effects, accurately identifying the real associations between data units, effectively eliminating the influence of interference factors, and making data association determination more consistent with the actual dissemination and association logic of internet information. This significantly improves the accuracy of the collection results, achieving an upgrade from explicit data capture to deep association analysis, and comprehensively expanding the coverage dimensions and depth of the collected content.

[0017] Second, this invention performs deep path traversal through nodes with significant causal relationships to uncover implicit correlations not directly presented on the original page. Combined with preset information type templates, it completes structured transformation, efficiently integrating explicit data with implicit correlations to generate complete and high-value collection results, significantly enhancing the practical value and analytical dimensions of information collection. Employing an event-driven incremental update mechanism, it processes only page-changing data units, maintains a causal effect cache table to reduce redundant calculations, adjusts the update rhythm based on page priority, and retains historical causal relationship data for trend tracking, greatly reducing system resource consumption and improving the real-time performance and response efficiency of information updates. The entire solution forms a closed-loop intelligent collection process, adapting to the high-frequency iteration characteristics of internet information, optimizing the operating efficiency of the collection system, and providing stable and reliable technical support for intelligent collection and in-depth analysis of internet information in multiple scenarios.

[0018] Other advantages, objectives and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination or study, or may be learned from the practice of the invention. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.

[0020] Figure 1 A step-by-step framework diagram of an artificial intelligence method for collecting internet information; Figure 2 A flowchart illustrating the multimodal data separation and standardized preprocessing of an artificial intelligence-based internet information acquisition method; Figure 3 This is a flowchart of an internet information artificial intelligence collection system. Detailed Implementation

[0021] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided below.

[0022] Example 1: S100 Multimodal Data Separation: This process acquires raw data from target internet public opinion pages and precisely separates it into four independent data units: text, images, videos, and structured tables. Simultaneously, standardized preprocessing is performed, ensuring that different types of public opinion data can enter a unified processing flow, laying a well-organized data foundation for subsequent feature extraction and correlation analysis. For the text modality, HTML tags, special characters, and duplicate content are removed. The data is segmented into independent text units at the sentence level and labeled with a publication timestamp, eliminating redundant and interfering information to keep the text-based public opinion content clean and clear. The publication time is accurately recorded, providing a reliable basis for subsequent time-based correlation analysis. For the image modality, perceptual hashing technology is used to remove duplicate images and filter low-quality images with a resolution below 72 dpi. The shooting time, device model, GPS location, and image size information are extracted from EXIF ​​metadata. Null values ​​are filled into metadata fields that fail to extract data, and anomalies are marked. Invalid image resources are filtered out, retaining high-definition, valid images with public opinion reference value. All image-related information is fully preserved, ensuring the authenticity and integrity of the image data. The video modality extracts keyframes at a frame rate of 1fps, removes blurry and duplicate frames using structural similarity technology, and extracts video duration, resolution, encoding format, and upload time information, streamlining the video data volume while retaining core video footage and basic attribute information. This avoids invalid video data consuming processing resources and accurately records video release-related information, adapting to the time-related analysis needs of public opinion data. The structured table modality identifies and merges tables with the same structure, removes blank rows and columns, converts table data to key-value pair format, and annotates data update times, standardizing the table data structure and making tabular public opinion data easier to read and analyze. Clearly annotating data update times accurately reflects the dynamic changes in public opinion data. All preprocessed data units retain the relative position coordinates information from the original page and are uniformly assigned a 128-bit unique identifier ID. A one-to-one mapping relationship is established between the ID and the original data source URL, collection time, and collection node number, ensuring that each public opinion data unit has a unique and identifiable identifier. This enables accurate data traceability and rapid location, ensuring that public opinion data is not confused or lost throughout the entire processing flow, comprehensively improving the standardization and accuracy of data management.

[0023] S200 Cross-Modal Feature Extraction: Features are extracted from data units of each modality, generating corresponding time-series feature vectors and semantic feature vectors. Alignment processing between features of different modalities is achieved, unifying the expression of features from four types of public opinion data: text, images, videos, and structured tables. This breaks down information barriers between modalities and supports the association and fusion of multimodal data. For the text modality, a bidirectional long short-term memory network combined with an attention mechanism is used to extract semantic and temporal features, accurately capturing the core semantic information and temporal change characteristics of text-based public opinion content, and deeply exploring the public opinion trends and temporal correlation logic behind the text. For the image modality, a visual Transformer is used to extract visual semantic and temporal features, efficiently parsing the visual content information in images and synchronously associating the image's temporal attributes, allowing for a complete fusion of visual and temporal features. For the video modality, a 3D convolutional neural network is used to extract spatiotemporal features, which are then concatenated with keyframe visual semantic features for dimensionality reduction, comprehensively capturing the spatial image features and temporal evolution features of the video, simplifying feature dimensions while retaining the core public opinion information of the video. The structured table modality employs a fully connected neural network to extract numerical and statistical time features, accurately analyzing the numerical patterns and temporal trends within the tables, making the structured data's feature representation more aligned with the needs of public opinion analysis. All four modalities generate a unified 768-dimensional feature vector, ensuring consistency in feature dimensions across different modalities and achieving standardized integration of multimodal features, avoiding feature matching failures due to dimensional differences. Time series features utilize a sliding window method with a 5-second window size and a 1-second step, extracting the changing features of the corresponding modal data and generating a 256-dimensional time series feature vector. All time series feature vectors undergo L2 normalization to accurately capture the dynamic changes in public opinion data over time. Normalization eliminates the impact of feature numerical differences, making the comparison and correlation of time series features more reasonable. Subsequently, alignment processing is performed between different modal features. The correlation weight of each data unit is determined through feature alignment weight calculation, using the following formula: ,in, The feature alignment weights between the i-th arbitrary modal data unit and the j-th arbitrary modal data unit; Cosine similarity between the semantic feature vectors of two data units; and These are the timestamps for two data units; The time decay coefficient, with a value of 86400s, is used to adjust the intensity of the influence of time difference on the association weight, and is determined by the statistics of the half-life of Internet information propagation. This represents the pixel distance between two data units on the page, in pixels (px). The modal difference coefficient is set as follows: text-to-text modality = 1, text-to-image modality = 1.2, text-to-video modality = 1.5, text-to-structured table modality = 1.3, image-to-video modality = 1.4, image-to-structured table modality = 1.6, and video-to-structured table modality = 1.7. This coefficient prioritizes multimodal public opinion data units that are close in time, semantically related, and geographically adjacent, accurately matching the intrinsic connections between different modalities. This makes the alignment results of cross-modal features more consistent with the actual correlation logic of public opinion data, providing a scientific weighting basis for subsequent causal graph construction.

[0024] S300 Initial Cause-and-Effect Graph Construction: An initial multimodal cause-and-effect graph is constructed based on multimodal feature vectors to determine the preliminary relationships between data units. This visualized graph structure presents the connections between multimodal public opinion data, allowing scattered data units to form an organic whole and clearly demonstrating the inherent logic of public opinion information. The initial multimodal cause-and-effect graph uses each independent data unit as a node. Node attributes include the data unit's modality type, timestamp, semantic feature vector, and page location information, comprehensively carrying the core attributes of the public opinion data unit and ensuring that node information fully supports subsequent correlation analysis and causal calculations. Edges between nodes are determined comprehensively based on the cross-modal feature alignment results of all modal pairwise combinations, cross-page co-occurrence, and semantic correlation. This multi-dimensional correlation assessment makes edge generation more objective and accurate, avoiding biases in correlation determination caused by a single criterion. The direction of the edges is determined by the order of the timestamps of the two data units, with the earlier data unit pointing to the later one. For data units released simultaneously, the direction of the edges is determined based on the semantic dominance relationship, following the propagation sequence and semantic logic of public opinion information to ensure that the direction of the causal graph aligns with the actual development of public opinion. Edge weights are categorized into high, medium, and low levels based on their correlation strength. Only edges with medium or higher correlations are retained, while weak, invalid edges are removed. This simplifies the causal graph structure, reduces the complexity of subsequent calculations, and ensures that the retained edges all possess practical value for public opinion analysis, significantly improving the effectiveness and practicality of the initial causal graph.

[0025] S400 Causal Effect Calculation: A structural causal model is used to calculate the causal effects of the initial multimodal causal graph. Node pairs with significant causal relationships are selected, and the causal strength between public opinion data units is quantified through scientific calculation methods. This accurately identifies the core causal relationships between public opinion information, eliminates false and weak correlations, and makes the determination of public opinion causal relationships more precise. The average causal effect is calculated during the process, using the following formula: ,in, The average causal effect of node u on node v; and These are two nodes in the initial multimodal causal graph; X is the intervention variable, representing whether the characteristics of node u are intervened. When X=1, it means that the characteristics of node u exist or have changed. When X=0, it means that the characteristics of node u do not exist or have not changed. The result variable represents the feature value of node v; The set of obfuscated variables, i.e., the set of nodes that simultaneously affect X and Y; The expected value of outcome variable Y is given when intervention variable X takes the value 1 and the set of confounding variables Z is fixed. The expected value of outcome variable Y is given when intervention variable X takes the value of 0 and the set of confounding variables Z is fixed. The initial edge weight levels are represented by quantized values ​​of 0.85, 0.50, and 0.15, respectively. These quantified values ​​visually reflect the degree of causal influence between nodes, providing a clear numerical basis for determining causal relationships. In the causal effect calculation, confounding variable identification employs a graph neural network-based method. The initial multimodal causal graph is input into a two-layer graph attention network. A multi-head attention mechanism assesses the influence of any node on other nodes. When the influence of any node on both node u and node v is greater than 0.5, that node is marked as a confounding variable and added to the confounding variable set. This accurately identifies confounding variables that interfere with causal calculations, eliminating interference from external factors in determining the causal relationship of public opinion and improving the accuracy of the calculation results. The maximum size of the confounding variable set is 10. When the number of identified confounding variables exceeds 10, only the top 10 nodes with the highest influence are retained. This controls the size of the confounding variable set, preventing computational efficiency reduction due to too many variables, while retaining core confounding variables to ensure effective interference removal. When calculating the causal effect of any node pair with a significant causal relationship, the corresponding set of confounding variables is substituted as conditional variables into the causal effect calculation formula to completely eliminate the interference of confounding variables, making the causal effect calculation results more consistent with the true correlation of public opinion data. Every 30 days, the system fine-tunes the graph attention network using new labeled data, continuously optimizing the accuracy of the confounding variable identification model, allowing the model to adapt to constantly changing public opinion data characteristics and maintain the accuracy and stability of confounding variable identification in the long term. When the average causal effect value reaches the judgment criteria, the node pair is determined to have a significant causal relationship. By using clear judgment criteria to screen core public opinion causal associations, precise analytical objects are provided for subsequent implicit information mining.

[0026] S500 Implicit Information Fusion: Based on the filtered causal relationship node pairs, it mines implicit correlation information not directly displayed on the original page. After fusing explicit collected information with implicit correlation information, it generates and stores complete collection results, supporting incremental updates. It deeply mines the potential connections behind the surface information of public opinion, upgrading public opinion collection from surface information gathering to deep information integration, comprehensively improving the completeness and depth of public opinion information collection. Implicit correlation information mining adopts a depth-first path traversal method. Starting from each node pair with a significant causal relationship, it traverses all causal paths with a length not exceeding 3, integrating all node information on each path to generate implicit correlation information. It deeply explores the hidden connections of public opinion information along the causal paths, leaving no potential public opinion logic unexplored, making the mining of implicit information more comprehensive and in-depth. The system pre-defines 12 common implicit information type templates, including supply chain risks, event causal relationships, user behavior preferences, product quality issues, market trend changes, public opinion dissemination paths, competitor dynamics, financial risk transmission, user churn causes, content dissemination patterns, security vulnerability associations, and compliance risk hazards. It matches the integrated implicit information with these templates to generate structured implicit information, transforming scattered implicit information into standardized, structured data that is easier to store, query, and analyze. Each structured implicit information piece generates a unique identifier, recording the corresponding causal path, causal effect value, and the number of data sources involved. This ensures that every piece of implicit information is traceable and verifiable, clearly indicating the basis for its generation and improving the credibility and usability of the implicit information. Incremental updates adopt an event-driven model. When the target page data changes, the system only extracts the changed data units and completes feature extraction, updating the affected nodes and edges in the initial multimodal causal graph. This eliminates the need for repeated processing of the entire dataset, significantly improving data update efficiency and reducing system resource consumption. The system maintains a causal effect cache table to store the causal effect values ​​of all generated node pairs. When the characteristics of any node change, only the causal effect values ​​of all node pairs directly connected to that node are regenerated, reducing redundant calculations and accelerating incremental updates. The incremental update interval is set to 5 minutes. For high-priority target pages, the interval can be adjusted to 1 minute, balancing the update efficiency of ordinary pages with the real-time requirements of high-priority pages, ensuring that public opinion data is always up-to-date. Historical causal relationship data is retained during incremental updates to track trends. By comparing historical and real-time data, the evolution of public opinion information is clearly displayed, providing a more comprehensive reference for public opinion analysis and decision-making. Figure 1 As shown.

[0027] Example 2: Multimodal Data Separation Module: This module is responsible for acquiring the raw data of the target e-commerce product page, separating it into independent data units of four modalities: text, images, videos, and structured tables. It performs standardized preprocessing to provide a clean and standardized foundation for subsequent processing of e-commerce product information, allowing different types of product information to enter a unified processing flow and eliminating processing barriers caused by data format differences. For the text modality, HTML tags, special characters, and duplicate content are removed. The data is segmented into independent text units at the sentence level and marked with a publication timestamp. Redundant and interfering information in product descriptions is eliminated, making product descriptions more concise and clear. The publication time of text information is accurately recorded, providing a basis for subsequent time-related analysis of product information. For the image modality, perceptual hashing technology is used to remove duplicate images, filter low-quality images with a resolution lower than 72 dpi, extract shooting time, device model, GPS location, and image size information from EXIF ​​metadata, fill null values ​​into metadata fields that fail to extract, and mark abnormal states. Duplicate and blurry invalid product display images are filtered out, retaining high-definition, high-quality product image resources and completely preserving the image's auxiliary information, making the product image data more valuable for reference while ensuring the integrity and traceability of image information. The video modality extracts keyframes at a frame rate of 1fps, removes blurry and duplicate frames using structural similarity technology, and extracts video duration, resolution, encoding format, and upload time information. This streamlines the data volume of product demonstration videos, retains the core display images and basic attributes, avoids invalid video data consuming system processing resources, and accurately records video upload time to meet the time-dimensional analysis needs of product information. The structured table modality identifies and merges tables with the same structure, removes blank rows and columns, converts table data to key-value pair format, and annotates data update times. This standardizes the data structure of product parameter tables, making product parameter information easier to read, parse, and analyze, clearly annotating data update times, and reflecting the dynamic changes of product parameters in real time. All preprocessed data units retain the relative position coordinates information from the original page, are uniformly assigned a 128-bit unique identifier ID, and establish a one-to-one mapping relationship between the ID and the original data source URL, collection time, and collection node number, such as... Figure 2 As shown, each e-commerce product data unit has a unique identifier, enabling accurate traceability, rapid location, and efficient management of product information. This avoids problems such as data confusion and loss during processing, and fully guarantees the standardization and accuracy of e-commerce product data.

[0028] Cross-modal feature extraction module: This module extracts features from product data units of each modality, generating corresponding time-series feature vectors and semantic feature vectors. It aligns features across different modalities, breaking down modal barriers between text, images, videos, and tables in e-commerce products. This allows for unified expression and precise matching of multimodal product features, laying a feature foundation for subsequent causal relationship analysis of product information. For text modality, a bidirectional long short-term memory network combined with an attention mechanism is used to extract semantic and temporal features. This deeply analyzes the core semantic information of product descriptions, simultaneously capturing the temporal changes in text information, and accurately extracting the product's core selling points and attributes. For image modality, a visual Transformer is used to extract visual semantic and temporal features, efficiently analyzing the visual content of product display images, associating the images with temporal attributes, and fully presenting the product's visual features and temporal relationships. For video modality, a 3D convolutional neural network is used to extract spatiotemporal features. After concatenation with keyframe visual semantic features, dimensionality is reduced, comprehensively capturing the spatial and temporal evolution features of product demonstration videos. While simplifying feature dimensions, the core product display information in the video is fully preserved. The structured table modality employs a fully connected neural network to extract numerical and statistical time features, accurately analyzing the numerical patterns and time trends in product parameter tables, making the feature representation of product parameters more aligned with the needs of e-commerce data analysis. All four modalities generate a unified 768-dimensional feature vector, unifying the feature dimensions of multimodal product data and achieving standardized integration of features from different modalities, ensuring smooth feature matching and fusion. Time series features utilize a sliding window method with a window size of 5 seconds and a step size of 1 second, extracting the changing features of the corresponding modal data and generating a 256-dimensional time series feature vector. All time series feature vectors undergo L2 normalization to accurately capture the dynamic changes of product information over time. Normalization eliminates differences in feature values, making the comparative analysis of time series features more reasonable. The module completes the alignment processing between features from different modalities through feature alignment weight calculation, giving higher priority to multimodal product data units that are time-prone, semantically related, and geographically adjacent. This accurately matches the inherent connections between product modal information, making the cross-modal feature alignment results more consistent with the actual association logic of e-commerce product information, providing scientific and reliable weight support for the initial causal graph construction.

[0029] Initial Cause-and-Effect Graph Construction Module: This module constructs an initial multimodal cause-and-effect graph based on multimodal feature vectors. It determines the preliminary relationships between product data units, clearly presenting the connections between multimodal information of e-commerce products in a graph structure. This integrates scattered product data units into an organic whole, intuitively showcasing the inherent logic of product information relationships. The initial multimodal cause-and-effect graph uses each independent product data unit as a graph node. Node attributes include the data unit's modality type, timestamp, semantic feature vector, and page location information, comprehensively carrying the core attributes of the product data and providing sufficient information support for subsequent association analysis and causal calculations. Edges between nodes are determined comprehensively based on the cross-modal feature alignment results of all modal pairwise combinations, cross-page co-occurrence, and semantic association degree. This multi-dimensional approach to determining node connection relationships makes edge generation more objective, avoiding bias in association judgments caused by a single criterion and accurately reflecting the true relationships between product information. The direction of the edges is determined by the order of the timestamps of the two data units, with the earlier data unit pointing to the later data unit. For data units published simultaneously, the direction of the edges is determined based on the semantic dominance relationship. The direction of the edges follows the publication sequence and semantic logic of e-commerce product information, ensuring that the direction of the causal graph conforms to the presentation rules of product information. Edge weights are divided into high, medium, and low levels based on the strength of the association. Only edges with medium or higher association strength are retained, while weak, invalid edges are removed. This simplifies the causal graph structure, reduces the complexity of subsequent causal effect calculations, and ensures that the retained association edges have practical e-commerce analysis value, significantly improving the practicality and effectiveness of the initial causal graph.

[0030] Causal Effect Calculation Module: This module uses a structural causal model to calculate causal effects on the initial multimodal causal graph, filtering out node pairs with significant causal relationships. It quantifies the causal strength between e-commerce product data units through scientific calculation methods, accurately identifying core causal associations between product information, and eliminating false and weak associations, making the determination of causal relationships in product information more accurate. The module performs average causal effect calculation, using quantitative values ​​to intuitively reflect the degree of causal influence between product data nodes, providing a clear numerical standard for determining significant causal relationships. During the causal effect calculation process, a graph neural network-based method is used to identify confounding variables. The initial multimodal causal graph is input into a two-layer graph attention network, and a multi-head attention mechanism is used to evaluate the influence of any node on other nodes, accurately identifying confounding variables that simultaneously affect the target node pair, eliminating the influence of interfering factors on the causal calculation of product information, and improving the accuracy of the calculation results. The maximum size of the confounding variable set is set to 10; if it exceeds this limit, the top 10 nodes with the highest influence are retained, ensuring the effectiveness of interference elimination while controlling the calculation scale and improving computational efficiency. During calculation, the set of confounding variables is substituted as conditional variables to completely eliminate their interference, making the causal effect calculation results more closely reflect the true correlation of e-commerce product information. Every 30 days, the system fine-tunes the graph attention network using newly labeled data, continuously optimizing the accuracy of the confounding variable identification model, allowing it to adapt to the dynamic changes in e-commerce product information and maintain long-term accuracy. When the average causal effect value reaches the judgment criteria, product data node pairs with significant causal relationships are accurately selected, providing precise analysis targets for subsequent implicit product information mining, thus improving the targeting and effectiveness of implicit information mining.

[0031] Implicit Information Fusion Module: This module, based on filtered causal relationship node pairs, mines implicit related information not directly displayed on the original product page. It merges explicit collected information with implicit related information to generate and store complete collection results. It also supports incremental updates of the collection results, deeply exploring the potential connections behind the surface information of e-commerce products. This upgrades product information collection from superficial gathering to deep integration, comprehensively improving the completeness and depth of e-commerce product information collection. The module uses a depth-first path traversal method to mine implicit related information. Starting from significant causal relationship node pairs, it traverses a causal path of a specified length and integrates path node information. It deeply explores potential connections along the causal path of product information, comprehensively mining the implicit attributes, relational logic, and market patterns of products. The module matches 12 preset implicit information type templates, generating structured data from the integrated implicit information. This standardizes scattered implicit product information into a format that is easier to store, query, analyze, and apply. The module generates a unique identifier for each structured implicit relationship, recording the corresponding causal path, causal effect value, and the number of data sources involved. This ensures that every piece of implicit product information is traceable and verifiable, clearly marking the basis for information generation and enhancing the credibility and commercial value of implicit information. Incremental updates adopt an event-driven model. When page data changes, only the updated data units are extracted and processed, updating only the affected nodes and edges in the causal graph, eliminating the need for full-scale reprocessing, significantly improving update efficiency and reducing system resource consumption. The module maintains a causal effect cache table, storing the causal effect values ​​of already generated node pairs. When node characteristics change, only the causal effects of connected node pairs are recalculated, reducing redundant calculations and accelerating update speed. The default interval for incremental updates is 5 minutes, which can be adjusted to 1 minute for high-priority product pages, balancing the update efficiency of ordinary product information with the real-time nature of core product information, ensuring that e-commerce product information is always up-to-date. Historical causal relationship data is retained during incremental updates. By comparing historical and real-time data, the evolution trend of product information is clearly displayed, providing comprehensive and accurate data support for e-commerce operations, product analysis, and market decision-making. Figure 3 As shown.

[0032] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.

Claims

1. A method for collecting internet information using artificial intelligence, characterized in that, The method includes: S100, Multimodal Data Separation: Obtain the raw data of the target Internet page, separate it into four independent data units of four modalities: text, images, videos, and structured tables, and complete standardized preprocessing; S200, Cross-modal feature extraction: Perform feature extraction on data units of each modality to generate corresponding time series feature vectors and semantic feature vectors, and complete the alignment processing between features of different modalities; S300, Initial Cause-Effect Graph Construction: Construct an initial multimodal cause-effect graph based on multimodal feature vectors to determine the preliminary correlation between data units; S400, Causal Effect Calculation: The structural causal model is used to calculate the causal effect of the initial multimodal causal graph and screen out node pairs with significant causal relationships; S500 Implicit Information Fusion: Based on the filtered causal relationship node pairs, it mines implicit related information that is not directly displayed in the original page, merges explicit collected information with implicit related information, generates complete collection results and stores them, and supports incremental updates of collection results.

2. The method for collecting internet information using artificial intelligence according to claim 1, characterized in that, In step S100, the text modality removes HTML tags, special characters, and duplicate content, and is segmented into independent text units at the sentence level and marked with a publication timestamp; The image modality uses perceptual hashing to remove duplicate images, filters low-quality images with a resolution below 72 dpi, and extracts shooting time, device model, GPS location, and image size information from EXIF ​​metadata. Data fields that fail to extract are filled with null values ​​and marked as abnormal. The video modality extracts keyframes at a frame rate of 1 fps, removes blurry and duplicate frames using structural similarity technology, and extracts video duration, resolution, encoding format, and upload time information. The structured table modality identifies and merges tables with the same structure, removes blank rows and columns, converts table data to key-value pair format, and labels the data update time. All preprocessed data units retain the relative position coordinates from the original page, are uniformly assigned a 128-bit unique identifier ID, and establish a one-to-one mapping relationship between the ID and the original data source URL, acquisition time, and acquisition node number.

3. The method for collecting internet information using artificial intelligence according to claim 1, characterized in that, In step S200, the alignment process between the different modal features utilizes the following calculation formula: in, The feature alignment weights between the i-th arbitrary modal data unit and the j-th arbitrary modal data unit; Cosine similarity between the semantic feature vectors of two data units; and These are the timestamps for the two data units respectively; The time decay coefficient, with a value of 86400s, is used to adjust the intensity of the influence of time difference on the association weight, and is determined by the statistics of the half-life of Internet information propagation. This represents the pixel distance between two data units on the page, in pixels (px). The modality difference coefficient is 1 for text-to-text modality, 1.2 for text-to-image modality, 1.5 for text-to-video modality, 1.3 for text-to-structured table modality, 1.4 for image-to-video modality, 1.6 for image-to-structured table modality, and 1.7 for video-to-structured table modality. The feature alignment weight The edge weights used for initial multimodal causal graph assignment ensure that multimodal data units that are temporally close, semantically related, and geographically adjacent receive higher association priority.

4. The method for collecting internet information using artificial intelligence according to claim 1, characterized in that, In step S200, during the cross-modal feature extraction: the text modality uses a bidirectional long short-term memory network combined with an attention mechanism to extract semantic and temporal features; Image modality uses a visual Transformer to extract visual semantics and temporal features; video modality uses a 3D convolutional neural network to extract spatiotemporal features, which are then concatenated with keyframe visual semantic features for dimensionality reduction; structured table modality uses a fully connected neural network to extract numerical and statistical temporal features; all four modalities generate a 768-dimensional unified feature vector; time series features are extracted using a sliding window method with a window size of 5 seconds and a step size of 1 second, extracting the change features of the corresponding modality data and generating a 256-dimensional time series feature vector, all of which are L2 normalized.

5. The method for artificial intelligence collection of Internet information according to claim 1, characterized in that, In step S300, the initial multimodal causal graph uses each independent data unit as a graph node. The node attributes include the modality type, timestamp, semantic feature vector, and page position information of the data unit. The edges between nodes are determined comprehensively based on the cross-modal feature alignment results of all pairs of modalities, cross-page co-occurrence, and semantic association. The direction of the edge is determined by the order of the timestamps of the two data units, with the earlier data unit pointing to the later data unit. For data units published at the same time, the direction of the edge is determined based on the semantic dominance relationship. The weight of the edge is divided into three levels: high, medium, and low, based on the association strength, and only the medium-level and above association edges are retained.

6. The method for collecting Internet information using artificial intelligence according to claim 1, characterized in that, In step S400, the structural causal model is used to calculate the causal effects of the initial multimodal causal graph, and the formula is as follows: in, The average causal effect of node u on node v; and These are two nodes in the initial multimodal causal graph; X is the intervention variable, representing whether the characteristics of node u are intervened. When X=1, it means that the characteristics of node u exist or have changed. When X=0, it means that the characteristics of node u do not exist or have not changed. The result variable represents the feature value of node v; The set of obfuscated variables, i.e., the set of nodes that simultaneously affect X and Y; The expected value of outcome variable Y is given when intervention variable X takes the value 1 and the set of confounding variables Z is fixed. The expected value of outcome variable Y is given when intervention variable X takes the value of 0 and the set of confounding variables Z is fixed. These are the quantized values ​​corresponding to the initial edge weight levels. The quantized values ​​for the high, medium, and low levels are 0.85, 0.50, and 0.15, respectively. when If the value is greater than 0.5, the node pair is considered to have a significant causal relationship.

7. The method for collecting internet information using artificial intelligence according to claim 6, characterized in that, In step S400, the identification of confounding variables in the causal effect calculation adopts a graph neural network-based method. The initial multimodal causal graph is input into a two-layer graph attention network. The influence of any node on other nodes is evaluated through a multi-head attention mechanism. When the influence of any node on both node u and node v is greater than 0.5, the node is marked as a confounding variable and added to the confounding variable set Z. The size of the confounding variable set Z is capped at 10. When the number of identified confounding variables exceeds 10, only the top 10 nodes with the highest influence are retained. When calculating the causal effect of any node pair with a significant causal relationship, the corresponding confounding variable set Z is substituted as a conditional variable into the causal effect calculation formula to exclude the influence of confounding variables. The system fine-tunes the graph attention network every 30 days using new labeled data.

8. The method for collecting internet information using artificial intelligence according to claim 1, characterized in that, In step S500, the latent association information mining adopts a depth-first path traversal method, starting from each node pair with a significant causal relationship, traversing all causal paths with a length not exceeding 3, and integrating all node information on each path to generate latent association information. Twelve common implicit information type templates are preset. The integrated implicit information is matched with the implicit information type templates to generate structured implicit information. A unique identifier is generated for each structured implicit information, and the corresponding causal path, causal effect value and number of data sources involved are recorded.

9. The method for collecting internet information using artificial intelligence according to claim 1, characterized in that, In step S500, the incremental update adopts an event-driven mode. When the target page data changes, the system only extracts the data units of the changed part and completes feature extraction, and updates the affected nodes and edges in the initial multimodal causal graph. The system maintains a causal effect cache table to store the causal effect values ​​of all generated node pairs. When the features of any node change, only the causal effect values ​​of all node pairs directly connected to that node are regenerated. The incremental update time interval is set to 5 minutes. For high-priority target pages, the time interval can be adjusted to 1 minute. During the incremental update process, historical causal relationship data is retained to track the change trend.

10. An Internet information artificial intelligence acquisition system, the system being applicable to the Internet information artificial intelligence acquisition method according to any one of claims 1-9, characterized in that, The system includes: Multimodal data separation module: used to acquire the raw data of the target Internet page, separate it into four independent data units of four modalities: text, images, videos and structured tables, and complete standardized preprocessing; Cross-modal feature extraction module: used to extract features from data units of each modality, generate corresponding time series feature vectors and semantic feature vectors, and complete the alignment processing between features of different modalities; Initial Cause-Effect Graph Construction Module: Used to construct an initial multimodal cause-effect graph based on multimodal feature vectors, and to determine the preliminary correlation between data units; Causal effect calculation module: used to calculate the causal effect of the initial multimodal causal graph using a structural causal model, and to screen out node pairs with significant causal relationships; Implicit Information Fusion Module: Based on the filtered causal relationship node pairs, it mines implicit related information that is not directly displayed in the original page, merges explicit collected information with implicit related information, generates complete collection results and stores them, and supports incremental updates of collection results.