A data acquisition method and device, electronic equipment and storage medium
Patent Information
- Application Number
- CN202610999156.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-06
- Publication Date
- 2026-09-25
AI Technical Summary
[0005]基于上述现有技术的不足,本申请提供了一种数据采集方法及装置、电子设备、存储介质,以解决现有技术无法获取网页的关键信息,导致无法生成准确结果的问题
[0064]本申请提供的一种数据采集方法,当接收目标网页的访问请求时,从所述访问请求的请求信息和行为信息中提取出多维访问特征,以便于可以基于请求信息和行为信息的特征信息,准确识别出目标爬虫。其中,所述目标爬虫可以为人工智能模型的数据爬虫,即AI爬虫。然后基于所述多维访问特征识别当前访问主体是否为目标爬虫,以能在识别出是AI爬虫时,执行后续的操作,以使AI爬虫爬取到目标网页中的关键信息。所以若识别出所述当前访问主体为目标爬虫,则识别出所述目标网页的各个页面实体,并从业务数据源中获取所述目标网页的各个页面实体的详情数据,生成结构化数据,从而通过识别页面实体,从业务数据源中获取到其详情数据,并生成结构化数据,便于AI爬虫可以爬取和理解结构化数据。最终将所述结构化数据注入所述目标网页的源码中,或返回所述结构化数据的数据接口,以使所述当前访问主体爬取到所述结构化数据,从而使得AI爬虫可以爬取到目标网页的页面实体的详情数据反馈给AI模型,进而可以保证其生成准确的结果。
Smart Images

Figure CN122817533A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data acquisition technology, and in particular to a data acquisition method and apparatus, electronic device, and storage medium. Background Technology
[0002] With the development of generative artificial intelligence (AI) question answering and AI search services, some users are beginning to obtain information directly through AI. Unlike traditional search engines that return webpage links for users to directly access and find information, generative AI models typically crawl webpage content and generate results directly based on that content, providing feedback to the user.
[0003] Therefore, current generative AI models mainly use web crawlers to extract information from one or more web pages related to user input, and then perform summarization, induction, and reorganization of the extracted web page information to generate results corresponding to user input and provide feedback to the user.
[0004] However, existing web pages are primarily designed for general users and traditional search engines, so the main information is often scattered across JavaScript, images, videos, asynchronous interfaces, business systems, and operational configurations. Web crawlers can only capture static text and traditional JSON-LD fields, which means AI models often fail to obtain crucial information such as video content, membership benefits, entity relationships, and update times, thus hindering the generation of accurate results for users. Summary of the Invention
[0005] In view of the shortcomings of the prior art, this application provides a data acquisition method and device, electronic device and storage medium to solve the problem that the prior art cannot obtain the key information of the webpage, resulting in the inability to generate accurate results.
[0006] To achieve the above objectives, this application provides the following technical solution:
[0007] The first aspect of this application provides a data acquisition method, including:
[0008] When an access request for a target webpage is received, multi-dimensional access features are extracted from the request information and behavior information of the access request.
[0009] Based on the multi-dimensional access features, it can be identified whether the current access subject is the target crawler;
[0010] If the current access subject is identified as the target crawler, then the various page entities of the target webpage are identified, and the detailed data of the various page entities of the target webpage are obtained from the business data source to generate structured data;
[0011] The structured data is injected into the source code of the target webpage, or the data interface of the structured data is returned, so that the current access subject can crawl the structured data.
[0012] Optionally, in the above data acquisition method, the step of extracting multi-dimensional access features from the request information and behavior information of the access request includes:
[0013] Multiple request information and multiple behavior information are collected from the information in the access request;
[0014] Extract one or more of the following from the multiple request information: request header features, network source features, and transport layer fingerprint features; and extract one or more of the following from the multiple behavior information: resource loading behavior features, access path sequence features, and historical access behavior features.
[0015] Optionally, in the above data acquisition method, the step of identifying whether the current access subject is a target crawler based on the multi-dimensional access features includes:
[0016] The multidimensional access features are used to generate access subject feature vectors;
[0017] The feature vector of the accessing entity is input into the rule engine to analyze whether the current accessing entity belongs to the target crawler;
[0018] If the analysis determines that the current access subject belongs to the target crawler, then the feature vector of the access subject is input into the crawler identification model, and the confidence score of the current access subject belonging to the target crawler is output.
[0019] If the confidence level that the current access subject belongs to the target crawler is greater than a preset threshold, then the current access subject is determined to be the target crawler.
[0020] Optionally, in the above data acquisition method, the step of identifying each page entity of the target webpage and obtaining detailed data of each page entity of the target webpage from the business data source to generate structured data includes:
[0021] Extract the identification criteria information from the information on the target webpage;
[0022] Based on the identification criteria, the various page entities of the target webpage are identified;
[0023] Obtain detailed data of each page entity of the target webpage from the business data source;
[0024] Based on the type, level, and structured data template of each page entity of the target webpage, the detailed data of each page entity of the target webpage is converted into structured data.
[0025] Optionally, in the above data collection method, before identifying each page entity of the target webpage, the method further includes:
[0026] Determine whether the pre-generated structured data of each page entity of the target webpage is cached;
[0027] If it is determined that the cache contains pre-generated structured data of each page entity of the target webpage, then the cached structured data is obtained, and the process of injecting the structured data into the source code of the target webpage is directly executed, or the data interface of the structured data is returned.
[0028] If it is determined that the structured data of each page entity of the target webpage is not cached beforehand, then the process of identifying each page entity of the target webpage is executed.
[0029] Optionally, in the above data collection method, injecting the structured data into the source code of the target webpage includes:
[0030] When rendering the target webpage, the structured data service is invoked to obtain the structured data, and the structured data is written to a specified location in the source code of the target webpage.
[0031] Optionally, in the above data acquisition method, the data interface that returns the structured data includes:
[0032] The data interface that provides feedback on the structured data in the response of the target webpage is through one of the following: response header, link tag, or agreed URL.
[0033] A second aspect of this application provides a data acquisition device, comprising:
[0034] The feature extraction unit is used to extract multi-dimensional access features from the request information and behavior information of the access request when receiving an access request for the target webpage.
[0035] A crawler identification unit is used to identify whether the current access subject is the target crawler based on the multi-dimensional access features.
[0036] The data generation unit is used to identify each page entity of the target webpage when the current access subject is identified as a target crawler, and to obtain detailed data of each page entity of the target webpage from the business data source to generate structured data.
[0037] An injection unit is used to inject the structured data into the source code of the target webpage, or to return the data interface of the structured data, so that the current access subject can crawl the structured data.
[0038] Optionally, in the above-described data acquisition device, the feature extraction unit includes:
[0039] An information collection unit is used to collect multiple request information and multiple behavior information from the information of the access request;
[0040] The multidimensional feature extraction unit is used to extract one or more of the following from the multiple request information: request header features, network source features, and transport layer fingerprint features; and to extract one or more of the following from the multiple behavioral information: resource loading behavior features, access path sequence features, and historical access behavior features.
[0041] Optionally, in the above-described data acquisition device, the crawler identification unit includes:
[0042] A vector generation unit is used to generate an access subject feature vector using the multidimensional access features;
[0043] The initial screening unit is used to input the feature vector of the accessing entity into the rule engine to analyze whether the current accessing entity belongs to the target crawler.
[0044] The model analysis unit is used to input the feature vector of the current access subject into the crawler identification model when it is determined that the current access subject belongs to the target crawler, and output the confidence score that the current access subject belongs to the target crawler.
[0045] The crawler determination unit is used to determine that the current access subject is the target crawler when the confidence level that the current access subject belongs to the target crawler is greater than a preset threshold.
[0046] Optionally, in the above-described data acquisition device, the data generation unit includes:
[0047] The extraction unit is used to extract various identification criteria information from the information of the target webpage;
[0048] An entity recognition unit is used to identify each page entity of the target webpage based on the recognition criteria information.
[0049] The detailed data acquisition unit is used to acquire detailed data of each page entity of the target webpage from the business data source;
[0050] The structured unit is used to convert the detailed data of each page entity of the target webpage into structured data based on the type, level and structured data template of each page entity of the target webpage.
[0051] Optionally, the data acquisition device described above further includes:
[0052] A cache determination unit is used to determine whether structured data of each page entity of the target webpage is cached in advance;
[0053] The cache acquisition unit is used to acquire the cached structured data when it is determined that there is pre-generated structured data of each page entity of the target webpage in the cache, and directly execute the injection unit.
[0054] If it is determined that the structured data of each page entity of the target webpage is not cached beforehand, the data generation unit is executed.
[0055] Optionally, in the data acquisition device described above, when the injection unit performs the action of injecting the structured data into the source code of the target webpage, it is used to:
[0056] When rendering the target webpage, the structured data service is invoked to obtain the structured data, and the structured data is written to a specified location in the source code of the target webpage.
[0057] Optionally, in the above-described data acquisition device, when the injection unit executes the data interface that returns the structured data, it is used to:
[0058] The data interface that provides feedback on the structured data in the response of the target webpage is through one of the following: response header, link tag, or agreed URL.
[0059] A third aspect of this application provides an electronic device, comprising:
[0060] Memory and processor;
[0061] The memory is used to store programs;
[0062] The processor is used to execute the program, which, when executed, is specifically used to implement the data acquisition method as described in any of the above.
[0063] A fourth aspect of this application provides a computer storage medium for storing a computer program, which, when executed by a processor, is used to implement the data acquisition method as described in any of the preceding claims.
[0064] This application provides a data acquisition method that, upon receiving an access request for a target webpage, extracts multi-dimensional access features from the request information and behavioral information of the access request. This allows for accurate identification of the target crawler based on the feature information of the request and behavioral information. The target crawler can be a data crawler based on an artificial intelligence model, i.e., an AI crawler. Then, based on the multi-dimensional access features, it identifies whether the current access subject is the target crawler. If it is identified as an AI crawler, subsequent operations are performed to enable the AI crawler to crawl key information from the target webpage. Therefore, if the current access subject is identified as the target crawler, the various page entities of the target webpage are identified, and detailed data of each page entity of the target webpage is obtained from the business data source to generate structured data. This allows the AI crawler to crawl and understand the structured data by identifying page entities and obtaining their detailed data from the business data source. Finally, the structured data is injected into the source code of the target webpage, or the data interface of the structured data is returned, so that the current access subject can crawl the structured data, thereby enabling the AI crawler to crawl the detailed data of the page entities of the target webpage and feed it back to the AI model, thus ensuring that it generates accurate results. Attached Figure Description
[0065] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0066] Figure 1 A flowchart illustrating a data acquisition method provided in this application embodiment;
[0067] Figure 2 A flowchart illustrating a method for extracting multidimensional access features provided in this application embodiment;
[0068] Figure 3 A flowchart illustrating a method for identifying AI crawlers provided in this application embodiment;
[0069] Figure 4 A flowchart illustrating a method for acquiring structured data provided in this application embodiment;
[0070] Figure 5 This is a schematic diagram of the architecture of a data acquisition device provided in an embodiment of this application;
[0071] Figure 6 This is a schematic diagram of the architecture of an electronic device provided in an embodiment of this application. Detailed Implementation
[0072] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0073] In this application, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0074] This application provides a data acquisition method, such as... Figure 1 As shown, it includes the following steps:
[0075] S101. When receiving an access request for the target webpage, extract multi-dimensional access features from the request information and behavior information of the access request.
[0076] The target webpage can be any webpage.
[0077] When any entity sends an access request to access a target webpage, the system receives the request. This request may originate from a regular browser, a traditional search engine crawler, or a generative artificial intelligence model crawler (AI crawler). For AI crawler requests, directly responding to their requests and allowing them to crawl data is insufficient to fully extract the crucial information needed by the AI; therefore, special processing is required. Thus, upon receiving an access request for a target webpage, it is necessary to first identify whether the entity sending the request is an AI crawler to determine whether further steps are required.
[0078] Current methods for identifying web crawlers primarily rely on user-agents, IP blacklists, and robots.txt access records. While relatively simple, these methods have low accuracy. User-agents can be spoofed, IP addresses can change, and many AI crawlers may not have document identifiers. Furthermore, current web crawler identification methods are mainly designed to prevent crawlers from scraping data. Therefore, they cannot guarantee the identification of AI crawlers.
[0079] To accurately identify AI web crawlers, this embodiment considers not only the request information of the access request, such as the basic access request information (e.g., user agent, IP address), but also the behavioral information of the access request to identify whether the accessing entity is an AI web crawler. Therefore, multi-dimensional access features are extracted from the request information and behavioral information of the access request.
[0080] Optionally, in another embodiment of this application, one specific implementation of step S101 is as follows: Figure 2 As shown, it includes:
[0081] S201. Collect multiple request information and multiple behavior information from the access request information.
[0082] Optionally, the following information can be collected from the access request, but is not limited to: request URL, request method, User-Agent, Accept, Accept-Language, Referer, Cookie, IP address, access time, TLS connection information, requested resource type, access path within the current session, whether static resources are loaded, whether front-end event tracking is triggered, and historical access records.
[0083] S202. Extract one or more of the following from multiple request information: request header features, network source features, and transport layer fingerprint features; and extract one or more of the following from multiple behavior information: resource loading behavior features, access path sequence features, and historical access behavior features.
[0084] To ensure accurate identification, features are typically extracted from multiple request information and multiple behavioral information. Therefore, the multidimensional features extracted in this application embodiment include: request header features, network source features, transport layer fingerprint features, resource loading behavior features, access path sequence features, and historical access behavior features.
[0085] Specifically, request header characteristics may include: User-Agent, Accept, Accept-Language, Referer, Cookie, etc.
[0086] Network origin characteristics can include: IP address, ASN, cloud service provider affiliation, region, IP reputation score, proxy characteristics, etc. Since AI crawlers may typically originate from data centers, cloud service providers, or specific network segments, network origin characteristics can be used to identify whether a crawler is an AI crawler.
[0087] Transport layer fingerprint features can include TLS version, Cipher Suite, JA3 / JA4 fingerprints, etc. Since different client libraries, web crawling frameworks, and browsers typically differ in their TLS handshake processes, this feature can be used to identify AI web crawlers.
[0088] Resource loading behavior characteristics can include whether CSS, JS, images, fonts, videos, and event tracking resources are loaded. Since ordinary browsers usually load the complete resources, while some crawlers only request HTML or a few interfaces, and AI crawlers need to crawl data such as images and videos, this characteristic can be used to identify AI crawlers.
[0089] Access path sequence characteristics may include: whether robots.txt, sitemap, details page, FAQ page, or structured data interface are accessed; whether multiple content pages are accessed in batches according to certain rules; and whether there is high-frequency deep traversal behavior.
[0090] Historical access behavior characteristics may include: the number of times the same IP, the same User-Agent, and the same TLS fingerprint are accessed within a certain time window, the distribution of page types, the access interval, the access failure rate, and the access depth.
[0091] S102. Identify whether the current access subject is the target crawler based on multi-dimensional access features.
[0092] Here, "target crawler" mainly refers to data crawlers based on artificial intelligence models, i.e., AI crawlers. If the current access subject is identified as a target crawler, appropriate processing is required, so step S103 is executed at this time. If the current access subject is identified as not a target crawler, a normal response can be performed directly.
[0093] Optionally, a rule engine can analyze multi-dimensional access characteristics based on set rules to determine whether the current accessing entity is the target crawler. Alternatively, a trained machine learning model can identify whether the current accessing entity is the target crawler based on multi-dimensional access characteristics. Another option is to use both a rule engine and a machine learning model simultaneously. Of course, other methods can also be used for identification.
[0094] Optionally, in another embodiment of this application, both a rule engine and a machine learning model are used for identification. Therefore, one specific implementation of step S102 is as follows: Figure 3 As shown, it includes:
[0095] S301. Generate access subject feature vector using multi-dimensional access features.
[0096] To analyze multi-dimensional access features using rule engines and machine learning models, these features are first converted into access subject feature vectors. For example, for the multi-dimensional access features mentioned above, the generated access subject feature vector is: V={H,N,T,R,P,B}. Here, H represents request header features; N represents network source features; T represents transport layer fingerprint features; R represents resource loading behavior features; P represents access path sequence features; and B represents historical access behavior features.
[0097] S302. Input the feature vector of the access subject into the rule engine to analyze whether the current access subject belongs to the target crawler.
[0098] It should be noted that in this embodiment of the application, a rule engine is first used for initial screening. Therefore, the feature vector of the access subject is first input into the rule engine so as to analyze whether the feature vector of the access subject matches the characteristics of the target crawler through the set rules, thereby analyzing whether the current access subject belongs to the target crawler.
[0099] If the analysis determines that the current access subject belongs to the target crawler, then step S303 is executed for further analysis.
[0100] S303. Input the feature vector of the accessing subject into the crawler identification model, and output the confidence level that the current accessing subject belongs to the target crawler.
[0101] S304. Determine whether the confidence level of the current access subject belonging to the target crawler is greater than the preset threshold.
[0102] If the confidence level that the current access subject belongs to the target crawler is greater than a preset threshold, then step S305 is executed.
[0103] S305. Determine that the current access subject is the target crawler.
[0104] Optionally, to not only execute subsequent steps on the target crawler (i.e., process it with appropriate strategies), but also to respond to access requests from traditional search engine crawlers and browsers with appropriate strategies, the confidence level that the current access subject belongs to the target crawler is not greater than the aforementioned preset threshold, but greater than the second threshold. This indicates that it is suspected to be an AI crawler, i.e., another type of crawler. In this case, it is only necessary to send basic data, such as links to the target network, normally. If the confidence level that the current access subject belongs to the target crawler is less than the second threshold, it means that it is a normal access subject, not a crawler. Therefore, the access request is responded to by directly returning the normal target page.
[0105] S103. Identify each page entity of the target webpage, obtain detailed data of each page entity of the target webpage from the business data source, and generate structured data.
[0106] Since the current access is handled by an AI crawler, in order to obtain key information from the target webpage, we first need to identify the various page entities contained within it. The specific page entities to be identified can be determined based on business requirements. For example, for long-form video services, the identified page entities may include, but are not limited to: drama series entities, variety show entities, movie entities, character entities, role entities, video clip entities, membership benefit entities, playback entry entities, event entities, and FAQ entities.
[0107] Then, it can access various business data sources, such as media asset systems, membership benefit systems, playback systems, search systems, recommendation systems, content operation systems, FAQ systems, and event configuration systems. Therefore, for each page entity, detailed data can be obtained from the corresponding business data sources. For example, for a TV series entity, the following detailed data can be obtained: title, alias, synopsis, actors, roles, number of episodes, category tags, release date, update date, playback entry point, membership restrictions, regional copyright, related Q&A, and related content recommendations.
[0108] Then, in order to facilitate the AI crawler to crawl and understand the acquired detailed data, the detailed data of each page entity is converted into unified structured data.
[0109] Optionally, if the page entity cannot be identified, the default structured data can be returned or no data injection can be performed.
[0110] Optionally, in another embodiment of this application, one specific implementation of step S103 is as follows: Figure 4 As shown, it includes:
[0111] S401. Extract the identification criteria information from the information on the target webpage.
[0112] In order to determine the page entities that need to be identified and to facilitate the identification of page entities, we first extract various identification criteria from the information of the target webpage. These criteria may include, but are not limited to: URL path, page template identifier, business object ID, page type, server-side rendering parameters, page tracking parameters, and object type returned by the backend interface.
[0113] S402. Based on the various identification criteria, identify the various page entities of the target webpage.
[0114] S403. Obtain detailed data of each page entity of the target webpage from the business data source.
[0115] S404. Based on the type, level, and structured data template of each page entity of the target webpage, convert the detailed data of each page entity of the target webpage into structured data.
[0116] Therefore, in this embodiment, structured data is dynamically generated from data obtained from various business data sources and templates, eliminating the need to manually maintain a fixed json-ld field on each page as in the prior art. This allows the AI crawler to promptly obtain the latest data from the target webpage.
[0117] Optionally, considering that some web pages have relatively stable data and high access volume, the page entities of such web pages can be identified in advance, and then their detailed data can be obtained, structured data can be generated and cached, so that the cached structured data can be directly returned for injection without real-time generation, thus allowing for rapid acquisition of structured data.
[0118] Therefore, in another embodiment of this application, before performing step S103, the following can be further performed:
[0119] Determine whether the structured data of each page entity of the pre-generated target webpage is cached.
[0120] If it is determined that there is structured data of each page entity of the pre-generated target webpage in the cache, then the cached structured data is obtained, and then step S104 is executed directly without executing step S103.
[0121] If it is determined that the structured data of each page entity of the pre-generated target webpage is not cached, then step S103 is executed.
[0122] S104. Inject structured data into the source code of the target webpage, or return a data interface for structured data, so that the current user can crawl the structured data.
[0123] To enable AI crawlers to extract the generated structured data, this data can be injected into the source code of the target webpage, allowing the AI crawler to extract it. Alternatively, a separate structured data interface can be provided, returning this interface to the AI crawler, which can then call the interface to extract the structured data. Finally, the AI crawler can provide the structured data containing details of each page entity on the target webpage to its AI model for analysis, generating accurate results to provide to the user.
[0124] Optionally, in another embodiment of this application, a specific implementation of injecting structured data into the source code of a target webpage includes:
[0125] When rendering the target webpage, the structured data service is called to obtain structured data, and the structured data is written to a specified location in the source code of the target webpage.
[0126] Specifically, when rendering the HTML of the target webpage, the page service calls the structured data service to obtain the structured data packets corresponding to each page entity of the current target page and writes them into the specified location in the HTML source code.
[0127] Optionally, in another embodiment of this application, a specific implementation of the data interface that returns structured data includes:
[0128] A data interface that provides structured data in the response of the target webpage via a response header, link tag, or a predefined URL.
[0129] Specifically, the system provides an independent structured data interface, which is exposed in the page response to the access request of the target page through response headers, Link tags, or conventional URLs.
[0130] This application provides a data acquisition method. When an access request for a target webpage is received, multi-dimensional access features are extracted from the request information and behavioral information of the access request. This allows for accurate identification of the target crawler based on the feature information of the request and behavioral information. The target crawler refers to a data crawler using an artificial intelligence model, i.e., an AI crawler. Then, based on the multi-dimensional access features, it is determined whether the current access subject is the target crawler. If it is identified as an AI crawler, subsequent operations are performed to enable the AI crawler to crawl key information from the target webpage. Therefore, if the current access subject is identified as the target crawler, the various page entities of the target webpage are identified, and detailed data of each page entity of the target webpage is obtained from the business data source to generate structured data. This allows the AI crawler to crawl and understand the structured data by identifying the page entities and obtaining their detailed data from the business data source. Finally, the structured data is injected into the source code of the target webpage, or the data interface of the structured data is returned, so that the current access subject can crawl the structured data, thereby enabling the AI crawler to crawl the detailed data of the page entities of the target webpage and feed it back to the AI model, thus ensuring that it generates accurate results.
[0131] Another embodiment of this application provides a data acquisition device, such as... Figure 5 As shown, it includes:
[0132] The feature extraction unit 501 is used to extract multi-dimensional access features from the request information and behavior information of the access request when receiving an access request for the target webpage.
[0133] The crawler identification unit 502 is used to identify whether the current access subject is the target crawler based on multi-dimensional access features.
[0134] The data generation unit 503 is used to identify each page entity of the target webpage when the current access subject is identified as the target crawler, and to obtain the detailed data of each page entity of the target webpage from the business data source to generate structured data.
[0135] Injection unit 504 is used to inject structured data into the source code of the target webpage or return a data interface for structured data so that the current access subject can crawl the structured data.
[0136] Optionally, in another embodiment of the data acquisition device provided in this application, the feature extraction unit includes:
[0137] The information collection unit is used to collect multiple request information and multiple behavioral information from the access request information.
[0138] The multidimensional feature extraction unit is used to extract one or more of the following from multiple request information: request header features, network source features, and transport layer fingerprint features; and to extract one or more of the following from multiple behavioral information: resource loading behavior features, access path sequence features, and historical access behavior features.
[0139] Optionally, in another embodiment of the data acquisition device provided in this application, the crawler identification unit includes:
[0140] The vector generation unit is used to generate access subject feature vectors using multidimensional access features.
[0141] The initial screening unit is used to input the feature vector of the access subject into the rule engine to analyze whether the current access subject belongs to the target crawler.
[0142] The model analysis unit is used to input the feature vector of the current access subject into the crawler identification model when it is determined that the current access subject belongs to the target crawler, and output the confidence score that the current access subject belongs to the target crawler.
[0143] The crawler determination unit is used to determine that the current access subject is the target crawler when the confidence level that the current access subject belongs to the target crawler is greater than a preset threshold.
[0144] Optionally, in another embodiment of the data acquisition device provided in this application, the data generation unit includes:
[0145] The extraction unit is used to extract various identification criteria information from the information on the target webpage.
[0146] The entity recognition unit is used to identify the various page entities of the target webpage based on various recognition criteria.
[0147] The detailed data acquisition unit is used to obtain detailed data of each page entity of the target webpage from the business data source.
[0148] Structured units are used to convert the detailed data of each page entity of a target webpage into structured data based on the type, level, and structured data template of each page entity.
[0149] Optionally, in another embodiment of the data acquisition device provided in this application, the device further includes:
[0150] The cache determination unit is used to determine whether the structured data of each page entity of the pre-generated target webpage is cached.
[0151] The cache retrieval unit is used to retrieve the cached structured data when it is determined that there is structured data of each page entity of the pre-generated target webpage in the cache, and directly execute the injection unit.
[0152] If it is determined that the structured data of each page entity of the pre-generated target webpage is not cached, then the data generation unit is executed.
[0153] Optionally, in another embodiment of the data acquisition device provided in this application, when the injection unit performs the function of injecting structured data into the source code of the target webpage, it is used to:
[0154] When rendering the target webpage, the structured data service is called to obtain structured data, and the structured data is written to a specified location in the source code of the target webpage.
[0155] Optionally, in another embodiment of the data acquisition device provided in this application, when the injection unit executes the data interface that returns structured data, it is used to:
[0156] A data interface that provides structured data in the response of the target webpage via a response header, link tag, or a predefined URL.
[0157] It should be noted that the specific working process of each unit provided in the above embodiments of this application can be referred to the implementation process of the corresponding steps in the above method embodiments, and will not be repeated here.
[0158] Another embodiment of this application provides an electronic device, such as... Figure 6 As shown, it includes:
[0159] Memory 601 and processor 602.
[0160] The memory 601 is used to store the program.
[0161] The processor 602 is used to execute the program stored in the memory 601. When the program is executed, it is specifically used to implement the data acquisition method provided in any of the above embodiments.
[0162] Another embodiment of this application provides a computer storage medium for storing a computer program, which, when executed by a processor, is used to implement the data acquisition method provided in any of the above embodiments.
[0163] Computer storage media, including both permanent and non-permanent, removable and non-removable media, can store information using any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0164] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0165] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A data acquisition method, characterized in that, include: When an access request for a target webpage is received, multi-dimensional access features are extracted from the request information and behavior information of the access request. Based on the multi-dimensional access features, it can be identified whether the current access subject is the target crawler; If the current access subject is identified as the target crawler, then the various page entities of the target webpage are identified, and the detailed data of the various page entities of the target webpage are obtained from the business data source to generate structured data; The structured data is injected into the source code of the target webpage, or the data interface of the structured data is returned, so that the current access subject can crawl the structured data.
2. The method according to claim 1, characterized in that, The extraction of multi-dimensional access features from the request information and behavior information of the access request includes: Multiple request information and multiple behavior information are collected from the information in the access request; Extract one or more of the following from the multiple request information: request header features, network source features, and transport layer fingerprint features; and extract one or more of the following from the multiple behavior information: resource loading behavior features, access path sequence features, and historical access behavior features.
3. The method according to claim 1, characterized in that, The step of identifying whether the current access subject is the target crawler based on the multi-dimensional access features includes: The multidimensional access features are used to generate access subject feature vectors; The feature vector of the accessing entity is input into the rule engine to analyze whether the current accessing entity belongs to the target crawler; If the analysis determines that the current access subject belongs to the target crawler, then the feature vector of the access subject is input into the crawler identification model, and the confidence score of the current access subject belonging to the target crawler is output. If the confidence level that the current access subject belongs to the target crawler is greater than a preset threshold, then the current access subject is determined to be the target crawler.
4. The method according to claim 1, characterized in that, The process of identifying each page entity of the target webpage and obtaining detailed data of each page entity of the target webpage from the business data source to generate structured data includes: Extract the identification criteria information from the information of the target webpage; Based on the identification criteria, the various page entities of the target webpage are identified; Obtain detailed data of each page entity of the target webpage from the business data source; Based on the type, level, and structured data template of each page entity of the target webpage, the detailed data of each page entity of the target webpage is converted into structured data.
5. The method according to claim 1, characterized in that, Before identifying the various page entities of the target webpage, the method further includes: Determine whether the pre-generated structured data of each page entity of the target webpage is cached; If it is determined that the cache contains pre-generated structured data of each page entity of the target webpage, then the cached structured data is obtained, and the process of injecting the structured data into the source code of the target webpage is directly executed, or the data interface of the structured data is returned. If it is determined that the structured data of each page entity of the target webpage is not cached beforehand, then the process of identifying each page entity of the target webpage is executed.
6. The method according to claim 1, characterized in that, The step of injecting the structured data into the source code of the target webpage includes: When rendering the target webpage, the structured data service is invoked to obtain the structured data, and the structured data is written to a specified location in the source code of the target webpage.
7. The method according to claim 1, characterized in that, The data interface that returns the structured data includes: The data interface that provides feedback on the structured data in the response of the target webpage is through one of the following: response header, link tag, or agreed URL.
8. A data acquisition device, characterized in that, include: The feature extraction unit is used to extract multi-dimensional access features from the request information and behavior information of the access request when receiving an access request for the target webpage. A crawler identification unit is used to identify whether the current access subject is the target crawler based on the multi-dimensional access features. The data generation unit is used to identify each page entity of the target webpage when the current access subject is identified as a target crawler, and to obtain detailed data of each page entity of the target webpage from the business data source to generate structured data. An injection unit is used to inject the structured data into the source code of the target webpage, or to return the data interface of the structured data, so that the current access subject can crawl the structured data.
9. An electronic device, characterized in that, include: Memory and processor; The memory is used to store programs; The processor is used to execute the program, which, when executed, is specifically used to implement the data acquisition method as described in any one of claims 1 to 7.
10. A computer storage medium, characterized in that, Used to store a computer program, which, when executed by a processor, is used to implement the data acquisition method as described in any one of claims 1 to 7.