Information acquisition method and device, electronic equipment and storage medium

By obtaining page format information and using the information collection rules in the rule engine, the problem of incomplete collection of page information in the mobile application is solved, and comprehensive and accurate collection of page information is achieved.

CN120469609APending Publication Date: 2025-08-12GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510593259.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-08-12

Smart Images

  • Figure CN120469609A_ABST
    Figure CN120469609A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an information collection method and device, electronic equipment and a storage medium. The method comprises the steps of obtaining format information of a to-be-collected page; obtaining at least one information collection rule corresponding to the format information of the page to be collected from a rule engine; and based on the at least one information collection rule, obtaining page information of the to-be-collected page. By means of the method, at least one information collection rule corresponding to the to-be-collected page can be accurately obtained through the format information of the to-be-collected page, and then the page information of the to-be-collected page can be comprehensively collected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of information collection technology, and specifically relates to an information collection method, device, electronic device and storage medium. Background Art

[0002] With the increasing popularity of mobile smart devices (such as smartphones and tablets), a growing number of mobile terminal applications, often referred to as mobile apps (APPs), have emerged. Mobile app page information is a crucial piece of data, and many processes require it. For example, international e-commerce apps need to detect untranslated content based on the page's content. However, related technologies often fail to fully extract this information when collecting it. Summary of the Invention

[0003] In view of the above problems, the present application proposes an information collection method, device, electronic device and storage medium to improve the above problems.

[0004] In a first aspect, an embodiment of the present application provides an information collection method, which includes: obtaining format information of a page to be collected; obtaining at least one information collection rule corresponding to the format information of the page to be collected from a rule engine; and obtaining page information of the page to be collected based on the at least one information collection rule.

[0005] In second aspect, an embodiment of the present application provides an information collection device, which includes: a format acquisition unit for acquiring format information of a page to be collected; a rule acquisition unit for acquiring at least one information collection rule corresponding to the format information of the page to be collected from a rule engine; and an information acquisition unit for acquiring page information of the page to be collected based on the at least one information collection rule.

[0006] In a third aspect, an embodiment of the present application provides an electronic device comprising one or more processors and a memory; one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs are configured to execute the above-mentioned method.

[0007] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which program code is stored, wherein the above method is executed when the program code is run.

[0008] In a fifth aspect, an embodiment of the present application provides a computer program product, comprising a computer program / instruction, which implements the steps of the above method when executed by a processor.

[0009] Embodiments of the present application provide an information collection method, apparatus, electronic device, and storage medium. First, the format information of a page to be collected is obtained. Then, at least one information collection rule corresponding to the format information of the page to be collected is obtained from a rule engine. Based on the at least one information collection rule, page information of the page to be collected is obtained. Through the above method, the at least one information collection rule corresponding to the page to be collected can be accurately obtained based on the format information of the page to be collected, thereby enabling comprehensive collection of page information of the page to be collected. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0011] Figure 1 A schematic diagram showing an application scenario of an information collection method proposed in one embodiment of the present application is shown;

[0012] Figure 2 A schematic diagram showing an application scenario of an information collection method proposed in one embodiment of the present application is shown;

[0013] Figure 3 The following is an architecture diagram of a fusion acquisition framework proposed in an embodiment of the present application;

[0014] Figure 4 The following is an architecture diagram of a picture and text collection framework proposed in one embodiment of the present application;

[0015] Figure 5 The following is an architectural diagram of a page parsing framework proposed in one embodiment of the present application;

[0016] Figure 6 A flowchart of an information collection method proposed in one embodiment of the present application is shown;

[0017] Figure 7 A flowchart of an information collection method proposed in another embodiment of the present application is shown;

[0018] Figure 8 A schematic diagram of key triggering in another embodiment of the present application is shown;

[0019] Figure 9 A schematic diagram showing a three-finger upward swipe trigger in another embodiment of the present application is shown;

[0020] Figure 10 A schematic diagram of control triggering in another embodiment of the present application is shown;

[0021] Figure 11 A flowchart of an information collection method proposed in another embodiment of the present application is shown;

[0022] Figure 12 A flowchart of an information collection method proposed in another embodiment of the present application is shown;

[0023] Figure 13 A schematic diagram of a desktop card in another embodiment of the present application is shown;

[0024] Figure 14 The following is a structural block diagram of an information collection device proposed in an embodiment of the present application;

[0025] Figure 15 The following is a structural block diagram of an information collection device proposed in an embodiment of the present application;

[0026] Figure 16 A structural block diagram of an electronic device or server for executing the information collection method according to an embodiment of the present application is shown in real time in the present application;

[0027] Figure 17 The present invention shows a storage unit in real time for storing or carrying program codes for implementing the information collection method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0028] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0029] The inventors discovered during their research on related information collection methods that the related information collection methods are still unable to fully collect all the page information when collecting page information.

[0030] Therefore, the inventors have proposed the information collection method, apparatus, electronic device, and storage medium of this application. First, the format information of the page to be collected is obtained. Then, at least one information collection rule corresponding to the format information of the page to be collected is obtained from a rule engine. Based on the at least one information collection rule, the page information of the page to be collected is obtained. Through the above method, the at least one information collection rule corresponding to the page to be collected can be accurately obtained based on the format information of the page to be collected, thereby enabling comprehensive collection of the page information of the page to be collected.

[0031] In the embodiment of the present application, the information collection method provided can be executed by an electronic device. In this way, all steps in the information collection method provided by the embodiment of the present application can be executed by the electronic device. For example, Figure 1 As shown, the format information of the page to be collected can be obtained through the format acquisition device of the electronic device 100, and then the obtained format information of the page to be collected is transmitted to the processor, so that the processor can obtain at least one information collection rule corresponding to the format information of the page to be collected from the rule engine, and thus can obtain the page information of the page to be collected based on at least one information collection rule.

[0032] Furthermore, the information collection method provided in the embodiments of the present application can also be executed by a server (cloud). Accordingly, in this server-based execution method, the electronic device can obtain the format information of the page to be collected and synchronously send the format information of the page to be collected to the server. The server then obtains at least one information collection rule corresponding to the format information of the page to be collected from the rule engine in real time, thereby obtaining the page information of the page to be collected based on the at least one information collection rule.

[0033] In addition, the electronic device and the server may collaborate to perform the information collection method. In this manner, some steps of the information collection method provided in the embodiment of the present application are performed by the electronic device, while other steps are performed by the server.

[0034] For example, Figure 2 As shown, the electronic device 100 can execute the information collection method including: obtaining the format information of the page to be collected, and then the server 200 executes at least one information collection rule corresponding to the format information of the page to be collected from the rule engine, and then the electronic device 100 executes based on the at least one information collection rule to obtain the page information of the page to be collected.

[0035] It should be noted that in this method of collaborative execution by the electronic device and the server, the steps respectively executed by the electronic device and the server are not limited to the methods introduced in the above examples. In actual applications, the steps respectively executed by the electronic device and the server can be dynamically adjusted according to actual conditions.

[0036] It should be noted that the electronic device 100 can be used for Figure 1 and Figure 2 In addition to the smartphone shown in FIG, it can also be a car device, a wearable device, a tablet computer, a laptop computer, a smart speaker, etc. The server 120 can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers.

[0037] like Figure 3 As shown, the information collection method provided in the embodiment of the present application can be applied to Figure 3 In the fusion acquisition framework 300 shown, the fusion acquisition framework 300 may include a picture and text extraction framework, a page parsing framework, AISDK (Application Intelligence SDK, a toolkit for artificial intelligence applications) and a cloud crawler. Among them, the image and text extraction framework is used to provide page information extraction capabilities, which can specifically include information collection rules for extracting page information from native, wbview, mini-programs, and quick applications. The page parsing framework is used to extract page information by instrumenting the application process, which can mainly include information collection rules for collecting the context environment (i.e., mContext), file path, large image path, and jump link (i.e., deeplink) in the page. AISDK is used to provide OCR (Optical Character Recognition), NER (Named Entity Recognition), and multimodal model capabilities, mainly including information collection rules for extracting information from self-rendering controls, screenshots, etc. in the page. The cloud crawler is used to crawl page image and text information from the web page through the page URL (Uniform Resource Locator), mainly including information collection rules for crawling page image and text information from the web page through the URL. This fusion collection framework can be understood as a constructed rule engine.

[0038] In such Figure 3 In the integrated collection framework shown, triggering refers to the method of triggering the information recording instruction, which can include page switching, control click triggering, gesture triggering, and voice command triggering. When the information recording instruction is triggered by any of these methods, the various information collection rules configured in the rule policy are used to collect page information, and then the collected page information is classified into specific business scenarios (for example, articles, orders, chats, travel, general, etc.).

[0039] In the embodiment of the present application, the overall logical view of the image and text extraction framework can be as follows: Figure 4 As shown, the image-text extraction framework may correspond to an image-text extraction framework interface, which may be used to call information collection rules in the image-text extraction framework.

[0040] The overall logical view of the page parsing framework can be as follows Figure 5 As shown, Figure 5The Rule Module in the framework is responsible for interacting with the cloud, pulling the latest page rule information, and storing it in the local database; the Client AI Module is responsible for interacting with the caller and the Framework layer, serving as the node for starting and ending the parsing process, and completing the role of connecting the upper and lower levels; the AIInterface Module is responsible for interacting with the AI platform and performing AI processing on the parsing results.

[0041] The embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0042] See also Figure 6 , an information collection method provided by the embodiment of the present application is applied to Figure 1 or Figure 2 The electronic device or server shown, the method includes:

[0043] Step S110: Acquire the format information of the page to be collected.

[0044] In an embodiment of the present application, the page to be collected may be a page that is determined to require information collection. The page to be collected may be a page that has been displayed, is currently being displayed, or is about to be displayed on the screen of an electronic device, or it may be a page that has been displayed, is currently being displayed, or is about to be displayed on the screen of another electronic device that has established a communication connection with the electronic device, and is not specifically limited here. Format information can be understood as information that characterizes the page format of the page to be collected. Format information refers to a collection of parameters that standardizes and controls the content presentation form through technical means, and may mainly include the following dimensions: structure definition, style control, interactive and dynamic functions, metadata and protocol specifications, file format compatibility, etc. In an embodiment of the present application, the format information corresponding to different types of pages to be collected may be different. The type of page to be collected may refer to the type of application to which the page to be collected belongs. Different types of applications (for example, native, webview, mini-programs, or quick applications, etc.) have different format information corresponding to the page to be collected.

[0045] In the embodiment of the present application, the pages to be collected can be determined in a variety of ways.

[0046] As one approach, the application page currently running in the electronic device can be determined as the page to be collected. When it is determined that page information collection is required, the application page currently running in the electronic device can be determined as the page to be collected. For example, when it is determined that page information collection is required, if the application page currently running in the electronic device is a chat page, then the chat page is determined as the page to be collected.

[0047] As another approach, when a designated operation is detected on the screen of the electronic device, the page displayed on the screen of the electronic device when the designated operation is performed is used as the page to be collected. The designated operation can be a pre-set operation for triggering information collection, for example, the designated operation can be a sliding operation, a clicking operation, etc. For example, if the designated operation is a double-click operation, then when a double-click operation is detected on the screen of the electronic device, the page displayed on the screen of the electronic device when the double-click operation is detected can be used as the page to be collected.

[0048] Optionally, when a designated page switch is detected, the page before the switch and / or the page after the switch are used as pages to be collected. A designated page is a pre-set page that can trigger page information collection. In an embodiment of the present application, when a switch from a first designated page to a second designated page is detected, the first designated page and / or the second designated page are used as pages to be collected.

[0049] Optionally, the type of application currently running on the electronic device is detected. If the application is an application of a preset type, the page currently displayed on the screen is obtained as the page to be collected. The preset type is a pre-set application type that can trigger page information collection. When it is detected that an application is running in the electronic device, the type of the running application is obtained. If the type of the running application is a preset type, the page displayed on the current screen is used as the page to be collected.

[0050] Optionally, when a preset operation is detected on a specified control on a page, the page containing the specified control is selected as the page to be collected. The specified control may be a pre-set virtual control that can trigger page information collection, such as a favorites button; the preset operation may be a pre-set operation for activating the function controlled by the virtual control, such as a click operation. For example, when a click operation is detected on a favorites button on a page, the page containing the favorites button may be selected as the page to be collected.

[0051] Optionally, the above-mentioned multiple methods for determining the pages to be collected may be combined to obtain a new method for determining the pages to be collected, which will not be described in detail here.

[0052] After the page to be collected is determined through the above-mentioned various methods, the format information of the page to be collected can be obtained. The format information of all pages in the electronic device can be pre-stored in a preset storage area. After the page to be collected is determined, the format information of the page to be collected can be obtained from the preset storage area. Optionally, in an embodiment of the present application, the page to be collected can include at least one page.

[0053] Step S120: obtaining at least one information collection rule corresponding to the format information of the page to be collected from the rule engine.

[0054] The information collection rules can be understood as pre-set rules for information collection. The at least one information collection rule corresponding to the page to be collected can be a pre-configured information collection rule that can comprehensively and quickly collect the information included in the page to be collected.

[0055] Optionally, in an embodiment of the present application, different format information may correspond to different information collection rules. The rule engine includes a variety of pre-set information collection rules, which may include at least text collection rules, image collection rules, file path collection rules, self-rendering control collection rules, OCR collection rules, cloud crawler collection rules, page link collection rules, and large image path collection rules; wherein, the text collection rule is used to obtain the text information of the page to be collected; the image collection rule is used to obtain the image information of the page to be collected; the file path collection rule is used to obtain the storage path of the file contained in the page to be collected; the self-rendering control collection rule is used to obtain the control information of the control that loads text in the self-rendering mode in the page to be collected; the OCR collection rule is used to obtain the text information of the picture in the page to be collected; the cloud crawler collection rule is used to obtain the graphic and text information of the page to be collected through the page link; the page link collection rule is used to obtain the jump link of the page to be collected; the large image path collection rule is used to obtain the large image path of the page to be collected.

[0056] In the embodiment of the present application, the text collection rules and the image collection rules are information collection rules in the image and text extraction framework. The text collection rules are used to collect text information in the page to be collected, and the image collection rules are used to collect image information in the page to be collected. Among them, text information can be understood as information such as text, text size, font, color, and the position of the text in the page to be collected, which is not specifically limited here. Image information can include pictures, picture size, picture color, the position of the picture in the page to be collected, object information included in the picture, etc., which is not specifically limited here.

[0057] It is known that some native controls use self-rendering to load text, which makes it impossible to collect information through text collection rules and image collection rules. Therefore, special information collection rules can be set for these self-rendering controls to achieve information collection for self-rendering controls.

[0058] Furthermore, for pages whose page information cannot be collected through the aforementioned information collection rules, corresponding OCR collection rules can be configured. The OCR collection rules are used to convert the printed or handwritten text in the to-be-collected page into editable text through image processing and pattern recognition technology. Its core tasks include text positioning, segmentation, recognition and structured output.

[0059] Alternatively, for some pages with special anti-theft mechanisms or specially designed image browsing frameworks, directly obtained page information may be incomplete. For such pages, you can adopt cloud crawler collection rules to crawl page image and text information from the web page according to the page URL to obtain complete page information for the page to be collected.

[0060] The page link collection rule is used to collect the page URL / deeplink of the page to be collected to obtain the page jump link and realize the page jump.

[0061] Page link collection rules typically share page URLs and deeplinks, such as https: / / www.xiaohongshu.com / discovery / item / xxx and xhsdiscover: / / item / xxx. Specifically, the page parsing framework extracts the page content ID by combining a fixed prefix with the page content ID. On the cloud side, link collection rules for the corresponding page are configured. After concatenating the page content ID with the link pre-set in the link collection rule, the page's redirect link is obtained, enabling page redirection. The page parsing framework extracts page information through application process instrumentation.

[0062] The file path collection rule is used to collect the storage paths of files included in the collection page. Specifically, you may sometimes receive a file sent by someone in a chat tool. In this case, after opening it, a copy will be downloaded to the local file manager. The page parsing framework can be used to extract the local path of the file.

[0063] Large image path collection rules are used to collect large image paths for pages to be collected. Large image paths refer to the storage location and access rules for high-resolution image resources within a project. For order pages that require collection of product covers and verification codes, large image collection rules can be configured on the cloud side, and the corresponding large image paths can be extracted through the page parsing framework.

[0064] Specifically, the corresponding page-side collection rules can be configured on the cloud side according to the different ways of page implementation. The different ways of page implementation can be understood as pages belonging to different types of applications. In the embodiment of the present application, the types of applications can be roughly divided into native, wbview, mini-programs and quick applications.

[0065] Step S130: Based on the at least one information collection rule, obtain page information of the page to be collected.

[0066] In an embodiment of the present application, the page information of the page to be collected may include text, pictures, jump links, file storage paths, large image paths, screenshots, controls, page layout information, etc. in the page to be collected, and no specific limitation is made here.

[0067] After determining at least one information collection rule corresponding to the page to be collected, the page information of the page to be collected can be fully acquired through the at least one information collection rule.

[0068] The present application provides an information collection method that can accurately obtain at least one information collection rule corresponding to a page to be collected through the format information of the page to be collected, thereby comprehensively collecting page information of the page to be collected.

[0069] See also Figure 7 , an information collection method provided by the embodiment of the present application is applied to Figure 1 or Figure 2 The electronic device or server shown, the method includes:

[0070] Step S210: Acquire the page identifier of the page to be collected, where the page identifier corresponds to the format information of the page to be collected.

[0071] In the embodiment of the present application, the page identifier can be understood as an identifier that uniquely identifies a page and the format information of the page. One page can correspond to one page identifier, and one page identifier corresponds to one format information.

[0072] A correspondence between the page identifier and the format information may be established in advance. After the page identifier of the page to be collected is obtained, the format information of the page to be collected may be determined based on the correspondence between the page identifier and the format information.

[0073] Step S220: Based on the page identifier, obtain at least one information collection rule corresponding to the page to be collected from a rule engine.

[0074] As one approach, based on the page identifier, determine whether the page to be collected is in a page whitelist; if so, obtain at least one information collection rule corresponding to the page to be collected from a rule engine based on the page identifier.

[0075] In an embodiment of the present application, a page whitelist may include a plurality of pre-set legal pages that a user is allowed to access. The page whitelist may store page identifiers of legal pages that the user is allowed to access, and then, when determining whether a page to be collected is in the page whitelist, the page identifier of the page to be collected may be used to determine whether the page to be collected is in the page whitelist. Specifically, if the page identifier of the page to be collected can be found in the page whitelist, it can be determined that the page to be collected is in the page whitelist. Otherwise, or if the page identifier of the page to be collected is not found in the page whitelist, it can be determined that the page to be collected is not in the page whitelist.

[0076] The corresponding relationship between page identifiers and information collection rules is stored in the rule engine.

[0077] After determining that the page to be collected is in the whitelist, at least one information collection rule corresponding to the page to be collected can be obtained from the rule engine based on the correspondence between the page identifier and the information collection rule.

[0078] Step S230: In response to the information recording instruction, based on the at least one information collection rule, obtain the page information of the page to be collected.

[0079] In an embodiment of the present application, the information recording instruction is a triggered instruction for automatically collecting, preprocessing, and recording the page information of the page to be collected. Among them, collecting refers to collecting page information; preprocessing refers to extracting target information from the page information; and recording refers to writing the extracted target information into a preset storage area. In an embodiment of the present application, no matter what application scenario, as long as the information recording instruction is triggered, the page information of the page to be collected will be automatically obtained. Among them, the application scenario can be understood as the current operating scenario of the electronic device, such as a call scenario, a navigation scenario, a chat dialogue scenario, an information browsing scenario, etc., which is not specifically limited here.

[0080] The triggering mode of the information recording instruction includes one of key triggering, gesture triggering, voice triggering, page switching triggering and control triggering.

[0081] Key triggering means that an electronic device may be provided with a key (which may be a physical key) specifically for triggering the information recording function. When it is detected that the user switches the key from the first state to the second state, it is determined that the information recording instruction has been triggered. The first state is the initial state of the key, that is, the state in which the information recording function cannot be triggered, and the second state is the state after switching from the initial state, that is, the state in which the information recording function can be triggered. Optionally, in order to facilitate the control of the information recording function, the key in the embodiment of the present application may only include the first state and the second state. For example, Figure 8 As shown, Figure 8A button 110 for controlling the information recording function is provided on the right side of the electronic device. When it is detected that the button 110 is in a pressed state, it can be determined that the information recording instruction is triggered.

[0082] Gesture triggering means that when a designated operation acting on the screen of an electronic device is detected, it is determined that an information recording instruction has been triggered. Among them, the designated operation is a pre-set gesture operation that can trigger an information recording instruction. For example, the designated operation can be a multi-finger swipe up operation (such as a three-finger swipe up), a multi-finger swipe left / right operation, etc., which are not specifically limited here; the designated operation acting on the screen of an electronic device can refer to a designated operation in which the finger touches the screen of the electronic device, or a designated operation within a preset distance range above the screen of the electronic device without touching the screen of the electronic device. For example, Figure 9 As shown, when the gesture trigger is set to when a three-finger swipe-up operation is detected on the screen of the electronic device, it is determined that the information recording instruction is triggered.

[0083] Voice triggering means that when a voice control command containing a preset keyword is detected, it can be determined that a message recording instruction has been triggered. The preset keyword is a pre-set keyword that can trigger the message recording function, for example, the preset keyword can be set to "message recording" or "open message recording", etc., without specific limitation here.

[0084] Page switching trigger means that when the switching of the designated page is detected, it can be determined that the information recording instruction is triggered. The designated page can be a pre-set page that can trigger the information recording function, for example, the designated page can be set to a page containing private information.

[0085] Control triggering means that when an operation is detected on a designated control displayed on the screen, it can be determined that an information recording instruction has been triggered. The designated control can be a pre-set virtual control that can trigger the information recording function. For example, Figure 10 As shown, the specified control can be set as a favorites button. Figure 10 In the example embodiment, when a click operation on the favorites button 120 is detected, it can be determined that an information recording instruction is triggered.

[0086] In the embodiments of the present application, the triggering of the above-mentioned information recording instruction can be monitored in various ways. For example, page switching triggering and control triggering can be achieved through control instrumentation technology, accessibility, AppSwitch, and other technologies; gesture triggering can be monitored by a touch sensor provided in the electronic device; voice triggering can be monitored by a voice recognition module provided in the electronic device; key triggering can be monitored by a listener provided in the electronic device, etc., without further specific limitations here.

[0087] When it is detected that the information recording instruction is triggered by any of the above methods, comprehensive and rapid information collection is performed on the page to be collected according to at least one determined information collection rule to obtain page information of the page to be collected.

[0088] The present application provides an information collection method that can accurately obtain information collection rules corresponding to the page to be collected through page identification, and then can comprehensively collect page information of the page to be collected.

[0089] See also Figure 11 , an information collection method provided by the embodiment of the present application is applied to Figure 1 or Figure 2 The electronic device or server shown, the method includes:

[0090] Step S310: Acquire the service window name of the page to be collected, where the service window name corresponds to the format information of the page to be collected.

[0091] In this embodiment of the present application, the activity name, also called the ActivityName, typically refers to the name of each tab opened in a browser. The activity name can help distinguish different pages or applications. Each page has only one unique activity name. Each activity name corresponds to a specific format information. Furthermore, the activity name of a page can be associated with the page's format information.

[0092] Specifically, a correspondence between the business window name and the format information may be pre-established. After obtaining the business window name of the page to be collected, the format information of the page to be collected may be determined based on the correspondence between the business window name and the format information.

[0093] It is known that the business window name is mainly displayed in the title bar area at the top of the window. Therefore, when it is necessary to obtain the business window name of the page to be collected, the business window name of the page to be collected can be obtained from the title bar area.

[0094] Step S320: Based on the business window name, obtain at least one information collection rule corresponding to the page to be collected from a rule engine.

[0095] As one approach, based on the business window name, determine whether the page to be collected is in a page whitelist; if so, obtain at least one information collection rule corresponding to the page to be collected from a rule engine based on the business window name.

[0096] In an embodiment of the present application, a page whitelist may include a plurality of pre-set legal pages that a user can access. The page whitelist may store the business window names of the legal pages that the user is allowed to access, and then, when determining whether the page to be collected is in the page whitelist, the business window name of the page to be collected may be used to determine whether the page to be collected is in the page whitelist. Specifically, if the business window name of the page to be collected can be found in the page whitelist, it can be determined that the page to be collected is in the page whitelist; otherwise, or if the business window name of the page to be collected is not found in the page whitelist, it can be determined that the page to be collected is not in the page whitelist.

[0097] The correspondence between business window names and information collection rules is stored in the rule engine.

[0098] After determining that the page to be collected is in the whitelist, at least one information collection rule corresponding to the page to be collected can be obtained from the rule engine based on the correspondence between the business window name and the information collection rule.

[0099] Step S330: Based on the at least one information collection rule, obtain the page information of the page to be collected.

[0100] The present application provides an information collection method that can accurately obtain the information collection rules corresponding to the page to be collected through the business window name, and then comprehensively collect the page information of the page to be collected.

[0101] See also Figure 12 , an information collection method provided by the embodiment of the present application is applied to Figure 1 or Figure 2 The electronic device or server shown, the method includes:

[0102] Step S410: obtaining the page type of the page to be collected, where the page type corresponds to the format information of the page to be collected.

[0103] In an embodiment of the present application, one page type may correspond to one format information, and a correspondence between the page type and the format information may be established in advance. After obtaining the page type of the page to be collected, the format information of the page to be collected may be determined based on the correspondence between the page type and the format information.

[0104] Step S420: Based on the page type, obtain at least one information collection rule corresponding to the page to be collected from a rule engine.

[0105] As one approach, based on the page type, determine whether the page to be collected is in a page whitelist; if so, obtain at least one information collection rule corresponding to the page to be collected from a rule engine based on the page type.

[0106] In an embodiment of the present application, a page whitelist may include a plurality of pre-set legal pages that a user is allowed to access. The page whitelist may store the page types of legal pages that the user is allowed to access, and then, when determining whether a page to be collected is in the page whitelist, the page type of the page to be collected may be used to determine whether the page to be collected is in the page whitelist. Specifically, if the page type of the page to be collected can be found in the page whitelist, then it can be determined that the page to be collected is in the page whitelist; otherwise, or if the page type of the page to be collected is not found in the page whitelist, then it can be determined that the page to be collected is not in the page whitelist.

[0107] The correspondence between page types and information collection rules is stored in the rule engine.

[0108] After determining that the page to be collected is in the whitelist, at least one information collection rule corresponding to the page to be collected can be obtained from the rule engine based on the correspondence between the page type and the information collection rule.

[0109] Step S430: Based on the at least one information collection rule, obtain the page information of the page to be collected.

[0110] Step S440: Generate a desktop card based on the page information.

[0111] In an embodiment of the present application, after obtaining the page information of the page to be collected, a desktop card can be generated based on partial information in the page information, wherein the partial information is information in the page information that can be used to generate the desktop card.

[0112] As one approach, generating a desktop card based on the page information includes: acquiring at least one target information from the page information; and if the at least one target information includes card data, generating a desktop card based on the card data.

[0113] In the embodiments of the present application, target information is information obtained by extracting target fields from page information and summarizing them according to preset rules. The extracted target information can be of various types, where the type refers to the category to which the extracted target information belongs, such as schedule, personal information, to-do list, text summary, structured information, etc. Preset rules are pre-set rules for obtaining target fields from page information, and target fields are pre-set fields from which the user wants to extract information.

[0114] Card data refers to target information that can generate cards. Desktop cards are a form of information presented by an electronic device to the user, and can include pictures, text, links, controls, and other information related to the same topic. For example, weather cards, stock cards, and news cards, etc. In an embodiment of the present application, a pre-trained data recognition model can be used to determine whether at least one target information contains card data. Specifically, at least one target information can be input into a pre-trained data recognition model, and the data recognition model can recognize the at least one target information and then output the card data in the at least one target information.

[0115] As a way, the desktop card can be the entrance to the application (application, APP) corresponding to the desktop card. The user can operate the desktop card to open the application corresponding to the desktop card, so that the electronic device presents the interface of the application corresponding to the desktop card to the user, so that the user can view more detailed information on the interface. Furthermore, after opening the application corresponding to the desktop card, the user can perform corresponding operations on the interface of the application to meet their own needs. For example, the application corresponding to the weather card is weather. The user can open the application weather by operating the weather card, so that the electronic device presents the weather interface to the user. The user can set the information displayed on the weather card, or view the weather conditions of a city, etc. by operating the weather interface to meet their own needs.

[0116] As another way, the desktop card can also be the entrance to one or more services provided by the application corresponding to the desktop card. The user can operate the desktop card to open the service provided by the application corresponding to the desktop card, so that the electronic device presents the service interface of the application corresponding to the desktop card to the user. Figure 13 As shown, when you click Figure 13 When the verification code in the desktop card 130 is displayed, the verification code can be jumped to be displayed for the merchant to verify.

[0117] In the embodiment of the present application, the desktop card can be displayed on the desktop or the negative one screen of the electronic device, which is not specifically limited here. Among them, the negative one screen refers to the user interface of the leftmost split screen when the screen is swiped to the right on the main screen of the electronic device.

[0118] In an embodiment of the present application, different types of card data can generate different types of desktop cards.

[0119] Different desktop cards have different card functions. Therefore, each desktop card has its corresponding card display rules. The card display rules include the rules for generating the graphic style and color of the card itself, as well as the display rules for the card display data. For example, when a desktop card is used to display the user's exercise status, different card colors are used to represent the user's current exercise intensity according to the length of the exercise time. The data displayed on the card also includes mapped user exercise items. When the user is walking, a walking person silhouette icon is displayed; when the user is running, a running person silhouette icon is displayed, and so on. The main purpose of the different display rules for desktop cards is to display the data that needs to be displayed on the card in an eye-catching manner, so as to facilitate user identification. Therefore, in this embodiment, the differentiated display that is conducive to user identification can be used in desktop cards that implement different functions, rather than being limited to the above-mentioned embodiments.

[0120] Optionally, the at least one target information is recorded.

[0121] In an embodiment of the present application, after obtaining at least one type of target information, the at least one type of target information may be recorded. Recording here means that the at least one type of target information may be saved.

[0122] When recording at least one type of target information, the at least one type of target information may be stored in different formats based on the category to which the target information belongs. For example, if the target information is schedule information, the target information may be recorded in the form of a schedule; if the target information is personal information, the target information may be recorded in an encrypted form, etc., without further limitation.

[0123] The page information is input into a plurality of pre-trained large models respectively, and the target information of the page information output by each of the plurality of large models is obtained to obtain at least one target information, wherein different large models are used to extract information from different types of page information.

[0124] In an embodiment of the present application, multiple pre-trained large models are used to extract target information from different types of page information, that is, different large models have different capabilities. For example, multiple pre-trained large models can be OCR models, NER (Named Entity Recognition) models, and image classification and summarization models. Specifically, the OCR model provides the ability to extract text from images; the NER model is used to extract schedule / personal identity information; the image classification and summarization model is used for structured extraction, such as extracting movie ticket information from an order page. Different types can refer to types such as images or texts, or they can refer to different extracted fields.

[0125] Furthermore, multiple pre-trained large models can be deployed in the cloud server. When the capabilities of multiple pre-trained large models need to be invoked, large models with different capabilities can be invoked by calling different AI (Artificial Intelligence) capability interfaces. One large model corresponds to one AI capability interface.

[0126] As a method, multiple pre-trained large models can be called at the same time. After obtaining the page information, the page information is input into the multiple pre-trained large models respectively to obtain at least one target information output by the multiple pre-trained large models.

[0127] Optionally, multiple pre-trained large models can also be deployed in the electronic device. When the capabilities of multiple pre-trained large models need to be called, the multiple pre-trained large models can be called directly in the electronic device.

[0128] After multiple large models output the target information of the page information, it is possible to further determine whether the target information output by the large model complies with security compliance. If it does, the target information is recorded; if it does not, the target information is discarded.

[0129] As one approach, a to-do list is generated based on the page information; or, an article summary is generated based on the page information.

[0130] To-do items can be generated based on text and date information in the page information; article summaries can be generated based on text information in the page information. Target information can be extracted from the page information to obtain at least one target information, and then the to-do items and article summaries can be generated based on the at least one target information.

[0131] This application provides an information collection method that can accurately obtain the information collection rules corresponding to the page to be collected based on the page type, and then comprehensively collect the page information of the page to be collected. The collected page information can be displayed on the desktop in the form of cards, which can facilitate users to quickly retrieve it.

[0132] See also Figure 14 , an embodiment of the present application provides an information collection device 500, the device 500 comprising:

[0133] The format acquisition unit 510 is used to acquire the format information of the page to be collected.

[0134] As one approach, the format acquiring unit 510 is specifically configured to acquire a page identifier of the page to be collected, where the page identifier corresponds to the format information of the page to be collected.

[0135] As another way, the format acquiring unit 510 is specifically configured to acquire the service window name of the page to be collected, where the service window name corresponds to the format information of the page to be collected.

[0136] Optionally, the format acquiring unit 510 is specifically configured to acquire a page type of the page to be collected, where the page type corresponds to format information of the page to be collected.

[0137] The rule acquisition unit 520 is configured to acquire at least one information collection rule corresponding to the format information of the page to be collected from a rule engine.

[0138] Among them, the rule engine includes at least text collection rules, image collection rules, file path collection rules, self-rendering control collection rules, OCR collection rules, cloud crawler collection rules, page link collection rules and large image path collection rules; wherein, the text collection rules are used to obtain the text information of the page to be collected; the image collection rules are used to obtain the image information of the page to be collected; the file path collection rules are used to obtain the storage path of the file contained in the page to be collected; the self-rendering control collection rules are used to obtain the control information of the control that loads text in the self-rendering manner in the page to be collected; the OCR collection rules are used to obtain the text information of the picture in the page to be collected; the cloud crawler collection rules are used to obtain the graphic and text information of the page to be collected through page links; the page link collection rules are used to obtain the jump link of the page to be collected; the large image path collection rules are used to obtain the large image path of the page to be collected.

[0139] As one approach, the rule acquisition unit 520 is specifically configured to acquire, based on the page identifier, from a rule engine at least one information collection rule corresponding to the page to be collected.

[0140] Furthermore, the rule acquisition unit 520 is specifically configured to determine whether the page to be collected is in a page whitelist based on the page identifier; if so, acquire at least one information collection rule corresponding to the page to be collected from a rule engine based on the page identifier.

[0141] As another way, the rule acquisition unit 520 is specifically configured to acquire at least one information collection rule corresponding to the page to be collected from a rule engine based on the business window name.

[0142] Optionally, the rule acquisition unit 520 is specifically configured to acquire, based on the page type, at least one information collection rule corresponding to the page to be collected from a rule engine.

[0143] The information acquisition unit 530 is configured to acquire the page information of the page to be collected based on the at least one information collection rule.

[0144] As one approach, the information acquisition unit 530 is specifically configured to acquire the page information of the page to be acquired based on the at least one information acquisition rule in response to the information recording instruction.

[0145] The triggering mode of the information recording instruction includes one of key triggering, gesture triggering, voice triggering, page switching triggering and control triggering.

[0146] Optional, see Figure 15 , the apparatus 500 further includes:

[0147] The generating unit 540 is configured to generate a desktop card based on the page information; or generate a to-do item based on the page information; or generate an article summary based on the page information.

[0148] As one approach, the generating unit 540 is specifically configured to obtain at least one target information from the page information; if the at least one target information includes card data, generate a desktop card based on the card data.

[0149] The recording unit 550 is configured to record the at least one type of target information.

[0150] It should be noted that the device embodiment in this application corresponds to the aforementioned method embodiment. The specific principles in the device embodiment can be found in the contents of the aforementioned method embodiment and will not be repeated here.

[0151] The following will be combined Figure 16 An electronic device or server provided in this application is described.

[0152] See also Figure 16Based on the above-mentioned information collection method and apparatus, the embodiments of the present application also provide another electronic device or server 800 that can execute the above-mentioned information collection method. The electronic device or server 800 includes one or more (only one is shown in the figure) processors 802, a memory 804, and a network module 806 that are coupled to each other. The memory 804 stores a program that can execute the content of the above-mentioned embodiments, and the processor 802 can execute the program stored in the memory 804.

[0153] The processor 802 may include one or more processing cores. The processor 802 utilizes various interfaces and circuits to connect various components within the server 800. It executes instructions, programs, code sets, or instruction sets stored in the memory 804, and accesses data stored in the memory 804 to perform various server 800 functions and process data. Optionally, the processor 802 may be implemented using at least one of the following hardware forms: a digital signal processing (DSP), a field-programmable gate array (FPGA), or a programmable logic array (PLA). The processor 802 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU primarily processes the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing display content; and the modem handles wireless communications. It is understood that the modem may not be integrated into the processor 802 and may be implemented separately via a communication chip.

[0154] The memory 804 may include a random access memory (RAM) or a read-only memory (ROM). The memory 804 may be used to store instructions, programs, codes, code sets, or instruction sets. The memory 804 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the following various method embodiments, etc. The data storage area may also store data (such as a phone book, audio and video data, chat history data) created by the electronic device or server 800 during use.

[0155] The network module 806 is used to receive and transmit electromagnetic waves, realize the mutual conversion between electromagnetic waves and electrical signals, and thus communicate with a communication network or other devices, such as communicating with an audio playback device. The network module 806 may include various existing circuit components for performing these functions, such as an antenna, a radio frequency transceiver, a digital signal processor, an encryption / decryption chip, a subscriber identity module (SIM) card, a memory, etc. The network module 806 can communicate with various networks such as the Internet, an enterprise intranet, a wireless network, or communicate with other devices via a wireless network. The above-mentioned wireless network may include a cellular telephone network, a wireless local area network, or a metropolitan area network. For example, the network module 806 can exchange information with a base station.

[0156] Please refer to Figure 17 , which shows a block diagram of a computer-readable storage medium provided in an embodiment of the present application. The computer-readable storage medium 900 stores program code, which can be called by a processor to execute the method described in the above method embodiment.

[0157] The computer-readable storage medium 900 can be an electronic memory such as a flash memory, an EEPROM (Electrically Erasable Programmable Read-Only Memory), an EPROM, a hard disk, or a ROM. Alternatively, the computer-readable storage medium 900 includes a non-transitory computer-readable storage medium. The computer-readable storage medium 900 has storage space for program code 910 for executing any of the method steps in the above method. These program codes can be read from or written to one or more computer program products. The program code 910 can be compressed, for example, in a suitable form.

[0158] This application provides an information collection method, apparatus, electronic device, and storage medium. The method first obtains format information of a page to be collected, then retrieves at least one information collection rule corresponding to the format information of the page to be collected from a rule engine, and then obtains page information of the page to be collected based on the at least one information collection rule. This method allows accurate retrieval of at least one information collection rule corresponding to the page to be collected based on the format information of the page to be collected, thereby enabling comprehensive collection of page information for the page to be collected.

[0159] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the present invention and the claims, all of which are protected by the present invention.

Claims

1. An information collection method, characterized in that: The method comprises: Get the format information of the page to be collected; Acquire at least one information collection rule corresponding to the format information of the page to be collected from a rule engine; Based on the at least one information collection rule, page information of the page to be collected is obtained.

2. The method according to claim 1, characterized in that Get the format information of the page to be collected, including: Obtaining a page identifier of the page to be collected, where the page identifier corresponds to format information of the page to be collected; The acquiring, from a rule engine, at least one information collection rule corresponding to the format information of the page to be collected includes: Based on the page identifier, at least one information collection rule corresponding to the page to be collected is obtained from a rule engine.

3. The method according to claim 2, characterized in that The acquiring, based on the page identifier, at least one information collection rule corresponding to the page to be collected from a rule engine includes: Based on the page identifier, determining whether the page to be collected is in a page whitelist; If yes, based on the page identifier, obtain at least one information collection rule corresponding to the page to be collected from a rule engine.

4. The method according to claim 1, wherein The step of obtaining the format information of the page to be collected includes: Acquire the business window name of the page to be collected, where the business window name corresponds to the format information of the page to be collected; The acquiring, from a rule engine, at least one information collection rule corresponding to the format information of the page to be collected includes: Based on the business window name, at least one information collection rule corresponding to the page to be collected is obtained from a rule engine.

5. The method according to claim 1, wherein The step of obtaining the format information of the page to be collected includes: Acquire the page type of the page to be collected, where the page type corresponds to the format information of the page to be collected; The acquiring, from a rule engine, at least one information collection rule corresponding to the format information of the page to be collected includes: Based on the page type, at least one information collection rule corresponding to the page to be collected is obtained from a rule engine.

6. The method according to claim 1, wherein After obtaining the page information of the page to be collected based on the at least one information collection rule, the method further includes: Generate a desktop card based on the page information; or Generate a to-do item based on the page information; or, Generate an article summary based on the page information.

7. The method according to claim 6, characterized in that Generating a desktop card based on the page information includes: Acquire at least one target information from the page information; If the at least one target information includes card data, a desktop card is generated based on the card data.

8. The method according to claim 7, characterized in that The method further comprises: The at least one target information is recorded.

9. The method according to any one of claims 1 to 8, characterized in that: The rule engine includes at least text collection rules, image collection rules, file path collection rules, self-rendering control collection rules, OCR collection rules, cloud crawler collection rules, page link collection rules and large image path collection rules; wherein, the text collection rules are used to obtain the text information of the page to be collected; the image collection rules are used to obtain the image information of the page to be collected; the file path collection rules are used to obtain the storage path of the file contained in the page to be collected; the self-rendering control collection rules are used to obtain the control information of the control that loads text in the self-rendering manner in the page to be collected; the OCR collection rules are used to obtain the text information of the picture in the page to be collected; the cloud crawler collection rules are used to obtain the graphic and text information of the page to be collected through page links; the page link collection rules are used to obtain the jump link of the page to be collected; the large image path collection rules are used to obtain the large image path of the page to be collected.

10. The method according to claim 1, characterized in that Acquiring the page information of the page to be collected based on the at least one information collection rule includes: In response to the information recording instruction, based on the at least one information collection rule, page information of the page to be collected is obtained.

11. The method according to claim 10, characterized in that The triggering mode of the information recording instruction includes one of key triggering, gesture triggering, voice triggering, page switching triggering and control triggering.

12. An information collection device, characterized in that: The device comprises: A format acquisition unit, used to acquire format information of the page to be collected; A rule acquisition unit, configured to acquire at least one information collection rule corresponding to the format information of the page to be collected from a rule engine; The information acquisition unit is configured to acquire the page information of the page to be collected based on the at least one information collection rule.

13. An electronic device, characterized in that: The method comprises one or more processors; one or more programs are stored in the memory and configured to execute the method according to any one of claims 1 to 11 by the one or more processors.

14. A computer-readable storage medium, characterized in that The computer-readable storage medium stores program code, wherein when the program code is executed by a processor, the method according to any one of claims 1 to 11 is executed.

Citation Information

Patent Citations

  • Crawler processing method and device, server and computer readable storage medium

    CN110851681A

  • Information collection method and device and computer storage medium

    CN113485963A

  • Method, device and equipment for collecting page information and storage medium

    CN114579856A