Information processing method, device, electronic device, and computer-readable storage medium

By aggregating and displaying effective information groups from industry reports, the problems of low search hit rate and difficulty in user identification are solved, achieving a higher hit rate and better user experience.

CN111666383BActive Publication Date: 2025-09-26TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010622216.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-06-30
Publication Date
2025-09-26
Estimated Expiration
2040-06-30

AI Technical Summary

Technical Problem

When searching for industry reports in the existing technology, the hit rate is low and users need to screen the content, especially in reports with diverse content formats and high complexity, which leads to poor user experience.

Method used

Obtain valid information groups by searching keywords, determine the report files they belong to, aggregate them, generate content boxes, display report file information and valid information groups, use the preset report content reader for interactive operations, and support excerpt and notebook management.

Benefits of technology

It improves the hit rate of search keywords, is compatible with more report content types, reduces user screening behavior when browsing, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111666383B_ABST
    Figure CN111666383B_ABST
Patent Text Reader

Abstract

The present application provides an information processing method, device, electronic device, and computer-readable storage medium, relating to the field of information processing. The method includes: obtaining at least one corresponding valid information group based on a search keyword using a preset search engine; determining the report file to which each valid information group belongs, and obtaining the report file information of each report file; aggregating each valid information group based on the report file to which it belongs, obtaining aggregated report files and valid information corresponding to each report file; generating a content box for each report file, obtaining at least one content box; the content box includes the report file information and corresponding valid information of the report file; and displaying each content box. The present application improves the hit rate of search keywords and reduces the user's screening behavior during browsing, thereby improving the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of information processing technology. Specifically, the present application relates to an information processing method, device, electronic device and computer-readable storage medium. Background Art

[0002] Industry reports refer to business information and competitive intelligence with strong timeliness. They are generally based on the latest statistical data and survey data from national government agencies and professional market research organizations, and are analyzed and studied by industry veterans based on the professional research models and specific analysis methods of cooperative institutions to make research, analysis and forecasts on the current industry and market.

[0003] In the prior art, the search for industry reports is done by searching for keywords entered by the user, hitting the chart titles in the report, extracting the relevant visual chart contents in the report, and displaying them in a waterfall flow on the search results page. The display results are as follows: Figure 1 shown.

[0004] However, this search method has the following disadvantages:

[0005] 1) Using keywords to hit the titles of visual charts in reports requires a high level of structured standards for the report content. This approach works well for brokerage reports with relatively simple content structures, but has a lower hit rate for institutional reports and other types of reports with diverse and complex content formats.

[0006] 2) The content matching the search keyword is displayed in a waterfall flow, and different content is independent of each other on the search results page. When the matching content is sorted in a disorderly manner, users need to identify the content, which results in a poor user experience. Summary of the Invention

[0007] This application provides an information processing method, device, electronic device, and computer-readable storage medium that can solve the problem of low hit rates in search industry reports and the need for user screening. The technical solution is as follows:

[0008] In a first aspect, an information processing method is provided, the method comprising:

[0009] Searching for a search keyword and obtaining at least one valid information group corresponding to the search keyword;

[0010] Determine the report file to which each valid information group belongs, and obtain the report file information of each report file;

[0011] Aggregate each valid information group based on the report file to which it belongs, to obtain an aggregated valid information group corresponding to each report file;

[0012] Generate a content box for each report file to obtain at least one content box; the content box includes the report file information and a corresponding valid information group of the report file;

[0013] The at least one content box is displayed separately.

[0014] Preferably, any valid information group in the at least one valid information group includes a valid information image, a valid information title, and a valid information keyword;

[0015] The method further comprises:

[0016] When a display instruction for any content box among the at least one content box is received, obtaining valid information titles in each valid information group corresponding to the any content box;

[0017] Each valid information title and a valid information group corresponding to a currently selected valid information title among the valid information titles are displayed through a preset report content reader.

[0018] Preferably, the report content reader is further provided with at least one interactive instruction for the currently displayed valid information group;

[0019] The method further comprises:

[0020] When any interactive instruction in the at least one content box is triggered, an interactive action corresponding to the interactive instruction is executed for the currently displayed valid information group.

[0021] Preferably, the interactive instruction includes an excerpt instruction;

[0022] When any interaction instruction is triggered, executing the interaction action corresponding to the interaction instruction for the currently displayed valid information group includes:

[0023] When the excerpt instruction is triggered, determining whether there is a generated notebook in the preset favorites;

[0024] If so, displaying a notebook list of generated notebooks, and when receiving a confirmation instruction for any notebook in the notebook list, copying the currently displayed valid information group to the notebook;

[0025] If not, a preset notebook creation interface is displayed, a new notebook is created based on the notebook creation interface, and the currently displayed valid information group is copied to the new notebook.

[0026] Preferably, it also includes:

[0027] When a display instruction for any notebook generated in the preset favorites is received, the valid information group in the notebook is displayed through the report content reader.

[0028] Preferably, the search obtains at least one valid information group corresponding to the search keyword, including:

[0029] Performing query analysis on the search keywords to obtain analyzed keywords;

[0030] Assembling the analyzed keywords based on Elasticsearch Query DSL syntax to obtain a query statement for a valid information group; the query statement includes a keyword field and a title field;

[0031] The query statement is used to query an index in a preset search engine to obtain at least one valid information group that matches the search keyword.

[0032] Preferably, the preset search engine is generated in the following manner:

[0033] When it is detected that any valid information group in the at least one valid information group stored in the preset valid information database has a data update, obtaining the valid information title and valid information keyword of the valid information group where the data update occurs; the data update includes at least one of adding, deleting, and modifying the valid information group;

[0034] An index is generated based on the valid information title and the valid information keyword, and a mapping relationship between the valid information title, the valid information keyword and the index is established; wherein the index includes a title field and a keyword field.

[0035] Preferably, the preset valid information database is generated by:

[0036] Obtain report files;

[0037] Performing document segmentation processing on the report file according to the number of pages to obtain at least one report file image;

[0038] Performing character block recognition on each report file image to obtain at least one character block corresponding to each report file image;

[0039] Taking the report document image in which the at least one character block meets the preset requirements as a valid information image, thereby obtaining at least one valid information image;

[0040] Extracting the valid information title and valid information keyword of each valid information image, and establishing an association relationship between each valid information image, the valid information title and the valid information keyword corresponding to each valid information image;

[0041] Each valid information image, the valid information title corresponding to each valid information image, the valid information keyword, and the association relationship are stored in the valid information database.

[0042] Preferably, the report document image in which the at least one character block meets the preset requirements is used as a valid information image, and at least one valid information image is obtained, including:

[0043] detecting whether the number of digital blocks in each report file image exceeds a first number threshold;

[0044] If so, the report document images exceeding the first quantity threshold in each report document image are taken as valid information images to obtain at least one valid information image.

[0045] Preferably, the report document image in which the at least one character block meets the preset requirements is used as a valid information image, and at least one valid information image is obtained, including:

[0046] detecting whether a ratio of the number of digital blocks in each report file image to the number of all word blocks in the corresponding report file image exceeds a ratio threshold;

[0047] If so, the report document image that exceeds the ratio threshold in each report document image is taken as a valid information image to obtain at least one valid information image.

[0048] Preferably, the report document image in which the at least one character block meets the preset requirements is used as a valid information image, and at least one valid information image is obtained, including:

[0049] Obtaining the height of at least one character block in each report file image, and determining a preset number of target character blocks with the largest heights;

[0050] Detect whether the target character block in each report file image contains a Chinese character block;

[0051] If so, detecting whether the number of Chinese characters in the target character block containing the Chinese block exceeds a third number threshold;

[0052] If so, the report document image in which the target character block includes a Chinese character block is used as a valid information image to obtain at least one valid information image.

[0053] In a second aspect, an information processing device is provided, the device comprising:

[0054] A search module, configured to search for a search keyword and obtain at least one valid information group corresponding to the search keyword;

[0055] A processing module, configured to determine the report file to which each valid information group belongs, and obtain report file information of each report file;

[0056] an aggregation module, configured to aggregate each valid information group based on the report file to which it belongs, to obtain an aggregated valid information group corresponding to each report file;

[0057] A generating module, configured to generate a content box for each report file to obtain at least one content box; the content box includes the report file information and a corresponding valid information group of the report file;

[0058] The display module is configured to display the at least one content box respectively.

[0059] Preferably, any valid information group in the at least one valid information group includes a valid information image, a valid information title, and a valid information keyword;

[0060] The device further comprises:

[0061] a receiving module, configured to receive a display instruction for any content box among the at least one content box;

[0062] an acquisition module, configured to acquire valid information titles in each valid information group corresponding to any one of the content boxes;

[0063] The display module is further configured to display each valid information title and a valid information group corresponding to a currently selected valid information title among the valid information titles through a preset report content reader.

[0064] Preferably, the report content reader is further provided with at least one interactive instruction for the currently displayed valid information group;

[0065] The device further comprises:

[0066] The execution module is configured to execute an interaction action corresponding to the interaction instruction for the currently displayed valid information group when any interaction instruction among the at least one interaction instruction is triggered.

[0067] Preferably, the interactive instruction includes an excerpt instruction;

[0068] The execution module is specifically used for:

[0069] When the excerpt instruction is triggered, determining whether there is a generated notebook in the preset favorites;

[0070] If so, displaying a notebook list of generated notebooks, and when receiving a confirmation instruction for any notebook in the notebook list, copying the currently displayed valid information group to the notebook;

[0071] If not, a preset notebook creation interface is displayed, a new notebook is created based on the notebook creation interface, and the currently displayed valid information group is copied to the new notebook.

[0072] Preferably, the receiving module is further configured to receive a display instruction for any notebook generated in the preset favorites;

[0073] The display module is further configured to display the valid information group in the notebook through a report content reader.

[0074] Preferably, the search module includes:

[0075] An analysis submodule, configured to perform query analysis on the search keywords to obtain analyzed keywords;

[0076] A statement assembly submodule is used to assemble the analyzed keywords based on the Elasticsearch Query DSL syntax to obtain a query statement for a valid information group; the query statement includes a keyword field and a title field;

[0077] The query submodule is used to use the query statement and the index in the preset search engine to perform a query to obtain at least one valid information group that matches the search keyword.

[0078] Preferably, the preset search engine is generated in the following manner:

[0079] When it is detected that any valid information group in the at least one valid information group stored in the preset valid information database has a data update, obtaining the valid information title and valid information keyword of the valid information group where the data update occurs; the data update includes at least one of adding, deleting, and modifying the valid information group;

[0080] An index is generated based on the valid information title and the valid information keyword, and a mapping relationship between the valid information title, the valid information keyword and the index is established; wherein the index includes a title field and a keyword field.

[0081] Preferably, the preset valid information database is generated by:

[0082] Obtain report files;

[0083] Performing document segmentation processing on the report file according to the number of pages to obtain at least one report file image;

[0084] Performing character block recognition on each report file image to obtain at least one character block corresponding to each report file image;

[0085] Taking the report document image in which the at least one character block meets the preset requirements as a valid information image, thereby obtaining at least one valid information image;

[0086] Extracting the valid information title and valid information keyword of each valid information image, and establishing an association relationship between each valid information image and the valid information title and valid information keyword corresponding to each valid information image;

[0087] Each valid information image, the valid information title corresponding to each valid information image, the valid information keyword, and the association relationship are stored in the valid information database.

[0088] Preferably, the report document image in which the at least one character block meets the preset requirements is used as a valid information image, and at least one valid information image is obtained, including:

[0089] detecting whether the number of digital blocks in each report file image exceeds a first number threshold;

[0090] If so, the report document images exceeding the first quantity threshold in each report document image are taken as valid information images to obtain at least one valid information image.

[0091] Preferably, the report document image in which the at least one character block meets the preset requirements is used as a valid information image, and at least one valid information image is obtained, including:

[0092] detecting whether a ratio of the number of digital blocks in each report file image to the number of all word blocks in the corresponding report file image exceeds a ratio threshold;

[0093] If so, the report document image that exceeds the ratio threshold in each report document image is taken as a valid information image to obtain at least one valid information image.

[0094] Preferably, the report document image in which the at least one character block meets the preset requirements is used as a valid information image, and at least one valid information image is obtained, including:

[0095] Obtaining the height of at least one character block in each report file image, and determining a preset number of target character blocks with the largest heights;

[0096] Detect whether the target character block in each report file image contains a Chinese character block;

[0097] If so, detecting whether the number of Chinese characters in the target character block containing the Chinese block exceeds a third number threshold;

[0098] If so, the report document image in which the target character block includes a Chinese character block is used as a valid information image to obtain at least one valid information image.

[0099] According to a third aspect, an electronic device is provided, comprising:

[0100] processor, memory, and bus;

[0101] The bus is used to connect the processor and the memory;

[0102] The memory is used to store operation instructions;

[0103] The processor is used to call the operation instruction, and the executable instruction enables the processor to perform the operation corresponding to the information processing method shown in the first aspect of the present application.

[0104] In a fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the information processing method shown in the first aspect of the present application is implemented.

[0105] The beneficial effects of the technical solution provided by this application are:

[0106] For the search keyword, at least one valid information group corresponding to the search keyword is searched and obtained, and then the report files to which each valid information group belongs are determined, and the report file information of each report file is obtained, and then each valid information group is aggregated based on the report file to which it belongs, to obtain the aggregated valid information group corresponding to each report file, and a content box is generated for each report file to obtain at least one content box; the content box includes the report file information and the corresponding valid information group of the report file; and the at least one content box is displayed separately. Through the above method, the embodiment of the present invention can comprehensively identify the content of all reports based on the search keyword, including but not limited to visual chart titles. Compared with the prior art that is limited to the identification of visual chart titles, resulting in a low search keyword hit rate for institutional reports and other types of reports with diverse and complex content formats, the embodiment of the present invention has a lower requirement for the standardization of report content and is compatible with more report content types, thereby improving the search keyword hit rate. At the same time, through comprehensive identification, each valid information group that matches the search keyword and belongs to different report files is obtained, and then each valid information group is displayed in an aggregated manner based on the report file, so that multiple valid information groups that match the search keyword in the same report file are related, reducing the user's screening behavior when browsing, thereby improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0107] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for describing the embodiments of the present application.

[0108] Figure 1 A schematic diagram of a search result page for searching for industry reports in the prior art;

[0109] Figure 2 A flowchart of an information processing method provided in one embodiment of the present application;

[0110] Figure 3 A flowchart of an information processing method provided in another embodiment of the present application;

[0111] Figure 4 This is a schematic diagram of the interface of the content box in this application;

[0112] Figure 5 This is a diagram of the search results page for searching for industry reports in this application;

[0113] Figures 6A-6B This is a diagram of the interface of the report content reader in this application Figure 1 and two ;

[0114] Figure 7This is a diagram of the interface of the report content reader in this application Figure 3 ;

[0115] Figures 8A-8B A diagram showing the effect of selecting a notebook for excerpting in this application;

[0116] Figure 9 A diagram showing the effect of creating an excerpt for a new notebook in this application;

[0117] Figure 10 This is a schematic diagram of the process extracted in this application;

[0118] Figure 11 This is a schematic diagram of the Favorites interface in this application;

[0119] Figure 12 This is a schematic diagram of the interface for browsing excerpts using the report content reader in this application;

[0120] Figure 13 This is a schematic diagram of the search process based on search keywords in this application;

[0121] Figure 14 This is a diagram of data processing for the ES search engine in this application;

[0122] Figure 15 This is a schematic diagram of the OCR effect in this application;

[0123] Figure 16 A schematic diagram of the process of extracting valid information images in this application;

[0124] Figure 17 A schematic structural diagram of an information processing device provided in yet another embodiment of the present application;

[0125] Figure 18 A schematic structural diagram of an electronic device for information processing provided in yet another embodiment of the present application. DETAILED DESCRIPTION

[0126] The following describes in detail embodiments of the present application, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and are not to be construed as limiting the present invention.

[0127] It will be understood by those skilled in the art that, unless expressly stated otherwise, the singular forms "a", "an", "said" and "the" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present application refers to the presence of the features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be intermediate elements. In addition, "connected" or "coupled" as used herein may include wireless connections or wireless couplings. The term "and / or" used herein includes all or any units and all combinations of one or more associated listed items.

[0128] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0129] First, several terms involved in this application are introduced and explained:

[0130] Cloud technology is a general term for network technologies, information technologies, integration technologies, management platform technologies, and application technologies based on the cloud computing business model. It can form a resource pool that can be used flexibly and conveniently on demand. Cloud computing technology will become a crucial support. Backend services for technical network systems, such as video websites, image websites, and more portals, require extensive computing and storage resources. With the rapid development and application of the internet industry, every item will likely have its own unique identifier, requiring transmission to backend systems for logical processing. Different levels of data will be processed separately, and data from various industries will require a strong system backend, which can only be achieved through cloud computing.

[0131] A database, in short, can be thought of as a digital filing cabinet—a place where electronic files are stored, allowing users to add, query, update, and delete data. A database is a collection of data stored in a specific way, shared by multiple users, with minimal redundancy, and independent of applications.

[0132] A database management system (DBMS) is a computer software system designed for managing databases, typically providing basic functions such as storage, retrieval, security, and backup. DBMSs can be categorized by the database model they support, such as relational or XML (Extensible Markup Language); by the type of computer they support, such as server clusters or mobile phones; by the query language they use, such as SQL (Structured Query Language) or XQuery; by performance priorities, such as maximum scale or maximum speed; or by other classification methods. Regardless of the classification method used, some DBMSs can cross categories, for example, supporting multiple query languages ​​simultaneously.

[0133] Cloud storage is a new concept that has been extended and developed from the concept of cloud computing. A distributed cloud storage system (hereinafter referred to as storage system) refers to a storage system that uses cluster applications, grid technology, and distributed storage file systems to bring together a large number of different types of storage devices (storage devices are also called storage nodes) in the network through application software or application interfaces to work together and provide external data storage and business access functions.

[0134] Currently, storage systems utilize a method for creating logical volumes. When creating a logical volume, physical storage space is allocated for each logical volume. This physical storage space may consist of disks on a specific storage device or several storage devices. When a client stores data on a logical volume, it stores the data on a file system. The file system divides the data into multiple parts, each of which is an object. An object contains not only the data but also additional information such as the data identifier (ID) of the data entity. The file system writes each object to the physical storage space of the logical volume and records the storage location information of each object. Therefore, when a client requests access to data, the file system can provide access to the data based on the storage location information of each object.

[0135] The storage system allocates physical storage space to logical volumes by pre-dividing the physical storage space into stripes based on the estimated capacity of the objects to be stored in the logical volume (this estimate often has a large margin relative to the actual capacity of the objects to be stored) and the Redundant Array of Independent Disks (RAID) groupings. A logical volume can be understood as a stripe, thereby allocating physical storage space to the logical volume.

[0136] In the present application, an information processing method can be executed in a server. The server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal can be a smart phone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, etc., but is not limited to this. The terminal and the server can be directly or indirectly connected via wired or wireless communication, and this application does not limit this.

[0137] Furthermore, users can interact with the server through the terminal to implement service requests. The terminal may have the following features:

[0138] (1) In terms of hardware system, the device has a central processing unit, memory, input components and output components. In other words, the device is often a microcomputer device with communication functions. In addition, it can also have multiple input methods, such as keyboard, mouse, touch screen, microphone and camera, and the input can be adjusted as needed. At the same time, the device often has multiple output methods, such as receiver, display screen, etc., which can also be adjusted as needed;

[0139] (2) In terms of software system, the device must have an operating system, such as Windows Mobile, Symbian, Palm, Android, iOS, etc. At the same time, these operating systems are becoming more and more open, and personalized applications developed based on these open operating system platforms are endless, such as address books, calendars, notepads, calculators, and various games, which greatly meet the needs of personalized users;

[0140] (3) In terms of communication capabilities, the device has flexible access methods and high-bandwidth communication performance, and can automatically adjust the selected communication method according to the selected service and the environment, so as to facilitate user use. The device can support GSM (Global System for Mobile Communication), WCDMA (Wideband Code Division Multiple Access), CDMA2000 (Code Division Multiple Access), TDSCDMA (Time Division-Synchronous Code Division Multiple Access), Wi-Fi (Wireless-Fidelity) and WiMAX (Worldwide Interoperability for Microwave Access), etc., so as to adapt to various network standards and support not only voice services but also various wireless data services;

[0141] (4) In terms of functional use, devices are more focused on humanization, personalization, and multifunctionality. With the development of computer technology, devices have moved from a "device-centric" model to a "people-centric" model, integrating embedded computing, control technology, artificial intelligence technology, and biometric authentication technology, fully embodying the people-oriented purpose. Due to the development of software technology, devices can adjust settings according to individual needs, making them more personalized. At the same time, the devices themselves integrate a lot of software and hardware, and their functions are becoming more and more powerful.

[0142] The information processing method, device, electronic device and computer-readable storage medium provided in this application are intended to solve the above technical problems in the prior art.

[0143] The following specific embodiments describe in detail the technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.

[0144] In one embodiment, an information processing method is provided, such as Figure 2 As shown, the method includes:

[0145] Step S201: searching for a search keyword to obtain at least one valid information group corresponding to the search keyword;

[0146] In an embodiment of the present invention, an application for browsing reports may be installed in the terminal. The application may include a search interface, and a search bar may be provided in the search interface. A user may enter a search keyword in the search bar and search. The application then obtains corresponding search results through the search and displays the search results to the user. In an embodiment of the present invention, a report may be any file containing relevant information. For example, an industry report may include an industry analysis report, an industry research report, an industry data report, and the like. For example, the "OTA Industry - Guosheng Securities Broker Report" is an industry report.

[0147] Furthermore, the search results may include at least one valid information group corresponding to the search keyword. A valid information group may include at least one valid piece of information, which may correspond to the search keyword. In some embodiments, valid information refers to report content within industry reports that matches the user's search intent and can directly provide high-value reference information for the user's research work. According to current report writing standards and practices among various institutions and teams in the market, high-value information within a report generally appears in the report's charts and graphs.

[0148] Valid information includes, but is not limited to, valid information images, valid information titles, and valid information keywords within a report. A valid information image refers to an image corresponding to a full page of report containing high-value valid information; a valid information title refers to the title of the valid information image; and valid information keywords refer to keywords contained within the valid information image and / or keywords surrounding the valid information image. In some embodiments, a valid information group may include at least one valid information corresponding to a search keyword and belonging to the same report file. This at least one valid information may be any one or more of the valid information image, valid information title, and valid information keyword corresponding to the search keyword.

[0149] Step S202, determining the report file to which each valid information group belongs, and obtaining the report file information of each report file;

[0150] Specifically, each valid information group can be stored in a valid information database. When a user searches using a search keyword, a preset search engine can query the preset valid information database to obtain at least one matching valid information group. Each valid information group obtained by the query can belong to a different report file. For example, three valid information groups are obtained through the query: group a, group b, and group c. Group a and group b belong to report file A, and group c belongs to report file B. A report file can be an industry report. For example, the industry report "OTA Industry - Guosheng Securities_Brokerage Report" is a report file.

[0151] Therefore, after searching for each valid information group, the report file to which each valid information group belongs can be further determined, and the report file information of each report file can be obtained. The report file information includes but is not limited to: report ID, creator, tags, introduction, summary, creation time, and industry type.

[0152] Step S203, aggregating each valid information group based on the report file to which it belongs, to obtain an aggregated valid information group corresponding to each report file;

[0153] After determining the report files to which each valid information group belongs, the valid information groups can be aggregated according to the report files to determine the aggregated valid information groups corresponding to each report file. For example, by aggregating valid information groups a, b, and c in the previous example, we can determine that report file A corresponds to valid information groups a and b, and report file B corresponds to valid information group c.

[0154] Step S204: Generate a content box for each report to obtain at least one content box; the content box includes report file information and a corresponding valid information group of the report file;

[0155] Specifically, a content box is generated for each report and corresponding valid information group, thereby obtaining the same number of content boxes as the report files, each content box including the report file information and the corresponding valid information group of the report file.

[0156] Step S205: display at least one content box respectively.

[0157] After obtaining multiple content boxes, you can display each content box separately in the application interface. For example, in content box 1, the report file information and a and b valid information groups of report file A are displayed, and in content box 2, the report file information and c valid information groups of report file B are displayed.

[0158] In an embodiment of the present invention, for a search keyword, at least one valid information group corresponding to the search keyword is searched and obtained, and then the report files to which each valid information group belongs are determined, and the report file information of each report file is obtained, and then each valid information group is aggregated based on the report file to which it belongs to obtain the aggregated valid information group corresponding to each report file, and a content box is generated for each report file to obtain at least one content box; the content box includes the report file information of the report file and the corresponding valid information group; and at least one content box is displayed respectively. Through the above method, the embodiment of the present invention can comprehensively identify the content of all reports based on the search keyword, including but not limited to visual chart titles. Compared with the prior art that is limited to the identification of visual chart titles, resulting in a low search keyword hit rate for institutional reports and other types of reports with diverse and complex content formats, the embodiment of the present invention has a lower requirement for the degree of standardization of report content and is compatible with more report content types, thereby improving the search keyword hit rate. At the same time, through comprehensive identification, each valid information group that matches the search keyword and belongs to different report files is obtained, and then each valid information group is displayed in an aggregated manner based on the report file, so that multiple valid information groups that match the search keyword in the same report file are related, reducing the user's screening behavior when browsing, thereby improving the user experience.

[0159] In another embodiment, an information processing method is provided. Figure 3 As shown, the method includes:

[0160] Step S301: searching for a search keyword and obtaining at least one corresponding valid information group;

[0161] In an embodiment of the present invention, an application for browsing industry reports may be installed in the terminal. The application may include a search interface, and a search bar may be provided in the search interface. A user may enter a search keyword in the search bar and search. The application then searches to obtain corresponding search results and displays the search results to the user. Among them, the industry report may be an industry analysis report, an industry research report, an industry data report, etc. For example, the "OTA Industry - Guosheng Securities Brokerage Report" is an industry report.

[0162] Furthermore, the search results can include at least one valid information group corresponding to the search keyword. Valid information refers to industry report content that matches the user's search intent and can directly provide high-value reference information for the user's research. Based on current report writing standards and practices among institutions and teams, high-value information within a report typically appears in the report's charts and graphs.

[0163] An effective information group includes, but is not limited to, effective information images, effective information titles, and effective information keywords within a report. An effective information image refers to an image corresponding to a full page of report containing high-value effective information; an effective information title refers to the title of the effective information image; and effective information keywords refer to the keywords contained within the effective information image and / or the keywords surrounding the effective information image.

[0164] When an application searches based on a keyword, it can call a preset valid information search interface to search. The valid information search interface includes:

[0165] Request method: GET

[0166] Request path: / api / search / modules

[0167] Request parameter: keyword, search keyword.

[0168] Step S302, determining the report file to which each valid information group belongs, and obtaining the report file information of each report file;

[0169] Specifically, each valid information group can be stored in a valid information database. When a user searches using a search keyword, a preset search engine can query the preset valid information database to obtain at least one matching valid information group. Each valid information group obtained by the query can belong to a different report file. A valid information group includes a valid information image, a valid information title, and a valid information keyword. For example, three valid information groups are obtained through the query: group a, group b, and group c. Group a and group b belong to report file A, and group c belongs to report file B. A report file can be an industry report. For example, the industry report "OTA Industry - Guosheng Securities_Brokerage Report" is a report file.

[0170] Therefore, after searching for each valid information group, the report file to which each valid information group belongs can be further determined, and the report file information of each report file can be obtained. The report file information includes but is not limited to: report ID, creator, tags, introduction, summary, creation time, and industry type.

[0171] Step S303: Aggregate each valid information group based on the report file to which it belongs, to obtain an aggregated valid information group corresponding to each report file;

[0172] After determining the report files to which each valid information group belongs, the valid information groups can be aggregated according to the report files to determine the aggregated valid information groups corresponding to each report file. For example, by aggregating valid information groups a, b, and c in the previous example, we can determine that report file A corresponds to valid information groups a and b, and report file B corresponds to valid information group c.

[0173] Step S304: Generate a content box for each report file to obtain at least one content box; the content box includes the report file information and the corresponding valid information group of the report file;

[0174] Specifically, a content box is generated for each report and corresponding valid information group, thereby obtaining the same number of content boxes as the report files, each content box including the report file information and the corresponding valid information group of the report file.

[0175] For example, Figure 4 As shown, the content box can include two areas: a first area and a second area. The first area can contain the report file information of the report file. Clicking it will jump to the details page of the report file. The second area can contain each valid information group that matches the search keyword. Clicking any valid information group will call out the report content reader. In this way, at least one valid information group belonging to the same report file is displayed in an aggregated manner, which improves the relevance of at least one valid information group belonging to the same report file. Readers can obtain the valid information group corresponding to each report file by simply screening the report file, avoiding the problem of needing to screen chaotic and disordered search results in the prior art.

[0176] It should be noted that the content box has the following formats: Figure 4 In addition to the above, other forms of content boxes are also applicable to the present application. Moreover, when there are a large number of valid information groups, a certain number of valid information groups can be displayed on the right side of the content box, and all valid information groups can be displayed in the report content reader. Alternatively, a scroll bar can be set on the right side of the content box to display all valid information groups. Of course, in actual applications, users can set the form and layout of the content box according to actual needs, and this application does not impose any restrictions on this.

[0177] Step S305, displaying at least one content box respectively;

[0178] After obtaining multiple content boxes, you can display each content box in the application interface. For example, display the report file information and valid information groups a and b of report file A in content box 1, and display the report file information and valid information group c of report file B in content box 2. For another example, if the search keyword is "OTA hotel agent commission", the following results are obtained: Figure 5 The individual content boxes shown.

[0179] Step S306: When a display instruction for any content box in the at least one content box is received, obtaining valid information titles in each valid information group corresponding to the content box;

[0180] Specifically, when the user clicks on any content box in at least one content box, a display instruction for the content box is initiated, and at this time, the valid information titles in each valid information group in the content box can be obtained.

[0181] Step S307: displaying each valid information title and the valid information group corresponding to the currently selected valid information title among the valid information titles through a preset report content reader;

[0182] The preset report content reader displays each valid information title and the valid information group corresponding to the currently selected valid information title in each valid information title, such as Figure 6A shown.

[0183] The report content reader can browse and manage all valid information groups in different report files corresponding to the current search keyword. Figure 6B As shown, the report content box can include four parts:

[0184] 1) Search keyword information area

[0185] Displays the search keywords entered by the user on the current search results page.

[0186] 2) Report and effective information title navigation area

[0187] Displays the report file to which the currently browsed content belongs, as well as other valid information titles belonging to the same report file. Users can switch other valid information titles in the navigation area by clicking on them or scrolling the cursor to browse continuously.

[0188] Furthermore, after the last valid information title of a report file, the navigation area can automatically load the title of the next report file, such as Figure 7 shown.

[0189] 3) Report content reading area

[0190] The effective information image corresponding to the effective information title, that is, the detailed report content, can be enlarged or reduced.

[0191] 4) Report content operation area

[0192] At least one interactive instruction is provided for the currently displayed effective information report content, which in the embodiment of the present invention includes but is not limited to: extract, download and original text.

[0193] The report content box can call a preset valid information acquisition interface to obtain valid information. The valid information acquisition interface includes:

[0194] Request method: GET

[0195] Request path: / api / report?modules=1

[0196] Request parameters: related parameters of the report file.

[0197] Step S308: When any one of the at least one interaction instruction is triggered, executing the interaction action corresponding to the interaction instruction for the currently displayed valid information group;

[0198] Among them, click "Extract" to extract the valid information group into the notebook; click "Download" to download the valid information image to the local computer; click "Original Text" to open a new page window, and display the original report file to which the valid information image belongs in the new page window, and locate the page with the same content as the valid information image.

[0199] When any of the at least one interaction instruction is triggered, the interaction action corresponding to the interaction instruction is executed for the currently displayed valid information group, including:

[0200] When the excerpt command is triggered, determine whether there is a generated notebook in the preset favorites;

[0201] If so, display the notebook list of the generated notebooks, and when a confirmation instruction is received for any notebook in the notebook list, copy the currently displayed valid information group to the notebook;

[0202] If not, the preset create notebook interface is displayed, a new notebook is created based on the create notebook interface, and the currently displayed valid information group is copied to the new notebook.

[0203] Specifically, when the user clicks "Extract", it can be determined whether there is a generated notebook in the preset favorites, that is, a notebook that the user has created. If so, a notebook list is displayed through a preset list window, and the notebook list can include all generated notebooks. When the user selects any of the notebooks and confirms, the currently displayed valid information group is copied to the notebook confirmed by the user, and then the "Extract" in the report content operation area can be changed to "Extracted", such as Figures 8A-8B shown.

[0204] If there is no generated notebook in the favorites, the new notebook window will be displayed directly. The user can set the name of the notebook in the new notebook window and generate the notebook after confirming it. Then, the user can change the "Excerpt" in the report content operation area to "Excerpted", such as Figure 9 shown.

[0205] Furthermore, in the list window, you can also set a "New Notebook" button. When the user clicks the button, the new notebook window can still be displayed, such as Figure 9 As shown, users can set the name of the notebook in the New Notebook window. After confirming, the notebook can be generated. At this time, the "Excerpt" in the report content operation area changes to "Excerpted", as shown in Figure 8B shown.

[0206] Notebooks are containers for recording valid information groups and can be used to manage and browse extracted valid information groups. Users can create, delete, and modify notebooks, and can also extract valid information groups into notebooks for easy viewing.

[0207] Reference Figure 10 , the detailed steps of the excerpt can be as follows:

[0208] 1) The user initiates a request to extract a valid information group and needs to select a notebook (including selecting one from the generated notebooks or creating a new notebook);

[0209] 2) The active information group will be completely cloned, not linked. Cloning can prevent the excerpt from being viewable when the active information group is deleted. After cloning, the excerpt is a copy of the active information group;

[0210] 3) Now associate the excerpt to the selected notebook;

[0211] 4) When the user needs to view the excerpt, he or she can initiate a request to view the notebook content.

[0212] Furthermore, the notebook interface may include:

[0213] 1) Notebook list GET / api / notebooks

[0214] 2) Notebook details (with excerpt list) GET / api / notebooks / {$notebook_id}

[0215] 3) Create a new notebook POST / api / notebooks

[0216] Parameter: required title, length 255.

[0217] 4) Update notebook PUT / PATCH / api / notebooks / {$notebook_id}

[0218] Parameter: required title, length 255.

[0219] 5) Delete a notebook DELETE / api / notebooks / {$notebook_id}

[0220] Parameter: Force is optional and can be set to 0 or 1. This parameter indicates whether to force deletion. If it is 0, no deletion will be performed. If it is 1, the excerpt in the notebook will be deleted together (no prompt will be given that there is an excerpt in the notebook).

[0221] When deleting a notebook, if there are excerpts in the notebook, a prompt and confirmation message will be generated, such as: "There are excerpts in the notebook, do you want to delete it?" The confirmation information includes "Yes" and "No". If the user clicks "Yes", the force value is 1; if the user clicks "No", the force value is 0.

[0222] 6)Excerpt content POST / api / notebooks / {$notebook_id} / excerpt

[0223] Parameters: required report_module_id, content module id.

[0224] 7) Delete the excerpt POST / api / notebooks / {$notebook_id} / unexcerpt

[0225] Parameters: Required report_module_id, content module id. You can delete multiple summaries in a notebook, separated by commas, for example: report_module_id=1,2,3.

[0226] Step S309: upon receiving a display instruction for any notebook generated in the preset favorites, displaying the valid information group in the notebook through the report content reader.

[0227] Specifically, all excerpted content can be browsed and managed uniformly in the favorites, such as Figure 11 In the embodiment of the present invention, the report content reader is a highly versatile control that can not only display valid information groups, but also be reused in more similar scenarios. For example, the excerpted content in the notebook can still be displayed using the report content reader control, such as Figure 12As shown, when the user clicks on any of the generated notebooks, the report content reader will be called out to browse the excerpt.

[0228] In a preferred embodiment of the present invention, the search obtains at least one valid information group corresponding to the search keyword, including:

[0229] Perform query analysis on the search keywords to obtain the analyzed keywords;

[0230] The analyzed keywords are assembled based on the Elasticsearch Query DSL syntax to obtain a query statement for the valid information group; the query statement includes the keyword field and the title field;

[0231] A query statement is used to query an index in a preset search engine to obtain at least one valid information group that matches the search keyword.

[0232] Specifically, in the search interface of the application, two search modes can be set: "Search content" and "Search report". Among them, search report can be based on the name of the report file, that is, ordinary search; search content is based on the content of the report.

[0233] Reference Figure 13 , which is a schematic diagram of the search process based on search keywords in an embodiment of the present invention. When searching for the search keywords input by the user, first determine whether the search mode is "search content". If not, perform a normal search (i.e., "search report") to obtain the result page of the normal search; if so, perform query analysis on the search keywords, including segmenting and expanding the search keywords to obtain the analyzed search keywords, and then use the Elasticsearch QueryDSL syntax to assemble the analyzed keywords to obtain a query statement for querying valid information groups, wherein the query statement includes a keyword field and a title field; then query the query statement through a preset search engine, including using the query statement to query the index in the search engine, so as to obtain at least one valid information group matching the search keyword, and then execute steps S201 to S205, or steps S302 to S305.

[0234] The analyzed search keywords are then segmented and expanded into synonyms. For Chinese word segmentation, the open-source Elasticsearch plugin IK Analysis for Elasticsearch can be used, while for synonym expansion, a pre-built synonym thesaurus can be used.

[0235] The analyzed search keywords are assembled based on the Elasticsearch Query DSL syntax to obtain the query statement for the valid information group. The query statement searches for the title and keywords of the content module, and the weights of the title and keywords are also set. For example, "market research" is a query statement, where "research" is a synonym of "investigation" and "study".

[0236] It should be noted that in addition to the above-mentioned plug-ins, the word segmentation plug-in can also be other word segmentation plug-ins, which can be set according to actual needs in actual applications, and the embodiments of the present invention do not limit this; in addition to the thesaurus obtained by the above-mentioned method, the synonym thesaurus can also be a thesaurus obtained by other means, which can be set according to actual needs in actual applications, and the embodiments of the present invention do not limit this.

[0237] Furthermore, the index in the search engine may also include the ID of the report, that is, the ID of the report file to which each valid information group belongs. In this way, when each valid information group is searched, the report file to which each valid information group belongs can also be determined.

[0238] In a preferred embodiment of the present invention, the index in the preset search engine is generated in the following manner:

[0239] When it is detected that any valid information group in at least one valid information group stored in a preset valid information database has data updated, the valid information title and valid information keyword of the valid information group having data updated are obtained; the data update includes at least one of adding, deleting, and modifying the valid information group;

[0240] An index is generated based on the valid information title and the valid information keyword, and a mapping relationship between the valid information title, the valid information keyword and the index is established; wherein the index includes a title field and a keyword field.

[0241] The search engine can be ES (ElasticSearch), a distributed full-text search engine. ES is document-oriented, meaning it can store entire objects or documents. However, it does more than just store data; it also indexes the content of each document to make it searchable. In ES, users can index and search documents or objects (rather than rows and columns of data).

[0242] Specifically, ES can obtain data from the valid information database based on asynchronous scripts. Figure 14As shown, when data in the valid information group in the valid information database MYSQL is updated, a data modification event is triggered. This data modification event enters the event processing queue and waits for ES to perform corresponding data processing on the valid information group. The data update includes at least one of adding, deleting, and modifying the valid information group. In this way, ES can update the valid information group in real time from the valid information database.

[0243] Furthermore, when ES updates the valid information group, it also needs to update the index based on the updated valid information group, including creating, modifying, and deleting the index. The index includes the title field and the keyword field, and determines the mapping between the index and the valid information group. The mapping can tell ES how to handle the various newly added fields. The fields that need to be processed in the valid information group are title (valid information title) and keyword (valid information keyword). The title is mapped to a text type field and will be segmented and inverted indexed during processing. The ik plug-in can be used for segmentation; the keyword is mapped to the keyword type and will only be matched exactly.

[0244] Furthermore, when the ES performs corresponding data processing on the valid information group, it can further obtain the ID of the report file to which the valid information group belongs (except for deleting the valid information).

[0245] In a preferred embodiment of the present invention, the preset valid information database is generated in the following manner:

[0246] Obtain report files;

[0247] Perform document segmentation processing on the report file according to the number of pages to obtain at least one report file image;

[0248] Performing character block recognition on each report file image to obtain at least one character block corresponding to each report file image;

[0249] Taking a report document image in which at least one character block meets preset requirements as a valid information image in each report document image, thereby obtaining at least one valid information image;

[0250] Extracting the valid information title and valid information keyword of each valid information image, and establishing an association relationship between each valid information image and the valid information title and valid information keyword corresponding to each valid information image;

[0251] Each valid information image, the valid information title corresponding to each valid information image, the valid information keyword, and the association relationship are stored in the valid information database.

[0252] Specifically, a complete report file is first obtained, and then each page of the report file is subjected to document slicing to obtain at least one report file image. Document slicing involves converting the obtained document page by page into images, such as PNG format images. Specifically, pdftopng in the xpdf toolkit can be used, which can convert PDF pages into PNG format images.

[0253] Then, each report document image is subjected to word block recognition to obtain at least one word block in each report document image. Among them, word block recognition can be performed using OCR (Optical Character Recognition). Each report document image must be processed by OCR. The text in different areas of the report document image can be called word blocks. After the word blocks are processed by OCR, information such as the content, position, confidence, and paragraph of the word blocks will be obtained. The word blocks obtained by OCR processing need to be filtered to delete non-text and digital word blocks. The effect of OCR is as follows: Figure 15 shown.

[0254] In a preferred embodiment of the present invention, it is detected whether the number of digital blocks in each report document image exceeds a first number threshold;

[0255] If so, the report document images exceeding the first quantity threshold in each report document image are taken as valid information images to obtain at least one valid information image.

[0256] In a preferred embodiment of the present invention, a report document image in which at least one character block meets preset requirements is used as a valid information image, and at least one valid information image is obtained, including:

[0257] detecting whether a ratio of the number of digital blocks in each report file image to the number of all word blocks in the corresponding report file image exceeds a ratio threshold;

[0258] If so, the report document image that exceeds the ratio threshold in each report document image is taken as a valid information image to obtain at least one valid information image.

[0259] In a preferred embodiment of the present invention, a report document image in which at least one character block meets preset requirements is used as a valid information image, and at least one valid information image is obtained, including:

[0260] Obtaining the height of at least one character block in each report file image, and determining a preset number of target character blocks with the largest heights;

[0261] Detect whether the target character block in each report file image contains a Chinese character block;

[0262] If so, detecting whether the number of Chinese characters in the target character block containing the Chinese block exceeds a third number threshold;

[0263] If so, the report document image in which the target character block includes a Chinese character block is used as a valid information image to obtain at least one valid information image.

[0264] Specifically, high-value effective information is generally in the form of graphs and tables, with generalized conclusions and mostly numerical information. For any report file image, the following rules can be used to judge:

[0265] 1) Whether the number of pure numeric blocks (including those containing "%", "-", and "+") exceeds a first threshold, such as 30;

[0266] 2) Whether the ratio of the number of pure digital blocks to the number of all blocks in the report file image exceeds a ratio threshold, such as 0.2;

[0267] 3) Obtain the height of at least one character block in each report file image, and determine a preset number of target character blocks with the largest height; detect whether the target character blocks in each report file image contain Chinese character blocks; if so, detect whether the number of Chinese characters in the target character blocks containing Chinese blocks exceeds a third quantity threshold; if so, use the report file image in which the target character blocks in each report file image contain Chinese character blocks as a valid information image; for example, sort all the character blocks in the report file image in descending order according to the character block height, then obtain the top three target character blocks in the sorting, detect whether the three target character blocks contain Chinese character blocks, and if so, detect whether the number of Chinese characters in the target character blocks containing Chinese character blocks exceeds 8.

[0268] Of course, the above rules and values ​​are derived from actual experiments and can be adjusted according to actual needs in actual applications, and the embodiments of the present invention do not limit this. Moreover, at least one of the above rules can be used during detection, or in addition to the above rules, other rules can be used. In actual applications, they can be set according to actual needs, and the embodiments of the present invention do not limit this.

[0269] Based on the above rules, all word blocks in each page of the report document image are judged, and the report document image that meets the above rules is used as a valid information image, thereby obtaining at least one valid information image.

[0270] Among them, the extraction process of a valid information image can be abstracted into a Job class, and the processing Job class will be placed in the queue for execution. The execution process is as follows Figure 16 shown.

[0271] For each valid information image, the valid information title and valid information keywords are extracted. The extraction of valid information titles and valid information keywords can be done using Tencent Cloud Natural Language Processing (NLP) services. This NLP service deeply integrates Tencent's internal NLP technology and, relying on hundreds of billions of Chinese language corpora, provides 18 intelligent text processing capabilities, including intelligent word segmentation, entity recognition, text error correction, sentiment analysis, text classification, sensitive review, word vectors, keyword extraction, automatic summarization, intelligent chat, and encyclopedia knowledge graph query.

[0272] Then, an association relationship is established between each valid information image, and the valid information title and valid information keyword corresponding to each valid information image. The valid information image, valid information title, valid information keyword and the association relationship are then stored in a preset valid information database. In addition to valid information, other data can also be stored, and a data table is generated to establish an association relationship between each valid information group and other data. The generated data table is shown in Table 1:

[0273]

[0274]

[0275] Table 1

[0276] In an embodiment of the present invention, any database can use Cloud Object Storage (COS). COS is a distributed storage service launched by Tencent Cloud that has no directory hierarchy, no data format restrictions, can accommodate massive data, and supports HTTP / HTTPS protocol access.

[0277] It should be noted that the complete report file can be stored in a preset complete report database. The complete report database and the valid information database can be two independent databases or two independent parts of a database. In actual applications, they can be set according to actual needs, and the embodiments of the present invention do not limit this.

[0278] In an embodiment of the present invention, for a search keyword, at least one valid information group corresponding to the search keyword is searched and obtained, and then the report files to which each valid information group belongs are determined, and the report file information of each report file is obtained, and then each valid information group is aggregated based on the report file to which it belongs to obtain the aggregated valid information group corresponding to each report file, and a content box is generated for each report file to obtain at least one content box; the content box includes the report file information of the report file and the corresponding valid information group; and at least one content box is displayed respectively. Through the above method, the embodiment of the present invention can comprehensively identify the content of all reports based on the search keyword, including but not limited to visual chart titles. Compared with the prior art that is limited to the identification of visual chart titles, resulting in a low search keyword hit rate for institutional reports and other types of reports with diverse and complex content formats, the embodiment of the present invention has a lower requirement for the standardization of report content and is compatible with more report content types, thereby improving the search keyword hit rate. At the same time, through comprehensive identification, each valid information group that matches the search keyword and belongs to different report files is obtained, and then each valid information group is displayed in an aggregated manner based on the report file, so that multiple valid information groups that match the search keyword in the same report file are related, reducing the user's screening behavior when browsing, thereby improving the user experience.

[0279] Figure 17 A structural diagram of an information processing device provided in another embodiment of the present application is shown in FIG. Figure 17 As shown, the device of this embodiment may include:

[0280] Search module 1701, configured to search for a search keyword and obtain at least one valid information group corresponding to the search keyword;

[0281] The processing module 1702 is used to determine the report file to which each valid information group belongs, and obtain the report file information of each report file;

[0282] Aggregation module 1703, configured to aggregate each valid information group based on the report file to which it belongs, to obtain an aggregated valid information group corresponding to each report file;

[0283] A generating module 1704 is configured to generate a content box for each report file, obtaining at least one content box; the content box includes report file information of the report file and a corresponding valid information group;

[0284] The display module 1705 is configured to display at least one content box.

[0285] In a preferred embodiment of the present invention, any valid information group in the at least one valid information group includes a valid information image, a valid information title, and a valid information keyword;

[0286] The device also includes:

[0287] a receiving module, configured to receive a display instruction for any content box in at least one content box;

[0288] An acquisition module, configured to acquire a valid information title from each valid information corresponding to any content box;

[0289] The display module is further used to display each valid information title and the valid information corresponding to the valid information title currently selected among the valid information titles through a preset report content reader.

[0290] In a preferred embodiment of the present invention, the report content reader is further provided with at least one interactive instruction for the currently displayed valid information group;

[0291] The device also includes:

[0292] The execution module is configured to execute an interaction action corresponding to the interaction instruction for the currently displayed valid information group when any interaction instruction among the at least one interaction instruction is triggered.

[0293] In a preferred embodiment of the present invention, the interactive instruction includes an excerpt instruction;

[0294] The execution module is specifically used to:

[0295] When the excerpt command is triggered, determine whether there is a generated notebook in the preset favorites;

[0296] If so, display the notebook list of the generated notebooks, and when a confirmation instruction is received for any notebook in the notebook list, copy the currently displayed valid information group to the notebook;

[0297] If not, the preset create notebook interface is displayed, a new notebook is created based on the create notebook interface, and the currently displayed valid information group is copied to the new notebook.

[0298] In a preferred embodiment of the present invention, the receiving module is further configured to receive a display instruction for any notebook generated in the preset favorites;

[0299] The display module is also used to display the valid information groups in the notebook through the report content reader.

[0300] In a preferred embodiment of the present invention, the search module includes:

[0301] The analysis submodule is used to perform query analysis on the search keywords and obtain the analyzed keywords;

[0302] The statement assembly submodule is used to assemble the analyzed keywords based on the Elasticsearch Query DSL syntax to obtain the query statement of the valid information group; the query statement includes the keyword field and the title field;

[0303] The query submodule is used to use a query statement and an index in a preset search engine to perform a query to obtain at least one valid information group that matches the search keyword.

[0304] In a preferred embodiment of the present invention, the preset search engine is generated in the following manner:

[0305] When it is detected that any valid information group in the at least one valid information group stored in the preset valid information database has data updated, obtaining the valid information title and valid information keyword of the valid information group where the data update occurs; the data update includes at least one of adding, deleting, and modifying the valid information group;

[0306] An index is generated based on the valid information title and the valid information keyword, and a mapping relationship between the valid information title, the valid information keyword and the index is established; wherein the index includes a title field and a keyword field.

[0307] In a preferred embodiment of the present invention, the preset valid information database is generated in the following manner:

[0308] Obtain report files;

[0309] Perform document segmentation processing on the report file according to the number of pages to obtain at least one report file image;

[0310] Performing character block recognition on each report file image to obtain at least one character block corresponding to each report file image;

[0311] Taking a report document image in which at least one character block meets preset requirements as a valid information image in each report document image, thereby obtaining at least one valid information image;

[0312] Extracting the valid information title and valid information keyword of each valid information image, and establishing an association relationship between each valid information image and the valid information title and valid information keyword corresponding to each valid information image;

[0313] Each valid information image, the valid information title corresponding to each valid information image, the valid information keyword, and the association relationship are stored in the valid information database.

[0314] Preferably, the report document image in which the at least one character block meets the preset requirements is used as a valid information image, and at least one valid information image is obtained, including:

[0315] detecting whether the number of digital blocks in each report file image exceeds a first number threshold;

[0316] If so, the report document images exceeding the first quantity threshold in each report document image are taken as valid information images to obtain at least one valid information image.

[0317] Preferably, the report document image in which the at least one character block meets the preset requirements is used as a valid information image, and at least one valid information image is obtained, including:

[0318] detecting whether a ratio of the number of digital blocks in each report file image to the number of all word blocks in the corresponding report file image exceeds a ratio threshold;

[0319] If so, the report document image that exceeds the ratio threshold in each report document image is taken as a valid information image to obtain at least one valid information image.

[0320] Preferably, the report document image in which the at least one character block meets the preset requirements is used as a valid information image, and at least one valid information image is obtained, including:

[0321] Obtaining the height of at least one character block in each report file image, and determining a preset number of target character blocks with the largest heights;

[0322] Detect whether the target character block in each report file image contains a Chinese character block;

[0323] If so, detecting whether the number of Chinese characters in the target character block containing the Chinese block exceeds a third number threshold;

[0324] If so, the report document image in which the target character block includes a Chinese character block is used as a valid information image to obtain at least one valid information image.

[0325] The information processing device of this embodiment can execute the information processing method shown in the first embodiment and the second embodiment of this application. The implementation principles are similar and will not be repeated here.

[0326] In an embodiment of the present invention, for a search keyword, at least one valid information group corresponding to the search keyword is searched and obtained, and then the report files to which each valid information group belongs are determined, and the report file information of each report file is obtained, and then each valid information group is aggregated based on the report file to which it belongs to obtain the aggregated valid information group corresponding to each report file, and a content box is generated for each report file to obtain at least one content box; the content box includes the report file information of the report file and the corresponding valid information group; and at least one content box is displayed respectively. Through the above method, the embodiment of the present invention can comprehensively identify the content of all reports based on the search keyword, including but not limited to visual chart titles. Compared with the prior art that is limited to the identification of visual chart titles, resulting in a low search keyword hit rate for institutional reports and other types of reports with diverse and complex content formats, the embodiment of the present invention has a lower requirement for the degree of standardization of report content and is compatible with more report content types, thereby improving the search keyword hit rate. At the same time, through comprehensive identification, each valid information group that matches the search keyword and belongs to different report files is obtained, and then each valid information group is displayed in an aggregated manner based on the report file, so that multiple valid information groups that match the search keyword in the same report file are related, reducing the user's screening behavior when browsing, thereby improving the user experience.

[0327] In another embodiment of the present application, an electronic device is provided, comprising: a memory and a processor; and at least one program stored in the memory for execution by the processor. Compared to the prior art, the electronic device can: search for a search keyword to obtain at least one valid information group corresponding to the search keyword, then determine the report files to which each valid information group belongs, obtain report file information for each report file, aggregate each valid information group based on the report file to obtain an aggregated valid information group corresponding to each report file, generate a content box for each report file to obtain at least one content box; the content box includes the report file information of the report file and the corresponding valid information group; and display the at least one content box. Through the above-described method, the embodiment of the present invention can comprehensively identify the content of all reports based on the search keyword, including but not limited to titles of visual charts. Compared to the prior art, which is limited to identifying titles of visual charts, resulting in a low search keyword hit rate for institutional reports and other types of reports with diverse and complex content formats, the embodiment of the present invention has lower requirements for the degree of standardization of report content and is compatible with more report content types, thereby improving the search keyword hit rate. At the same time, through comprehensive identification, each valid information group that matches the search keyword and belongs to different report files is obtained, and then each valid information group is displayed in an aggregated manner based on the report file, so that multiple valid information groups that match the search keyword in the same report file are related, reducing the user's screening behavior when browsing, thereby improving the user experience.

[0328] In an alternative embodiment, an electronic device is provided, such as Figure 18 As shown, Figure 18 The electronic device 18000 shown includes: a processor 18001 and a memory 18003. The processor 18001 and the memory 18003 are connected, for example, via a bus 18002. Optionally, the electronic device 18000 may further include a transceiver 18004. It should be noted that in actual applications, the number of transceivers 18004 is not limited to one, and the structure of the electronic device 18000 does not constitute a limitation on the embodiments of the present application.

[0329] Processor 18001 may be a CPU, a general-purpose processor, a DSP, an ASIC, an FPGA, or other programmable logic device, a transistor logic device, a hardware component, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 18001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like.

[0330] The bus 18002 may include a path for transmitting information between the above components. The bus 18002 may be a PCI bus or an EISA bus. The bus 18002 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 18 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0331] Memory 18003 can be a ROM or other type of static storage device that can store static information and instructions, a RAM or other type of dynamic storage device that can store information and instructions, or an EEPROM, CD-ROM or other optical disk storage, optical disc storage (including compact disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited to these.

[0332] The memory 18003 is used to store application code for executing the solution of the present application, and the execution is controlled by the processor 18001. The processor 18001 is used to execute the application code stored in the memory 18003 to implement the content shown in any of the above method embodiments.

[0333] Among them, electronic devices include but are not limited to: mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc.

[0334] Another embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon. When executed on a computer, the computer is enabled to execute the corresponding contents of the aforementioned method embodiments. Compared to the prior art, for a search keyword, at least one valid information group corresponding to the search keyword is searched for, then the report files to which each valid information group belongs are determined, and report file information for each report file is obtained. Each valid information group is then aggregated based on the report file to which it belongs, obtaining aggregated valid information groups corresponding to each report file. A content box is generated for each report file, obtaining at least one content box; the content box includes the report file information of the report file and the corresponding valid information group; and at least one content box is displayed. Through the above-described method, the embodiment of the present invention can comprehensively identify the content of all reports based on the search keyword, including but not limited to titles of visual charts. Compared to the prior art, which is limited to identifying titles of visual charts, resulting in a low search keyword hit rate for institutional reports and other types of reports with diverse and complex content formats, the embodiment of the present invention has lower requirements for the standardization of report content and is compatible with more report content types, thereby improving the search keyword hit rate. At the same time, through comprehensive identification, each valid information group that matches the search keyword and belongs to different report files is obtained, and then each valid information group is displayed in an aggregated manner based on the report file, so that multiple valid information groups that match the search keyword in the same report file are related, reducing the user's screening behavior when browsing, thereby improving the user experience.

[0335] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown in sequence as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the flowcharts of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.

[0336] The above descriptions are only partial embodiments of the present invention. It should be pointed out that ordinary technicians in this technical field can make several improvements and modifications without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.

Claims

1. An information processing method, characterized in that: include: For a search keyword, at least one valid information group corresponding to the search keyword is searched and obtained; the valid information group includes a valid information image in a report file, the valid information image is an image of a report file page in the report file, and the character blocks in the report file image meet preset requirements, the preset requirements including: at least one target character block of maximum height in the report file image contains a Chinese character block, and the number of Chinese characters in the target character block containing the Chinese character block exceeds a third quantity threshold; Determine the report file to which each valid information group belongs, and obtain the report file information of each report file; Aggregate each valid information group based on the report file to which it belongs, to obtain an aggregated valid information group corresponding to each report file; Generate a corresponding content box for each report file, obtaining at least one content box; the content box includes the report file information and a corresponding valid information group of the report file; the report file information includes a summary of the report file; The at least one content box is displayed separately.

2. The information processing method according to claim 1, wherein: Any valid information group in the at least one valid information group further includes a valid information title and a valid information keyword; The method further comprises: When a display instruction for any content box among the at least one content box is received, obtaining valid information titles in each valid information group corresponding to the any content box; Each valid information title and a valid information group corresponding to a currently selected valid information title among the valid information titles are displayed through a preset report content reader.

3. The information processing method according to claim 2, wherein: The report content reader is further provided with at least one interactive instruction for the currently displayed valid information; The method further comprises: When any one of the at least one interaction instruction is triggered, an interaction action corresponding to the interaction instruction is executed for the currently displayed valid information.

4. The information processing method according to claim 3, wherein: The interactive instructions include excerpt instructions; When any one of the at least one interaction instruction is triggered, executing the interaction action corresponding to the interaction instruction for the currently displayed valid information group includes: When the excerpt instruction is triggered, determining whether there is a generated notebook in the preset favorites; If so, displaying a notebook list of generated notebooks, and when receiving a confirmation instruction for any notebook in the notebook list, copying the currently displayed valid information group to the notebook; If not, a preset notebook creation interface is displayed, a new notebook is created based on the notebook creation interface, and the currently displayed valid information group is copied to the new notebook.

5. The information processing method according to any one of claims 1 to 4, characterized in that: Also includes: When a display instruction for any notebook generated in the preset favorites is received, valid information in the notebook is displayed through the report content reader.

6. The information processing method according to claim 1, wherein: The search obtains at least one valid information group corresponding to the search keyword, including: Performing query analysis on the search keywords to obtain analyzed keywords; Assembling the analyzed keywords based on Elasticsearch Query DSL syntax to obtain a query statement for a valid information group; the query statement includes a keyword field and a title field; The query statement is used to query an index in a preset search engine to obtain at least one valid information group that matches the search keyword.

7. The information processing method according to claim 1 or 6, characterized in that: The index in the default search engine is generated as follows: When it is detected that any valid information group in the at least one valid information group stored in the preset valid information database has a data update, obtaining the valid information title and valid information keyword of the valid information group where the data update occurs; the data update includes at least one of adding, deleting, and modifying the valid information group; An index is generated based on the valid information title and the valid information keyword, and a mapping relationship between the valid information title, the valid information keyword and the index is established; wherein the index includes a title field and a keyword field.

8. The information processing method according to claim 7, wherein: The preset valid information database is generated in the following manner: Obtain report files; Performing document segmentation processing on the report file according to the number of pages to obtain at least one report file image; Performing character block recognition on each report file image to obtain at least one character block corresponding to each report file image; Taking the report document image in which at least one character block meets the preset requirements as a valid information image in each report document image, thereby obtaining at least one valid information image; Extracting the valid information title and valid information keyword of each valid information image, and establishing an association relationship between each valid information image and the valid information title and valid information keyword corresponding to each valid information image; Each valid information image, the valid information title corresponding to each valid information image, the valid information keyword, and the association relationship are stored in the valid information database.

9. The information processing method according to claim 8, characterized in that In each report document image, the report document image in which the at least one character block meets the preset requirements is used as a valid information image, and at least one valid information image is obtained, comprising: detecting whether the number of digital blocks in each report file image exceeds a first number threshold; If so, the report document images exceeding the first quantity threshold in each report document image are taken as valid information images to obtain at least one valid information image.

10. The information processing method according to claim 8, wherein: In each report document image, the report document image in which the at least one character block meets the preset requirements is used as a valid information image, and at least one valid information image is obtained, comprising: detecting whether a ratio of the number of digital blocks in each report file image to the number of all word blocks in the corresponding report file image exceeds a ratio threshold; If so, the report document image that exceeds the ratio threshold in each report document image is taken as a valid information image to obtain at least one valid information image.

11. The information processing method according to claim 8, wherein: In each report document image, the report document image in which the at least one character block meets the preset requirements is used as a valid information image, and at least one valid information image is obtained, comprising: Obtaining the height of at least one character block in each report file image, and determining a preset number of target character blocks with the largest heights; Detect whether the target character block in each report file image contains a Chinese character block; If so, detecting whether the number of Chinese characters in the target character block containing the Chinese block exceeds a third number threshold; If so, the report document image in which the target character block includes a Chinese character block is used as a valid information image to obtain at least one valid information image.

12. An information processing device, characterized in that: include: a search module configured to search for a search keyword and obtain at least one valid information group corresponding to the search keyword; the valid information group includes a valid information image in a report file, the valid information image being an image of a page of the report file, and the character blocks in the report file image satisfying preset requirements, the preset requirements including: at least one target character block of maximum height in the report file image contains a Chinese character block, and the number of Chinese characters in the target character block containing Chinese characters exceeds a third threshold; a processing module configured to determine the report file to which each valid information group belongs, and to obtain report file information for each report file; an aggregation module, configured to aggregate each valid information group based on the report file to which it belongs, to obtain an aggregated valid information group corresponding to each report file; a generating module, configured to generate a corresponding content box for each report file, obtaining at least one content box; the content box including the report file information and a corresponding valid information group of the report file; the report file information including a summary of the report file; The display module is configured to display the at least one content box respectively.

13. An electronic device, characterized in that: It includes: processor, memory, and bus; The bus is used to connect the processor and the memory; The memory is used to store operation instructions; The processor is configured to execute the information processing method according to any one of claims 1 to 11 by calling the operation instruction.

14. A computer-readable storage medium, characterized in that The computer storage medium is used to store computer instructions, and when the computer storage medium is run on a computer, the computer can execute the information processing method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Image element searching

    CN102483747A

  • Search result integration method and apparatus

    CN105786852A