Information processing method, device and equipment and computer readable storage medium
Through various scanning methods and language models, the similarity is calculated and the results are not repeated, and searched based on keywords and categories is solved, and the problems of high search complexity and low accuracy in the existing technology are solved, achieving efficient and accurate search results.
Patent Information
- Application Number
- CN202510611282.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-08-26
AI Technical Summary
In the prior art, the search complexity of industry research reports is high and the search accuracy is low. Especially when processing ultra-long text, lack of index establishment, chart adaptation, semantic understanding and content classification of the full text content, resulting in errors in the search results and poor readability.
By receiving user requests, multiple scanning methods and target language models are used to process pending reports, calculate the similarity between the report and the database, store non-repeated identification results, and determine the search content from the database based on keywords and categories, support complex search of multiple keywords, understand user intentions, and improve search accuracy.
It improves database quality, reduces retrieval complexity, enhances retrieval efficiency and accuracy, supports multi-keyword retrieval, and improves the readability and accuracy of search results.
Smart Images

Figure CN120541241A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to an information processing method, apparatus, device, and computer-readable storage medium. Background Art
[0002] In order to facilitate industry analysts to retrieve research reports on the target industry from numerous industry research reports and obtain the required content fragments therefrom, it is usually necessary to finely structure the content in the industry research reports, identify as much effective information as possible, and organize it; currently, related technologies usually use large model technology to perform targeted optimization on very long industry reports, and add key information extraction and classification models for paragraphs and charts to improve search accuracy; however, during the implementation process, the inventors found that there are at least the following problems in the existing technology: high retrieval complexity and low retrieval accuracy. Summary of the Invention
[0003] To solve the above technical problems, the embodiments of the present application provide an information processing method, apparatus, device and computer-readable storage medium, which solve the problems of high retrieval complexity and low retrieval accuracy in the prior art.
[0004] The technical solution of this application is achieved as follows:
[0005] Receive target requests sent by users;
[0006] When the target request is a processing request, each to-be-processed page in the to-be-processed report corresponding to the processing request is processed based on multiple scanning methods and target language models to obtain a target recognition result for each to-be-processed page;
[0007] determining, based on the plurality of target recognition results, a similarity between the report to be processed and each target report in the first database;
[0008] If each of the similarities is less than the target similarity, storing the target recognition result in the first database;
[0009] When the target request is a first search request, determining a plurality of first keywords in a to-be-searched sentence corresponding to the first search request and first categories corresponding to the plurality of first keywords;
[0010] determining a first search content of each to-be-selected report from the first database based on the plurality of first keywords and the first category;
[0011] Based on the user identifier and the identifier of the first search request, the first search content and the report name and access information corresponding to the first search content are sent to the user.
[0012] In the above solution, each to-be-processed page in the to-be-processed report corresponding to the processing request is processed to obtain a target recognition result for each to-be-processed page, including:
[0013] For each of the pages to be processed, extracting content from the page to be processed using a first scanning method and a second scanning method, respectively, to obtain a first recognition result and a second recognition result;
[0014] Based on the target language model, determining a first value for the first recognition result and a second value for the second recognition result, respectively; wherein the first value represents the content fluency of the first recognition result, and the second value represents the content fluency of the second recognition result;
[0015] The target recognition result is obtained by filtering the first recognition result and the second recognition result based on the first value and the second value.
[0016] In the above solution, determining the similarity between the report to be processed and each target report in the first database based on the plurality of target recognition results includes:
[0017] For each of the to-be-processed pages, determining a plurality of second keywords from the target recognition results based on the target classification model, and determining a first hash value for each of the second keywords;
[0018] determining a second hash value for the report to be processed based on a plurality of first hash values corresponding to a plurality of pages to be processed;
[0019] Each of the similarities is determined based on the second hash value and a target hash value corresponding to each of the target reports.
[0020] In the above solution, after determining the first Hash value for each second keyword, the method further includes:
[0021] determining, based on the target classification model, a third keyword for the report name of the report to be processed;
[0022] determining a target keyword for the report to be processed based on the plurality of second keywords and the third keyword corresponding to the plurality of pages to be processed;
[0023] determining a target category for the report to be processed based on the second category of each of the second keywords;
[0024] For each of the pending reports, the target keyword, the target identification content and the report name are stored in the first database, and the basic attribute information of the pending report is stored in the second database, and the pending report is stored in the third database; wherein the basic attribute information includes the access information of the pending report.
[0025] In the above solution, determining the first search content of each report to be selected from the first database includes:
[0026] Determining search intent information for the sentence to be searched based on the multiple first keywords and the first category;
[0027] determining the plurality of target paragraphs of each of the candidate reports based on the search intent information and the target keywords of each of the target reports;
[0028] For each of the candidate reports, retrieval enhancement generation is performed on the multiple target paragraphs based on the target language model to obtain the first retrieval content.
[0029] In the above solution, determining the multiple target paragraphs of each of the candidate reports includes:
[0030] Matching the search intent information with each target keyword to obtain a matching result;
[0031] If the matching result indicates that there is no first report matching the search intent information among the multiple candidate reports, determining the multiple target paragraphs based on the multiple first keywords and the report content of each report corresponding to the first category;
[0032] If the matching result indicates that the first report exists in the plurality of reports to be selected, the plurality of target paragraphs are determined based on the plurality of first keywords and the report content of each of the first reports.
[0033] In the above solution, after sending the first search content and the report name and access information corresponding to the first search content to the user, the method further includes:
[0034] Storing the first search content and the plurality of first keywords in a temporary memory;
[0035] When receiving a second search request sent by the user, determining a relevance between each first keyword and each third keyword corresponding to the second search request;
[0036] When the relevance satisfies the target relevance, second search content for the second search request is determined from the temporary memory based on a plurality of third keywords and a plurality of the first search content.
[0037] An information processing device, comprising:
[0038] A receiving unit, configured to receive a target request sent by a user;
[0039] A first processing unit is configured to process each to-be-processed page in the to-be-processed report corresponding to the target request, based on multiple scanning modes and target language models, to obtain a target recognition result for each to-be-processed page, when the target request is a processing request;
[0040] The first processing unit is further configured to determine a similarity between the report to be processed and each target report in the first database based on the plurality of target recognition results;
[0041] a storage unit, configured to store the target recognition result in the first database if each of the similarities is less than the target similarity;
[0042] a second processing unit, configured to determine, when the target request is a first search request, a plurality of first keywords in a to-be-searched statement corresponding to the first search request and a first category corresponding to the plurality of first keywords;
[0043] The second processing unit is further configured to determine a first search content of each to-be-selected report from the first database based on the plurality of first keywords and the first category;
[0044] A sending unit is configured to send the first search content and a report name and access information corresponding to the first search content to the user based on the user identifier and the identifier of the first search request.
[0045] An information processing device, comprising: a processor, a memory, and a communication bus;
[0046] The communication bus is used to realize the communication connection between the processor and the memory;
[0047] The processor is used to execute the information processing program in the memory to implement the steps of the above-mentioned information processing method.
[0048] A computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the above-mentioned information processing method.
[0049] Because after receiving the target request sent by the user, the target request is a processing request, and based on multiple scanning methods and target language models, each to-be-processed page in the to-be-processed report corresponding to the processing request is processed to obtain a target recognition result for each to-be-processed page, and based on multiple target recognition results, the similarity between the to-be-processed report and each target report in the first database is determined. If each similarity is less than the target similarity, multiple target recognition results are stored in the first database. In this way, by increasing the similarity calculation of the report content, only reports with a similarity less than the target similarity are stored when storing the target recognition results. In other words, only non-duplicate reports are stored, which improves the quality of the database, thereby improving the subsequent retrieval through the database. The retrieval efficiency is improved, which overcomes the problem of high retrieval complexity in the prior art; then, for the target request being the first retrieval request, multiple first keywords in the sentence to be retrieved corresponding to the first retrieval request and the first categories corresponding to the multiple first keywords are determined, and based on the multiple first keywords and the first categories, the first retrieval content of each report to be selected is determined from the first database, and then based on the user's identifier and the identifier of the first retrieval request, the first retrieval content and the report name and access information corresponding to the first retrieval content are sent to the user. In this way, not only complex retrieval of multiple keywords is supported, but also the intention of the user's sentence to be retrieved is understood through multiple first keywords and the first categories, which overcomes the problem of low retrieval accuracy in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 A flowchart of an information processing method provided in an embodiment of the present application;
[0051] Figure 2 A flowchart of another information processing method provided in an embodiment of the present application;
[0052] Figure 3 A schematic diagram of report processing in an information processing method provided in an embodiment of the present application;
[0053] FIG4( a ) is a schematic diagram of a content scanning method provided in an embodiment of the present application;
[0054] FIG4( b ) is a schematic diagram of another content scanning method provided in an embodiment of the present application;
[0055] Figure 5 A schematic diagram of report retrieval in an information processing method provided in an embodiment of the present application;
[0056] Figure 6 An architectural diagram of report processing and retrieval in an information processing method provided in an embodiment of the present application;
[0057] Figure 7A schematic diagram of the structure of an information processing device provided in an embodiment of the present application;
[0058] Figure 8 A schematic diagram of the structure of an information processing device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0059] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application.
[0060] It should be understood that the “embodiments of the present application” or “the aforementioned embodiments” mentioned throughout the specification mean that the specific features, structures or characteristics related to the embodiments are included in at least one embodiment of the present application. Therefore, “in the embodiments of the present application” or “in the aforementioned embodiments” appearing throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. In the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages and disadvantages of the embodiments.
[0061] Unless otherwise specified, when an electronic device performs any step in the embodiments of the present application, the processor of the electronic device may perform the step. It is also worth noting that the embodiments of the present application do not limit the order in which the electronic device performs the following steps. In addition, the methods used to process data in different embodiments may be the same method or different methods. It should also be noted that any step in the embodiments of the present application can be independently executed by the electronic device, that is, when the electronic device performs any step in the following embodiments, it can be independent of the execution of other steps.
[0062] It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application.
[0063] It should be noted that in response to the relevant demands for industry report retrieval, the current large model technology is already very mature. By integrating text processing technology, large language model technology, intelligent search technology, etc., it lowers the user threshold and improves the accuracy of retrieval. The results after retrieval can also be summarized and key information extracted through the large model, which improves user readability. However, in the case of industry reports with dozens of pages of content, more than 3,000 words, and mixed charts and content, the existing technology usually extracts information from the document content and builds a document index. When using this method, in order to improve efficiency, the document title, author, and a few hundred words before and after the document are usually indexed. This does not: 1. Lack of indexing of the full text content; 2. Lack of targeted adaptation of charts; 3. Lack of semantic and contextual understanding of the content, resulting in erroneous retrieval results; 4. Lack of classification of reports or long text content.
[0064] In response to the above-mentioned problems, there are currently many new methods for targeted optimization of very long texts such as document reports, and key information extraction and classification models for paragraphs and charts have been added to improve search accuracy. However, there are also some shortcomings: 1. The readability of the results is poor. The solution extracts and classifies the content according to industry keywords, but does not abbreviate and summarize the full text to form a core copy to further improve search accuracy and result readability. 2. The search complexity is high. The search method is still the traditional expert database retrieval model, and the retrieval and result sorting of multiple keywords have not been optimized in a targeted manner. There may be low matching accuracy in the retrieval results. 3. Key information coverage is incomplete, the retrieval accuracy is low, and there is a lack of recognition and classification of the characteristic information of the document itself.
[0065] At present, in the content data processing and retrieval of long texts such as industry reports, with the continuous increase in the length and complexity of document content and the continuous development of current natural language technology, users usually describe their needs as a paragraph of text to retrieve related content, which is closer to the form of human-computer dialogue questioning and searching, and also puts higher requirements on the accuracy of the search return results. In addition, the returned results also need to have better readability to facilitate quick identification of whether the results meet the query requirements.
[0066] Based on this, in view of the above-mentioned demand changes and the problems that long texts such as reports have high content complexity, large amount of text content information, high retrieval complexity, and low retrieval accuracy, the embodiment of the present application provides an information processing method, which can be applied to information processing equipment, with reference to Figure 1 As shown, the method includes the following steps:
[0067] Step 101: Receive a target request sent by a user.
[0068] In an embodiment of the present application, the target request may include a processing request and a retrieval request; after receiving the target request sent by the user, the type of request is first determined, that is, whether the request is a processing request or a retrieval request, and then different processing can be performed for different types of requests.
[0069] It should be noted that steps 102 to 104 may be executed after step 101 , and steps 105 to 107 may be executed after step 101 .
[0070] Step 102 : For a target request being a processing request, each to-be-processed page in the to-be-processed report corresponding to the processing request is processed based on multiple scanning methods and target language models to obtain a target recognition result for each to-be-processed page.
[0071] In an embodiment of the present application, when the request is a processing request, the processing request may carry a report to be processed, that is, a series of operations may be performed on the report to be processed subsequently; the report to be processed may include multiple pages to be processed, and the report to be processed may be an extra-long text including more pages to be processed and more words; the multiple scanning methods may include a first scanning method and a second scanning method; the target language model may refer to a generative pre-trained transformer (GPT), which can be used to determine the smoothness of each scanned content of the report to be processed; the target recognition result may represent the scanned content of the page to be processed, and each page to be processed will obtain a target recognition result.
[0072] In an embodiment of the present application, a report to be processed can be obtained first. For each page to be processed in the report to be processed, the first scanning method and the second scanning method can be used to scan the report to be processed respectively to obtain multiple scanning results. Then, the multiple scanning results can be input into the target language model to obtain multiple processing results. Then, one scanning result can be selected from the multiple scanning results as the target recognition result based on the multiple processing results.
[0073] Step 103: Based on the multiple target recognition results, determine the similarity between the report to be processed and each target report in the first database.
[0074] In an embodiment of the present application, the first database may include multiple target reports, and the target reports are reports already stored in the first database, and the first database may be a vector database; for each target recognition result, the hash value of each to-be-processed page may be determined, and then the hash value of the to-be-processed report may be determined based on the multiple hash values of the multiple to-be-processed pages, and then the hash value of each target report may be determined, and then the similarity between the to-be-processed report and each target report may be determined based on the hash value of the to-be-processed report and the hash value of each target report. In a feasible implementation, the first database may specifically be an ES vector database.
[0075] Step 104: If each similarity is less than the target similarity, store the multiple target recognition results in the first database.
[0076] In an embodiment of the present application, the target similarity can be determined based on historical experiments; each similarity can be compared with the target similarity. If each similarity is less than the target similarity, it means that there is no report similar to the report to be processed in the first database. At this time, multiple target recognition results of the report to be processed can be added to the first database.
[0077] Step 105 : For a target request that is a first search request, determine a plurality of first keywords in the to-be-searched sentence corresponding to the first search request and first categories corresponding to the plurality of first keywords.
[0078] In an embodiment of the present application, the first search request may carry a statement to be searched, that is, the statement to be searched (i.e., query) may be obtained from the first search request; the first search request may refer to the search request sent by the user for the first time; the statement to be searched may refer to the search information entered by the user; the first category may refer to the industry category to which multiple keywords belong.
[0079] In an embodiment of the present application, when the target request is a first search request, the obtained statement to be searched can be input into the target search model to obtain multiple first keywords, and the categories to which the multiple keywords belong can be determined based on the multiple keywords; in a feasible implementation method, when the multiple keywords can be millet and rice, the first category can be determined to be a food industry report; when the multiple keywords can be millet and mobile phones, the first category can be determined to be an electronic consumer products industry report.
[0080] Step 106: Determine the first search content of each report to be selected from the first database based on the multiple first keywords and the first category.
[0081] In an embodiment of the present application, the report to be selected may refer to a report selected from the first database based on multiple first keywords and the first category; the first search content may be the search content retrieved in response to the user's first search request, and for each report to be selected, the first search content may refer to a coherent content summary description of the report to be selected. In an embodiment of the present application, the report to be selected may be first selected from the first database based on multiple first keywords and the first category, and then a target number of paragraphs may be selected from the report to be selected, and then the target number of paragraphs may be enhanced and optimized through the target language model to generate a coherent content summary description.
[0082] Step 107: Based on the user identifier and the identifier of the first search request, the first search content and the report name and access information corresponding to the first search content are sent to the user.
[0083] In an embodiment of the present application, the report name corresponding to the first search content may refer to the report name of the source report of the first search content; the access information may refer to the access address where the source report can be accessed; the signature may be verified based on the user's identifier in the search statement entered by the user and the identifier of the first search request, and the results of the user's search at that time may be returned, including at least the first search content, the report name corresponding to the first search content, and the access address information; it should be noted that by returning information based on the user's identifier and the identifier of the first search request, it can be ensured that the returned information is retrieved by the user and is retrieved by the user this time, thereby improving the accuracy of the returned information. In a feasible implementation method, the access information may refer to a full-text download link for the report.
[0084] The information processing method provided in the embodiment of the present application increases the similarity calculation of the report content, thereby only storing reports with a similarity less than the target similarity when storing the target recognition results. That is, only non-duplicate reports are stored, thereby improving the quality of the database, thereby improving the retrieval efficiency during subsequent retrieval through the database, and overcoming the problem of high retrieval complexity in the prior art. At the same time, it not only supports complex retrieval of multiple keywords, but also uses multiple first keywords and first categories to jointly understand the intention of the user's statement to be retrieved, thereby overcoming the problem of low retrieval accuracy in the prior art.
[0085] Based on the above embodiments, the present application provides an information processing method, referring to Figure 2 As shown, the method includes the following steps:
[0086] Step 201: The information processing device receives a target request sent by a user.
[0087] It should be noted that steps 202 to 212 may be performed after step 201, and steps 213 to 220 may also be performed after step 201;
[0088] It should be noted that the target request is a processing request. For the specific processing flow, please refer to Figure 2 and Figure 3 As shown, the details are as follows:
[0089] Step 202: For each pending page in the pending report corresponding to the processing request, the information processing device uses a first scanning method and a second scanning method to extract content from the pending page, respectively, to obtain a first recognition result and a second recognition result.
[0090] In an embodiment of the present application, the first scanning mode may refer to ① row traversal scanning from left to right as shown in Figure 4(a), and in accordance with ② scanning order from top to bottom; the second scanning mode may refer to ① vertical traversal scanning from top to bottom as shown in Figure 4(b), and in accordance with ② scanning order from left to right.
[0091] In the embodiment of the present application, after obtaining the report to be processed, the file format of the report to be processed can be converted into a pdf file, that is, the word and ppt files can be converted into pdf files, and then Figure 3 As shown, the pdfplumber module can be used to perform content recognition (such as optical character recognition (OCR)) on the page to be processed in the PDF format. To extract content information, a first scanning method can be used to extract the content of the page to be processed to obtain a first recognition result, and a second scanning method can be used to extract the content of the page to be processed to obtain a second recognition result.
[0092] It should be noted that if Figure 3 As shown, after obtaining the report to be processed, a preliminary duplication check can be performed on the report to be processed and the target report in the first database, that is, the report to be processed can be encoded using the message-digest algorithm 5 (MD5) and matched with the target report already in the first database for duplication verification. If there is no duplication, the subsequent steps can be executed. If there is a duplication, the duplicate identifier of the report to be processed can be returned to the user. In this way, the waste of computing resources and storage resources can be reduced through duplication verification, and multiple identical results can be avoided from being returned when the user searches.
[0093] Step 203: The information processing device determines a first value for the first recognition result and a second value for the second recognition result based on the target language model.
[0094] The first value represents the content fluency of the first recognition result, and the second value represents the content fluency of the second recognition result.
[0095] Step 204: The information processing device filters the first recognition result and the second recognition result based on the first value and the second value to obtain a target recognition result.
[0096] In an embodiment of the present application, the first recognition result and the second recognition result can be respectively input into the target language model for content fluency screening to obtain a first value for the first recognition result and a second value for the second recognition result. The first value and the second value can then be compared, and the recognition result corresponding to the larger of the first and second values is selected as the target recognition result, that is, if the first value is greater than the second value, the first recognition result is determined to be the target recognition result; if the second value is greater than the first value, the second recognition result is determined to be the target recognition result. In a feasible implementation, the target language model can refer to a GPT model.
[0097] Step 205: For each page to be processed, the information processing device determines a plurality of second keywords from the target recognition result based on the target classification model, and determines a first hash value for each second keyword.
[0098] In an embodiment of the present application, the first hash value may refer to the minimum hash value of each second keyword (i.e., the LeanMinHash value); each target recognition result may be input into a target classification model for keyword extraction, i.e., multiple second keywords for each page to be processed may be obtained, and each second keyword of each page to be processed may then be converted into a LeanMinHash value. In one feasible implementation, the target classification model may refer to a Detection Transformer (DERT) model.
[0099] Step 206: The information processing device determines a second hash value for the report to be processed based on the multiple first hash values corresponding to the multiple pages to be processed.
[0100] Step 207: The information processing device determines each similarity based on the second hash value and the target hash value reported by each target.
[0101] In an embodiment of the present application, the second hash value may refer to the hash value of the report to be processed; after obtaining the LeanMinHash value of each page to be processed in step 204, the LeanMinHash values of multiple pages to be processed can be merged into the LeanMinHash of the report to be processed (i.e., the second hash value). Thereafter, the similarity between the report to be processed and each target report can be calculated by comparing the second hash value of the report to be processed with the hash value of each target report in the first database; it should be noted that a similarity association of the reports to be processed can also be established based on the identification of similar reports and stored in the first database to facilitate the recommendation of similar documents in subsequent searches.
[0102] Step 208: The information processing device determines the third keyword of the report name of the report to be processed based on the target classification model.
[0103] In an embodiment of the present application, the third keyword may refer to the keyword of the report name of the report to be processed; the target keyword may refer to the core keyword of the report to be processed; the report name of the report to be processed may be input into the target classification model for keyword extraction to obtain the keyword of the report name.
[0104] Step 209: The information processing device determines a target keyword for the report to be processed based on the multiple second keywords and the third keywords corresponding to the multiple pages to be processed.
[0105] In an embodiment of the present application, the multiple second keywords corresponding to the multiple pages to be processed can be processed first, that is, the multiple second keywords of each page to be processed can be counted, the frequency of occurrence of each second keyword can be calculated, and based on the LeanMinHash value of each second keyword, similar keywords can be merged and calculated, and the frequency of occurrence of similar keywords can be merged to obtain high-frequency content words, and then the high-frequency content words and the third keywords can be merged into the target keywords of the document to be processed.
[0106] Step 210: The information processing device determines a target category for the report to be processed based on the second category of each second keyword.
[0107] In an embodiment of the present application, while the target classification model is used to determine multiple second keywords from the target recognition results in the above step 205, the target classification model can also be used to determine the second category of each keyword. At the same time, the reports to be processed are marked according to the second categories of the multiple second keywords to obtain the target categories of the reports to be processed, such as the relevant industries, report types, report authors or institutions, and report times of the reports to be processed.
[0108] Step 211: For each report to be processed, the information processing device stores the target keyword, target identification content and report name in the first database, stores the basic attribute information of the report to be processed in the second database, and stores the report to be processed in the third database.
[0109] The basic attribute information includes access information of the report to be processed.
[0110] In the embodiment of the present application, the second database may refer to a relational database; the third database may refer to an object storage service (OSS) database; in the embodiment of the present application, Figure 3 As shown, the target keyword, target identification content and report name can be stored in the first database. At the same time, the category of the target to be reported can also be stored in the first database. Figure 3 As shown, basic attribute information of pending reports can be extracted and stored in a relational database. This basic attribute information can include information such as report name, size, author, format, time, identifier, access address, category, and identifiers of similar reports. This not only covers a variety of document types, supporting the processing of very long documents in Word, PPT, and PDF formats, as well as complex content, but also establishes a vector retrieval model library for the basic attribute information of reports, resolving issues where complex and diverse content leads to incomplete retrieval information coverage, further improving the coverage and accuracy of retrieval results. In one feasible implementation, the relational database can specifically be a MySQL relational database.
[0111] In an embodiment of the present application, the pending report may be stored in the OSS database, and if there are images and / or charts in the pending report, the images and / or charts may be captured from the pending report and stored in the OSS database.
[0112] It should be noted that for the pending reports uploaded by users, the super-long text is disassembled and calculated by executing the various processing nodes of the above report processing process (i.e., steps 202 to 211), so that the corresponding report content and documents can be quickly retrieved when the report is retrieved, and the batch calculation process for multiple industry reports is the same as above.
[0113] It should be noted that the target request is a retrieval request. For the specific processing flow, please refer to Figure 2 and Figure 5 As shown, the details are as follows:
[0114] Step 212: When the target request is a first search request, the information processing device determines a plurality of first keywords in the to-be-searched sentence corresponding to the first search request and a first category corresponding to the plurality of first keywords.
[0115] It should be noted that if Figure 5 As shown, after obtaining the sentence to be searched, the sentence to be searched can be checked for illegal words first to determine whether the sentence to be searched includes illegal words and blacklisted words, to ensure that only compliant content is subject to subsequent search processing, and the word "not searchable" is directly returned to the user for non-compliant content.
[0116] Step 213: The information processing device determines the search intention information for the sentence to be searched based on the multiple first keywords and the first category.
[0117] In an embodiment of the present application, the part of speech and contextual semantics of multiple first keywords and the first category can be analyzed together to determine the search intent information of the sentence to be retrieved. For example, the keyword combination of millet and rice is a food industry report, which means that the user's search intention is to search for a report on the food industry; the keyword combination of millet and mobile phone is an electronic consumer product industry report, which means that the user's search intention is to search for a report on electronic consumer products. In this way, by searching multiple keywords, the part of speech and contextual semantics are understood, complex searches of multiple keywords are supported, and the user's search intention is accurately understood, thereby improving the accuracy of the search results.
[0118] Step 214 : The information processing device determines a plurality of target paragraphs for each candidate report based on the search intention information and the target keywords of each target report.
[0119] In an embodiment of the present application, the report to be selected may be a report screened out from multiple target reports, and the retrieval intent information may be matched with the target keywords of each target report. When there is a matching first report, the first matching report is determined to be the report to be selected; when there is no matching report, the report in the first category among the multiple target reports is determined to be the report to be selected; and then multiple target paragraphs may be selected from each report to be selected.
[0120] It should be noted that step 214 can be implemented in the following ways:
[0121] Step 214A1: The information processing device matches the search intent information with the target keyword of each target report to obtain a matching result.
[0122] In the embodiment of the present application, the matching results may include reports that are matched and reports that are not matched.
[0123] It should be noted that step 214A2 may be executed after step 214A1, and step 214A3 may be executed after step 214A1.
[0124] Step 214A2: If the matching result indicates that there is no first report matching the search intent information among the multiple reports to be selected, the information processing device determines multiple target paragraphs based on the report content of each report corresponding to the multiple first keywords and the first category.
[0125] The reports to be selected include each report corresponding to the first category.
[0126] In an embodiment of the present application, if the matching result indicates that there is no first report matching the retrieval intent information among multiple candidate reports, indicating that there is no report matching the target keyword with the retrieval intent information among multiple target reports, then all reports under the first category can be determined from multiple target reports, and multiple first keywords can be vector matched with the report content of each report under the first category, and each report under the first category can be arranged in reverse order according to the number of keyword hits, and each report under the first category can be screened and matched according to the keyword hit rate to a target number of paragraphs with different keyword hit numbers and the highest keyword repetition rate; in a feasible implementation method, the target number can be 3.
[0127] Step 214A3: If the matching result indicates that the first report exists in the plurality of candidate reports, the information processing device determines a plurality of target paragraphs based on the plurality of first keywords and the report content of each first report.
[0128] Among them, the selected reports include the first report.
[0129] In an embodiment of the present application, the matching result indicates that the existence of the first report in multiple candidate reports indicates that there is a report that matches the search intent information in multiple target reports; in the case that there is a first report that matches the search intent information in multiple target reports, multiple first keywords can be vector matched with the report content of each first report, and then the first reports can be arranged in reverse order according to the number of keyword hits, and the ranking weights can be weighted according to whether the keywords hit the title, author / institution, year and time of the first report to optimize the result sorting, and then each first report is screened according to the keyword hit rate to match the target number of paragraphs with different keyword hits and the highest keyword repetition rate; in this way, by increasing the similarity calculation of the report content, the similarity calculation can be used to identify duplicate content under the same report, assist in improving content quality and keyword weight configuration, and can help identify duplicates between different reports, improve database quality, and further improve retrieval efficiency. Search recommendations can be performed between similar reports to improve retrieval quality.
[0130] In one possible implementation, the target number may be 3.
[0131] Step 215 : The information processing device performs search enhancement generation on multiple target paragraphs based on the target language model for each candidate report to obtain first search content.
[0132] In an embodiment of the present application, for each report to be selected, multiple target paragraphs corresponding to the report can be input into the target language model to generate an optimized and fluent content summary description of the report (i.e., the first search content). In this way, by performing retrieval-augmented generation (RAG) on the multiple paragraphs retrieved, a document summary of each report to be selected is obtained, thereby improving the accuracy of the retrieval and the readability of the retrieval results, and quickly understanding the core content of ultra-long documents, helping to identify the intended retrieval results.
[0133] Step 216: The information processing device sends the first search content and the report name and access information corresponding to the first search content to the user based on the user identifier and the identifier of the first search request.
[0134] In the embodiment of the present application, through the above search, multiple candidate reports may be obtained, that is, multiple first search contents will be obtained. Later, when sending the first search contents and the report names and access information corresponding to the first search contents to the user, not only the following steps should be performed: Figure 5 The signature verification shown is to send information according to the user's identification and the identification of the first search request, and can return the first search content, report name and access information corresponding to each selected report in sequence according to the report order determined in step 215A2 or step 215A3.
[0135] It should be noted that the above embodiment may further include the following steps:
[0136] Step 217: The information processing device stores the first search content and the plurality of first keywords in a temporary memory.
[0137] In the embodiment of the present application, after obtaining the first search content and the plurality of first keywords through each search, the first search content and the plurality of first keywords may be stored in a temporary memory.
[0138] Step 218: When the information processing device receives the second search request sent by the user, it determines the relevance between each first keyword and each third keyword corresponding to the second search request.
[0139] Step 219 : When there is a correlation that satisfies the target correlation, the information processing device determines a second search content for the second search request from the temporary memory based on the plurality of third keywords and the plurality of first search content.
[0140] In an embodiment of the present application, a second search request may refer to a search initiated by a user after a search request has been made; relevance may refer to the degree of correlation between keywords; target relevance may be a relevance threshold set based on business needs; relevance meeting the target relevance may mean that the relevance is greater than or equal to the target relevance, and the relevance being greater than or equal to the target relevance indicates that there is relevance between the user's first search and the second search; when the user performs a new round of search (i.e., a second search request), the above steps 212 and 213 may be repeated, and the determined new search keyword (i.e., the third keyword) and the previous search records recorded in the data buffer (i.e., multiple first keywords and multiple first search contents) may be judged based on the relevance whether to merge the queries; if the relevance meets the target relevance, the queries may be merged, and it is only necessary to execute the steps corresponding to the user's search intention for the previous return results recorded in the buffer to provide the user with more accurate query results (i.e., the second search content).
[0141] In a feasible implementation, when the user performs a new round of search, if the user intends to query more results, the relevance query step is executed, that is, based on the similar report identifier stored in the buffer according to the previous result report recorded in the buffer, all similar report data are queried (excluding the report results that have been returned), and then the above steps 214 to 217 are executed to return recommended results with high relevance (that is, the second search content).
[0142] It should be noted that, through the above steps 212 to 219 , the user can quickly and accurately search for a report document with an extremely long text using a descriptive statement of the report (ie, the statement to be searched).
[0143] In other embodiments of the present application, the overall system link architecture of the present application is as follows: Figure 6As shown, it specifically includes a report processing module, a report search module, a message distribution and security gateway module, a processing center module, a temporary register module and three different types of databases (i.e., a first database, a second database and a third database). The report to be processed is processed by the report processing module, and different contents are written into different types of databases (i.e., steps 202 to 211 above); the statement to be retrieved is processed by the report search module (i.e., steps 212 to 219 above); the message distribution / security gateway module performs security verification on the transmitted data and distributes different types of messages (processing requests, retrieval requests, etc.). Specifically, different types of messages can be managed and distributed, pointing to different processing links, mainly including report processing requests, report data writing requests, report data editing requests, report search / result push requests, etc.; and a communication security gateway is integrated therein to perform authentication verification on the message, including: user account verification, request Token verification, data reading Write security checks, etc., to ensure the security of data transmission and data storage; the main function of the data processing center is to allocate and process different requests, perform multi-task concurrent orchestration and scheduling of report processing tasks and report retrieval tasks, and perform orchestration and control of the link nodes of each task to improve execution efficiency. At the same time, it combines computing resource scheduling to optimize resources and improve the efficiency of long text processing and report retrieval; and through the mutual cooperation with three different types of databases, it realizes the read and write management and query optimization management of different types of data; and through the mutual cooperation with the cache, it realizes the caching of retrieval data; supports users to conduct continuous queries on the results of the first query, provides more accurate query results, and also supports users to query more similar results through the data cache. In this way, after the report data completed by the report processing module is stored in the database, the user can query the report data in the database through the report search module, forming a closed loop from the processing of report documents to the retrieval of report documents.
[0144] It should be noted that, for the description of the same steps and contents in this embodiment as those in other embodiments, reference can be made to the description in other embodiments and will not be repeated here.
[0145] The information processing method provided in the embodiment of the present application increases the similarity calculation of the report content, thereby only storing reports with a similarity less than the target similarity when storing the target recognition results. That is, only non-duplicate reports are stored, thereby improving the quality of the database, thereby improving the retrieval efficiency during subsequent retrieval through the database, and overcoming the problem of high retrieval complexity in the prior art. At the same time, it not only supports complex retrieval of multiple keywords, but also uses multiple first keywords and first categories to jointly understand the intention of the user's statement to be retrieved, thereby overcoming the problem of low retrieval accuracy in the prior art.
[0146] Based on the above embodiments, the present invention provides an information processing device, which can be applied to Figure 1 and Figure 2 In the information processing method provided in the corresponding embodiment, refer to Figure 7 As shown, the information processing device 3 may include: a receiving unit 31, a first processing unit 32, a storage unit 33, a second processing unit 34 and a sending unit 35, wherein:
[0147] The receiving unit 31 is configured to receive a target request sent by a user;
[0148] The first processing unit 32 is configured to process each to-be-processed page in the to-be-processed report corresponding to the target request, based on multiple scanning methods and target language models, to obtain a target recognition result for each to-be-processed page;
[0149] The first processing unit 32 is further configured to determine a similarity between the report to be processed and each target report in the first database based on the plurality of target recognition results;
[0150] A storage unit 33, configured to store the target recognition result in the first database if each similarity is less than the target similarity;
[0151] The second processing unit 34 is configured to determine, when the target request is a first search request, a plurality of first keywords in the to-be-searched sentence corresponding to the first search request and first categories corresponding to the plurality of first keywords;
[0152] The second processing unit 34 is further configured to determine a first search content of each candidate report from the first database based on the plurality of first keywords and the first category;
[0153] The sending unit 35 is configured to send the first search content and the report name and access information corresponding to the first search content to the user based on the user identifier and the identifier of the first search request.
[0154] In other embodiments of the present application, the first processing unit 32 is specifically configured to perform the following steps:
[0155] For each page to be processed, extracting content from the page to be processed using a first scanning method and a second scanning method, respectively, to obtain a first recognition result and a second recognition result;
[0156] Determining, based on the target language model, a first value for the first recognition result and a second value for the second recognition result, respectively; wherein the first value represents the content fluency of the first recognition result, and the second value represents the content fluency of the second recognition result;
[0157] Based on the first value and the second value, a target recognition result is obtained by filtering the first recognition result and the second recognition result.
[0158] In other embodiments of the present application, the first processing unit 32 is specifically configured to perform the following steps:
[0159] For each page to be processed, determine a plurality of second keywords from the target recognition result based on the target classification model, and determine a first hash value for each second keyword;
[0160] Determining a second hash value for the report to be processed based on a plurality of first hash values corresponding to a plurality of pages to be processed;
[0161] Each similarity is determined based on the second hash value and the target hash value corresponding to each target report.
[0162] In other embodiments of the present application, the first processing unit 32 is specifically configured to perform the following steps:
[0163] determining a third keyword for a report name of a report to be processed based on the target classification model;
[0164] Determining target keywords for the report to be processed based on the plurality of second keywords and third keywords corresponding to the plurality of pages to be processed;
[0165] determining a target category for the report to be processed based on the second category of each second keyword;
[0166] For each report to be processed, the target keyword, target identification content and report name are stored in the first database, and the basic attribute information of the report to be processed is stored in the second database, and the report to be processed is stored in the third database; wherein the basic attribute information includes access information of the report to be processed.
[0167] In other embodiments of the present application, the second processing unit 34 is specifically configured to perform the following steps:
[0168] Determining search intent information for the sentence to be searched based on the plurality of first keywords and the first category;
[0169] determining a plurality of target paragraphs of each candidate report based on the search intent information and the target keywords in each target report;
[0170] For each candidate report, multiple target paragraphs are searched and enhanced based on the target language model to obtain the first search content.
[0171] In other embodiments of the present application, the second processing unit 34 is specifically configured to perform the following steps:
[0172] Match the search intent information with each target keyword to obtain matching results;
[0173] If the matching result indicates that there is no first report matching the search intent information among the multiple candidate reports, determining multiple target paragraphs based on the report content of each report corresponding to the multiple first keywords and the first category;
[0174] If the matching result indicates that the first report exists in the plurality of candidate reports, a plurality of target paragraphs are determined based on the plurality of first keywords and the report content of each first report.
[0175] In other embodiments of the present application, the second processing unit 34 is specifically configured to perform the following steps:
[0176] Storing the first search content and the plurality of first keywords in a temporary memory;
[0177] When receiving a second search request sent by the user, determining a relevance between each first keyword and each third keyword corresponding to the second search request;
[0178] When there is a relevance that satisfies the target relevance, second search content for the second search request is determined from the temporary memory based on the plurality of third keywords and the plurality of first search content.
[0179] It should be noted that the specific description of the steps performed by each unit can be referred to Figure 1 and Figure 2 The information processing method provided in the corresponding embodiment will not be described in detail here.
[0180] The information processing device provided in the embodiment of the present application increases the similarity calculation of the report content, thereby only storing reports with a similarity less than the target similarity when storing the target recognition results. That is, only non-duplicate reports are stored, thereby improving the quality of the database, thereby improving the retrieval efficiency during subsequent retrieval through the database, and overcoming the problem of high retrieval complexity in the prior art. At the same time, it not only supports complex retrieval of multiple keywords, but also uses multiple first keywords and first categories to jointly understand the intention of the user's statement to be retrieved, thereby overcoming the problem of low retrieval accuracy in the prior art.
[0181] Based on the above embodiments, the embodiments of the present application provide an information processing device, which can be applied to Figure 1 and Figure 2 In the information processing method provided in the corresponding embodiment, refer to Figure 8 As shown, the information processing device 4 may include: a processor 41, a memory 42 and a communication bus 43, wherein:
[0182] The communication bus 43 is used to realize the communication connection between the processor 41 and the memory 42;
[0183] The processor 41 is used to execute the information processing program in the memory 42 to implement the following steps:
[0184] Receive target requests sent by users;
[0185] For a target request being a processing request, each to-be-processed page in the to-be-processed report corresponding to the processing request is processed based on multiple scanning methods and target language models to obtain a target recognition result for each to-be-processed page;
[0186] Determining, based on the plurality of target recognition results, a similarity between the report to be processed and each target report in the first database;
[0187] If each similarity is less than the target similarity, storing the target recognition result in the first database;
[0188] For a target request that is a first search request, determining a plurality of first keywords in a to-be-searched sentence corresponding to the first search request and first categories corresponding to the plurality of first keywords;
[0189] Determining first search content of each to-be-selected report from the first database based on a plurality of first keywords and the first category;
[0190] Based on the user identifier and the identifier of the first search request, the first search content and the report name and access information corresponding to the first search content are sent to the user.
[0191] In other embodiments of the present application, the processor 41 is configured to execute the information processing program in the memory 42 based on multiple scanning modes and target language models, process each to-be-processed page in the to-be-processed report corresponding to the processing request, and obtain a target recognition result for each to-be-processed page, so as to implement the following steps:
[0192] For each page to be processed, extracting content from the page to be processed using a first scanning method and a second scanning method, respectively, to obtain a first recognition result and a second recognition result;
[0193] Determining, based on the target language model, a first value for the first recognition result and a second value for the second recognition result, respectively; wherein the first value represents the content fluency of the first recognition result, and the second value represents the content fluency of the second recognition result;
[0194] Based on the first value and the second value, a target recognition result is obtained by filtering the first recognition result and the second recognition result.
[0195] In other embodiments of the present application, the processor 41 is configured to execute the information processing program in the memory 42 to determine the similarity between the report to be processed and each target report in the first database based on the multiple target recognition results, so as to implement the following steps:
[0196] For each page to be processed, determine a plurality of second keywords from the target recognition result based on the target classification model, and determine a first hash value for each second keyword;
[0197] Determining a second hash value for the report to be processed based on a plurality of first hash values corresponding to a plurality of pages to be processed;
[0198] Each similarity is determined based on the second hash value and the target hash value corresponding to each target report.
[0199] In other embodiments of the present application, the processor 41 is configured to execute the information processing method of the information processing program in the memory 42 to implement the following steps:
[0200] determining a third keyword for a report name of a report to be processed based on the target classification model;
[0201] Determining target keywords for the report to be processed based on the plurality of second keywords and third keywords corresponding to the plurality of pages to be processed;
[0202] determining a target category for the report to be processed based on the second category of each second keyword;
[0203] For each report to be processed, the target keyword, target identification content and report name are stored in the first database, and the basic attribute information of the report to be processed is stored in the second database, and the report to be processed is stored in the third database; wherein the basic attribute information includes access information of the report to be processed.
[0204] In other embodiments of the present application, the processor 41 is configured to execute the information processing program in the memory 42 to determine the first search content of each candidate report from the first database based on the multiple first keywords and the first category, so as to implement the following steps:
[0205] Determining search intent information for the sentence to be searched based on the plurality of first keywords and the first category;
[0206] determining a plurality of target paragraphs for each candidate report based on the search intent information and the target keywords of each target report;
[0207] For each candidate report, multiple target paragraphs are searched and enhanced based on the target language model to obtain the first search content.
[0208] In other embodiments of the present application, the processor 41 is configured to execute the information processing program in the memory 42 to determine multiple target paragraphs of each of the candidate reports based on the search intent information and the target keywords of the multiple candidate reports corresponding to the first category in the multiple target reports, so as to implement the following steps:
[0209] Match the search intent information with each target keyword to obtain matching results;
[0210] If the matching result indicates that there is no first report matching the search intent information among the multiple candidate reports, determining multiple target paragraphs based on the report content of each report corresponding to the multiple first keywords and the first category;
[0211] If the matching result indicates that the first report exists in the plurality of candidate reports, a plurality of target paragraphs are determined based on the plurality of first keywords and the report content of each first report.
[0212] In other embodiments of the present application, the processor 41 is configured to execute the information processing program in the memory 42 based on the user identifier and the identifier of the first search request, and send the first search content and the report name and access information corresponding to the first search content to the user, so as to implement the following steps:
[0213] Storing the first search content and the plurality of first keywords in a temporary memory;
[0214] When receiving a second search request sent by the user, determining a relevance between each first keyword and each third keyword corresponding to the second search request;
[0215] When there is a relevance that satisfies the target relevance, second search content for the second search request is determined from the temporary memory based on the plurality of third keywords and the plurality of first search content.
[0216] It should be noted that the specific description of the steps performed by the processor can be referred to Figure 1 and Figure 2 The information processing method provided in the corresponding embodiment will not be described in detail here.
[0217] The information processing device provided in the embodiment of the present application increases the similarity calculation of the report content, thereby only storing reports with a similarity less than the target similarity when storing the target recognition results. That is, only non-duplicate reports are stored, thereby improving the quality of the database, thereby improving the retrieval efficiency during subsequent retrieval through the database, and overcoming the problem of high retrieval complexity in the prior art. At the same time, it not only supports complex retrieval of multiple keywords, but also uses multiple first keywords and first categories to jointly understand the intention of the user's statement to be retrieved, thereby overcoming the problem of low retrieval accuracy in the prior art.
[0218] Based on the above embodiments, the embodiments of the present application provide a computer-readable storage medium, which stores one or more programs, and the one or more programs can be executed by one or more processors to implement Figure 1 and Figure 2 The corresponding embodiments provide steps of the information processing method.
[0219] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of hardware embodiments, software embodiments, or embodiments combining software and hardware. Furthermore, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage and optical storage, etc.) containing computer-usable program code.
[0220] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0221] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0222] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0223] The above description is merely a preferred embodiment of the present application and is not intended to limit the scope of protection of the present application.
Claims
1. An information processing method, characterized in that: The method comprises: Receive target requests sent by users; When the target request is a processing request, each to-be-processed page in the to-be-processed report corresponding to the processing request is processed based on multiple scanning methods and target language models to obtain a target recognition result for each to-be-processed page; determining, based on the plurality of target recognition results, a similarity between the report to be processed and each target report in the first database; If each of the similarities is less than the target similarity, storing the plurality of target recognition results in the first database; When the target request is a first search request, determining a plurality of first keywords in a to-be-searched sentence corresponding to the first search request and first categories corresponding to the plurality of first keywords; determining a first search content of each to-be-selected report from the first database based on the plurality of first keywords and the first category; Based on the identifier of the user and the identifier of the first search request, the first search content and the report name and access information corresponding to the first search content are sent to the user.
2. The method according to claim 1, characterized in that Processing each to-be-processed page in the to-be-processed report corresponding to the processing request to obtain a target recognition result for each to-be-processed page includes: For each of the pages to be processed, extracting content from the page to be processed using a first scanning method and a second scanning method, respectively, to obtain a first recognition result and a second recognition result; Based on the target language model, determining a first value for the first recognition result and a second value for the second recognition result, respectively; wherein the first value represents the content fluency of the first recognition result, and the second value represents the content fluency of the second recognition result; The target recognition result is obtained by filtering the first recognition result and the second recognition result based on the first value and the second value.
3. The method according to claim 1, characterized in that The determining, based on the plurality of target recognition results, the similarity between the report to be processed and each target report in the first database includes: For each of the to-be-processed pages, determining a plurality of second keywords from the target recognition results based on the target classification model, and determining a first hash value for each of the second keywords; determining a second hash value for the report to be processed based on a plurality of first hash values corresponding to a plurality of pages to be processed; Each of the similarities is determined based on the second hash value and the target hash value of each target report.
4. The method according to claim 3, characterized in that After determining the first Hash value for each second keyword, the method further includes: determining, based on the target classification model, a third keyword for the report name of the report to be processed; determining a target keyword for the report to be processed based on the plurality of second keywords and the third keyword corresponding to the plurality of pages to be processed; determining a target category for the report to be processed based on the second category of each of the second keywords; For each of the pending reports, the target keyword, the target identification content and the report name are stored in the first database, and the basic attribute information of the pending report is stored in the second database, and the pending report is stored in the third database; wherein the basic attribute information includes the access information of the pending report.
5. The method according to claim 1, characterized in that Determining first search content of each report to be selected from the first database includes: Determining search intent information for the sentence to be searched based on the multiple first keywords and the first category; determining the plurality of target paragraphs of each of the candidate reports based on the search intent information and the target keywords of each of the target reports; For each of the candidate reports, retrieval enhancement generation is performed on the multiple target paragraphs based on the target language model to obtain the first retrieval content.
6. The method according to claim 5, characterized in that Determining the plurality of target paragraphs of each of the candidate reports includes: Matching the search intent information with each target keyword to obtain a matching result; If the matching result indicates that there is no first report matching the search intent information among the multiple candidate reports, determining the multiple target paragraphs based on the multiple first keywords and the report content of each report corresponding to the first category; If the matching result indicates that the first report exists in the plurality of reports to be selected, the plurality of target paragraphs are determined based on the plurality of first keywords and the report content of each of the first reports.
7. The method according to claim 1, characterized in that After sending the first search content and the report name and access information corresponding to the first search content to the user, the method further includes: Storing the first search content and the plurality of first keywords in a temporary memory; When receiving a second search request sent by the user, determining a relevance between each first keyword and each third keyword corresponding to the second search request; When the relevance satisfies the target relevance, second search content for the second search request is determined from the temporary memory based on a plurality of third keywords and a plurality of the first search content.
8. An information processing device, characterized in that The device comprises: A receiving unit, configured to receive a target request sent by a user; A first processing unit is configured to process each to-be-processed page in the to-be-processed report corresponding to the target request, based on multiple scanning modes and target language models, to obtain a target recognition result for each to-be-processed page, when the target request is a processing request; The first processing unit is further configured to determine a similarity between the report to be processed and each target report in the first database based on the plurality of target recognition results; a storage unit, configured to store the target recognition result in the first database if each of the similarities is less than the target similarity; a second processing unit, configured to determine, when the target request is a first search request, a plurality of first keywords in a to-be-searched statement corresponding to the first search request and a first category corresponding to the plurality of first keywords; The second processing unit is further configured to determine a first search content of each to-be-selected report from the first database based on the plurality of first keywords and the first category; A sending unit is configured to send the first search content and a report name and access information corresponding to the first search content to the user based on the user identifier and the identifier of the first search request.
9. An information processing device, characterized in that The device includes: a processor, a memory and a communication bus; The communication bus is used to realize the communication connection between the processor and the memory; The processor is configured to execute the information processing program in the memory to implement the steps of the information processing method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the information processing method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Method and equipment for realizing full-text retrieval of character picture and image type scanning copy
CN113806472A
Document analysis system for integration of paper records into a searchable electronic database
US20070168382A1