Report generation method, apparatus, device, storage medium and program product
By receiving target report generation requests from user terminals and using large language models to generate reports, the problems of low efficiency and poor accuracy caused by manual extraction and analysis of information in the prior art are solved, and more efficient and accurate report generation is achieved, improving user experience.
Patent Information
- Application Number
- CN202411024952.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-29
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-07-29
AI Technical Summary
In the prior art, manual extraction and analysis of information are relied on inexpensive report generation efficiency and poor content accuracy, which affects the user experience.
Provide a report generation method, by receiving target report generation requests from user terminals, obtaining preset dimension information and problem information, and generating reports using a large language model. The method includes obtaining a text data set, searching based on preset dimension information, merging and reordering processing, and splicing the processed documents and problem information into input information, and generating a final target report.
Improve the efficiency of report generation and the accuracy of report content, improve user experience, and reduce the need for manual intervention.
Smart Images

Figure CN119003692B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a report generation method, apparatus, device, storage medium, and program product. Background Art
[0002] With the continuous development of the financial field and computer technology, major financial enterprises need to provide professional consulting services to customers, which involves the writing of reports so that customers can obtain relevant information based on the reports.
[0003] Currently, when generating a report, generally, researchers manually extract relevant information corresponding to each dimension from a large number of files based on different dimensions, and conduct manual analysis based on the extraction results to obtain analysis information corresponding to each dimension, and then generate a report based on the analysis information of each dimension.
[0004] However, this method of generating reports relies on manual extraction and analysis of information, which reduces the efficiency of report generation and the accuracy of report content, affecting the user experience. Summary of the Invention
[0005] This application provides a report generation method, apparatus, device, storage medium, and program product to solve the technical problem in the prior art that relying on manual extraction and analysis of information reduces the efficiency of report generation and the accuracy of report content.
[0006] In a first aspect, this application provides a report generation method, including: receiving a target report generation request sent by a user terminal; the target report generation request includes preset dimension information and problem information, and the number of the preset dimension information is at least one;
[0007] Obtaining a first text data set and a second text data set, and respectively retrieving in the first text data set and the second text data set based on the preset dimension information using corresponding retrieval strategies to obtain a preset number of preset text data;
[0008] Successively performing a merging process and a preset reordering process on the preset number of preset text data to obtain a target number of target documents;
[0009] Concatenating the target number of target documents and the problem information into input information, and using a preset large language model to generate a target report according to the input information.
[0010] In a second aspect, this application provides a report generation apparatus, including: a receiving module, configured to receive a target report generation request sent by a user terminal; the target report generation request includes preset dimension information and problem information, and the number of the preset dimension information is at least one;
[0011] An acquisition module, configured to acquire a first text data set and a second text data set;
[0012] A retrieval module, configured to retrieve in the first text data set and the second text data set respectively based on the preset dimension information by using corresponding retrieval strategies, so as to obtain a preset number of preset text data;
[0013] A processing module, configured to perform a merging process and a preset reordering process in sequence based on the preset number of preset text data, so as to obtain a target number of target documents;
[0014] A splicing module, configured to splice the target number of target documents and the problem information into input information;
[0015] A generation module, configured to generate a target report by using a preset large language model according to the input information.
[0016] In a third aspect, the present application provides an electronic device, including: a processor, a memory and a transceiver communicatively connected to the processor;
[0017] The memory stores computer execution instructions; the transceiver is configured to transmit and receive data;
[0018] The processor executes the computer execution instructions stored in the memory to implement the method described in the first aspect.
[0019] In a fourth aspect, the present application provides a computer-readable storage medium, in which computer execution instructions are stored, and when the computer execution instructions are executed by a processor, they are used to implement the method described in the first aspect.
[0020] In a fifth aspect, the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the method described in the first aspect.
[0021] The report generation method, device, equipment, storage medium and program product provided by this application receive a target report generation request sent by a user terminal; the target report generation request includes preset dimension information and problem information, and the number of the preset dimension information is at least one; obtain a first text data set and a second text data set, and respectively retrieve in the first text data set and the second text data set based on the preset dimension information by using corresponding retrieval strategies to obtain a preset number of preset text data; perform a merging process and a preset reordering process on the preset number of preset text data in sequence to obtain a target number of target documents; splice the target number of target documents and the problem information into input information, and use a preset large language model to generate a target report according to the input information. Since multiple preset dimension information are preset in advance, and a first text data set and a second text data set are established in advance, by receiving a target report generation request including at least one preset dimension information and problem information sent by a user terminal, it is possible to respectively retrieve in the first text data set and the second text data set based on each preset dimension information by using corresponding retrieval strategies, and obtain a preset number of preset text data associated with each preset dimension information. And by sequentially performing a merging process and a preset reordering process on the preset number of preset text data, it is possible to obtain a target number of target documents associated with each preset dimension information. Thus, by splicing the target number of target documents and the problem information into input information, it is possible to obtain input information associated with each preset dimension information, and further use a preset large language model to generate a target report according to the input information associated with each preset dimension information. There is no need for manual extraction and analysis of information, which improves the efficiency of report generation and the accuracy of report content, and improves the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with this application, and are used together with the specification to explain the principles of this application.
[0023] Figure 1 It is a diagram of the application scenario of the report generation method provided by an embodiment of this application;
[0024] Figure 2 It is the flow of the report generation method provided by an embodiment of this application Figure 1 ;
[0025] Figure 3 It is the flow of the report generation method provided by an embodiment of this application Figure 2 ;
[0026] Figure 4 It is the flow of the report generation method provided by an embodiment of this application Figure 3 ;
[0027] Figure 5 The flowchart of the report generation method provided by the embodiment of the present application Figure 4 ;
[0028] Figure 6 The flowchart of the report generation method provided by the embodiment of the present application Figure 5 ;
[0029] Figure 7 The complete flowchart of the report generation method provided by the embodiment of the present application;
[0030] Figure 8 The structural schematic diagram of the report generation device provided by the embodiment of the present application;
[0031] Figure 9 The structural schematic diagram of the electronic device provided by the embodiment of the present application.
[0032] Through the above-mentioned drawings, the clear embodiments of the present application have been shown, and there will be more detailed descriptions hereinafter. These drawings and textual descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. Detailed Description of the Embodiment
[0033] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0034] In the technical solution of the present application, the processing of the collection, storage, use, processing, transmission, provision, and disclosure of financial data, user data, and other information all complies with the provisions of relevant laws and regulations and does not violate public order and good customs.
[0035] It should be noted that in the embodiments of the present application, some industry-existing solutions such as certain software, components, models, etc. may be mentioned. They should be considered exemplary, and their purpose is only to illustrate the feasibility in the implementation of the technical solution of the present application, but it does not mean that the applicant has already or necessarily used this solution.
[0036] In the prior art, when generating a report, generally, researchers manually extract relevant information in different dimensions from a vast amount of documents according to specific requirements to obtain relevant information corresponding to at least one dimension. Then, further analyze the relevant information corresponding to each of the above dimensions manually to obtain the analysis information corresponding to each dimension. Finally, write a report based on the analysis information of each dimension, thereby generating the report required by the researchers. This requires manual extraction and analysis of information and manual writing of the report, reducing the efficiency of report generation and the accuracy of the report content, and affecting the user experience.
[0037] To address the above technical problems, the present application proposes the following technical concept: In order to improve the efficiency of report generation and the accuracy of the report content, instead of manually extracting and analyzing information and further manually writing the report, multiple dimensions are preset. Researchers can select the required dimensions from multiple dimensions through a terminal and input the questions corresponding to the required dimensions. Then, based on the selection results and questions, trigger the report generation component. Then, in response to the triggering of the above component, automatically obtain the required dimensions and corresponding questions, and obtain multiple pre-processed text data. Then, according to the required dimensions, find the text data that is closest to the required dimensions in each text data. Further, perform merging and re-sorting processing on the closest text data to obtain the final text data. Finally, form a prompt message according to the final text data and the corresponding questions and input it into the large language model, and output a report based on the above large language model, improving the efficiency of report generation and the accuracy of the report content, and improving the user experience.
[0038] Figure 1 It is an application scenario diagram of the report generation method provided by an embodiment of the present application. As Figure 1As shown in the figure, the application scenario includes: a user terminal 1, an electronic device 2, and a database 3. Among them, the user terminal 1 is the terminal where the user with the report generation requirement is located, the electronic device 2 is the device that executes report generation, and the database 3 is the database corresponding to the electronic device 2, which can specifically be a vector database or a preset database. The preset database is a database for storing general data, and the vector database is a database for storing vector data. Among them, the electronic device 2 is communicatively connected to the user terminal 1 and the database 3 respectively. First, the user selects preset dimension information according to the required dimensions through the user terminal, and based on the selection result, raises a question to be asked, and then triggers the target report generation component. Then, in response to the triggering of the above component, the user terminal generates a target report generation request based on the preset dimension information and the question information, and sends it to the electronic device 2. Then, the electronic device 2 receives the target report generation request sent by the user terminal 1, and obtains a first text data set and a second text data set based on the database 3. Based on the preset dimension information, corresponding retrieval strategies are used to retrieve in the first text data set and the second text data set respectively to obtain a preset number of preset text data, and based on the preset number of preset text data, merging processing and preset reordering processing are sequentially performed to obtain a target number of target documents. Finally, the target number of target documents and the question information are spliced into input information, and a preset large language model is used to generate a target report according to the input information. Among them, the number of preset dimension information is at least one.
[0039] The report generation method provided by this application aims to solve the above technical problems in the prior art.
[0040] The technical solution of this application and how the technical solution of this application solves the above technical problems will be described in detail below with specific embodiments. These specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of this application will be described below with reference to the accompanying drawings.
[0041] Figure 2 is the flow of the report generation method provided by the embodiments of this application Figure 1 , as Figure 2 shown, the execution subject of this embodiment is a report generation device, and this report generation device is located in the electronic device, then the report generation method includes:
[0042] S201. Receive the target report generation request sent by the user terminal. The target report generation request includes preset dimension information and question information, and the number of preset dimension information is at least one.
[0043] Among them, the user terminal is the terminal corresponding to the researcher with the report generation requirement.
[0044] Among them, the target report generation request is a request indicating the generation of a target report, and this target report is the report required by the researcher.
[0045] Among them, the preset dimension information is pre-set information used to characterize a dimension, and specifically can be the field where a certain dimension is located. The problem information is the problem information input by the researcher and required to be answered by the target report, and specifically can be the field where the problem is located.
[0046] It can be understood that the problem information input by the researcher may cover multiple dimensions, that is, it is required that the target report cover the content of multiple dimensions. Therefore, the number of preset dimension information is at least one.
[0047] In this embodiment, the researcher can communicate and negotiate with experts in the field in advance to determine multiple preset dimension information and display each preset dimension information in a tree structure, so as to convert each preset dimension information into a selectable component.
[0048] Exemplarily, if the researcher is a researcher in the financial field, communicating and negotiating with experts in the financial field can determine the following nine preset dimension information: operating income and profit performance, cost control ability, operation quality and growth potential, R & D investment and innovation ability, development strategy and business layout, order quality and sales ability, organizational construction and employee motivation, business structure adjustment and layout, industry opportunities and development expectations.
[0049] Based on this, when the researcher has a report generation requirement, the problem information is input through the operation interface corresponding to the user terminal, and at least one required preset dimension information is selected from multiple preset dimension information based on the dimension involved in the problem information, and then the target report generation request is triggered.
[0050] Exemplarily, based on the nine preset dimension information given in the previous example, if the problem information is "Please analyze the strategic development direction for the next year based on the operating income, profit performance, order quality, and sales ability in this quarter", then the required at least one preset dimension information can be operating income and profit performance, order quality and sales ability, and development strategy and business layout.
[0051] Further, in response to receiving the target report generation request sent by the user terminal, the problem information and at least one preset dimension information are obtained based on this request.
[0052] It can be understood that the multiple preset dimension information determined in advance can be increased, deleted, or modified based on specific requirements to improve the flexibility of the required dimensions.
[0053] It can be understood that the above exemplary description is only for illustration and should not constitute any limitation to this application.
[0054] S202. Obtain the first text data set and the second text data set, and respectively retrieve in the first text data set and the second text data set based on the preset dimension information using corresponding retrieval strategies to obtain a preset number of preset text data.
[0055] Among them, the first text data set is a set formed by text data parsed from multiple documents, specifically including a plurality of first text data. The first text data is the text data after document parsing.
[0056] Among them, the second text data set is a set formed by text data parsed and vectorized from multiple documents, specifically including a plurality of second text data. The second text data is the text data after document parsing and vectorization processing.
[0057] In this embodiment, relevant documents in the field can be obtained in advance, and the above documents can be parsed to obtain at least one first text data. Then, each first text data is vectorized to obtain at least one second text data. Further, a first text data set is formed based on each first text data, a second text data set is formed based on each second text data, and the above first text data set and second text data set are stored.
[0058] It can be understood that when storing, it can be specifically stored in a corresponding database or other storage media, and this embodiment does not limit this.
[0059] Based on this, after obtaining the problem information and the preset dimension information, the pre-stored first text data set and second text data set are obtained, and corresponding retrieval strategies are used to respectively retrieve in the first text data set and the second text data set based on the preset dimension information to obtain a preset number of preset text data.
[0060] Among them, the corresponding retrieval strategy is a retrieval strategy set based on different text data. That is to say, for the first text data set and the second text data set, there are corresponding retrieval strategies for retrieval.
[0061] Among them, the preset number is the number of text data that needs to be obtained in the retrieval result set in advance, and the preset text data is the text data retrieved from the first text data set and the second text data set.
[0062] Specifically, retrieve in the first text data set using the corresponding retrieval strategy of the first text data set based on the preset dimension information. At the same time, retrieve in the second text data set using the corresponding retrieval strategy of the second text data set. Then obtain a preset quantity, and obtain preset text data according to the preset quantity from the above two retrieval results, resulting in a preset quantity of preset text data.
[0063] Exemplarily, if the preset quantity is 10, then 10 preset text data are obtained.
[0064] It can be understood that since the number of preset dimension information is at least one, for each preset dimension information, a corresponding preset quantity of preset text data will be obtained.
[0065] It can be understood that the above exemplary description is only for illustration and should not constitute any limitation to this application.
[0066] S203. Perform merging processing and preset reordering processing in sequence based on the preset quantity of preset text data to obtain a target quantity of target documents.
[0067] Among them, the merging processing is a processing method of merging based on the documents to which the preset text data belongs.
[0068] Among them, the reordering processing is a processing method of reordering the merging result.
[0069] Among them, the target quantity is the quantity of documents that need to be obtained in the processing result and is preset, and the target document is the document determined from the documents to which the preset quantity of preset text data belongs.
[0070] It can be understood that the first text data is obtained after parsing the document, and the second text data is further vectorized based on the first text data. Therefore, among the preset quantity of preset text data obtained after retrieving the first text data set and the second text data set, there may be two preset text data belonging to the same document.
[0071] Based on this, after obtaining the preset quantity of preset text data, obtain the documents to which each preset text data belongs, and perform merging processing based on the same belonging documents to obtain at least one belonging document. Further, reorder each belonging document, then obtain the target quantity of belonging documents from the sorted belonging documents, and determine the above target quantity of belonging documents as the target quantity of target documents.
[0072] It can be understood that since the number of preset dimension information is at least one, and for each piece of preset dimension information, there will be a preset number of preset text data corresponding to it. Therefore, after this step is executed, for each piece of preset dimension information, there will also be a target number of target documents corresponding to it.
[0073] S204. Concatenate the target number of target documents and the problem information into input information, and use a preset large language model to generate a target report based on the input information.
[0074] Among them, the input information is information recognizable by the large language model, which can specifically be prompt information. The preset large language model is a large language model pre-trained using natural language text.
[0075] Specifically, after obtaining the target number of target documents, obtain the problem information, and concatenate the target number of target documents and the problem information to obtain the input information.
[0076] Exemplarily, if the problem information is "Please analyze the strategic development direction for the next year based on the operating income, profit performance, order quality, and sales ability in this quarter", and the target number of target documents is the documents related to the preset dimension information of "operating income and profit performance", then concatenate the content of the above documents with "Please analyze the strategic development direction for the next year based on the operating income, profit performance, order quality, and sales ability in this quarter" to obtain the input information, such as "xxx (document content), please analyze the strategic development direction for the next year based on the above content, according to the operating income, profit performance, order quality, and sales ability in this quarter".
[0077] Based on this, obtain the pre-established preset large language model, input the input information into the above preset large language model, and generate a target report based on the association between the problem and the document content in the input information by the above preset large language model.
[0078] It can be understood that since the number of preset dimension information is at least one, and for each piece of preset dimension information, there will be a target number of target documents corresponding to it. Therefore, each piece of preset dimension information will have corresponding input information. Then when generating the target report, if the number of preset dimension information is one, directly generate the target report according to the input information corresponding to this preset dimension information. If the number of preset dimension information is multiple, generate the target report according to the input information corresponding to each preset dimension information.
[0079] The report generation method provided in this embodiment, since multiple preset dimension information is preset in advance and a first text data set and a second text data set are established in advance, by receiving a target report generation request sent by a user terminal that includes at least one preset dimension information and question information, it is possible to retrieve in the first text data set and the second text data set respectively based on each preset dimension information by using corresponding retrieval strategies, and obtain a preset number of preset text data associated with each preset dimension information. And by sequentially performing a merging process and a preset reordering process on the preset number of preset text data, it is possible to obtain a target number of target documents associated with each preset dimension information. Thus, by splicing the target number of target documents and the question information into input information, it is possible to obtain input information associated with each preset dimension information, and further use a preset large language model to generate a target report according to the input information associated with each preset dimension information. There is no need for manual extraction and analysis of information, which improves the efficiency of report generation and the accuracy of report content, and improves the user experience.
[0080] As an alternative embodiment, on the basis of the above embodiment, when obtaining the first text data set and the second text data set, S202 specifically includes the following steps:
[0081] Obtain the first text data set based on a preset database, and obtain the second text data set based on a vector database.
[0082] Among them, the preset database is a pre-set database for storing general data, and specifically can be the enterprise-level distributed database TDSQL. The vector database is a pre-set database for storing vector data, and specifically can be the vector data storage function provided by the analysis engine Elasticsearch.
[0083] In this embodiment, the first text data set can be pre-stored in the preset database, and the second text data can be stored in the vector database. Based on this, in response to obtaining the first text data set and the second text data set, obtain the first text data set in the preset database, and obtain the second text data set in the vector database.
[0084] Correspondingly, before obtaining the first text data set and the second text data set, the following steps are also included:
[0085] Step a1: Obtain a preset document set.
[0086] It can be understood that the prerequisite for obtaining the first text data set and the second text data set is that the first text data set and the second text data set have been stored currently. Therefore, the process of forming and storing the first text data set and the second text data set is described.
[0087] Among them, the preset document set is a set formed by multiple relevant documents in the field, and the preset document is a relevant document in the field.
[0088] Exemplarily, if the above-mentioned field refers to the financial field, the preset documents in the preset document set can be financial report data, listing announcement data, and the corresponding pdf data of the listing announcement data.
[0089] Specifically, according to the location of at least one required preset document, based on the communication connection with each location, each preset document is obtained, and a preset document set is formed based on each preset document, thereby obtaining the preset document set.
[0090] Exemplarily, if at least one required preset document is financial report data, listing announcement data, and the corresponding pdf data of the listing announcement data, then based on the location of the above-mentioned preset documents, the financial report data pre-stored by business personnel can be obtained according to the communication connection with the preset database, and the listing announcement data and the corresponding pdf data of the listing announcement data can be pulled according to the network interface, thereby obtaining each preset document.
[0091] It can be understood that in order to further improve the accuracy of the report content, business personnel can also pre-store the historical reports that have been output in the preset database in advance. When selecting preset documents, historical reports can also be selected, so that the historical reports can also be used as the basis for the target report output this time, increasing the reference content and further improving the accuracy of the report content.
[0092] It can be understood that the above exemplary description is only for illustration and should not constitute any limitation to this application.
[0093] Step a2: Parse and process each preset document in the preset document set to obtain a first text data set.
[0094] Among them, the parsing process is a processing method of parsing the document into corresponding system-recognizable data.
[0095] In this embodiment, after obtaining the preset document set, each preset document is parsed to obtain the text corresponding to each preset document, and the text corresponding to each preset document is determined as each first text data, thereby obtaining the first text data set.
[0096] Among them, the above text can include the main text, title, table, etc., and this embodiment does not make any limitations thereto.
[0097] Among them, when parsing, specifically, the pdfplumber library and PaddleOCR library of Python can be called for parsing, or other methods can be used for parsing, and this embodiment does not make any limitations thereto.
[0098] Step a3: Perform segmentation processing and vectorization processing on each first text data in the first text data set in sequence to obtain a second text data set.
[0099] Among them, the segmentation processing is a processing method for segmenting text. The vectorization processing is a processing method for vectorizing the segmented text.
[0100] In this embodiment, after obtaining the first text data set, the text in each first text data is further obtained, the text except the chapter name is segmented, and spliced with a text length of 128 to obtain the segmentation result corresponding to each first text data. Further, vectorization processing is performed on each segmentation result to obtain each second text data, and thus a second text data set is obtained.
[0101] Among them, when segmenting the text except the chapter name, it can specifically be segmented based on the title, segmented based on the table content, segmented based on the paragraphs included in the text body, and further segmented based on the punctuation marks of sentences in the paragraphs, such as ",", "!", "?", etc. This embodiment does not limit the specific segmentation method.
[0102] Among them, when performing vectorization processing, the open-source Multilingual-E5-large model can specifically be used to perform vectorization processing on the segmented text. Specifically, the segmented text is input into the above model, and vectorization representation of the segmented text is performed based on the above model.
[0103] Step a4: Store the first text data set in a preset database, and store the second text data set in a vector database.
[0104] In this embodiment, after obtaining the first text data set and the second text data set, the first text data set is stored in a preset database based on the general format of each first text data in the first text data set, and the first text data set is stored in a vector database based on the vectorized format of each second text data in the second text data set.
[0105] The report generation method provided in this embodiment, since the prerequisite for obtaining the first and second text data sets is to form and store the first and second text data sets, by obtaining a preset document set, each preset document in it can be parsed to obtain the first text data set, and the first text data can be further segmented and vectorized in sequence to obtain the second text data set. Then, the first text data set is stored in the preset database, and the second text data set is stored in the vector database, ensuring that the first and second text data sets can be successfully obtained subsequently and improving the acquisition success rate. And on this basis, when obtaining the first and second text data sets subsequently, the first text data set can be directly obtained based on the preset database, and the second text data set can be obtained based on the vector database, improving the acquisition efficiency.
[0106] Figure 3 is the flow of the report generation method provided in the embodiment of the present application Figure 2 , such as Figure 3 shown. Based on the above embodiment, in this embodiment, the corresponding retrieval strategy is further refined, and the corresponding retrieval strategy is used to retrieve in the first text data set and the second text data set respectively based on the preset dimension information to obtain a preset number of preset text data for further refinement. In this embodiment, the corresponding retrieval strategy includes a keyword retrieval strategy and a vector retrieval strategy. When using the corresponding retrieval strategy to retrieve in the first text data set and the second text data set respectively based on the preset dimension information to obtain a preset number of preset text data, the specific steps are as follows:
[0107] S301. Use the keyword retrieval strategy to retrieve in the first text data set according to the preset dimension information to obtain a first number of first text data.
[0108] In this embodiment, the corresponding retrieval strategy includes a keyword retrieval strategy and a vector retrieval strategy. Among them, the keyword retrieval strategy is a strategy for retrieving according to the keywords associated with the preset dimension information, and the vector retrieval strategy is a strategy for retrieving according to the vector text associated with the preset dimension information.
[0109] It can be understood that this solution provides two retrieval methods for retrieving the preset document set, that is, using the keyword retrieval strategy to retrieve in the first text data set obtained by converting the preset document set, and using the vector retrieval strategy to retrieve in the second text data set obtained by converting the preset document set, thereby obtaining two retrieval results, and the final retrieval result can be comprehensively obtained based on the two retrieval results, improving the retrieval accuracy.
[0110] Specifically, when retrieving in the first text data set according to the preset dimension information using the keyword retrieval strategy, at least one keyword is determined in the field corresponding to the preset dimension information, and then each keyword is searched for in each first text data, and the matching degree of each first text data and the preset dimension information is determined based on the number of search results. Based on this, each first text data is sorted according to the matching degree, and the first quantity is obtained, and the first quantity of first text data with a higher matching degree is obtained from each first text data.
[0111] Among them, when performing the retrieval, specifically, the term frequency-inverse document frequency TF-IDF technology can be used for the retrieval, and the specific method used during the retrieval in this embodiment is not limited.
[0112] Among them, the first quantity is the number of first text data that needs to be obtained from the first text data set and is set in advance.
[0113] S302. Use the vector retrieval strategy to retrieve in the second text data set according to the preset dimension information to obtain the second quantity of second text data.
[0114] Among them, the second quantity is the number of second text data that needs to be obtained from the second text data set and is set in advance.
[0115] Specifically, the field corresponding to the preset dimension information is vectorized to obtain the corresponding vectorized text. The specific implementation method is similar to the method of vectorizing the first text data in step a3 and will not be elaborated here. Then, the cosine similarity between each second text data and the vectorized text corresponding to the preset dimension information is calculated respectively. Based on this, each second text data is sorted according to the similarity, and the second quantity is obtained, and the second quantity of second text data with a higher similarity is obtained from each second text data.
[0116] Among them, when calculating the cosine similarity between each second text data and the vectorized text corresponding to the preset dimension information, specifically: calculate the product of the dot product and the Euclidean length between each second text data and the vectorized text corresponding to the preset dimension information, take the ratio of the above dot product to the above product of the Euclidean lengths, and determine this ratio as the cosine similarity.
[0117] S303. Determine the first quantity of first text data and the second quantity of second text data as the preset quantity of preset text data to obtain the preset quantity of preset text data.
[0118] In this embodiment, based on obtaining the first quantity of first text data and the second quantity of second text data, the above-mentioned quantity of text data is determined as the preset quantity of preset text data.
[0119] The report generation method provided in this embodiment, since two retrieval strategies are preset in advance, namely the keyword retrieval strategy and the vector retrieval strategy, so by adopting the keyword retrieval strategy to retrieve in the first text data set according to the preset dimension information, the first text data of the first quantity is obtained, and by adopting the vector retrieval strategy to retrieve in the second text data set according to the preset dimension information, the second text data of the second quantity is obtained, the first text data of the first quantity and the second text data of the second quantity can be determined as the preset text data of the preset quantity, and thus two retrieval results are obtained, improving the comprehensiveness of the retrieval.
[0120] As an alternative embodiment, on the basis of the above embodiment, the first text data includes at least one parsed title identifier and the parsed text corresponding to each parsed title. When retrieving in the first text data set according to the preset dimension information by adopting the keyword retrieval strategy, S301 specifically includes the following steps:
[0121] S3011. Extract keywords from the preset dimension information, and obtain at least one parsed title identifier corresponding to each first text data in the first text data set.
[0122] In this embodiment, the first text data includes at least one parsed title identifier and the parsed text corresponding to each parsed title. The parsed title identifier is any identifier representing the parsed title, such as the name, number, etc. of the parsed title. This embodiment does not limit the specific identifier. The parsed title is the title obtained by parsing the title in the preset document. The parsed text is the text under the parsed title.
[0123] It can be understood that since the parsed title is a summary description of the corresponding parsed text, therefore, when retrieving in the first text data set according to the preset dimension information by adopting the keyword retrieval strategy, the retrieval can be first performed based on each parsed title identifier, and when a retrieval result exists, the corresponding parsed text can be further retrieved.
[0124] Based on this, obtain the fields corresponding to the preset dimension information, and extract at least one keyword that can fully reflect the preset dimension information from the fields. At the same time, obtain at least one parsed title identifier pre-marked in each first text data.
[0125] S3012. Retrieve the keywords in at least one parsed title identifier corresponding to each first text data.
[0126] In this embodiment, after obtaining the keywords and at least one parsed title identifier corresponding to each first text data, retrieve in at least one parsed title identifier corresponding to each first text data respectively. The specific execution method is similar to S301 and will not be elaborated here. Thus, the retrieval results corresponding to each first text data are obtained.
[0127] S3013. In response to the retrieval result corresponding to at least one first text data being non-zero, retrieve in the corresponding parsed text based on the parsed title identifier associated with the keyword, and output the matching degree corresponding to the first text data with a non-zero retrieval result.
[0128] In this embodiment, it is judged whether the retrieval result corresponding to each of the above first text data is zero. Then, in response to the retrieval result corresponding to at least one first text data being non-zero, it indicates that there is at least one parsed title identifier in the above first text data that is associated with the keyword, and the corresponding parsed text of these parsed title identifiers may have content related to the preset dimension information.
[0129] Based on this, further retrieve the keyword in the parsed text corresponding to the parsed title identifier associated with the keyword, obtain the final retrieval result corresponding to the first text data with a non-zero retrieval result, and output the corresponding matching degree according to the amount of the above final retrieval result.
[0130] It can be understood that for the first text data with a retrieval result of zero, its corresponding matching degree is also zero.
[0131] S3014. Obtain the first text data with the highest matching degree according to the first quantity to complete the retrieval.
[0132] It can be understood that the higher the matching degree, the higher the relevance between the preset document to which the corresponding first text data belongs and the preset dimension information, that is, the above-mentioned preset document has more content related to the preset dimension information. Therefore, after obtaining the matching degrees corresponding to each first text data, obtain the first text data with the highest matching degree according to the first quantity to complete the retrieval.
[0133] For the report generation method provided in this embodiment, since at least one parsed title identifier and the parsed text corresponding to each parsed title are included in the first text data in advance, by extracting keywords from the preset dimension information and obtaining at least one parsed title identifier corresponding to each first text data, the keyword can be initially retrieved in at least one parsed title identifier corresponding to each first text data, and when it is determined that there is a first text data with a non-zero retrieval result, further retrieve in the corresponding parsed text based on the parsed title identifier associated with the keyword, and then the matching degree corresponding to the first text data with a non-zero retrieval result can be output, so as to obtain the first text data with the highest matching degree according to the first quantity, without retrieving based on all the content of the first text data, thereby improving the retrieval efficiency.
[0134] As an alternative embodiment, based on the above previous embodiment, the second text data includes at least one vector title identifier and vector texts corresponding to each vector title. When retrieving in the second text data set according to the preset dimension information using the vector retrieval strategy, S302 specifically includes the following steps:
[0135] S3021. Vectorize the preset dimension information and obtain at least one vector title identifier corresponding to each second text data in the second text data set.
[0136] It can be understood that since the second text data is converted from the first text data, therefore, based on at least one parsed title identifier included in the corresponding first text data and the parsed texts corresponding to each title, the obtained second text data will include at least one vector title identifier and vector texts corresponding to each vector title.
[0137] Among them, the vector title identifier is any identifier representing the vectorized title, such as the name, number, etc. of the vectorized title. This embodiment does not limit the specific identifier. The vector title is the title after vectorizing the parsed title, and the vector text is the text under the vector title.
[0138] Similarly, since the vector title is a summary description of the corresponding vector text, therefore, when retrieving in the first text data set according to the preset dimension information using the vector retrieval strategy, it is possible to first retrieve based on each vector title identifier, and then further retrieve the corresponding vector text when it is determined that there is a retrieval result.
[0139] Based on this, the vectorization process of the preset dimension information is similar to the specific processing method and steps of a2, which will not be elaborated here. At the same time, at least one pre-marked vector title identifier is obtained in each second text data.
[0140] S3022. Calculate the cosine similarity between at least one vector title identifier corresponding to each second text data and the vectorized preset dimension information respectively.
[0141] In this embodiment, after the vectorized preset dimension information and at least one vector title identifier corresponding to each second text data, calculate the cosine similarity between at least one vector title identifier corresponding to each second text data and the vectorized preset dimension information respectively. The specific calculation method is similar to S302, which will not be elaborated here, and the retrieval results corresponding to each second text data are obtained.
[0142] S3023. In response to there being at least one retrieval result corresponding to the second text data not being zero, retrieve in the corresponding vector text based on the vector title identifier associated with the preset dimension information, and output the similarity corresponding to the second text data with the retrieval result not being zero.
[0143] In this embodiment, it is determined whether the retrieval results corresponding to the above second text data are zero. Then, in response to the fact that there is at least one second text data whose retrieval result is not zero, it indicates that there is at least one vector title identifier in the above second text data that is associated with the vectorized preset dimension information, and the vectorized text corresponding to these vector title identifiers may have content related to the preset dimension information.
[0144] Based on this, further retrieve in the vector text corresponding to the vector title identifier associated with the preset dimension information, that is, calculate the cosine similarity between the above vector text and the vectorized preset dimension information to obtain the final retrieval result corresponding to the second text data with a non-zero retrieval result, that is, the corresponding similarity.
[0145] It can be understood that for the second text data with a zero retrieval result, its corresponding similarity is also zero.
[0146] S3024. Obtain the second text data with the highest similarity according to the second quantity to complete the retrieval.
[0147] It can be understood that the higher the similarity, the higher the relevance between the corresponding preset document of the second text data and the preset dimension information, that is, the above-mentioned preset document has more content related to the preset dimension information. Therefore, after obtaining the similarities corresponding to the second text data, obtain the second text data with the highest similarity according to the second quantity to complete the retrieval.
[0148] For the report generation method provided in this embodiment, since at least one vector title identifier and the vector text corresponding to each vector title are included in the second text data in advance, by vectorizing the preset dimension information and obtaining at least one vector title identifier corresponding to each second text data, the cosine similarity between at least one vector title identifier corresponding to each second text data and the vectorized preset dimension information can be calculated respectively. And when it is determined that there is second text data with a non-zero retrieval result, further retrieve in the corresponding vector text based on the vector title identifier associated with the preset dimension information, and then the similarity corresponding to the second text data with a non-zero retrieval result can be output, so as to obtain the second text data with the highest similarity according to the second quantity, without retrieving based on all the content of the second text data, further improving the retrieval efficiency.
[0149] As an alternative embodiment, based on the above embodiment or the previous embodiment, before retrieving in the first text data set and the second text data set respectively using the corresponding retrieval strategy based on the preset dimension information, the following steps are further included:
[0150] Identify at least one title identifier included in each preset document and the paragraphs belonging to each title identifier, and mark each paragraph with the corresponding title identifier.
[0151] Among them, the title identifier is an identifier representing the title identity, such as the title name, number, etc. This embodiment does not limit the specific identifier. The title is used to summarize the title of the body text in the preset document.
[0152] Among them, the paragraph is a paragraph that constitutes the body text.
[0153] Specifically, at least one title identifier included in each preset document can be identified in advance, and each included paragraph can be identified, then the title identifier to which each paragraph belongs can be determined, and then the corresponding paragraph can be marked with the belonging title identifier.
[0154] It can be understood that after marking the preset document, the corresponding first text data can be obtained by parsing the marked preset document. Then, in the first text data obtained thereby, there will be included at least one parsed title identifier and the parsed text corresponding to each parsed title.
[0155] For the report generation method provided in this embodiment, since both the second and first text data are obtained by converting the preset document, by identifying at least one title identifier included in each preset document and the paragraphs belonging to each title identifier, each paragraph can be marked with the belonging title identifier, so that the corresponding feature title identifier and the corresponding feature text can be included in the subsequent obtained first and second text data, thereby successfully executing the corresponding retrieval process and improving the success rate of the retrieval.
[0156] Figure 4 This is the flow of the report generation method provided in the embodiments of the present application Figure 3 , such as Figure 4 shown. Based on any of the above embodiments, this embodiment further refines the sequential merging process and preset reordering process for a preset number of preset text data. Then, when this embodiment sequentially performs the merging process and preset reordering process on a preset number of preset text data, it specifically includes the following steps:
[0157] S401. Obtain the preset documents to which the first text data of the first quantity and the second text data of the second quantity belong, and merge the same preset documents to obtain the preset documents of the third quantity.
[0158] Among them, the third quantity is the quantity of the preset documents after merging.
[0159] In this embodiment, based on obtaining a preset number of preset text data, among the preset text data, determine the preset documents to which the first text data of the first number belongs, and determine the preset documents to which the second text data of the second number belongs, and obtain the above-mentioned preset documents.
[0160] It can be understood that since both the first text data and the second text data are converted from the preset documents, therefore, among the obtained above-mentioned preset documents, there may be at least two identical preset documents, that is to say, there is a certain preset document with a relatively high matching degree and similarity to the preset dimension information.
[0161] Based on this, merge the identical preset documents to obtain the preset documents of the third number.
[0162] Exemplarily, if the preset number of preset text data is A, B, C, D, E, F, and among them, the first text data of the first number is A, B, C, the second text data of the second number is D, E, F, and the preset documents to which A and D belong are the same, then after merging A and D, the preset documents of the third number obtained can be the preset document to which A&D belongs, the preset document to which B belongs, the preset document to which C belongs, the preset document to which E belongs, and the preset document to which F belongs.
[0163] S402. Use a preset reordering algorithm to perform a reordering process on the preset documents of the third number to obtain the target documents of the target number.
[0164] Among them, the preset reordering algorithm is an algorithm preset according to the priorities of the preset documents for sorting.
[0165] Among them, the target number is the number preset to be obtained from the reordered preset documents.
[0166] Among them, the target document is the preset document with a higher priority.
[0167] In this embodiment, after obtaining the preset documents of the third number, obtain the preset reordering algorithm stored in advance, and use the preset reordering algorithm to determine the priorities corresponding to the above-mentioned preset documents, and reorder the preset documents of the third number according to the priorities to obtain the target documents of the target number.
[0168] The report generation method provided in this embodiment, due to the pre-set reordering algorithm, by obtaining the first number of first text data and the preset documents to which the second number of second text data belong among the preset number of preset text data, the same preset documents can be merged to obtain the third number of preset documents, avoiding duplicate documents. And by using the preset reordering algorithm to reorder the above-mentioned preset documents, the target number of target documents with higher priority can be obtained, thereby determining the documents with stronger relevance to the preset dimension information, making the subsequent analysis based on these documents improve the accuracy of report analysis.
[0169] As an alternative embodiment, based on the above embodiment, when using the preset reordering algorithm to reorder the third number of preset documents to obtain the target number of target documents, S402 specifically includes the following steps:
[0170] S4021. Obtain the preset documents corresponding to the third number with a similarity greater than the preset similarity threshold to obtain the fourth number of preset documents.
[0171] Among them, the fourth number is the number of preset documents with a higher similarity.
[0172] Among them, the preset similarity threshold is pre-set, and it is the minimum similarity value that the preset document needs to reach when meeting the similarity requirement.
[0173] Specifically, first, based on the third number of preset documents, obtain the matching degree corresponding to each of the above-mentioned preset documents according to the matching degree of the corresponding first text data, and obtain the similarity corresponding to each of the above-mentioned preset documents according to the similarity of the corresponding second text data. Then compare the similarity corresponding to each of the above-mentioned preset documents with the preset similarity threshold, and obtain the preset documents with a similarity greater than the preset similarity threshold, and determine these preset documents as the fourth number of preset documents and obtain them.
[0174] S4022. Calculate the length of the longest common substring between each preset document and the problem information at the fourth number. The length of the longest common substring is used to characterize the relevance between the document and the problem information.
[0175] Among them, the length of the longest common substring is the length of the longest identical continuous character sequence found in two strings. In this embodiment, it is specifically used to characterize the relevance between the document and the problem information, that is, the longer the length of the longest common substring, the stronger the relevance between the document and the problem information.
[0176] In this embodiment, based on obtaining the fourth quantity of preset documents, problem information is acquired. The string lengths of each of the above-mentioned preset documents and the string length of the problem information are calculated respectively, and traversal is performed between the two string lengths. Based on the traversal result, the maximum common substring length between the two strings is obtained, and this maximum common substring length is determined as the maximum common substring length between each of the above-mentioned preset documents and the problem information.
[0177] S4023. Obtain the preset documents with the longest maximum common substring length according to the fifth quantity, and calculate the weighted similarity and weighted matching degree corresponding to each preset document under the fifth quantity.
[0178] Among them, the fifth quantity is the number of preset documents with a relatively long maximum common substring length.
[0179] Among them, the weighted similarity is the similarity with corresponding weights set. The weighted matching degree is the matching degree with corresponding weights set. The corresponding weights can be determined and configured by researchers themselves, and this embodiment does not limit this.
[0180] In this embodiment, after obtaining the maximum common substring length between each preset document and the problem information under the fourth quantity, the maximum common substring lengths corresponding to the fourth quantity of preset documents are compared, and based on the comparison result, the fifth quantity of preset documents with a relatively long maximum common substring length is obtained, and further the weighted similarity and weighted matching degree corresponding to each preset document under the fifth quantity are calculated.
[0181] Among them, when obtaining the fifth quantity of preset documents with a relatively long maximum common substring length based on the comparison result, specifically: the fourth quantity of preset documents is sorted based on the comparison result, such as sorting the fourth quantity of preset documents according to the rule of from large to small or from small to large of the maximum common substring length, and the preset documents located in the first fifth quantity or the last fifth quantity are obtained from the sorting result.
[0182] Among them, when calculating the weighted similarity and weighted matching degree corresponding to each preset document under the fifth quantity, specifically: the pre-configured corresponding weights are obtained, the corresponding weights are added on the basis of the similarity corresponding to each preset document, and the corresponding weights are added on the basis of the corresponding matching degree, to obtain the weighted similarity and weighted matching degree corresponding to each preset document under the fifth quantity.
[0183] S4024. Obtain the preset documents with the largest weighted similarity according to the sixth quantity, and obtain the preset documents with the largest weighted matching degree according to the seventh quantity under the sixth quantity.
[0184] Among them, the sixth quantity is the number of documents with a relatively large weighted similarity. The seventh quantity is the number of documents with a relatively large weighted matching degree.
[0185] In this embodiment, based on the weighted similarity and weighted matching degree corresponding to each preset document under the fifth quantity, the weighted similarities of the preset documents corresponding to the fifth quantity are compared, and based on the comparison result, the sixth quantity of preset documents with larger weighted similarities is obtained. Further, the weighted matching degrees corresponding to the sixth quantity of preset documents are compared, and based on the comparison result, the seventh quantity of preset documents with larger weighted matching degrees is obtained.
[0186] Among them, the specific method of obtaining the sixth quantity of preset documents with larger weighted similarities based on the comparison result, and the specific method of obtaining the seventh quantity of preset documents with larger weighted matching degrees based on the comparison result are similar to the method of obtaining the fifth quantity of preset documents with longer maximum common substring lengths based on the comparison result in S4023, and will not be elaborated here.
[0187] S4025. Determine the preset documents corresponding to the seventh quantity as the target documents corresponding to the target quantity.
[0188] In this embodiment, after obtaining the preset documents corresponding to the seventh quantity, these preset documents are determined as the target documents corresponding to the target quantity.
[0189] The report generation method provided in this embodiment can obtain the fourth quantity of preset documents by obtaining the preset documents with similarity greater than the preset similarity threshold under the third quantity, and by further determining the relevance between each preset document and the question information under the fourth quantity, that is, calculating the maximum common substring length, the preset document with the longest maximum common substring length can be obtained according to the fifth quantity. Then, by further calculating the weighted similarity and weighted matching degree corresponding to each preset document under the fifth quantity, the preset document with the largest weighted similarity can be obtained according to the sixth quantity, and the preset document with the largest weighted matching degree can be obtained according to the seventh quantity under the sixth quantity. Thus, the preset documents corresponding to the seventh quantity are determined as the target documents corresponding to the target quantity. In this way, the priority of each preset document is determined layer by layer, improving the accuracy of the determined target documents.
[0190] Figure 5 is the flow of the report generation method provided in the embodiments of the present application Figure 4 , as Figure 5 shown, on the basis of the corresponding embodiment in Figure 2 this embodiment, the quantity of preset dimension information is further limited, and the generation of the target report according to the input information using the preset large language model is further limited. Then, in this embodiment, the quantity of preset dimension information is one. When generating the target report according to the input information using the preset large language model, the following steps are specifically included:
[0191] S501. Obtain problem information and each target document from the input information based on a preset large language model, and analyze each target document according to the problem information to obtain at least one content area associated with the problem information.
[0192] Among them, the content area is a rectangular area covered by some text in the document, such as the rectangular area covered by a certain paragraph, the rectangular area covered by a certain line of text, etc. This embodiment does not limit this.
[0193] In this embodiment, when the number of preset dimension information is one, only output the target report based on this preset dimension information.
[0194] Based on this, obtain the input information based on the preset large language model, and further determine the problem information contained therein and each target document corresponding to the preset dimension information. Then, search for the content related to the problem information in each target document according to the problem information, obtain the area where the above content is located, and obtain at least one content area associated with the problem information.
[0195] S502. Determine the target document to which each content area belongs and the corresponding coordinate range, and generate a traceability component corresponding to each content area based on the identifier of the target document to which it belongs and the corresponding coordinate range.
[0196] It can be understood that when parsing each preset document, the corresponding coordinates of the text contained in each preset document can also be further parsed.
[0197] Among them, the corresponding coordinate range is the corresponding coordinate range covered by the content area.
[0198] Among them, the traceability component is a component used to lock the corresponding coordinate range.
[0199] Among them, the identifier of the target document is an identifier representing the identity of the target document, such as the general title of the target document. This embodiment does not limit this. This identifier can be specifically obtained after obtaining the target number of target documents, and each target document is marked with the corresponding identifier, such as obtaining the corresponding general title and marking each target document with the field where the general title is located.
[0200] Based on this, after obtaining each content area, determine the target document to which each content area belongs, and determine the corresponding coordinate range covered by it in the target document to which it belongs according to the specific position of each content area in the target document to which it belongs. Then, construct a mapping relationship between the target document to which it belongs, the identifier of the target document to which it belongs, and the corresponding coordinate range, and generate a traceability component corresponding to each content area based on the above mapping relationship.
[0201] S503. Generate a target report based on each content area and the corresponding traceability component according to the first report template.
[0202] Among them, the first report template is a preset report template applicable to single preset dimension information, and this template can be customized based on the preset dimension information according to the needs of researchers.
[0203] In this embodiment, after obtaining the traceability components corresponding to each content area, the first report template is obtained, and further the text corresponding to each content area is obtained. The above text is filled into the corresponding position of the first report template, and the corresponding traceability components are embedded into the corresponding positions, thereby generating a target report.
[0204] For the report generation method provided in this embodiment, since a large prediction model is established and trained in advance, based on this large language model, when the preset dimension information is one, by obtaining the problem information in the input information and each target document, each target document can be analyzed based on the problem information to obtain at least one content area associated with the problem information. And by determining the target document to which each content area belongs and the corresponding coordinate range, a traceability component that can locate each content area can be generated based on the identifier of the target document to which it belongs and the corresponding coordinate range. Thus, through the pre-customized first report template, a target report can be generated based on each content area and the corresponding traceability components, without manual writing, further improving the efficiency of report generation under one preset dimension information, and on this basis, adding a traceability component, making the report content well-founded and improving the reliability of the report.
[0205] Figure 6 is the flow of the report generation method provided in the embodiments of the present application Figure 5 , as Figure 6 shown, on the basis of the above embodiment, the number of preset dimension information is further limited, and the generation of the target report according to the input information by using the preset large language model is further limited. Then, in this embodiment, the number of preset dimension information is multiple. When generating the target report according to the input information by using the preset large language model, it specifically includes the following steps:
[0206] S601. Obtain the problem information corresponding to each preset dimension information and each target document in each input information based on the preset large language model.
[0207] In this embodiment, when the number of preset dimension information is multiple, a target report needs to be output based on each preset dimension information.
[0208] Based on this, the input information corresponding to each preset dimension information is obtained based on the preset large language model, and further the problem information included in each input information and each target document under the corresponding preset dimension information are obtained, so as to obtain the problem information corresponding to each preset dimension information and each target document.
[0209] S602. Determine at least one content area corresponding to each preset dimension information according to the corresponding problem information and each target document, and generate a traceability component corresponding to each content area.
[0210] In this embodiment, based on the problem information corresponding to each preset dimension information and each target document, search for the content related to the problem information corresponding to each preset dimension information in the corresponding target documents according to the problem information corresponding to each preset dimension information, obtain the area where the above content is located, and determine the area as at least one content area corresponding to each preset dimension information.
[0211] Furthermore, determine the target document to which each content area belongs and the corresponding coordinate range, and generate a traceability component corresponding to each content area based on the identifier of the target document to which it belongs and the corresponding coordinate range. The specific implementation method is similar to S502 and will not be elaborated here.
[0212] It can be understood that through the above execution process, at least one content area corresponding to each preset dimension information and a traceability component corresponding to each content area can be obtained.
[0213] S603. Generate an initial report corresponding to each preset dimension information based on the corresponding at least one content area and the traceability component corresponding to each content area, and splice each preset dimension information and the corresponding initial report according to the second report template to generate a target report.
[0214] Among them, the initial report is a report corresponding to each preset dimension information under multiple preset dimension information.
[0215] Among them, the second report template is a pre-set report template applicable to multiple preset dimension information, and this template can also be customized based on different preset dimension information according to the needs of researchers.
[0216] In this embodiment, after obtaining at least one content area corresponding to each preset dimension information and the traceability component corresponding to each content area, generate an initial report corresponding to each preset dimension information based on the corresponding at least one content area and the traceability component corresponding to each content area. The specific implementation method is similar to S503 and will not be elaborated here. Then obtain the second report template, and splice each preset dimension information and the corresponding initial report at the positions specified in the second report template to obtain the target report.
[0217] The report generation method provided in this embodiment can, when there are multiple preset dimension information, obtain the problem information in each input information and each target document, and then determine at least one content area corresponding to each preset dimension information according to the corresponding problem information and each target document, and generate a traceability component corresponding to each content area. By generating an initial report corresponding to each preset dimension information based on the corresponding at least one content area and the traceability component corresponding to each content area, the initial reports corresponding to each preset dimension information can be spliced according to a pre-customized second report template, so as to generate a target report under multiple preset dimension information. There is no need for manual writing, which improves the efficiency of report generation under multiple preset dimension information, and on this basis, a traceability component is added, making the report content well-founded and further improving the reliability of the report.
[0218] As an alternative embodiment, on the basis of the above embodiment or the previous embodiment, after generating the target report, the following steps are further included:
[0219] Step c1: Receive a traceability request sent by a user terminal. The traceability request is triggered by the user based on the traceability component corresponding to the content area to be traced.
[0220] Among them, the traceability request is a request for indicating to find the source of the content area, specifically triggered by the user based on the traceability component corresponding to the content area to be traced.
[0221] Among them, the content area to be traced is the content area for which the source is to be found.
[0222] In this embodiment, a researcher can, based on the traceability requirement, open a certain target report on the user terminal and click on the traceability component corresponding to the content area to be traced in the target report, then the user terminal responds to the above click operation trigger and generates a traceability request based on the traceability component.
[0223] Based on this, receive the above traceability request.
[0224] Step c2: Obtain the document identifier to be traced and the corresponding coordinate range corresponding to the content area to be traced according to the traceability request.
[0225] Among them, the document identifier to be traced is the identifier representing the document to which the content area to be traced belongs, such as the general title of the document. The document to be traced is the document to which the content area to be traced belongs.
[0226] It can be understood that the traceability component is generated based on the identifier of the target document to which each content area belongs and the corresponding coordinate range. Therefore, after receiving the request triggered by the traceability component, the document identifier to be traced and the corresponding coordinate range corresponding to the content area to be traced can be obtained based on this request.
[0227] Step c3: Obtain the document to be traced according to the document identifier to be traced, and determine the content area to be traced in the document to be traced according to the corresponding coordinate range.
[0228] In this embodiment, after obtaining the document identifier corresponding to the content area to be traced and the corresponding coordinate range, according to the document identifier to be traced, search for the document with the same identifier in the preset document set, and determine the document as the document to be traced. Obtain the document to be traced, further determine the coverage area of the corresponding coordinate range in the document to be traced, and determine the above coverage area as the content area to be traced, thus completing the tracing.
[0229] For the report generation method provided in this embodiment, since the target report includes a tracing component, by receiving a tracing request triggered by the tracing component corresponding to the content area to be traced from the user terminal, the document identifier corresponding to the content area to be traced and the corresponding coordinate range can be obtained according to the request. And by obtaining the document to be traced according to the document identifier to be traced, the corresponding content area to be traced can be determined in the document according to the corresponding coordinate range, thereby quickly determining the source of the report content and improving the tracing efficiency.
[0230] As an alternative embodiment, this embodiment is based on Figure 2 the corresponding embodiment and further includes the following steps:
[0231] Step d1: Receive a target document update instruction sent by the user terminal. The target document update instruction includes target document update information and the corresponding coordinate range.
[0232] Among them, the target document update instruction is an instruction for instructing to update the target document, and specifically includes target document update information and the corresponding coordinate range. The target document update information is the relevant information required to update the target document.
[0233] In this embodiment, a researcher can open the target document on the user terminal and update the target document based on requirements, such as adding, deleting, or adjusting some text in the target document. Then, in response to the above update operation being triggered, the user terminal obtains the identifier of the target document and the updated text, and further obtains the corresponding coordinate range of the above text. Generate target document update information based on the identifier of the target document and the updated text, and generate a target document update instruction based on the target document update information and the corresponding coordinate range.
[0234] Based on this, receive the above target document update instruction.
[0235] Step d2: Obtain the target report according to the target document update instruction, and obtain the corresponding content area in the target report according to the corresponding coordinate range.
[0236] In this embodiment, based on the target document update instruction, the target document update information and the corresponding coordinate range are obtained, and further, the identifier of the target document and the updated text are obtained from the target document update information. Based on the identifier of the target document, the target report applied thereto is obtained, and the content area indicated by the corresponding coordinate range is obtained from the above target report.
[0237] Step d3: Update the corresponding content area with the target document update information to update the target report.
[0238] In this embodiment, the updated text in the target document update information is used to update the corresponding content area, thereby updating the target report. For example, the corresponding content area is deleted with the deleted text, or the corresponding content area is added with the added text, or the corresponding content area is adjusted with the adjusted text.
[0239] For the report generation method provided in this embodiment, since an association is built between the target document and the applied target report based on the traceability component, by receiving the target document update instruction including the target document update information and the corresponding coordinate range sent by the user terminal, the target report can be obtained according to this instruction, and the corresponding content area can be obtained from the target report according to the corresponding coordinate range, so as to update the corresponding content area with the target document update information, realizing real-time update of the target report, improving the update efficiency, and improving the user experience.
[0240] Figure 7 is the complete flowchart of the report generation method provided in the embodiment of the present application, as Figure 7 shown, which schematically illustrates the complete execution process of this solution.
[0241] S701. Obtain a preset document set.
[0242] S702. Parse each preset document in the preset document set to obtain a first text data set.
[0243] S703. Perform segmentation processing and vectorization processing on each first text data in the first text data set in sequence to obtain a second text data set.
[0244] S704. Store the first text data set in a preset database, and store the second text data set in a vector database.
[0245] S705. Receive a target report generation request sent by a user terminal.
[0246] Among them, the target report generation request includes preset dimension information and problem information, and the number of preset dimension information is at least one.
[0247] S706. Obtain the first text data set and the second text data set.
[0248] Specifically, obtain the first text data set based on a preset database, and obtain the second text data set based on a vector database.
[0249] S707. Identify at least one title identifier included in each preset document and the paragraphs belonging to each title identifier, and mark each paragraph with the corresponding title identifier.
[0250] S708. Retrieve in the first text data set and the second text data set respectively based on the preset dimension information using corresponding retrieval strategies to obtain a preset number of preset text data.
[0251] Among them, the corresponding retrieval strategies include a keyword retrieval strategy and a vector retrieval strategy.
[0252] Specifically, use the keyword retrieval strategy to retrieve in the first text data set according to the preset dimension information to obtain a first number of first text data. Use the vector retrieval strategy to retrieve in the second text data set according to the preset dimension information to obtain a second number of second text data. Determine the first number of first text data and the second number of second text data as the preset number of preset text data to obtain the preset number of preset text data.
[0253] Among them, the first text data includes at least one parsed title identifier and the parsed text corresponding to each parsed title. When using the keyword retrieval strategy to retrieve in the first text data set according to the preset dimension information: Extract keywords from the preset dimension information, and obtain at least one parsed title identifier corresponding to each first text data in the first text data set. Retrieve the keywords respectively in at least one parsed title identifier corresponding to each first text data. In response to there being at least one first text data with a non-zero retrieval result, retrieve in the corresponding parsed text based on the parsed title identifier associated with the keyword, and output the matching degree corresponding to the first text data with a non-zero retrieval result. Obtain the first text data with the highest matching degree according to the first number to complete the retrieval.
[0254] Among them, the second text data includes at least one vector title identifier and vector texts corresponding to each vector title. When retrieving in the second text data set according to the preset dimension information using the vector retrieval strategy: vectorize the preset dimension information, and obtain at least one vector title identifier corresponding to each second text data in the second text data set. Calculate the cosine similarity between at least one vector title identifier corresponding to each second text data and the vectorized preset dimension information respectively. In response to the retrieval result corresponding to at least one second text data being non-zero, retrieve in the corresponding vector text based on the vector title identifier associated with the preset dimension information, and output the similarity corresponding to the second text data with a non-zero retrieval result. Obtain the second text data with the highest similarity according to the second quantity to complete the retrieval.
[0255] S709. Perform a merging process and a preset reordering process on the preset text data based on a preset quantity in sequence to obtain a target quantity of target documents.
[0256] Specifically, perform a merging process and a preset reordering process on the preset text data based on a preset quantity in sequence, obtain the first text data of the first quantity and the preset documents to which the second text data of the second quantity belongs, and merge the same preset documents to obtain the preset documents of the third quantity. Use a preset reordering algorithm to perform a reordering process on the preset documents of the third quantity to obtain a target quantity of target documents.
[0257] Among them, when using a preset reordering algorithm to perform a reordering process on the preset documents of the third quantity to obtain a target quantity of target documents: obtain the preset documents with a similarity greater than the preset similarity threshold corresponding to the third quantity to obtain the preset documents of the fourth quantity. Calculate the maximum common substring length between each preset document and the question information in the fourth quantity. The maximum common substring length is used to characterize the relevance between the document and the question information. Obtain the preset document with the longest maximum common substring length according to the fifth quantity, and calculate the weighted similarity and weighted matching degree corresponding to each preset document in the fifth quantity. Obtain the preset document with the largest weighted similarity according to the sixth quantity, and obtain the preset document with the largest weighted matching degree according to the seventh quantity in the sixth quantity. Determine the preset document of the seventh quantity as the target document of the target quantity.
[0258] S710. Concatenate the target documents of the target quantity and the question information into input information.
[0259] S711. Use a preset large language model to generate a target report according to the input information.
[0260] Among them, the number of preset dimension information is one.
[0261] Specifically, based on a preset large language model, problem information and each target document are obtained from the input information, and each target document is analyzed according to the problem information to obtain at least one content area associated with the problem information. Determine the target document to which each content area belongs and the corresponding coordinate range, and generate a traceability component corresponding to each content area based on the identifier of the target document to which it belongs and the corresponding coordinate range. Generate a target report based on each content area and the corresponding traceability component according to the first report template.
[0262] Among them, the number of preset dimension information is multiple.
[0263] Specifically, based on a preset large language model, problem information corresponding to each preset dimension information and each target document are obtained from each input information. At least one content area corresponding to each preset dimension information is determined according to the corresponding problem information and each target document, and a traceability component corresponding to each content area is generated. An initial report corresponding to each preset dimension information is generated based on the corresponding at least one content area and the traceability component corresponding to each content area, and each preset dimension information and the corresponding initial report are spliced according to the second report template to generate a target report.
[0264] S712. Receive a traceability request sent by the user terminal.
[0265] Among them, the traceability request is triggered by the user based on the traceability component corresponding to the content area to be traced.
[0266] S713. Obtain the identifier of the document to be traced and the corresponding coordinate range corresponding to the content area to be traced according to the traceability request.
[0267] S714. Obtain the document to be traced according to the identifier of the document to be traced.
[0268] S715. Determine the content area to be traced in the document to be traced according to the corresponding coordinate range.
[0269] S716. Receive a target document update instruction sent by the user terminal.
[0270] Among them, the target document update instruction includes target document update information and the corresponding coordinate range.
[0271] S717. Obtain the target report according to the target document update instruction.
[0272] S718. Obtain the corresponding content area in the target report according to the corresponding coordinate range.
[0273] S719. Update the corresponding content area with the target document update information to update the target report.
[0274] Figure 8The structural schematic diagram of the report generation device provided by the embodiment of the present application is as follows. Figure 5 As shown, the execution subject of the present application is the report generation device 80, and the report generation device 80 is specifically located in an electronic device. Then, the report generation device provided in this embodiment includes: a receiving module 81, an obtaining module 82, a retrieving module 83, a processing module 84, a splicing module 85, and a generating module 86.
[0275] Among them, the receiving module 81 is used to receive a target report generation request sent by a user terminal; the target report generation request includes preset dimension information and problem information, and the number of preset dimension information is at least one; the obtaining module 82 is used to obtain a first text data set and a second text data set; the retrieving module 83 is used to retrieve based on the preset dimension information by using corresponding retrieval strategies in the first text data set and the second text data set respectively to obtain a preset number of preset text data; the processing module 84 is used to perform merging processing and preset reordering processing on the preset number of preset text data in sequence to obtain a target number of target documents; the splicing module 85 is used to splice the target number of target documents and the problem information into input information; the generating module 86 is used to generate a target report by using a preset large language model according to the input information.
[0276] Optionally, when the obtaining module 82 obtains the first text data set and the second text data set, it is specifically used for:
[0277] Obtaining the first text data set based on a preset database and obtaining the second text data set based on a vector database;
[0278] Correspondingly, the report generation device provided in this embodiment further includes: a storage module.
[0279] Among them, the obtaining module 82 is further used to obtain a preset document set before obtaining the first text data set and the second text data set; the processing module 84 is further used to perform parsing processing on each preset document in the preset document set to obtain the first text data set; perform segmentation processing and vectorization processing on each first text data in the first text data set in sequence to obtain the second text data set; the storage module is used to store the first text data set in the preset database and store the second text data set in the vector database.
[0280] Optionally, the corresponding retrieval strategies include a keyword retrieval strategy and a vector retrieval strategy;
[0281] Correspondingly, when the retrieving module 83 retrieves based on the preset dimension information by using corresponding retrieval strategies in the first text data set and the second text data set respectively to obtain a preset number of preset text data, it is specifically used for:
[0282] Adopt a keyword retrieval strategy to retrieve in the first text data set according to preset dimension information to obtain a first quantity of first text data; adopt a vector retrieval strategy to retrieve in the second text data set according to preset dimension information to obtain a second quantity of second text data; determine the first quantity of first text data and the second quantity of second text data as a preset quantity of preset text data to obtain a preset quantity of preset text data.
[0283] Optionally, the first text data includes at least one parsed title identifier and the parsed text corresponding to each parsed title;
[0284] Correspondingly, when the retrieval module 83 retrieves in the first text data set according to the preset dimension information by adopting the keyword retrieval strategy, it is specifically used for:
[0285] Extract keywords from the preset dimension information, and obtain at least one parsed title identifier corresponding to each first text data in the first text data set; retrieve the keywords in at least one parsed title identifier corresponding to each first text data respectively; in response to the retrieval result corresponding to at least one first text data being non-zero, retrieve in the corresponding parsed text based on the parsed title identifier associated with the keyword, and output the matching degree corresponding to the first text data with a non-zero retrieval result; obtain the first text data with the highest matching degree according to the first quantity to complete the retrieval.
[0286] Optionally, the second text data includes at least one vector title identifier and the vector text corresponding to each vector title;
[0287] Correspondingly, when the retrieval module 83 retrieves in the second text data set according to the preset dimension information by adopting the vector retrieval strategy, it is specifically used for:
[0288] Perform vectorization processing on the preset dimension information, and obtain at least one vector title identifier corresponding to each second text data in the second text data set; calculate the cosine similarity between at least one vector title identifier corresponding to each second text data and the vectorized preset dimension information respectively; in response to the retrieval result corresponding to at least one second text data being non-zero, retrieve in the corresponding vector text based on the vector title identifier associated with the preset dimension information, and output the similarity corresponding to the second text data with a non-zero retrieval result; obtain the second text data with the highest similarity according to the second quantity to complete the retrieval.
[0289] The report generation device provided in this embodiment further includes: an identification module, a marking module.
[0290] Among them, an identification module is used to identify at least one title identifier included in each preset document and paragraphs belonging to each title identifier, and a marking module is used to mark each paragraph with the belonging title identifier.
[0291] Optionally, when the processing module 84 performs merging processing and preset reordering processing on a preset number of preset text data in sequence, it is specifically used for:
[0292] Obtain the preset documents to which the first quantity of first text data and the second quantity of second text data belong, and merge the same preset documents to obtain a third quantity of preset documents; perform reordering processing on the third quantity of preset documents using a preset reordering algorithm to obtain a target quantity of target documents.
[0293] Optionally, when the processing module 84 performs reordering processing on the third quantity of preset documents using a preset reordering algorithm to obtain a target quantity of target documents, it is specifically used for:
[0294] Obtain the preset documents with a similarity greater than a preset similarity threshold under the third quantity to obtain a fourth quantity of preset documents; calculate the maximum common substring length between each preset document and the question information under the fourth quantity; the maximum common substring length is used to characterize the relevance between the document and the question information; obtain the preset documents with the longest maximum common substring length according to the fifth quantity, and calculate the weighted similarity and weighted matching degree corresponding to each preset document under the fifth quantity; obtain the preset document with the largest weighted similarity according to the sixth quantity, and obtain the preset document with the largest weighted matching degree according to the seventh quantity under the sixth quantity; determine the preset document of the seventh quantity as the target document of the target quantity.
[0295] Optionally, the number of preset dimension information is one;
[0296] Correspondingly, when the generation module 86 generates a target report according to the input information using a preset large language model, it is specifically used for:
[0297] Obtain question information and each target document from the input information based on the preset large language model, and analyze each target document according to the question information to obtain at least one content area associated with the question information; determine the target document to which each content area belongs and the corresponding coordinate range, and generate a traceability component corresponding to each content area based on the identifier of the belonging target document and the corresponding coordinate range; generate a target report based on each content area and the corresponding traceability component according to the first report template.
[0298] Optionally, the number of preset dimension information is multiple;
[0299] Correspondingly, when the generation module 86 generates a target report according to the input information using a preset large language model, it is specifically used for:
[0300] Obtain the question information corresponding to each preset dimension information and each target document from each input information based on a preset large language model; determine at least one content area corresponding to each preset dimension information according to the corresponding question information and each target document, and generate a traceability component corresponding to each content area; generate an initial report corresponding to each preset dimension information based on the corresponding at least one content area and the traceability component corresponding to each content area, and splice each preset dimension information and the corresponding initial report according to the second report template to generate a target report.
[0301] The report generation device provided in this embodiment further includes: a determination module.
[0302] Among them, the receiving module 81 is further configured to, after the generating module 86 generates the target report, receive a traceability request sent by the user terminal; the traceability request is triggered by the user based on the traceability component corresponding to the content area to be traced; the obtaining module 82 is further configured to obtain the document identifier to be traced and the corresponding coordinate range corresponding to the content area to be traced according to the traceability request; obtain the document to be traced according to the document identifier to be traced, and the determination module is configured to determine the content area to be traced in the document to be traced according to the corresponding coordinate range.
[0303] The report generation device provided in this embodiment further includes: an update module.
[0304] Among them, the receiving module 81 is further configured to receive a target document update instruction sent by the user terminal; the target document update instruction includes target document update information and the corresponding coordinate range; the obtaining module 82 is further configured to obtain the target report according to the target document update instruction, and obtain the corresponding content area in the target report according to the corresponding coordinate range; the update module is configured to update the corresponding content area with the target document update information to update the target report.
[0305] The report generation device provided in the embodiment of the present application can be used to execute the technical solution of the report generation method in the above embodiment, and its implementation principle and technical effect are similar, and will not be described in detail here.
[0306] It should be noted that it should be understood that the division of each module of the above device is only a division of logical functions. In actual implementation, it can be fully or partially integrated into a physical entity, or physically separated. And these modules can all be implemented in the form of software called by a processing element; they can also all be implemented in the form of hardware; or some modules can be implemented in the form of software called by a processing element, and some modules can be implemented in the form of hardware. For example, the receiving module 81 can be a separately established processing element, or can be integrated in a certain chip of the above device. In addition, it can also be stored in the memory of the above device in the form of program code, and called and executed by a certain processing element of the above device to perform the functions of the above receiving module 81. The implementation of other modules is similar. In addition, all or part of these modules can be integrated together or can be independently implemented. Here, the processing element can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above modules can be completed by the integrated logic circuit of the hardware in the processor element or the instructions in the form of software.
[0307] Figure 9 Schematic structural diagram of the electronic device provided by the embodiment of the present application. As Figure 9 shown, the electronic device 90 may include: a processor 91, a memory 92, and a transceiver 93.
[0308] The processor 91 executes the computer execution instructions stored in the memory, so that the processor 91 executes the solutions in the above embodiments. The processor 91 can be a general-purpose processor, including a central processing unit CPU, a network processor (NP), etc.; it can also be a digital signal processor DSP, an application-specific integrated circuit ASIC, a field programmable gate array FPGA, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0309] The memory 92 is connected to the processor 91 through the system bus and completes the communication therebetween. The memory 92 is used to store computer program instructions. The transceiver 93 is used to transmit and receive data.
[0310] The system bus can be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The system bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience in representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus. The transceiver is used to implement the communication between the database access device and other computers (such as clients, read-write libraries, and read-only libraries). The memory may include a Random Access Memory (RAM), and may also include a non-volatile memory.
[0311] An embodiment of the present application also provides a chip for running instructions, and the chip is used to execute the technical solution of the report generation method in the above embodiment.
[0312] An embodiment of the present application also provides a computer-readable storage medium, in which computer instructions are stored. When the computer instructions run on a computer, the computer is enabled to execute the technical solution of the report generation method in the above embodiment.
[0313] An embodiment of the present application also provides a computer program product, which includes a computer program stored in a computer-readable storage medium. At least one processor can read the computer program from the computer-readable storage medium, and when at least one processor executes the computer program, the technical solution of the report generation method in the above embodiment can be implemented.
[0314] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of devices or modules can be in an electrical, mechanical or other form.
[0315] The modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to implement the solution of this embodiment.
[0316] In addition, in each embodiment of the present application, each functional module can be integrated in a processing unit, or each module can exist physically alone, or two or more modules can be integrated in one unit. The unit formed by the above modules can be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.
[0317] The above integrated module implemented in the form of a software functional module can be stored in a computer-readable storage medium. The above software functional module stored in a storage medium includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute some steps of the methods in the various embodiments of the present application.
[0318] It should be understood that the above processor can be a Central Processing Unit (CPU for short), or other general-purpose processors, Digital Signal Processors (DSP for short), Application Specific Integrated Circuits (ASIC for short), etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the invention can be directly implemented by the execution of the hardware processor, or can be implemented by the combination of hardware and software modules in the processor.
[0319] The memory may include high-speed RAM memory, and may also include non-volatile storage NVM, such as at least one disk memory, and can also be a USB flash drive, a mobile hard disk, a read-only memory, a magnetic disk, or an optical disc, etc.
[0320] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, the buses in the drawings of the present application are not limited to only one bus or one type of bus.
[0321] The above storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. The storage medium can be any available medium accessible by a general or special purpose computer.
[0322] An exemplary storage medium is coupled to a processor such that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an application specific integrated circuit (ASIC). Of course, the processor and the storage medium can also exist as discrete components in an electronic control unit or a master control device.
[0323] Those of ordinary skill in the art will understand that all or part of the steps for implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disk or optical disk that can store program codes.
[0324] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include common general knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present application are pointed out by the following claims.
[0325] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. A report generation method, characterized in that: The method comprises: receiving a target report generation request sent by a user terminal; the target report generation request includes preset dimension information and question information, and the number of the preset dimension information is at least one; Acquire a first text data set based on a preset database, and acquire a second text data set based on a vector database; and search the first text data set and the second text data set respectively using a corresponding search strategy based on the preset dimension information to obtain a preset number of preset text data; Determine a matching degree between each first text data and preset dimension information, and obtain a first number of first text data with a higher matching degree from each first text data; Calculating the cosine similarity between each second text data and the vectorized text corresponding to the preset dimension information, and obtaining a second number of second text data with higher similarity from each second text data; Acquire preset documents to which the first quantity of first text data and the second quantity of second text data belong, and merge the same preset documents to obtain a third quantity of preset documents; Reordering the third number of preset documents using a preset reordering algorithm to obtain a fourth number of preset documents; Calculating the maximum common substring length between each preset document and the question information under the fourth quantity; the maximum common substring length is used to characterize the relevance between the document and the question information; According to the fifth quantity, the preset documents with the longest maximum common substring length are obtained, and the weighted similarity and weighted matching degree corresponding to each preset document under the fifth quantity are calculated; Obtaining a preset document with the largest weighted similarity according to the sixth quantity, and obtaining a preset document with the largest weighted matching degree according to the seventh quantity under the sixth quantity; determining a seventh number of preset documents as a target number of target documents; The target number of target documents and the problem information are spliced into input information, and a preset large language model is used to determine at least one content area related to the problem information in the target document and its corresponding traceability component according to the input information to generate a target report.
2. The method according to claim 1, characterized in that: Before acquiring the first text data set and the second text data set, the method further includes: Get the preset document collection; Parsing each preset document in the preset document set to obtain a first text data set; Sequentially performing segmentation and vectorization processing on each first text data set in the first text data set to obtain a second text data set; The first text data set is stored in a preset database, and the second text data set is stored in a vector database.
3. The method according to claim 2, characterized in that The corresponding search strategy includes a keyword search strategy and a vector search strategy; The step of using a corresponding search strategy based on the preset dimension information to search in the first text data set and the second text data set respectively to obtain a preset amount of preset text data includes: Using a keyword search strategy to search the first text data set according to the preset dimension information to obtain a first quantity of first text data; Using a vector search strategy to search in the second text data set according to the preset dimension information to obtain a second amount of second text data; The first quantity of first text data and the second quantity of second text data are determined as a preset quantity of preset text data to obtain a preset quantity of preset text data.
4. The method according to claim 3, characterized in that: The first text data includes at least one parsed title identifier and parsed text corresponding to each parsed title; The step of using a keyword search strategy to search the first text data set according to the preset dimension information includes: Extracting keywords from the preset dimension information, and obtaining at least one parsed title identifier corresponding to each of the first text data in the first text data set; Retrieving keywords from at least one parsed title identifier corresponding to each of the first text data; In response to the existence of a non-zero search result corresponding to at least one first text data, searching in the corresponding parsed text based on the parsed title identifier of the associated keyword, and outputting the matching degree corresponding to the first text data whose search result is non-zero; The first text data with the highest matching degree is obtained according to the first quantity to complete the search.
5. The method according to claim 3, characterized in that: The second text data includes at least one vector title identifier and vector text corresponding to each vector title; The adopting a vector search strategy to search in the second text data set according to the preset dimension information includes: Performing vectorization processing on the preset dimension information, and obtaining at least one vector title identifier corresponding to each of the second text data in the second text data set; respectively calculating the cosine similarity between at least one vector title identifier corresponding to each of the second text data and the vectorized preset dimension information; In response to the existence of a non-zero search result corresponding to at least one second text data, searching in the corresponding vector text based on the vector title identifier associated with the preset dimension information, and outputting the similarity corresponding to the second text data whose search result is non-zero; The second text data with the highest similarity is obtained according to the second number to complete the retrieval.
6. The method according to claim 4 or 5, characterized in that: Before searching the first text data set and the second text data set respectively by using a corresponding search strategy based on the preset dimension information, the method further includes: At least one title tag included in each preset document and paragraphs belonging to each title tag are identified, and each paragraph is marked with the corresponding title tag.
7. The method according to claim 1, characterized in that The number of the preset dimension information is one; The using a preset large language model to generate a target report according to the input information includes: Acquire question information and each target document from the input information based on the preset large language model, and analyze each target document according to the question information to obtain at least one content area associated with the question information; Determine the target document and the corresponding coordinate range of each content area, and generate the traceability component corresponding to each content area based on the identifier of the target document and the corresponding coordinate range; Generate a target report based on each content area and the corresponding traceability component according to the first report template.
8. The method according to claim 7, characterized in that The number of the preset dimension information is multiple; The using a preset large language model to generate a target report according to the input information includes: Based on the preset large language model, obtaining question information corresponding to each preset dimension information and each target document in each input information; Determine at least one content area corresponding to each preset dimension information according to the corresponding question information and each target document, and generate a traceability component corresponding to each content area; Based on at least one corresponding content area and the traceability component corresponding to each content area, an initial report corresponding to each preset dimensional information is generated, and each preset dimensional information and the corresponding initial report are spliced according to a second report template to generate a target report.
9. The method according to claim 7 or 8, characterized in that: After generating the target report, the following steps are also included: Receiving a traceability request sent by a user terminal; the traceability request is triggered by the user based on a traceability component corresponding to the content area to be traced; Obtaining the to-be-traced document identifier and the corresponding coordinate range corresponding to the to-be-traced content area according to the traceability request; The document to be traced is obtained according to the document identifier to be traced, and the content area to be traced is determined in the document to be traced according to the corresponding coordinate range.
10. The method according to claim 1, characterized in that Also includes: Receiving a target document update instruction sent by a user terminal; the target document update instruction includes target document update information and a corresponding coordinate range; Acquire a target report according to the target document update instruction, and acquire a corresponding content area in the target report according to a corresponding coordinate range; The target document update information is used to update the corresponding content area to update the target report.
11. An electronic device, comprising: A processor, a memory and a transceiver communicatively connected to the processor; The memory stores computer-executable instructions; the transceiver is used to send and receive data; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 10.
12. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 10 when executed by a processor.
13. A computer program product, characterized in that The invention comprises a computer program, which implements the method according to any one of claims 1 to 10 when being executed by a processor.
Citation Information
Patent Citations
Method and system for generating intelligent analysis report based on large language model and process mining data
CN116955597A
Multi-document intelligent question and answer method and system based on large language model
CN118394897A