Document benchmarking method and device, equipment, storage medium and program product
Through the input of identification information and target paragraph information of large models, the automated output of document benchmarking is achieved, which solves the problems of low and inaccurate document benchmarking efficiency in the existing technology, improves the efficiency and accuracy of document benchmarking, and provides new document management and benchmarking solutions for different industries.
Patent Information
- Application Number
- CN202510571354.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-05-06
AI Technical Summary
In the prior art, the document benchmarking efficiency is low and not accurate enough to effectively meet the efficient document management and benchmarking needs in the vehicle industry and other fields.
By obtaining relevant information of multiple objects to be compared, using identification information and target paragraph information for large-scale input, we can automatically output target benchmarking results, including vectorization processing and paragraph summary matching to filter unrelated data and optimize server resource utilization.
It greatly improves the efficiency and accuracy of document benchmarking, improves the reliability and professionalism of document benchmarking, and provides new document management and benchmarking solutions for different industries.
Smart Images

Figure CN120087353A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of information processing, and particularly relates to a method, device, equipment, storage medium and program product for document alignment. Background Art
[0002] Aligning relevant documents of different objects can help identify the differences and advantages between different objects, optimize market competition strategies, assist in decision-making, and enhance transparency and trust, etc. Taking the alignment of vehicle model documents as an example, vehicle model document alignment is to conduct comparative analysis on the technical documents (such as design specifications, configuration parameters, function descriptions, etc.) of different vehicle models during the vehicle design and development process to identify differences, optimize designs, or meet specific requirements. This process can help understand the similarities and differences as well as the product capabilities between different vehicle models, and is crucial in aspects such as the research and development, procurement, and formulation of market strategies in the vehicle industry.
[0003] Currently, manual comparative analysis is usually adopted for document alignment, but the alignment efficiency is relatively low. Therefore, there is an urgent need for an effective solution for document alignment. Summary of the Invention
[0004] The purpose of the present invention is to provide a method, device, equipment, storage medium and program product for document alignment to more efficiently and accurately align documents.
[0005] To achieve the above purpose, the technical solution adopted by the present invention is as follows:
[0006] A method for document alignment includes: obtaining relevant information corresponding to multiple objects to be aligned, where the relevant information includes identification information corresponding to each of the multiple objects to be aligned; for each object to be aligned among the multiple objects to be aligned, if the relevant information includes a target alignment dimension, then according to the identification information and the target alignment dimension, obtaining target paragraph information corresponding to the object to be aligned; if the relevant information does not include the target alignment dimension, then performing vectorization processing on the relevant information to obtain a target vector, and according to the target vector and multiple paragraph summary vectors, obtaining target paragraph information corresponding to the object to be aligned; the target paragraph information includes a target paragraph summary, context paragraph information of the original paragraph corresponding to the target paragraph summary, and the alignment dimension to which the target paragraph summary belongs, and the multiple paragraph summary vectors, alignment dimension, paragraph summary, original paragraph, and context paragraph information are obtained by preprocessing the relevant documents corresponding to different objects respectively; inputting the identification information and the target paragraph information into a large model to obtain target alignment results corresponding to the multiple objects to be aligned output by the large model.
[0007] According to the above technical means, for each of the multiple objects to be benchmarked, if the relevant information corresponding to the multiple objects to be benchmarked includes the target benchmarking dimension, then according to the identification information and the target benchmarking dimension, obtain the target paragraph information corresponding to the object to be benchmarked; if the relevant information does not include the target benchmarking dimension, then perform vectorization processing on the relevant information to obtain a target vector, and according to the target vector and multiple paragraph summary vectors, obtain the target paragraph information corresponding to the object to be benchmarked; the above method of obtaining the target paragraph information can effectively filter out irrelevant data items, thereby reducing the context length required by the large model and optimizing the utilization rate of the video memory resources of the server; thus, input the identification information and the target paragraph information into the large model to obtain the target benchmarking results corresponding to the multiple objects to be benchmarked output by the large model, realizing the automatic output of the target benchmarking results corresponding to the multiple objects to be benchmarked through the large model, which can greatly improve the efficiency of document benchmarking and can more accurately obtain the target benchmarking results corresponding to the multiple objects to be benchmarked, enhancing the reliability and professionalism of document benchmarking.
[0008] Further, input the identification information and the target paragraph information into the large model to obtain the target benchmarking results corresponding to the multiple objects to be benchmarked output by the large model, including: determining the target prompt word information corresponding to the large model according to the identification information, the object evaluation dimension score information, and the target paragraph information, where the object evaluation dimension score information is obtained by the large model based on the paragraph summary to evaluate multiple benchmarking dimensions; input the target prompt word information into the large model to obtain the target benchmarking results corresponding to the multiple objects to be benchmarked output by the large model.
[0009] Further, the benchmarking dimension includes multiple first-level dimensions and second-level dimensions included in each first-level dimension, and the object evaluation dimension score information is obtained through the following method: for each second-level dimension included in the first-level dimension, extract the target data related to the scoring index from the paragraph summary through the large model according to the scoring index corresponding to the second-level dimension; score the second-level dimension according to the target data and the scoring standard corresponding to the second-level dimension to obtain the score corresponding to the second-level dimension; add up the scores corresponding to the second-level dimensions through the large model to obtain the score corresponding to the first-level dimension, and the object evaluation dimension score information includes the first-level dimension and the score corresponding to the first-level dimension.
[0010] Further, obtaining the relevant information corresponding to multiple objects to be benchmarked includes: in response to the first selection operation for the target benchmarking dimension and the second selection operation for multiple objects to be benchmarked, or in response to the second selection operation for multiple objects to be benchmarked, obtain the relevant information corresponding to multiple objects to be benchmarked; or, obtain the user question information, identify the intention corresponding to the user question information, and if the intention is to perform document benchmarking, then obtain the relevant information corresponding to multiple objects to be benchmarked according to the user question information.
[0011] Further, according to the identification information and the target benchmarking dimension, obtain the target paragraph information corresponding to the object to be benchmarked, including: if the target benchmarking dimension is a first-level dimension, obtain the target paragraph information corresponding to the object to be benchmarked according to the target benchmarking dimension, the second-level dimensions included in the target benchmarking dimension, and the identification information; if the target benchmarking dimension is a second-level dimension, obtain the target paragraph information corresponding to the object to be benchmarked according to the target benchmarking dimension, the first-level dimension to which the target benchmarking dimension belongs, and the identification information.
[0012] Further, according to the target vector and multiple paragraph summary vectors, obtain the target paragraph information corresponding to the object to be benchmarked, including: matching the target vector with each paragraph summary vector among the multiple paragraph summary vectors to obtain the similarity between the target vector and each paragraph summary vector; determining the top pre-set number of target paragraph summary vectors with relatively high similarity among the multiple paragraph summary vectors whose similarity is less than the similarity threshold; and obtaining the target paragraph information corresponding to the object to be benchmarked according to the target paragraph summary vectors.
[0013] Further, preprocess the relevant documents corresponding to different objects respectively, including: splitting the document content of the relevant documents into paragraphs to obtain the original paragraphs, the paragraph sequence numbers of the original paragraphs, and the context paragraph information of the original paragraphs; obtaining the paragraph summary and the benchmarking dimension to which the paragraph summary belongs from the original paragraphs through a large model; performing vectorization processing on multiple paragraph summaries through a document vectorization model to obtain multiple paragraph summary vectors, where the document vectorization model is obtained through unsupervised training based on the paragraph summaries; storing the identification information, benchmarking dimension, original paragraphs, paragraph sequence numbers of the original paragraphs, context paragraph information of the original paragraphs, paragraph summaries, and paragraph summary vectors corresponding to different objects respectively.
[0014] Further, the document content also includes images. Splitting the document content of the relevant documents into paragraphs to obtain the original paragraphs includes: using the optical character recognition (OCR) method to extract the text information in the images; and merging the text information with the paragraphs corresponding to the text information to obtain the original paragraphs.
[0015] Further, perform vectorization processing on the relevant information to obtain a target vector, including: performing vectorization processing on the relevant information through a document vectorization model to obtain the target vector.
[0016] A document benchmarking device, including:
[0017] An acquisition module, configured to acquire the relevant information corresponding to multiple objects to be benchmarked, where the relevant information includes the identification information corresponding to multiple objects to be benchmarked respectively;
[0018] A processing module, for each object to be benchmarked among multiple objects to be benchmarked, if the relevant information contains the target benchmarking dimension, then according to the identification information and the target benchmarking dimension, obtain the target paragraph information corresponding to the object to be benchmarked; if the relevant information does not contain the target benchmarking dimension, then perform vectorization processing on the relevant information to obtain a target vector, and according to the target vector and multiple paragraph summary vectors, obtain the target paragraph information corresponding to the object to be benchmarked; the target paragraph information includes the target paragraph summary, the context paragraph information of the original paragraph corresponding to the target paragraph summary, and the benchmarking dimension to which the target paragraph summary belongs. The multiple paragraph summary vectors, benchmarking dimensions, paragraph summaries, original paragraphs, and context paragraph information are obtained by preprocessing the relevant documents corresponding to different objects respectively.
[0019] An output module, for inputting the identification information and the target paragraph information into a large model to obtain the target benchmarking results corresponding to multiple objects to be benchmarked output by the large model.
[0020] Further, the output module is specifically used for: determining the target prompt word information corresponding to the large model according to the identification information, the object evaluation dimension score information, and the target paragraph information, where the object evaluation dimension score information is obtained by the large model evaluating multiple benchmarking dimensions based on the paragraph summary; inputting the target prompt word information into the large model to obtain the target benchmarking results corresponding to multiple objects to be benchmarked output by the large model.
[0021] Further, the benchmarking dimension includes multiple first-level dimensions and second-level dimensions included in each first-level dimension. The acquisition module is also used to obtain the object evaluation dimension score information in the following way: for each second-level dimension included in the first-level dimension, extract the target data related to the scoring index from the paragraph summary through the large model according to the scoring index corresponding to the second-level dimension; score the second-level dimension according to the target data and the scoring standard corresponding to the second-level dimension to obtain the score corresponding to the second-level dimension; add up the scores corresponding to the second-level dimensions through the large model to obtain the score corresponding to the first-level dimension. The object evaluation dimension score information includes the first-level dimension and the score corresponding to the first-level dimension.
[0022] Further, the acquisition module is specifically used for: in response to the first selection operation for the target benchmarking dimension and the second selection operation for multiple objects to be benchmarked, or in response to the second selection operation for multiple objects to be benchmarked, obtain the relevant information corresponding to multiple objects to be benchmarked; or, obtain the user question information, identify the intent corresponding to the user question information, and if the intent is to perform document benchmarking, then obtain the relevant information corresponding to multiple objects to be benchmarked according to the user question information.
[0023] Further, when the processing module is used to obtain the target paragraph information corresponding to the object to be benchmarked according to the identification information and the target benchmarking dimension, it is specifically used for: if the target benchmarking dimension is a first-level dimension, obtaining the target paragraph information corresponding to the object to be benchmarked according to the target benchmarking dimension, the second-level dimensions included in the target benchmarking dimension, and the identification information; if the target benchmarking dimension is a second-level dimension, obtaining the target paragraph information corresponding to the object to be benchmarked according to the target benchmarking dimension, the first-level dimension to which the target benchmarking dimension belongs, and the identification information.
[0024] Further, when the processing module is used to obtain the target paragraph information corresponding to the object to be benchmarked according to the target vector and multiple paragraph summary vectors, it is specifically used for: matching the target vector with each paragraph summary vector among the multiple paragraph summary vectors to obtain the similarity between the target vector and each paragraph summary vector; determining the top preset number of target paragraph summary vectors with higher similarity among the multiple paragraph summary vectors whose similarity is less than the similarity threshold; and obtaining the target paragraph information corresponding to the object to be benchmarked according to the target paragraph summary vectors.
[0025] Further, the document benchmarking device further includes a preprocessing module, which is used for: splitting the document content of the relevant document into paragraphs to obtain the original paragraphs, the paragraph sequence numbers of the original paragraphs, and the context paragraph information of the original paragraphs; obtaining the paragraph summary and the benchmarking dimension to which the paragraph summary belongs from the original paragraphs through a large model; performing vectorization processing on multiple paragraph summaries through a document vectorization model to obtain multiple paragraph summary vectors, and the document vectorization model is obtained through unsupervised training based on the paragraph summaries; storing the identification information, benchmarking dimension, original paragraphs, paragraph sequence numbers of the original paragraphs, context paragraph information of the original paragraphs, paragraph summaries, and paragraph summary vectors corresponding to different objects respectively.
[0026] Further, the document content further includes images. When the preprocessing module is used to split the document content of the relevant document into paragraphs to obtain the original paragraphs, it is specifically used for: extracting the text information in the images by using the optical character recognition (OCR) method; and merging the text information with the paragraphs corresponding to the text information to obtain the original paragraphs.
[0027] Further, when the processing module is used to perform vectorization processing on relevant information to obtain a target vector, it is specifically used for: performing vectorization processing on the relevant information through a document vectorization model to obtain the target vector.
[0028] An electronic device includes: a processor, and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the document benchmarking method as described above in the present invention.
[0029] A computer-readable storage medium stores computer program instructions, and when the computer program instructions are executed, the document alignment method as described above in the present invention is implemented.
[0030] A computer program product includes a computer program, and when the computer program is executed, the document alignment method as described above in the present invention is implemented.
[0031] Advantages of the present invention: The document alignment method, device, equipment, storage medium and program product provided by the present invention obtain relevant information corresponding to a plurality of objects to be aligned, and the relevant information includes identification information respectively corresponding to the plurality of objects to be aligned; for each object to be aligned among the plurality of objects to be aligned, if the relevant information includes a target alignment dimension, then according to the identification information and the target alignment dimension, obtain the target paragraph information corresponding to the object to be aligned; if the relevant information does not include the target alignment dimension, then perform vectorization processing on the relevant information to obtain a target vector, and according to the target vector and a plurality of paragraph summary vectors, obtain the target paragraph information corresponding to the object to be aligned; the above way of obtaining the target paragraph information can effectively filter out irrelevant data items, thereby reducing the context length required by the large model and optimizing the utilization rate of the video memory resources of the server; the target paragraph information includes a target paragraph summary, context paragraph information of the original paragraph corresponding to the target paragraph summary, and the alignment dimension to which the target paragraph summary belongs, and the plurality of paragraph summary vectors, alignment dimensions, paragraph summaries, original paragraphs and context paragraph information are obtained by preprocessing relevant documents respectively corresponding to different objects; input the identification information and the target paragraph information into the large model to obtain the target alignment results corresponding to the plurality of objects to be aligned output by the large model, and realize automatically outputting the target alignment results corresponding to the plurality of objects to be aligned through the large model, which can greatly improve the efficiency of document alignment, and can more accurately obtain the target alignment results corresponding to the plurality of objects to be aligned, and improve the reliability and professionalism of document alignment. The present invention can not only improve the automation and intelligence level of document processing, but also provide a new document management and alignment solution for different industries. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0033] Figure 1 It is a schematic diagram of an application scenario provided by an embodiment of the present invention;
[0034] Figure 2Flowchart of the document alignment method provided by an embodiment of the present invention;
[0035] Figure 3 Flowchart of the method for preprocessing relevant documents corresponding to different objects provided by an embodiment of the present invention;
[0036] Figure 4 Flowchart of the document alignment method provided by another embodiment of the present invention;
[0037] Figure 5 Schematic diagram of the prompt template provided by an embodiment of the present invention;
[0038] Figure 6 Flowchart of the method for obtaining the score information of the object evaluation dimension provided by an embodiment of the present invention;
[0039] Figure 7 Schematic diagram of the structure of the document alignment device provided by an embodiment of the present invention;
[0040] Figure 8 Schematic diagram of the structure of the document alignment device provided by another embodiment of the present invention;
[0041] Figure 9 Schematic diagram of the structure of the electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0042] The following will illustrate the implementation manners of the present invention with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred embodiments are only for illustrating the present invention, rather than for limiting the protection scope of the present invention.
[0043] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Therefore, only the components related to the present invention are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and ratio of each component in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.
[0044] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present invention are all information and data that have been authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards, and corresponding operation entrances are provided for users to choose to authorize or reject.
[0045] Comparing relevant documents of different objects can help identify the differences and advantages between different objects, optimize market competition strategies, assist in decision-making, and enhance transparency and trust, etc. Taking the comparison of vehicle model documents as an example, the comparison of vehicle model documents is to conduct a comparative analysis of the technical documents (such as design specifications, configuration parameters, function descriptions, etc.) of different vehicle models during the vehicle design and development process to identify differences, optimize the design, or meet specific requirements. This process can help understand the similarities and differences as well as the product capabilities between different vehicle models, and is crucial in aspects such as the research and development, procurement, and formulation of market strategies in the vehicle industry.
[0046] Currently, the method of manual comparative analysis is usually adopted for document comparison, but the comparison efficiency is low and the comparison results are not accurate enough. Therefore, there is an urgent need for an effective solution for document comparison.
[0047] In addition, with the increasingly wide application of large models, document processing can be carried out based on large models. For example, parsing documents, constructing a knowledge base according to the knowledge obtained from the parsing, and sending the knowledge of the knowledge base into the large model for question answering. Among them, the public basic quantization model is usually adopted for document quantization storage, and the vectorization model exclusive to this application field is not trained according to the application field materials, which will pose challenges to the semantic similarity recall of the large model and ultimately affect the accuracy of the output results of the large model. However, there is currently no implementation solution for comparing documents through large models.
[0048] Based on the above problems, the present invention provides a document alignment method. By preprocessing the document contents of the relevant documents respectively corresponding to different objects, the preprocessed paragraph information respectively corresponding to different objects is obtained, realizing efficient document management and facilitating quick retrieval and matching. Obtain the relevant information corresponding to multiple objects to be aligned, where the relevant information includes the identification information respectively corresponding to multiple objects to be aligned. If the relevant information contains the target alignment dimension, then according to the identification information and the target alignment dimension, directly obtain the target paragraph information corresponding to the object to be aligned. If the relevant information does not contain the target alignment dimension, then perform vectorization processing on the relevant information to obtain a target vector, and according to the target vector and multiple paragraph summary vectors, obtain the target paragraph information corresponding to the object to be aligned. Based on the preprocessed paragraph information, the target paragraph information and the identification information, output the target alignment results corresponding to multiple objects to be aligned through a large model, which can greatly improve the efficiency of document alignment and can more accurately obtain the target alignment results corresponding to multiple objects to be aligned, providing a new document management and alignment solution for the vehicle industry.
[0049] Hereinafter, an example of the application scenario of the solution provided by the present invention will be described first.
[0050] Figure 1 It is a schematic diagram of an application scenario provided by an embodiment of the present invention. As Figure 1 shown, the application scenario may include: a server cluster 11 and a terminal 12; wherein, the server cluster 11 includes multiple servers 111 and a memory 112, and the terminal 12 may be a mobile phone, a tablet computer, a notebook computer, a desktop computer, a smart home appliance, etc. Taking the vehicle model document alignment as an example, the object to be aligned is the vehicle model to be aligned. The user inputs the target alignment dimension and the identification information (i.e., vehicle model information) respectively corresponding to multiple vehicle models to be aligned through the terminal 12. For example, the user inputs "Help me compare the intelligent cockpit dimensions of vehicle model A and vehicle model B" through the terminal 12, where the intelligent cockpit dimension is the target alignment dimension. Correspondingly, the server 111 obtains the target alignment dimension input by the user through the terminal 12 and the vehicle model information respectively corresponding to multiple vehicle models to be aligned. The server 111 obtains the target alignment results corresponding to multiple vehicle models to be aligned according to the document alignment method provided by the embodiment of the present invention, such as outputting the target alignment results corresponding to vehicle model A and vehicle model B. The server 111 sends the target alignment results to the terminal 12 to display the target alignment results to the user through the terminal 12. Among them, the server 111 obtains relevant data from the memory 112 and stores the generated data in the memory 112. In addition, the server 111 and the terminal 12 communicate through a wireless network or a wired network.
[0051] It should be noted that Figure 1 is only a schematic diagram of an application scenario provided by an embodiment of the present invention, and the embodiments of the present invention do not Figure 1does not limit the equipment included in Figure 1 The positional relationship between the devices is limited.
[0052] The document benchmarking method provided in the embodiment of the present invention can be applied to at least: (1) Industry benchmarking: used to analyze the relevant competitiveness of benchmarking enterprises in the industry to which it belongs, which helps to cultivate its own relevant competitiveness; (2) Patent benchmarking: by analyzing the patent applications and authorizations of competitors or in the industry, to understand the technology development trends, identify potential technology cooperation opportunities, and avoid possible intellectual property risks; (3) Product benchmarking: used to analyze the product specifications of competitors or industry leaders to understand the latest trends and technological advances in the market; (4) Technology benchmarking: used to analyze the performance of competitors or leaders in core technologies, R&D investment, innovation methods, and technical standard participation; pay attention to their underlying technical capabilities, R&D system efficiency, and future technology roadmap; help to judge the advancement and sustainability of their own technologies, and find opportunities for technological breakthroughs or cooperation; (5) Customer benchmarking: used to analyze the customer group characteristics, customer satisfaction, customer loyalty, and customer feedback of competitors; understand customers' evaluation of competitors' products and services, as well as their unmet needs, which helps to optimize their own product design, service experience, and customer relationship management.
[0053] The document benchmarking results obtained by the document benchmarking method provided by the embodiment of the present invention can be applied to recommendation scenarios. For example, taking vehicle model document benchmarking as an example, assuming that the models to be benchmarked are model A and model B, and the target benchmarking dimension is the smart cockpit dimension, the target benchmarking results corresponding to model A and model B can be obtained by the document benchmarking method provided by the embodiment of the present invention, and then it can be determined that model A is better than model B in the smart cockpit dimension based on the target benchmarking results, so that model A can be recommended to the user.
[0054] The technical solution of the present invention is described in detail below through specific embodiments. It should be noted that the following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.
[0055] Figure 2 The flowchart of the document matching method provided by one embodiment of the present invention. The document matching method can be executed by software and / or hardware devices. For example, the hardware device can be a document matching device, and the document matching device can be an electronic device or a processing chip in the electronic device. Figure 2 As shown, the method of the embodiment of the present invention includes:
[0056] S201 . Obtain relevant information corresponding to a plurality of objects to be benchmarked, where the relevant information includes identification information corresponding to the plurality of objects to be benchmarked.
[0057] In the embodiments of the present invention, the relevant information corresponding to multiple objects to be benchmarked includes the identification information corresponding to each of the multiple objects to be benchmarked, and may also include the target benchmarking dimension to be benchmarked. It can be understood that the benchmarking dimension may include multiple first-level dimensions, and each first-level dimension may contain multiple second-level dimensions. The benchmarking dimension can be defined as needed, and the embodiments of the present invention do not limit this. Exemplarily, taking the benchmarking of vehicle type documents as an example, the object to be benchmarked is the vehicle type to be benchmarked. The first-level dimensions may include, for example, the intelligent cockpit dimension, the intelligent vehicle control dimension, the intelligent parking dimension, and the intelligent driving dimension. Each first-level dimension may contain multiple second-level dimensions. Specifically, the intelligent cockpit dimension may contain multiple second-level dimensions such as intelligent voice, navigation system, multimedia, imaging, settings, Bluetooth phone, system stability, system fluency, interface aesthetics, display quality, sound quality, and extended functions; the intelligent vehicle control dimension may contain multiple second-level dimensions such as physical key, greeting and sending off guests, lighting control, seat control, mobile phone remote control, digital key, and expansion function; the intelligent parking dimension may contain multiple second-level dimensions such as vision system and automatic parking; the intelligent driving dimension may contain multiple second-level dimensions such as cruise function, lane assistance, traffic sign recognition, and emergency braking.
[0058] Optionally, obtaining the relevant information corresponding to multiple objects to be benchmarked may include: in response to a first selection operation for the target benchmarking dimension and a second selection operation for the multiple objects to be benchmarked, or in response to the second selection operation for the multiple objects to be benchmarked, obtaining the relevant information corresponding to the multiple objects to be benchmarked; or, obtaining user question information, identifying the intention corresponding to the user question information, and if the intention is to perform document benchmarking, obtaining the relevant information corresponding to the multiple objects to be benchmarked according to the user question information.
[0059] It can be understood that the embodiments of the present invention provide two ways to obtain the document alignment results. One is to directly select the identification information of the object to be aligned for document alignment, and the other is to perform document alignment through natural language communication. Taking the vehicle model document alignment as an example, the object to be aligned is the vehicle model to be aligned. In one example, the user can select the target alignment dimension and multiple vehicles to be aligned through the terminal, where the multiple vehicles to be aligned are two or more vehicles to be aligned. For example, the multiple vehicles to be aligned are vehicle model A and vehicle model B. Correspondingly, the electronic device executing the embodiment of this method responds to the first selection operation for the target alignment dimension and the second selection operation for the multiple vehicles to be aligned, and obtains the vehicle information corresponding to the target alignment dimension and the multiple vehicles to be aligned respectively. The vehicle information is, for example, the vehicle model names corresponding to vehicle model A and vehicle model B respectively. In another example, the user can select multiple vehicles to be aligned through the terminal without selecting the target alignment dimension. Correspondingly, the electronic device executing the embodiment of this method responds to the second selection operation for the multiple vehicles to be aligned and obtains the vehicle information corresponding to the multiple vehicles to be aligned respectively. Since the user does not select the target alignment dimension, the default target alignment dimension is all first-level dimensions. In yet another example, the user can input a question sentence "Help me compare the intelligent cockpit dimensions of vehicle model A and vehicle model B" through the terminal, where the intelligent cockpit dimension is the target alignment dimension. Correspondingly, the electronic device executing the embodiment of this method can obtain the user's question sentence information, identify the intention corresponding to the user's question sentence information. If the intention is to perform vehicle model document alignment, then perform slot extraction on the user's question sentence information to obtain the vehicle information corresponding to the target alignment dimension and the multiple vehicles to be aligned respectively. Among them, when identifying the intention corresponding to the user's question sentence information, an intention recognition module can be built using the Rasa framework (a machine learning framework for building conversational artificial intelligence applications), and the intention recognition module is used to identify the intention corresponding to the user's question sentence information. For specific reference, please refer to the subsequent embodiments. The user's question sentence information may not contain the target alignment dimension and contains other relevant information about the vehicle model, then the entire user's question sentence information can be used as the relevant information corresponding to the multiple vehicles to be aligned.
[0060] S202. For each object to be aligned among the multiple objects to be aligned, if the relevant information includes the target alignment dimension, then according to the identification information and the target alignment dimension, obtain the target paragraph information corresponding to the object to be aligned; if the relevant information does not include the target alignment dimension, then perform vectorization processing on the relevant information to obtain a target vector, and according to the target vector and the multiple paragraph summary vectors, obtain the target paragraph information corresponding to the object to be aligned; the target paragraph information includes the target paragraph summary, the context paragraph information of the original paragraph corresponding to the target paragraph summary, and the alignment dimension to which the target paragraph summary belongs. The multiple paragraph summary vectors, alignment dimensions, paragraph summaries, original paragraphs, and context paragraph information are obtained by preprocessing the relevant documents corresponding to different objects respectively.
[0061] In this step, multiple passage abstract vectors, alignment dimensions, passage abstracts, original passages, and context passage information are obtained by preprocessing relevant documents corresponding to different objects respectively. For how to preprocess relevant documents corresponding to different objects specifically, reference can be made to the subsequent embodiments and will not be elaborated here. After preprocessing relevant documents corresponding to different objects, for example, the identification information, original passages, context passage information of the original passages, passage abstracts, alignment dimensions to which the passage abstracts belong, and passage abstract vectors corresponding to different objects can be stored in a vector database for obtaining the target passage information corresponding to the object to be aligned in this step. Among them, the alignment dimension to which the passage abstract belongs is the first-level dimension to which the target passage abstract belongs and all second-level dimensions included under this first-level dimension. And the target alignment dimension is obtained based on relevant information corresponding to multiple objects to be aligned. For example, after obtaining the user question information input by the user, it can be determined whether the user question information contains the target alignment dimension, and this target alignment dimension is the dimension that the user wants to use for document alignment, which can be a first-level dimension, or a second-level dimension, or a first-level dimension and a second-level dimension.
[0062] When the relevant information contains the target alignment dimension, optionally, according to the identification information and the target alignment dimension, obtaining the target passage information corresponding to the object to be aligned may include: if the target alignment dimension is a first-level dimension, obtaining the target passage information corresponding to the object to be aligned according to the target alignment dimension, the second-level dimensions included in the target alignment dimension, and the identification information; if the target alignment dimension is a second-level dimension, obtaining the target passage information corresponding to the object to be aligned according to the target alignment dimension, the first-level dimension to which the target alignment dimension belongs, and the identification information.
[0063] Exemplarily, after obtaining the target alignment dimension and the identification information corresponding to multiple objects to be aligned respectively, if the target alignment dimension is a first-level dimension, all second-level dimensions included in the target alignment dimension can be obtained, and then, according to the target alignment dimension, all second-level dimensions included in the target alignment dimension, and the identification information, directly query the vector database to obtain the target passage information corresponding to the object to be aligned. If the target alignment dimension is a second-level dimension, the first-level dimension to which the target alignment dimension belongs can be obtained, and then, according to the target alignment dimension, the first-level dimension to which the target alignment dimension belongs, and the identification information, directly query the vector database to obtain the target passage information corresponding to the object to be aligned.
[0064] In the case where the relevant information does not include the target benchmarking dimension, for example, if the relevant information corresponding to multiple objects to be benchmarked is the user's question information, the relevant information can be directly vectorized to obtain the target vector. The target vector is matched with each paragraph summary vector among the multiple paragraph summary vectors to obtain the similarity between the target vector and each paragraph summary vector, so as to obtain the target paragraph information corresponding to the object to be benchmarked according to the similarity.
[0065] Obtaining the target paragraph information based on the above method can effectively filter out irrelevant data items, and further reduce the context length required by the large model, optimizing the video memory resource utilization rate of the electronic device implementing the embodiment of this method. For how to specifically obtain the target paragraph information corresponding to the object to be benchmarked according to the target vector and multiple paragraph summary vectors, reference can be made to the subsequent embodiments. Taking the vehicle model document benchmarking as an example, the object to be benchmarked is the vehicle model to be benchmarked. Assuming the vehicle models to be benchmarked are vehicle model A and vehicle model B, the target paragraph information corresponding to vehicle model A and the target paragraph information corresponding to vehicle model B can be obtained respectively. Among them, the target paragraph information includes the target paragraph summary, the context paragraph information of the original paragraph corresponding to the target paragraph summary, and the benchmarking dimension to which the target paragraph summary belongs.
[0066] S203. Input the identification information and the target paragraph information into the large model to obtain the target benchmarking results corresponding to multiple objects to be benchmarked output by the large model.
[0067] Exemplarily, in the embodiment of the present invention, the specific large model can be selected from the currently mainstream large models. The large model is, for example, a large model containing at least 7 billion (7B) parameters. After obtaining the target paragraph information corresponding to the object to be benchmarked, the identification information and the target paragraph information corresponding to multiple objects to be benchmarked can be input into the large model respectively to obtain the target benchmarking results corresponding to multiple objects to be benchmarked output by the large model. For how to specifically obtain the target benchmarking results through the large model, reference can be made to the subsequent embodiments.
[0068] The document alignment method provided by the embodiments of the present invention obtains relevant information corresponding to multiple objects to be aligned. The relevant information includes identification information corresponding to multiple objects to be aligned respectively. For each object to be aligned among the multiple objects to be aligned, if the relevant information includes a target alignment dimension, the target paragraph information corresponding to the object to be aligned is obtained according to the identification information and the target alignment dimension. If the relevant information does not include the target alignment dimension, the relevant information is vectorized to obtain a target vector. The target paragraph information corresponding to the object to be aligned is obtained according to the target vector and multiple paragraph summary vectors. The above way of obtaining the target paragraph information can effectively filter out irrelevant data items, thereby reducing the context length required by the large model and optimizing the utilization rate of the video memory resources of the server. The target paragraph information includes a target paragraph summary, context paragraph information of the original paragraph corresponding to the target paragraph summary, and the alignment dimension to which the target paragraph summary belongs. The multiple paragraph summary vectors, alignment dimensions, paragraph summaries, original paragraphs, and context paragraph information are obtained by preprocessing the relevant documents corresponding to different objects respectively. The identification information and the target paragraph information are input into the large model to obtain the target alignment results corresponding to multiple objects to be aligned output by the large model, realizing the automatic output of the target alignment results corresponding to multiple objects to be aligned through the large model, which can greatly improve the efficiency of document alignment and can more accurately obtain the target alignment results corresponding to multiple objects to be aligned, enhancing the reliability and professionalism of document alignment. The embodiments of the present invention can not only improve the automation and intelligence level of document processing, but also provide a new document management and alignment solution for different industries.
[0069] Based on the above embodiments, Figure 3 is a flowchart of a method for preprocessing relevant documents corresponding to different objects provided by an embodiment of the present invention. As Figure 3 shown, the method of the embodiments of the present invention may include:
[0070] S301. Split the document content of the relevant documents corresponding to different objects into paragraphs respectively to obtain the original paragraphs, the paragraph sequence numbers of the original paragraphs, and the context paragraph information of the original paragraphs.
[0071] Exemplarily, taking the vehicle model document alignment as an example, different objects are different vehicle models. Taking the vehicle model as vehicle model A as an example, assuming that the document content of the vehicle model-related document corresponding to vehicle model A contains 3 paragraphs, then the document content of the vehicle model-related document is split into paragraphs respectively to obtain 3 original paragraphs, the paragraph sequence numbers corresponding to the 3 original paragraphs respectively, and the context paragraph information of the original paragraphs. The context paragraph information of the original paragraphs may include, for example, the previous original paragraph and the next original paragraph adjacent to the current original paragraph.
[0072] Optionally, the document content also includes an image, and splitting the document content of the relevant document into paragraphs to obtain original paragraphs may include: extracting text information from the image using an optical character recognition (OCR) method; and merging the text information with the paragraph corresponding to the text information to obtain the original paragraph.
[0073] Exemplarily, when the document content also includes an image, an OCR method may be used to extract text information from the image, and the text information may be merged with the paragraph corresponding to the text information, such as inserting the text information into the paragraph corresponding to the text information, thereby obtaining the final original paragraph.
[0074] It can be understood that by splitting the document contents of the related documents corresponding to different objects into paragraphs through this step, a paragraph set can be finally formed. The document formats of the related documents can include PDF, DOCX, PPTX, TXT, etc. The embodiment of the present invention can finely read and divide the document contents of related documents in various formats, and use the OCR method to extract text information in the image, which can effectively improve the integrity of the document content.
[0075] S302: Obtain a paragraph summary and a benchmarking dimension to which the paragraph summary belongs from the original paragraph through a large model.
[0076] For example, after obtaining the paragraph set corresponding to different objects, the key information can be extracted from the original paragraph through the big model to obtain the paragraph summary and the benchmarking dimension to which the paragraph summary belongs, and finally a paragraph summary set can be formed. By extracting the paragraph summary and the benchmarking dimension to which the paragraph summary belongs from the paragraph set through the prompt word (Prompt) of the big model, the vectorization of the document can be made more accurate. For example, taking the vehicle model document benchmarking as an example, assuming that the user inputs "help me compare the intelligent voice of vehicle model A and vehicle model B", it can be determined that the target benchmarking dimension is "intelligent voice". The specific acquisition process of the target benchmarking dimension is as follows: the original paragraph is: Voice: supports 5-zone recognition; Voice control ability: supports visible and speaking, cross-zone inheritance, supports multiple intentions for one language, and joins the big model dialogue. It also supports fast vehicle control adjustment, such as voice to raise the rearview mirror angle, adjust the steering wheel height, and support "pause" interruption during the adjustment process; the voice of the mobile phone and the car machine is connected: the car machine voice calls the mobile phone application. According to the original paragraph, the big model can be used to obtain the benchmark dimension to which the paragraph summary belongs from the original paragraph: smart cockpit#intelligent voice, where smart cockpit is the first-level dimension and intelligent voice is the second-level dimension.
[0077] S303, vectorizing the multiple paragraph summaries through a document vectorization model to obtain multiple paragraph summary vectors, wherein the document vectorization model is obtained by performing unsupervised training based on the paragraph summaries.
[0078] In this step, a document summary database can be constructed using the paragraph summary set, and the document vectorization model can be unsupervised trained based on the paragraph summaries in the document summary database, so as to obtain a trained document vectorization model that is more suitable for document benchmarking. The document vectorization model can accurately convert the document content of relevant documents corresponding to different objects into vector representations in a high-dimensional space, facilitating subsequent querying and analysis. Exemplarily, multiple paragraph summary vectors can be obtained by vectorizing multiple paragraph summaries through the document vectorization model.
[0079] S304. Store the identification information, benchmarking dimension, original paragraph, paragraph sequence number of the original paragraph, context paragraph information of the original paragraph, paragraph summary, and paragraph summary vector corresponding to different objects respectively.
[0080] Exemplarily, the identification information, benchmarking dimension, original paragraph, paragraph sequence number of the original paragraph, context paragraph information of the original paragraph, paragraph summary, and paragraph summary vector corresponding to different objects can be batch stored in a vector database to improve the efficiency and accuracy of matching the target vector and multiple paragraph summary vectors in the above embodiments. Among them, the vector database is, for example, a Milvus vector database.
[0081] In the embodiments of the present invention, by splitting the document content of relevant documents corresponding to different objects paragraph by paragraph, the original paragraph, the paragraph sequence number of the original paragraph, and the context paragraph information of the original paragraph are obtained. By using a large model to obtain the paragraph summary and the benchmarking dimension to which the paragraph summary belongs from the original paragraph, the paragraph summary and the benchmarking dimension to which the paragraph summary belongs can be accurately extracted; by vectorizing multiple paragraph summaries through the document vectorization model, multiple paragraph summary vectors are obtained; among them, the document vectorization model is obtained through unsupervised training based on paragraph summaries, which can reduce the manual annotation cost and improve the generality of model training, so as to accurately convert the document content into vector representations, making subsequent vector matching and querying more efficient and accurate; storing the identification information, benchmarking dimension, original paragraph, paragraph sequence number of the original paragraph, context paragraph information of the original paragraph, paragraph summary, and paragraph summary vector corresponding to different objects respectively realizes efficient document management and facilitates fast retrieval and matching.
[0082] Figure 4 It is a flowchart of a document benchmarking method provided by another embodiment of the present invention. On the basis of the above embodiments, the embodiments of the present invention further illustrate the document benchmarking method, where the method of document benchmarking through natural language communication is taken as an example. As Figure 4 shown, the method of the embodiments of the present invention may include:
[0083] S401. Obtain the user's question information, identify the intent corresponding to the user's question information. If the intent is to perform document alignment, obtain the relevant information corresponding to multiple objects to be aligned according to the user's question information. The relevant information includes the identification information corresponding to each of the multiple objects to be aligned.
[0084] Exemplarily, taking the alignment of vehicle type documents as an example, the objects to be aligned are the vehicle types to be aligned. The user can input a question "Help me compare the intelligent cockpits of vehicle type A and vehicle type B" through the terminal. Correspondingly, the electronic device executing the embodiment of this method can obtain the user's question information and identify the intent corresponding to the user's question information. If the intent is to perform vehicle type document alignment, perform slot extraction on the user's question information to obtain the relevant information corresponding to multiple vehicle types to be aligned. This relevant information may include, for example, the target alignment dimension and the identification information (i.e., vehicle type information) corresponding to each of the multiple vehicle types to be aligned. Among them, when identifying the intent corresponding to the user's question information, an intent recognition module can be built using the Rasa framework, and the intent recognition module is used to identify the intent corresponding to the user's question information. Specifically, building an intent recognition module using the Rasa framework may include the following modules:
[0085] (1) SpacyNLP module, which is used to perform in-depth semantic analysis on Chinese text using the spaCy library (a high-level natural language processing library), including part-of-speech tagging, dependency parsing, etc.; in the embodiment of the present invention, for example, the "zh_core_web_sm" model that supports Chinese semantics can be selected.
[0086] (2) SpacyTokenizer module, which is used to decompose the input text string into tokens. These tokens may include words, characters, sub-words, etc., laying a solid foundation for subsequent language processing and analysis steps, and ensuring the accuracy and operability of text data in the preprocessing stage.
[0087] (3) SpacyFeaturizer module, which is used to extract a series of detailed and in-depth linguistic features from text data, covering multiple dimensions such as part-of-speech tagging, named entity recognition, dependency syntactic analysis, and lexical feature extraction. These features provide rich semantic information for machine learning models and play a crucial role in deeply understanding and interpreting the inherent meaning of the text.
[0088] (4) RegexEntityExtractor module, which is used to accurately identify and extract specific entities in the text, such as structured information like dates and phone numbers; especially when implementing object information extraction, for example, the method of "use_lookup_tables" can be adopted, which can not only improve the accuracy of entity recognition but also enhance the adaptability and robustness of the system in the face of diverse text inputs.
[0089] (5) LexicalSyntacticFeaturizer module, which is used to deeply explore the lexical and syntactic levels of the text. By extracting key features such as morphological changes of words, word order, and phrase structure, it provides a window for the model to deeply understand the deep structure and semantic relationships of the text, so as to more accurately grasp the meaning of information and the overall framework of the text in complex contexts.
[0090] (6) CountVectorsFeaturizer module, which is used to convert the text into numerical vectors. Adopting the bag-of-words model technology, it converts the original text content into a set of quantitative representations. These numerical vectors can effectively capture the keyword and phrase frequencies in the text, providing a simple and powerful feature expression method for subsequent machine learning models, enabling the model to focus on the most discriminative language elements in the text.
[0091] (7) DIETClassifier module, that is, the intent classifier, which is used to accurately classify the user's intent during the training process through an end-to-end learning method, and can also synchronously identify and extract key entity information. It can greatly improve the generalization ability of the model and the flexibility in dealing with complex dialogue scenarios, providing a more accurate interaction experience for users.
[0092] (8) EntitySynonymMapper module, which acts as a role of entity normalization, and is used to cleverly map various synonyms, near-synonyms or variants in the user input to the predefined standard entity forms. This process can significantly enhance the model's understanding and recognition ability of different expression ways, thereby improving the accuracy and user-friendliness of the entire dialogue system.
[0093] (9) ResponseSelector module, which is used to select the most appropriate answer from the predefined response options to build a rule-based dialogue system.
[0094] It can be understood that by identifying the intent corresponding to the user's question information through the intent recognition module, the user's query can be made more intelligent, realizing the precise matching of the user's intent and the rapid retrieval of identification information, improving the user experience, and thus helping to improve the accuracy of document alignment.
[0095] S402. For each of the multiple objects to be aligned, if the relevant information contains the target alignment dimension, then according to the identification information and the target alignment dimension, obtain the target paragraph information corresponding to the object to be aligned; wherein, the target paragraph information includes the target paragraph summary, the context paragraph information of the original paragraph corresponding to the target paragraph summary, and the alignment dimension to which the target paragraph summary belongs.
[0096] For the specific description of this step, reference can be made to Figure 2The relevant description of S202 in the illustrated embodiment will not be elaborated here.
[0097] In the embodiment of the present invention, Figure 2 Step S202 may further include the following four steps of S403 to S406:
[0098] S403. For each of the multiple objects to be benchmarked, if the relevant information does not include the target benchmark dimension, the relevant information is vectorized through a document vectorization model to obtain a target vector.
[0099] Exemplarily, the document vectorization model is obtained through unsupervised training, and the relevant information can be vectorized through the document vectorization model to obtain a target vector.
[0100] S404. Match the target vector with each of the multiple passage summary vectors to obtain the similarity between the target vector and each passage summary vector.
[0101] Exemplarily, referring to Figure 3 the example of the step, the multiple passage summary vectors corresponding to different objects are stored in a vector database. After obtaining the target vector, it can be retrieved in the vector database, and the target vector is matched with each passage summary vector in the vector database, so as to obtain the similarity between the target vector and each passage summary vector. It can be understood that since the target vector is obtained based on the target benchmark dimension and the secondary dimensions included in the target benchmark dimension, when retrieving in the vector database according to the target vector, the retrieval of irrelevant information can be reduced, thereby shortening the content (context) length of the large model and reducing the video memory occupancy of the server.
[0102] S405. Determine the top pre-set number of target passage summary vectors with a similarity less than the similarity threshold among the multiple passage summary vectors.
[0103] Exemplarily, if the similarity threshold is, for example, 1 and the pre-set number is, for example, N, the similarities can be sorted, for example, in ascending order, to determine the top N (topN) target passage summary vectors with a similarity less than the similarity threshold among the multiple passage summary vectors.
[0104] S406. Obtain the target paragraph information corresponding to the object to be benchmarked according to the target paragraph summary vector; wherein, the target paragraph information includes the target paragraph summary, the context paragraph information of the original paragraph corresponding to the target paragraph summary, and the benchmarking dimension to which the target paragraph summary belongs. For example, taking the benchmarking of vehicle type documents as an example, assuming the user inputs "Help me compare the human-machine voice interaction functions of vehicle type A and vehicle type B", the target benchmarking dimension cannot be determined. The specific process of obtaining the benchmarking dimension is as follows: The original paragraph is: Voice: Supports 5-zone recognition; Voice control ability: Supports "what you see is what you say", cross-zone inheritance, supports multiple intents in one sentence, and joins the large model conversation. It also supports fast vehicle control adjustment, such as adjusting the rearview mirror angle and the steering wheel height by voice, and supports "pause" interruption during the adjustment process; The voice connection between the mobile phone and the in-vehicle system: Supports calling mobile phone applications by in-vehicle system voice. The benchmarking dimension to which it belongs: Intelligent cockpit#Intelligent voice, where intelligent cockpit is the first-level dimension and intelligent voice is the second-level dimension. According to the original paragraph summary vector matched by the original input, the corresponding target paragraph summary, the context paragraph information of the original paragraph corresponding to the target paragraph summary, and the benchmarking dimension to which the target paragraph summary belongs, that is, intelligent cockpit#Intelligent voice, are brought out, so as to realize that in the case of missing the target benchmarking dimension, the benchmarking ability can uniformly descend to the benchmarking dimension.
[0105] Exemplarily, the vector database can be queried according to the target paragraph summary vector to obtain the target paragraph summary corresponding to the target paragraph summary vector, the context paragraph information of the original paragraph corresponding to the target paragraph summary, and the benchmarking dimension to which the target paragraph summary belongs, that is, obtain the target paragraph information corresponding to the object to be benchmarked.
[0106] In the embodiments of the present invention, Figure 2 Step S203 in can further include the following two steps of S407 and S408:
[0107] S407. Determine the target prompt word information corresponding to the large model according to the identification information, the object evaluation dimension score information, and the target paragraph information, and the object evaluation dimension score information is obtained by the large model evaluating multiple benchmarking dimensions based on the paragraph summary.
[0108] In this step, the object evaluation dimension score information is obtained by the large model evaluating multiple benchmarking dimensions based on the passage summary, and includes multiple first-level dimensions and the scores corresponding to the first-level dimensions. For the specific method of obtaining the object evaluation dimension score information, reference can be made to the subsequent embodiments, which will not be elaborated here. Exemplarily, after obtaining the target passage information corresponding to the object to be benchmarked, the target prompt word information corresponding to the large model can be determined according to the identification information, object evaluation dimension score information, and target passage information respectively corresponding to multiple objects to be benchmarked. The target prompt word information includes the identification information, object evaluation dimension score information, and target passage information respectively corresponding to multiple objects to be benchmarked. Among them, the target passage information includes the target passage summary, the context passage information of the original passage corresponding to the target passage summary, and the benchmarking dimension to which the target passage summary belongs.
[0109] S408. Input the target prompt word information into the large model to obtain the target benchmarking results corresponding to multiple objects to be benchmarked output by the large model.
[0110] Exemplarily, after determining the target prompt word information corresponding to the large model, a prompt template can be spliced according to the target prompt word information. Figure 5 The following is a schematic diagram of the prompt template provided by an embodiment of the present invention. As Figure 5 shown, taking the benchmarking of vehicle type documents as an example, the object to be benchmarked is the vehicle type to be benchmarked. Taking the vehicle types to be benchmarked as vehicle type A and vehicle type B as an example, the scoring information of vehicle type A is the object evaluation dimension score information of vehicle type A, and the scoring information of vehicle type B is the object evaluation dimension score information of vehicle type B; the relevant information of vehicle type A is the target passage information corresponding to vehicle type A, and the relevant information of vehicle type B is the target passage information corresponding to vehicle type B. Among them, the target passage information is obtained based on the target benchmarking dimension. Therefore, the target passage information can correspond to different target benchmarking dimensions and the secondary dimensions included in the target benchmarking dimension. For example, "Intelligent Cockpit#Intelligent Voice" represents the intelligent cockpit dimension and the intelligent voice, which is a secondary dimension included in the intelligent cockpit dimension. In this step, inputting the spliced prompt template into the large model can obtain the target benchmarking results corresponding to multiple objects to be benchmarked output by the large model, and can more accurately obtain the target benchmarking results corresponding to multiple objects to be benchmarked, improving the reliability and professionalism of document benchmarking.
[0111] The document alignment method provided by the embodiments of the present invention can make the user query more intelligent, achieve precise matching of user intentions and rapid retrieval of identification information, and improve the user experience by obtaining user question information, identifying the intention corresponding to the user question information, and if the intention is to perform document alignment, obtaining relevant information corresponding to multiple objects to be aligned according to the user question information, where the relevant information includes identification information corresponding to each of the multiple objects to be aligned; for each object to be aligned among the multiple objects to be aligned, if the relevant information includes the target alignment dimension, obtaining the target paragraph information corresponding to the object to be aligned according to the identification information and the target alignment dimension; if the relevant information does not include the target alignment dimension, performing vectorization processing on the relevant information through a document vectorization model to obtain a target vector; matching the target vector with each paragraph summary vector among the multiple paragraph summary vectors, obtaining the similarity between the target vector and each paragraph summary vector, and determining the first preset number of target paragraph summary vectors with relatively high similarity and a similarity less than the similarity threshold among the multiple paragraph summary vectors, which can quickly and accurately obtain the target paragraph summary vectors; obtaining the target paragraph information corresponding to the object to be aligned according to the target paragraph summary vectors, where the target paragraph information includes the target paragraph summary, the context paragraph information of the original paragraph corresponding to the target paragraph summary, and the alignment dimension to which the target paragraph summary belongs; determining the target prompt word information corresponding to the large model according to the identification information, the object evaluation dimension score information, and the target paragraph information corresponding to each of the multiple objects to be aligned, where the object evaluation dimension score information is obtained by the large model based on the evaluation of multiple alignment dimensions for the paragraph summary; inputting the target prompt word information into the large model to obtain the target alignment results corresponding to the multiple objects to be aligned output by the large model, realizing the automatic output of the target alignment results corresponding to the multiple objects to be aligned through the large model, which can greatly improve the efficiency of document alignment and can more accurately obtain the target alignment results corresponding to the multiple objects to be aligned.
[0112] Based on the above embodiments, Figure 6 is a flowchart of a method for obtaining object evaluation dimension score information provided by an embodiment of the present invention. The alignment dimension includes multiple first-level dimensions and second-level dimensions included in each first-level dimension. As Figure 6 shown, the method of the embodiment of the present invention may include:
[0113] S601. For each second-level dimension included in the first-level dimension, extract target data related to the scoring index from the paragraph summaries corresponding to different objects through the large model according to the scoring index corresponding to the second-level dimension; score the second-level dimension according to the target data and the scoring standard corresponding to the second-level dimension to obtain the score corresponding to the second-level dimension.
[0114] Exemplarily, the large model uses the paragraph summaries corresponding to different objects to evaluate the preset first-level dimensions. Each first-level dimension contains multiple second-level dimensions. The large model scores each second-level dimension according to the paragraph summary, and the scoring basis is the predefined scoring criteria and scoring metrics. Specifically, the scoring process for each second-level dimension is as follows: Determine the scoring criteria and scoring metrics for this second-level dimension, then analyze the relevant information in the paragraph summary set, extract the target data related to the scoring metrics, and according to the extracted target data, compare with the scoring criteria to score each second-level dimension, thus completing the scoring of all second-level dimensions.
[0115] S602. Add up the scores corresponding to the second-level dimensions through the large model to obtain the score corresponding to the first-level dimension. The object evaluation dimension score information includes the first-level dimension and the score corresponding to the first-level dimension.
[0116] Exemplarily, after obtaining the scores corresponding to each second-level dimension included in each first-level dimension, the large model can add up the scores corresponding to all second-level dimensions under each first-level dimension to obtain the total score (score) corresponding to the first-level dimension, thereby obtaining a detailed object evaluation dimension score information. The object evaluation dimension score information includes the first-level dimension and the score corresponding to the first-level dimension. The object evaluation dimension score information can be managed through a table. Correspondingly, a detailed object evaluation dimension score table can be obtained.
[0117] The method for obtaining object evaluation dimension score information provided by the embodiments of the present invention, for each second-level dimension included in the first-level dimension, extracts the target data related to the scoring metrics from the paragraph summaries corresponding to different objects through the large model according to the scoring metrics corresponding to the second-level dimension; scores the second-level dimension according to the target data and the scoring criteria corresponding to the second-level dimension to obtain the score corresponding to the second-level dimension; adds up the scores corresponding to the second-level dimensions through the large model to obtain the score corresponding to the first-level dimension. The object evaluation dimension score information includes the first-level dimension and the score corresponding to the first-level dimension. Through the large model, the embodiments of the present invention can accurately obtain the object evaluation dimension score information, so as to use the object evaluation dimension score information to accurately obtain the target benchmarking results corresponding to multiple objects to be benchmarked.
[0118] Based on the above embodiments, the large model in the embodiments of the present invention for outputting the target benchmarking results corresponding to multiple objects to be benchmarked has the following differences in technical implementation from the large models in the current related technologies: The large models in the related technologies only trigger the pre-trained knowledge base inside the large model through prompts (Prompts), and rely on the semantic understanding and reasoning ability of the large model itself to generate answers; the output dimensions are automatically extracted by the large model according to the keywords of the input question, lacking explicit constraints on the dimension hierarchy; only relying on the historical corpus during the training of the large model, it is impossible to dynamically introduce external knowledge or preset structured dimensions; it may ignore specific dimensions required by users, or result in incorrect information due to training data bias; if the input question involves a field that the large model has not been fully trained in (such as emerging technologies), it is easy to cause the "hallucination" problem of the large model and output false information. However, the large model in the embodiments of the present invention forcibly constrains the input structure of the large model through a preset multi-level dimension framework (such as first-level dimensions and second-level dimensions), clearly specifies dimension information in the prompt, ensures that the large model must analyze and output for each preset dimension, avoiding omission or random selection, and combines the preset domain knowledge base (i.e., the data in the vector database), and calls it in real time during the process of generating the target benchmarking results, which can ensure the consistency of the output content with the predefined dimensions and knowledge.
[0119] Correspondingly, in terms of technical effects, the output dimensions of the large models in the related technologies are determined by the large models themselves, and it is impossible to ensure coverage of, for example, the second-level dimensions specified by users; due to relying on the internal knowledge of the large model, it is easy to cause incorrect output. However, the large model in the embodiments of the present invention ensures that the output content of the large model covers the preset dimensions 100% through the preset dimensions (i.e., first-level dimensions and second-level dimensions), and all output content is based on the preset domain knowledge base. The large model outputs content based on the domain knowledge base, which can avoid the "hallucination" problem, that is, the large model will not attempt to output content not recorded in the domain knowledge base.
[0120] In addition, the large models in the related technologies have certain limitations in application. If it is necessary to switch to a new field for application, it completely depends on the large model to re-learn, and it is impossible to quickly adapt to the new field. Moreover, when a user newly adds dimensions temporarily, it is necessary to re-train the large model or substantially modify the prompt. However, the large model in the embodiments of the present invention can quickly adapt to the new field or requirement changes by modifying the preset dimensions, and supports regular updates of the preset domain knowledge base (such as introducing the latest industry standards) to ensure that the output content is always consistent with the latest technical specifications.
[0121] The following is an embodiment of the device of the present invention, which can be used to execute the embodiment of the method of the present invention. For the details not disclosed in the embodiment of the device of the present invention, please refer to the embodiment of the method of the present invention.
[0122] Figure 7 It is a schematic structural diagram of a document benchmarking device provided by an embodiment of the present invention. AsFigure 7 As shown in Figure 7 , the document alignment device 700 according to an embodiment of the present invention includes: an acquisition module 701, a processing module 702, and an output module 703. Among them:
[0123] The acquisition module 701 is configured to acquire relevant information corresponding to a plurality of objects to be aligned, and the relevant information includes identification information corresponding to the plurality of objects to be aligned respectively.
[0124] The processing module 702 is configured to, for each object to be aligned among the plurality of objects to be aligned, if the relevant information includes a target alignment dimension, obtain the target paragraph information corresponding to the object to be aligned according to the identification information and the target alignment dimension; if the relevant information does not include the target alignment dimension, perform vectorization processing on the relevant information to obtain a target vector, and obtain the target paragraph information corresponding to the object to be aligned according to the target vector and a plurality of paragraph summary vectors; the target paragraph information includes a target paragraph summary, context paragraph information of the original paragraph corresponding to the target paragraph summary, and the alignment dimension to which the target paragraph summary belongs, and the plurality of paragraph summary vectors, alignment dimensions, paragraph summaries, original paragraphs, and context paragraph information are obtained by preprocessing the relevant documents corresponding to different objects respectively.
[0125] The output module 703 is configured to input the identification information and the target paragraph information into a large model to obtain the target alignment results corresponding to the plurality of objects to be aligned output by the large model.
[0126] Further, the output module 703 may specifically be configured to: determine the target prompt word information corresponding to the large model according to the identification information, object evaluation dimension score information, and target paragraph information, where the object evaluation dimension score information is obtained by the large model based on the paragraph summary for evaluating a plurality of alignment dimensions; input the target prompt word information into the large model to obtain the target alignment results corresponding to the plurality of objects to be aligned output by the large model.
[0127] Further, the alignment dimension includes a plurality of first-level dimensions and second-level dimensions included in each first-level dimension. The acquisition module 701 may also be configured to obtain the object evaluation dimension score information in the following manner: for each second-level dimension included in the first-level dimension, extract target data related to the scoring index from the paragraph summary through the large model according to the scoring index corresponding to the second-level dimension; score the second-level dimension according to the target data and the scoring standard corresponding to the second-level dimension to obtain the score corresponding to the second-level dimension; sum up the scores corresponding to the second-level dimensions through the large model to obtain the score corresponding to the first-level dimension, and the object evaluation dimension score information includes the first-level dimension and the score corresponding to the first-level dimension.
[0128] Further, the obtaining module 701 can be specifically configured to: in response to a first selection operation for a target benchmarking dimension and a second selection operation for multiple objects to be benchmarked, or in response to the second selection operation for multiple objects to be benchmarked, obtain relevant information corresponding to the multiple objects to be benchmarked; or, obtain user question information, identify the intent corresponding to the user question information, and if the intent is to perform document benchmarking, obtain relevant information corresponding to the multiple objects to be benchmarked according to the user question information.
[0129] Further, when the processing module 702 is configured to obtain target paragraph information corresponding to an object to be benchmarked according to identification information and a target benchmarking dimension, it can be specifically configured to: if the target benchmarking dimension is a first-level dimension, obtain the target paragraph information corresponding to the object to be benchmarked according to the target benchmarking dimension, the second-level dimensions included in the target benchmarking dimension, and the identification information; if the target benchmarking dimension is a second-level dimension, obtain the target paragraph information corresponding to the object to be benchmarked according to the target benchmarking dimension, the first-level dimension to which the target benchmarking dimension belongs, and the identification information.
[0130] Further, when the processing module 702 is configured to obtain target paragraph information corresponding to an object to be benchmarked according to a target vector and multiple paragraph summary vectors, it can be specifically configured to: match the target vector with each paragraph summary vector among the multiple paragraph summary vectors to obtain the similarity between the target vector and each paragraph summary vector; determine the top preset number of target paragraph summary vectors with higher similarity whose similarity is less than a similarity threshold among the multiple paragraph summary vectors; and obtain the target paragraph information corresponding to the object to be benchmarked according to the target paragraph summary vectors.
[0131] Further, the document benchmarking apparatus 700 may further include a preprocessing module ( Figure 7 not shown in the figure) for: splitting the document content of a relevant document into paragraphs to obtain the original paragraphs, the paragraph sequence numbers of the original paragraphs, and the context paragraph information of the original paragraphs; obtaining paragraph summaries and the benchmarking dimensions to which the paragraph summaries belong from the original paragraphs through a large model; performing vectorization processing on multiple paragraph summaries through a document vectorization model, where the document vectorization model is obtained through unsupervised training based on paragraph summaries; and storing the identification information, benchmarking dimensions, original paragraphs, paragraph sequence numbers of the original paragraphs, context paragraph information of the original paragraphs, paragraph summaries, and paragraph summary vectors corresponding to different objects respectively.
[0132] Further, the document content further includes images. When the preprocessing module is configured to split the document content of a relevant document into paragraphs to obtain the original paragraphs, it can be specifically configured to: extract the text information in the images by using the optical character recognition (OCR) method; and merge the text information with the paragraphs corresponding to the text information to obtain the original paragraphs.
[0133] Further, when the processing module 702 is used to perform vectorization processing on relevant information to obtain a target vector, it can specifically be used to: perform vectorization processing on relevant information through a document vectorization model to obtain a target vector.
[0134] The device according to the embodiment of the present invention can be used to execute the technical solutions of any of the above - shown method embodiments. The implementation principles and technical effects are similar and will not be elaborated here.
[0135] Figure 8 It is a structural schematic diagram of a document alignment device provided by another embodiment of the present invention. As Figure 8 shown, the document alignment device 800 according to the embodiment of the present invention includes: a document content processing module 801, a document content quantization module 802, a document content storage module 803, an object index scoring module 804, and an object comprehensive comparison module 805. Among them:
[0136] The document content processing module 801 is used to split and segment the document content of relevant documents respectively corresponding to different objects, construct a large - scale structured database based on the generated paragraph set, and then use a large model to generate paragraph summaries. The document content processing module 801 is specifically designed to efficiently filter and process relevant documents uploaded by users. Specifically, the document content processing module 801 can accurately split relevant documents in PDF, DOCX, and TXT formats by paragraph, split relevant documents in PPTX format by page, extract text information in images using the OCR method, and integrate the text information into the original paragraphs, thereby forming a structured document paragraph set; this paragraph set is then sent to a large model to execute the paragraph summary generation service, and then converted into a paragraph summary set containing key information, providing a basis for subsequent data processing; using the generated paragraph set to construct a large - scale, structured database, that is, an information summary database, which not only contains rich information summaries corresponding to objects, but also provides the necessary data basis for the training of the subsequent document vectorization model.
[0137] The document content quantization module 802 is used to adopt an unsupervised learning method, through contrastive learning training, to obtain the embedding vector of the object description; perform vector generation on the paragraph summary and output a paragraph summary vector set; the document content quantization module 802 adopts unsupervised learning technology, through the data in the information summary database, trains a document vectorization model specific to the object domain, and the document vectorization model can effectively convert text information into a high - dimensional vector representation; using the pre - trained document vectorization model to extract features from the document paragraph set, the identification information of the object can be converted into a vector representation in a high - dimensional space.
[0138] The document content storage module 803 is used to store the paragraph summary vector, identification information, benchmark dimension, paragraph summary, original paragraph, paragraph sequence number of the original paragraph, and context paragraph information of the original paragraph into the vector database; for example, the document content storage module 803 stores the relevant information of relevant documents in the Milvus vector database, providing a basis for the efficient retrieval of the vector database.
[0139] The object index scoring module 804 is used to evaluate multiple benchmark dimensions based on the paragraph summary through a large model, obtain the object evaluation dimension score information, and form an object evaluation dimension score table; the large model uses the paragraph summary to evaluate the preset first-level dimensions, and each first-level dimension includes multiple second-level dimensions. The large model will score each second-level dimension according to the paragraph summary, and the scoring basis is the predefined scoring criteria and scoring indicators; specifically, the scoring process for each second-level dimension is as follows: determine the scoring criteria and scoring indicators for this second-level dimension; then analyze the relevant information in the paragraph summary set, extract the target data related to the scoring indicators, and according to the extracted target data, compare with the scoring criteria to score each second-level dimension. After completing the scoring of all second-level dimensions, add up the scores of all second-level dimensions under each first-level dimension to obtain the total score of this first-level dimension, and form a detailed object evaluation dimension score table.
[0140] The object comprehensive comparison module 805 includes intention recognition, semantic matching, and document generation. First, the object comprehensive comparison module 805 performs intention recognition on the user's question information, extracts the identification information and target comparison dimensions in the user's question information through semantic matching for conditional query; then combines the user's question information and the object evaluation dimension score table, matches with the data in the vector database, and outputs the top N matching results; finally, uses the matching results to generate a comparison document. Specifically, by classifying the intention of the user's question information, when the user wants to perform object comparison, the target comparison dimension and the identification information of the object to be compared are captured through semantic matching; if the user's question information contains the target comparison dimension, according to the identification information and the target comparison dimension, the target paragraph information corresponding to the object to be compared is obtained as the target prompt word information for the large model; if the user's question information does not contain the target comparison dimension, the document vectorization model is called to convert the user's question information into a vector representation, and efficiently matches with the data stored in the document content storage module 803, and finally returns the top N most relevant matching results; after determining the target prompt word information corresponding to the large model according to the top N matching results, the prompt (Prompt) template is spliced according to the target prompt word information, and the spliced prompt template is input into the large model, and a detailed and accurate document comparison result output by the large model can be obtained. The large model service module is reused in the document content processing module 801, the object index scoring module 804, and the object comprehensive comparison module 805, which can effectively improve the quality and efficiency of generating the document comparison result.
[0141] The device according to the embodiment of the present invention can be used to execute the technical solutions of any of the above method embodiments, and its implementation principle and technical effects are similar, which will not be elaborated here.
[0142] Figure 9 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. As Figure 9 shown, the electronic device 900 may include: at least one processor 901 and a memory 902.
[0143] The memory 902 is used to store programs. Specifically, the program may include program code, and the program code includes computer execution instructions.
[0144] The memory 902 may include a high-speed random access memory (Random Access Memory, RAM), and may also include a non-volatile memory, such as at least one disk memory.
[0145] The processor 901 is used to execute the computer-executable instructions stored in the memory 902 to implement the document alignment method described in the foregoing method embodiments. Among them, the processor 901 may be a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention. Specifically, when implementing the document alignment method described in the foregoing method embodiments, the electronic device may be, for example, a server.
[0146] Optionally, the electronic device 900 may further include a communication interface 903. In a specific implementation, if the communication interface 903, the memory 902, and the processor 901 are implemented independently, the communication interface 903, the memory 902, and the processor 901 may be interconnected through a bus and communicate with each other. The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc., but it does not mean that there is only one bus or one type of bus.
[0147] Optionally, in a specific implementation, if the communication interface 903, the memory 902, and the processor 901 are integrated on a chip, the communication interface 903, the memory 902, and the processor 901 may communicate through an internal interface.
[0148] The present invention also provides a computer-readable storage medium, in which computer program instructions are stored. When the processor executes the computer program instructions, the solution of the above document alignment method is implemented.
[0149] The present invention also provides a computer program product, including a computer program, which implements the solution of the above document alignment method when executed by the processor.
[0150] The above-mentioned computer-readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc. The readable storage medium can be any available medium accessible by a general-purpose or special-purpose computer.
[0151] An exemplary readable storage medium is coupled to the processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can be located in an application-specific integrated circuit. Of course, the processor and the readable storage medium can also exist as discrete components in the document alignment device.
[0152] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including those of the above method embodiments; and the aforementioned storage medium includes various media such as ROM, RAM, magnetic disks, or optical discs that can store program codes.
[0153] Finally, it should be noted that the above embodiments are only preferred embodiments given to fully illustrate the present invention, and the protection scope of the present invention is not limited thereto. Equivalent substitutions or transformations made by those skilled in the art on the basis of the present invention are all within the protection scope of the present invention.
Claims
1. A document benchmarking method, characterized in that: include: Acquire relevant information corresponding to a plurality of objects to be benchmarked, wherein the relevant information includes identification information respectively corresponding to the plurality of objects to be benchmarked; For each of the multiple objects to be benchmarked, if the relevant information includes a target benchmarking dimension, then according to the identification information and the target benchmarking dimension, the target paragraph information corresponding to the object to be benchmarked is obtained; if the relevant information does not include a target benchmarking dimension, then the relevant information is vectorized to obtain a target vector, and according to the target vector and multiple paragraph summary vectors, the target paragraph information corresponding to the object to be benchmarked is obtained; the target paragraph information includes a target paragraph summary, context paragraph information of an original paragraph corresponding to the target paragraph summary, and the benchmarking dimension to which the target paragraph summary belongs, and the multiple paragraph summary vectors, benchmarking dimension, paragraph summary, original paragraph, and context paragraph information are obtained by pre-processing relevant documents corresponding to different objects respectively; The identification information and the target paragraph information are input into a large model to obtain target benchmarking results corresponding to the multiple to-be-benchmarked objects output by the large model.
2. The document benchmarking method according to claim 1, characterized in that: The step of inputting the identification information and the target paragraph information into a large model to obtain target benchmarking results corresponding to the plurality of to-be-benchmarked objects output by the large model includes: Determine the target prompt word information corresponding to the large model according to the identification information, the object evaluation dimension score information and the target paragraph information, wherein the object evaluation dimension score information is obtained by the large model evaluating multiple benchmarking dimensions based on the paragraph summary; The target prompt word information is input into the large model to obtain the target matching results corresponding to the multiple objects to be matched output by the large model.
3. The document benchmarking method according to claim 2, characterized in that: The benchmarking dimension includes multiple primary dimensions and secondary dimensions contained in each primary dimension, and the object evaluation dimension score information is obtained in the following manner: For each secondary dimension included in the primary dimension, the target data related to the scoring index corresponding to the secondary dimension is extracted from the paragraph summary by the large model according to the scoring index corresponding to the secondary dimension; the secondary dimension is scored according to the target data and the scoring standard corresponding to the secondary dimension to obtain the score corresponding to the secondary dimension; The scores corresponding to the secondary dimensions are summed up through the large model to obtain the scores corresponding to the primary dimensions, and the object evaluation dimension score information includes the primary dimensions and the scores corresponding to the primary dimensions.
4. The document benchmarking method according to claim 3, characterized in that: The obtaining of relevant information corresponding to the plurality of objects to be benchmarked includes: In response to a first selection operation on the target benchmarking dimension and a second selection operation on the multiple objects to be benchmarked, or in response to a second selection operation on the multiple objects to be benchmarked, obtaining relevant information corresponding to the multiple objects to be benchmarked; Alternatively, user question information is obtained, and the intention corresponding to the user question information is identified. If the intention is to perform document matching, relevant information corresponding to multiple objects to be matched is obtained according to the user question information.
5. The document benchmarking method according to claim 3, characterized in that: The acquiring, according to the identification information and the target benchmarking dimension, target paragraph information corresponding to the object to be benchmarked includes: If the target benchmarking dimension is the primary dimension, acquiring target paragraph information corresponding to the to-be-benchmarked object according to the target benchmarking dimension, the secondary dimension included in the target benchmarking dimension, and the identification information; If the target benchmarking dimension is the secondary dimension, the target paragraph information corresponding to the object to be benchmarked is acquired according to the target benchmarking dimension, the primary dimension to which the target benchmarking dimension belongs, and the identification information.
6. The document benchmarking method according to claim 1, characterized in that: The step of acquiring target paragraph information corresponding to the object to be benchmarked according to the target vector and the plurality of paragraph summary vectors includes: Matching the target vector with each of the multiple paragraph summary vectors to obtain a similarity between the target vector and each of the paragraph summary vectors; Determine a preset number of target paragraph summary vectors with relatively high similarity among the plurality of paragraph summary vectors whose similarity is less than a similarity threshold; According to the target paragraph summary vector, target paragraph information corresponding to the object to be matched is obtained.
7. The document benchmarking method according to any one of claims 1 to 6, characterized in that: The preprocessing of the relevant documents corresponding to different objects includes: Splitting the document content of the relevant document according to paragraphs to obtain the original paragraphs, the paragraph sequence numbers of the original paragraphs, and the context paragraph information of the original paragraphs; Acquire the paragraph summary and the benchmarking dimension to which the paragraph summary belongs from the original paragraph by using the large model; Vectorizing the plurality of paragraph summaries using a document vectorization model to obtain the plurality of paragraph summary vectors, wherein the document vectorization model is obtained by unsupervised training based on the paragraph summaries; The identification information, benchmarking dimension, original paragraph, paragraph sequence number of the original paragraph, context paragraph information of the original paragraph, paragraph summary and paragraph summary vector corresponding to different objects are stored respectively.
8. The document benchmarking method according to claim 7, characterized in that: The document content further includes an image, and the document content of the related document is split into paragraphs to obtain the original paragraphs, including: Extracting text information from the image using an optical character recognition (OCR) method; The text information is merged with the paragraph corresponding to the text information to obtain the original paragraph.
9. The document benchmarking method according to claim 7, characterized in that: The vectorizing the relevant information to obtain a target vector includes: The relevant information is vectorized by the document vectorization model to obtain the target vector.
10. A document benchmarking device, characterized in that: include: An acquisition module is used to acquire relevant information corresponding to a plurality of objects to be benchmarked, wherein the relevant information includes identification information respectively corresponding to the plurality of objects to be benchmarked; A processing module is used for obtaining, for each of the multiple objects to be benchmarked, target paragraph information corresponding to the object to be benchmarked according to the identification information and the target benchmarking dimension if the relevant information includes a target benchmarking dimension; if the relevant information does not include a target benchmarking dimension, vectorizing the relevant information to obtain a target vector, and obtaining target paragraph information corresponding to the object to be benchmarked according to the target vector and multiple paragraph summary vectors; the target paragraph information includes a target paragraph summary, context paragraph information of an original paragraph corresponding to the target paragraph summary, and a benchmarking dimension to which the target paragraph summary belongs, and the multiple paragraph summary vectors, benchmarking dimension, paragraph summary, original paragraph, and context paragraph information are obtained by preprocessing relevant documents corresponding to different objects respectively; The output module is used to input the identification information and the target paragraph information into the large model to obtain the target benchmarking results corresponding to the multiple objects to be benchmarked output by the large model.
11. An electronic device, characterized in that: include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the document alignment method according to any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer program instructions, and when the computer program instructions are executed, the document alignment method according to any one of claims 1 to 9 is implemented.
13. A computer program product, comprising a computer program, characterized in that When the computer program is executed, the document alignment method according to any one of claims 1 to 9 is implemented.
Citation Information
Patent Citations
Conversational data analysis method and device, equipment and storage medium
CN118069820A
Document information extraction method and device based on question and answer model, equipment and medium
CN118133815A
User intention response method, device and equipment of intelligent robot and storage medium
CN118171658A
Document processing method, system and equipment and storage medium
CN118245574A
Automobile product benchmarking analysis method and device
CN118278665A