Document benchmarking methods, devices, equipment, storage media and program products
Through the big model benchmarking method, document information is obtained and vectorized, which solves the problem of low document benchmarking efficiency, realizes efficient and accurate document benchmarking, and improves the level of document management and benchmarking automation.
Patent Information
- Application Number
- CN202510571354.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2045-05-06
AI Technical Summary
In the prior art, document benchmarking efficiency is low and the results are not accurate enough. Especially in the process of vehicle design and development, manual comparison analysis cannot meet the needs of efficient and accurate document benchmarking.
By obtaining relevant information of multiple objects to be benchmarked, using the big model benchmarking method, if the relevant information contains the target benchmarking dimension, the target paragraph information is directly obtained. If it is not included, vectorization is performed, combining identification information and paragraph summary vectors to automatically output benchmarking results and optimize server resource utilization.
It improves the efficiency and accuracy of document benchmarking, improves the reliability and professionalism of document benchmarking, realizes the automation and intelligence of document processing, and provides a brand new document management and benchmarking solution.
Smart Images

Figure CN120087353B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information processing technology, and in particular to a document alignment method, apparatus, device, storage medium, and program product. Background Art
[0002] Benchmarking relevant documents across different entities can help identify differences and strengths, optimize market competition strategies, aid decision-making, and enhance transparency and trust. For example, vehicle model document benchmarking involves comparative analysis of technical documentation (such as design specifications, configuration parameters, and functional specifications) across different vehicle models during the vehicle design and development process to identify differences, optimize designs, or meet specific needs. This process helps understand the similarities and differences between different models, as well as their product strengths, and is crucial in the automotive industry's R&D, procurement, and market strategy development.
[0003] Currently, manual comparative analysis is often used to perform document comparison, but this method is inefficient. Therefore, an effective solution for document comparison is urgently needed. Summary of the Invention
[0004] The object of the present invention is to provide a document benchmarking method, apparatus, device, storage medium and program product to benchmark documents more efficiently and accurately.
[0005] In order to achieve the above object, the technical solution adopted by the present invention is as follows:
[0006] A document benchmarking method includes: obtaining relevant information corresponding to multiple objects to be benchmarked, the relevant information including identification information corresponding to the multiple objects to be benchmarked respectively; for each object to be benchmarked, if the relevant information includes a target benchmarking dimension, obtaining target paragraph information corresponding to the object to be benchmarked based on the identification information and the target benchmarking dimension; if the relevant information does not include the target benchmarking dimension, vectorizing the relevant information to obtain a target vector, and obtaining target paragraph information corresponding to the object to be benchmarked based on the target vector and multiple paragraph summary vectors; the target paragraph information includes a target paragraph summary, context paragraph information of an original paragraph corresponding to the target paragraph summary, and a benchmarking dimension to which the target paragraph summary belongs, the multiple paragraph summary vectors, benchmarking dimensions, paragraph summaries, original paragraphs, and context paragraph information being obtained by pre-processing relevant documents corresponding to different objects respectively; inputting the identification information and the target paragraph information into a large model to obtain target benchmarking results corresponding to the multiple objects to be benchmarked output by the large model.
[0007] According to the above technical means, for each of the multiple objects to be benchmarked, if the relevant information corresponding to the multiple objects to be benchmarked contains the target benchmarking dimension, the target paragraph information corresponding to the object to be benchmarked is obtained based on the identification information and the target benchmarking dimension; if the relevant information does not contain the target benchmarking dimension, the relevant information is vectorized to obtain the target vector, and the target paragraph information corresponding to the object to be benchmarked is obtained based on the target vector and the multiple paragraph summary vectors; the above method of obtaining the target paragraph information can effectively filter out irrelevant data items, thereby reducing the context length required by the large model and optimizing the server's video memory resource utilization; thereby inputting the identification information and the target paragraph information into the large model to obtain the target benchmarking results corresponding to the multiple objects to be benchmarked output by the large model, and realizing the automatic output of the target benchmarking results corresponding to the multiple objects to be benchmarked through the large model, which can greatly improve the efficiency of document benchmarking, and can more accurately obtain the target benchmarking results corresponding to the multiple objects to be benchmarked, thereby improving the reliability and professionalism of document benchmarking.
[0008] Furthermore, the identification information and target paragraph information are input into the big model to obtain target benchmarking results corresponding to multiple objects to be benchmarked output by the big model, including: determining the target prompt word information corresponding to the big model based on the identification information, the object evaluation dimension score information and the target paragraph information, the object evaluation dimension score information is obtained by the big model evaluating multiple benchmarking dimensions based on the paragraph summary; the target prompt word information is input into the big model to obtain target benchmarking results corresponding to multiple objects to be benchmarked output by the big model.
[0009] Furthermore, the benchmarking dimension includes multiple first-level dimensions and second-level dimensions contained in each first-level dimension. The object evaluation dimension score information is obtained in the following manner: for each second-level dimension contained in the first-level dimension, the large model is used to extract target data related to the scoring index from the paragraph summary according to the scoring index corresponding to the second-level dimension; the second-level dimension is scored according to the target data and the scoring criteria corresponding to the second-level dimension to obtain the score corresponding to the second-level dimension; the score corresponding to the second-level dimension is added up by the large model to obtain the score corresponding to the first-level dimension. The object evaluation dimension score information includes the first-level dimension and the score corresponding to the first-level dimension.
[0010] Furthermore, relevant information corresponding to multiple objects to be benchmarked is obtained, including: in response to a first selection operation for a target benchmarking dimension and a second selection operation for multiple objects to be benchmarked, or in response to a second selection operation for multiple objects to be benchmarked, obtaining relevant information corresponding to multiple objects to be benchmarked; or, obtaining user question information, identifying the intention corresponding to the user question information, and if the intention is to perform document benchmarking, obtaining relevant information corresponding to multiple objects to be benchmarked based on the user question information.
[0011] Furthermore, based on the identification information and the target benchmarking dimension, the target paragraph information corresponding to the object to be benchmarked is obtained, including: if the target benchmarking dimension is a first-level dimension, then based on the target benchmarking dimension, the second-level dimension contained in the target benchmarking dimension and the identification information, the target paragraph information corresponding to the object to be benchmarked is obtained; if the target benchmarking dimension is a second-level dimension, then based on the target benchmarking dimension, the first-level dimension to which the target benchmarking dimension belongs and the identification information, the target paragraph information corresponding to the object to be benchmarked is obtained.
[0012] Furthermore, target paragraph information corresponding to the object to be benchmarked is obtained based on the target vector and the multiple paragraph summary vectors, including: matching the target vector with each paragraph summary vector in the multiple paragraph summary vectors to obtain the similarity between the target vector and each paragraph summary vector; determining a first preset number of target paragraph summary vectors with higher similarity among the multiple paragraph summary vectors whose similarity is less than a similarity threshold; and obtaining the target paragraph information corresponding to the object to be benchmarked based on the target paragraph summary vector.
[0013] Furthermore, the relevant documents corresponding to different objects are preprocessed, including: splitting the document content of the relevant documents according to paragraphs to obtain the original paragraphs, the paragraph sequence numbers of the original paragraphs, and the contextual paragraph information of the original paragraphs; obtaining the paragraph summary and the benchmarking dimension to which the paragraph summary belongs from the original paragraph through the large model; vectorizing multiple paragraph summaries through the document vectorization model to obtain multiple paragraph summary vectors, and the document vectorization model is obtained by unsupervised training based on the paragraph summary; storing the identification information, benchmarking dimensions, original paragraphs, the paragraph sequence numbers of the original paragraphs, the contextual paragraph information of the original paragraphs, paragraph summaries, and paragraph summary vectors corresponding to different objects.
[0014] Furthermore, the document content also includes images, and the document content of the relevant document is split into paragraphs to obtain original paragraphs, including: extracting text information from the image using an optical character recognition (OCR) method; and merging the text information with the paragraph corresponding to the text information to obtain the original paragraph.
[0015] Furthermore, the relevant information is vectorized to obtain a target vector, including: vectorizing the relevant information through a document vectorization model to obtain a target vector.
[0016] A document benchmarking device, comprising:
[0017] An acquisition module, configured to acquire relevant information corresponding to a plurality of objects to be benchmarked, wherein the relevant information includes identification information corresponding to the plurality of objects to be benchmarked;
[0018] a processing module configured to, for each of the multiple objects to be benchmarked, obtain target paragraph information corresponding to the object to be benchmarked based on the identification information and the target benchmarking dimension if the relevant information includes a target benchmarking dimension; and, if the relevant information does not include a target benchmarking dimension, perform vectorization processing on the relevant information to obtain a target vector, and obtain target paragraph information corresponding to the object to be benchmarked based on the target vector and multiple paragraph summary vectors; the target paragraph information includes a target paragraph summary, contextual paragraph information of an original paragraph corresponding to the target paragraph summary, and the benchmarking dimension to which the target paragraph summary belongs. The multiple paragraph summary vectors, benchmarking dimension, paragraph summary, original paragraph, and contextual paragraph information are obtained by preprocessing relevant documents corresponding to different objects.
[0019] The output module is used to input the identification information and target paragraph information into the large model to obtain the target benchmarking results corresponding to multiple objects to be benchmarked output by the large model.
[0020] Furthermore, the output module is specifically used to: determine the target prompt word information corresponding to the large model based on the identification information, object evaluation dimension score information and target paragraph information, the object evaluation dimension score information is obtained by the large model by evaluating multiple benchmarking dimensions based on the paragraph summary; input the target prompt word information into the large model to obtain the target benchmarking results corresponding to multiple objects to be benchmarked output by the large model.
[0021] Furthermore, the benchmarking dimension includes multiple first-level dimensions and second-level dimensions included in each first-level dimension. The acquisition module is also used to obtain the object evaluation dimension score information in the following manner: for each second-level dimension included in the first-level dimension, the large model is used to extract the target data related to the scoring index from the paragraph summary according to the scoring index corresponding to the second-level dimension; the second-level dimension is scored according to the target data and the scoring criteria corresponding to the second-level dimension to obtain the score corresponding to the second-level dimension; the score corresponding to the second-level dimension is added up by the large model to obtain the score corresponding to the first-level dimension. The object evaluation dimension score information includes the first-level dimension and the score corresponding to the first-level dimension.
[0022] Furthermore, the acquisition module is specifically used to: obtain relevant information corresponding to multiple objects to be benchmarked in response to a first selection operation for the target benchmarking dimension and a second selection operation for multiple objects to be benchmarked, or in response to a second selection operation for multiple objects to be benchmarked; or, obtain user question information, identify the intention corresponding to the user question information, and if the intention is to perform document benchmarking, obtain relevant information corresponding to multiple objects to be benchmarked based on the user question information.
[0023] Furthermore, when the processing module is used to obtain the target paragraph information corresponding to the object to be benchmarked based on the identification information and the target benchmarking dimension, it is specifically used to: if the target benchmarking dimension is a first-level dimension, then according to the target benchmarking dimension, the second-level dimension contained in the target benchmarking dimension and the identification information, obtain the target paragraph information corresponding to the object to be benchmarked; if the target benchmarking dimension is a second-level dimension, then according to the target benchmarking dimension, the first-level dimension to which the target benchmarking dimension belongs and the identification information, obtain the target paragraph information corresponding to the object to be benchmarked.
[0024] Furthermore, when the processing module is used to obtain target paragraph information corresponding to the object to be benchmarked based on the target vector and multiple paragraph summary vectors, it is specifically used to: match the target vector with each paragraph summary vector in the multiple paragraph summary vectors to obtain the similarity between the target vector and each paragraph summary vector; determine the first preset number of target paragraph summary vectors with higher similarity among the multiple paragraph summary vectors whose similarity is less than a similarity threshold; and obtain the target paragraph information corresponding to the object to be benchmarked based on the target paragraph summary vector.
[0025] Furthermore, the document benchmarking device also includes a preprocessing module, which is used to: split the document content of the relevant documents according to paragraphs to obtain the original paragraphs, the paragraph sequence numbers of the original paragraphs and the contextual paragraph information of the original paragraphs; obtain the paragraph summary and the benchmarking dimension to which the paragraph summary belongs from the original paragraph through a large model; vectorize multiple paragraph summaries through a document vectorization model to obtain multiple paragraph summary vectors, and the document vectorization model is obtained by unsupervised training based on the paragraph summary; store the identification information, benchmarking dimensions, original paragraphs, the paragraph sequence numbers of the original paragraphs, the contextual paragraph information of the original paragraphs, paragraph summaries and paragraph summary vectors corresponding to different objects.
[0026] Furthermore, the document content also includes images. When the preprocessing module is used to split the document content of the relevant document into paragraphs to obtain the original paragraphs, it is specifically used to: extract text information from the image using the optical character recognition (OCR) method; merge the text information with the paragraph corresponding to the text information to obtain the original paragraph.
[0027] Furthermore, when the processing module is used to perform vectorization processing on the relevant information to obtain the target vector, it is specifically used to: perform vectorization processing on the relevant information through a document vectorization model to obtain the target vector.
[0028] An electronic device includes: a processor and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the document alignment method as described above in the present invention.
[0029] A computer-readable storage medium stores computer program instructions, which, when executed, implement the document alignment method described above in the present invention.
[0030] A computer program product includes a computer program, and when the computer program is executed, it implements the document matching method as described above in the present invention.
[0031] Beneficial effects of the present invention: The document benchmarking method, apparatus, device, storage medium and program product provided by the present invention obtain relevant information corresponding to multiple objects to be benchmarked, and the relevant information includes identification information corresponding to multiple objects to be benchmarked respectively; for each object to be benchmarked among the multiple objects to be benchmarked, if the relevant information contains a target benchmarking dimension, the target paragraph information corresponding to the object to be benchmarked is obtained according to the identification information and the target benchmarking dimension; if the relevant information does not contain the target benchmarking dimension, the relevant information is vectorized to obtain a target vector, and the target paragraph information corresponding to the object to be benchmarked is obtained according to the target vector and multiple paragraph summary vectors; the above-mentioned method of obtaining target paragraph information can effectively filter out irrelevant data items, thereby reducing the upper limit required for large models. The length of the following text is used to optimize the utilization rate of the server's video memory resources; the target paragraph information includes the target paragraph summary, the context paragraph information of the original paragraph corresponding to the target paragraph summary, and the benchmarking dimension to which the target paragraph summary belongs. Multiple paragraph summary vectors, benchmarking dimensions, paragraph summaries, original paragraphs, and context paragraph information are obtained by pre-processing the relevant documents corresponding to different objects; the identification information and the target paragraph information are input into the large model to obtain the target benchmarking results corresponding to the multiple objects to be benchmarked output by the large model, and the target benchmarking results corresponding to the multiple objects to be benchmarked are automatically output by the large model, which can greatly improve the efficiency of document benchmarking, and can more accurately obtain the target benchmarking results corresponding to the multiple objects to be benchmarked, thereby improving the reliability and professionalism of document benchmarking. The present invention can not only improve the automation and intelligence level of document processing, but also provide a new document management and benchmarking solution for different industries. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0033] Figure 1 A schematic diagram of an application scenario provided by an embodiment of the present invention;
[0034] Figure 2A flowchart of a document benchmarking method provided by one embodiment of the present invention;
[0035] Figure 3 A flow chart of a method for preprocessing related documents corresponding to different objects provided by one embodiment of the present invention;
[0036] Figure 4 A flowchart of a document benchmarking method provided in another embodiment of the present invention;
[0037] Figure 5 A schematic diagram of a prompt word template provided by an embodiment of the present invention;
[0038] Figure 6 A flowchart of a method for obtaining object evaluation dimension score information provided by one embodiment of the present invention;
[0039] Figure 7 A schematic diagram of the structure of a document alignment device provided by an embodiment of the present invention;
[0040] Figure 8 A schematic structural diagram of a document alignment device provided in another embodiment of the present invention;
[0041] Figure 9 A schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0042] The following describes the embodiments of the present invention with reference to the accompanying drawings and preferred embodiments. Those skilled in the art will readily appreciate the other advantages and benefits of the present invention from the disclosure herein. The present invention may also be implemented or applied through various other specific embodiments, and the various details in this specification may be modified or altered based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred embodiments are intended only to illustrate the present invention and are not intended to limit the scope of protection of the present invention.
[0043] It should be noted that the illustrations provided in the following embodiments are merely schematic illustrations of the basic concept of the present invention. Therefore, the illustrations only show components related to the present invention and are not drawn according to the number, shape, and size of components in actual implementation. In actual implementation, the type, quantity, and proportion of each component may be changed arbitrarily, and the component layout may also be more complex.
[0044] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the present invention are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards, and corresponding operation entrances must be provided for users to choose to authorize or refuse.
[0045] Benchmarking relevant documents across different entities can help identify differences and strengths, optimize market competition strategies, aid decision-making, and enhance transparency and trust. For example, vehicle model document benchmarking involves comparative analysis of technical documentation (such as design specifications, configuration parameters, and functional specifications) across different vehicle models during the vehicle design and development process to identify differences, optimize designs, or meet specific needs. This process helps understand the similarities and differences between different models, as well as their product strengths, and is crucial in the automotive industry's R&D, procurement, and market strategy development.
[0046] Currently, manual comparison and analysis are often used to compare documents, but this method is inefficient and the results are not accurate enough. Therefore, an effective solution for document comparison is urgently needed.
[0047] Furthermore, with the increasing application of large models, document processing can be performed based on these models. For example, documents can be parsed, knowledge bases can be constructed based on the knowledge gained from the parsing, and this knowledge base can be fed into the large model for question-and-answering. However, this often involves using a common basic quantitative model for quantitative document storage, without training a domain-specific vectorization model based on application domain data. This poses challenges for semantic similarity recall in large models, ultimately impacting the accuracy of their output. However, there is currently no solution for benchmarking documents using large models.
[0048] Based on the above problems, the present invention provides a document benchmarking method, which preprocesses the document contents of relevant documents corresponding to different objects to obtain preprocessed paragraph information corresponding to different objects, thereby realizing efficient document management and facilitating rapid retrieval and matching; obtaining relevant information corresponding to multiple objects to be benchmarked, and the relevant information includes identification information corresponding to multiple objects to be benchmarked; if the relevant information contains a target benchmarking dimension, then according to the identification information and the target benchmarking dimension, the target paragraph information corresponding to the object to be benchmarked is directly obtained; if the relevant information does not contain the target benchmarking dimension, the relevant information is vectorized to obtain a target vector, and according to the target vector and multiple paragraph summary vectors, the target paragraph information corresponding to the object to be benchmarked is obtained; based on the preprocessed paragraph information, target paragraph information and identification information, a large model is used to output target benchmarking results corresponding to multiple objects to be benchmarked, which can greatly improve the efficiency of document benchmarking, and can more accurately obtain target benchmarking results corresponding to multiple objects to be benchmarked, providing a new document management and benchmarking solution for the vehicle industry.
[0049] Hereinafter, the application scenarios of the solution provided by the present invention are first described with examples.
[0050] Figure 1 This is a schematic diagram of an application scenario provided by an embodiment of the present invention. Figure 1 As shown, this application scenario may include: a server cluster 11 and a terminal 12. The server cluster 11 includes multiple servers 111 and a storage device 112. The terminal 12 may be a mobile phone, tablet computer, laptop computer, desktop computer, smart home appliance, or the like. Taking vehicle model document benchmarking as an example, the target vehicle model is the vehicle model to be benchmarked. A user enters target benchmarking dimensions and identification information (i.e., vehicle model information) corresponding to multiple vehicle models through terminal 12. For example, the user enters "Help me compare the smart cockpit dimensions of vehicle models A and B" through terminal 12, where the smart cockpit dimension is the target benchmarking dimension. Accordingly, server 111 obtains the target benchmarking dimensions and vehicle model information corresponding to the multiple vehicle models entered by the user through terminal 12. Based on the document benchmarking method provided by an embodiment of the present invention, server 111 obtains target benchmarking results corresponding to the multiple vehicle models to be benchmarked, such as outputting the target benchmarking results for vehicle models A and B. Server 111 sends the target benchmarking results to terminal 12, which then displays the results to the user through terminal 12. The server 111 obtains relevant data from the memory 112 and stores the generated data in the memory 112. In addition, the server 111 and the terminal 12 communicate via a wireless network or a wired network.
[0051] It should be noted that Figure 1 This is only a schematic diagram of an application scenario provided by an embodiment of the present invention. Figure 1The equipment included in the Figure 1 The positional relationship between the devices is limited.
[0052] The document benchmarking method provided by the embodiment of the present invention can be applied to at least: (1) Industry benchmarking: used to analyze the relevant competitiveness of benchmarking companies in the industry to which it belongs, which helps to cultivate its own relevant competitiveness; (2) Patent benchmarking: by analyzing the patent applications and authorization status of competitors or in the industry, to understand the technology development trend, identify potential technology cooperation opportunities, and avoid possible intellectual property risks; (3) Product benchmarking: used to analyze the product specifications of competitors or industry leaders to understand the latest trends and technological advances in the market; (4) Technology benchmarking: used to analyze the performance of competitors or leaders in core technology, R&D investment, innovation methods, and participation in technology standards; pay attention to their underlying technical capabilities, R&D system efficiency and future technology roadmap; help to judge the advancement and sustainability of their own technology, and find opportunities for technological breakthroughs or cooperation; (5) Customer benchmarking: used to analyze the customer group characteristics, customer satisfaction, customer loyalty and customer feedback of competitors; understand customers' evaluation of competitors' products and services, as well as their unmet needs, which helps to optimize their own product design, service experience and customer relationship management.
[0053] The document benchmarking results obtained by the document benchmarking method provided by the embodiment of the present invention can be applied to recommendation scenarios. For example, taking vehicle model document benchmarking as an example, assuming that the models to be benchmarked are Model A and Model B, and the target benchmarking dimension is the smart cockpit dimension, the document benchmarking method provided by the embodiment of the present invention can obtain the target benchmarking results corresponding to Model A and Model B. Based on the target benchmarking results, it can be determined that Model A is superior to Model B in the smart cockpit dimension, so that Model A can be recommended to the user.
[0054] The technical solution of the present invention is described in detail below through specific embodiments. It should be noted that the following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.
[0055] Figure 2 This is a flowchart of a document alignment method provided by an embodiment of the present invention. The document alignment method can be executed by software and / or hardware devices. For example, the hardware device can be a document alignment device, and the document alignment device can be an electronic device or a processing chip in an electronic device. Figure 2 As shown, the method of the embodiment of the present invention includes:
[0056] S201: Obtain relevant information corresponding to a plurality of objects to be benchmarked, where the relevant information includes identification information corresponding to the plurality of objects to be benchmarked.
[0057] In an embodiment of the present invention, the relevant information corresponding to the multiple objects to be benchmarked includes identification information corresponding to the multiple objects to be benchmarked, and may also include target benchmarking dimensions to be benchmarked. It can be understood that the benchmarking dimensions can include multiple first-level dimensions, each first-level dimension can include multiple second-level dimensions, and the benchmarking dimensions can be defined as needed, which is not limited by the embodiment of the present invention. For example, taking the vehicle model document benchmarking as an example, the object to be benchmarked is the vehicle model to be benchmarked, and the first-level dimensions can include, for example, the smart cockpit dimension, the smart vehicle control dimension, the smart parking dimension, and the smart driving dimension. Each first-level dimension can contain multiple second-level dimensions. Specifically, the smart cockpit dimension can include multiple second-level dimensions such as intelligent voice, navigation system, multimedia, imaging, settings, Bluetooth phone, system stability, system fluency, interface aesthetics, display quality, sound quality and extended functions; the smart vehicle control dimension can include multiple second-level dimensions such as physical keys, welcoming guests, lighting control, seat control, mobile phone remote control, digital keys and expanded functions; the smart parking dimension can include multiple second-level dimensions such as vision system and automatic parking; the smart driving dimension can include multiple second-level dimensions such as cruise function, lane assistance, traffic sign recognition and emergency braking.
[0058] Optionally, obtaining relevant information corresponding to multiple objects to be benchmarked may include: obtaining relevant information corresponding to multiple objects to be benchmarked in response to a first selection operation for a target benchmarking dimension and a second selection operation for multiple objects to be benchmarked, or in response to a second selection operation for multiple objects to be benchmarked; or, obtaining user question information, identifying the intention corresponding to the user question information, and if the intention is to perform document benchmarking, obtaining relevant information corresponding to multiple objects to be benchmarked based on the user question information.
[0059] It will be appreciated that embodiments of the present invention provide two methods for obtaining document benchmarking results: one is to directly select the identification information of the target object for document benchmarking, and the other is to conduct document benchmarking through natural language communication. Taking vehicle model document benchmarking as an example, the target object is a vehicle model. In one example, a user can select a target benchmarking dimension and multiple vehicle models through a terminal. The multiple vehicle models are two or more vehicle models to be benchmarked, for example, Model A and Model B. Accordingly, an electronic device executing this embodiment of the method, in response to a first selection operation for the target benchmarking dimension and a second selection operation for the multiple vehicle models to be benchmarked, obtains vehicle model information corresponding to the target benchmarking dimension and the multiple vehicle models to be benchmarked, such as the vehicle names corresponding to Model A and Model B. In another example, a user can select multiple vehicle models to be benchmarked through a terminal without selecting a target benchmarking dimension. Accordingly, the electronic device executing this embodiment of the method, in response to the second selection operation for the multiple vehicle models to be benchmarked, obtains vehicle model information corresponding to the multiple vehicle models to be benchmarked. Since the user did not select a target benchmarking dimension, the target benchmarking dimension is assumed to be all primary dimensions. In another example, the user can input the question "Help me compare the smart cockpit dimensions of model A and model B" through the terminal, where the smart cockpit dimension is the target benchmarking dimension. Accordingly, the electronic device executing the embodiment of this method can obtain the user question information, identify the intention corresponding to the user question information, and if the intention is to perform vehicle model document benchmarking, then the user question information is subjected to word slot extraction to obtain the target benchmarking dimension and the vehicle model information corresponding to the multiple models to be benchmarked. In particular, when identifying the intention corresponding to the user question information, the Rasa framework (a machine learning framework for building conversational artificial intelligence applications) can be used to build an intent recognition module, and the intent corresponding to the user question information can be identified through the intent recognition module. For details, please refer to the subsequent embodiments. The user question information may also not contain the target benchmarking dimension, but contain other relevant information about the vehicle model. In this case, the entire user question information can be used as relevant information corresponding to the multiple models to be benchmarked.
[0060] S202. For each of the multiple objects to be benchmarked, if the relevant information includes a target benchmarking dimension, the target paragraph information corresponding to the object to be benchmarked is obtained based on the identification information and the target benchmarking dimension; if the relevant information does not include the target benchmarking dimension, the relevant information is vectorized to obtain a target vector, and the target paragraph information corresponding to the object to be benchmarked is obtained based on the target vector and multiple paragraph summary vectors; the target paragraph information includes the target paragraph summary, the context paragraph information of the original paragraph corresponding to the target paragraph summary, and the benchmarking dimension to which the target paragraph summary belongs. The multiple paragraph summary vectors, benchmarking dimension, paragraph summary, original paragraph, and context paragraph information are obtained by pre-processing the relevant documents corresponding to different objects.
[0061] In this step, multiple paragraph summary vectors, benchmarking dimensions, paragraph summaries, original paragraphs, and contextual paragraph information are obtained by pre-processing the relevant documents corresponding to different objects. For specific pre-processing of relevant documents corresponding to different objects, please refer to the subsequent embodiments and will not be repeated here. After pre-processing the relevant documents corresponding to different objects, for example, the identification information, original paragraphs, contextual paragraph information of the original paragraphs, paragraph summaries, benchmarking dimensions to which the paragraph summaries belong, and paragraph summary vectors corresponding to different objects can be stored in a vector database for use in this step to obtain the target paragraph information corresponding to the object to be benchmarked. Among them, the benchmarking dimension to which the paragraph summary belongs is the first-level dimension to which the target paragraph summary belongs and all the second-level dimensions contained under the first-level dimension. The target benchmarking dimension is obtained based on the relevant information corresponding to multiple objects to be benchmarked. For example, after obtaining the user question information input by the user, it can be determined whether the user question information contains the target benchmarking dimension. The target benchmarking dimension is the dimension that the user wants to use for document benchmarking, which can be a first-level dimension, a second-level dimension, or a first-level dimension and a second-level dimension.
[0062] In the case where the relevant information includes a target benchmarking dimension, optionally, obtaining target paragraph information corresponding to the object to be benchmarked based on the identification information and the target benchmarking dimension may include: if the target benchmarking dimension is a first-level dimension, obtaining the target paragraph information corresponding to the object to be benchmarked based on the target benchmarking dimension, the second-level dimension included in the target benchmarking dimension, and the identification information; if the target benchmarking dimension is a second-level dimension, obtaining the target paragraph information corresponding to the object to be benchmarked based on the target benchmarking dimension, the first-level dimension to which the target benchmarking dimension belongs, and the identification information.
[0063] For example, after obtaining the identification information corresponding to the target benchmarking dimension and multiple objects to be benchmarked, if the target benchmarking dimension is a primary dimension, all secondary dimensions contained in the target benchmarking dimension can be obtained. Then, based on the target benchmarking dimension, all secondary dimensions contained in the target benchmarking dimension, and the identification information, the vector database can be directly queried to obtain the target paragraph information corresponding to the objects to be benchmarked. If the target benchmarking dimension is a secondary dimension, the primary dimension to which the target benchmarking dimension belongs can be obtained. Then, based on the target benchmarking dimension, the primary dimension to which the target benchmarking dimension belongs, and the identification information, the vector database can be directly queried to obtain the target paragraph information corresponding to the objects to be benchmarked.
[0064] If the relevant information doesn't include the target benchmarking dimension, for example, if the relevant information corresponding to multiple benchmarking objects is user question information, the relevant information can be directly vectorized to obtain a target vector. The target vector is then matched with each of the multiple paragraph summary vectors to obtain the similarity between the target vector and each paragraph summary vector. Based on this similarity, the target paragraph information corresponding to the benchmarking object can be obtained.
[0065] Obtaining the target paragraph information based on the above method can effectively filter out irrelevant data items, thereby reducing the context length required for the large model and optimizing the memory resource utilization of the electronic device executing the embodiment of this method. For specific information on how to obtain the target paragraph information corresponding to the object to be benchmarked based on the target vector and multiple paragraph summary vectors, please refer to the subsequent embodiments. Taking the vehicle model document benchmarking as an example, the object to be benchmarked is the vehicle model to be benchmarked. Assuming that the vehicle models to be benchmarked are model A and model B, the target paragraph information corresponding to model A and the target paragraph information corresponding to model B can be obtained respectively. Among them, the target paragraph information includes the target paragraph summary, the context paragraph information of the original paragraph corresponding to the target paragraph summary, and the benchmarking dimension to which the target paragraph summary belongs.
[0066] S203: Input the identification information and the target paragraph information into the large model to obtain target alignment results corresponding to multiple to-be-aligned objects output by the large model.
[0067] For example, the specific large model in the embodiments of the present invention can be a currently mainstream large model, for example, one containing at least 7 billion (7B) parameters. After obtaining target paragraph information corresponding to the target object to be benchmarked, the identification information and target paragraph information corresponding to multiple target objects can be input into the large model, and the large model outputs target benchmarking results corresponding to the multiple target objects to be benchmarked. For details on how to obtain target benchmarking results using the large model, please refer to the subsequent embodiments.
[0068] The document benchmarking method provided by the embodiment of the present invention obtains relevant information corresponding to multiple objects to be benchmarked, and the relevant information includes identification information corresponding to the multiple objects to be benchmarked respectively; for each object to be benchmarked among the multiple objects to be benchmarked, if the relevant information contains a target benchmarking dimension, the target paragraph information corresponding to the object to be benchmarked is obtained according to the identification information and the target benchmarking dimension; if the relevant information does not contain the target benchmarking dimension, the relevant information is vectorized to obtain a target vector; according to the target vector and multiple paragraph summary vectors, the target paragraph information corresponding to the object to be benchmarked is obtained; the above-mentioned method of obtaining the target paragraph information can effectively filter out irrelevant data items, thereby reducing the context length required for the large model and optimizing the server. Memory resource utilization; target paragraph information includes a target paragraph summary, context paragraph information of the original paragraph corresponding to the target paragraph summary, and the benchmarking dimension to which the target paragraph summary belongs. Multiple paragraph summary vectors, benchmarking dimensions, paragraph summaries, original paragraphs, and context paragraph information are obtained by preprocessing relevant documents corresponding to different objects; inputting the identification information and the target paragraph information into the large model to obtain target benchmarking results corresponding to multiple objects to be benchmarked output by the large model, and realizing automatic output of target benchmarking results corresponding to multiple objects to be benchmarked through the large model, which can greatly improve the efficiency of document benchmarking, and can more accurately obtain target benchmarking results corresponding to multiple objects to be benchmarked, thereby improving the reliability and professionalism of document benchmarking. The embodiments of the present invention can not only improve the automation and intelligence level of document processing, but also provide a new document management and benchmarking solution for different industries.
[0069] Based on the above embodiments, Figure 3 A flowchart of a method for preprocessing related documents corresponding to different objects provided by an embodiment of the present invention. Figure 3 As shown, the method of the embodiment of the present invention may include:
[0070] S301 : Split the document contents of the relevant documents corresponding to different objects into paragraphs to obtain original paragraphs, paragraph sequence numbers of the original paragraphs, and context paragraph information of the original paragraphs.
[0071] For example, taking the vehicle model document matching as an example, different objects are different vehicle models. Taking the vehicle model A as an example, assuming that the document content of the vehicle model-related document corresponding to vehicle model A contains 3 paragraphs, the document content of the vehicle model-related document is split according to paragraphs, and 3 original paragraphs, the paragraph sequence numbers corresponding to the 3 original paragraphs, and the context paragraph information of the original paragraphs are obtained. The context paragraph information of the original paragraph, for example, includes the previous original paragraph and the next original paragraph adjacent to the current original paragraph.
[0072] Optionally, the document content also includes an image. Splitting the document content of the relevant document into paragraphs to obtain original paragraphs may include: extracting text information from the image using an optical character recognition (OCR) method; and merging the text information with the paragraph corresponding to the text information to obtain the original paragraph.
[0073] For example, when the document content also includes an image, the OCR method can be used to extract the text information in the image, and merge the text information with the paragraph corresponding to the text information, such as inserting the text information into the paragraph corresponding to the text information, thereby obtaining the final original paragraph.
[0074] It can be understood that through this step, the document content of the related documents corresponding to different objects is divided into paragraphs, and finally a paragraph set can be formed. The document formats of the related documents can include, for example, PDF, DOCX, PPTX, and TXT. The embodiment of the present invention can accurately read and divide the document content of related documents in various formats, and use the OCR method to extract text information from the image, which can effectively improve the integrity of the document content.
[0075] S302: Obtain a paragraph summary and the benchmarking dimension to which the paragraph summary belongs from the original paragraph through the large model.
[0076] For example, after obtaining a set of paragraphs corresponding to different objects, the big model can be used to extract key information from the original paragraphs, resulting in a paragraph summary and the corresponding benchmarking dimension, ultimately forming a set of paragraph summaries. Extracting the paragraph summary and the corresponding benchmarking dimension from the paragraph collection using the big model's prompts can further improve document vectorization accuracy. For example, taking vehicle model document benchmarking as an example, if the user inputs "Help me compare the intelligent voice control of vehicle models A and B," the target benchmarking dimension can be determined to be "intelligent voice." The specific process for obtaining this target benchmarking dimension is as follows: The original paragraphs are: Voice: supports recognition in five voice zones; Voice control capabilities: supports see-and-speak, cross-voice inheritance, and multiple intents per sentence, and incorporates big model dialogue. It also supports quick vehicle control adjustments, such as adjusting the rearview mirror angle and steering wheel height with voice commands, and supports "pause" interruption during adjustments; and Voice communication between mobile phone and vehicle: supports calling mobile apps from the vehicle's voice control. Based on the original paragraph, the big model can be used to obtain the benchmark dimension to which the paragraph summary belongs: smart cockpit#smart voice, where smart cockpit is the first-level dimension and smart voice is the second-level dimension.
[0077] S303 , vectorizing the multiple paragraph summaries using a document vectorization model to obtain multiple paragraph summary vectors, where the document vectorization model is obtained by performing unsupervised training based on the paragraph summaries.
[0078] In this step, a document summary database can be constructed using a collection of paragraph summaries, and a document vectorization model can be unsupervisedly trained based on the paragraph summaries in the document summary database, thereby obtaining a trained document vectorization model that is more suitable for document matching. The document vectorization model can accurately convert the document content of related documents corresponding to different objects into vector representations in a high-dimensional space, thereby facilitating subsequent query and analysis. Exemplarily, multiple paragraph summaries can be vectorized by the document vectorization model to obtain multiple paragraph summary vectors.
[0079] S304: Store identification information, alignment dimensions, original paragraphs, paragraph sequence numbers of original paragraphs, contextual paragraph information of original paragraphs, paragraph summaries, and paragraph summary vectors corresponding to different objects.
[0080] For example, the identification information, alignment dimensions, original paragraphs, paragraph sequence numbers of the original paragraphs, contextual paragraph information of the original paragraphs, paragraph summaries, and paragraph summary vectors corresponding to different objects can be stored in batches in a vector database to improve the efficiency and accuracy of matching the target vector and multiple paragraph summary vectors in the above embodiment. The vector database is, for example, the Milvus vector database.
[0081] In an embodiment of the present invention, the document contents of the relevant documents corresponding to different objects are split according to paragraphs to obtain the original paragraphs, the paragraph sequence numbers of the original paragraphs and the contextual paragraph information of the original paragraphs. The paragraph summary and the benchmarking dimension to which the paragraph summary belongs are obtained from the original paragraphs through a large model, and the paragraph summary and the benchmarking dimension to which the paragraph summary belongs can be accurately extracted; multiple paragraph summaries are vectorized by a document vectorization model to obtain multiple paragraph summary vectors; wherein, the document vectorization model is obtained by unsupervised training based on the paragraph summary, which can reduce the cost of manual labeling and improve the versatility of model training, so that the document content can be accurately converted into a vector representation, making subsequent vector matching and query more efficient and accurate; the identification information, benchmarking dimensions, original paragraphs, the paragraph sequence numbers of the original paragraphs, the contextual paragraph information of the original paragraphs, paragraph summaries and paragraph summary vectors corresponding to different objects are stored to achieve efficient document management and facilitate rapid retrieval and matching.
[0082] Figure 4 This is a flowchart of a document matching method provided by another embodiment of the present invention. Based on the above embodiment, the embodiment of the present invention further illustrates the document matching method, wherein the document matching method through natural language communication is used as an example. Figure 4 As shown, the method of the embodiment of the present invention may include:
[0083] S401. Obtain user question information and identify the intention corresponding to the user question information. If the intention is to perform document matching, obtain relevant information corresponding to multiple objects to be matched based on the user question information. The relevant information includes identification information corresponding to the multiple objects to be matched.
[0084] Exemplarily, taking vehicle model document benchmarking as an example, the object to be benchmarked is the vehicle model to be benchmarked, and the user can input the question "Help me compare the smart cockpit dimensions of vehicle model A and vehicle model B" through the terminal. Accordingly, the electronic device executing the embodiment of this method can obtain the user's question information and identify the intention corresponding to the user's question information. If the intention is to perform vehicle model document benchmarking, word slot extraction is performed on the user's question information to obtain relevant information corresponding to multiple vehicle models to be benchmarked. The relevant information includes, for example, the target benchmarking dimension and identification information corresponding to multiple vehicle models to be benchmarked (i.e., vehicle model information). Among them, when identifying the intention corresponding to the user's question information, the Rasa framework can be used to build an intention recognition module, and the intention recognition module is used to identify the intention corresponding to the user's question information. Specifically, the Rasa framework is used to build an intention recognition module, which can include the following modules:
[0085] (1) SpacyNLP module, which is used to perform deep semantic analysis on Chinese text using the spaCy library (an advanced natural language processing library), including part-of-speech tagging, dependency parsing, etc.; for example, the embodiment of the present invention can select the "zh_core_web_sm" model that supports Chinese semantics.
[0086] (2) SpacyTokenizer module, which is used to decompose the input text string into tokens. These tokens may include words, characters, sub-words, etc., laying a solid foundation for subsequent language processing and analysis steps, and ensuring the accuracy and operability of text data in the preprocessing stage.
[0087] (3) The SpacyFeaturizer module is used to extract a series of detailed and in-depth linguistic features from text data, covering multiple dimensions such as part-of-speech tagging, named entity recognition, dependency syntax analysis, and lexical feature extraction. These features provide rich semantic information for machine learning models and play a vital role in deeply understanding and interpreting the intrinsic meaning of the text.
[0088] (4) RegexEntityExtractor module, which is used to accurately identify and extract specific entities in text, such as dates, phone numbers and other structured information; especially when implementing object information extraction, for example, the "use_lookup_tables" method can be used, which can not only improve the accuracy of entity recognition, but also enhance the adaptability and robustness of the system when facing diverse text inputs.
[0089] (5) The LexicalSyntacticFeaturizer module is used to deeply explore the lexical and syntactic levels of the text. By extracting key features such as morphological changes, word order, and phrase structure of vocabulary, it provides the model with a window to deeply understand the deep structure and semantic relationships of the text, thereby more accurately grasping the meaning of information and the overall framework of the text in a complex context.
[0090] (6) CountVectorsFeaturizer module, which is used to convert text into numerical vectors. It uses the bag-of-words model technology to convert the original text content into a set of quantitative representations. These numerical vectors can effectively capture the frequency of key words and phrases in the text, providing a concise and powerful feature expression for subsequent machine learning models, allowing the model to focus on the most discriminative language elements in the text.
[0091] (7) The DIETClassifier module, namely the intent classifier, is used to accurately classify the user's intent through end-to-end learning during the training process, and can also simultaneously identify and extract key entity information. It can greatly improve the generalization ability of the model and the flexibility of handling complex dialogue scenarios, providing users with a more accurate interactive experience.
[0092] (8) The EntitySynonymMapper module plays the role of entity normalization and is used to cleverly map various synonyms, near-synonyms, or variants in user input to predefined standard entity forms. This process can significantly enhance the model's ability to understand and recognize different expressions, thereby improving the accuracy and user-friendliness of the entire dialogue system.
[0093] (9) ResponseSelector module, which is used to select the most appropriate answer from predefined response options to build a rule-based dialogue system.
[0094] It can be understood that identifying the intent corresponding to the user's question information through the intent recognition module can make user queries more intelligent, achieve accurate matching of user intent and rapid retrieval of identification information, enhance user experience, and thus help improve the accuracy of document matching.
[0095] S402. For each of the multiple objects to be benchmarked, if the relevant information includes a target benchmarking dimension, obtain the target paragraph information corresponding to the object to be benchmarked based on the identification information and the target benchmarking dimension; wherein the target paragraph information includes the target paragraph summary, the context paragraph information of the original paragraph corresponding to the target paragraph summary, and the benchmarking dimension to which the target paragraph summary belongs.
[0096] The detailed description of this step can be found in Figure 2The relevant description of S202 in the illustrated embodiment will not be repeated here.
[0097] In the embodiment of the present invention, Figure 2 Step S202 may further include the following four steps S403 to S406:
[0098] S403 : For each of the multiple objects to be benchmarked, if the relevant information does not include a target benchmarking dimension, vectorize the relevant information using a document vectorization model to obtain a target vector.
[0099] Exemplarily, the document vectorization model is obtained through unsupervised training, and the document vectorization model can be used to vectorize relevant information to obtain a target vector.
[0100] S404: Match the target vector with each of the multiple paragraph summary vectors to obtain a similarity between the target vector and each paragraph summary vector.
[0101] For example, refer to Figure 3 In this example, multiple paragraph summary vectors corresponding to different objects are stored in a vector database. After obtaining a target vector, the vector database can be searched and matched with each paragraph summary vector in the database to obtain the similarity between the target vector and each paragraph summary vector. It can be understood that because the target vector is obtained based on the target benchmarking dimension and the secondary dimensions contained in the target benchmarking dimension, when searching the vector database based on the target vector, the retrieval of irrelevant information can be reduced, thereby shortening the content (context) length of the large model and reducing server memory usage.
[0102] S405 : Determine a preset number of target paragraph summary vectors with relatively high similarity among the multiple paragraph summary vectors whose similarity is less than a similarity threshold.
[0103] For example, the similarity threshold is 1, and the preset number is N. Then, the similarities can be sorted, for example, in ascending order, to determine the top N (topN) target paragraph summary vectors with higher similarity whose similarity is less than the similarity threshold among multiple paragraph summary vectors.
[0104] S406. Based on the target paragraph summary vector, obtain the target paragraph information corresponding to the target object. The target paragraph information includes the target paragraph summary, the contextual paragraph information of the original paragraph corresponding to the target paragraph summary, and the benchmarking dimension to which the target paragraph summary belongs. For example, taking vehicle model document benchmarking as an example, suppose a user inputs, "Help me compare the human-machine voice interaction features of vehicle models A and B." The target benchmarking dimension cannot be determined. The specific process for obtaining the benchmarking dimension is as follows: The original paragraph is: Voice: Supports five-zone recognition; Voice control capabilities: Supports see-and-speak, cross-zone inheritance, multiple intents per sentence, and inclusion in large-scale model dialogues. It also supports rapid vehicle control adjustments, such as voice-activated rearview mirror angle adjustment and steering wheel height adjustment, with "pause" interruption support during adjustments; Voice communication between mobile phone and vehicle: Supports voice-activated calls to mobile apps. The benchmarking dimension to which it belongs is: Smart Cockpit # Smart Voice, where Smart Cockpit is the first-level dimension and Smart Voice is the second-level dimension. According to the original paragraph summary vector matched to the original input, the corresponding target paragraph summary, the contextual paragraph information of the original paragraph corresponding to the target paragraph summary, and the benchmarking dimension of the benchmarking dimension to which the target paragraph summary belongs are brought out: Smart Cockpit # Smart Voice. In the case of missing target benchmarking dimensions, the benchmarking capability can be uniformly lowered to the benchmarking dimension.
[0105] Exemplarily, the vector database can be queried based on the target paragraph summary vector to obtain the target paragraph summary corresponding to the target paragraph summary vector, the context paragraph information of the original paragraph corresponding to the target paragraph summary, and the benchmarking dimension to which the target paragraph summary belongs, that is, to obtain the target paragraph information corresponding to the object to be benchmarked.
[0106] In the embodiment of the present invention, Figure 2 Step S203 may further include the following two steps S407 and S408:
[0107] S407. Determine the target prompt word information corresponding to the large model based on the identification information, the object evaluation dimension score information, and the target paragraph information. The object evaluation dimension score information is obtained by the large model by evaluating multiple benchmarking dimensions based on the paragraph summary.
[0108] In this step, the object evaluation dimension score information is obtained by evaluating multiple benchmarking dimensions based on the paragraph summary of the large model, and includes multiple first-level dimensions and scores corresponding to the first-level dimensions. For specific information on how to obtain the object evaluation dimension score information, please refer to the subsequent embodiments, which will not be repeated here. For example, after obtaining the target paragraph information corresponding to the object to be benchmarked, the target prompt word information corresponding to the large model can be determined based on the identification information, object evaluation dimension score information and target paragraph information corresponding to the multiple objects to be benchmarked. The target prompt word information includes the identification information, object evaluation dimension score information and target paragraph information corresponding to the multiple objects to be benchmarked, wherein the target paragraph information includes the target paragraph summary, the context paragraph information of the original paragraph corresponding to the target paragraph summary and the benchmarking dimension to which the target paragraph summary belongs.
[0109] S408: Input the target prompt word information into the large model to obtain target matching results corresponding to multiple to-be-matched objects output by the large model.
[0110] For example, after the target prompt word information corresponding to the large model is determined, a prompt word template may be spliced according to the target prompt word information. Figure 5 A schematic diagram of a prompt word template provided by an embodiment of the present invention is shown as follows: Figure 5 As shown, taking vehicle model document benchmarking as an example, the benchmarking object is the vehicle model to be benchmarked. Taking the vehicle models A and B as examples, the scoring information of vehicle model A is the object evaluation dimension score information of vehicle model A, and the scoring information of vehicle model B is the object evaluation dimension score information of vehicle model B. The relevant information of vehicle model A is the target paragraph information corresponding to vehicle model A, and the relevant information of vehicle model B is the target paragraph information corresponding to vehicle model B. Among them, the target paragraph information is obtained based on the target benchmarking dimension. Therefore, the target paragraph information can correspond to different target benchmarking dimensions and the secondary dimensions contained in the target benchmarking dimension. For example, "smart cockpit#smart voice" represents the smart cockpit dimension and the secondary dimension of smart voice contained in the smart cockpit dimension. In this step, the spliced prompt word template is input into the large model, and the target benchmarking results corresponding to multiple objects to be benchmarked output by the large model can be obtained. The target benchmarking results corresponding to multiple objects to be benchmarked can be obtained more accurately, thereby improving the reliability and professionalism of document benchmarking.
[0111] The document matching method provided by the embodiment of the present invention obtains user question information and identifies the intention corresponding to the user question information. If the intention is to perform document matching, then the relevant information corresponding to multiple objects to be matched is obtained according to the user question information. The relevant information includes identification information corresponding to the multiple objects to be matched, which can make user queries more intelligent, achieve accurate matching of user intentions and fast retrieval of identification information, and improve user experience; for each object to be matched among the multiple objects to be matched, if the relevant information includes a target matching dimension, then the target paragraph information corresponding to the object to be matched is obtained according to the identification information and the target matching dimension; if the relevant information does not include the target matching dimension, the relevant information is vectorized by a document vectorization model to obtain a target vector; the target vector is matched with each paragraph summary vector in the multiple paragraph summary vectors to obtain the similarity between the target vector and each paragraph summary vector, and it is determined that the similarity among the multiple paragraph summary vectors is less than A preset number of target paragraph summary vectors with high similarity before the similarity threshold can quickly and accurately obtain the target paragraph summary vector; according to the target paragraph summary vector, the target paragraph information corresponding to the object to be benchmarked is obtained, and the target paragraph information includes the target paragraph summary, the context paragraph information of the original paragraph corresponding to the target paragraph summary, and the benchmarking dimension to which the target paragraph summary belongs; according to the identification information corresponding to multiple objects to be benchmarked, the object evaluation dimension score information and the target paragraph information, the target prompt word information corresponding to the large model is determined, and the object evaluation dimension score information is obtained by the large model through evaluation of multiple benchmarking dimensions based on the paragraph summary; the target prompt word information is input into the large model to obtain the target benchmarking results corresponding to multiple objects to be benchmarked output by the large model, so as to realize automatic output of the target benchmarking results corresponding to multiple objects to be benchmarked through the large model, which can greatly improve the efficiency of document benchmarking and can more accurately obtain the target benchmarking results corresponding to multiple objects to be benchmarked.
[0112] Based on the above embodiments, Figure 6 This is a flow chart of a method for obtaining object evaluation dimension score information provided by an embodiment of the present invention, wherein the benchmarking dimension includes multiple first-level dimensions and each first-level dimension includes a second-level dimension. Figure 6 As shown, the method of the embodiment of the present invention may include:
[0113] S601. For each secondary dimension included in the primary dimension, extract target data related to the scoring index from the paragraph summaries corresponding to different objects based on the scoring index corresponding to the secondary dimension through the large model; score the secondary dimension based on the target data and the scoring criteria corresponding to the secondary dimension to obtain the score corresponding to the secondary dimension.
[0114] Exemplarily, the large model uses paragraph summaries corresponding to different objects to evaluate preset first-level dimensions. Each first-level dimension contains multiple second-level dimensions, and the large model scores each second-level dimension based on the paragraph summaries, using predefined scoring criteria and scoring indicators. Specifically, the scoring process for each second-level dimension is as follows: the scoring criteria and scoring indicators for the second-level dimension are determined, and then the relevant information in the paragraph summary collection is analyzed to extract target data related to the scoring indicators. Based on the extracted target data and the scoring criteria, a score is assigned to each second-level dimension, thereby completing the scoring of all second-level dimensions.
[0115] S602. The scores corresponding to the secondary dimensions are summed up through the large model to obtain the scores corresponding to the primary dimensions. The object evaluation dimension score information includes the primary dimensions and the scores corresponding to the primary dimensions.
[0116] For example, after obtaining the scores corresponding to each secondary dimension contained in each first-level dimension, the large model can be used to sum the scores corresponding to all secondary dimensions under each first-level dimension to obtain the total score (score) corresponding to the first-level dimension, thereby obtaining detailed object evaluation dimension score information. The object evaluation dimension score information includes the first-level dimension and the scores corresponding to the first-level dimension. The object evaluation dimension score information can be managed in a table, and accordingly, a detailed object evaluation dimension score table can be obtained.
[0117] The method for obtaining object evaluation dimension score information provided by an embodiment of the present invention is to extract target data related to the scoring index from paragraph summaries corresponding to different objects based on the scoring index corresponding to the secondary dimension for each secondary dimension included in the primary dimension through a large model; score the secondary dimension based on the target data and the scoring criteria corresponding to the secondary dimension to obtain the score corresponding to the secondary dimension; and sum the scores corresponding to the secondary dimension through the large model to obtain the score corresponding to the primary dimension. The object evaluation dimension score information includes the scores corresponding to the primary dimension and the primary dimension. The embodiment of the present invention can accurately obtain the object evaluation dimension score information through the large model, thereby using the object evaluation dimension score information to accurately obtain target benchmarking results corresponding to multiple objects to be benchmarked.
[0118] Based on the above embodiments, the big model used to output target benchmarking results corresponding to multiple objects to be benchmarked in the embodiments of the present invention has the following differences in technical implementation from the big models of current related technologies: the big model of related technologies only triggers the trained knowledge base inside the big model through prompts, and relies on the big model's own semantic understanding and reasoning capabilities to generate answers; the output dimensions are automatically extracted by the big model based on the keywords of the input question, and lacks explicit constraints on the dimension hierarchy; it only relies on the historical corpus during the training of the big model, and cannot dynamically introduce external knowledge or preset structured dimensions; it may ignore specific dimensions required by the user, or cause erroneous information due to training data deviation; if the input question involves a field that the big model has not been fully trained in (such as emerging technologies), it is easy for the big model to have "hallucination" problems and output false information. The large model of the embodiment of the present invention uses a preset multi-level dimensional framework (such as the first-level dimension and the second-level dimension) to forcibly constrain the input structure of the large model, and clearly specifies the dimension information in the prompt word to ensure that the large model must analyze and output each preset dimension to avoid omissions or random selections. In addition, it is combined with the preset domain knowledge base (i.e., the data in the vector database) and called in real time during the process of generating the target benchmarking results, which can ensure the consistency of the output content with the predefined dimensions and knowledge.
[0119] Accordingly, in terms of technical effectiveness, the output dimensions of the large model in related technologies are determined independently by the large model, and cannot guarantee coverage of, for example, user-specified secondary dimensions. This reliance on the large model's internal knowledge can easily lead to output errors. However, the large model in the embodiments of the present invention uses preset dimensions (i.e., primary and secondary dimensions) to ensure that the output content of the large model fully covers the preset dimensions. All output content is based on a preset domain knowledge base. By basing the output of the large model on the domain knowledge base, the "hallucination" problem can be avoided, i.e., the large model will not attempt to output content that is not recorded in the domain knowledge base.
[0120] Furthermore, the large models of related technologies have certain limitations in their application. Switching to a new domain requires complete reliance on the large model for relearning, making it difficult to quickly adapt to the new domain. Furthermore, when users temporarily add new dimensions, the large model must be retrained or the prompt word must be significantly modified. However, the large model of the embodiments of the present invention can quickly adapt to new domains or changing requirements by modifying preset dimensions. It also supports regular updates to the pre-set domain knowledge base (for example, by introducing the latest industry standards), ensuring that the output content is always consistent with the latest technical specifications.
[0121] The following are embodiments of the apparatus of the present invention, which can be used to implement the method embodiments of the present invention. For details not disclosed in the apparatus embodiments of the present invention, please refer to the method embodiments of the present invention.
[0122] Figure 7 This is a schematic diagram of the structure of a document alignment device provided by an embodiment of the present invention. Figure 7 As shown, the document alignment device 700 of the embodiment of the present invention includes: an acquisition module 701, a processing module 702 and an output module 703.
[0123] The acquisition module 701 is configured to acquire relevant information corresponding to a plurality of objects to be benchmarked, where the relevant information includes identification information corresponding to the plurality of objects to be benchmarked.
[0124] Processing module 702 is used to obtain, for each of the multiple objects to be benchmarked, target paragraph information corresponding to the object to be benchmarked based on the identification information and the target benchmarking dimension if the relevant information includes the target benchmarking dimension; if the relevant information does not include the target benchmarking dimension, vectorize the relevant information to obtain a target vector, and obtain the target paragraph information corresponding to the object to be benchmarked based on the target vector and multiple paragraph summary vectors; the target paragraph information includes the target paragraph summary, the context paragraph information of the original paragraph corresponding to the target paragraph summary, and the benchmarking dimension to which the target paragraph summary belongs. The multiple paragraph summary vectors, benchmarking dimension, paragraph summary, original paragraph, and context paragraph information are obtained by pre-processing the relevant documents corresponding to different objects.
[0125] The output module 703 is used to input the identification information and the target paragraph information into the large model to obtain target alignment results corresponding to multiple to-be-aligned objects output by the large model.
[0126] Furthermore, the output module 703 can be specifically used to: determine the target prompt word information corresponding to the large model based on the identification information, the object evaluation dimension score information and the target paragraph information, the object evaluation dimension score information is obtained by the large model by evaluating multiple benchmarking dimensions based on the paragraph summary; input the target prompt word information into the large model to obtain the target benchmarking results corresponding to the multiple objects to be benchmarked output by the large model.
[0127] Furthermore, the benchmarking dimension includes multiple first-level dimensions and second-level dimensions included in each first-level dimension. The acquisition module 701 can also be used to obtain object evaluation dimension score information in the following manner: for each second-level dimension included in the first-level dimension, the large model is used to extract target data related to the scoring index from the paragraph summary according to the scoring index corresponding to the second-level dimension; the second-level dimension is scored according to the target data and the scoring criteria corresponding to the second-level dimension to obtain the score corresponding to the second-level dimension; the score corresponding to the second-level dimension is added up by the large model to obtain the score corresponding to the first-level dimension. The object evaluation dimension score information includes the first-level dimension and the score corresponding to the first-level dimension.
[0128] Furthermore, the acquisition module 701 can be specifically used to: obtain relevant information corresponding to multiple objects to be benchmarked in response to a first selection operation for the target benchmarking dimension and a second selection operation for multiple objects to be benchmarked, or in response to a second selection operation for multiple objects to be benchmarked; or, obtain user question information, identify the intention corresponding to the user question information, and if the intention is to perform document benchmarking, obtain relevant information corresponding to multiple objects to be benchmarked based on the user question information.
[0129] Furthermore, when the processing module 702 is used to obtain the target paragraph information corresponding to the object to be benchmarked based on the identification information and the target benchmarking dimension, it can be specifically used to: if the target benchmarking dimension is a first-level dimension, then according to the target benchmarking dimension, the second-level dimension contained in the target benchmarking dimension and the identification information, obtain the target paragraph information corresponding to the object to be benchmarked; if the target benchmarking dimension is a second-level dimension, then according to the target benchmarking dimension, the first-level dimension to which the target benchmarking dimension belongs and the identification information, obtain the target paragraph information corresponding to the object to be benchmarked.
[0130] Furthermore, when the processing module 702 is used to obtain target paragraph information corresponding to the object to be benchmarked based on the target vector and multiple paragraph summary vectors, it can be specifically used to: match the target vector with each paragraph summary vector in the multiple paragraph summary vectors to obtain the similarity between the target vector and each paragraph summary vector; determine the first preset number of target paragraph summary vectors with higher similarity among the multiple paragraph summary vectors whose similarity is less than a similarity threshold; and obtain the target paragraph information corresponding to the object to be benchmarked based on the target paragraph summary vector.
[0131] Furthermore, the document matching device 700 may further include a pre-processing module ( Figure 7 ), which is not shown in the figure, is used to: split the document content of the relevant documents into paragraphs to obtain the original paragraphs, the paragraph sequence numbers of the original paragraphs, and the contextual paragraph information of the original paragraphs; obtain the paragraph summary and the benchmarking dimension to which the paragraph summary belongs from the original paragraph through the large model; vectorize multiple paragraph summaries through the document vectorization model to obtain multiple paragraph summary vectors, and the document vectorization model is obtained by unsupervised training based on the paragraph summary; store the identification information, benchmarking dimensions, original paragraphs, the paragraph sequence numbers of the original paragraphs, the contextual paragraph information of the original paragraphs, paragraph summaries, and paragraph summary vectors corresponding to different objects.
[0132] Furthermore, the document content also includes images. When the preprocessing module is used to split the document content of the relevant document into paragraphs to obtain the original paragraphs, it can be specifically used to: extract text information from the image using the optical character recognition (OCR) method; merge the text information with the paragraph corresponding to the text information to obtain the original paragraph.
[0133] Furthermore, when the processing module 702 is used to perform vectorization processing on the relevant information to obtain the target vector, it can be specifically used to: perform vectorization processing on the relevant information through a document vectorization model to obtain the target vector.
[0134] The device of the embodiment of the present invention can be used to execute the technical solution of any of the above-mentioned method embodiments. Its implementation principles and technical effects are similar and will not be repeated here.
[0135] Figure 8 This is a schematic diagram of the structure of a document alignment device provided by another embodiment of the present invention. Figure 8 As shown, the document benchmarking device 800 of the embodiment of the present invention includes: a document content processing module 801, a document content quantification module 802, a document content storage module 803, an object index scoring module 804 and an object comprehensive comparison module 805.
[0136] The document content processing module 801 is used to split and segment the document contents of the relevant documents corresponding to different objects, build a large-scale structured database based on the generated paragraph set, and then use the large model to generate paragraph summaries. The document content processing module 801 is specially designed to efficiently filter and process relevant documents uploaded by users. Specifically, the document content processing module 801 can accurately split relevant documents in PDF, DOCX and TXT formats by paragraphs, and split relevant documents in PPTX format by pages, use the OCR method to extract text information from the image, and integrate the text information into the original paragraph, thereby forming a structured document paragraph set; the paragraph set is then sent to the large model to perform the paragraph summary generation service, and then converted into a paragraph summary set containing key information, providing a basis for subsequent data processing; using the generated paragraph set, a large-scale, structured database is constructed, namely the information summary database, which not only contains a rich information summary corresponding to the object, but also provides the necessary data foundation for the subsequent training of the document vectorization model.
[0137] The document content quantification module 802 is used to adopt an unsupervised learning method to obtain the embedded vector of the object description through comparative learning training; generate vectors for paragraph summaries and output a set of paragraph summary vectors; the document content quantification module 802 adopts unsupervised learning technology to train a document vectorization model specific to the object field through data in the information summary database. The document vectorization model can effectively convert text information into a high-dimensional vector representation; using the pre-trained document vectorization model, feature extraction is performed on the document paragraph set, and the identification information of the object can be converted into a vector representation in a high-dimensional space.
[0138] The document content storage module 803 is used to store the paragraph summary vector, identification information, benchmarking dimension, paragraph summary, original paragraph, paragraph sequence number of the original paragraph, and context paragraph information of the original paragraph into the vector database; the document content storage module 803, for example, stores the relevant information of the relevant documents in the Milvus vector database, providing a basis for efficient retrieval of the vector database.
[0139] The object indicator scoring module 804 is used to evaluate multiple benchmarking dimensions based on paragraph summaries through the large model, obtain object evaluation dimension score information, and form an object evaluation dimension score table; the large model uses paragraph summaries to evaluate preset first-level dimensions, each first-level dimension includes multiple second-level dimensions, and the large model will score each second-level dimension based on the paragraph summary, and the scoring is based on pre-defined scoring standards and scoring indicators; specifically, the scoring process for each second-level dimension is as follows: determine the scoring standards and scoring indicators for the second-level dimension; then analyze the relevant information in the paragraph summary set, extract target data related to the scoring indicators, and score each second-level dimension based on the extracted target data and the scoring standards. After completing the scoring of all second-level dimensions, add up the scores of all second-level dimensions under each first-level dimension to obtain the total score of the first-level dimension, forming a detailed object evaluation dimension score table.
[0140] The object comprehensive comparison module 805 includes intent recognition, semantic matching and document generation. First, the object comprehensive comparison module 805 performs intent recognition on the user question information, and extracts the identification information and target benchmarking dimensions in the user question information through semantic matching to perform conditional query; then, the user question information and the object evaluation dimension score table are combined to match the data in the vector database, and the topN matching results are output; finally, the matching results are used to generate the benchmarking document; specifically, by classifying the user question information, when the user wants to perform object benchmarking, the target benchmarking dimensions and the identification information of the object to be compared are captured through semantic matching; if the user question information contains the target benchmarking dimension, the target benchmarking dimension ... If the user question information does not contain the target alignment dimension, the target paragraph information corresponding to the target object is obtained based on the identification information and the target alignment dimension, which serves as the target prompt word information for the large model. If the user question information does not contain the target alignment dimension, the document vectorization model is called to convert the user question information into a vector representation, and then efficiently matches it with the data stored in the document content storage module 803, ultimately returning the top N most relevant matching results. After determining the target prompt word information corresponding to the large model based on the top N matching results, a prompt word template is spliced based on the target prompt word information. This spliced prompt word template is input into the large model, resulting in a detailed and accurate document alignment result output by the large model. The large model service module is reused in the document content processing module 801, the object indicator scoring module 804, and the object comprehensive comparison module 805, effectively improving the quality and efficiency of document alignment result generation.
[0141] The device of the embodiment of the present invention can be used to execute the technical solution of any of the above-mentioned method embodiments. Its implementation principles and technical effects are similar and will not be repeated here.
[0142] Figure 9 This is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. Figure 9 As shown, the electronic device 900 may include: at least one processor 901 and a memory 902 .
[0143] The memory 902 is used to store programs. Specifically, the programs may include program codes, and the program codes include computer-executable instructions.
[0144] The memory 902 may include a high-speed random access memory (RAM), and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.
[0145] Processor 901 is configured to execute computer-executable instructions stored in memory 902 to implement the document alignment method described in the aforementioned method embodiment. Processor 901 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention. Specifically, when implementing the document alignment method described in the aforementioned method embodiment, the electronic device may be, for example, a server.
[0146] Optionally, the electronic device 900 may further include a communication interface 903. In a specific implementation, if the communication interface 903, memory 902, and processor 901 are implemented independently, the communication interface 903, memory 902, and processor 901 may be interconnected via a bus and communicate with each other. The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. Buses can be categorized as address buses, data buses, control buses, and so on, but this does not mean there is only one bus or only one type of bus.
[0147] Optionally, in a specific implementation, if the communication interface 903, the memory 902 and the processor 901 are integrated on a chip, the communication interface 903, the memory 902 and the processor 901 can complete communication through an internal interface.
[0148] The present invention also provides a computer-readable storage medium, in which computer program instructions are stored. When a processor executes the computer program instructions, the above document alignment method is implemented.
[0149] The present invention also provides a computer program product, including a computer program, which implements the above document alignment method when executed by a processor.
[0150] The computer-readable storage medium can be implemented by any type of volatile or non-volatile memory device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The computer-readable storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0151] An exemplary readable storage medium is coupled to a processor, such that the processor can read information from and write information to the readable storage medium. Of course, the readable storage medium may also be an integral part of the processor. The processor and the readable storage medium may be located in an application-specific integrated circuit. Of course, the processor and the readable storage medium may also be present as discrete components in the document alignment device.
[0152] Those skilled in the art will appreciate that all or part of the steps in the above-described method embodiments can be implemented using hardware associated with program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0153] Finally, it should be noted that the above embodiments are only preferred embodiments for fully illustrating the present invention, and the scope of protection of the present invention is not limited thereto. Equivalent substitutions or modifications made by those skilled in the art based on the present invention are all within the scope of protection of the present invention.
Claims
1. A document benchmarking method, characterized in that: include: Obtaining relevant information corresponding to a plurality of objects to be benchmarked input by a user, wherein the relevant information includes identification information corresponding to the plurality of objects to be benchmarked respectively; The object to be benchmarked is the vehicle model to be benchmarked, and the identification information is vehicle model information; For each of the plurality of objects to be benchmarked, If the relevant information further includes a target benchmarking dimension input by the user, then obtaining target paragraph information corresponding to the object to be benchmarked from a vector database according to the identification information and the target benchmarking dimension; If the relevant information does not include the target benchmarking dimension input by the user, the relevant information is vectorized to obtain a target vector, and the target paragraph information corresponding to the object to be benchmarked is obtained from the vector database according to the target vector; wherein the target paragraph information includes a target paragraph summary, context paragraph information of the original paragraph corresponding to the target paragraph summary, and the benchmarking dimension to which the target paragraph summary belongs; the benchmarking dimension includes multiple first-level dimensions and a second-level dimension included in each first-level dimension; the vector database stores identification information of different objects, original paragraphs, context paragraph information of the original paragraphs, paragraph summaries, benchmarking dimensions to which the paragraph summaries belong, and paragraph summary vectors obtained after preprocessing the relevant documents corresponding to different objects; Inputting the identification information, object evaluation dimension score information, and target paragraph information corresponding to the multiple objects to be benchmarked into the large model, and obtaining the target benchmarking results corresponding to the multiple objects to be benchmarked output by the large model; the object evaluation dimension score information includes: the first-level dimension of the object to be benchmarked and the score corresponding to the first-level dimension of the object to be benchmarked; The object evaluation dimension score information is obtained in the following way: For each secondary dimension included in the primary dimension of the object to be benchmarked, extracting target data related to the scoring indicator from the paragraph summary set corresponding to the object to be benchmarked using the large model according to the scoring indicator corresponding to the secondary dimension; Scoring the secondary dimension according to the target data and the scoring criteria corresponding to the secondary dimension to obtain a score corresponding to the secondary dimension; The scores corresponding to the secondary dimensions are summed up by the large model to obtain the scores corresponding to the primary dimensions; Among them, obtaining the target paragraph information corresponding to the object to be benchmarked from the vector database based on the target vector includes: obtaining the target paragraph information corresponding to the object to be benchmarked based on the similarity between the target vector and multiple paragraph summary vectors corresponding to the identification information stored in the vector database.
2. The document benchmarking method according to claim 1, characterized in that: The obtaining of relevant information corresponding to the multiple objects to be benchmarked input by the user includes: In response to a first selection operation of the user on the target benchmarking dimension and a second selection operation on the multiple objects to be benchmarked, or in response to a second selection operation of the user on the multiple objects to be benchmarked, obtaining relevant information corresponding to the multiple objects to be benchmarked; Alternatively, user question information input by the user is obtained, and the intention corresponding to the user question information is identified. If the intention is to perform document matching, relevant information corresponding to multiple objects to be matched is obtained based on the user question information.
3. The document benchmarking method according to claim 1, characterized in that: The acquiring, from a vector database according to the identification information and the target benchmarking dimension, target paragraph information corresponding to the object to be benchmarked includes: If the target benchmarking dimension is the primary dimension, then obtaining target paragraph information corresponding to the to-be-benchmarked object from a vector database according to the target benchmarking dimension, the secondary dimension included in the target benchmarking dimension, and the identification information; If the target benchmarking dimension is the secondary dimension, target paragraph information corresponding to the object to be benchmarked is acquired from a vector database according to the target benchmarking dimension, the primary dimension to which the target benchmarking dimension belongs, and the identification information.
4. The document benchmarking method according to claim 1, characterized in that: The acquiring target paragraph information corresponding to the object to be matched according to the similarity between the target vector and a plurality of paragraph summary vectors corresponding to the identification information stored in the vector database includes: Matching the target vector with each of the paragraph summary vectors corresponding to the identification information stored in the vector database to obtain a similarity between the target vector and each of the paragraph summary vectors; Determining a preset number of target paragraph summary vectors with relatively high similarity among the plurality of paragraph summary vectors whose similarity is less than a similarity threshold; According to the target paragraph summary vector, target paragraph information corresponding to the object to be aligned is obtained.
5. The document benchmarking method according to any one of claims 1 to 4, characterized in that: Also includes: Splitting the document content of the relevant document into paragraphs to obtain the original paragraphs, the paragraph sequence numbers of the original paragraphs, and the context paragraph information of the original paragraphs; Obtaining the paragraph summary and the benchmarking dimension to which the paragraph summary belongs from the original paragraph through the large model; Vectorizing the plurality of paragraph summaries using a document vectorization model to obtain the plurality of paragraph summary vectors, wherein the document vectorization model is obtained by unsupervised training based on the paragraph summaries; The identification information, benchmarking dimension, original paragraph, paragraph sequence number of the original paragraph, context paragraph information of the original paragraph, paragraph summary and paragraph summary vector corresponding to different objects are stored in the vector database.
6. The document benchmarking method according to claim 5, characterized in that: The document content further includes an image, and the document content of the related document is split into paragraphs to obtain the original paragraphs, including: Extracting text information from the image using an optical character recognition (OCR) method; The text information is merged with the paragraph corresponding to the text information to obtain the original paragraph.
7. The document benchmarking method according to claim 5, characterized in that: The vectorizing the relevant information to obtain a target vector includes: The relevant information is vectorized using the document vectorization model to obtain the target vector.
8. A document benchmarking device, characterized in that: include: An acquisition module, configured to acquire relevant information corresponding to a plurality of objects to be benchmarked input by a user, wherein the relevant information includes identification information corresponding to the plurality of objects to be benchmarked respectively; The object to be benchmarked is the vehicle model to be benchmarked, and the identification information is vehicle model information; a processing module configured to, for each of the plurality of objects to be benchmarked, if the relevant information further includes a target benchmarking dimension input by the user, obtain target paragraph information corresponding to the object to be benchmarked from a vector database based on the identification information and the target benchmarking dimension; if the relevant information does not include the target benchmarking dimension input by the user, perform vectorization processing on the relevant information to obtain a target vector, and obtain target paragraph information corresponding to the object to be benchmarked from a vector database based on the target vector; wherein the target paragraph information includes a target paragraph summary, contextual paragraph information of an original paragraph corresponding to the target paragraph summary, and the benchmarking dimension to which the target paragraph summary belongs; the benchmarking dimension includes multiple first-level dimensions and a second-level dimension included in each of the first-level dimensions; The vector database stores identification information of different objects obtained after preprocessing relevant documents corresponding to different objects, original paragraphs, contextual paragraph information of the original paragraphs, paragraph summaries, benchmarking dimensions to which the paragraph summaries belong, and paragraph summary vectors; An output module is configured to input identification information, object evaluation dimension score information, and target paragraph information corresponding to the multiple objects to be benchmarked, respectively, into a large model, and obtain target benchmarking results corresponding to the multiple objects to be benchmarked, output by the large model; the object evaluation dimension score information includes: the first-level dimension of the object to be benchmarked and the score corresponding to the first-level dimension of the object to be benchmarked; The processing module is further configured to extract, for each secondary dimension included in the primary dimension of the object to be benchmarked, target data related to the scoring indicator corresponding to the secondary dimension from the paragraph summary set corresponding to the object to be benchmarked using the large model; score the secondary dimension according to the target data and the scoring standard corresponding to the secondary dimension to obtain a score corresponding to the secondary dimension; and sum the scores corresponding to the secondary dimensions using the large model to obtain a score corresponding to the primary dimension; The processing module obtains target paragraph information corresponding to the object to be benchmarked from a vector database based on the target vector, and is specifically used to obtain the target paragraph information corresponding to the object to be benchmarked based on the similarity between the target vector and multiple paragraph summary vectors corresponding to the identification information stored in the vector database.
9. An electronic device, characterized in that: include: a processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the document alignment method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer program instructions, and when the computer program instructions are executed, the document alignment method according to any one of claims 1 to 7 is implemented.
11. A computer program product comprising a computer program, characterized in that When the computer program is executed, the document alignment method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
User intention response method, device and equipment of intelligent robot and storage medium
CN118171658A
Document processing method, system and equipment and storage medium
CN118245574A
Method and device for evaluating novelty and credibility of novelty search point under multi-model cooperation
CN119166746A