A method and system for generating pre-trip reports based on a vector model
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-29
- Publication Date
- 2026-08-11
AI Technical Summary
[0005]本发明为了解决现有技术中存在的问题,创新提出了一种基于向量模型的巡前报告生成方法及系统,有效解决由于现有技术造成巡前报告生成过程中准确度以及训练成本无法兼顾的问题,不仅提高了巡前报告生成准确性,而且降低了训练成本
Smart Images

Figure CN121235086B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of report generation, and in particular to a method and system for generating pre-trip reports based on a vector model. Background Technology
[0002] In the process of generating pre-inspection reports, traditional methods often rely on manual writing. Staff need to collect data from multiple channels, including internal enterprise systems, external databases, and historical documents, manually organize inspection points, summarize problems and risks, and write recommendations. This process is not only time-consuming and inefficient, but also prone to errors such as data omissions and semantic discrepancies due to manual operation. With the advancement of enterprise digitalization, the amount of various types of data involved in pre-inspection inspections is growing exponentially, and traditional manual methods are unable to meet the demand for efficient and accurate report generation.
[0003] While methods for automatically generating reports exist in related technologies, they generally rely on fixed templates and preset rules. For example, simple checklists are generated by filling structured data into Excel templates, or reports are pieced together by extracting information from documents based on keyword matching. These methods have significant drawbacks: they largely depend on simple templates and rules, lacking a deep understanding and processing capability of pre-inspection data, and are unsuitable for the application scenario of pre-inspection reports. Although large models can be used, they often employ general-purpose or specialized pre-inspection report models for question-and-answer retrieval. However, general-purpose models suffer from low professionalism and accuracy, while specialized pre-inspection report models have high training costs, failing to meet practical application needs and unable to balance accuracy and training cost in the pre-inspection report generation process.
[0004] To address this problem, the present invention provides a method and system for generating pre-circuit reports based on a vector model, thereby solving the aforementioned issues. Summary of the Invention
[0005] To address the problems existing in the prior art, this invention innovatively proposes a pre-cycle report generation method and system based on a vector model. This effectively solves the problem that the accuracy and training cost cannot be balanced in the pre-cycle report generation process due to the prior art. It not only improves the accuracy of pre-cycle report generation but also reduces the training cost.
[0006] The first aspect of this invention provides a method for generating pre-circuit reports based on a vector model, comprising: Acquire pre-inspection report document data and establish the first training dataset; the first training dataset includes pre-inspection report document fragments and text descriptions. Based on the first training dataset, train and adjust the vector model; based on the adjusted vector model, obtain the first text description vector and the first pre-circuit report document fragment vector, and store them in a pre-established knowledge base; Using the semantic description in the user's needs as the text description, the k first pre-circuit report document fragments with the highest similarity ranking are obtained based on the similarity between the first text description vector corresponding to the user's needs and the first pre-circuit report document fragment vector in the knowledge base. The user requirements, k fragments of the first pre-cycle report document, and optional original pre-cycle document data are used as inputs, and the pre-cycle report is used as output to construct a second training dataset. The large model is then trained and adjusted based on the second training dataset. The large model is a large language model. Using the semantic description in the current user needs as the text description, based on the similarity between the current text description vector corresponding to the current user needs and the first pre-circuit report document fragment vector in the knowledge base, the k second pre-circuit report document fragments with the highest similarity ranking are obtained. The current user requirements, k fragments of the second pre-circuit report document, and optional original pre-circuit document data are used as inputs to the adjusted large model, and the pre-circuit report is output.
[0007] Optionally, the pre-inspection report document data is obtained to establish the first training dataset, specifically as follows: Pre-inspection document data is collected from various data sources, including inspection records, historical reports, and expert opinions in different structural formats. Preprocessing of pre-inspection document data; The pre-inspection document data is segmented according to different parts of the report content. The segmented pre-inspection report document fragments include background information document fragments, inspection content document fragments, risk and problem discovery document fragments, rectification suggestion document fragments, and public opinion summary document fragments. The obtained pre-inspection report document fragments are associated with the original pre-inspection document data one by one through metadata or ID. The segmented pre-circuit report document fragments were labeled to obtain the first training set.
[0008] Furthermore, based on the first training dataset, training the adjusted vector model specifically includes: The labeled pre-patrol report document fragments and labeled text descriptions in the first training dataset are used as positive samples; The pre-process report document fragments and randomly generated text descriptions, or pre-process report document fragments randomly generated from labeled text descriptions in the first training dataset, are used as the first negative samples. In each training round, the second negative sample is dynamically selected using the vector model under adjustment; wherein, the similarity between the second pre-cycle report document fragment vector corresponding to the first negative sample and the second text description vector is less than the similarity between the second pre-cycle report document fragment vector corresponding to the second negative sample and the second text description vector. In the early stage of training, the vector model is pre-trained using the first negative sample, and in the later stage of training, the vector model is adjusted using the second negative sample.
[0009] Furthermore, during the training process of the vector model, the contrastive loss function is specifically as follows: , Where L is the contrastive loss function, q is the text description vector, p is the pre-pass report document vector in the positive samples, and τ is the temperature coefficient for training the vector model. Let n be the cosine similarity between the text description vector and the pre-patrol report document vector in the positive samples, where j is the j-th first negative sample or second negative sample, k is the sum of the first negative sample and the second negative sample, and n is the cosine similarity between the two vectors. j This refers to the pre-circuit report document vector in the j-th first negative sample or second negative sample. The weight is the weight corresponding to the j-th first negative sample or the second negative sample.
[0010] Furthermore, if the negative sample is the first negative sample, then The value is 1; if the negative sample is the second negative sample, then Greater than 1.
[0011] Optionally, the first text description vector and the first pre-circuit report document fragment vector are obtained based on the adjusted vector model and stored in a pre-established knowledge base, specifically as follows: Based on the adjusted vector model, the first text description vector and the first pre-circuit report document fragment vector are obtained, and stored in a pre-established knowledge base respectively. When storing the first pre-circuit report document fragment vector in the pre-established knowledge base, the ID or metadata, the pre-circuit report document fragment, and the generated first pre-circuit report document fragment vector are stored as a complete record in the vector database. When performing a similarity search based on the current text description vector corresponding to the current user's needs, the k most similar first pre-circuit report document fragment vectors are returned, along with the ID and original pre-circuit document data corresponding to each first pre-circuit report document fragment vector.
[0012] Optionally, a second training dataset is constructed by taking user requirements, k fragments of the first pre-cycle report document, and optional original pre-cycle document data as input and the pre-cycle report as output. The large model is then trained and adjusted based on this second training dataset, specifically including: The user requirements, k first pre-circuit report document fragments, and optional original pre-circuit document data are taken as input; wherein, the optional original pre-circuit document data are the original pre-circuit document data corresponding to each first pre-circuit report document fragment vector in the k first pre-circuit report document fragments, and other original pre-circuit document data that form a context with the original pre-circuit document data corresponding to each first pre-circuit report document fragment vector. The pre-inspection report, which includes risk point discovery, rectification suggestions, and public opinion summary, is used as the output to construct a second training dataset. The large model is then trained and adjusted based on the second training dataset. The public opinion summary is expressed in the form of a triplet consisting of the event subject, event behavior, and impact result.
[0013] Furthermore, in the large model training and adjustment, the specific method for updating the weight coefficients is as follows: W new = W original Learning rate Gradient Among them, W new For adjusting the updated weight parameters during the training of a large model, W original Weight parameters loaded for pre-training large models, learning rate The learning rate is adjusted during training of a large model, where Gradient is the gradient of the loss function with respect to each weight parameter of the large model, calculated via backpropagation; and the learning rate is... rate Specifically: learning rate = , in, The initial learning rate, To control the decay rate, To reduce the learning rate to The number of steps at / 2, where t is the current training step during the large model training adjustment.
[0014] Alternatively, the main model can be the qwen-72B series of open-source models.
[0015] A second aspect of the present invention provides a pre-circuit report generation system based on a vector model, comprising: The acquisition module acquires pre-inspection report document data and establishes a first training dataset; the first training dataset includes pre-inspection report document fragments and text descriptions. The first training module trains and adjusts the vector model based on the first training dataset; based on the adjusted vector model, it obtains the first text description vector and the first pre-circuit report document fragment vector, and stores them in a pre-established knowledge base. The first module uses the semantic description in the user's requirements as the text description. Based on the similarity between the first text description vector corresponding to the user's requirements and the first pre-circuit report document fragment vector in the knowledge base, it obtains the k first pre-circuit report document fragments with the highest similarity ranking. The second training module takes user requirements, k fragments of the first pre-cycle report document, and optional original pre-cycle document data as input, and the pre-cycle report as output to construct a second training dataset. The large model is then trained and adjusted based on the second training dataset. The large model is a large language model. The second module uses the semantic description in the current user demand as the text description, and obtains the k second pre-circuit report document fragments with the highest similarity ranking based on the similarity between the current text description vector corresponding to the current user demand and the first pre-circuit report document fragment vector in the knowledge base. The output module takes the current user requirements, k fragments of the second pre-circuit report document, and optional original pre-circuit document data as input to the adjusted large model, and outputs the pre-circuit report.
[0016] The technical solution adopted in this invention has the following technical effects: 1. This invention establishes a first training dataset based on the acquisition of pre-inspection report document data; trains and adjusts a vector model based on the first training dataset; uses the semantic description in the user's requirements as the text description, and obtains the k first pre-inspection report document fragments with the highest similarity ranking based on the similarity between the first text description vector corresponding to the user's requirements and the first pre-inspection report document fragment vector in the knowledge base; constructs a second training dataset using the user's requirements, the k first pre-inspection report document fragments, and optional original pre-inspection document data as input, and the pre-inspection report as output, and trains and adjusts the large model based on the second training dataset; uses the current user's requirements, the k second pre-inspection report document fragments, and optional original pre-inspection document data as input to the adjusted large model, and outputs the pre-inspection report, effectively solving the problem that the accuracy and training cost cannot be balanced in the pre-inspection report generation process due to existing technologies, not only improving the accuracy of pre-inspection report generation, but also reducing the training cost.
[0017] 2. The technical solution of this invention collects pre-inspection document data from various data sources. This data includes inspection records, historical reports, and expert opinions in different structural formats. It supports structured, semi-structured, and unstructured data formats, thus ensuring the accuracy of the generated pre-inspection report. Furthermore, the pre-processed pre-inspection document data is segmented according to different parts of the report content. These segmented segments include background information, inspection content, risk and problem findings, rectification suggestions, and public opinion summary. The obtained pre-inspection report segments are associated with the original pre-inspection document data through metadata or IDs, ensuring a one-to-one correspondence between the pre-inspection report segments, pre-inspection report segment vectors, and the original pre-inspection document data during the pre-inspection report generation process.
[0018] 3. In the technical solution of this invention, the labeled pre-circuit report document fragments and labeled text descriptions in the first training dataset are used as positive samples; the pre-circuit report document fragments and randomly generated text descriptions in the first training dataset, or pre-circuit report document fragments randomly generated from labeled text descriptions, are used as first negative samples; in each round of training, the second negative sample is dynamically selected using the vector model under adjustment; the vector model is pre-trained using the first negative sample in the early stage of training, and the vector model is adjusted using the second negative sample in the later stage of training, so that the vector model can better adapt to the generation of pre-circuit reports and has higher accuracy after adjustment.
[0019] 4. In the technical solution of this invention, when the first pre-circuit report document fragment vector is stored in the pre-established knowledge base, the ID or metadata, the pre-circuit report document fragment, and the generated first pre-circuit report document fragment vector are stored as a complete record in the vector database. When performing a similarity search based on the current text description vector corresponding to the current user's needs, the k most similar first pre-circuit report document fragment vectors are returned, along with the unique ID and original pre-circuit document data corresponding to each first pre-circuit report document fragment vector. The original pre-circuit document data can be the original pre-circuit document data corresponding to each first pre-circuit report document fragment vector in the k first pre-circuit report document fragments, as well as other original pre-circuit document data that form the context with the original pre-circuit document data corresponding to each first pre-circuit report document fragment vector. This ensures the correlation and correspondence between the pre-circuit report document fragments, the pre-circuit report document fragment vectors, and the original pre-circuit document data during the pre-circuit report generation process.
[0020] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit the invention. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a flowchart illustrating the method of Embodiment 1 in the present invention; Figure 2 This is a flowchart illustrating step S1 in the method of Embodiment 1 of the present invention; Figure 3 This is a flowchart illustrating step S2 in the method of Embodiment 1 of the present invention; Figure 4 This is a schematic diagram illustrating the architecture of the vector model and the large model adjusted in the method of Embodiment 1 of the present invention; Figure 5 This is a schematic diagram of the system structure in Embodiment 2 of the present invention. Detailed Implementation
[0023] To clearly illustrate the technical features of this solution, the invention will be described in detail below through specific embodiments and in conjunction with the accompanying drawings. The following disclosure provides many different embodiments or examples for implementing different structures of the invention. To simplify the disclosure of the invention, components and arrangements of specific examples are described below. Furthermore, reference numerals and / or letters may be repeated in different examples. This repetition is for simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed. It should be noted that the components illustrated in the drawings are not necessarily drawn to scale. Descriptions of well-known components, processing techniques, and processes are omitted in this invention to avoid unnecessarily limiting the invention.
[0024] Example 1 like Figure 1 As shown, this invention provides a method for generating pre-circuit reports based on a vector model, including: S1, Obtain pre-inspection report document data and establish the first training dataset; the first training dataset includes pre-inspection report document fragments and text descriptions; S2, based on the first training dataset, train and adjust the vector model; based on the adjusted vector model, obtain the first text description vector and the first pre-circuit report document fragment vector, and store them in the pre-established knowledge base; S3, using the semantic description in the user's needs as the text description, and based on the similarity between the first text description vector corresponding to the user's needs and the first pre-circuit report document fragment vector in the knowledge base, obtain the k first pre-circuit report document fragments with the highest similarity ranking. S4 takes user requirements, k fragments of the first pre-cycle report document, and optional original pre-cycle document data as input, and the pre-cycle report as output to construct a second training dataset, and trains and adjusts the large model based on the second training dataset; where the large model is a large language model. S5: Using the semantic description in the current user needs as the text description, based on the similarity between the current text description vector corresponding to the current user needs and the first pre-circuit report document fragment vector in the knowledge base, the k second pre-circuit report document fragments with the highest similarity ranking are obtained. S6 takes the current user requirements, k second pre-trip report document fragments, and optional original pre-trip document data as input to the adjusted large model and outputs the pre-trip report.
[0025] Among them, such as Figure 2 As shown, in step S1, obtaining the pre-inspection report document data and establishing the first training dataset specifically involves: S11, collect pre-inspection document data from various data sources, including inspection records, historical reports and expert opinions in different structural forms; Collect pre-inspection data from various data sources (such as internal enterprise systems, external databases, questionnaires, etc.) to ensure data diversity and coverage. Collected data may include inspection records, historical reports, relevant documents, and expert opinions.
[0026] The pre-inspection document data mainly covers the following three categories: First, structured data, such as tabular financial statements or inspection records in Excel, CSV, etc., with clearly defined fields and formats; second, semi-structured data, such as expert opinions in JSON, XML, HTML, etc., which, although lacking a strict table structure, usually contain parsable tags or hierarchies; and third, unstructured data, such as pre-inspection historical report documents in PDF, Word, etc., whose content format is relatively flexible and the text semantics are rich. Regarding data content, accuracy, completeness, and timeliness must be ensured, and desensitization processing must be completed. Furthermore, to facilitate the construction of an automated data processing pipeline, similar data should maintain a consistent format as much as possible. That is, both semi-structured and unstructured data are converted into structured data. Existing methods or tools can be used for conversion, and this embodiment of the invention does not impose any limitations.
[0027] S12, Preprocessing of pre-inspection document data; Preprocessing primarily involves cleaning the collected pre-inspection document data, removing redundant and noisy data to ensure the accuracy and consistency of the pre-inspection document data. The cleaning process includes data format standardization, missing value imputation, and outlier handling.
[0028] For example, format standardization: unify the "date" field in the pre-inspection document data to "YYYY-MM-DD", and retain two decimal places in the "value" field; missing value imputation: use the KNN algorithm (k=5) or other methods to imput missing values in the pre-inspection document data; outlier handling: based on the 3σ principle, numerical data exceeding the preset range can be identified as outlier data. Outlier data can be discarded or manually corrected.
[0029] S13, the pre-inspection document data is segmented according to different parts of the report content. The segmented pre-inspection report document fragments include background information document fragments, inspection content document fragments, risk and problem discovery document fragments, rectification suggestion document fragments, and public opinion summary document fragments. The obtained pre-inspection report document fragments are associated with the original pre-inspection document data one-to-one through metadata or ID. First, the original pre-inspection document data is split into a list of text blocks (background information document fragments, inspection content document fragments, risk and issue discovery document fragments, rectification suggestion document fragments, and public opinion summary document fragments). Then, text embedding technology is used to extract vector representations of each text block, forming a vector list which is stored in a vector database, while simultaneously associating it with the corresponding original pre-inspection document data and metadata. The generated pre-inspection report document fragment vector data and the original pre-inspection document data are linked using an ID or metadata. The latter uses a unique identifier; specifically, before feeding the text block (pre-inspection report document fragment, e.g., a paragraph, an article) into the text embedding model (vector model) to generate vectors, a globally unique ID is generated for each text block. This (unique ID, pre-inspection report document fragment, generated vector) is treated as a complete record, facilitating the storage of the first text description vector and the first pre-inspection report document fragment vector output by the adjusted vector model into the vector database after model adjustment.
[0030] S14, annotate the segmented pre-pattern report document fragments to obtain the first training set.
[0031] Use manual or semi-automated tools to annotate pre-inspection report document fragments to provide high-quality labeled data for vector model training. If the pre-inspection report document fragments explicitly contain words such as "background information," "inspection content," "risk issues discovered," "rectification suggestions," or "public opinion summary," regular expression matching or string search can be used. If the pre-inspection report document fragments do not have fixed titles, they can be segmented by paragraphs (\n\n or period) first, and then identified using a rule + semantic classification model. Specifically, paragraph segmentation or sentence division can be done using "\n\n," for example, paragraphs = [p.strip() for p in text.split("\n\n") if p.strip()].
[0032] Then, combining rules and a semantic model: a sliding window is used, moving a distance of (chunk_size - chunk_overlap) each time for segmentation. Here, chunk_size is the maximum length of the input text sequence to be segmented, chunk_overlap is the number of overlapping tokens between two adjacent text sequence chunks, and the semantic model can be RecursiveCharacterTextSplitter (a class in the LangChain library for text segmentation, supporting recursive segmentation of text into blocks of a specified size).
[0033] Among them, such as Figure 3 As shown, in step S2, training the adjusted vector model based on the first training dataset specifically includes: S21, take the labeled pre-patrol report document fragments and labeled text descriptions in the first training dataset as positive samples; The vector model is trained using labeled pre-trip report document fragments and their corresponding text descriptions as positive samples. The training sample data includes labeled pre-trip report document fragments and their corresponding text descriptions.
[0034] The vector model plays a role in the pre-trip report generation process by performing knowledge retrieval, report template filling, problem discovery and summarization, and personalized suggestion generation.
[0035] Retrieval: Given a query vector (a vector generated from the semantic description of the user's requirements, i.e., a vector generated from the labeled text description), quickly find the k most similar first-round report document fragment vectors from a large vector knowledge base. Calculation methods can include cosine similarity, dot product, Euclidean distance, etc.
[0036] Top-bottom concatenation: Before embedding text into blocks, you can enhance the context by concatenating surrounding text blocks using IDs or metadata. For example, you can concatenate left and right blocks, such as the previous block + the current block + the next block. The relationships between different text blocks within the same report can be established using unique IDs or metadata.
[0037] S22, take the pre-process report document fragments and randomly generated text descriptions, or the pre-process report document fragments randomly generated from the labeled text descriptions in the first training dataset, as the first negative sample; Negative sample mining methods are used to generate negative sample data, which refers to data fragments that are irrelevant or have low relevance to the target task. Specifically, pre-trip report document fragments and randomly generated text descriptions from the first training dataset, or pre-trip report document fragments randomly generated from labeled text descriptions, are used as the first negative samples (ordinary negative samples). These negative samples improve the model's classification and retrieval capabilities. Positive samples have manually labeled pairs. The first negative sample can be generated through the context of the same pre-trip report document fragment, data augmentation, etc., or by using methods such as exclusion of similar data.
[0038] S23, in each round of training, the second negative sample is dynamically selected using the vector model under adjustment; wherein, the similarity between the second pre-cycle report document fragment vector corresponding to the first negative sample and the second text description vector is less than the similarity between the second pre-cycle report document fragment vector corresponding to the second negative sample and the second text description vector. Second Negative Sample (Hard Negative Sample) Selection / Generation: Hard negative samples are dynamically selected in each batch or training round. Specifically, BM25 / old vector model / SimCSE can be used for initial retrieval, and candidates with high similarity scores can be used as second negative samples. Selection criteria: high similarity but not positive samples; high confusion within the same domain; "false positive samples" dynamically retrieved during training, i.e., the similarity between the second pre-process report document fragment vector and the second text description vector corresponding to the first negative sample is less than the similarity between the second pre-process report document fragment vector and the second text description vector corresponding to the second negative sample. Specifically, the similarity range between the second pre-process report document fragment vector and the second text description vector corresponding to the second negative sample can be 0.5-0.8, and the similarity range between the second pre-process report document fragment vector and the second text description vector corresponding to the first negative sample can be 0.2-0.5.
[0039] S24, in the early stage of training, the vector model is pre-trained using the first negative sample, and in the later stage of training, the vector model is adjusted using the second negative sample.
[0040] Training methods: Weighted contrastive loss function (Loss) is used, giving higher weight to difficult negative samples. Sampling ratio control: For example, a positive:negative ratio of 1:3 is commonly used in detection to avoid overwhelming positive samples with too many negative samples. Phased training: Pre-training is performed using random negative samples in the early stages, and then fine-tuned using difficult negative samples in the later stages. The division between the early and later training stages can be based on the number of training steps or on the contrastive loss function and its threshold relative to a preset target. Adjustment is based on: similarity (higher similarity, higher weight), retrieval ranking (higher ranking, higher weight), and loss magnitude (higher loss, similar to Focal Loss).
[0041] The weighted contrastive loss function used in the training process of the vector model is as follows: , Where L is the contrastive loss function, q is the text description vector, p is the pre-pass report document vector in the positive samples, and τ is the temperature coefficient for training the vector model. Let n be the cosine similarity between the text description vector and the pre-patrol report document vector in the positive samples, where j is the j-th first negative sample or second negative sample, k is the sum of the first negative sample and the second negative sample, and n is the cosine similarity between the two vectors. j This refers to the pre-circuit report document vector in the j-th first negative sample or second negative sample. Let be the weight corresponding to the j-th first negative sample or the second negative sample. The value is 1, indicating it is the first negative sample; If the value is greater than 1, it is the second negative sample.
[0042] Specifically, if the negative sample is the first negative sample, then The value is 1; if the negative sample is the second negative sample, then Greater than 1.
[0043] Initial training phase: The vector model is almost random. In positive samples, the similarity between the pre-process report document fragment vector and the text description vector (i.e., the semantic description vector in the user's requirements) is low. In the first negative sample, the similarity between the pre-process report document fragment vector and the text description vector query may also be randomly high. The second negative sample might be treated as a positive sample by the vector model. Mid-training phase: The vector model can now correctly approximate the text description vector query and the positive sample pre-process report document fragment vector. Ordinary negative samples are easily distinguishable and no longer provide much information. Difficult negative samples still cause significant interference. Late-training phase: The model can basically distinguish most ordinary negative samples. Difficult negative samples have a larger weight and dominate the contrastive loss function. Training is similar to "difficult example-driven learning." The semantic representation ability of the vector model gradually improves, and the generalization effect is better.
[0044] How it works in practice: The vector model learns the "discrimination boundary" through the first negative sample. Difficult negative samples (second negative samples) allow the vector model to focus more on "subtle differences" rather than simple distinctions.
[0045] Vector model fine-tuning: Select a suitable vector model (bge-large) for fine-tuning. Utilize training data to adjust model parameters via backpropagation to improve the understanding of pre-prompt reports. In scenarios with limited training data or requiring deployment on low-computing devices, parameter-efficient fine-tuning methods can be chosen. The basis for this fine-tuning is the semantic alignment requirements of the vector model in downstream tasks; that is, the general representational capability of the pre-trained vector model needs to be adapted to the semantic distribution of a specific task. For example, LoRA (low-rank adaptation) reduces the number of parameters by inserting low-rank updates into the attention matrix; Prefix-tuning / Prompt-tuning only optimizes the cue vectors without changing the original vector model parameters.
[0046] Fine-tuning is based on differences in task semantics, distribution, and relevance. Vector model fine-tuning methods include supervised fine-tuning and parameter fine-tuning. Adapting to the semantic distribution of a specific task involves distribution alignment, continuing unsupervised contrastive learning on the target domain corpus to make the embedding distribution closer to the domain text. Task relevance modeling is crucial, defining "what is relevance." During fine-tuning, positive and negative sample pairs must meet the task's relevance criteria. Hard negative sample adaptation involves selecting the most easily confused samples in the task distribution as the best hard negatives. In-batch negatives for retrieval tasks are also important; retrieval task data is generally query-document pairs, and other documents within a batch can be used as negative samples. Backpropagation involves calculating the gradient of the parameters using the loss function L and then using optimization algorithms (such as SGD and Adam) to update the parameters.
[0047] The pre-inspection report allows for the querying of specific business text blocks, such as those related to corporate auditing, quality control, and compliance checks. These blocks are generated using the same model and encoded using the same vector model (e.g., text-embedding-ada-002, bge-m3). The relationship between the query vector and the text block is as follows: the query vector is a text description vector, representing the semantic description of the user's needs or questions; the text block vector is a vector generated from or corresponding to a fragment of the pre-inspection report document. The matching degree is calculated using cosine similarity and dot product.
[0048] Specifically, in step S2, the first text description vector and the first pre-circuit report document fragment vector are obtained based on the adjusted vector model and stored in the pre-established knowledge base. Based on the adjusted vector model, the first text description vector and the first pre-circuit report document fragment vector are obtained, and stored in a pre-established knowledge base. When storing the first pre-circuit report document fragment vector in the pre-established knowledge base, the ID or metadata, the pre-circuit report document fragment, and the generated first pre-circuit report document fragment vector are stored as a complete record in the vector database. When performing a similarity search based on the current text description vector corresponding to the current user's needs, the k most similar first pre-circuit report document fragment vectors are returned, along with the unique ID and original pre-circuit document data corresponding to each first pre-circuit report document fragment vector.
[0049] The knowledge base is constructed by first splitting the original pre-inspection document data into a list of text blocks (background information document fragments, inspection content document fragments, risk and issue discovery document fragments, rectification suggestion document fragments, and public opinion summary document fragments). Then, text embedding technology is used to extract vector representations of each text block, forming a vector list which is stored in a vector database, while simultaneously associating it with the corresponding original pre-inspection document data and metadata. The generated pre-inspection report document fragment vector data and the original pre-inspection document data are linked using IDs or metadata. Thus, the complete storage system consists of two parts: a vector database and an original pre-inspection document database. The original pre-inspection document database uses unique identifiers. Specifically, before feeding each text block (pre-inspection report document fragment, e.g., a paragraph or an article) into the text embedding model (vector model) to generate vectors, a globally unique ID is generated for each text block. The sequence (unique ID, pre-inspection report document fragment, generated first pre-inspection report document fragment vector) is treated as a complete record. This facilitates storing the first text description vector and the first pre-inspection report document fragment vector output by the adjusted vector model into the vector database after the vector model is adjusted.
[0050] The completed knowledge base is primarily based on the Retrieval Augmentation Generation (RAG) framework. Retrieval: Given a query vector (a vector generated from the semantic description of user requirements, i.e., a vector generated from the labeled text description), quickly find the k most similar first-round report document fragment vectors from the vast vector knowledge base. Calculation methods can include cosine similarity, dot product, Euclidean distance, etc. Contextualization: Before embedding the text into blocks, surrounding text blocks can be concatenated using IDs or metadata to enhance the context. For example, left-right concatenation: the previous segment + the current section + the next segment. The correlation between different text blocks of the same report can be established using unique IDs or metadata.
[0051] In step S3, the user requirements, k fragments of the first pre-trip report document, and optional original pre-trip document data are taken as input, and the pre-trip report is taken as output to construct a second training dataset. The training and adjustment of the large model based on the second training dataset specifically includes: S31, take user requirements, k first pre-circuit report document fragments and optional original pre-circuit document data as input; wherein, the optional original pre-circuit document data are the original pre-circuit document data corresponding to each first pre-circuit report document fragment vector in the k first pre-circuit report document fragments, and other original pre-circuit document data that form a context with the original pre-circuit document data corresponding to each first pre-circuit report document fragment vector. S32, with the pre-inspection report including at least risk point discovery, rectification suggestions and public opinion summary as output, construct a second training dataset, and train and adjust the large model based on the second training dataset; wherein, the public opinion summary is expressed in the form of a triplet of event subject, event behavior and impact result.
[0052] In the second training dataset, the results of fine-tuning the vector model are combined with the original data to generate training samples for pre-inspection reports. These samples include at least risk point discovery data, suggestion generation data (which forms a corresponding question-and-answer relationship with the risk point discovery), and public opinion information extraction data. The public opinion information extraction data is generated by extracting triples ([Event Subject] - [Event Behavior] - [Impact Result]) and labeled data. Specifically, the triple format could be "[A Company] - [Accounts Receivable Over 1 Year Aging] - [Cash Flow Tight Risk]". Among them, positive samples can be text based on user needs, k fragments of the first pre-inspection report document, and optional original pre-inspection document data; pre-inspection reports that are manually compiled and include at least risk point discovery, rectification suggestions, and public opinion summary can be used as positive samples. Negative sample data can be pre-inspection reports that are randomly generated based on user needs, k fragments of the first pre-inspection report document, and optional original pre-inspection document data and include at least risk point discovery, rectification suggestions, and public opinion summary, or pre-inspection reports that are generated by ordinary large models and include at least risk point discovery, rectification suggestions, and public opinion summary (and may also include background information, inspection content, etc.) can be used as negative samples, until a preset number of training samples are formed.
[0053] Specifically, in the large model training and adjustment, the update method for the weight coefficients is as follows: W new = W original Learning rate Gradient Among them, W new For adjusting the updated weight parameters during the training of a large model, W originalWeight parameters loaded for pre-training large models, learning rate The learning rate is adjusted during training of a large model, where Gradient is the gradient of the loss function with respect to each weight parameter of the large model, calculated via backpropagation; and the learning rate is... rate Specifically: learning rate = , in, The initial learning rate, To control the decay rate, To reduce the learning rate to The number of steps at / 2, where t is the current training step during the large model training adjustment.
[0054] The core form of large model parameter updates is achieved through a large model fine-tuning process, the specific process of which can be summarized as follows: First, initialize the model: load all parameters of the pre-trained large model (i.e., the weight matrix W). original Next, forward propagation is performed: new training data (e.g., sentence pairs "poor battery life" and "turn on power saving mode") is input into the large model to obtain the corresponding predicted output. Then, the loss is calculated: the predicted output of the large model is compared with the true labels (e.g., whether two sentences are similar), and the loss value (which can be the loss function of the original large model or the contrastive loss function in the vector model in this embodiment) is calculated to quantify the error of the model on the current task. Subsequently, backpropagation and parameter updates are performed: the gradient of the loss function with respect to each model parameter is calculated using the backpropagation algorithm. This gradient represents how each parameter should be fine-tuned to improve the model's performance on the new data. Finally, the iterative optimization process is performed: using an optimizer (e.g., AdamW) with a low learning rate, all parameters are updated in the opposite direction of the gradient. This process is run iteratively for multiple training epochs on the new data to achieve gradual optimization of the parameters.
[0055] The learning rate is dynamically adjusted based on: training progress (epochs / steps), with a smaller learning rate in later stages; validation set performance, automatically reducing the learning rate (LR) if the validation set performance no longer improves; and task difficulty or data distribution—a larger learning rate can be used to accelerate convergence when data differences are large; a smaller learning rate is needed for detailed optimization. In the early stages, a larger learning rate (e.g., 5e-4) is used, allowing the large model to quickly learn the format and terminology of pre-inspection reports. In the mid-stages, the learning rate gradually decreases, allowing the model to more stably learn key information such as "equipment name, inspection time, and risk warnings." In the later stages, a smaller learning rate (e.g., 1e-5) is used to prevent the large model from excessively modifying existing knowledge; only minor adjustments to terminology and industry standards are made to the original large model. Dynamic learning rate adjustment can also be set as follows: early stages: 5e-4; mid-stages: 1e-4; late stages: 1e-5.
[0056] Large-scale model updates primarily utilize annotation platforms (such as Label Studio and Prodigy) to allow annotators to label sentences according to predefined schemas. The annotation results can be sequence labels (BIO format), triples, or JSON, or suitable publicly available evaluation datasets can be downloaded and used directly.
[0057] Large model fine-tuning: Fine-tuning is performed based on the prepared data. Through training, the large model is enabled to generate pre-trip reports that meet actual needs. Model selection requires a comprehensive consideration of the performance metrics required by the task and the computational resource overhead associated with the model size.
[0058] Fine-tuning primarily targets two types of objects: one is full fine-tuning of all parameters of a large model, which optimizes the model by updating the weights of each layer. This method usually achieves the best performance, but it has extremely high requirements for computing power and memory, and requires a powerful GPU cluster to complete. The other is partial parameter fine-tuning (i.e., parameter efficient fine-tuning, PEFT), which has become the mainstream strategy. Its basic idea is to keep the original model parameters in a "frozen" state and not update them, while only introducing and training a small number of new parameters. The most representative PEFT method is LoRA.
[0059] In essence, fine-tuning belongs to the supervised learning paradigm. It requires pre-constructing a dataset consisting of "input-output" paired samples, and then allowing a large model to learn the mapping relationship between them, so that it can adapt to the knowledge structure and target behavior style required for a specific task.
[0060] Large-scale model optimization: Optimizing the large language model to adjust the quality and contextual consistency of the generated text. This includes using additional training data and adjusting generation parameters to improve model performance. On one hand, there's training-level optimization: additional data (e.g., continued training / fine-tuning / instruction fine-tuning / RLHF) allows the large model to "learn more." On the other hand, there's inference-level optimization: adjusting generation parameters (e.g., temperature, top-k, top-p, length control) and enhancement methods (RAG, Prompt) allows the large model to "use better." Adjustment rationale: Insufficient business domain-specific knowledge; inspection industry terminology, equipment models, and regulatory clauses may not be adequately covered in the pre-training corpus. Examples include "high-voltage electrical equipment insulation testing process" and "fire lane inspection specifications." Task format requirements: pre-inspection reports require structured content and professional expression: risk points, recommendations, and public opinion information. General large models may generate free text that does not conform to report specifications. Accuracy and reliability of generated content: generated reports must cite regulatory clauses or historical records, and cannot fabricate facts. Reliable content is generated based on retrieved vector data.
[0061] The requirements are to be met in three aspects: Professionalism / Accuracy: Risk points, recommendations, and regulatory citations must be correct. Evaluation metrics can include retrieval coverage, human review accuracy, and information extraction accuracy. Readability / Fluency: The text must be natural and readable. Evaluation metrics can include BLEU, ROUGE, and human scoring. Retrieval Information Utilization: The model must appropriately use retrieved fragments. Evaluation metrics can include hit rate / coverage or a reinforcement learning reward model.
[0062] In step S5, the semantic description in the current user requirement is used as the text description. Based on the similarity between the current text description vector corresponding to the current user requirement and the first pre-circuit report document fragment vector in the knowledge base, the top-k second pre-circuit report document fragments with the highest similarity ranking are obtained.
[0063] In step S6, the current user requirements, k second pre-inspection report document fragments, and optional original pre-inspection document data are used as input to the adjusted large model. The output is a complete pre-inspection report including risk point discovery, rectification suggestions, and public opinion summary. The optional original pre-inspection document data can be the original pre-inspection document data corresponding to each second pre-inspection report document vector in the k second pre-inspection report document fragments, as well as other original pre-inspection document data that form the context with the original pre-inspection document data corresponding to each second pre-inspection report document vector. The context can be enhanced by concatenating surrounding text blocks (pre-inspection report document fragments). For example, concatenating left and right, the previous paragraph + the current end + the next paragraph. The correlation between different text blocks in the same report can be achieved using intelligent text segmentation, using IDs, or using metadata for association.
[0064] It should be noted that the large model in this solution is the qwen-72B series open source model, but other open source large models can also be used. This embodiment does not impose any restrictions here.
[0065] A fine-tuned Large Language Model (LLM) is used to generate pre-inspection reports. The model receives preprocessed data and contextual information to generate structured report content. In the pre-inspection report scenario, the vector model is responsible for encoding text such as pre-inspection report document fragments into first pre-inspection report document fragment vectors, generating a corresponding current text description vector query for the current user requirement, and retrieving the most relevant top-k second pre-inspection report document fragments by calculating similarity. The retrieved second pre-inspection report document fragments, along with user requirement instructions and optional original pre-inspection document data, are used as input to the Large Language Model (LLM) to generate a complete pre-inspection report containing risk points, rectification suggestions, and a public opinion summary, ensuring that the content is professional, accurate, and compliant with regulations.
[0066] The adjusted large model can also perform post-processing on the generated report, including formatting, language correction, and information completion, to ensure the accuracy and readability of the report; and output the final pre-trip report to a specified format (such as PDF, Word, etc.) and store it in the system database for subsequent retrieval and review.
[0067] like Figure 4 As shown, the implementation structure of the technical solution of the present invention mainly includes four stages: 1. Training data preparation stage (step S1): (1) Data collection: collect pre-inspection document data from various data sources, including inspection records, historical reports and expert opinions in different structural forms.
[0068] (2) Document segmentation or splitting: The original pre-inspection document data is split to generate a list of text blocks (background information document fragments, inspection content document fragments, risk problem discovery document fragments, rectification suggestion document fragments, and public opinion summary document fragments).
[0069] (3) Text extraction: The segmented pre-patrol report document fragments are labeled with corresponding text descriptions.
[0070] (4) Manual verification: The text description is manually verified. The main rule is to confirm that the text description is relevant to the original text, so as to ensure the accuracy and quality of the problem.
[0071] 2. Vector model fine-tuning stage (step S2): (1) Data preparation: The pre-trip report document fragments corresponding to the annotated text descriptions are used as positive samples in the training data.
[0072] (2) Negative sample mining: Negative sample data is generated using negative sample mining methods. These negative samples refer to data fragments that are unrelated to or have low relevance to the target task. Specifically, the pre-trip report document fragments and randomly generated text descriptions in the first training dataset, or pre-trip report document fragments randomly generated from labeled text descriptions, are used as the first negative samples (ordinary negative samples).
[0073] In each training round, the second negative sample is dynamically selected using the vector model under adjustment; wherein, the similarity between the second pre-cycle report document fragment vector corresponding to the first negative sample and the second text description vector is less than the similarity between the second pre-cycle report document fragment vector corresponding to the second negative sample and the second text description vector.
[0074] (3) bge-large fine-tuning: During training, the contrastive loss function is weighted, giving higher weights to difficult negative samples. Sampling ratio control: For example, in detection, a positive:negative ratio of 1:3 is often used to avoid overwhelming positive samples with too many negative samples. Phased training: In the early stage of training, random negative samples are used for pre-training, and in the later stage of training, difficult negative samples are used for fine-tuning. The division between the early stage and the later stage of training can be based on the number of training steps, or it can be based on the contrastive loss function and the threshold of the preset target contrastive loss function. The adjustment basis is: similarity (the more similar, the higher the weight), retrieval ranking (the higher the ranking, the higher the weight), and loss size (the larger the loss, the higher the weight, similar to Focal Loss).
[0075] 3. Knowledge base construction phase (step S2): (1) Document splitting: The construction of the knowledge base first involves splitting the original pre-inspection document data to generate a list of text blocks (background information document fragments, inspection content document fragments, risk problem discovery document fragments, rectification suggestion document fragments, and public opinion summary document fragments). (2) Text embedding: Then, the vector representation of each text block is extracted through text embedding technology, forming a vector list and storing it in the vector database, while associating it with the corresponding original pre-inspection document data and metadata.
[0076] (3) Vector Storage: The generated pre-inspection report document fragment vector data and the original pre-inspection inspection document data are associated with each other using IDs or metadata. That is, the complete storage system consists of two parts: a vector database and an original pre-inspection inspection document database. The original pre-inspection inspection document database uses a unique identifier. Specifically, before the text block (pre-inspection report document fragment, for example, a paragraph or an article) is fed into the text embedding model (vector model) to generate vectors, a globally unique ID is generated for each text block. The (unique ID, pre-inspection report document fragment, generated first pre-inspection report document fragment vector) is treated as a complete record, which facilitates storing the first text description vector and the first pre-inspection report document fragment vector output by the adjusted vector model into the vector database after the vector model is adjusted.
[0077] Four major model adjustment phases (steps S3-S4): (1) Data preparation: The user requirements, k fragments of the first pre-inspection report document and optional original pre-inspection document data are used as input; the pre-inspection report, which includes at least risk point discovery, rectification suggestions and public opinion summary, is used as output to construct the second training dataset.
[0078] (2) Question-answer pair extraction: Combine the results of the vector model fine-tuning with the original data to generate training samples for the pre-inspection report, including at least risk point discovery data, suggestion generation data (which form a corresponding question-answer pair relationship with the risk point discovery), and public opinion information extraction data.
[0079] (3) Data augmentation: Positive samples can be text based on user needs, k fragments of the first pre-inspection report document, and optional original pre-inspection document data; pre-inspection reports that are manually compiled and include at least risk point discovery, rectification suggestions, and public opinion summary can be used as positive samples. Negative sample data can be pre-inspection reports that are randomly generated based on user needs, k fragments of the first pre-inspection report document, and optional original pre-inspection document data, and include at least risk point discovery, rectification suggestions, and public opinion summary, or pre-inspection reports that are generated by ordinary large models and include at least risk point discovery, rectification suggestions, and public opinion summary (and may also include background information, inspection content, etc.) as negative samples, until a preset number of training samples are formed.
[0080] (4) LoRa Fine-tuning: Large Model Fine-tuning: Fine-tuning based on prepared data. Through training, the large model is able to generate pre-trip report content that meets actual needs. The selection of the model needs to comprehensively consider the performance indicators required by the task and the computational resource overhead associated with the model size.
[0081] Fine-tuning primarily targets two types of objects: one is full fine-tuning of all parameters of a large model, which optimizes the model by updating the weights of each layer. This method usually achieves the best performance, but it has extremely high requirements for computing power and memory, and requires a powerful GPU cluster to complete. The other is partial parameter fine-tuning (i.e., parameter efficient fine-tuning, PEFT), which has become the mainstream strategy. Its basic idea is to keep the original model parameters in a "frozen" state and not update them, while only introducing and training a small number of new parameters. The most representative PEFT method is LoRA.
[0082] Knowledge base retrieval: The user question is optimized and encoded (using the same encoding method as the document) to retrieve the top k most relevant knowledge points in the knowledge base. The k knowledge data points with the highest similarity ranking are then concatenated with the user question and used as input to the first model for subsequent answer generation.
[0083] To better understand this solution, an example is provided below: The user's current requirement is "a pre-audit report for a certain company in 20xx". The search yields the second pre-audit report document fragment from the top-5 list (containing "accounts receivable turnover rate of 20xx was 3.25 (industry average 3.5)" and "Revenue recognition clause of Accounting Standard No. 14"). Context construction: The second pre-audit report document fragment from the top-5 list, the user's requirement, and the original pre-audit document data are concatenated into the adjusted large model to generate a pre-audit report containing "background information (financial overview of the company in 20xx), inspection content (accounts receivable, inventory), problem found (low accounts receivable turnover rate), rectification suggestions (aging analysis + collection), and public opinion summary (no major negative public opinion)".
[0084] This invention innovatively combines vector models and large language models and applies them to pre-trip report generation, improving the overall performance of pre-trip report generation through fine-tuning and optimization. This combined application method can effectively handle complex report generation tasks and produce high-quality report content.
[0085] Application scenarios: These can include areas or scenarios such as enterprise inspections, quality control, and compliance checks. Specifically, in enterprise inspections: Pre-inspection reports are automatically generated during internal inspections, improving inspection efficiency and report quality. In quality control: Inspection reports are automatically generated to help enterprises identify and resolve quality issues promptly. In compliance checks: Reports are automatically generated during compliance checks to ensure enterprises comply with regulatory requirements, improving compliance and management levels.
[0086] This invention establishes a first training dataset based on pre-inspection report document data. A vector model is trained and adjusted using this dataset. Semantic descriptions in user requirements are used as text descriptions. The k most similar first pre-inspection report document fragments are obtained based on the similarity between the first text description vector corresponding to the user requirements and the first pre-inspection report document fragment vectors in the knowledge base. A second training dataset is constructed using the user requirements, the k first pre-inspection report document fragments, and optional original pre-inspection document data as input, with the pre-inspection report as output. The large model is then trained and adjusted based on this second training dataset. Finally, the current user requirements, the k second pre-inspection report document fragments, and optional original pre-inspection document data are used as input to the adjusted large model, outputting the pre-inspection report. This effectively solves the problem of balancing accuracy and training cost in the pre-inspection report generation process caused by existing technologies, improving both the accuracy of pre-inspection report generation and reducing training costs.
[0087] The technical solution of this invention collects pre-inspection document data from various data sources, including inspection records, historical reports, and expert opinions in different structural forms. It supports multiple data types, including structured, semi-structured, and unstructured data, thus ensuring the accuracy of the generated pre-inspection report. Furthermore, the pre-processed pre-inspection document data is segmented according to different parts of the report content. These segmented pre-inspection report document fragments include background information fragments, inspection content fragments, risk and problem discovery fragments, rectification suggestion fragments, and public opinion summary fragments. The obtained pre-inspection report document fragments are associated one-to-one with the original pre-inspection document data through metadata or IDs. This facilitates the selection of the original pre-inspection document data when training and adjusting the large model later, ensuring the correlation and correspondence between the pre-inspection report document fragments, pre-inspection report document fragment vectors, and the original pre-inspection document data during the pre-inspection report generation process.
[0088] In this invention, the labeled pre-circuit report document fragments and labeled text descriptions in the first training dataset are used as positive samples; the pre-circuit report document fragments and randomly generated text descriptions in the first training dataset, or pre-circuit report document fragments randomly generated from labeled text descriptions, are used as first negative samples; in each training round, a second negative sample is dynamically selected using the vector model under adjustment; the vector model is pre-trained using the first negative sample in the early training stage, and adjusted using the second negative sample in the later training stage, so that the adjusted vector model can better adapt to the generation of pre-circuit reports and has higher accuracy.
[0089] In the technical solution of this invention, when the first pre-travel report document fragment vector is stored in a pre-established knowledge base, the ID or metadata, the pre-travel report document fragment, and the generated first pre-travel report document fragment vector are stored as a complete record in the vector database. When performing a similarity search based on the current text description vector corresponding to the current user's needs, the k most similar first pre-travel report document fragment vectors are returned, along with the unique ID and original pre-travel document data corresponding to each first pre-travel report document fragment vector. The original pre-travel document data can be the original pre-travel document data corresponding to each first pre-travel report document fragment vector in the k first pre-travel report document fragments, as well as other original pre-travel document data that form a context with the original pre-travel document data corresponding to each first pre-travel report document fragment vector. This ensures the correlation and correspondence between the pre-travel report document fragments, the pre-travel report document fragment vectors, and the original pre-travel document data during the pre-travel report generation process.
[0090] Example 2 like Figure 5 As shown, the present invention also provides a pre-circuit report generation system based on a vector model, comprising: Module 101 acquires pre-inspection report document data and establishes a first training dataset; the first training dataset includes pre-inspection report document fragments and text descriptions. The first training module 102 trains and adjusts the vector model based on the first training dataset; and obtains the first text description vector and the first pre-circuit report document fragment vector based on the adjusted vector model, and stores them in a pre-established knowledge base. The first module 103 uses the semantic description in the user's requirements as the text description, and obtains the k first pre-circuit report document fragments with the highest similarity ranking based on the similarity between the first text description vector corresponding to the user's requirements and the first pre-circuit report document fragment vector in the knowledge base. The second training module 104 takes user requirements, k fragments of the first pre-cycle report document, and optional original pre-cycle document data as input, and takes the pre-cycle report as output to construct a second training dataset, and trains and adjusts the large model based on the second training dataset; wherein, the large model is a large language model. The second module 105 uses the semantic description in the current user demand as the text description, and obtains the k second pre-circuit report document fragments with the highest similarity ranking based on the similarity between the current text description vector corresponding to the current user demand and the first pre-circuit report document fragment vector in the knowledge base. Output module 106 takes the current user requirements, k second pre-circuit report document fragments, and optional original pre-circuit document data as input to the adjusted large model and outputs the pre-circuit report.
[0091] It should be noted that the implementation methods of the acquisition module 101, the first training module 102, the first obtaining module 103, the second training module 104, the second obtaining module 105, and the output module 106 in this embodiment are the same as the method steps in Embodiment 1, and will not be repeated here.
[0092] This invention establishes a first training dataset based on pre-inspection report document data. A vector model is trained and adjusted using this dataset. Semantic descriptions in user requirements are used as text descriptions. The k most similar first pre-inspection report document fragments are obtained based on the similarity between the first text description vector corresponding to the user requirements and the first pre-inspection report document fragment vectors in the knowledge base. A second training dataset is constructed using the user requirements, the k first pre-inspection report document fragments, and optional original pre-inspection document data as input, with the pre-inspection report as output. The large model is then trained and adjusted based on this second training dataset. Finally, the current user requirements, the k second pre-inspection report document fragments, and optional original pre-inspection document data are used as input to the adjusted large model, outputting the pre-inspection report. This effectively solves the problem of balancing accuracy and training cost in the pre-inspection report generation process caused by existing technologies, improving both the accuracy of pre-inspection report generation and reducing training costs.
[0093] The technical solution of this invention collects pre-inspection document data from various data sources, including inspection records, historical reports, and expert opinions in different structural forms. It supports multiple data types, including structured, semi-structured, and unstructured data, thus ensuring the accuracy of the generated pre-inspection report. Furthermore, the pre-processed pre-inspection document data is segmented according to different parts of the report content. These segmented pre-inspection report document fragments include background information fragments, inspection content fragments, risk and problem discovery fragments, rectification suggestion fragments, and public opinion summary fragments. The obtained pre-inspection report document fragments are associated one-to-one with the original pre-inspection document data through metadata or IDs. This facilitates the selection of the original pre-inspection document data when training and adjusting the large model later, ensuring the correlation and correspondence between the pre-inspection report document fragments, pre-inspection report document fragment vectors, and the original pre-inspection document data during the pre-inspection report generation process.
[0094] In this invention, the labeled pre-circuit report document fragments and labeled text descriptions in the first training dataset are used as positive samples; the pre-circuit report document fragments and randomly generated text descriptions in the first training dataset, or pre-circuit report document fragments randomly generated from labeled text descriptions, are used as first negative samples; in each training round, a second negative sample is dynamically selected using the vector model under adjustment; the vector model is pre-trained using the first negative sample in the early training stage, and adjusted using the second negative sample in the later training stage, so that the adjusted vector model can better adapt to the generation of pre-circuit reports and has higher accuracy.
[0095] In the technical solution of this invention, when the first pre-travel report document fragment vector is stored in a pre-established knowledge base, the ID or metadata, the pre-travel report document fragment, and the generated first pre-travel report document fragment vector are stored as a complete record in the vector database. When performing a similarity search based on the current text description vector corresponding to the current user's needs, the k most similar first pre-travel report document fragment vectors are returned, along with the unique ID and original pre-travel document data corresponding to each first pre-travel report document fragment vector. The original pre-travel document data can be the original pre-travel document data corresponding to each first pre-travel report document fragment vector in the k first pre-travel report document fragments, as well as other original pre-travel document data that form a context with the original pre-travel document data corresponding to each first pre-travel report document fragment vector. This ensures the correlation and correspondence between the pre-travel report document fragments, the pre-travel report document fragment vectors, and the original pre-travel document data during the pre-travel report generation process.
[0096] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.
Claims
1. A method for generating pre-trip reports based on a vector model, characterized in that, include: Obtain pre-inspection report document data and establish a first training dataset; the first training dataset includes pre-inspection report document fragments and text descriptions; specifically, obtaining pre-inspection report document data and establishing the first training dataset involves: Pre-inspection document data is collected from various data sources, including inspection records, historical reports, and expert opinions in different structural formats. Preprocessing of pre-inspection document data; The pre-inspection document data is segmented according to different parts of the report content. The segmented pre-inspection report document fragments include background information document fragments, inspection content document fragments, risk and problem discovery document fragments, rectification suggestion document fragments, and public opinion summary document fragments. The obtained pre-inspection report document fragments are associated with the original pre-inspection document data one by one through metadata or ID. The segmented pre-pattern report document fragments were labeled to obtain the first training dataset; Based on the first training dataset, an adjusted vector model is trained; based on the adjusted vector model, a first text description vector and a first pre-circuit report document fragment vector are obtained and stored in a pre-established knowledge base; wherein, training the adjusted vector model based on the first training dataset specifically includes: The labeled pre-patrol report document fragments and labeled text descriptions in the first training dataset are used as positive samples; The pre-process report document fragments and randomly generated text descriptions, or pre-process report document fragments randomly generated from labeled text descriptions in the first training dataset, are used as the first negative samples. In each training round, the second negative sample is dynamically selected using the vector model under adjustment; wherein, the similarity between the second pre-cycle report document fragment vector corresponding to the first negative sample and the second text description vector is less than the similarity between the second pre-cycle report document fragment vector corresponding to the second negative sample and the second text description vector. In the early stage of training, the vector model is pre-trained using the first negative sample, and in the later stage of training, the vector model is adjusted using the second negative sample. Using the semantic description in the user's needs as the text description, the k first pre-circuit report document fragments with the highest similarity ranking are obtained based on the similarity between the first text description vector corresponding to the user's needs and the first pre-circuit report document fragment vector in the knowledge base. Using user requirements, k first pre-circuit report document fragments, and optional original pre-circuit document data as input, and pre-circuit reports as output, a second training dataset is constructed, and the large model is trained and adjusted based on the second training dataset; wherein, the large model is a large language model; specifically, it includes: using user requirements, k first pre-circuit report document fragments, and optional original pre-circuit document data as input; wherein, the optional original pre-circuit document data consists of the original pre-circuit document data corresponding to each first pre-circuit report document fragment vector in the k first pre-circuit report document fragments, and other original pre-circuit document data that form the context with the original pre-circuit document data corresponding to each first pre-circuit report document fragment vector; The second training dataset is constructed using a pre-inspection report that includes at least risk point identification, rectification suggestions, and public opinion summary as output. The large model is then trained and adjusted based on the second training dataset. The public opinion summary is expressed in the form of a triplet consisting of event subject, event behavior, and impact result. Using the semantic description in the current user needs as the text description, based on the similarity between the current text description vector corresponding to the current user needs and the first pre-circuit report document fragment vector in the knowledge base, the k second pre-circuit report document fragments with the highest similarity ranking are obtained. The current user requirements, k fragments of the second pre-circuit report document, and optional original pre-circuit document data are used as inputs to the adjusted large model, and the pre-circuit report is output.
2. The method for generating pre-circuit reports based on a vector model according to claim 1, characterized in that, During the training process of the vector model, the contrastive loss function is as follows: , Where L is the contrastive loss function, q is the text description vector, p is the pre-pass report document vector in the positive samples, and τ is the temperature coefficient for training the vector model. Let be the cosine similarity between the text description vector and the pre-patrol report document vector in the positive samples, where j is the j-th first negative sample or second negative sample, k1 is the sum of the first negative sample and the second negative sample, and n is the cosine similarity between the two vectors. j This refers to the pre-circuit report document vector in the j-th first negative sample or second negative sample. The weight is the weight corresponding to the j-th first negative sample or the second negative sample.
3. The method for generating pre-circuit reports based on a vector model according to claim 2, characterized in that, If the negative sample is the first negative sample, then The value is 1; if the negative sample is the second negative sample, then Greater than 1.
4. The method for generating pre-circuit reports based on a vector model according to claim 1, characterized in that, Based on the adjusted vector model, the first text description vector and the first pre-circuit report document fragment vector are obtained and stored in a pre-established knowledge base as follows: Based on the adjusted vector model, the first text description vector and the first pre-circuit report document fragment vector are obtained, and stored in a pre-established knowledge base respectively. When storing the first pre-circuit report document fragment vector in the pre-established knowledge base, the ID or metadata, the pre-circuit report document fragment, and the generated first pre-circuit report document fragment vector are stored as a complete record in the vector database. When performing a similarity search based on the current text description vector corresponding to the current user's needs, the k most similar first pre-circuit report document fragment vectors are returned, along with the ID and original pre-circuit document data corresponding to each first pre-circuit report document fragment vector.
5. The method for generating pre-circuit reports based on a vector model according to claim 1, characterized in that, In large model training and adjustment, the weight coefficients are updated in the following way: W new = W original Learning rate Gradient, Among them, W new To adjust the updated weight parameters during the training of a large model, W original Weight parameters loaded for pre-training large models, Learning rate The learning rate is adjusted during training of a large model, where Gradient is the gradient of the loss function calculated via backpropagation with respect to each weight parameter of the large model; and the learning rate is... rate Specifically: Learning rate = , in, The initial learning rate, To control the decay rate, To reduce the learning rate to The number of steps at / 2, where t is the current training step number during the large model training adjustment.
6. A method for generating pre-trip reports based on a vector model according to any one of claims 1-5, characterized in that, The main model is the qwen-72B series of open-source models.
7. A pre-trip report generation system based on a vector model, characterized in that, include: The acquisition module acquires pre-inspection report document data and establishes a first training dataset. The first training dataset includes pre-inspection report document fragments and text descriptions. Specifically, acquiring pre-inspection report document data and establishing the first training dataset involves: Pre-inspection document data is collected from various data sources, including inspection records, historical reports, and expert opinions in different structural formats. Preprocessing of pre-inspection document data; The pre-inspection document data is segmented according to different parts of the report content. The segmented pre-inspection report document fragments include background information document fragments, inspection content document fragments, risk and problem discovery document fragments, rectification suggestion document fragments, and public opinion summary document fragments. The obtained pre-inspection report document fragments are associated with the original pre-inspection document data one by one through metadata or ID. The segmented pre-pattern report document fragments were labeled to obtain the first training dataset; The first training module trains and adjusts the vector model based on the first training dataset; based on the adjusted vector model, it obtains the first text description vector and the first pre-circuit report document fragment vector, and stores them in a pre-established knowledge base; wherein, training and adjusting the vector model based on the first training dataset specifically includes: The labeled pre-patrol report document fragments and labeled text descriptions in the first training dataset are used as positive samples; The pre-process report document fragments and randomly generated text descriptions, or pre-process report document fragments randomly generated from labeled text descriptions in the first training dataset, are used as the first negative samples. In each training round, the second negative sample is dynamically selected using the vector model under adjustment; wherein, the similarity between the second pre-cycle report document fragment vector corresponding to the first negative sample and the second text description vector is less than the similarity between the second pre-cycle report document fragment vector corresponding to the second negative sample and the second text description vector. In the early stage of training, the vector model is pre-trained using the first negative sample, and in the later stage of training, the vector model is adjusted using the second negative sample. The first module uses the semantic description in the user's requirements as the text description. Based on the similarity between the first text description vector corresponding to the user's requirements and the first pre-circuit report document fragment vector in the knowledge base, it obtains the k first pre-circuit report document fragments with the highest similarity ranking. The second training module takes user requirements, k first pre-circuit report document fragments, and optional original pre-circuit document data as input, and takes the pre-circuit report as output to construct a second training dataset, and trains and adjusts the large model based on the second training dataset; wherein, the large model is a large language model; specifically, it includes: taking user requirements, k first pre-circuit report document fragments, and optional original pre-circuit document data as input; wherein, the optional original pre-circuit document data consists of the original pre-circuit document data corresponding to each first pre-circuit report document fragment vector in the k first pre-circuit report document fragments, and other original pre-circuit document data that form the context with the original pre-circuit document data corresponding to each first pre-circuit report document fragment vector; The second training dataset is constructed using a pre-inspection report that includes at least risk point identification, rectification suggestions, and public opinion summary as output. The large model is then trained and adjusted based on the second training dataset. The public opinion summary is expressed in the form of a triplet consisting of event subject, event behavior, and impact result. The second module uses the semantic description in the current user demand as the text description, and obtains the k second pre-circuit report document fragments with the highest similarity ranking based on the similarity between the current text description vector corresponding to the current user demand and the first pre-circuit report document fragment vector in the knowledge base. The output module takes the current user requirements, k fragments of the second pre-circuit report document, and optional original pre-circuit document data as input to the adjusted large model, and outputs the pre-circuit report.
Citation Information
Patent Citations
Intelligent document duplicate checking system and method based on vector database and large language model
CN120373283A
Computer-generated content based on text classification, semantic relevance, and activation of deep learning large language models
US11748577B1