Aviation Multi-Domain Data Adaptive Extraction Method and System Based on Large Models

Through the adaptive extraction method of aviation multi-domain data based on large models, the problems of low intelligence in traditional aviation data retrieval and neglected semantic information are solved, and more accurate and rich search results are achieved.

CN119474344BActive Publication Date: 2025-05-30CHINA AERO POLYTECH ESTAB
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411703323.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-26
Publication Date
2025-05-30
Estimated Expiration
2044-11-26

AI Technical Summary

Technical Problem

Traditional aviation data retrieval methods have problems such as low intelligence, search results deviating from user expectations, and semantic information being ignored.

Method used

Adaptive extraction method of aviation multi-domain data based on large models is adopted, and semantic feature extraction and comparison are combined with keyword extraction and semantic extraction results to generate more accurate and rich search results.

Benefits of technology

It realizes automatic selection of search fields based on user query problems, automatically generating search weights and statements, and improves the accuracy and richness of search results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119474344B_ABST
    Figure CN119474344B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for adaptively extracting multi-domain data in the aviation field based on a large model, which relates to the technical field of data retrieval. The method includes: obtaining a multi-domain dataset in the aviation field and performing preprocessing to obtain a preprocessed dataset; constructing a plurality of inverted index tables based on the preprocessed dataset; extracting semantic feature vectors for each paragraph of each text in the preprocessed dataset based on the trained BGE model to obtain a semantic feature vector library; extracting semantic features for the text input by the user based on the trained BGE model to obtain an input feature vector; calculating a cosine similarity set and a morpheme similarity set based on the large model; and performing normalization processing on the morpheme similarity set to obtain a normalized value set; fusing and sorting the normalized value set and the cosine similarity set to obtain an extraction result. The present invention fuses and sorts the keyword extraction result and the semantic extraction result simultaneously to ensure the richness of extraction and the accuracy of sorting.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data retrieval, and particularly to a method and system for adaptively extracting multi-domain aviation data based on a large model. Background Art

[0002] During the use of an aviation knowledge service platform, a key requirement of users is to retrieve and use various types of aviation data. Due to different query purposes and data types of users, the fixed retrieval statements and weights used in traditional retrieval methods are difficult to meet the retrieval experience and requirements of users. The problems mainly include the following aspects:

[0003] (1) During the user's retrieval process, it is necessary to manually select a specific field for retrieval, with a low level of intelligence;

[0004] (2) The fixed retrieval statements and retrieval field weights are inconsistent with the user's retrieval purpose, resulting in the retrieval results deviating from the user's expectations;

[0005] (3) Traditional keyword retrieval only focuses on character matching, ignoring semantic information and resulting in inaccurate retrieval results. Summary of the Invention

[0006] The purpose of the present invention is to provide a method and system for adaptively extracting multi-domain aviation data based on a large model, which performs semantic feature extraction and comparison on the retrieval content and text information, and at the same time fuses and sorts the keyword extraction and semantic extraction results to ensure the richness of extraction and the accuracy of sorting.

[0007] A method for adaptively extracting multi-domain aviation data based on a large model, which includes:

[0008] Obtain a multi-domain dataset in the aviation field, and preprocess the multi-source dataset in the aviation field to obtain a preprocessed dataset;

[0009] Construct a plurality of inverted index tables based on the preprocessed dataset; each of the inverted index tables corresponds to a field; the inverted index table includes Json key-value pairs and a plurality of morphemes, and each morpheme corresponds to a text in the preprocessed dataset; the Json key-value pairs include an index name, index data content, and usage;

[0010] Split a plurality of texts in the preprocessed dataset into sentences, and divide each sentence based on a large model to obtain a first training set, and perform unsupervised training on the BGE model based on the first training set to obtain the initially trained BGE model;

[0011] Use the paragraphs of several texts in the preprocessed dataset as the second training set, and ask questions about the paragraphs of several texts in the preprocessed dataset based on a large model to obtain a long and short text annotation dataset;

[0012] Use the long and short text annotation dataset as the label set and combine it with the second training set to perform supervised training on the initially trained BGE model to obtain the trained BGE model;

[0013] Extract semantic feature vectors for each paragraph of each text in the preprocessed dataset based on the trained BGE model to obtain a semantic feature vector library;

[0014] Extract semantic features of the text input by the user based on the trained BGE model to obtain an input feature vector;

[0015] Calculate the cosine similarity between the input feature vector and each semantic feature vector in the semantic feature vector library based on a large model to obtain an initial cosine similarity set;

[0016] Analyze the text input by the user based on a large model using the Json key-value pairs of each inverted index table to obtain a demand inverted index table; the demand inverted index table is the inverted index table corresponding to the field of the text input by the user;

[0017] Based on each morpheme in the demand inverted index table, use a large model to form a query statement and the weights of each query field in the query statement according to the text input by the user;

[0018] Calculate the similarity between the query statement and each morpheme in the demand inverted index table based on a large model and the weights of each query field in the query statement to obtain a morpheme similarity set; and normalize the morpheme similarity set to obtain an initial normalization value set;

[0019] If the normalization value in the initial normalization value set and the cosine similarity in the initial cosine similarity set correspond to the same text in the preprocessed dataset, use this text as the current text, perform weighted summation on the corresponding normalization value and cosine similarity pair of the current text to obtain a fusion value, and delete the cosine similarity corresponding to the current text from the initial cosine similarity set and delete the normalization value corresponding to the current text from the initial normalization value set, traverse the initial normalization value set and the initial cosine similarity set to obtain a normalization value set, a cosine similarity set and a fusion value set; the fusion value set includes several of the fusion values;

[0020] The fusion values ​​in the fusion value set, the normalized values ​​in the normalized value set, and the cosine similarities in the cosine similarity set are sorted from large to small, and corresponding texts are given as extraction results according to the required quantity.

[0021] Optionally, constructing an inverted index table based on the preprocessed data set includes:

[0022] Extract aviation-related terms from the Air Force Aviation Engineering Dictionary, the Commercial Aircraft Professional Terminology Dictionary, the China Aviation Encyclopedia Dictionary, and aviation standards to obtain a term list;

[0023] Splitting the preprocessed data set into words according to the term list and the search engine's own word list to obtain a plurality of morphemes;

[0024] The morpheme is used as the key value, and the text corresponding to the morpheme is used as the value value to construct a key-value pair to obtain the inverted index table.

[0025] Optionally, the cosine similarity expression is:

[0026]

[0027] Where A is the input feature vector, B is the semantic feature vector, n is the dimension of the feature vector, A i is the value of the i-th dimension of the input feature vector, B i Represents the value of the i-th dimension of the semantic feature vector, Sim cos AB Represents the cosine similarity between the input feature vector and the semantic feature vector.

[0028] Optionally, the morpheme similarity expression is:

[0029]

[0030]

[0031] Where: w i Represents a query statement, w j represents the jth morpheme in the demand inverted index table, j = [1,,2,…,N], N is the total number of morphemes in the demand inverted index table, n is the total number of query fields in the query statement, and w ik represents the kth query field in the query statement, IDF(w ik ) indicates w ik The weight of f(w ik ,w j ) indicates that w j Medium ik The number of occurrences, k 1a is the first adjustment factor, b is the second adjustment factor, |w j | represents the length of the text corresponding to w j , and avgdl represents the average length of all texts in the demand inverted index table.

[0032] Optionally, the fusion value expression is:

[0033] score final = W * sigmoid(score bm25 , c, a) + d * score cosine ;

[0034]

[0035] In the formula: score final represents the fusion value, W is the weight of the normalization value, d is the weight of the cosine similarity, c is the first parameter of the sigmoid function, a is the second parameter of the sigmoid function, score cosine is the cosine similarity, score bm25 is the morpheme similarity, sigmoid(score bm25 , c, a) represents the normalized value of the morpheme similarity.

[0036] Optionally, the loss function in the unsupervised training and supervised training processes is:

[0037]

[0038] In the formula: F is the loss function value, p and q represent text pairs, is the negative sample, τ is the BGE model temperature coefficient, e p is the feature vector of p, e q is the feature vector of q, is 's feature vector, ∑ (p,q) is the sum of the loss function values of all text pairs, min·Σ (p,q) is the minimum value of the sum of the loss function values of all text pairs.

[0039] Optionally, the large model is Qwen2.5 - 72B.

[0040] The present invention also provides an aviation multi - source data adaptive extraction system based on a large model, which includes:

[0041] A data acquisition and processing module, configured to acquire multi - domain data sets in the aviation field and pre - process the aviation field multi - source data sets to obtain pre - processed data sets;

[0042] An index table module for constructing a number of inverted index tables based on the preprocessed dataset; each of the inverted index tables corresponds to a field; the inverted index table includes Json key-value pairs and a number of morphemes, and each morpheme corresponds to a text in the preprocessed dataset; the Json key-value pairs include an index name, index data content, and usage;

[0043] A first training module for splitting a number of texts in the preprocessed dataset into sentences, dividing each sentence based on a large model to obtain a first training set, and performing unsupervised training on the BGE model based on the first training set to obtain the initially trained BGE model;

[0044] A labeling module for using the paragraphs of a number of texts in the preprocessed dataset as a second training set, and asking questions about the paragraphs of a number of texts in the preprocessed dataset based on a large model to obtain a long and short text labeling dataset;

[0045] A second training module for using the long and short text labeling dataset as a label set and combining it with the second training set to perform supervised training on the initially trained BGE model to obtain the trained BGE model;

[0046] A vector library module for extracting semantic feature vectors for each paragraph of each text in the preprocessed dataset based on the trained BGE model to obtain a semantic feature vector library;

[0047] A feature extraction module for extracting semantic features of the text input by the user based on the trained BGE model to obtain an input feature vector;

[0048] A cosine similarity module for calculating the cosine similarity between the input feature vector and each semantic feature vector in the semantic feature vector library based on a large model to obtain an initial cosine similarity set;

[0049] A requirement index table module for analyzing the text input by the user based on the Json key-value pairs of each of the inverted index tables using a large model to obtain a requirement inverted index table; the requirement inverted index table is the inverted index table corresponding to the field corresponding to the text input by the user;

[0050] A weight module for forming a query statement and weights for each query field in the query statement based on the morphemes in the requirement inverted index table using a large model according to the text input by the user;

[0051] A morpheme similarity module, which is used to calculate the similarity between the query statement and each morpheme in the demand inverted index table based on the large model and the weights of each query field in the query statement, so as to obtain a morpheme similarity set; and normalize the morpheme similarity set to obtain an initial normalized value set;

[0052] A fusion module, which is used to, if the normalized value in the initial normalized value set and the cosine similarity in the initial cosine similarity set correspond to the same text in the preprocessed dataset, take this text as the current text, perform weighted summation on the corresponding normalized value and the cosine similarity pair of the current text to obtain a fusion value, delete the cosine similarity corresponding to the current text from the initial cosine similarity set, delete the normalized value corresponding to the current text from the initial normalized value set, traverse the initial normalized value set and the initial cosine similarity set to obtain a normalized value set, a cosine similarity set and a fusion value set; the fusion value set includes a number of the fusion values;

[0053] A text extraction module, which is used to sort each of the fusion values in the fusion value set, each of the normalized values in the normalized value set and each of the cosine similarities in the cosine similarity set from large to small and give the corresponding text as the extraction result according to the required quantity.

[0054] The effects of the present invention are as follows:

[0055] The method for adaptively extracting aviation multi-domain data based on a large model of the present invention can automatically select the retrieval domain according to the user's query problem.

[0056] The method for adaptively extracting aviation multi-domain data based on a large model of the present invention automatically generates retrieval weights and retrieval statements according to the relationship between the user's query problem and the retrieval fields.

[0057] The method for adaptively extracting aviation multi-domain data based on a large model of the present invention performs semantic feature extraction and comparison on the retrieval content and text information, and at the same time fuses and sorts the keyword extraction and semantic extraction results to ensure the richness of extraction and the accuracy of sorting. Description of the Drawings

[0058] Figure 1 is the flowchart of the method for adaptively extracting aviation multi-domain data based on a large model of the present invention. Detailed Embodiments

[0059] Hereinafter, the embodiments of the present invention will be described with reference to the drawings.

[0060] Figure 1 is the flowchart of the method for adaptively extracting aviation multi-domain data based on a large model of the present invention. As Figure 1As shown, the present invention provides a method for adaptively extracting aviation multi-domain data based on a large model, which includes:

[0061] S1, obtain multi-domain datasets in the aviation field, and preprocess the multi-source datasets in the aviation field to obtain preprocessed datasets.

[0062] Specifically, the papers in the field of aviation are downloaded and the metadata related to the papers are obtained, including title, author, publication date, keywords and abstract. Since most of the data obtained are PDF files, it is necessary to perform text recognition and layout analysis on some files to form chapter titles and the text corresponding to the chapter titles to obtain the preprocessed data set.

[0063] S2, build several inverted index tables based on the preprocessed data set. Each inverted index table corresponds to a field. The inverted index table includes a Json key-value pair and several morphemes, each of which corresponds to a text in the preprocessed data set. The Json key-value pair includes the index name, index data content and purpose.

[0064] Preferably, S2 includes:

[0065] S21, extract aviation-related terms from the Air Force Aviation Engineering Dictionary, the Commercial Aircraft Professional Terminology Dictionary, the China Aviation Encyclopedia Dictionary and aviation standards to obtain a term list.

[0066] S22, splitting the preprocessed data set into words according to the term list and the search engine's own word list to obtain a number of morphemes.

[0067] S23, taking the morpheme as the key value and the text corresponding to the morpheme as the value value to construct a key-value pair, and obtain an inverted index table.

[0068] S3, splitting several texts in the preprocessed data set into sentences, and dividing each sentence based on the large model to obtain a first training set, and performing unsupervised training on the BGE model based on the first training set to obtain an initially trained BGE model. In this embodiment, the large model is a large language model, preferably Qwen2.5-72B.

[0069] S4, taking several text paragraphs in the preprocessed data set as the second training set, and asking questions to several text paragraphs in the preprocessed data set based on the large model to obtain a long and short text annotation data set.

[0070] S5, using the long and short text annotation dataset as a label set and combining it with the second training set to perform supervised training on the initially trained BGE model to obtain a trained BGE model.

[0071] The loss function during unsupervised training and supervised training is:

[0072]

[0073] Where: F is the value of the loss function, p and q represent text pairs, is a negative sample, τ is the temperature coefficient of the BGE model, e p is the feature vector of p, e q is the feature vector of q, is 's feature vector, Σ (p,q) is the sum of the loss function values of all text pairs, min·Σ (p,q) is the minimum value of the sum of the loss function values of all text pairs.

[0074] S6. Based on the trained BGE model, semantic feature extraction is performed on each paragraph of each text in the preprocessed dataset to obtain a semantic feature vector library.

[0075] S7. Based on the trained BGE model, semantic feature extraction is performed on the text input by the user to obtain an input feature vector. In the present invention, the BGE model is used to fine-tune the multi-source data corpus in the aviation field, making the vector extraction and sorting accuracy higher.

[0076] S8. Based on the large model, the cosine similarity is calculated between the input feature vector and each semantic feature vector in the semantic feature vector library to obtain an initial cosine similarity set. The BGE model is a fine-tuned semantic model used to fine-tune the text input by the user so that the length of the semantic feature vector is consistent with the input feature vector.

[0077] The cosine similarity expression is:

[0078]

[0079] Where: A is the input feature vector, B is the semantic feature vector, n is the dimension of the feature vector, A i is the value of the i-th dimension of the input feature vector, B i represents the value of the i-th dimension of the semantic feature vector, Sim cos AB represents the cosine similarity between the input feature vector and the semantic feature vector.

[0080] S9. Based on the Json key-value pairs of each inverted index table, the large model is used to analyze the text input by the user to obtain a demand inverted index table. The demand inverted index table is the inverted index table corresponding to the field of the text input by the user.

[0081] S10. Based on each morpheme in the demand inverted index table, the large model is used to form a query statement and the weights of each query field in the query statement according to the text input by the user.

[0082] S11. Calculate the similarity between the query statement and each morpheme in the demand inverted index table based on the large model and the weights of each query field in the query statement, obtaining a set of morpheme similarities; and normalize the set of morpheme similarities to obtain an initial set of normalized values.

[0083] The morpheme similarity expression is:

[0084]

[0085]

[0086] In the formula: w i represents the query statement, w j represents the j-th morpheme in the demand inverted index table, j = [1, 2,..., N], N is the total number of morphemes in the demand inverted index table, n is the total number of query fields in the query statement, w ik represents the k-th query field in the query statement, IDF(w ik ) represents the weight of w ik , f(w ik , w j ) represents the number of occurrences of w j in w ik , k 1 is the first adjustment factor, b is the second adjustment factor, |w j | represents the length of the text corresponding to w j , and avgdl represents the average length of all texts in the demand inverted index table.

[0087] S12. If the normalized value in the initial set of normalized values and the cosine similarity in the initial set of cosine similarities correspond to the same text in the preprocessed dataset, take this text as the current text, perform a weighted sum on the normalized value and the cosine similarity pair corresponding to the current text to obtain a fusion value, and delete the cosine similarity corresponding to the current text from the initial set of cosine similarities and delete the normalized value corresponding to the current text from the initial set of normalized values. Traverse the initial set of normalized values and the initial set of cosine similarities to obtain a set of normalized values, a set of cosine similarities, and a set of fusion values. The set of fusion values includes several fusion values.

[0088] Specifically, the fusion value expression is:

[0089] score final = W * sigmoid(score bm25 , c, a) + d * score cosine ;

[0090]

[0091] where: score final represents the fusion value, W is the weight of the normalized value, d is the weight of the cosine similarity, c is the first parameter of the sigmoid function, a is the second parameter of the sigmoid function, and score cosine is the cosine similarity, score bm25 is the morpheme similarity, and sigmoid(score bm25 , c, a) represents the normalized value of the morpheme similarity.

[0092] S13. Sort each fusion value in the fusion value set, each normalized value in the normalized value set, and each cosine similarity in the cosine similarity set from largest to smallest, and give the corresponding text as the extraction result according to the required quantity.

[0093] Specifically, when the user inputs: What are the applicable standards for turbines, compressors, fans, and turbocharger rotors, it corresponds to the aviation standard field, and the demand inverted index table is shown in Table 1.

[0094] Table 1 Inverted Index Table

[0095] Index Document Engine 4 Rotor 15、13、5 Compressor 14、3、6 Fan 13、7、9 … …

[0096] The input feature vector is a 1024-dimensional feature vector. The cosine similarities obtained in documents 1, 8, and 9 are 0.7981214340483506, 0.7781211345683576, and 0.7651324545987272 respectively. The scores of the morpheme similarities obtained in documents 1, 7, and 9 are 74, 52, and 41 respectively. The values obtained after sigmoid normalization are 0.999997617294756, 0.999929166980283, and 0.9992376924918762 respectively. When the weights of both the morpheme similarity and the cosine similarity are 1, the final scores of documents 1, 7, 8, and 9 are: 1.7981190513431065, 0.999929166980283, 0.7781211345683576, and 1.7643701470906035 respectively. The final sorting result is documents 1, 9, 7, 8.

[0097] Preferably, display the metadata corresponding to the text in the extraction result according to actual needs.

[0098] The method of the present invention is implemented based on the Elasticsearch search engine.

[0099] The present invention also provides an aviation multi-source data adaptive extraction system based on a large model, which includes:

[0100] A data acquisition and processing module for acquiring multi-domain datasets in the aviation field and preprocessing the multi-source datasets in the aviation field to obtain preprocessed datasets.

[0101] An index table module for constructing a number of inverted index tables based on the preprocessed datasets. Each inverted index table corresponds to a domain. The inverted index table includes Json key-value pairs and a number of morphemes, and each morpheme corresponds to a text in the preprocessed dataset. The Json key-value pairs include index names, index data contents, and uses.

[0102] A first training module for splitting a number of texts in the preprocessed dataset into sentences, partitioning each sentence based on a large model to obtain a first training set, and performing unsupervised training on the BGE model based on the first training set to obtain an initially trained BGE model.

[0103] A labeling module for using the paragraphs of a number of texts in the preprocessed dataset as a second training set and asking questions about the paragraphs of a number of texts in the preprocessed dataset based on a large model to obtain a long and short text labeling dataset.

[0104] A second training module for using the long and short text labeling dataset as a label set and combining it with the second training set to perform supervised training on the initially trained BGE model to obtain a trained BGE model.

[0105] A vector library module for extracting semantic feature vectors for each paragraph of each text in the preprocessed dataset based on the trained BGE model to obtain a semantic feature vector library.

[0106] A feature extraction module for extracting semantic features from the text input by the user based on the trained BGE model to obtain input feature vectors.

[0107] A cosine similarity module for calculating the cosine similarity between the input feature vectors and each semantic feature vector in the semantic feature vector library based on a large model to obtain an initial cosine similarity set.

[0108] A requirement index table module for analyzing the text input by the user based on the Json key-value pairs of each inverted index table using a large model to obtain a requirement inverted index table. The requirement inverted index table is the inverted index table corresponding to the domain corresponding to the text input by the user.

[0109] A weight module for forming query statements and the weights of each query field in the query statements according to the text input by the user based on each morpheme in the requirement inverted index table using a large model.

[0110] A morpheme similarity module, which is used to calculate the similarity between the query statement and each morpheme in the demand inverted index table based on the large model and the weights of each query field in the query statement, so as to obtain a set of morpheme similarities; and normalize the set of morpheme similarities to obtain an initial set of normalized values.

[0111] A fusion module, which is used to, if the normalized value in the initial set of normalized values and the cosine similarity in the initial set of cosine similarities correspond to the same text in the preprocessed dataset, take this text as the current text, perform a weighted sum of the normalized value and the cosine similarity pair corresponding to the current text to obtain a fusion value, and delete the cosine similarity corresponding to the current text from the initial set of cosine similarities and delete the normalized value corresponding to the current text from the initial set of normalized values, and traverse the initial set of normalized values and the initial set of cosine similarities to obtain a set of normalized values, a set of cosine similarities, and a set of fusion values. The set of fusion values includes several fusion values.

[0112] A text extraction module, which is used to sort each fusion value in the set of fusion values, each normalized value in the set of normalized values, and each cosine similarity in the set of cosine similarities from largest to smallest and give the corresponding text as the extraction result according to the required quantity.

[0113] The embodiments described above are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.

Claims

1. A method for adaptively extracting aviation multi-domain data based on a large model, characterized in that: It includes: Acquire a multi-domain dataset in the aviation field, and preprocess the multi-domain dataset in the aviation field to obtain a preprocessed dataset; Constructing a plurality of inverted index tables based on the preprocessed data set; Each of the inverted index tables corresponds to a field; the inverted index table includes a Json key-value pair and a plurality of morphemes, each of which corresponds to a text in the preprocessed data set; the Json key-value pair includes an index name, index data content and purpose; Splitting a number of texts in the preprocessed data set into sentences, and dividing each sentence based on the large model to obtain a first training set, and performing unsupervised training on the BGE model based on the first training set to obtain the initially trained BGE model; Using several text paragraphs in the preprocessed data set as a second training set, and asking questions to the several text paragraphs in the preprocessed data set based on the large model to obtain a long and short text annotated data set; Using the long and short text annotation dataset as a label set and combining it with the second training set to perform supervised training on the initially trained BGE model to obtain the trained BGE model; Extracting semantic features from each paragraph of each text in the preprocessed data set based on the trained BGE model to obtain a semantic feature vector library; Extracting semantic features of the text input by the user based on the trained BGE model to obtain an input feature vector; Based on the large model, cosine similarity calculation is performed on the input feature vector and each semantic feature vector in the semantic feature vector library to obtain an initial cosine similarity set; Based on the Json key-value pairs of each inverted index table, a large model is used to analyze the text input by the user to obtain a demand inverted index table; The demand inverted index table is an inverted index table corresponding to the field corresponding to the text input by the user; Based on each morpheme in the demand inverted index table, a large model is used to form a query statement and a weight of each query field in the query statement according to the text input by the user; Based on the large model and the weight of each query field in the query statement, similarity calculation is performed between the query statement and each morpheme in the demand inverted index table to obtain a morpheme similarity set; and normalizing the morpheme similarity set to obtain an initial normalized value set; If the normalized value in the initial normalized value set and the cosine similarity in the initial cosine similarity set correspond to the same text in the preprocessed data set, take this text as the current text, perform weighted summation on the normalized value and the cosine similarity corresponding to the current text to obtain a fusion value, delete the cosine similarity corresponding to the current text from the initial cosine similarity set, delete the normalized value corresponding to the current text from the initial normalized value set, traverse the initial normalized value set and the initial cosine similarity set, and obtain a normalized value set, a cosine similarity set, and a fusion value set; The fusion value set includes a plurality of fusion values; The fusion values ​​in the fusion value set, the normalized values ​​in the normalized value set, and the cosine similarities in the cosine similarity set are sorted from large to small, and corresponding texts are given as extraction results according to the required quantity.

2. The method for adaptively extracting aviation multi-domain data based on a large model according to claim 1 is characterized in that: Constructing an inverted index table based on the preprocessed data set includes: Extract aviation-related terms from the Air Force Aviation Engineering Dictionary, the Commercial Aircraft Professional Terminology Dictionary, the China Aviation Encyclopedia Dictionary, and aviation standards to obtain a term list; Splitting the preprocessed data set into words according to the term list and the search engine's own word list to obtain a plurality of morphemes; The morpheme is used as the key value, and the text corresponding to the morpheme is used as the value value to construct a key-value pair to obtain the inverted index table.

3. The method for adaptively extracting aviation multi-domain data based on a large model according to claim 1 is characterized in that: The cosine similarity expression is: Where A is the input feature vector, B is the semantic feature vector, n is the dimension of the feature vector, A i is the value of the i-th dimension of the input feature vector, B i Represents the value of the i-th dimension of the semantic feature vector, Sim cos AB Represents the cosine similarity between the input feature vector and the semantic feature vector.

4. The method for adaptively extracting aviation multi-domain data based on a large model according to claim 1 is characterized in that: The morpheme similarity expression is: Where: w i Represents a query statement, w j represents the jth morpheme in the demand inverted index table, j = [1,,2,…,N], N is the total number of morphemes in the demand inverted index table, n is the total number of query fields in the query statement, and w ik represents the kth query field in the query statement, IDF(w ik ) indicates w ik The weight of f(w ik ,w j ) indicates w j Medium ik The number of occurrences, k1 is the first adjustment factor, b is the second adjustment factor, |w j | indicates w j The length of the corresponding text, avgdl represents the average length of all texts in the required inverted index table.

5. The method for adaptively extracting aviation multi-domain data based on a large model according to claim 1, characterized in that: The fusion value expression is: score final =W*sigmoid(score bm25 ,c,a)+d*score cosine ; Where: score final represents the fusion value, W is the weight of the normalized value, d represents the weight of the cosine similarity, c is the first parameter of the simoid function, a is the second parameter of the simoid function, and score cosine is the cosine similarity, score bm25 is the morpheme similarity, sigmoid(score bm25 ,c,a) represents the normalized value of morpheme similarity.

6. The method for adaptively extracting aviation multi-domain data based on a large model according to claim 1, characterized in that: The loss function during unsupervised training and supervised training is: Where: F is the loss function value, p and q represent text pairs, is a negative sample, τ is the BGE model temperature coefficient, e p is the eigenvector of p, e q is the eigenvector of q, yes The characteristic vector of (p,q) is the sum of the loss function values ​​of all text pairs, min·∑ (p,q) It is the minimum value of the sum of loss function values ​​of all text pairs.

7. The method for adaptively extracting aviation multi-domain data based on a large model according to claim 1, characterized in that: The large model is Qwen2.5-72B.

8. An adaptive extraction system for aviation multi-domain data based on a large model, characterized in that: It includes: A data acquisition processing module, used for acquiring a multi-field data set in the aviation field, and preprocessing the multi-field data set in the aviation field to obtain a preprocessed data set; An index table module, used to construct a plurality of inverted index tables based on the preprocessed data set; each of the inverted index tables corresponds to a field; the inverted index table includes a Json key-value pair and a plurality of morphemes, each of the morphemes corresponds to a text in the preprocessed data set; the Json key-value pair includes an index name, index data content and purpose; A first training module is used to split the plurality of texts in the preprocessed data set into sentences, and divide each sentence based on the large model to obtain a first training set, and perform unsupervised training on the BGE model based on the first training set to obtain the initially trained BGE model; A labeling module, used for taking the paragraphs of several texts in the preprocessed data set as a second training set, and asking questions to the paragraphs of several texts in the preprocessed data set based on the large model to obtain a long and short text labeling data set; A second training module is used to use the long and short text annotation dataset as a label set and combine it with the second training set to perform supervised training on the initially trained BGE model to obtain the trained BGE model; A vector library module is used to extract semantic features from each paragraph of each text in the preprocessed data set based on the trained BGE model to obtain a semantic feature vector library; A feature extraction module is used to extract semantic features of the text input by the user based on the trained BGE model to obtain an input feature vector; A cosine similarity module, used for calculating the cosine similarity between the input feature vector and each semantic feature vector in the semantic feature vector library based on a large model to obtain an initial cosine similarity set; The demand index table module is used to analyze the text input by the user using a large model based on the Json key-value pairs of each inverted index table to obtain a demand inverted index table; The demand inverted index table is an inverted index table corresponding to the field corresponding to the text input by the user; A weight module, for forming a query statement and a weight of each query field in the query statement using a large model according to the text input by the user based on each morpheme in the demand inverted index table; A morpheme similarity module, used to calculate the similarity between the query statement and each morpheme in the demand inverted index table based on the large model and the weight of each query field in the query statement, so as to obtain a morpheme similarity set; and normalizing the morpheme similarity set to obtain an initial normalized value set; A fusion module, for, if the normalized value in the initial normalized value set and the cosine similarity in the initial cosine similarity set correspond to the same text in the preprocessed data set, taking the text as the current text, performing weighted summation on the normalized value and the cosine similarity corresponding to the current text to obtain a fusion value, deleting the cosine similarity corresponding to the current text from the initial cosine similarity set, deleting the normalized value corresponding to the current text from the initial normalized value set, traversing the initial normalized value set and the initial cosine similarity set to obtain a normalized value set, a cosine similarity set and a fusion value set; the fusion value set includes a plurality of the fusion values; The text extraction module is used to sort the fusion values ​​in the fusion value set, the normalized values ​​in the normalized value set and the cosine similarities in the cosine similarity set from large to small and provide corresponding texts as extraction results according to the required quantity.

Citation Information

Patent Citations

  • Aviation system knowledge graph construction method based on fusion and semi-supervised information extraction

    CN116127090A

  • Aviation literature keyword similarity judgment method fused with multi-modal semantic association map

    CN116362221A