Training data set construction method and device based on large model and RAG and related components

Through the method based on large models and RAG, an efficient, accurate and time-sensitive training data set is constructed, which solves the problems of low efficiency, high cost and poor quality of existing data set construction methods, and achieves better performance in question-and-answer applications in specific fields.

CN120196715APending Publication Date: 2025-06-24SHENZHEN ALL THINGS CLOUD TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510261877.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-24

Smart Images

  • Figure CN120196715A_ABST
    Figure CN120196715A_ABST
Patent Text Reader

Abstract

The invention discloses a training data set construction method and device based on a large model and RAG and related components. The method comprises the steps that input instruction information is acquired; performing cue word engineering on the input instruction information to obtain engineering instruction information; analyzing the engineering instruction information and extracting a task keyword; the task keywords are input into a large language model for generalization processing, and generalization instruction information is obtained; returning the generalization instruction information to the user, and performing multi-round supplementation and correction according to feedback information of the user to obtain final instruction information; performing fragmentation and vectorization processing on the final instruction information to obtain vector information; performing retrieval matching on the vector information through a vector database to obtain matching information; and screening from the vector database according to the matching information to obtain a training data set. According to the method, the target knowledge is recalled by using the retrieval capability of the RAG, and then the data set is generated by using the large language model and engineering processing, so that the construction speed and accuracy are improved, and the cost is also reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly relates to a method, apparatus, and related components for constructing a training data set based on a large model and RAG. Background Art

[0002] In recent years, significant progress has been made in the field of artificial intelligence, and the rise of large language models has been particularly remarkable. Through pre-training on massive amounts of data, these models have accumulated profound language knowledge and powerful generation and understanding capabilities, bringing revolutionary changes to the field of natural language processing (NLP). However, although large language models perform well in general tasks, their accuracy and information integrity are still insufficient in question-and-answer applications in specific fields. This is mainly because large models that solely rely on pre-training data are difficult to fully capture professional terms, industry backgrounds, and the latest developments within a specific field, resulting in a deviation between the question-and-answer results and actual requirements.

[0003] To overcome this challenge, researchers have proposed the Retrieval-Augmented Generation (RAG) technology. The RAG technology combines an information retrieval module with a generation model, aiming to enhance the accuracy and relevance of the generation model by retrieving relevant documents from an external database. Under the RAG framework, the system first uses an efficient retrieval algorithm to screen out highly relevant documents from a vast knowledge base, and then inputs these documents as context information into the generation model. This process can not only effectively improve the accuracy and relevance of the question-and-answer results but also reduce the dependence on training data to a certain extent, making the model more flexible and adaptable.

[0004] However, the RAG technology still faces some challenges in practical applications. Since the performance of the retrieval algorithm and the quality of the database directly affect the overall performance of the RAG system, how to ensure the accuracy and timeliness of the retrieval results has become a key issue. In addition, even if relevant documents are retrieved, how to effectively combine these documents with the generation model to generate answers that are both accurate and meet the user's needs is also a major difficulty in current research.

[0005] In terms of data set construction, traditional methods such as manual annotation and web scraping can meet the training requirements to a certain extent, but they often suffer from problems such as low efficiency, high cost, and unstable data quality. Although manual annotation can ensure the high quality and accuracy of data, the labor cost is high and it takes a long time; while web scraping can quickly obtain a large amount of data, but the data quality and relevance are often difficult to guarantee.

[0006] On the other hand, although using open-source datasets can save certain time and costs, open-source datasets usually contain a large amount of general data, which deviate from the requirements of specific fields. In addition, the data in open-source datasets often need to be strictly screened and cleaned before being used to train models, and this process also requires a large amount of manpower and material resources. Summary of the Invention

[0007] The object of the present invention is to provide a method, device and related components for constructing a training dataset based on a large model and RAG, aiming to solve problems such as low construction efficiency, high cost and poor quality of existing datasets.

[0008] In a first aspect, an embodiment of the present invention provides a method for constructing a training dataset based on a large model and RAG, including:

[0009] Obtain the input instruction information of the user, where the input instruction information includes text information or language information;

[0010] Perform prompt engineering on the input instruction information to obtain engineering instruction information;

[0011] Parse the engineering instruction information and extract task keywords therefrom;

[0012] Input the task keywords into a large language model for generalization processing to obtain generalized instruction information;

[0013] Return the generalized instruction information to the user and perform multi-round supplementation and correction according to the user's feedback information to obtain the final instruction information;

[0014] Perform sharding and vectorization processing on the final instruction information to obtain vector information;

[0015] Retrieve and match the vector information through a vector database to obtain matching information;

[0016] Screen according to the matching information from the vector database to obtain a training dataset, where the training dataset includes a picture training dataset and a video training dataset.

[0017] In a second aspect, an embodiment of the present invention provides a device for constructing a training dataset based on a large model and RAG, including:

[0018] An acquisition unit, configured to acquire the input instruction information of the user, where the input instruction information includes text information or language information;

[0019] An engineering unit, configured to perform prompt engineering on the input instruction information to obtain engineering instruction information;

[0020] A parsing unit for parsing the engineering instruction information and extracting task keywords therefrom;

[0021] A generalization unit for inputting the task keywords into a large language model for generalization processing to obtain generalized instruction information;

[0022] A correction unit for returning the generalized instruction information to the user and performing multi-round supplementation and correction according to the user's feedback information to obtain final instruction information;

[0023] A sharding unit for sharding and vectorizing the final instruction information to obtain vector information;

[0024] A matching unit for retrieving and matching the vector information through a vector database to obtain matching information;

[0025] A screening unit for screening from the vector database according to the matching information to obtain a training data set, where the training data set includes an image training data set and a video training data set.

[0026] In a third aspect, an embodiment of the present invention further provides a computer device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method for constructing a training data set based on a large model and RAG described in the first aspect above is implemented.

[0027] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for constructing a training data set based on a large model and RAG described in the first aspect above is implemented.

[0028] The present invention discloses a method, apparatus, and related components for constructing a training dataset based on a large model and RAG. The method includes: obtaining input instruction information of a user; performing prompt engineering on the input instruction information to obtain engineered instruction information; parsing the engineered instruction information and extracting task keywords therefrom; inputting the task keywords into a large language model for generalization processing to obtain generalized instruction information; returning the generalized instruction information to the user and performing multiple rounds of supplementation and correction based on the user's feedback information to obtain final instruction information; performing sharding and vectorization processing on the final instruction information to obtain vector information; performing retrieval matching on the vector information through a vector database to obtain matching information; and screening from the vector database according to the matching information to obtain a training dataset. By using the retrieval ability of RAG to recall target knowledge and then using a large language model and engineering processing to generate a dataset, the present invention significantly improves the construction speed and accuracy of the dataset. Secondly, by dynamically retrieving the latest information, the timeliness and relevance of the dataset are ensured. In addition, it reduces the dependence on manual participation, reduces the introduction of human biases, and also reduces costs. Finally, by combining with an external knowledge base, RAG helps the large language model to be customized and optimized in a specific field, enabling it to handle more complex and professional tasks. Embodiments of the present invention also provide a device for constructing a training dataset based on a large model and RAG, a computer-readable storage medium, and a computer device, which have the above beneficial effects and will not be elaborated herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0030] Figure 1 It is a schematic flowchart of a method for constructing a training dataset based on a large model and RAG;

[0031] Figure 2 It is a schematic sub - flowchart of a method for constructing a training dataset based on a large model and RAG;

[0032] Figure 3 It is another schematic sub - flowchart of a method for constructing a training dataset based on a large model and RAG;

[0033] Figure 4 It is a schematic block diagram of a device for constructing a training dataset based on a large model and RAG. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0034] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0035] It should be understood that when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0036] It should also be understood that the terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in this specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.

[0037] It should be further understood that the term " / and" used in this specification of the present invention and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0038] Please refer to Figures 1-3 , this embodiment provides a method for constructing a training dataset based on a large model and RAG, including:

[0039] S101: Obtain the input instruction information of the user, where the input instruction information includes text information or language information;

[0040] In this embodiment, the input instruction information includes one or more of a task type, a scenario limitation, a dataset format requirement, a label included in the scenario, the format of the label (Chinese / English), the ratio division of the training set and the test set, and the picture resolution.

[0041] For example, when the user inputs an instruction of "help me generate a YOLO format dataset for community pet object detection" through the platform. This instruction clearly indicates a specific detection task, that is, an object detection task, limits the application scenario to the community pet scenario, and at the same time determines the dataset format requirement to be the YOLO format.

[0042] Next, the user has the function of adding more requirements in the input instruction. For example, the user can specify the tags included in the scenario, such as clearly indicating the pet category tags to be included in the community pet scenario, like specific pet species tags such as cats and dogs. At the same time, the user can require the format of the tags, and can choose Chinese format tags, such as "cat", "dog", or English format tags, such as "cat", "dog".

[0043] In addition, the user can set the ratio division of the training set and the test set in the instruction. For example, set the proportion of the training set to 80% and the proportion of the test set to 20% to meet the different data volume requirements for subsequent model training and testing. Moreover, the user can also require the picture resolution. For example, specify the picture resolution as 1920×1080 to ensure that the dataset pictures meet specific quality and size standards, so as to better adapt to the subsequent object detection model training and application process.

[0044] S102: Perform prompt engineering on the input instruction information to obtain engineering instruction information;

[0045] In this embodiment, before performing prompt engineering on the input instruction information to obtain engineering instruction information, it includes: collecting picture information and device information of the upstream data system; wherein, the picture information includes Chinese and English label information, annotation information, picture-text pair information, and picture description information, and the device information includes regional scenario information and device label information; then perform data collation and standardization processing on the picture information and device information to obtain standardized information; then perform vectorization processing on the standardized information to obtain vector information; and then store the vector information in the vector database.

[0046] By collecting the picture information and device information of the upstream data system, the comprehensiveness and integrity of the data are ensured. Performing data collation and standardization processing on the collected picture information and device information can effectively eliminate redundancy, errors, and inconsistencies in the data, and improve the accuracy and usability of the data. Vectorized data can perform operations such as similarity calculation, classification, and clustering faster, thereby improving the efficiency and accuracy of data processing.

[0047] In some embodiments, performing data collation and standardization processing on the picture information and device information to obtain standardized information includes: clarifying the specific structure and types of the picture information and device information. Specifically, clarify the resolution, format, label information, and timestamp in the picture information, and clarify the device type, device parameters, and geographical location information in the device information, etc., and perform data collation and standardization processing on the picture information and device information (such as label unification, association of regional scenarios and devices, and picture-text alignment).

[0048] More specifically, taking image information as an example, it receives upstream data. Then, during the process of data reception and storage, it tags each device. The tag specifies which targets this image contains. At the same time, it formulates the classification of regional scenarios, classifies the image or device into a specific scenario (for example: residential lobby, community road, parking awning, etc.), and then annotates the description information of the image by an annotation engineer.

[0049] Furthermore, the standardized information is vectorized to obtain vector information including:

[0050] A dictionary is directly established for the Chinese and English label information, annotation information, regional scenario information, and device label information after data collation and standardization processing for vectorization, obtaining the first vector information;

[0051] After text chunking, keyword extraction, and relationship extraction on the text-image pair information and image description information after data collation and standardization processing, vectorization is performed to obtain the second vector information.

[0052] Specifically, in this embodiment, the Chinese and English label information, annotation information, regional scenario information, and device label information are vectorized by directly establishing a dictionary. This method avoids complex feature engineering and can quickly convert text information into numerical vectors, facilitating subsequent calculations and analyses. At the same time, through text chunking, long text information can be processed more effectively to avoid information loss; keyword extraction and relationship extraction can deeply mine the key information in the text and the relationships between them, making the vectorization result more accurately reflect the text content.

[0053] Specifically, the Chinese and English label information is sorted out to establish a label dictionary. For example, for the pet category label, cat / cat is stored in the dictionary as a key-value pair, with the key being cat and the value being cat; similarly, dog / dog is also stored in the form of a key-value pair, with the key being dog and the value being dog. For the scenario label, such as residential lobby / residential lobby, the key is residential lobby and the value is residential lobby, and so on. All Chinese and English label information is established into a complete dictionary in this form of key-value pairs for subsequent query and processing. Then, using a pre-trained word vector model, such as the openAi's text-embedding-ada-002 model, etc., each label in the Chinese and English label information dictionary is converted into a corresponding vector representation. For example, the vector corresponding to cat is [0.1, 0.2, 0.3,...], and the vector corresponding to cat is [0.4, 0.5, 0.6,...]. The vectors of all labels are combined in a certain order into a label vector matrix, and the vector of each label is used as a row in the matrix, thus obtaining the vectorized representation of the label information.

[0054] Then, for regional scene information such as residential lobbies, community roads, parking canopies, etc., a scene dictionary is established. The keys are the names of the scenes, and the values can be some descriptive information of the scene or the association identifiers with other relevant information. For example, taking the residential lobby as the key, the value can be a structured data containing the common features of the scene, relevant equipment information, etc., such as {common features: having a front desk, sofas, decorative paintings, etc., relevant equipment information: [entrance and exit cameras of the community, guard booth cameras, equipment room cameras,...]}. In this way, all regional scene information is established into a dictionary. Then, for the regional scene information dictionary, a similar method can be used to vectorize the scene names and descriptive information. The scene names can be represented by word vectors, and the keywords in the descriptive information can also be extracted and vectorized, and then these vectors are combined into a vector representation of the regional scene information. For example, the word vector of the residential lobby is [0.7, 0.8, 0.9,...], and the keywords in the descriptive information such as the front desk, sofas, and decorative paintings correspond to vectors [0.2, 0.3, 0.4,...], [0.5, 0.6, 0.7,...], [0.8, 0.9, 1.0,...] respectively. By performing weighted summation or other fusion operations on these vectors, a vector representation of the regional scene information is obtained. Processing all regional scene information results in a vectorized matrix of regional scene information.

[0055] Then, based on the device label information, including device type, model, location, etc., a device dictionary is established. For example, for a device with a type of surveillance camera, a model of Hikvision DS-2CD2143G0-I, and a location of Room 101, Unit 1, Building 1, XX Community, these information are combined into key-value pairs and stored in the dictionary. The key can be the unique identifier of the device, and the value is a structured data containing information such as type, model, location, etc., such as {type: surveillance camera, model: Hikvision DS-2CD2143G0-I, location: Room 101, Unit 1, Building 1, XX Community}, and a dictionary is established for all device label information. For the device label information dictionary, the information such as device type, model, and location are numerically processed. The device type can be represented by one-hot encoding or an embedding vector. The model can extract the key numbers and letters for numerical conversion, and the location information can be converted into a latitude and longitude numerical vector. For example, the device type surveillance camera is one-hot encoded as [1, 0,...] (assuming there are m device types in total, and the surveillance camera is the first type). The model Hikvision DS-2CD2143G0-I extracts the key numbers and letters and converts them into a numerical vector [1, 2, 3, 4, 5,...]. The location Room 101, Unit 1, Building 1, XX Community is converted into a latitude and longitude numerical vector [39.9042, 116.4074]. These vectors are concatenated to obtain the vector representation of the device label information, and the same processing is performed on all device label information to obtain the vectorized matrix of the device label information.

[0056] Then, the vectorized matrices of the above label information, annotation information, regional scene information, and device label information are concatenated or fused in a certain order to obtain the first vector information. For example, the vectorized matrix of the label information, the vectorized matrix of the annotation information, the vectorized matrix of the regional scene information, and the vectorized matrix of the device label information can be horizontally concatenated in sequence to form a high-dimensional first vector information matrix. This matrix contains the vectorized representations of various types of information after data collation and standardization processing, and can provide a unified and computable vector representation form for subsequent data analysis, model training, and other tasks, thereby improving the efficiency and accuracy of data processing.

[0057] Specifically, text chunking of the image-text pair information and image description information after data collation and standardization processing includes: segmenting according to the correspondence between images and texts, with natural paragraphs or semantically complete sentences as units. Then divide by description object or functional module (such as: main body description, background description, action description). Next, use a pre-trained language model (such as BERT) or a rule base (regular expression) for semantic boundary detection. Then, for extremely long texts (such as complex image descriptions), use the sliding window method (SlidingWindow) to segment, and adjust the window length according to task requirements (such as 256 tokens), with an overlap rate of 10%-20% to retain context. Immediately afterwards, for texts containing lists and tables, segment according to punctuation marks (semicolons, full stops) or indentation formats, retaining the structural information.

[0058] In this embodiment, the sliding window method can efficiently process long texts, avoid information loss, and at the same time maintain the coherence and integrity of text descriptions.

[0059] In this embodiment, prompt engineering is performed on the input instruction information to obtain engineering instruction information, including:

[0060] Identify the task type to which the natural language instruction input by the user belongs, such as text generation, question answering, translation, summarization, etc. This step can be achieved through a pre-trained classification model or rule matching.

[0061] Clarify the specific requirements of the user for the output, including but not limited to output format, length, style, language, etc. For example, the user may require generating a formal business report or a colloquial dialogue text.

[0062] Design a structured Prompt template to decompose the natural language instruction input by the user into multiple parts, such as task description, input information, output requirements, etc. For example, for a question answering task, the Prompt template can be designed as: Question: {question content}. Please answer according to the following information: {background information}. Answer requirements: {answer format}.

[0063] Then extract the key information in the natural language instruction input by the user as parameters in the Prompt template. For example, extract the question content, background information, answer format, etc., and fill them into the template to generate a specific Prompt input.

[0064] Then, according to the natural language instruction input by the user, retrieve the relevant device information. For example, if the user asks about the functions of a certain mobile phone, the system can retrieve information such as the specifications and user manuals of the mobile phone and add it as context information to the Prompt.

[0065] Then, during the process of generating the Prompt input, test different words and sentence patterns to find the expression that can most accurately convey the user's intention. For example, for the same question, different question structures can be tried, such as "Please explain...", "Can you illustrate...", etc., and observe the accuracy of the generated results.

[0066] Adjust the length of the Prompt according to the accuracy and efficiency of the generated results. An overly long Prompt may lead to redundancy in the generated results, while an overly short Prompt may lead to omission of information. Through experiments and user feedback, find the optimal Prompt length.

[0067] After completing the above steps, convert the natural language instructions input by the user into an accurate Prompt input. This step can be achieved through natural language processing technologies such as semantic understanding and syntactic analysis.

[0068] After receiving the user's instructions, use some strategies to try to match the constructed structured template (i.e., engineering instruction information) as much as possible. These strategies include, but are not limited to, semantic matching, keyword matching, template matching, etc., to ensure that the generated Prompt input can accurately reflect the user's intention.

[0069] S103: Parse the engineering instruction information and extract task keywords therefrom;

[0070] Among them, the task keywords include task type, device type, scenario, output requirements, and context information, etc.

[0071] S104: Input the task keywords into the large language model for generalization processing to obtain generalized instruction information;

[0072] In some embodiments, combine the context information of the task, the extracted keywords, and relevant information (such as: types of labels, Chinese and English of labels, image descriptions, and scenario and device-related information), and generalize other possibly relevant information through the large language model, so as to ensure that the content scope covers relevant content not explicitly mentioned by the user. Furthermore, it can ensure wide matching of different types of data to improve the accuracy of dataset construction.

[0073] S105: Return the generalized instruction information to the user and perform multiple rounds of supplementation and correction according to the user's feedback information to obtain the final instruction information;

[0074] After the large language model completes the generalization process and returns the corresponding generalization instruction information, the user can carry out subsequent operations based on the result, specifically including confirming or adjusting relevant content. Among them, the label information confirmed by the user covers multiple aspects, such as labels related to pictures (including specific content of image annotation and annotation objects, etc.), device labels (such as the scene where the device is located, detailed description of the device, etc.), and picture-related content (such as the corresponding relationship between text and pictures, specific description of pictures, etc.). The user has the right to supplement or modify the above label information according to their own needs and system feedback.

[0075] If the user has additional requirements, such as removing specific labels, adding new labels, specifying output format requirements, setting dataset division ratios, etc., the system will regenerate and return the relevant generalization instruction information according to these adjustments made by the user. During this process, the user can use the method of multi-round dialogue to continuously supplement and correct information to ensure that the finally generated instructions accurately meet the user's expected goals. At the same time, the system will also provide some template instructions for common needs for the user to select, aiming to reduce the number of times the user conducts conversations, so as to more efficiently meet the general needs of the user.

[0076] S106: Perform sharding and vectorization processing on the final instruction information to obtain vector information;

[0077] Specifically, after obtaining the final instruction information generated by the large language model, it can be tokenized or chunked, and the obtained chunk information is vectorized (vectorize label information, content description, device area, device scene), and then the vectorized information is used for the next step of vector data retrieval.

[0078] S107: Retrieve and match the vector information through a vector database to obtain matching information;

[0079] Specifically, retrieving and matching the vector information through a vector database to obtain matching information includes:

[0080] Retrieve and match the vector information in an exact matching manner to obtain the first matching information;

[0081] Then retrieve and match the vector information in a similarity matching manner (i.e., the vectorization matching in Figure 2 ) to obtain the second matching information.

[0082] Among them, when retrieving and matching the vector information in a similarity matching manner, when the data volume is extremely large and the search cost is very high, the approximate nearest neighbor search (ANN) method is used to retrieve and match the vector information.

[0083] In this embodiment, by first retrieving and matching vector information in a precise matching manner, it is possible to quickly locate matching information that is exactly the same as the query vector. This method can greatly improve the retrieval speed and reduce unnecessary computational effort when dealing with clear and specific query requirements. Then, a similarity matching method is used to retrieve and match the vector information. Similarity matching can capture the degree of similarity between vectors, thereby finding matching information similar to the query vector. This method broadens the retrieval scope and increases the possibility of finding relevant information.

[0084] Furthermore, for tag information and region information, a precise matching method can be adopted, while for image-text pair information and image description information, a similarity calculation method is used for matching.

[0085] Tag information and region information often have clear and fixed meanings and scopes. Adopting a precise matching method for these information can quickly and accurately locate results that exactly match the query. This method avoids unnecessary similarity calculations, thus greatly improving the retrieval efficiency. Image-text pair information and image description information usually contain rich visual and text content, and the similarity between these contents is often difficult to describe with simple rules or tags. Using a similarity calculation method for matching can capture the complex relationships between these information, thereby finding the results closest to the query. This flexibility enables the system to adapt to more complex and variable information matching requirements.

[0086] Among them, the similarity matching method can be matching methods such as Euclidean distance and cosine similarity.

[0087] Common ANN methods include: locality-sensitive hashing, random projection, KD-tree, and vector index structure. In the case of a large amount of data, the random projection method is adopted in this embodiment. Its core idea is: dimensionality reduction of high-dimensional data (by randomly projecting high-dimensional data into a low-dimensional space, the computational effort can be reduced), and maintaining distance relationships (the random projection method uses mathematical theories to ensure that the distances between data points can be approximately maintained to a certain extent in the low-dimensional space).

[0088] Specifically, the data point is represented as a vector x = {x1, x2, …, x n} ∈ R d , where d is the dimension of the data. Project it into a k-dimensional space R k (usually k << d). The formula for random projection is as follows:

[0089] y = Px

[0090] P ∈ R k×dis a random projection matrix, where each element is typically an independent and identically distributed random variable, usually following the standard normal distribution N(0,1);

[0091] y = {y1, y2, …, y n} ∈ R k is a vector in the low-dimensional space (the data point after dimensionality reduction);

[0092] Assume the query vector q ∈ R k , and the distance function dist(q, y) is used to measure the similarity between two vectors q and y (such as Euclidean distance, cosine similarity, etc.). The problem of vector nearest neighbor search is to find the vector y in the vector set D that minimizes the distance function m (where 1 ≤ m ≤ n), that is:

[0093]

[0094] Using this method to reduce the dimensionality of data without destroying the original data content can reduce the amount of calculation and thus improve the retrieval speed.

[0095] Next is the construction of the graph. Specifically, after fragmenting and vectorizing the final instruction information to obtain vector information, it includes: defining nodes, edges, and the dimensions of nodes according to the vector information to obtain an initial graph structure; calculating the distances between each pair of nodes in the initial graph structure, generally using Euclidean distance or cosine similarity to calculate the distances, and generating edge connections between each pair of nodes according to the calculated distances to obtain a neighbor graph (KNN); adjusting the edge connections between each pair of nodes through a greedy algorithm; randomly assigning levels to each node to form a sparse high-level structure; iteratively adjusting the number of node levels and the number of edges according to a preset threshold until the convergence condition is reached to obtain the final graph structure; storing the final graph structure in a graph database (Neo4j) or a vector database (Milvus).

[0096] In this embodiment, by calculating the distances between each pair of nodes in the initial graph structure and generating edge connections according to the distances, the degree of closeness of the relationship between nodes can be accurately described, providing a basis for subsequent graph structure optimization. At the same time, by adjusting the edge connections through a greedy algorithm, and randomly assigning levels to each node and forming a sparse high-level structure, the graph structure can be further optimized, thus ensuring the navigability and connectivity of the graph, and at the same time improving the search efficiency. Iteratively adjusting the number of node levels and the number of edges according to a preset threshold until the convergence condition is reached, this process ensures the stability and accuracy of the graph structure, avoiding overfitting or underfitting. In addition, storing the final graph structure in a graph database or a vector database facilitates subsequent rapid retrieval and analysis. The graph database and the vector database are respectively good at handling complex relationships and efficient vector matching, and can meet the requirements in different scenarios.

[0097] In this embodiment, an evaluation function needs to be defined before adjusting the edge connections between each node through the greedy algorithm. The evaluation function is used to measure the quality of the edge connections in the knowledge graph. Common evaluation metrics can be the accuracy, integrity, consistency, etc. of the graph. For example, an evaluation function F = w1A + w2C + w3U can be defined, where A represents accuracy, that is, whether the edge connection conforms to the actual knowledge relationship; C represents integrity, measuring whether the graph contains enough necessary edges; U represents consistency, checking whether there are contradictory or unreasonable edge connections in the graph. w1, w2, and w3 are the corresponding weights, which are set according to specific requirements and importance.

[0098] When it is necessary to adjust the edge connections between each node through the greedy algorithm, first select a node from the knowledge graph as the starting point. The selection can be based on factors such as the degree of the node (the number of edges connected to the node), the importance of the node, etc. For example, in a social knowledge graph, a core person node with a higher degree may be selected as the starting node because these nodes have an important impact on the structure and information dissemination of the graph.

[0099] Then, for the selected starting node, traverse all its neighbor nodes. For each neighbor node, calculate the change in the score of the edge connected to this neighbor node under the current evaluation function. Assume that the current starting node is i, the neighbor node is j, and the weight of the edge (i, j) is w ij , when considering adjusting this edge, calculate the change value ΔF of the evaluation function F after adjustment. For example, if the weight of the edge (i, j) is increased, calculate the change in F; or if this edge is deleted, calculate the change in F again.

[0100] After traversing all the operation possibilities of neighbor nodes and their related edges, select the operation that maximally improves (or minimally decreases) the evaluation function. If increasing the weight of a certain edge can maximize the value, then perform the operation of increasing the weight; if deleting a certain edge is more optimal, then delete that edge. For example, in an academic knowledge graph, if increasing the edge weight between two research field nodes can significantly improve the accuracy of the graph in reflecting the relevance of the fields and has little impact on integrity and consistency, then increase the weight of this edge.

[0101] Then, according to the selected optimal operation, update the knowledge graph accordingly. If it is to increase the edge weight, increase the weight of the edge according to the set rules; if it is to delete an edge, remove that edge from the graph. At the same time, update the attributes of related nodes and other edge-related information.

[0102] Then select the next node as the new starting node and repeat the above steps. The next node can be selected in a certain order, such as in the order of breadth-first search or depth-first search, to ensure a comprehensive adjustment of the entire knowledge graph. For example, in breadth-first search, starting from the starting node, traverse the nodes layer by layer and perform edge connection adjustment operations on each node.

[0103] In this embodiment, when the graph needs to be incrementally updated, the graph structure can be quickly restored and dynamically updated locally.

[0104] Specifically, clarify the attribute information of the new node, such as name, type, etc. At the same time, determine the relationship between the new node and the existing nodes (if any). For example, if node A and the new node B are in a "friend" relationship, then determine information such as the type and direction of the relationship.

[0105] Then use the corresponding client or programming language (such as connecting to Neo4j using the py2neo library in Python) to establish a connection with the graph storage system for data operations.

[0106] Then perform the operation of inserting a new node. Taking Neo4j as an example, the Cypher language can be used to create a new node through the CREATE statement.

[0107] Among them, if there is a relationship between the new node and the existing nodes, use the MATCH statement to find the corresponding existing nodes, and then use the CREATE statement to establish the relationship.

[0108] In this embodiment, the description information is directly vectorized, retrieved using a vector database, and the most similar description information and its associated pictures are selected for return. If more detailed structural information is desired, the graph database can be queried for return.

[0109] S108: Screen from the vector database according to the matching information to obtain a training data set, where the training data set includes a picture training data set and a video training data set.

[0110] After obtaining the matching information, in the vector database, filter the device data packet IDs that meet the user requirements according to the device-related tags and the matching information. Then, under the device data packet, further filter out the picture IDs through the image feature vectors to ensure that the image data meets the requirements of resolution, scene, annotation content, etc. of the user.

[0111] In this embodiment, screening from the vector database according to the matching information to obtain a training data set includes: assisting in screening according to the matching information through a graph database (KonwGraph) and a distributed search analysis engine (ElasticSearch).

[0112] Specifically, for ElasticSearch, the query results are returned by keyword matching. For knowledge graphs, the query is performed based on the extracted entity and relationship information, and the top N (skip) results are returned.

[0113] With the assistance of ElasticSearch and KonwGraph and the supplementation of additional information, the diversity of the returned result samples can be ensured.

[0114] Furthermore, after obtaining the image IDs, query the storage location information, image annotation information, image description information, image text-image pair information, image tag information, etc. of these images in the database. Then assemble the information in the format required by the user, such as YOLO format, COCO format, Qianwen format, etc., and upload the assembled information (i.e., the training dataset) to the object storage for preservation, and provide a download link for the user to download.

[0115] In this embodiment, after screening from the vector database according to the matching information to obtain the training dataset, it includes: obtaining the recalled data and inputting the recalled data into the vector retrieval model for tuning.

[0116] Tuning can further improve the retrieval speed and matching accuracy, and ensure the screening of high-quality and highly relevant datasets.

[0117] Among them, for some data, the operation personnel need to select according to the optimization goal for targeted optimization.

[0118] In this embodiment, by using the retrieval ability of RAG to recall the target knowledge, and then using the large language model and engineering processing to generate the dataset, the construction speed and accuracy of the dataset are significantly improved. Secondly, by dynamically retrieving the latest information, the timeliness and relevance of the dataset are ensured. In addition, the dependence on manual participation is reduced, the introduction of human bias is reduced, and the cost is also reduced. Finally, by combining with external knowledge bases, RAG helps the large language model to be customized and optimized in specific fields, enabling it to handle more complex and professional tasks.

[0119] Please refer to Figure 4 , this embodiment provides a training dataset construction device 400 based on a large model and RAG, including:

[0120] An acquisition unit 401, configured to acquire the input instruction information of the user, where the input instruction information includes text information or language information;

[0121] An engineering unit 402, configured to perform prompt engineering on the input instruction information to obtain engineering instruction information;

[0122] The parsing unit 403 is configured to parse the engineering instruction information and extract task keywords therefrom;

[0123] The generalization unit 404 is configured to input the task keywords into a deep learning model for generalization processing to obtain generalized instruction information;

[0124] The correction unit 405 is configured to return the generalized instruction information to the user and perform multi-round supplementation and correction according to the user's feedback information to obtain the final instruction information;

[0125] The sharding unit 406 is configured to perform sharding and vectorization processing on the final instruction information to obtain vector information;

[0126] The matching unit 407 is configured to retrieve and match the vector information through a vector database to obtain matching information;

[0127] The screening unit 408 is configured to screen from the vector database according to the matching information to obtain a training data set, and the training data set includes a picture training data set and a video training data set.

[0128] Further, it further includes:

[0129] The collection unit is configured to collect picture information and device information of an upstream data system; wherein, the picture information includes Chinese and English label information, annotation information, picture-text pair information, and picture description information, and the device information includes regional scene information and device label information;

[0130] The standardization unit is configured to perform data collation and standardization processing on the picture information and device information to obtain standardized information;

[0131] The vectorization unit is configured to perform vectorization processing on the standardized information to obtain vector information;

[0132] The storage unit is configured to store the vector information into the vector database.

[0133] Further, the vectorization unit includes:

[0134] The first vector information acquisition sub-unit is configured to perform dictionary direct vectorization processing on the Chinese and English label information, annotation information, regional scene information, and device label information after data collation and standardization processing to obtain first vector information;

[0135] The second vector information acquisition sub-unit is configured to perform text chunking, keyword extraction, and relationship extraction on the picture-text pair information and picture description information after data collation and standardization processing, and then perform vectorization processing to obtain second vector information.

[0136] Further, the matching unit 407 includes:

[0137] A first matching information acquisition subunit, configured to retrieve and match the vector information in an exact matching manner to obtain first matching information;

[0138] A second matching information acquisition subunit, configured to retrieve and match the vector information in a similarity matching manner to obtain second matching information.

[0139] Further, the sharding unit 406 includes:

[0140] A definition subunit, configured to define nodes, edges, and the dimensions of the nodes according to the vector information to obtain an initial graph structure;

[0141] A distance calculation subunit, configured to calculate the distances between the respective nodes in the initial graph structure, and generate edge connections between the respective nodes according to the calculated distances to obtain a neighbor graph;

[0142] An adjustment subunit, configured to adjust the edge connections between the respective nodes through a greedy algorithm;

[0143] An assignment subunit, configured to randomly assign levels to each node to form a sparse high-level structure;

[0144] An iteration subunit, configured to iteratively adjust the number of node layers and the number of edges according to a preset threshold until a convergence condition is reached to obtain a final graph structure;

[0145] A storage subunit, configured to store the final graph structure in a graph database or a vector database.

[0146] Further, the screening unit 408 includes:

[0147] An auxiliary screening subunit, configured to perform auxiliary screening according to the matching information through a graph database and a distributed search analysis engine.

[0148] Further, the screening unit 408 further includes:

[0149] An optimization subunit, configured to obtain recall data and input the recall data into a vector retrieval model for optimization.

[0150] Those skilled in the art can clearly understand that for the sake of convenience and brevity of description, the specific working processes of the above-mentioned devices and units can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0151] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed, the methods provided in the above embodiments can be implemented. The storage medium may include: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0152] The present invention also provides a computer device, which may include a memory and a processor. A computer program is stored in the memory. When the processor calls the computer program in the memory, the methods provided in the above embodiments can be implemented. Of course, the computer device may also include various network interfaces, power supplies, and other components.

[0153] The various embodiments in the specification are described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method section. It should be noted that for those of ordinary skill in the art in the technical field of the present invention, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.

[0154] It should also be noted that in this specification, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion.

[0155] Comprise, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such a process, method, article or device. Without more limitations, the element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.

Claims

1. A method for constructing a training data set based on a large model and RAG, characterized in that: include: Acquire input instruction information of the user, wherein the input instruction information includes text information or language information; Performing prompt word engineering on the input instruction information to obtain engineering instruction information; Parsing the engineering instruction information and extracting task keywords therefrom; Inputting the task keywords into a large language model for generalization processing to obtain generalized instruction information; The generalized instruction information is returned to the user and supplemented and corrected multiple times according to the user's feedback information to obtain the final instruction information; Slicing and vectorizing the final instruction information to obtain vector information; Search and match the vector information through a vector database to obtain matching information; The vector database is screened according to the matching information to obtain a training data set, wherein the training data set includes a picture training data set and a video training data set.

2. The method for constructing a training data set based on a large model and RAG according to claim 1, characterized in that: The step of engineering the input instruction information into prompt words to obtain the engineering instruction information includes: Collect image information and device information from the upstream data system; wherein the image information includes Chinese and English label information, annotation information, image-text pair information and image description information, and the device information includes regional scene information and device label information; Performing data sorting and standardization processing on the image information and device information to obtain standardized information; Vectorizing the standardized information to obtain vector information; The vector information is stored in the vector database.

3. The method for constructing a training data set based on a large model and RAG according to claim 2, characterized in that: The vectorizing the standardized information to obtain the vector information includes: After data sorting and standardization, a dictionary is established to directly vectorize the Chinese and English label information, annotation information, regional scene information and device label information to obtain the first vector information; After data sorting and standardization, the text-image information and picture description information are vectorized after text segmentation, keyword extraction and relationship extraction to obtain the second vector information.

4. The method for constructing a training data set based on a large model and RAG according to claim 1, characterized in that: The searching and matching of the vector information by the vector database to obtain the matching information includes: The vector information is searched and matched by using an accurate matching method to obtain first matching information; The vector information is searched and matched by similarity matching to obtain second matching information.

5. The method for constructing a training data set based on a large model and RAG according to claim 1, characterized in that: The slicing and vectorizing of the final instruction information to obtain vector information includes: Defining nodes, edges, and dimensions of nodes according to the vector information to obtain an initialized graph structure; Calculating the distance between each node in the initialization graph structure, and generating edge connections between each node according to the calculated distance to obtain a neighbor graph; Adjust the edge connections between nodes through a greedy algorithm; Randomly assign levels to each node to form a sparse high-level structure; Iteratively adjust the number of node layers and edges according to the preset threshold until the convergence condition is reached to obtain the final graph structure; The final graph structure is stored in a graph database or a vector database.

6. The method for constructing a training data set based on a large model and RAG according to claim 1, characterized in that: The obtaining of a training data set by screening from the vector database according to the matching information comprises: Auxiliary screening is performed based on the matching information through a graph database and a distributed search analysis engine.

7. The method for constructing a training data set based on a large model and RAG according to claim 1, characterized in that: The step of screening the vector database according to the matching information to obtain a training data set includes: Recall data is obtained and input into a vector retrieval model for tuning.

8. A training data set construction device based on a large model and RAG, characterized in that: include: An acquisition unit, used to acquire input instruction information of a user, wherein the input instruction information includes text information or language information; An engineering unit, used for engineering the input instruction information with prompt words to obtain engineering instruction information; A parsing unit, used to parse the engineering instruction information and extract task keywords therefrom; A generalization unit, used for inputting the task keywords into a large language model for generalization processing to obtain generalized instruction information; A correction unit, used for returning the generalized instruction information to the user and performing multiple rounds of supplementation and correction according to the user's feedback information to obtain final instruction information; A slicing unit, used for slicing and vectorizing the final instruction information to obtain vector information; A matching unit, used to search and match the vector information through a vector database to obtain matching information; A screening unit is used to screen the vector database according to the matching information to obtain a training data set, wherein the training data set includes a picture training data set and a video training data set.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method for constructing a training data set based on a large model and RAG is implemented as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, enables the processor to execute the method for constructing a training data set based on a large model and RAG as described in any one of claims 1 to 7.