Method, device, medium and electronic equipment for generating training data for large models
By generating high-quality and diverse visual detection question-answer pairs, and utilizing the detection target library and target description information, the training dataset is automatically extracted and constructed, which solves the problem of insufficient recognition accuracy of visual language models in specific industries and improves the detection effect.
Patent Information
- Application Number
- CN202511432827.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-30
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2045-09-30
AI Technical Summary
Existing visual language models lack the ability to accurately identify subtle and complex defects or violations in specific industry applications, making them difficult to train and test effectively, resulting in poor performance of AI visual detection.
By generating high-quality and diverse visual detection question-answer pairs, utilizing the detection target library and target description information, training data is automatically extracted and generated. The training dataset is constructed by combining detection strategies and question-answer pairs, achieving efficient data acquisition.
It improves the recognition accuracy and defect detection capability of visual language models in complex detection scenarios, and promotes the upgrade of visual language models from general perception to automated intelligent detection in professional scenarios.
Smart Images

Figure CN120913010B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of computer, in particular, to a large model training data generation method and device, medium and electronic equipment. BACKGROUND
[0002] With the breakthrough progress of artificial intelligence, large models have been increasingly applied in automation and intelligentization in various industries, especially in visual detection. However, visual detection requires a large number of labeled special data sets for training and testing of vision language models (VLM). Therefore, how to generate data sets special for VLM is a technical problem to be solved. SUMMARY
[0003] This summary is provided to introduce a selection of concepts that are further described below in the detailed description. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it used to limit the scope of the claimed subject matter's scope.
[0004] In a first aspect, the present disclosure provides a large model training data generation method, comprising:
[0005] receiving an initial image, determining a detection target corresponding to the initial image based on a detection target library;
[0006] extracting target description information of the initial image, and generating a detection strategy corresponding to the initial image according to the target description information and the detection target;
[0007] generating a question and answer pair corresponding to the initial image based on the detection strategy, and constructing training data according to the question and answer pair.
[0008] In a second aspect, the present disclosure provides a large model training data generation device, comprising:
[0009] a determination module configured to receive an initial image, and determine a detection target corresponding to the initial image based on a detection target library;
[0010] a rule acquisition module configured to extract target description information of the initial image, and generate a detection strategy corresponding to the initial image according to the target description information and the detection target;
[0011] a construction module configured to generate a question and answer pair corresponding to the initial image based on the detection strategy, and construct training data according to the question and answer pair.
[0012] In a third aspect, the present disclosure provides a computer readable medium having stored thereon a computer program which, when executed by a processing apparatus, implements the steps of the method of the first aspect.
[0013] In a fourth aspect, the present disclosure provides an electronic device comprising:
[0014] a storage device having stored thereon a computer program;
[0015] a processing apparatus configured to execute the computer program in the storage device to implement the steps of the method of the first aspect.
[0016] In a fifth aspect, the present disclosure provides a computer program product comprising a computer program which, when executed by a processor, implements the steps of the method of the first aspect.
[0017] Through the above technical solution, in the case of receiving a plurality of initial images, for each initial image, the present disclosure can automatically obtain the corresponding labeled data thereof, and in this process, the corresponding detection target of each initial image can be obtained based on the initial image and the detection target library, that is, the corresponding detection target of the initial image can be obtained by understanding the initial image, and after the target description information of each initial image is extracted, the detection strategy specific to each initial image can be comprehensively generated by combining the target description information with the detection target, and based on the detection strategy, a large amount of labeled training data can be automatically and efficiently generated. Since the training data is obtained by combining the detection target and the target description information of the image, the quality of the training data obtained can be greatly guaranteed.
[0018] Other features and advantages of the present disclosure will be described in detail in the following detailed description. BRIEF DESCRIPTION OF DRAWINGS
[0019] The above and other features, advantages, and aspects of embodiments of the present disclosure will become more apparent by describing in detail the following specific embodiments thereof with reference to the attached drawings. The same or similar components have the same reference numbers throughout the drawings. It is to be understood that the drawings are schematic, and the components and elements are not necessarily drawn to scale. In the drawings:
[0020] Figure 1 is a flowchart of a method for generating large model training data according to an embodiment of the present disclosure.
[0021] Figure 2 is an example diagram of a first initial image in a method for generating large model training data according to an embodiment of the present disclosure.
[0022] Figure 3is an example diagram of training data constructed based on the first initial image in a large model training data generation method according to an embodiment of the present disclosure.
[0023] Figure 4 is an example diagram of a second initial image in a large model training data generation method according to an embodiment of the present disclosure.
[0024] Figure 5 is an example diagram of training data constructed based on the second initial image in a large model training data generation method according to an embodiment of the present disclosure.
[0025] Figure 6 is a specific example diagram of obtaining training data in a large model training data generation method according to an embodiment of the present disclosure.
[0026] Figure 7 is a block diagram of a large model training data generation apparatus according to an embodiment of the present disclosure.
[0027] Figure 8 is a structural schematic diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0028] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. While certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be interpreted as being limited to the embodiments set forth herein; rather, these embodiments are provided so that the present disclosure can be more thoroughly and completely understood. It should be understood that the drawings and embodiments of the present disclosure are only for illustrative purposes and are not intended to limit the scope of protection of the present disclosure.
[0029] It should be understood that each of the steps recited in the method embodiments of the present disclosure can be executed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit the execution of the steps shown. The scope of the present disclosure is not limited in this respect.
[0030] The term “comprising” and variations thereof as used herein are open-ended, that is, “including but not limited to”. The term “based on” is “based, at least in part, on”. The term “one embodiment” means “at least one embodiment”; the term “another embodiment” means “at least one additional embodiment”; the term “some embodiments” means “at least some embodiments”. Related definitions of other terms will be given in the description below.
[0031] It should be noted that the terms “first”, “second”, and the like mentioned in the present disclosure are only used to distinguish different apparatuses, modules or units, and do not imply the order or interdependence of the functions performed by these apparatuses, modules or units.
[0032] It should be noted that the modification of "one", "multiple" mentioned in the present disclosure is illustrative but not restrictive, and those skilled in the art should understand that "one or more" should be understood unless the context clearly indicates otherwise.
[0033] The names of the messages or information exchanged between the plurality of devices in the embodiments of the present disclosure are only for illustrative purposes, and are not used to limit the scope of the messages or information.
[0034] It can be understood that, before using the technical solutions disclosed in the embodiments of the present disclosure, the type, use range, use scenario, etc. of the personal information involved in the present disclosure should be informed to the user and the authorization of the user should be obtained through appropriate means in accordance with relevant laws and regulations.
[0035] For example, in response to receiving the active request of the user, prompt information is sent to the user to explicitly prompt the user that the operation requested to be performed will require obtaining and using the personal information of the user. Thus, the user can voluntarily choose whether to provide personal information to the software or hardware such as electronic device, application program, server or storage medium, etc. that performs the operation of the technical solution of the present disclosure according to the prompt information.
[0036] As an optional but non-limiting implementation manner, in response to receiving the active request of the user, the manner of sending prompt information to the user may, for example, be a pop-up window manner, and the prompt information may be presented in the form of text in the pop-up window. In addition, the pop-up window may also carry selection controls for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0037] It can be understood that the above notification and user authorization process is only illustrative, and does not limit the implementation manner of the present disclosure, and other manners that meet the relevant laws and regulations can also be applied to the implementation manner of the present disclosure.
[0038] At the same time, it can be understood that the data involved in the present technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the relevant laws and regulations and relevant provisions.
[0039] In related technologies, as the core technology of multi-modal artificial intelligence, visual language models have significant advantages in understanding complex scenes and performing fine-grained tasks. However, general visual language models lack deep understanding of specific industry knowledge and cannot accurately identify subtle and complex defects or violations in various vertical fields, i.e., the application effect is not good. In order to effectively land the artificial intelligence visual detection, it is necessary to train or fine-tune the VLM detection model that belongs to a specific field and is highly customized. Therefore, obtaining massive, high-quality and diversified labeled training data is a core problem of artificial intelligence visual detection.
[0040] Therefore, to solve the above problems, the embodiment of the present disclosure provides a large model training data generation method, device, medium and electronic equipment, which can intelligently convert user requirements and visual materials into large-scale, high-quality and diversified question and answer pairs (QA) through a highly structured and automated process, that is, the visual detection related training data set can be more efficiently obtained through the embodiment of the present disclosure.
[0041] The present disclosure will be further explained and described below in conjunction with the accompanying drawings.
[0042] Figure 1 FIG. 1 is a flowchart of a large model training data generation method according to an embodiment of the present disclosure. The large model training data generation method can be applied to an electronic device with processing capability, such as a terminal or a server. The large model training data generation method can be executed by a large model training data generation device, wherein the large model training data generation device can be implemented by software and / or hardware, and the software and / or hardware can be configured in the electronic device. Referring to FIG. 1, the large model training data generation method can include the following steps. Figure 1 The large model training data generation method can include the following steps.
[0043] In step S110, an initial image is received, and a detection target corresponding to the initial image is determined based on a detection target library.
[0044] In the embodiment of the present disclosure, the initial image can be a material image, which can be a picture or video material to be detected, that is, when the user uploads a material image, the embodiment of the present disclosure can directly take it as an initial image; when the user uploads a video material, the embodiment of the present disclosure can sample and frame process the video material to obtain a plurality of video frames, and take these video frames as initial images. Here, the picture or video material to be detected uploaded by the user can be original visual material.
[0045] In the embodiment of the present disclosure, the original visual material can be uploaded by the user, or can be inspection data extracted by searching a database, which can be extracted based on the user's requirements. For example, the user's requirement is to "generate detection data about urban road traffic scenes", and the inspection data is visual data related to urban road traffic.
[0046] As an optional way, in the case of receiving a plurality of initial images, the embodiment of the present disclosure can pre-process the initial images, that is, screen the initial images. In this process, the embodiment of the present disclosure can screen out high-quality pictures or videos highly related to the task scene through the way of large model coarse screening and visual model fine screening.
[0047] For example, in the coarse screening process, the embodiments of the present disclosure can use a large model to quickly and preliminarily screen the data, and eliminate samples that do not obviously meet the requirements, so as to improve the efficiency of subsequent training data generation, such as eliminating images that are blurred, unclear, or have obvious defect problems. Fine screening can be performed on the basis of coarse screening, that is, the embodiments of the present disclosure can perform more in-depth and fine analysis on the data through a visual model to achieve higher screening accuracy. Through the cooperative work of coarse screening and fine screening, the embodiments of the present disclosure can ensure that the data processing is both efficient and accurate.
[0048] In the process of fine screening the initial images, the embodiments of the present disclosure can screen the initial images based on the requirement information to obtain candidate images related to the requirement information. On this basis, the detection target corresponding to the candidate images is determined based on the detection target library. That is, the embodiments of the present disclosure can screen a plurality of initial images according to the requirement input by the user to select images that meet the actual requirements.
[0049] Taking the above example, the requirement information input by the user is "generate detection data about urban road traffic scenes", and in the process of screening the initial images, the embodiments of the present disclosure can remove data unrelated to urban roads to screen out high-quality visual materials related to the task scene.
[0050] It should be noted that in the process of fine screening the initial images, the embodiments of the present disclosure can determine the correlation degree between the initial images and the requirement information, and when the correlation degree is less than a preset threshold, the initial images are eliminated; when the correlation degree is greater than the preset threshold, the embodiments of the present disclosure can grade the initial images, and the higher the correlation degree, the higher the grade of the initial images. Subsequently, the detection strategy corresponding to the initial images can be generated in combination with the target description information, the detection target and the correlation degree grading, and the detection strategy can also be referred to as a detection rule. For example, the first correlation degree between the first initial image and the target requirement is greater than the second correlation degree between the second initial image and the target requirement, and therefore the information involved in the first initial image is more detailed when generating the corresponding detection strategy, that is, the first initial image can have a higher level of detail than the second initial image.
[0051] As another optional way, the embodiments of the present disclosure can determine the detection target corresponding to each initial image based on the detection target library, wherein the detection target library can be generated based on the user requirement, that is, in response to the received requirement information, the embodiments of the present disclosure can construct a detection target library related to the requirement information.
[0052] Upon receiving the requirement information, the embodiments of the present disclosure can analyze the requirement information, that is, receive the core task requirement of the user in the form of natural language. For example, the requirement information can be "generate detection data about urban road traffic scenes", or can be "generate data for identifying the state of goods on the warehouse shelves".
[0053] Exemplarily, the embodiment of the present disclosure can utilize a large language model (LLM) to perform semantic analysis on the demand information, and based on the semantic analysis result, the macro target and boundary of the task can be accurately captured.
[0054] Based on this, the embodiment of the present disclosure can perform a core element extraction operation, that is, extract a plurality of key elements from the semantic obtained by analysis, and through these key elements, a complete detection task can be constructed. By extracting the core elements in the demand information, the fuzzy detection concept can be converted into machine understandable and operable semantic architecture, thereby providing clear guidance for the construction of the subsequent detection target library.
[0055] Among them, the key elements can include at least one of a detection scene, a detection subject, a detection behavior, and a detection target. Exemplarily, the detection scene can be a construction site, a road, etc., the detection subject can be a worker, a vehicle, etc., the detection behavior can be wearing, illegal parking, etc., and the detection target can be a safety helmet, a no-parking area, etc.
[0056] In addition, in the process of constructing the detection target library, the embodiment of the present disclosure finds and obtains the specification text associated with the semantic analysis according to the semantic analysis result of the demand information, and through processing, extracting, understanding, etc. of the specification text, the above-mentioned plurality of key elements related to the demand information can be obtained.
[0057] In summary, the detection target library can be constructed based on the demand information input by the user. When receiving new demand information, the embodiment of the present disclosure can update the detection target library based on the new demand information, that is, add the hierarchical information related to the demand information. Conversely, when the received demand information does not involve new demand, the embodiment of the present disclosure can reuse the hierarchical information of the detection target library corresponding to the demand information.
[0058] As can be seen, the detection target library can include a plurality of sets of hierarchical information, and each set of hierarchical information can include at least one of detection scene information, detection field information, and detection target information. Among them, the detection scene information refers to the specific detection scene, that is, which scene is detected, that is, according to the user demand, the embodiment of the present disclosure can define the macro detection scene, such as the detection scene can be a city road, an indoor scene, an outdoor factory area, etc.
[0059] The detection field information is to detect the category of the target in each scene. For example, for the "urban road" scene, the detection type (detection field) can be divided into "road equipment detection", "road traffic detection", "personnel safety detection", etc. That is, the detection field information can be used to limit which type of detection is performed on the specified scene, and in order to better implement the detection of the detection scene, the detection field can be further divided into a first-level label and a second-level classification.
[0060] The detection target information can be a further refinement of the detection field, that is, further refining the detection type into specific and operable detection targets / monitoring purposes. For example, the detection target information can be "check whether the motorcycle and electric vehicle passengers wear safety helmets", "detect whether the warning signs in the road construction area are set according to the specification", etc.
[0061] In order to better understand the structure of the detection target library, the present embodiment of the disclosure gives Table 1 as shown below:
[0062] Table 1
[0063]
[0064] Based on Table 1, for the same scene "outdoor factory area", the fields are different, and the corresponding inspection purposes (detection targets) are also different. The above detection scene information, field information and target information can be structured information obtained by analyzing the demand information.
[0065] After the above structured information is obtained by carding, the present embodiment of the disclosure can store these structured information into the detection target library to form a standard and reusable knowledge system. It can be seen that by carding the present embodiment of the disclosure, the multi-level classification and scene of each requirement can be obtained, such as obtaining first-level, second-level and third-level classification and scene. On this basis, the inspection purposes under each scene are constructed to obtain the detection target library.
[0066] It should be noted that the detection target library in the present embodiment of the disclosure can be generated in real time according to the initial image and the demand information, that is, when the demand information input by the user is received, the multi-group hierarchical information of the detection target library corresponding to the demand information is generated.
[0067] Alternatively, the detection target library can also be pre-constructed, that is, the user can generate a complete detection target library by inputting a large amount of demand information, and in the subsequent process of generating training data, the hierarchical information related to the demand information can be directly searched from the detection target library.
[0068] The embodiment of the present disclosure can extract structured information such as scenes, types and detection targets based on demand information, so as to convert scattered implicit knowledge and unstructured documents into structured instructions that can be accurately executed by devices. In the conversion process, the embodiment of the present disclosure can convert fuzzy concepts into a series of explicit and operable detection points to solve the problem of unclear target or unclear definition in the model training process from the source, so as to provide standard, traceable, high-quality and unambiguous input for the entire data set construction.
[0069] As another optional way, when determining the detection target corresponding to the initial image based on the detection target library, the embodiment of the present disclosure can identify the scene of the initial image to obtain the scene information corresponding to each initial image. On this basis, the scene information is used to determine the detection target matching the initial image in the detection target library. That is, the embodiment of the present disclosure can find the detection target matching the initial image from the detection target library by identifying the scene of the initial image, so as to match the structured knowledge with the unstructured visual content.
[0070] For example, the embodiment of the present disclosure can analyze each initial image based on the VLM model to obtain the scene information corresponding to each initial image, which can be used as the label of the corresponding initial image. For example, by identifying the first initial image, it is determined that the scene information of the first initial image is a city road scene, and the scene label of the first initial image can be "city road", so as to realize the image scene identification and labeling operation.
[0071] After identifying the scene information corresponding to the initial image, the embodiment of the present disclosure can perform a detection target matching operation, that is, determine the detection target matching the initial image in the detection target library based on the scene information. In other words, in the identified scene, the specific detection target defined in the detection target library is further matched in the image.
[0072] For example, by identifying the first initial image, it is determined that the scene matching the first initial image is a city road, and by searching the detection target library, it is determined that the target information corresponding to the scene "city road" is multiple. On this basis, the embodiment of the present disclosure can determine the detection target matching the initial image by performing semantic recognition on the initial image, such as "checking whether the motorcycle and electric vehicle passengers are in violation of carrying multiple people".
[0073] The embodiment of the present disclosure can obtain the detection target corresponding to each initial image by matching the image scene purpose library. Here, the detection target can also be referred to as the rule of the detection target library, that is, based on the detection target library, the embodiment of the present disclosure can determine the reference detection strategy (detection target) corresponding to each initial image.
[0074] In step S120, target description information of the initial image is extracted, and a detection strategy corresponding to the initial image is generated based on the target description information and the detection target.
[0075] As an optional approach, after receiving the initial image, this embodiment of the present disclosure can, on the one hand, determine the detection target matching each initial image through a detection target library, and on the other hand, extract the target description information of the initial image, that is, perform descriptive analysis on the initial image. Here, the target description information can also be called objective information, which can be a pure and inviolable visual fact in the initial image. This target description information can provide a realistic basis for subsequent generation of context-related detection tasks.
[0076] It is understood that, in the process of extracting target description information from the initial image, embodiments of this disclosure can generate initial description information for the initial image, wherein the initial description information may be textual information obtained to describe the initial image. Based on this, the initial description information is processed to obtain target description information, that is, unique and objective information is selected from the initial image and used as the target description information.
[0077] For example, embodiments of this disclosure can generate a purely descriptive caption for each initial image based on a VLM model, without subjective inference. This caption can serve as initial descriptive information. In other words, the initial descriptive information is a description of the things included in the initial image.
[0078] For example, for such Figure 2 The initial description information obtained by extracting the image information of the first initial image 201 shown can be: "This image is an indoor work scene under renovation and construction. There are three workers in the image, and construction materials are scattered on the ground. The three workers are installing building materials. They seem to be observing or pointing at the building materials piled up next to them. The one on the right is a woman wearing a baseball cap, a safety vest, and a short-sleeved shirt; the man in the middle is dressed neatly and seems to be smoking; the man on the left is wearing a safety helmet, gloves, and holding tools."
[0079] In order to ensure the uniqueness and authenticity of the extracted target description information, the embodiment of the disclosure can process (check) the visual information on the basis of generating the initial description information of the initial image, that is, detect the speculative information and ambiguity information in the initial description information, and delete these information to obtain the target description information. In the process of processing the visual information, the embodiment of the disclosure can determine the uncertain description information in the visual information, such as the above examples of "they seem to be observing or pointing to the building boards piled aside" and "he seems to be smoking", and "seem" and "seem" all represent uncertain meanings. Such expressions may be the result obtained by the large model through speculation on the semantics of the initial image, and are not the target description information of the image, that is, they have speculative meanings. In order to ensure the uniqueness and authenticity of the target description information, the embodiment of the disclosure can screen out ambiguous expressions or speculative expressions.
[0080] Based on the above examples, the initial description information is processed to obtain the target description information, which can be "an indoor working scene under renovation and construction, there are three workers in the picture, and there are renovation materials scattered on the ground. Among them, the three workers are installing boards, the female on the right side is wearing a baseball cap, a safety vest and a short sleeve; the male in the middle is wearing a neat shirt; the male on the left side is wearing a safety helmet, wearing gloves and holding a tool". Figure 2 The target description information obtained by processing the initial description information can be "an indoor working scene under renovation and construction, there are three workers in the picture, and there are renovation materials scattered on the ground. Among them, the three workers are installing boards, the female on the right side is wearing a baseball cap, a safety vest and a short sleeve; the male in the middle is wearing a neat shirt; the male on the left side is wearing a safety helmet, wearing gloves and holding a tool".
[0081] Through the processing of the initial description information, the unique target description information of each initial image can be screened. It can be seen that by purifying the generated initial description information (caption), key and clear objective fact elements can be screened out.
[0082] It should be noted that in the process of processing the initial description information, the embodiment of the disclosure can determine whether there is an obviously unconventional description in the initial description information. If such a description exists, it can be screened out. For example, by extracting the second initial image, the initial description information obtained includes "there are two people in the picture, and the two people are walking in the rain", but by identifying the ground, it is determined that it is in a dry state. It can be seen that the initial description information at this time has a description that violates the convention, that is, the "rain" in the picture may be dust. In order to ensure the uniqueness and accuracy of the target description information, it can be screened out.
[0083] In the embodiment of the disclosure, the extracted target description information can be used as the fact anchor point of the initial image, that is, it can be used as the unforgeable visual basis. For example, the target description information of the first initial image is "there are potholes, water and snow on the road", and "there are potholes, water and snow on the road" can be used as the target description information of the first initial image.
[0084] In summary, on the one hand, the present disclosure can guide and match the visual material (initial image) to obtain the detection target, and on the other hand, the present disclosure can generate the corresponding initial description information (caption) by describing the visual material, and further extract the fact anchor point that cannot be violated. Here, the detection target and the target description information as the guidance of subsequent training data generation can not only make the subsequent generated content have a clear business orientation, but also ensure that the generated content is rooted in real visual evidence, which can to some extent prevent the common groundless speculation and content hallucination of large language models from the source. That is, the detection target and the target description information are the core cornerstone of guaranteeing the generation of high-quality and high-credibility data.
[0085] As another optional way, after obtaining the detection target and the target description information corresponding to each initial image, the present disclosure can generate the detection strategy corresponding to the initial image according to the target description information and the detection target. For example, the present disclosure can generate the detection strategy corresponding to the initial image by imitation and expansion, wherein the imitation and expansion can be based on the target description information and the detection target.
[0086] In other words, the present disclosure can enrich the detection target dynamically and finely based on the detection target and the objective fact, and generate the actual detection strategy. For example, the detection target obtained by detecting the first initial image is "check whether the road surface has water", and the fact anchor point obtained by extracting the target description information of the first initial image is "the road has potholes, water, and snow". Based on this, the detection strategy generated based on "check whether the road surface has water" and "the road has potholes, water, and snow" can be "check whether the road has water, check whether the road has potholes, and check whether the road has snow".
[0087] Optionally, during the process of generating the detection strategy corresponding to the initial image according to the target description information and the detection target, the present disclosure can determine the key information of the detection target, and expand the target description information based on the key information to obtain the detection strategy corresponding to the initial image. Here, the key information can be the risk point of the detection target, and the present disclosure can generate the instantiation and diversification of the detection strategy by taking the risk point as a reference.
[0088] The determination of the detection target information can be obtaining structured points of the detection target, i.e., abstracting the detection target into a structured data template, which can include inherent attributes and relationship attributes. In addition, the target description information can be multi-dimensional, such as the target description information at least including object recognition information, spatial information, state information, and environmental context information. The embodiments of the present disclosure can generate a detection strategy based on logical matching and instantiation of the template attributes and the target description information. That is, the abstract attributes in the template are combined with the specific information in the image, and a specific and executable detection instruction is generated through a predefined logical relationship.
[0089] For example, the embodiments of the present disclosure can obtain a detection target template (detection target) defining inherent attributes and / or relationship attributes of at least one object to be detected, obtain target description information of an initial image, which can include object information, spatial relationship information, and state information extracted by image recognition technology, and perform coupling analysis on the attributes of the detection target template and the target description information of the initial image to obtain a detection strategy.
[0090] Specifically, the inherent attributes are compared with the state information of the corresponding object in the initial image to generate a first type of detection strategy about object integrity or state compliance, the relationship attributes are matched with the spatial relationship information and the object information in the image to generate a second type of detection strategy about object and surrounding environment relationship compliance, and the initial image corresponding detection strategy is constructed by the first type of detection strategy and the second type of detection strategy.
[0091] As an example, the detection target is to detect whether the fire hydrant has a pressure gauge and whether the paint color is red (inherent attribute), and whether there is an obstacle within 1.2 meters (relationship attribute), and the target description information of the initial image is that the pressure gauge pointer of the fire hydrant is in the red area, the paint color is partially rusted, and a wooden tray is identified 0.8 meters in front of it. The first type of rule generated by combining the inherent attributes and the state information can be whether the pressure gauge pointer of the fire hydrant is in the red area, and the second type of rule generated by combining the relationship attributes, the spatial information, and the object information can be whether there is an article within 1.2 meters around the fire hydrant.
[0092] The embodiments of the present disclosure can perform a copy operation by taking the detection target as a template and taking the target description information as filling content, so as to automatically expand a small amount of detection target into a large number of specific and instantiated detection strategies, thereby ensuring that each generated detection strategy is highly bound to the visual material in terms of semantics and details. That is, the embodiments of the present disclosure can realize intelligent expansion from "one detection target" to "N scene-based instance detection strategies" through the detection target and the target description information, while suppressing model recognition errors and hallucinations, the embodiments of the present disclosure can greatly improve the quality and diversity of generated data.
[0093] It should be noted that in the process of generating the detection strategy corresponding to the initial image, the embodiments of the present disclosure can directly follow or copy the rules, that is, the detection target can be directly used to generate the detection strategy, or the detection target can be copied and learned to obtain a detection strategy unique to the initial image. At the same time, the large model can follow or refer to the prompt example to generate the prompt information corresponding to the initial image.
[0094] In addition, in the process of copying the detection strategy, the detection target can carry label information, which can be the hierarchical information in the detection target library. As introduced above, the same scene can correspond to multiple detection categories, such as the "outdoor factory area" scene, which can correspond to the cleaning state category under environmental health, the behavior category under rules and regulations, and the behavior category under safety monitoring, etc. By configuring label information for the detection target, a detection strategy unique to the initial image can be more accurately generated.
[0095] By copying and learning the detection targets of various categories in the scene corresponding to the initial image, a detection strategy unique to the initial image can be generated based on the target description information, such as copying and learning each detection target of the outdoor factory area scene in the detection target library, and combining the target description information of the initial image, the final generated detection strategy is "detect whether there are garbage and sundries around the gatekeeper room in the factory area; detect whether there are garbage and sundries on the ground at the factory gate".
[0096] In step S130, the detection strategy is used to generate a question and answer pair corresponding to the initial image, and the training data is constructed according to the question and answer pair.
[0097] As an optional way, after generating the detection strategy corresponding to the initial image, the embodiments of the present disclosure can generate a question and answer pair corresponding to the initial image based on the detection strategy, and construct training data according to the question and answer pair, so as to convert the processed information into formatted training data. Here, the initial image is not the same, and the training data constructed according to the initial image is also not the same.
[0098] In order to ensure that the generated question and answer pair is uniform in format and natural in style, the embodiments of the present disclosure can randomly extract a specified number of question and answer pair examples (prompt cases) from the high-quality data provided by the user or the preset excellent examples, and learn these question and answer pair examples to obtain the ability to output the specified paradigm. That is, the embodiments of the present disclosure can perform a random learning (Few-shot Learning) operation to dynamically master the prompt (prompt information) output paradigm corresponding to the initial image.
[0099] On this basis, the embodiment of the disclosure can generate a target question of a specified paradigm according to the detection strategy, and detect the initial image for the target question to obtain a target answer of the specified paradigm. That is, the embodiment of the disclosure can generate a target question (Question, Q) according to the detection strategy and the question and answer pair example, and based on this, the target answer (Answer, A) can be generated through interaction with a large model.
[0100] In the data synthesis process, in order to avoid the problem of rigid and generalized output content format, the embodiment of the disclosure introduces a random learning method, that is, the embodiment of the disclosure does not use a fixed output template, but dynamically learns a small amount of high-quality data samples provided by the user or dynamically learns preset excellent examples before generating data. Through the imitation learning of these examples, the embodiment of the disclosure can automatically master and reproduce the specific question and answer style and data format expected by the current image, so as to ensure that the generated training data has diversity.
[0101] In other words, the embodiment of the disclosure can make the batch-generated question and answer pairs not only accurate in content, but also diverse in language paradigm by introducing a random learning method, that is, the automatically generated training data can directly meet the strict requirements of downstream model training without manual adjustment or manual post-tuning.
[0102] In the process of generating a question and answer pair corresponding to the initial image based on the detection strategy, the embodiment of the disclosure can generate prompt information (prompt) according to the detection strategy, which can include the detection strategy, output format, instruction information and matters needing attention, etc. On this basis, the embodiment of the disclosure can input the prompt information and the initial image into the large model to obtain the answer of the specified paradigm.
[0103] In summary, the embodiment of the disclosure can combine the detection strategy corresponding to the initial image with the output format (specified paradigm) learned by the model to automatically generate a large number of high-quality detection question and answer pairs. Based on this, the embodiment of the disclosure can bind the initial image, the prompt information and the automatically labeled detection result to generate a complete training data triple, and store it in the final training data set in batches.
[0104] As an example, when receiving a first initial image 201 as shown in Figure 2 When receiving a first initial image 201 as shown in Figure 2 , the embodiment of the disclosure can determine the detection target corresponding to the first initial image 201 based on the detection target library, and at the same time, the embodiment of the disclosure can extract the target description information of the first initial image 201, and comprehensively generate the detection strategy corresponding to the first initial image 201 according to the target description information and the detection target. On this basis, the question and answer pair corresponding to the first initial image 201 can be generated based on the detection strategy, and the training data of the first initial image 201 can be constructed according to the question and answer pair.
[0105] As known from the above introduction, the training data can include images, prompt information and detection results, so the finally obtained training data of the first initial image 201 can include the first initial image 201, the prompt information corresponding to the first initial image 201 and the corresponding detection result, as shown in detail in Figure 3 Figure 3 It can be known that the prompt information (inspection prompt) of the first initial image 201 can include role description information in addition to the inspection rule (detection strategy), output format and description information. After obtaining the prompt information, the embodiment of the disclosure can input it into the large model together with the first initial image 201, and the detection result (answer) can be obtained.
[0106] By comparing Figure 2 It can be known that Figure 2 There is a violation, and the violation is because someone does not wear a safety helmet, and the specific number of violators is 2, so Figure 3 The generated inspection result is completely correct.
[0107] As another example, when receiving the second initial image 401 as shown in Figure 4 The embodiment of the disclosure can determine the detection target corresponding to the second initial image 401 based on the detection target library, and at the same time, the embodiment of the disclosure can extract the target description information of the second initial image 401, and the detection strategy corresponding to the second initial image 401 can be generated according to the target description information and the detection target. On this basis, the question and answer pair corresponding to the second initial image 401 can be generated based on the detection strategy, and the training data of the second initial image 401 can be constructed according to the question and answer pair.
[0108] As above, as shown in Figure 5 The training data of the second initial image 401 can include the second initial image 401, the prompt information corresponding to the second initial image 401 and the corresponding detection result. Based on Figure 5 It can be known that the prompt information (inspection prompt) of the second initial image 401 can include role description information and output requirements in addition to the detection rule (detection strategy) and output format. After obtaining the prompt information, the embodiment of the disclosure can input it into the large model together with the second initial image 401, and the detection result (answer) can be obtained.
[0109] By comparing Figure 4 It can be known that Figure 4 The personnel equipment in Figure 5 The generated inspection result is completely correct.
[0110] The format of the detection result is generated based on the output format in the prompt information, and the output format is generated through a random learning operation. The output format of the image at different times is not the same. Even for the same image, the output format obtained by the corresponding acquisition may also be different. In this way, the generalization ability of the large model in actual application can be improved, and the recognition rate for low-frequency events is high. To some extent, the diversity of the training data can be ensured.
[0111] As another optional mode, after obtaining the training data, the embodiment of the disclosure can evaluate the training data. Specifically, the embodiment of the disclosure can determine whether the detection strategy in the training data matches the content of the corresponding initial image, that is, the matching degree of the two is obtained. When the matching degree exceeds the preset matching degree, it is determined that the generated data is high-quality data. That is, when it is determined that the detection strategy conforms to the scene in the picture and can obtain a unique objective answer according to the clues in the picture, it is determined that the detection strategy matches the image content.
[0112] Optionally, the embodiment of the disclosure can determine whether the detection strategy conforms to the actual application requirements of the detection scene, that is, when it is determined that the detection strategy is applicable to the corresponding image and can also be applied to the same scene of the image, it is determined that the training data meets the requirements.
[0113] Optionally, the embodiment of the disclosure can determine whether the detection strategy in the prompt information is limited to the detection target library, that is, whether it can be dynamically expanded according to the material. If it can be dynamically expanded, it is determined to meet the requirements, otherwise, if it is limited to the detection target library, it is determined not to meet the requirements. Optionally, the embodiment of the disclosure can determine whether the prompt information is directly available or can be directly delivered through local fine-tuning (modification time cost <= 5 minutes). If it is satisfied, it is determined to meet the requirements.
[0114] The above-mentioned multiple evaluation modes can be used comprehensively, or used individually according to requirements. The specific mode for evaluating the training data is not limited here, and can be selected according to the actual situation.
[0115] As another optional mode, after obtaining the multiple training data, the embodiment of the disclosure can train the large model based on the training data. Here, the large model can be used to detect the to-be-detected image based on the requirement information to obtain the detection result and the detection basis.
[0116] Since the large model is generated based on high-quality data and diversified instructions, the recognition accuracy, defect detection capability and environmental adaptability of the visual language model in complex detection scenarios can be significantly enhanced, thereby promoting the upgrade of the visual language model from general perception to automatic intelligent detection in multiple professional scenarios.
[0117] As a specific implementation, as shown in Figure 6 When receiving a detection requirement described in natural language (requirement information), the embodiment of the disclosure can understand the detection requirement, and a detection target library can be constructed through the understanding of the detection requirement. In addition, after obtaining picture / video material to be detected (initial image), on the one hand, the embodiment of the disclosure can perform scene recognition on the picture / video material to be detected, so as to match a detection target in the detection target library through the recognition result, and on the other hand, the embodiment of the disclosure can perform visual understanding on the picture / video material to be detected, so as to extract target description information thereof.
[0118] On this basis, the embodiment of the disclosure can generate an actual detection strategy in combination with the detection target and the target description information, and perform large-scale production of data based on a random learning question and answer pair example (prompt sample) in combination with the detection strategy, wherein the random learning question and answer pair example is obtained through a high-quality question and answer pair example library. Based on this, the embodiment of the disclosure can obtain a training data triple, which can include an image / video, a detection prompt, and a detection result.
[0119] Through a set of highly structured and automated processing links, the embodiment of the disclosure can accurately match abstract and complex detection task requirements with specific and diverse visual content, and creatively expand it, so as to generate detection QA data sets in a large scale, high quality, and scene customization to a certain extent.
[0120] Based on the same inventive concept, the disclosure also provides a large model training data generation device, Figure 7 is a block diagram of a large model training data generation device 700 according to an example embodiment, as shown in Figure 7 The large model training data generation device 700 can include a determination module 710, a rule acquisition module 720, and a construction module 730.
[0121] The determination module 710 is configured to receive an initial image, and determine a detection target corresponding to the initial image based on a detection target library;
[0122] The rule acquisition module 720 is configured to extract target description information of the initial image, and generate a detection strategy corresponding to the initial image according to the target description information and the detection target;
[0123] The construction module 730 is configured to generate a question and answer pair corresponding to the initial image based on the detection strategy, and construct the training data according to the question and answer pair.
[0124] Optionally, the rule obtaining module 720 can also be configured to generate initial description information of the initial image, the initial description information being textual information obtained by describing the initial image; and process the initial description information to obtain the target description information.
[0125] Optionally, the rule obtaining module 720 can also be configured to filter out speculative information and ambiguous information in the initial description information to obtain the target description information.
[0126] Optionally, the rule obtaining module 720 can also be configured to determine key information of the detection target, and perform copy expansion on the target description information based on the key information to obtain a detection strategy corresponding to the initial image.
[0127] Optionally, the large model training data generation apparatus 700 can further include:
[0128] A learning module configured to randomly obtain a specified number of question and answer pair examples, and learn the question and answer pair examples to obtain the ability to output a specified paradigm.
[0129] Optionally, the construction module 730 can also be configured to generate a target question of the specified paradigm based on the detection strategy, and detect the initial image for the target question to obtain a target answer of the specified paradigm.
[0130] Optionally, the determination module 710 can also be configured to identify a scene of the initial image to obtain scene information, and determine a detection target matching the initial image in the detection target library based on the scene information.
[0131] Optionally, the construction module 730 can also be configured to construct the detection target library in response to received requirement information.
[0132] Optionally, the determination module 710 can also be configured to filter at least one initial image based on the requirement information to obtain a candidate image related to the requirement information, and determine the detection target corresponding to the candidate image based on the detection target library.
[0133] Optionally, the detection target library includes a plurality of sets of hierarchical information, and each set of hierarchical information includes at least one of:
[0134] detection scene information;
[0135] detection field information;
[0136] detection target information.
[0137] Optionally, the large model training data generation apparatus 700 can further include:
[0138] The training module is configured to train a large model based on the training data, and the large model is configured to detect a to-be-detected image based on demand information to obtain a detection result and a detection basis.
[0139] The detection target library in the embodiments of the present disclosure can support detection requirements of multiple industries and can quickly respond to customization of training data, realize efficient development, iterative optimization and highly personalized deployment of visual language models, and meet the needs of multiple segmented markets. In addition, the large-scale, high-quality and diversified training data synthesized above can provide key data support for research fields such as few-shot learning, explainability and human-machine collaboration of visual language models, and therefore can create conditions for constructing industry-level detection data sets and evaluation benchmarks to a certain extent.
[0140] The embodiments of each module in the large model training data generation apparatus 700 can refer to the related embodiments of the above method, and the present embodiment will not be repeated here.
[0141] The embodiments of the present disclosure also provide a computer readable medium having a computer program stored thereon, and the computer program is executed by a processing apparatus to implement the steps of the above large model training data generation method.
[0142] The embodiments of the present disclosure also provide a computer program product comprising a computer program, and the computer program is executed by a processor to implement the steps of the above large model training data generation method.
[0143] The embodiments of the present disclosure also provide an electronic device comprising:
[0144] a storage device having a computer program stored thereon;
[0145] a processing device configured to execute the computer program in the storage device to implement the steps of the above large model training data generation method.
[0146] Reference will now be made to the following description Figure 8 which shows a structural schematic diagram of an electronic device 800 (terminal device or server) suitable for implementing the embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure can include but is not limited to mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Personal Computers), PMPs (Portable Multimedia Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), and the like, and fixed terminals such as digital TVs, desktop computers, and the like. Figure 8 The electronic device shown is only an example and should not impose any limitation on the functions and use range of the embodiments of the present disclosure.
[0147] As Figure 8As shown, the electronic device 800 can include a processing device (e.g., a central processor, a graphics processor, etc.) 801 that can perform various suitable actions and processes in accordance with programs stored in a read-only memory (ROM) 802 or loaded into a random access memory (RAM) 803 from a storage device 808. Various programs and data required by the electronic device 800 for operation are also stored in the RAM 803. The processing device 801, the ROM 802, and the RAM 803 are connected to each other by a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0148] Generally, the following devices can be connected to the I / O interface 805: input devices 806 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 807 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 808 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 809. The communication devices 809 can allow the electronic device 800 to communicate wirelessly or wired with other devices to exchange data. Although Figure 8 The electronic device 800 is shown with various devices, but it should be understood that not all of the shown devices are required to be implemented or present. More or fewer devices can alternatively be implemented or present.
[0149] In particular, the processes described above with reference to the flowcharts can be implemented as a computer software program according to embodiments of the present disclosure. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network through the communication devices 809, or installed from the storage devices 808, or installed from the ROM 802. When the computer program is executed by the processing device 801, the above-mentioned functions defined in the methods of embodiments of the present disclosure are performed.
[0150] It is noted that the aforementioned computer-readable medium of the present disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium can be, for example and without limitation, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program used by or in connection with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include a computer-readable program code transmitted by a computer-readable signal medium, in a baseband or as a part of a carrier wave. Such a computer-readable signal medium can take a variety of forms, including but not limited to, electro-magnetic, optical, or any suitable combination of the foregoing. The computer-readable signal medium can also be any computer-readable medium that is not a computer-readable storage medium and that can be used to carry or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained in the computer-readable medium can be transmitted by any suitable medium, including but not limited to, wire, cable, RF, etc., or any suitable combination of the foregoing.
[0151] In some embodiments, the terminal device, the server can communicate using any currently known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication of any form or medium (e.g., a communication network). Examples of the communication network include a local area network ("LAN"), a wide area network ("WAN"), an internetwork (e.g., the Internet), and an end-to-end network (e.g., an ad hoc end-to-end network), as well as any currently known or future developed network.
[0152] The aforementioned computer-readable medium can be included in the aforementioned electronic device; or can exist separately from the electronic device and not be assembled into the electronic device.
[0153] The computer readable medium described above carries one or more programs, when the one or more programs are executed by the electronic device, cause the electronic device to: receive an initial image, determine a detection target corresponding to the initial image based on a detection target library; extract target description information of the initial image, generate a detection strategy corresponding to the initial image according to the target description information and the detection target; generate a question and answer pair corresponding to the initial image based on the detection strategy, and construct the training data according to the question and answer pair.
[0154] Computer program code for carrying out operations of the present disclosure can be written in one or more programming languages or combinations of languages including object oriented programming languages such as Java, Smalltalk, C++ as well as conventional procedural programming languages such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0155] The flow diagrams and the block diagrams in the drawings are illustrations of architectures, functionalities, and operations of possible implementations of systems, methods, and computer program products according to various embodiments of present disclosure. In this regard, each block in the flow diagrams or block diagrams can represent a module, a procedure, or a part of code, which comprises one or more executable instructions for implementing the specified functions. It should also be noted in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks depicted in succession can in fact be executed substantially concurrently or the blocks can sometimes be executed in the reverse order, depending upon the functionality involved. It is also noted that each block in the block diagrams and / or flow diagrams and combinations of blocks in the block diagrams and / or flow diagrams can be implemented by dedicated hardware-based systems that perform the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0156] The modules involved in the embodiments of the present disclosure can be implemented in a software manner or in a hardware manner. In some cases, the name of the module does not constitute a limitation on the module itself.
[0157] The functionality described above can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, illustrative types of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Application-specific Integrated Circuits (ASICs), Application-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.
[0158] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more of: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0159] The above description is only that of preferred embodiments of the present disclosure and a description of principles of technology applied. It should be understood by those skilled in the art that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and should also cover other technical solutions formed by any combinations of the above technical features or equivalent features without departing from the above disclosed concept. For example, the technical solutions formed by replacing the above features with the technical features disclosed in the present disclosure (but not limited to) having similar functions.
[0160] Further, while operations are depicted in a particular order, this should not be understood as requiring these operations to be performed in the particular order shown or in sequential order, as some can be performed in parallel or in any suitable order. Also, while a number of specific implementation details have been discussed, these should not be construed as limiting the scope of the disclosure. Certain features that are described in the context of separate embodiments can also be implemented in combination. Conversely, various features that are described in the context of a single embodiment can also be implemented separately or in any suitable subcombination. It will be appreciated that details of the foregoing embodiments, given only for purposes of illustration, are not intended to limit the scope of the present disclosure. As one of ordinary skill in the art will readily appreciate, the concepts, as generally taught herein, can be applied to a wide variety of alternative designs, structures, applications, and procedures. Therefore, the specific embodiments discussed above are not intended as being limiting; other embodiments can readily be devised in accordance with these teachings while remaining within the scope of the present disclosure.
[0161] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims. With respect to the devices in the above-described embodiments, in which various modules perform operations, the specific manner in which the various modules perform the operations has been described in detail in the embodiments relating to the method. Here, no detailed explanation will be given.
Claims
1. A method for generating training data for a large model, characterized in that, The method includes: In response to the received request information, construct a target detection library; Receive an initial image, determine the detection target corresponding to the initial image based on the detection target library, and the detection target is a reference detection strategy; Generate initial description information for the initial image, filter out speculative and ambiguous information in the initial description information, and obtain target description information for the initial image; Determine the key information of the target to be detected, and expand the target description information based on the key information to obtain the detection strategy corresponding to the initial image; Based on the detection strategy, question-answer pairs corresponding to the initial image are generated, and training data is constructed based on the question-answer pairs.
2. The method for generating large model training data according to claim 1, characterized in that, The initial description information is the textual information obtained by describing the initial image.
3. The method for generating large model training data according to claim 1, characterized in that, The method further includes: Randomly generate a specified number of question-answer pair examples; By learning from the question-and-answer pairs of examples, the ability to output a specified paradigm can be obtained.
4. The method for generating large model training data according to claim 3, characterized in that, The step of generating the question-answer pair corresponding to the initial image based on the detection strategy includes: Based on the detection strategy, a target problem of the specified paradigm is generated; The initial image is detected in response to the target problem to obtain the target answer of the specified paradigm.
5. The method for generating large model training data according to any one of claims 1 to 4, characterized in that, Determining the detection target corresponding to the initial image based on the detection target library includes: The scene of the initial image is identified to obtain scene information; Based on the scene information, a detection target matching the initial image is determined in the detection target library.
6. The method for generating large model training data according to claim 1, characterized in that, Determining the detection target corresponding to the initial image based on the detection target library includes: Based on the demand information, at least one of the initial images is filtered to obtain candidate images related to the demand information; The detection target corresponding to the candidate image is determined based on the detection target library.
7. The method for generating large model training data according to any one of claims 1 to 4, characterized in that, The detection target library includes multiple sets of hierarchical information, and each set of hierarchical information includes at least one of the following: Detection scene information; Information on the testing field; Detect target information.
8. The method for generating large model training data according to any one of claims 1 to 4, characterized in that, The method further includes: The large model is trained based on the training data. The large model is used to detect the image to be detected based on the requirement information, so as to obtain the detection result and the detection basis.
9. A device for generating large model training data, characterized in that, The device includes: The module builds a target detection library in response to the received requirement information; A determination module is used to receive an initial image and determine the detection target corresponding to the initial image based on the detection target library, wherein the detection target is a reference detection strategy; The rule acquisition module is used to generate initial description information of the initial image, filter out speculative and ambiguous information in the initial description information to obtain target description information of the initial image; determine the key information of the detection target, and perform imitation and expansion on the target description information based on the key information to obtain the detection strategy corresponding to the initial image; A construction module is used to generate question-answer pairs corresponding to the initial image based on the detection strategy, and to construct training data based on the question-answer pairs.
10. A computer-readable medium having a computer program stored thereon, characterized in that, When executed by a processing device, the computer program performs the steps of the method according to any one of claims 1-8.
11. An electronic device, characterized in that, include: A storage device on which computer programs are stored; A processing device for executing the computer program in the storage device to implement the steps of the method according to any one of claims 1-8.
12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Method and device for multi-stage generation of medical image question and answer thinking chain data
CN120317386A