Large model training data generation method and device, medium and electronic equipment
By receiving initial images, determining detection targets based on a target detection library, and generating question-answer pairs, the problem of insufficient recognition ability of visual language models in specific industries is solved. This enables the efficient generation of high-quality training data and improves the recognition accuracy and adaptability of the model.
Patent Information
- Application Number
- CN202511432827.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-30
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-09-30
AI Technical Summary
Existing visual language models lack the ability to accurately identify subtle and complex defects or violations in specific industry applications, resulting in poor detection performance. High-quality and diverse training data are needed to improve recognition capabilities.
By receiving an initial image, the system determines the target to be detected based on a target detection library, extracts target description information, generates a detection strategy, and constructs question-answer pairs to automatically generate training data. By combining the target to be detected with the target description information of the image, the quality and diversity of the training data are ensured.
It has enabled the efficient generation of high-quality training data, improved the recognition accuracy and defect detection capability of visual language models in complex scenarios, and promoted the upgrade of models from general perception to automated intelligent detection in professional scenarios.
Smart Images

Figure CN120913010A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of computer, in particular, to a large model training data generation method and device, medium and electronic equipment. BACKGROUND
[0002] With the breakthrough progress of artificial intelligence, large models have been increasingly applied in automation and intelligentization in various industries, especially in visual detection. However, visual detection requires a large number of labeled special data sets for training and testing of vision language models (VLM). Therefore, how to generate data sets special for VLM is a technical problem to be solved. SUMMARY
[0003] This summary is provided to introduce a selection of concepts that are further described below in the detailed description. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it used to limit the scope of the claimed subject matter's scope.
[0004] In a first aspect, the present disclosure provides a large model training data generation method, comprising: receiving an initial image, determining a detection target corresponding to the initial image based on a detection target library; extracting target description information of the initial image, and generating a detection strategy corresponding to the initial image according to the target description information and the detection target; generating a question and answer pair corresponding to the initial image based on the detection strategy, and constructing training data according to the question and answer pair.
[0005] In a second aspect, the present disclosure provides a large model training data generation device, comprising: a determination module configured to receive an initial image, and determine a detection target corresponding to the initial image based on a detection target library; a rule acquisition module configured to extract target description information of the initial image, and generate a detection strategy corresponding to the initial image according to the target description information and the detection target; a construction module configured to generate a question and answer pair corresponding to the initial image based on the detection strategy, and construct training data according to the question and answer pair.
[0006] In a third aspect, the present disclosure provides a computer readable medium having a computer program stored thereon, wherein the computer program is executed by a processing device to implement the steps of the method of the first aspect.
[0007] In a fourth aspect, the present disclosure provides an electronic device, comprising: a storage device having stored thereon a computer program; a processing device configured to execute the computer program in the storage device to implement the steps of the method of the first aspect.
[0008] In a fifth aspect, the present disclosure provides a computer program product comprising a computer program which, when executed by a processor, implements the steps of the method of the first aspect.
[0009] Through the above technical solution, in the case of receiving a plurality of initial images, for each initial image, the present disclosure can automatically obtain its corresponding labeled data. In this process, based on the initial image and the detection target library, the detection target corresponding to each initial image can be obtained, that is, by understanding the initial image, the corresponding detection target can be obtained. Subsequently, after extracting the target description information of each initial image, the detection strategy specific to each initial image can be comprehensively generated by combining the target description information with the detection target. Based on the detection strategy, a large amount of labeled training data can be automatically and efficiently generated. Since the training data is obtained by combining the detection target and the target description information of the image, the quality of the training data obtained can be greatly guaranteed.
[0010] Other features and advantages of the present disclosure will be described in detail in the following detailed description. BRIEF DESCRIPTION OF DRAWINGS
[0011] The above and other features, advantages, and aspects of embodiments of the present disclosure will become more apparent by describing in detail the following specific embodiments thereof with reference to the attached drawings. The same or similar components have the same reference numbers throughout the drawings. It is to be understood that the drawings are schematic, and the components and elements are not necessarily drawn to scale. In the drawings: Figure 1 is a flowchart of a large model training data generation method according to an embodiment of the present disclosure.
[0012] Figure 2 is an example diagram of a first initial image in a large model training data generation method according to an embodiment of the present disclosure.
[0013] Figure 3 is an example diagram of training data constructed based on a first initial image in a large model training data generation method according to an embodiment of the present disclosure.
[0014] Figure 4 is an example diagram of a second initial image in a large model training data generation method according to an embodiment of the present disclosure.
[0015] Figure 5is an example diagram of training data constructed based on the second initial image in a large model training data generation method according to an embodiment of the present disclosure.
[0016] Figure 6 is a specific example diagram of obtaining training data in a large model training data generation method according to an embodiment of the present disclosure.
[0017] Figure 7 is a block diagram of a large model training data generation apparatus according to an embodiment of the present disclosure.
[0018] Figure 8 is a structural schematic diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0019] Embodiments of the present disclosure will be described more fully hereinafter with reference to the accompanying drawings. While several embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be construed as being limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and fully convey the scope of the present disclosure to those skilled in the art. It should be understood that the drawings and embodiments of the present disclosure are only for illustrative purposes and are not intended to limit the scope of protection of the present disclosure.
[0020] It should be understood that each of the steps recited in the method embodiments of the present disclosure can be executed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit the execution of the steps shown. The scope of the present disclosure is not limited in this respect.
[0021] The term “comprising” and variations thereof as used herein are open-ended, that is, “including but not limited to”. The term “based on” is “based, at least in part, on”. The term “one embodiment” means “at least one embodiment”; the term “another embodiment” means “at least one additional embodiment”; the term “some embodiments” means “at least some embodiments”. Related definitions will be given in the description below.
[0022] It should be noted that the terms “first”, “second”, and the like mentioned in the present disclosure are only used to distinguish different apparatuses, modules or units, and are not intended to limit the order or interdependence of the functions performed by these apparatuses, modules or units.
[0023] It should be noted that the adjectives “one”, “more than one” mentioned in the present disclosure are illustrative and not limiting, and those skilled in the art should understand that “one or more” should be understood unless the context clearly indicates otherwise.
[0024] Names of messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information.
[0025] It can be understood that, before using the technical solutions disclosed in the embodiments of the present disclosure, the type, use range, use scenario, etc. of the personal information involved in the present disclosure should be informed to the user and the authorization of the user should be obtained through appropriate means according to relevant laws and regulations.
[0026] For example, in response to receiving the active request of the user, prompt information is sent to the user to explicitly prompt the user that the operation requested to be performed will require obtaining and using the personal information of the user. Thus, the user can voluntarily choose whether to provide the personal information to the software or hardware such as electronic devices, application programs, servers or storage media that perform the operation of the technical solutions of the present disclosure according to the prompt information.
[0027] As an optional but non-limiting implementation manner, in response to receiving the active request of the user, the manner of sending prompt information to the user may, for example, be a pop-up window manner, and the prompt information may be presented in the form of text in the pop-up window. In addition, the pop-up window may also carry selection controls for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0028] It can be understood that the above notification and user authorization process is only illustrative, and does not limit the implementation manners of the present disclosure, and other manners that meet the relevant laws and regulations can also be applied to the implementation manners of the present disclosure.
[0029] At the same time, it can be understood that the data involved in the present technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the relevant laws and regulations and relevant provisions.
[0030] In related technologies, as the core technology of multi-modal artificial intelligence, visual language models have significant advantages in understanding complex scenes and performing fine-grained tasks. However, general visual language models lack deep understanding of specific industry knowledge and cannot accurately identify subtle and complex defects or violations in various vertical fields, i.e., the application effect is not good. In order to realize the effective landing of artificial intelligence visual detection, it is necessary to train or fine-tune VLM detection models that belong to specific fields and are highly customized. Therefore, obtaining massive, high-quality and diversified labeled training data is a core problem of artificial intelligence visual detection.
[0031] In view of this, to solve the above problems, the embodiment of the disclosure provides a large model training data generation method, device, medium and electronic equipment, so as to intelligently convert user demand and visual material into large-scale, high-quality and diversified detection question and answer pairs (QA) through a highly structured and automated process, that is, through the embodiment of the disclosure, visual detection related training data set can be more efficiently obtained.
[0032] The disclosure will be further explained and described below in conjunction with the drawings.
[0033] Figure 1 It is a flowchart of a large model training data generation method according to the embodiment of the disclosure. The large model training data generation method can be applied to an electronic device with processing capability, such as a terminal or a server. And the large model training data generation method can be executed by a large model training data generation device, wherein the large model training data generation device can be realized by software and / or hardware, and the software and / or hardware can be configured in the electronic device. Refer to Figure 1 The large model training data generation method can include the following steps.
[0034] In step S110, an initial image is received, and a detection target corresponding to the initial image is determined based on a detection target library.
[0035] In the embodiment of the disclosure, the initial image can be a material image, which can be a picture or video material to be detected, that is, when the user uploads a material image, the embodiment of the disclosure can directly take it as an initial image; when the user uploads a video material, the embodiment of the disclosure can sample and frame process the video material to obtain a plurality of video frames, and take these video frames as initial images. Here, the picture or video material to be detected uploaded by the user can be original visual material.
[0036] In the embodiment of the disclosure, the original visual material can be uploaded by the user, or can be inspection data extracted by searching a database, which can be extracted based on the user's demand. For example, the user's demand is "generate detection data about urban road traffic scene", and the inspection data is visual data related to urban road traffic.
[0037] As an optional way, in the case of receiving a plurality of initial images, the embodiment of the disclosure can pre-process the initial images, that is, screen the initial images. In this process, the embodiment of the disclosure can screen out high-quality pictures or videos highly related to the task scene through the way of large model coarse screening and visual model fine screening.
[0038] For example, in the coarse screening process, the embodiments of the present disclosure can use a large model to quickly and preliminarily screen the data, and eliminate samples that obviously do not meet the requirements, so as to improve the efficiency of subsequent training data generation, such as eliminating images that are blurred, unclear, or have obvious defect problems. Fine screening can be performed on the basis of coarse screening, that is, the embodiments of the present disclosure can perform more in-depth and fine analysis on the data through a visual model to achieve higher screening accuracy. Through the cooperative work of coarse screening and fine screening, the embodiments of the present disclosure can ensure that the data processing is both efficient and accurate.
[0039] In the process of fine screening the initial images, the embodiments of the present disclosure can screen the initial images based on the requirement information to obtain candidate images related to the requirement information. On this basis, the detection target corresponding to the candidate images is determined based on the detection target library. That is, the embodiments of the present disclosure can screen a plurality of initial images according to the requirement input by the user to select images that meet the actual requirements.
[0040] Taking the above example, the requirement information input by the user is "generate detection data about urban road traffic scenes", and in the process of screening the initial images, the embodiments of the present disclosure can remove data unrelated to urban roads to screen out high-quality visual materials related to the task scene.
[0041] It should be noted that in the process of fine screening the initial images, the embodiments of the present disclosure can determine the correlation degree between the initial images and the requirement information, and when the correlation degree is less than a preset threshold, the initial images are eliminated; when the correlation degree is greater than the preset threshold, the embodiments of the present disclosure can grade the initial images, and the higher the correlation degree, the higher the grade of the initial images. Subsequently, the detection strategy corresponding to the initial images can be generated in combination with the target description information, the detection target and the correlation degree grading, and the detection strategy can also be referred to as a detection rule. For example, the first correlation degree between the first initial image and the target requirement is greater than the second correlation degree between the second initial image and the target requirement, and therefore the information involved in the first initial image is more detailed when generating the corresponding detection strategy, that is, the first initial image can have a higher level of detail than the second initial image.
[0042] As another optional way, the embodiments of the present disclosure can determine the detection target corresponding to each initial image based on the detection target library, wherein the detection target library can be generated based on the user requirement, that is, in response to the received requirement information, the embodiments of the present disclosure can construct a detection target library related to the requirement information.
[0043] Upon receiving the requirement information, the embodiments of the present disclosure can analyze the requirement information, that is, receive the core task requirement of the user in the form of natural language. For example, the requirement information can be "generate detection data about urban road traffic scenes", or can be "generate data for identifying the state of goods on the warehouse shelves".
[0044] Exemplarily, the embodiment of the present disclosure can utilize a large language model (LLM) to perform semantic analysis on the demand information, and based on the semantic analysis result, the macro target and boundary of the task can be accurately captured.
[0045] Based on this, the embodiment of the present disclosure can perform a core element extraction operation, that is, extract a plurality of key elements from the semantic obtained by analysis, and through these key elements, a complete detection task can be constructed. By extracting the core elements in the demand information, the fuzzy detection concept can be converted into machine understandable and operable semantic architecture, thereby providing clear guidance for the construction of the subsequent detection target library.
[0046] Among them, the key elements can include at least one of a detection scene, a detection subject, a detection behavior, and a detection target. Exemplarily, the detection scene can be a construction site, a road, etc., the detection subject can be a worker, a vehicle, etc., the detection behavior can be wearing, illegal parking, etc., and the detection target can be a safety helmet, a no-parking area, etc.
[0047] In addition, in the process of constructing the detection target library, the embodiment of the present disclosure finds and obtains the specification text associated with the semantic analysis according to the semantic analysis result of the demand information, and through processing, extracting, understanding, etc. of the specification text, the above-mentioned plurality of key elements related to the demand information can be obtained.
[0048] In summary, the detection target library can be constructed based on the demand information input by the user. When receiving new demand information, the embodiment of the present disclosure can update the detection target library based on the new demand information, that is, add the hierarchical information related to the demand information. Conversely, when the received demand information does not involve new demand, the embodiment of the present disclosure can reuse the hierarchical information of the detection target library corresponding to the demand information.
[0049] As can be seen, the detection target library can include a plurality of sets of hierarchical information, and each set of hierarchical information can include at least one of detection scene information, detection field information, and detection target information. Among them, the detection scene information refers to the specific detection scene, that is, which scene is detected, that is, according to the user demand, the embodiment of the present disclosure can define the macro detection scene, such as the detection scene can be a city road, an indoor scene, an outdoor factory area, etc.
[0050] The detection field information is to detect the category of the target in each scene. For example, for the "urban road" scene, the detection type (detection field) can be divided into "road equipment detection", "road traffic detection", "personnel safety detection", etc. That is, the detection field information can be used to limit which type of detection is performed on the specified scene, and in order to better implement the detection of the detection scene, the detection field can be further divided into a first-level label and a second-level classification.
[0051] The detection target information can be a further refinement of the detection field, that is, further refining the detection type into specific and operable detection targets / monitoring purposes. For example, the detection target information can be "check whether the motorcycle and electric vehicle passengers wear safety helmets", "detect whether the warning signs in the road construction area are set according to the specification", etc.
[0052] In order to better understand the structure of the detection target library, the present embodiment of the disclosure gives the following Table 1: Table 1
[0053] Based on Table 1, for the same scene "outdoor factory area", the fields are different, and the corresponding inspection purposes (detection targets) are also different. The above detection scene information, field information and target information can be structured information obtained by analyzing the demand information.
[0054] After the above structured information is obtained by carding, the present embodiment of the disclosure can store these structured information into the detection target library to form a standard and reusable knowledge system. It can be seen that by carding the present embodiment of the disclosure, the multi-level classification and scene of each requirement can be obtained, such as obtaining first-level, second-level and third-level classification and scene. On this basis, the inspection purposes under each scene are constructed to obtain the detection target library.
[0055] It should be noted that the detection target library in the present embodiment of the disclosure can be generated in real time according to the initial image and the demand information, that is, when the demand information input by the user is received, the multi-level information of the detection target library corresponding to the demand information is generated.
[0056] Alternatively, the detection target library can also be pre-constructed, that is, the user can generate a complete detection target library by inputting a large amount of demand information, and in the subsequent process of generating training data, the level information related to the demand information can be directly searched from the detection target library.
[0057] The embodiment of the present disclosure can extract structured information such as scenes, types and detection targets based on demand information, so as to convert scattered implicit knowledge and unstructured documents into structured instructions that can be accurately executed by devices. In the conversion process, the embodiment of the present disclosure can convert fuzzy concepts into a series of explicit and operable detection points to solve the problem of unclear target or unclear definition in the model training process from the source, so as to provide standard, traceable, high-quality and unambiguous input for the entire data set construction.
[0058] As another optional way, when determining the detection target corresponding to the initial image based on the detection target library, the embodiment of the present disclosure can identify the scene of the initial image to obtain the scene information corresponding to each initial image. On this basis, the scene information is used to determine the detection target matching the initial image in the detection target library. That is, the embodiment of the present disclosure can find the detection target matching the initial image from the detection target library by identifying the scene of the initial image, so as to match the structured knowledge with the unstructured visual content.
[0059] For example, the embodiment of the present disclosure can analyze each initial image based on the VLM model to obtain the scene information corresponding to each initial image, which can be used as the label of the corresponding initial image. For example, by identifying the first initial image, it is determined that the scene information of the first initial image is a city road scene, and the scene label of the first initial image can be "city road", so as to realize the image scene identification and labeling operation.
[0060] After identifying the scene information corresponding to the initial image, the embodiment of the present disclosure can perform a detection target matching operation, that is, determine the detection target matching the initial image in the detection target library based on the scene information. In other words, in the identified scene, the specific detection target defined in the detection target library is further matched in the image.
[0061] For example, by identifying the first initial image, it is determined that the scene matching the first initial image is a city road, and by searching the detection target library, it is determined that the target information corresponding to the scene "city road" is multiple. On this basis, the embodiment of the present disclosure can determine the detection target matching the initial image by performing semantic recognition on the initial image, such as "checking whether the driver of the motorcycle and electric vehicle violates the rule of carrying multiple people".
[0062] The embodiment of the present disclosure can obtain the detection target corresponding to each initial image by matching the image scene purpose library. Here, the detection target can also be referred to as the rule of the detection target library, that is, based on the detection target library, the embodiment of the present disclosure can determine the reference detection strategy (detection target) corresponding to each initial image.
[0063] In step S120, target description information of the initial image is extracted, and a detection strategy corresponding to the initial image is generated based on the target description information and the detection target.
[0064] As an optional approach, after receiving the initial image, this embodiment of the present disclosure can, on the one hand, determine the detection target matching each initial image through a detection target library, and on the other hand, extract the target description information of the initial image, that is, perform descriptive analysis on the initial image. Here, the target description information can also be called objective information, which can be a pure and inviolable visual fact in the initial image. This target description information can provide a realistic basis for subsequent generation of context-related detection tasks.
[0065] It is understood that, in the process of extracting target description information from the initial image, embodiments of this disclosure can generate initial description information for the initial image, wherein the initial description information may be textual information obtained to describe the initial image. Based on this, the initial description information is processed to obtain target description information, that is, unique and objective information is selected from the initial image and used as the target description information.
[0066] For example, embodiments of this disclosure can generate a purely descriptive caption for each initial image based on a VLM model, without subjective inference. This caption can serve as initial descriptive information. In other words, the initial descriptive information is a description of the things included in the initial image.
[0067] For example, for such Figure 2 The initial description information obtained by extracting the image information of the first initial image 201 shown can be: "This image is an indoor work scene under renovation and construction. There are three workers in the image, and construction materials are scattered on the ground. The three workers are installing building materials. They seem to be observing or pointing at the building materials piled up next to them. The one on the right is a woman wearing a baseball cap, a safety vest, and a short-sleeved shirt; the man in the middle is dressed neatly and seems to be smoking; the man on the left is wearing a safety helmet, gloves, and holding tools."
[0068] To ensure the uniqueness and authenticity of the extracted target description information, the embodiments of the present disclosure can process (check) the visual information on the basis of the initial description information of the initial image, that is, detect the speculative information and ambiguity information in the initial description information, and delete these information to obtain the target description information. In the process of processing the visual information, the embodiments of the present disclosure can determine the uncertain description information in the visual information, such as the above examples of “they seem to be observing or pointing to the building boards piled aside” and “he seems to be smoking”, and “seem” and “seem” all represent uncertain meanings. Such expressions may be the result of the large model acquiring the semantics of the initial image by speculation, rather than the target description information of the image, that is, with speculative meaning. In order to ensure the uniqueness and authenticity of the target description information, the embodiments of the present disclosure can screen out ambiguous expressions or speculative expressions.
[0069] Based on the above examples, the initial description information is processed to obtain the target description information, which can be “an indoor working scene under renovation and construction, there are three workers in the picture, and there are renovation materials scattered on the ground. Among them, the three workers are installing boards, the female on the right side is wearing a baseball cap, a safety vest and a short sleeve; the male in the middle is wearing a neat shirt; the male on the left side is wearing a safety helmet, wearing gloves and holding a tool”. Figure 2 The target description information obtained by processing the initial description information can be “an indoor working scene under renovation and construction, there are three workers in the picture, and there are renovation materials scattered on the ground. Among them, the three workers are installing boards, the female on the right side is wearing a baseball cap, a safety vest and a short sleeve; the male in the middle is wearing a neat shirt; the male on the left side is wearing a safety helmet, wearing gloves and holding a tool”.
[0070] Through the processing of the initial description information, the unique target description information of each initial image can be screened. It can be seen that by purifying the generated initial description information (caption), key and clear objective fact elements can be screened out.
[0071] It should be noted that in the process of processing the initial description information, the embodiments of the present disclosure can determine whether there is an obviously unconventional description in the initial description information. If such a description exists, it can be screened out. For example, by extracting the second initial image, the initial description information obtained includes “there are two people in the picture, and the two people are walking in the rain”, but by identifying the ground, it is determined that it is in a dry state. It can be seen that the initial description information at this time has a description that violates the convention, that is, the “rain” in the picture may be dust. In order to ensure the uniqueness and accuracy of the target description information, it can be screened out.
[0072] In the embodiments of the present disclosure, the extracted target description information can be used as the fact anchor point of the initial image, that is, as an unforgeable visual basis. For example, the target description information of the first initial image is “there are potholes, water and snow on the road”, and “there are potholes, water and snow on the road” can be used as the target description information of the first initial image.
[0073] In summary, on the one hand, the present disclosure can guide and match the visual material (initial image) to obtain the detection target, and on the other hand, the present disclosure can generate the corresponding initial description information (caption) by describing the visual material, and further extract the fact anchor point that cannot be violated. Here, the detection target and the target description information as the guidance of subsequent training data generation can not only make the subsequent generated content have a clear business orientation, but also ensure that the generated content is rooted in real visual evidence, which can to some extent prevent the common groundless speculation and content hallucination of large language models from the source. That is, the detection target and the target description information are the core cornerstone of guaranteeing the generation of high-quality and high-credibility data.
[0074] As another optional way, after obtaining the detection target and the target description information corresponding to each initial image, the present disclosure can generate the detection strategy corresponding to the initial image according to the target description information and the detection target. For example, the present disclosure can generate the detection strategy corresponding to the initial image by imitation and expansion, wherein the imitation and expansion can be based on the target description information and the detection target.
[0075] In other words, the present disclosure can generate the actual detection strategy by dynamically and finely expanding and enriching the detection target based on the detection target and the objective fact. For example, the detection target obtained by detecting the first initial image is "check whether the road surface has water", and the fact anchor point obtained by extracting the target description information of the first initial image is "the road has potholes, water, and snow". Based on this, the detection strategy generated based on "check whether the road surface has water" and "the road has potholes, water, and snow" can be "check whether the road has water, check whether the road has potholes, and check whether the road has snow".
[0076] Optionally, during the process of generating the detection strategy corresponding to the initial image according to the target description information and the detection target, the present disclosure can determine the key information of the detection target, and expand the target description information based on the key information to obtain the detection strategy corresponding to the initial image. Here, the key information can be the risk point of the detection target, and the present disclosure can generate the instantiation and diversification of the detection strategy by taking the risk point as a reference.
[0077] The determination of the detection target information can be obtaining a structured key point of the detection target, that is, abstracting the detection target into a structured data template, which can include inherent attributes and relationship attributes. In addition, the target description information can be multi-dimensional, such as the target description information at least including object recognition information, spatial information, state information, and environmental context information. The embodiments of the present disclosure can generate a detection strategy based on logical matching and instantiation of the template attributes and the target description information. That is, combining the abstract attributes in the template with the specific information in the image, a specific and executable detection instruction is generated through a predefined logical relationship.
[0078] For example, the embodiments of the present disclosure can obtain a detection target template (detection target), which defines at least one inherent attribute and / or relationship attribute of an object to be detected; obtain target description information of an initial image, which can include object information, spatial relationship information, and state information extracted through image recognition technology; and perform coupling analysis on the attributes of the detection target template and the target description information of the initial image to obtain a detection strategy.
[0079] Specifically, the inherent attributes are compared with the state information of the corresponding object in the initial image to generate a first type of detection strategy about the integrity or state compliance of the object; the relationship attributes are matched with the spatial relationship information and object information in the image to generate a second type of detection strategy about the compliance of the relationship between the object and the surrounding environment; and the initial image corresponding detection strategy is constructed by the first type of detection strategy and the second type of detection strategy.
[0080] As an example, the detection target is to detect whether the fire hydrant has a pressure gauge and whether the paint color is red (inherent attribute), and whether there is an obstacle within 1.2 meters (relationship attribute). The target description information of the initial image is that the pressure gauge pointer of the fire hydrant is in the red area, the paint color is partially rusted, and a wooden tray is identified 0.8 meters in front of it. The first type of rule generated by combining the inherent attributes and the state information can be whether the pressure gauge pointer of the fire hydrant is in the red area, and the second type of rule generated by combining the relationship attributes, the spatial information, and the object information can be whether there is an article within 1.2 meters around the fire hydrant.
[0081] The embodiments of the present disclosure can perform a copy operation by taking the detection target as a template and taking the target description information as filling content, so as to automatically expand a small amount of detection target into a large number of specific and instantiated detection strategies, thereby ensuring that each generated detection strategy is highly bound to the visual material in terms of semantics and details. That is, the embodiments of the present disclosure can realize intelligent expansion from “one detection target” to “N scene-based instance detection strategies” through the detection target and the target description information, while suppressing model recognition errors and hallucinations, the embodiments of the present disclosure can greatly improve the quality and diversity of the generated data.
[0082] It should be noted that in the process of generating the detection strategy corresponding to the initial image, the embodiments of the present disclosure can directly follow or copy the rules, that is, the detection target can be directly used to generate the detection strategy, or the detection target can be copied and learned to obtain a detection strategy unique to the initial image. At the same time, the large model can follow or refer to the prompt example to generate the prompt information corresponding to the initial image.
[0083] In addition, in the process of copying the detection strategy, the detection target can carry label information, which can be the hierarchical information in the detection target library. As introduced above, the same scene can correspond to multiple detection categories, such as the "outdoor factory area" scene, which can correspond to the cleaning state category under environmental health, the behavior category under rules and regulations, and the behavior category under safety monitoring, etc. By configuring label information for the detection target, a detection strategy unique to the initial image can be more accurately generated.
[0084] By copying and learning the detection targets of various categories in the scene corresponding to the initial image, a detection strategy unique to the initial image can be generated based on the target description information, such as copying and learning each detection target of the outdoor factory area scene in the detection target library, and combining the target description information of the initial image, the final generated detection strategy is "detect whether there are garbage and sundries around the gatekeeper room in the factory area; detect whether there are garbage and sundries on the ground at the factory gate".
[0085] In step S130, the detection strategy is used to generate a question and answer pair corresponding to the initial image, and the training data is constructed according to the question and answer pair.
[0086] As an optional way, after generating the detection strategy corresponding to the initial image, the embodiments of the present disclosure can generate the question and answer pair corresponding to the initial image based on the detection strategy, and construct the training data according to the question and answer pair, so as to convert the processed information into formatted training data. Here, the initial image is not the same, and the training data constructed according to the initial image is also not the same.
[0087] In order to ensure that the generated question and answer pair is uniform in format and natural in style, the embodiments of the present disclosure can randomly extract a specified number of question and answer pair examples (prompt cases) from the high-quality data provided by the user or the preset excellent examples, and learn these question and answer pair examples to obtain the ability to output the specified paradigm. That is, the embodiments of the present disclosure can perform a random learning (Few-shot Learning) operation to dynamically master the prompt (prompt information) output paradigm corresponding to the initial image.
[0088] On this basis, the embodiment of the disclosure can generate a target question of a specified paradigm according to the detection strategy, and detect the initial image for the target question to obtain a target answer of the specified paradigm. That is, the embodiment of the disclosure can generate a target question (Question, Q) according to the detection strategy and the question and answer pair example, and based on this, the target answer (Answer, A) can be generated through interaction with a large model.
[0089] In the data synthesis process, in order to avoid the problem of rigid and generalized output content format, the embodiment of the disclosure introduces a random learning method, that is, the embodiment of the disclosure does not use a fixed output template, but dynamically learns a small amount of high-quality data samples provided by the user or dynamically learns preset excellent examples before generating data. Through the imitation learning of these examples, the embodiment of the disclosure can automatically master and reproduce the specific question and answer style and data format expected by the current image, so as to ensure that the generated training data has diversity.
[0090] In other words, the embodiment of the disclosure can make the batch-generated question and answer pairs not only accurate in content, but also diverse in language paradigm by introducing a random learning method, that is, the automatically generated training data can directly meet the strict requirements of downstream model training without manual adjustment or manual post-tuning.
[0091] In the process of generating a question and answer pair corresponding to the initial image based on the detection strategy, the embodiment of the disclosure can generate prompt information (prompt) according to the detection strategy, which can include the detection strategy, output format, instruction information and matters needing attention, etc. On this basis, the embodiment of the disclosure can input the prompt information and the initial image into the large model to obtain the answer of the specified paradigm.
[0092] In summary, the embodiment of the disclosure can combine the detection strategy corresponding to the initial image with the output format (specified paradigm) learned by the model to automatically generate a large number of high-quality detection question and answer pairs. Based on this, the embodiment of the disclosure can bind the initial image, the prompt information and the automatically labeled detection result to generate a complete training data triple, and store it in the final training data set in batches.
[0093] As an example, when receiving a first initial image 201 as shown in Figure 2 The embodiment of the disclosure can determine the detection target corresponding to the first initial image 201 based on the detection target library, and at the same time, the embodiment of the disclosure can extract the target description information of the first initial image 201, and comprehensively generate the detection strategy corresponding to the first initial image 201 according to the target description information and the detection target. On this basis, the question and answer pair corresponding to the first initial image 201 can be generated based on the detection strategy, and the training data of the first initial image 201 can be constructed according to the question and answer pair.
[0094] It can be known from the above introduction that the training data can include images, prompt information and detection results, so the finally obtained training data of the first initial image 201 can include the first initial image 201, the prompt information corresponding to the first initial image 201 and the corresponding detection result, as shown in detail in Figure 3 Figure 3 It can be known that the prompt information (inspection prompt) of the first initial image 201 can include role description information in addition to the inspection rule (detection strategy), output format and description information. After obtaining the prompt information, the embodiment of the disclosure can input it into the large model together with the first initial image 201, and the detection result (answer) can be obtained.
[0095] By comparing Figure 2 It can be known that Figure 2 There is a violation, and the violation is because someone does not wear a safety helmet, and the specific number of violators is 2, so Figure 3 The generated inspection result is completely correct.
[0096] As another example, when receiving the second initial image 401 as shown in Figure 4 The embodiment of the disclosure can determine the detection target corresponding to the second initial image 401 based on the detection target library, and at the same time, the embodiment of the disclosure can extract the target description information of the second initial image 401, and the detection strategy corresponding to the second initial image 401 can be generated according to the target description information and the detection target. On this basis, the question and answer pair corresponding to the second initial image 401 can be generated based on the detection strategy, and the training data of the second initial image 401 can be constructed according to the question and answer pair.
[0097] As above, the training data of the second initial image 401 as shown in Figure 5 may include the second initial image 401, the prompt information corresponding to the second initial image 401 and the corresponding detection result. Based on Figure 5 It can be known that the prompt information (inspection prompt) of the second initial image 401 can include role description information and output requirements in addition to the detection rule (detection strategy) and output format. After obtaining the prompt information, the embodiment of the disclosure can input it into the large model together with the second initial image 401, and the detection result (answer) can be obtained.
[0098] By comparing Figure 4 It can be known that Figure 4 The personnel equipment in the image conforms to the regulations, that is, two high-altitude operation personnel wear safety helmets, safety ropes, work gloves and work clothes, and in addition, there are rust marks on the surface of the outdoor equipment, so Figure 5 The generated inspection result is completely correct.
[0099] The format of the detection result is generated based on the output format in the prompt information, and the output format is generated through a random learning operation. The output format of the image at different times is not the same. Even for the same image, the output format obtained by the corresponding acquisition may also be different. In this way, the generalization ability of the large model in actual application can be improved, and the recognition rate for low-frequency events is high. To some extent, the diversity of the training data can be ensured.
[0100] As another optional mode, after obtaining the training data, the embodiment of the disclosure can evaluate the training data. Specifically, the embodiment of the disclosure can determine whether the detection strategy in the training data matches the content of the corresponding initial image, that is, the matching degree of the two is obtained. When the matching degree exceeds the preset matching degree, it is determined that the generated data is high-quality data. That is, when it is determined that the detection strategy conforms to the scene in the picture and can obtain a unique objective answer according to the clues in the picture, it is determined that the detection strategy matches the image content.
[0101] Optionally, the embodiment of the disclosure can determine whether the detection strategy conforms to the actual application requirements of the detection scene, that is, when it is determined that the detection strategy is applicable to the corresponding image and can also be applied to the same scene of the image, it is determined that the training data meets the requirements.
[0102] Optionally, the embodiment of the disclosure can determine whether the detection strategy in the prompt information is limited to the detection target library, that is, whether it can be dynamically expanded according to the material. If it can be dynamically expanded, it is determined to meet the requirements, otherwise, if it is limited to the detection target library, it is determined not to meet the requirements. Optionally, the embodiment of the disclosure can determine whether the prompt information is directly available or can be directly delivered through local fine-tuning (modification time cost <= 5 minutes). If it is satisfied, it is determined to meet the requirements.
[0103] The above-mentioned multiple evaluation modes can be used comprehensively, or can be used individually according to requirements. The specific mode for evaluating the training data is not limited here, and can be selected according to the actual situation.
[0104] As another optional mode, after obtaining the multiple training data, the embodiment of the disclosure can train the large model based on the training data. Here, the large model can be used to detect the to-be-detected image based on the requirement information to obtain the detection result and the detection basis.
[0105] Since the large model is generated based on high-quality data and diversified instructions, the recognition accuracy, defect detection capability and environmental adaptability of the visual language model in a complex detection scene can be significantly enhanced, thereby promoting the visual language model from general perception to automatic intelligent detection in multiple professional scenes.
[0106] As a specific implementation method, such as Figure 6 As shown, upon receiving a detection request (request information) described in natural language, embodiments of this disclosure can understand the detection request and construct a detection target library based on this understanding. Furthermore, upon acquiring the image / video material to be detected (initial image), embodiments of this disclosure can perform scene recognition on the image / video material to be detected, matching the detection target in the detection target library based on the recognition results. Additionally, embodiments of this disclosure can perform visual understanding on the image / video material to be detected, extracting its target description information.
[0107] Based on this, the present disclosure can generate an actual detection strategy by combining the detection target and target description information, and perform large-scale production of data based on randomly learned question-answer pair examples (prompt samples) combined with the detection strategy. The randomly learned question-answer pair examples are obtained through a high-quality question-answer pair example library. Based on this, the embodiments of the present disclosure can obtain training data triples, which may include an image / video, a detection prompt, and a detection result.
[0108] The embodiments of this disclosure, through a highly structured and automated processing chain, can accurately match abstract and complex detection task requirements with specific and diverse visual content, and creatively extend them. In this way, detection QA datasets can be automatically generated on a large scale, with high quality and customized for specific scenarios to a certain extent.
[0109] Based on the same inventive concept, this disclosure also provides an apparatus for generating large model training data. Figure 7 This is a block diagram illustrating a large model training data generation apparatus 700 according to an exemplary embodiment, such as... Figure 7 As shown, the large model training data generation device 700 may include a determination module 710, a rule acquisition module 720, and a construction module 730.
[0110] The determining module 710 is used to receive an initial image and determine the detection target corresponding to the initial image based on the detection target library; The rule acquisition module 720 is used to extract target description information from the initial image and generate a detection strategy corresponding to the initial image based on the target description information and the detection target. The construction module 730 is used to generate question-answer pairs corresponding to the initial image based on the detection strategy, and to construct the training data based on the question-answer pairs.
[0111] Optionally, the rule acquisition module 720 can also be used to generate initial description information of the initial image, wherein the initial description information is textual information obtained to describe the initial image; and to process the initial description information to obtain the target description information.
[0112] Optionally, the rule obtaining module 720 can also be configured to filter out the speculative information and the ambiguous information in the initial description information to obtain the target description information.
[0113] Optionally, the rule obtaining module 720 can also be configured to determine the key information of the detection target, and perform copy expansion on the target description information based on the key information to obtain the detection strategy corresponding to the initial image.
[0114] Optionally, the large model training data generation apparatus 700 can further include: The learning module is configured to randomly obtain a specified number of question and answer pair examples, and learn the question and answer pair examples to obtain the ability to output a specified paradigm.
[0115] Optionally, the construction module 730 can also be configured to generate a target question of the specified paradigm based on the detection strategy, and detect the initial image for the target question to obtain a target answer of the specified paradigm.
[0116] Optionally, the determination module 710 can also be configured to identify the scene of the initial image to obtain scene information, and determine the detection target matched with the initial image in the detection target library based on the scene information.
[0117] Optionally, the construction module 730 can also be configured to construct the detection target library in response to the received requirement information.
[0118] Optionally, the determination module 710 can also be configured to filter at least one initial image based on the requirement information to obtain a candidate image related to the requirement information, and determine the detection target corresponding to the candidate image based on the detection target library.
[0119] Optionally, the detection target library includes multiple sets of hierarchical information, and each set of hierarchical information includes at least one of: detection scene information; detection field information; detection target information.
[0120] Optionally, the large model training data generation apparatus 700 can further include: The training module is configured to train a large model based on the training data, and the large model is configured to detect a to-be-detected image based on requirement information to obtain a detection result and a detection basis.
[0121] The detection target library in the embodiments of the present disclosure can support detection requirements of multiple industries and can quickly respond to customization of training data, realize efficient development, iterative optimization and highly personalized deployment of visual language models, to meet the needs of multiple segmented markets. In addition, the large-scale, high-quality and diversified training data synthesized above can provide key data support for visual language models in the research fields of few-shot learning, explainability and human-machine collaboration, and therefore can create conditions for constructing industry-level detection data sets and evaluation benchmarks to a certain extent.
[0122] The embodiments of each module in the generation apparatus 700 of the large model training data can refer to the related embodiments of the above method, and the present embodiment will not be repeated here.
[0123] The embodiments of the present disclosure also provide a computer readable medium having a computer program stored thereon, which, when executed by a processing apparatus, implements the steps of the above generation method of large model training data.
[0124] The embodiments of the present disclosure also provide a computer program product comprising a computer program, which, when executed by a processor, implements the steps of the above generation method of large model training data.
[0125] The embodiments of the present disclosure also provide an electronic device comprising: a storage device having a computer program stored thereon; a processing apparatus configured to execute the computer program in the storage device to implement the steps of the above generation method of large model training data.
[0126] Reference will now be made to Figure 8 which shows a structural schematic diagram of an electronic device 800 (terminal device or server) suitable for implementing the embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure can include but is not limited to mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Personal Computers), PMPs (Portable Multimedia Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), and the like, and fixed terminals such as digital TVs, desktop computers, and the like. Figure 8 The electronic device shown is only an example and should not impose any limitation on the functions and use range of the embodiments of the present disclosure.
[0127] As Figure 8As shown, the electronic device 800 can include a processing device (e.g., a central processor, a graphics processor, etc.) 801 that can perform various suitable actions and processes in accordance with programs stored in a read-only memory (ROM) 802 or loaded into a random access memory (RAM) 803 from a storage device 808. Various programs and data required by the electronic device 800 for operation are also stored in the RAM 803. The processing device 801, the ROM 802, and the RAM 803 are connected to each other by a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0128] Generally, the following devices can be connected to the I / O interface 805: input devices 806 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 807 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 808 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 809. The communication devices 809 can allow the electronic device 800 to communicate wirelessly or wired with other devices to exchange data. Although Figure 8 The electronic device 800 is shown with various devices, but it should be understood that not all of the shown devices are required to be implemented or present. More or fewer devices can alternatively be implemented or present.
[0129] In particular, the processes described above with reference to the flowcharts can be implemented as a computer software program according to embodiments of the present disclosure. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network through the communication devices 809, or installed from the storage devices 808, or installed from the ROM 802. When the computer program is executed by the processing device 801, the above-mentioned functions defined in the methods of embodiments of the present disclosure are performed.
[0130] It is noted that the aforementioned computer-readable medium of the present disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium can be, for example and without limitation, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program used by or in connection with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include a computer-readable program code transmitted by a computer-readable signal medium, in a baseband or as a part of a carrier wave. Such a computer-readable signal medium can take a variety of forms, including but not limited to, electro-magnetic, optical, or any suitable combination of the foregoing. The computer-readable signal medium can also be any computer-readable medium that is not a computer-readable storage medium and that can be used to carry or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained in the computer-readable medium can be transmitted by any suitable medium, including but not limited to, wire, cable, RF, etc., or any suitable combination of the foregoing.
[0131] In some embodiments, the terminal device, the server can communicate using any currently known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication of any form or medium (e.g., a communication network). Examples of the communication network include a local area network ("LAN"), a wide area network ("WAN"), an internetwork (e.g., the Internet), and an end-to-end network (e.g., an ad hoc end-to-end network), as well as any currently known or future developed network.
[0132] The aforementioned computer-readable medium can be included in the aforementioned electronic device; or can exist separately from the electronic device and not be assembled into the electronic device.
[0133] The computer readable medium described above carries one or more programs, when the one or more programs are executed by the electronic device, cause the electronic device to: receive an initial image, determine a detection target corresponding to the initial image based on a detection target library; extract target description information of the initial image, generate a detection strategy corresponding to the initial image according to the target description information and the detection target; generate a question and answer pair corresponding to the initial image based on the detection strategy, and construct the training data according to the question and answer pair.
[0134] Computer program code for carrying out operations of the present disclosure can be written in one or more programming languages or combinations of languages including object oriented programming languages such as Java, Smalltalk, C++ as well as conventional procedural programming languages such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0135] The flow diagrams and the block diagrams in the drawings are illustrations of architectures, functionalities, and operations of possible implementations of systems, methods, and computer program products according to various embodiments of present disclosure. In this regard, each block in the flow diagrams or block diagrams can represent a module, a procedure, or a part of code, which comprises one or more executable instructions for implementing the specified functions. It should also be noted in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks depicted in succession can in fact be executed substantially concurrently or the blocks can sometimes be executed in the reverse order, depending upon the functionality involved. It is also noted that each block in the block diagrams and / or flow diagrams and combinations of blocks in the block diagrams and / or flow diagrams can be implemented by dedicated hardware-based systems that perform the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0136] The modules involved in the embodiments of the present disclosure can be implemented in a software manner or in a hardware manner. In some cases, the name of the module does not constitute a limitation on the module itself.
[0137] The functionality described above can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, an example type of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Application-specific Integrated Circuits (ASICs), Application-specific Standard Products (ASSPs), System-on-a-chip (SOCs), Complex Programmable Logic Devices (CPLDs), etc.
[0138] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more of: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0139] The above description is only that of preferred embodiments of the present disclosure and a description of principles of technology applied. It should be understood by those skilled in the art that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and should also cover other technical solutions formed by any combinations of the above technical features or equivalent features without departing from the above disclosed concept. For example, the technical solutions formed by replacing the above features with the technical features disclosed in the present disclosure (but not limited to) having similar functions.
[0140] Further, while operations are depicted in a particular order, this should not be understood as requiring these operations to be performed in the particular order shown or in sequential order, as some can be performed in parallel or in any suitable order. Also, while a number of specific implementation details have been discussed, these should not be construed as limiting the scope of the disclosure. Certain features that are described in the context of separate embodiments can also be implemented in combination. Conversely, various features that are described in the context of a single embodiment can also be implemented separately or in any suitable subcombination. It will be appreciated that details of the foregoing embodiments, given only for purposes of illustration, are not intended to limit the scope of the present disclosure. As one of ordinary skill in the art will readily appreciate, the concepts, as generally taught herein, can be applied to a wide variety of alternative designs, structures, applications, and procedures. Therefore, the specific embodiments discussed above are not intended as being limiting; other embodiments can readily be devised in accordance with these teachings while remaining within the scope of the present disclosure.
[0141] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims. With respect to the devices in the above-described embodiments, in which various modules perform operations, the specific manner in which the operations are performed by the various modules has been described in detail in the embodiments relating to the method. Here, no detailed explanation will be given.
Claims
1. A method for generating large model training data, characterized in that, The method comprises: receiving an initial image, determining a detection target corresponding to the initial image based on a detection target library; extracting target description information of the initial image, and generating a detection strategy corresponding to the initial image according to the target description information and the detection target; generating a question and answer pair corresponding to the initial image based on the detection strategy, and constructing training data according to the question and answer pair.
2. The method of claim 1, wherein, The extraction of the target description information of the initial image comprises: generating initial description information of the initial image, wherein the initial description information is textual information obtained by describing the initial image; processing the initial description information to obtain the target description information.
3. The method of claim 2, wherein, The processing of the initial description information to obtain the target description information comprises: screening out speculative information and ambiguous information in the initial description information to obtain the target description information.
4. The method of claim 1, wherein, The generation of the detection strategy corresponding to the initial image according to the target description information and the detection target comprises: determining key information of the detection target, and expanding the key information to obtain the detection strategy corresponding to the initial image.
5. The method of claim 1, wherein, The method further comprises: randomly obtaining a specified number of question and answer pair examples; learning the question and answer pair examples to obtain the ability to output a specified paradigm.
6. The method of claim 5, wherein, The generation of the question and answer pair corresponding to the initial image based on the detection strategy comprises: generating a target question of the specified paradigm based on the detection strategy; detecting the initial image for the target question to obtain a target answer of the specified paradigm.
7. The method of claim 1 to 6, wherein, The determination of the detection target corresponding to the initial image based on the detection target library comprises: recognizing a scene of the initial image to obtain scene information; determining a detection target matching the initial image in the detection target library based on the scene information.
8. The method of claim 1 to 6, wherein, The method further comprises: constructing the detection target library in response to received requirement information.
9. The method of claim 8, wherein, The determination of the detection target corresponding to the initial image based on the detection target library comprises: screening at least one initial image based on the requirement information to obtain a candidate image related to the requirement information; determining the detection target corresponding to the candidate image based on the detection target library.
10. The method of claim 1 to 6, wherein, The detection target library comprises multiple sets of hierarchical information, and each set of hierarchical information comprises at least one of: detection scene information; detection field information; detection target information.
11. The method of claim 1 to 6, wherein, The method further comprises: training a large model based on the training data, wherein the large model is used to detect a to-be-detected image based on requirement information to obtain a detection result and a detection basis.
12. A large model training data generation apparatus, comprising: The device comprises: a determination module configured to receive an initial image, and determine a detection target corresponding to the initial image based on a detection target library; a rule extraction module configured to extract target description information of the initial image, and generate a detection strategy corresponding to the initial image according to the target description information and the detection target; a construction module configured to generate a question and answer pair corresponding to the initial image based on the detection strategy, and construct training data according to the question and answer pair.
13. A computer readable medium having stored thereon a computer program, characterized in that The computer program, which is executed by the processing means, implements the steps of the method according to any one of claims 1-11.
14. An electronic device, comprising: comprising: a storage device having stored thereon a computer program; processing means for executing the computer program in the storage device to implement the steps of the method according to any one of claims 1-11.
15. A computer program product comprising a computer program, characterized in that, The computer program, which is executed by the processing means, implements the steps of the method according to any one of claims 1-11. The computer program, which is executed by the processing means, implements the steps of the method according to any one of claims 1-11.
Citation Information
Patent Citations
Visual language large model training method and system for rich text image question and answer and rich text image question and answer method
CN119066178A
Retrieval method and device based on large model and intelligent agent
CN119202192A
Implementation method of visual question-answering system
CN119988659A
Method and device for multi-stage generation of medical image question and answer thinking chain data
CN120317386A
Logical text passage generation and retrieval for retrieval-augmented generation
US20250298962A1