Method and system for generating data set of sparse scene

Through object detection and adaptive image fusion technology, the problem of data acquisition difficulties in sparse scenes is solved, high-quality data sets are generated, data acquisition costs are reduced, and data acquisition costs are provided with rich data resources for model training.

CN119942164APending Publication Date: 2025-05-06CHINA TELECOM CLOUD TECH CO LTD
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202411698683.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-25
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In sparse scenarios, data acquisition is difficult, traditional data acquisition methods are expensive and are limited by user data privacy issues, making it difficult to achieve comprehensive data collection. Existing image fusion technology uses fixed parameter methods when processing images of different sources and different quality, resulting in the generated data set being inaccurate or comprehensive in some cases.

Method used

Through object detection technology, the target object in the image is identified and positioned, providing prompt words for the segmentation of the large model, and precisely segmenting the target object and background are combined with the large model segmentation technology. Adaptive image fusion technology is used to adjust the fusion method according to images of different sources and different quality to generate a more comprehensive and accurate data set.

Benefits of technology

Only a small number of target objects and background images can be used to generate a large number of new high-quality images, solving the problem of scarcity of scene data, providing rich data resources for subsequent model training, and reducing data acquisition costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942164A_ABST
    Figure CN119942164A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data generation of sparse scenes, and discloses a data set generation method and system of a sparse scene, and the method comprises the steps: inputting an image which is actually collected by a target sparse scene into a target detection model, and obtaining a detection frame of a target foreground and a background; using a preset segmentation large model to take the detection frames of the target foreground and the background as cue words to perform model reasoning, respectively obtaining masks of the target foreground and the background, and obtaining a target foreground image and background area coordinate information based on the masks of the target foreground and the background; and integrating the target foreground image and the background area coordinate information by using an adaptive image fusion technology to generate a new target sparse scene image, and adding the newly generated target sparse scene image into the training data set of the target detection model. According to the method, only a small number of target objects and background images are needed, a large number of new high-quality images can be generated, the problem of scarcity of scene data is solved, and rich data resources are provided for subsequent model training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data generation for sparse scenes, and in particular to a method and system for generating a data set for sparse scenes. Background Art

[0002] In actual production and life, data acquisition in some scenarios becomes difficult due to the difficulty in sensor deployment, high data collection costs, unstable data quality, etc. At the same time, traditional data collection methods are often costly and subject to user data privacy issues, making it difficult to achieve comprehensive data collection.

[0003] To solve this problem, a potential solution is to use image fusion technology to generate new data. Image fusion technology can integrate image data from different sources to generate an image with richer information and higher definition. This method can not only improve the quality of data, but also reduce the cost of data acquisition. Existing image fusion technology usually uses fixed parameters to fuse images of different sources and qualities, which may cause the generated data set to be inaccurate or incomplete in some cases. For example, when the quality of the source images varies greatly or the feature differences are obvious, the fixed parameter fusion method may not be able to fully take these differences into account, resulting in unsatisfactory fusion results. Summary of the invention

[0004] In view of this, the present invention provides a method and system for generating a data set for a sparse scene, so as to solve the problem of how to obtain rich and high-quality image data when there is less data for the sparse scene.

[0005] In a first aspect, the present invention provides a method for generating a data set for a sparse scene, the method comprising:

[0006] Input the real collected images of the target sparse scene into the target detection model to obtain the detection boxes of the target foreground and background;

[0007] Using a preset segmentation model, the detection frames of the target foreground and background are used as prompt words to perform model reasoning, and masks of the target foreground and background are obtained respectively, and coordinate information of the target foreground image and background area is obtained based on the masks of the target foreground and background;

[0008] Adaptive image fusion technology is used to integrate the target foreground image and background area coordinate information to generate a new target sparse scene image, and the newly generated target sparse scene image is added to the training data set of the target detection model.

[0009] The embodiment of the present invention provides a method for generating a data set for sparse scenes. Through target detection technology, it can accurately identify and locate targets in images, provide effective prompts for the segmentation of large models, and combine large model segmentation technology to accurately segment targets and backgrounds. According to images of different sources and different qualities, the fusion method can be adjusted adaptively and flexibly, thereby improving the authenticity of the fused image. Only a small number of target and background images are needed to generate a large number of new high-quality images, which solves the problem of scarce scene data and provides rich data resources for subsequent model training.

[0010] In an optional implementation, the target detection model is trained by the following process, including:

[0011] Annotate the target foreground and background of a preset small number of images collected from scenes with sparse targets;

[0012] Normalize the labeled data and generate a training set using data augmentation techniques;

[0013] Based on the training set, training is performed using a preset machine learning model, and the trained model is used as a target detection model. The target detection model outputs a detection frame for detecting the target foreground and background.

[0014] The embodiment of the present invention only needs to process a small number of images of the target sparse scene. Compared with large-scale data collection, it greatly reduces the manpower, material and time costs required in the data collection stage. By marking the target foreground and background, it provides a clear learning goal for the model. The model can accurately understand which areas are the target foreground and which are the background based on the marked information, which can effectively improve the accuracy and efficiency of target detection and the generalization ability of the model.

[0015] In an optional implementation, the method of using the preset segmentation model to use the detection boxes of the target foreground and background as prompt words to perform model reasoning, obtain masks of the target foreground and background respectively, and obtain target foreground image and background area coordinate information based on the masks of the target foreground and background, including:

[0016] Obtaining the background single channel, using a preset large segmentation model to use the detection frames of the target foreground and background as prompt words to perform model reasoning, and obtaining a mask of the target foreground single channel image and a mask of the background single channel image respectively;

[0017] The mask of the single-channel image of the target foreground is expanded to mask data with the same number of channels as the original real collected image, and the segmented target foreground image is obtained based on the original real collected image and the mask data with the same number of channels as the original real collected image, and the width and height values ​​of the target object in the target foreground image are calculated;

[0018] Get the minimum bounding rectangle corresponding to the background single-channel image, and get the center coordinates, width, and height of the minimum bounding rectangle.

[0019] The embodiment of the present invention uses the detection boxes of the target foreground and background as prompt words for the reasoning of the preset segmentation large model, which can provide clear location information for the model and effectively combine the advantages of target detection and image segmentation. Target detection provides the approximate location of the target (detection box), while image segmentation can be further refined to the pixel level to generate a mask to accurately divide the target foreground and background. This combination allows for a deeper understanding of the image and can be better applied to subsequent processing tasks.

[0020] In an optional implementation, when the width value or height value of the target object is greater than the width value or height value of the background, the size of the target object is scaled based on the scaling factor, and the width value and height value of the target object are updated after scaling.

[0021] In the embodiment of the present invention, when the size (width or height) of the target object is larger than the background size, the target object is scaled based on the scaling factor to ensure that the target object can be reasonably placed in the background area. Scaling the target object can make the ratio between the target and the background more coordinated. In a data set containing multiple targets and background combinations, if the size of the target object is too large, it may cause data inconsistency. By scaling the target object, all data can be made more uniform in size, which is convenient for unified processing and analysis.

[0022] In an optional implementation, the adaptive image fusion technology is used to integrate the target foreground image and the background area coordinate information to generate a new target sparse scene image, and the newly generated target sparse scene image is added to the training data set of the target detection model, including:

[0023] Performing data enhancement processing on the target foreground image to introduce new features, obtaining a target foreground image set, and determining a fusion position range in a random manner based on the center coordinates of a minimum circumscribed rectangle;

[0024] Based on the fusion position range, the target foreground image set is superimposed and fused with the real collected background image to obtain a fused image set, and the divergence value of each image in the fusion area of ​​the fused image set is calculated;

[0025] The images whose divergence values ​​in the fused image set are less than or equal to the first threshold are taken as the first fused image set;

[0026] The images whose divergence values ​​in the fused image set are greater than the first threshold and less than the second threshold are processed by a preset fusion method as the second fused image set;

[0027] Input the images in the fused image set whose divergence values ​​are greater than the second threshold into the fine-tuned diffusion model to obtain a third fused data set;

[0028] The first fused image set, the second fused image set and the second fused image set are added as newly generated target sparse scene images to the training data set of the target detection model.

[0029] The embodiment of the present invention implements the entire process from image fusion and divergence value calculation to adaptive image fusion technology processing of images according to different divergence value ranges, and finally adds the processed image set to the training data set, which helps to further optimize the image fusion effect and enrich the content of the training data set.

[0030] In an optional implementation, the process of fine-tuning the diffusion model includes:

[0031] Annotate the target foreground and background of a preset small number of images collected from scenes with sparse targets;

[0032] Gradually add Gaussian noise to the labeled image until the final pure noise image is obtained;

[0033] The noisy image is input into the preset noise prediction network, and the model is trained using the preset loss function to obtain a fine-tuned diffusion model.

[0034] The embodiment of the present invention uses a preset small number of images actually collected from the target sparse scene for annotation, so that the training data is closely centered around the target scene. This ensures that the diffusion model focuses on learning the characteristics of the target foreground and background in the specific scene, avoiding the model from being disturbed by irrelevant data, and thus accurately adapting to the needs of the target sparse scene. The preset loss function provides a clear optimization goal for model training. It can measure the difference between the noise predicted by the model and the actual noise added, and guide the model to adjust parameters through the back-propagation algorithm to minimize this difference. It can ensure that the model learns in the right direction during the fine-tuning process and effectively updates the model parameters, thereby improving the performance and generation quality of the model.

[0035] In a second aspect, the present invention provides a data set generation system for sparse scenes, comprising:

[0036] The detection frame output module is used to input the real collected images of the target sparse scene into the target detection model to obtain the detection frames of the target foreground and background;

[0037] An image segmentation module is used to use a preset segmentation model to use the detection frames of the target foreground and background as prompt words to perform model reasoning, obtain masks of the target foreground and background respectively, and obtain target foreground image and background area coordinate information based on the masks of the target foreground and background;

[0038] The adaptive fusion module is used to integrate the target foreground image and background area coordinate information using adaptive image fusion technology to generate a new target sparse scene image, and add the newly generated target sparse scene image to the training data set of the target detection model.

[0039] In a third aspect, the present invention provides a computer device, comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the data set generation method for sparse scenes of the first aspect or any corresponding embodiment thereof by executing the computer instructions.

[0040] In a fourth aspect, the present invention provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the method for generating a data set for a sparse scene according to the first aspect or any corresponding embodiment thereof.

[0041] In a fifth aspect, the present invention provides a computer program product, comprising computer instructions for causing a computer to execute the method for generating a data set for a sparse scene according to the first aspect or any corresponding embodiment thereof. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0043] Figure 1 This is a schematic diagram of a scenario where an indoor fire escape passage is blocked;

[0044] Figure 2 A schematic diagram of a flow chart of a method for generating a data set for a sparse scene according to an embodiment of the present invention;

[0045] Figure 3 is a flowchart of a method for generating a data set for another sparse scenario according to an embodiment of the present invention;

[0046] Figure 4 A structural block diagram of a data set generation system for sparse scenes according to an embodiment of the present invention;

[0047] Figure 5 It is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0048] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0049] With the acceleration of urbanization, there are more and more high-rise buildings, and fire safety issues are becoming increasingly prominent. Indoor fire escape routes are important passages for evacuation and firefighting and rescue in case of fire, and their smoothness is directly related to people's life safety. However, due to various reasons, such as the accumulation of debris, damage to firefighting facilities, etc., Figure 1 As shown in the figure, the blockage of indoor fire passages occurs from time to time, which poses a serious hidden danger to fire safety. At present, the means of addressing the problem of indoor fire passage blockage mainly include regular inspections, publicity and education, and video surveillance. However, these methods still have certain limitations in practical applications. For example, inspectors may not be able to discover all blockages in time; the effectiveness of publicity and education is affected by many factors, and it is difficult to ensure the actual effect. In addition, due to the difficulty in deploying sensors, high data collection costs, and unstable data quality, it becomes difficult to obtain data in some key areas. At the same time, traditional data collection methods are often costly and are subject to user data privacy issues, making it difficult to achieve comprehensive data collection. Therefore, how to obtain data on indoor fire passages at low cost and high efficiency has become an urgent problem to be solved in the field of urban governance.

[0050] Image fusion technology can integrate image data from different sources to generate an image with richer information and higher definition. This method can not only improve the quality of data, but also reduce the cost of data acquisition. When processing images of different sources and different qualities, existing image fusion technology usually adopts a fixed parameter method for fusion, which may cause the generated data set to be inaccurate or comprehensive in some cases. For example, when the quality difference of the source images is large or the feature difference is obvious, the fixed parameter fusion method may not fully take these differences into account, resulting in unsatisfactory fusion results. In order to ensure the quality of the generated picture, the adaptive image fusion technology proposed in the embodiment of the present invention can automatically adjust the fusion parameters according to the characteristics of images of different sources and different qualities, so as to generate a more comprehensive and accurate data set. This method can not only improve the data quality, but also solve the problem of data acquisition in key areas, and provide strong support for urban governance.

[0051] According to an embodiment of the present invention, an embodiment of a method for generating a data set for a sparse scene is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in an order different from that shown here.

[0052] In this embodiment, a method for generating a data set for a sparse scene is provided. Figure 2 is a flow chart of a method for generating a data set for a sparse scene according to an embodiment of the present invention. Figure 2 As shown, the process includes the following steps:

[0053] S101, input the real captured image of the target sparse scene into the target detection model to obtain the detection frame of the target foreground and background.

[0054] In the embodiment of the present invention, the scene of indoor fire escape blocked is taken as an example. The image of indoor fire escape blocked is input into the target detection model to obtain the target obstruction (target foreground) and indoor stairs (background). The above target detection model is trained through the following process, including:

[0055] A1, target foreground and background are annotated for a preset small number of images collected from target sparse scenes.

[0056] Specifically, for example, in an existing real-world collected data set, 100 pictures are selected for annotation, and the annotation categories include indoor stairs and target obstructions, which is only an example and is not limited to this.

[0057] A2, normalize the labeled data and combine it with data augmentation technology to generate a training set.

[0058] Specifically, for example, the mosaic method is used to increase the diversity of training data and expand the data set. The mosaic method generates new images by splicing together parts of images of different categories, which allows the model to learn the relationship between the features of objects of different categories, thereby improving the accuracy of classification. Since the new image is composed of parts of different images, the model can be exposed to more different combinations of object features, allowing the model to see combinations of objects that are unlikely to appear at the same time in natural scenes, thereby improving the model's ability to recognize complex scenes.

[0059] A3, based on the training set, uses the preset machine learning model for training, and uses the trained model as the target detection model. The target detection model outputs a detection box for detecting the target foreground and background.

[0060] The embodiment of the present invention uses the YOLO model as the initial target detection model, and obtains the target detection model after iterative training. YOLO has multiple versions, for example, YOLOv5 can be selected, which has the characteristics of small model, fast training speed, high accuracy, etc. Before performing target detection, the input labeled image is preprocessed, usually including image scaling, normalization and other operations. The preprocessed image is input into the model for detection, and the model will output the category, detection box coordinates, confidence and other information of the detected target object. Reliable target detection results can be screened out according to the confidence threshold, for example, only targets with a confidence greater than 0.5 are retained.

[0061] The embodiment of the present invention only needs to process a small number of images of the target sparse scene. Compared with large-scale data collection, it greatly reduces the manpower, material and time costs required in the data collection stage. By marking the target foreground and background, it provides a clear learning goal for the model. The model can accurately understand which areas are the target foreground and which are the background based on the marked information, which can effectively improve the accuracy and efficiency of target detection and the generalization ability of the model.

[0062] S102, using a preset large segmentation model to use the detection boxes of the target foreground and background as prompt words to perform model reasoning, obtain masks of the target foreground and background respectively, and obtain the target foreground image and background area coordinate information based on the masks of the target foreground and background.

[0063] Specifically, there are currently a variety of excellent image segmentation models to choose from, such as Mask R-CNN, U-Net and its variants, FastSAM model, DeepLab series, etc. The embodiment of the present invention uses the FastSAM model as the segmentation model, provides the prepared prompt words containing the target foreground and background detection box information as input to the segmentation model, and extracts the target foreground and background masks from the output results of the model. To obtain the target foreground image, the mask information can be used to process the original image, and the original image and the foreground mask are multiplied element by element so that only the pixels in the foreground area are retained, and then the result is converted into a suitable image data type. The coordinate information of the background area can be obtained based on the background mask by traversing the pixels in the mask and finding the pixel coordinates whose mask value is a specific value (for example, 1, indicating the background area).

[0064] The embodiment of the present invention utilizes the target detection model to provide the approximate position of the target by outputting the detection box. The image segmentation can be further refined to the pixel level through the segmentation model, and a mask is generated to accurately divide the target foreground and background. This combination enables a deeper understanding of the image and can be better applied to subsequent processing tasks.

[0065] S103, using adaptive image fusion technology to integrate the target foreground image and background area coordinate information to generate a new target sparse scene image, and adding the newly generated target sparse scene image to the training data set of the target detection model.

[0066] Specifically, adaptive image fusion technology aims to dynamically adjust the fusion strategy according to the characteristics and requirements of different image regions to achieve a more natural and target-oriented image synthesis effect. When integrating the target foreground image and the background area coordinate information, it can reasonably determine the fusion method based on the characteristics of the background area and the characteristics of the foreground image, perform image fusion to generate a new target sparse scene image, and add the newly generated target sparse scene image according to the format and storage method of the training data set. If it is stored in the image folder and annotation file, first ensure that the file name of the new image conforms to the naming convention of the dataset, and then copy or move it to the corresponding training image folder. At the same time, it is necessary to generate a corresponding annotation file for the new image according to the specific situation, and the annotation content can be determined according to the role and characteristics of the new image in the entire training task.

[0067] The method for generating a data set for a sparse scene provided by an embodiment of the present invention can accurately identify and locate the target in the image through target detection technology, provide effective prompt words for the segmentation of the large model, and at the same time, combined with the large model segmentation technology, can accurately segment the target and the background, and can adaptively and flexibly adjust the fusion method according to images of different sources and different qualities, thereby improving the authenticity of the fused image. Only a small number of target and background images are needed to generate a large number of new high-quality images, which solves the problem of scarce scene data and provides rich data resources for subsequent model training.

[0068] In this embodiment, a method for generating a data set for a sparse scene is provided. Figure 3 is a flow chart of a method for generating a data set for a sparse scene according to an embodiment of the present invention. Figure 3 As shown, the process includes the following steps:

[0069] S201, input the real captured image of the target sparse scene into the target detection model to obtain the detection frame of the target foreground and background; see the relevant description of step S101 for details, which will not be repeated here.

[0070] S202, using a preset large segmentation model to use the detection frames of the target foreground and background as prompt words to perform model reasoning, respectively obtaining masks of the target foreground and background, and obtaining target foreground image and background area coordinate information based on the masks of the target foreground and background; specifically comprising the following steps:

[0071] S2021, using the preset large segmentation model to use the detection boxes of the target foreground and background as prompt words to perform model inference, and obtain the mask of the target foreground single-channel image and the mask of the background single-channel image respectively.

[0072] Specifically, the detection frames of the target foreground and background are obtained through the previous target detection steps (such as using the YOLO model, etc.). These detection frame information is usually given in the form of coordinates. For example, for a detection frame, there may be upper left corner coordinates (x1, y1) and lower right corner coordinates (x2, y2), as well as corresponding target category information. The detection frame coordinate information and related categories are organized into a suitable tensor or data structure form and input into the FastSAM model to obtain the mask m of the target obstruction image. obj and the mask m for the interior stairs bkg . It should be noted that a mask is a binary image (usually pixel values ​​are only 0 and 1) or a grayscale image (pixel values ​​are within a certain range, such as 0 to 255), and each pixel corresponds to a corresponding position in the original image. In a binary mask, an area with a pixel value of 1 usually represents an area of ​​interest, that is, an area that requires specific processing (such as extraction, analysis, modification, etc.); an area with a pixel value of 0 represents an area that does not require such processing. The mask obtained in the embodiment of the present invention is a binary image.

[0073] S2022, expand the mask of the single-channel image of the target foreground into mask data with the same number of channels as the original real captured image, and obtain a segmented target foreground image based on the original real captured image and the mask data with the same number of channels as the original real captured image, and calculate the width value and height value of the target object in the target foreground image.

[0074] Specifically, the single-channel mask data m of the target obstruction is obj Make a copy and splice it to the original mask m obj The mask data m with the same number of channels as the original real collected image is obtained. ′ obj , the pixel value of the target blockage at each pixel point is obtained by the following formula:

[0075] Target blockage = m ′ obj ×Original real collected image

[0076] Then the target obstruction is segmented from the original image to obtain the segmented target foreground image, and the width w of the target object is calculated. o and high h o .

[0077] The embodiment of the present invention expands the single-channel image mask of the target foreground into mask data with the same number of channels as the original real captured image, and then obtains the segmented target foreground image based on the original image and the mask data, so that the target foreground can be accurately extracted from the original image, and the original channel information of the image is maintained. For example, for an RGB color image, the expanded mask can ensure that the RGB pixel values ​​of the target foreground part are correctly extracted, thereby obtaining a complete target foreground image, which is convenient for subsequent independent analysis and processing of the target foreground.

[0078] S2023, obtaining the minimum bounding rectangle corresponding to the background single-channel image, and obtaining the center coordinates, width value, and height value of the minimum bounding rectangle.

[0079] Specifically, the minimum bounding rectangle information helps analyze the spatial relationship between the background and the target foreground. For example, by comparing the position and size relationship of the minimum bounding rectangle of the target foreground and the minimum bounding rectangle of the background, the relative position (center or edge) and proportion (size ratio of the target to the background) of the target in the background can be evaluated.

[0080] The embodiment of the present invention obtains the background single channel to obtain the indoor staircase mask m bkg The coordinate set (X, Y) with a median value of 1, the area enclosed by this coordinate set is the area of ​​the stairs in the original image, and then the minimum circumscribed rectangle of the coordinate set is calculated to obtain the center coordinates of the rectangle (x c ,y c ) and width w b and high h b If the width or height of the target obstruction is larger than the width or height of the indoor stairs, the fused obstruction will exceed the original image, and the size of the target obstruction needs to be scaled. After scaling, the width w is updated. o and high h o , the scaling factor is calculated as follows:

[0081] scale=max(w b / w o ,w o / w b )

[0082] The embodiment of the present invention scales the target based on the scaling factor, which can ensure that the target can be reasonably placed in the background area. Scaling the target can make the ratio between the target and the background more coordinated. In a data set containing multiple targets and background combinations, if the size of the target is too large, it may cause data inconsistency. By scaling the target, all data can be made more uniform in size, which is convenient for unified processing and analysis.

[0083] S203, using adaptive image fusion technology, integrating the target foreground image and the background area coordinate information to generate a new target sparse scene image, and adding the newly generated target sparse scene image to the training data set. Specifically, the following steps are included:

[0084] S2031, performing data enhancement processing on the target foreground image to introduce new features, thereby obtaining a target foreground image set; specifically, for example, performing data enhancement processing operations such as horizontal flipping, adjusting brightness, and adding artificial noise on the target image to introduce new features into the image, thereby enriching the data of the target blockage.

[0085] S2032, determining a fusion position range in a random manner based on the center coordinates of the minimum circumscribed rectangle;

[0086] In the embodiment of the present invention, the coordinates of the fusion position (the center coordinates of the minimum bounding rectangle) (x, y) are randomly selected in the following range:

[0087] x b +w o / 2≤x≤x b +w b -w o / 2

[0088] y b +h o / 2≤y≤y b +h b -h o / 2

[0089] Among them, x b and b They respectively represent the horizontal and vertical coordinates of the background area. The above random selection method is only an example and is not limited to this.

[0090] S2033, based on the fusion position range, the target foreground image set is superimposed and fused with the actually collected background image to obtain a fused image set, and the divergence value of each image in the fused image set in the fusion area is calculated.

[0091] Specifically, before performing superposition fusion, it may be necessary to perform some preprocessing operations on the target foreground image set and the background image to ensure that they are more matched in terms of size, color space, etc., so as to facilitate better fusion. According to the determined fusion position range, each image of the target foreground image set is superimposed and fused with the background image. Taking the rectangular fusion area as an example, assuming that the target foreground image is I f , the background image is I b , the coordinates of the upper left corner of the fusion area are (x1, y1), and the coordinates of the lower right corner are (x2, y2). The following steps can be used to achieve superposition fusion:

[0092] First, create an all-zero image of the same size as the background image as a temporary image I for fusion temp ; Then the target foreground image I f The pixel values ​​of the corresponding fusion area in I are copied to temp Finally, it will be placed at the same position as the background image I b Perform element-by-element addition (or perform other operations according to a specific fusion strategy, such as weighted addition, etc.) to obtain a fused image f.

[0093] Furthermore, the divergence value of each fused image f in the fused area S in the fused image set is calculated by the following formula:

[0094] △f(x,y)=f(x-1,y)+f(x+1,y)+f(x,y-1)+f(x,y+1)-4f(x,y)

[0095] The divergence value is a measure of the degree of divergence of the vector field. In the context of fused images, it reflects the change in pixel values ​​in the fused area. The size and distribution of the divergence value play an important role in indicating the image fusion effect. When the divergence value of the fused image in the fused area is low, it usually means that the change in pixel values ​​is relatively flat, indicating that the transition between the target foreground and background in the fused area is relatively natural.

[0096] S2034: Use images in the fused image set whose divergence values ​​are less than or equal to the first threshold as a first fused image set.

[0097] In one embodiment, for example, the first threshold is 1. If the divergence value of the fused image is less than or equal to 1, it indicates that the fusion effect is relatively good and no further processing is required.

[0098] S2035, using a preset fusion method to process the images in the fused image set whose divergence values ​​are greater than the first threshold and less than the second threshold, as the second fused image set.

[0099] In one embodiment, for example, if the second threshold is 10, it means that the quality of the fused image in step S2033 is poor and needs to be further processed to improve the image quality. In the embodiment of the present invention, Poisson fusion is used to make the transition between the fusion area and the surrounding environment very natural in terms of color and texture, and there will be no obvious splicing, color difference and other problems. Poisson fusion is based on the Poisson equation to achieve image fusion. The core idea is to make the fused image match the source image and the target image as much as possible in terms of gradient field under the premise of ensuring that the boundary conditions of the fusion area remain unchanged, so as to achieve a natural fusion effect.

[0100] S2036, input the images with divergence values ​​greater than the second threshold in the fused image set into the fine-tuned diffusion model to obtain a third fused data set.

[0101] In the embodiment of the present invention, the images with divergence values ​​greater than 10 in the fused image set are input into the fine-tuned diffusion model to obtain a third fused data set. The process of fine-tuning the diffusion model includes:

[0102] B1, annotate a preset small number of images (for example, 100 real annotated images) of the target sparse scene for the target foreground and background; gradually add Gaussian noise to the annotated images until the final pure noise image is obtained.

[0103] This step is the forward process of the diffusion model: the forward diffusion process is defined as q(x t |x t-1 ), that is, gradually add Gaussian noise to the image until pure noise is finally obtained. Define a series of coefficients 0<β1<…<1:

[0104]

[0105] in, is a normal distribution.

[0106] B2, input the noisy image into the preset noise prediction network, train the model using the preset loss function, and obtain the fine-tuned diffusion model.

[0107] Specifically, this process is the reverse process of the diffusion model. In the embodiment of the present invention, the UNet model is used as the noise prediction network. The UNet network needs to predict noise data. The noise predicted by the network is subtracted from the image with noise applied to generate a real image. The noise output by the UNet network is ∈ θ (x t ,t), the real noise is taken from the normal distribution The loss function is defined as:

[0108]

[0109] The training mechanism of the diffusion model is usually based on the learning and restoration of noise. By adding Gaussian noise until a pure noise image is obtained, a complete noise learning path is provided for the model. This helps the model understand the process of an image from its original state to being completely corrupted by noise, so as to better learn how to reversely generate high-quality images. It should be noted that the diffusion model is a relatively mature model, and the embodiment of the present invention fine-tunes it for high-quality image generation in sparse scenes.

[0110] The embodiment of the present invention uses a preset small number of images actually collected from the target sparse scene for annotation, so that the training data is closely centered around the target scene. This ensures that the diffusion model focuses on learning the characteristics of the target foreground and background in this specific scene, avoiding the model from being disturbed by irrelevant data, and thus accurately adapting to the needs of the target sparse scene. The preset loss function provides a clear optimization goal for model training, which can measure the difference between the noise predicted by the model and the actual noise added, and guide the model to adjust parameters through the back propagation algorithm to minimize this difference. This ensures that the model learns in the right direction during the fine-tuning process, effectively updates the model parameters, and thus improves the performance and generation quality of the model.

[0111] S2037 , adding the first fused image set, the second fused image set, and the third fused image set as newly generated target sparse scene images to the training data set.

[0112] Specifically, it needs to be determined according to the format and storage method of the training dataset. For example, if the training dataset is stored in image folders and annotation files, the new image needs to be copied or moved to the corresponding folder, and a suitable annotation file needs to be generated for it. If a specific data format is used (such as TensorFlow's TFRecord format, PyTorch's Dataset format, etc.), the new image and its related information need to be encoded and added to the existing data file in accordance with the corresponding specifications.

[0113] Through the above steps, the entire process from image fusion, divergence value calculation to adaptive image fusion technology processing of images according to different divergence value ranges, and finally adding the processed image set to the training data set is completed, which helps to further optimize the image fusion effect and enrich the content of the training data set.

[0114] In this embodiment, a data set generation system for sparse scenes is also provided, which is used to implement the above-mentioned embodiments and preferred implementation modes, and will not be repeated hereafter. As used below, the term "module" can implement a combination of software and / or hardware for a predetermined function. Although the systems described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.

[0115] This embodiment provides a data set generation system for sparse scenes, such as Figure 4 As shown, including:

[0116] The detection frame output module 401 is used to input the real captured image of the target sparse scene into the target detection model to obtain the detection frame of the target foreground and background;

[0117] The image segmentation module 402 is used to use the preset segmentation model to perform model reasoning using the detection boxes of the target foreground and background as prompt words, obtain masks of the target foreground and background respectively, and obtain the target foreground image and background area coordinate information based on the masks of the target foreground and background;

[0118] The adaptive fusion module 403 is used to integrate the target foreground image and background area coordinate information using adaptive image fusion technology to generate a new target sparse scene image, and add the newly generated target sparse scene image to the training data set of the target detection model.

[0119] In some optional implementations, the object detection model in the detection frame output module 401 is trained by the following process, including:

[0120] Annotate the target foreground and background of a preset small number of images collected from scenes with sparse targets;

[0121] Normalize the labeled data and generate a training set using data augmentation techniques;

[0122] Based on the training set, training is performed using a preset machine learning model, and the trained model is used as a target detection model. The target detection model outputs a detection frame for detecting the target foreground and background.

[0123] In some optional implementations, the image segmentation module 402 includes:

[0124] A segmentation model inference unit is used to use the preset segmentation model to perform model inference using the detection frames of the target foreground and background as prompt words to obtain a mask of the target foreground single-channel image and a mask of the background single-channel image respectively;

[0125] A target object size acquisition unit is used to expand the mask of the single-channel image of the target foreground into mask data with the same number of channels as the original real collected image, obtain a segmented target foreground image based on the original real collected image and the mask data with the same number of channels as the original real collected image, and calculate the width value and height value of the target object in the target foreground image;

[0126] The background area coordinate information acquisition unit is used to acquire the minimum bounding rectangle corresponding to the background single-channel image, and acquire the center coordinates, width value and height value of the minimum bounding rectangle.

[0127] In some optional implementations, when the width value or height value of the target object is greater than the width value or height value of the background, the size of the target object is scaled based on the scaling factor, and the width value and height value of the target object are updated after scaling.

[0128] In some optional implementations, the adaptive fusion module 403 includes:

[0129] a target foreground image set acquisition unit, configured to perform data enhancement processing on the target foreground image to introduce new features, thereby obtaining a target foreground image set;

[0130] A fusion position range acquisition unit, used for determining the fusion position range in a random manner based on the center coordinates of the minimum circumscribed rectangle;

[0131] A divergence value calculation unit is used to superimpose and fuse the target foreground image set with the real collected background image based on the fusion position range to obtain a fused image set, and calculate the divergence value of each image in the fused image set in the fusion area;

[0132] A first fused image set acquisition unit, configured to take images in the fused image set whose divergence values ​​are less than or equal to a first threshold as a first fused image set;

[0133] A second fused image set acquisition unit, which processes the images in the fused image set whose divergence values ​​are greater than the first threshold and less than the second threshold by using a preset fusion method, as the second fused image set;

[0134] A third fused image set acquisition unit, inputting images in the fused image set whose divergence values ​​are greater than a second threshold into the fine-tuned diffusion model to obtain a third fused data set;

[0135] The training data set expansion unit is used to add the first fused image set, the second fused image set and the third fused image set as newly generated target sparse scene images into the training data set.

[0136] In some optional implementations, the process of fine-tuning the diffusion model includes:

[0137] Annotate the target foreground and background of a preset small number of images collected from scenes with sparse targets;

[0138] Gradually add Gaussian noise to the labeled image until the final pure noise image is obtained;

[0139] The noisy image is input into the preset noise prediction network, and the model is trained using the preset loss function to obtain a fine-tuned diffusion model.

[0140] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.

[0141] The data set generation system for sparse scenarios in this embodiment is presented in the form of functional units, where the units refer to ASIC (Application Specific Integrated Circuit) circuits, processors and memories that execute one or more software or fixed programs, and / or other devices that can provide the above functions.

[0142] The embodiment of the present invention also provides a computer device having the above Figure 4 The dataset generation system for the sparse scene shown.

[0143] See also Figure 5 , Figure 5 is a schematic diagram of the structure of a computer device provided by an optional embodiment of the present invention, such as Figure 5 As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components are connected to each other using different buses for communication, and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the computer device, including instructions stored in or on the memory to display the graphical information of the GUI on an external input / output device (such as, a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 5 A processor 10 is taken as an example.

[0144] The processor 10 may be a central processing unit, a network processor or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.

[0145] The memory 20 stores instructions executable by at least one processor 10, so that at least one processor 10 executes the method shown in the above embodiment.

[0146] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely arranged relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0147] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 20 may also include a combination of the above types of memory.

[0148] The computer device further comprises a communication interface 30 for the computer device to communicate with other devices or a communication network.

[0149] The embodiment of the present invention also provides a computer-readable storage medium. The method according to the embodiment of the present invention can be implemented in hardware, firmware, or can be implemented as a computer code that can be recorded in a storage medium, or can be implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and will be stored in a local storage medium through a network download, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state hard disk, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor central control system or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor or hardware, the method shown in the above embodiment is implemented.

[0150] A part of the present invention may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. Those skilled in the art should understand that the existence of the computer program instruction in a computer-readable medium includes, but is not limited to, a source file, an executable file, an installation package file, etc., and accordingly, the way in which the computer program instruction is executed by the computer includes, but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium may be any available computer-readable storage medium or communication medium accessible to the computer.

[0151] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations are all within the scope defined by the appended claims.

Claims

1. A method for generating a dataset of sparse scenes, characterized in that: include: Input the real collected images of the target sparse scene into the target detection model to obtain the detection boxes of the target foreground and background; Using a preset segmentation model, the detection frames of the target foreground and background are used as prompt words to perform model reasoning, and masks of the target foreground and background are obtained respectively, and coordinate information of the target foreground image and background area is obtained based on the masks of the target foreground and background; Adaptive image fusion technology is used to integrate the target foreground image and background area coordinate information to generate a new target sparse scene image, and the newly generated target sparse scene image is added to the training data set of the target detection model.

2. The method according to claim 1, characterized in that The target detection model is trained through the following process, including: Annotate the target foreground and background of a preset small number of images collected from scenes with sparse targets; Normalize the labeled data and generate a training set using data augmentation techniques; Based on the training set, training is performed using a preset machine learning model, and the trained model is used as a target detection model. The target detection model outputs a detection frame for detecting the target foreground and background.

3. The method according to claim 1, characterized in that The method uses the preset segmentation model to use the detection frames of the target foreground and background as prompt words to perform model reasoning, obtain the masks of the target foreground and background respectively, and obtain the target foreground image and background area coordinate information based on the masks of the target foreground and background, including: Using a preset large segmentation model, the detection frames of the target foreground and background are used as prompt words to perform model reasoning, and a mask of a target foreground single-channel image and a mask of a background single-channel image are obtained respectively; The mask of the single-channel image of the target foreground is expanded to mask data with the same number of channels as the original real collected image, and the segmented target foreground image is obtained based on the original real collected image and the mask data with the same number of channels as the original real collected image, and the width and height values ​​of the target object in the target foreground image are calculated; Get the minimum bounding rectangle corresponding to the background single-channel image, and get the center coordinates, width, and height of the minimum bounding rectangle.

4. The method according to claim 3, characterized in that When the width value or the height value of the target object is greater than the width value or the height value of the background, the size of the target object is scaled based on the scaling factor, and the width value and the height value of the target object are updated after scaling.

5. The method according to claim 3, characterized in that: The method of using the adaptive image fusion technology to integrate the target foreground image and the background area coordinate information to generate a new target sparse scene image, and adding the newly generated target sparse scene image to the training data set of the target detection model includes: Performing data enhancement processing on the target foreground image to introduce new features, obtaining a target foreground image set, and determining a fusion position range in a random manner based on the center coordinates of a minimum circumscribed rectangle; Based on the fusion position range, the target foreground image set is superimposed and fused with the real collected background image to obtain a fused image set, and the divergence value of each image in the fusion area of ​​the fused image set is calculated; The images whose divergence values ​​in the fused image set are less than or equal to the first threshold are taken as the first fused image set; The images whose divergence values ​​in the fused image set are greater than the first threshold and less than the second threshold are processed by a preset fusion method as the second fused image set; Input the images in the fused image set whose divergence values ​​are greater than the second threshold into the fine-tuned diffusion model to obtain a third fused data set; The first fused image set, the second fused image set and the second fused image set are added as newly generated target sparse scene images to the training data set of the target detection model.

6. The method according to claim 5, characterized in that The process of fine-tuning the diffusion model involves: Annotate the target foreground and background of a preset small number of images collected from scenes with sparse targets; Gradually add Gaussian noise to the labeled image until the final pure noise image is obtained; The noisy image is input into the preset noise prediction network, and the model is trained using the preset loss function to obtain a fine-tuned diffusion model.

7. A data set generation system for sparse scenes, characterized in that: include: The detection frame output module is used to input the real collected images of the target sparse scene into the target detection model to obtain the detection frames of the target foreground and background; An image segmentation module is used to use a preset segmentation model to use the detection frames of the target foreground and background as prompt words to perform model reasoning, obtain masks of the target foreground and background respectively, and obtain target foreground image and background area coordinate information based on the masks of the target foreground and background; The adaptive fusion module is used to integrate the target foreground image and background area coordinate information using adaptive image fusion technology to generate a new target sparse scene image, and add the newly generated target sparse scene image to the training data set of the target detection model.

8. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the method for generating a data set for a sparse scene according to any one of claims 1 to 6 by executing the computer instructions.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the method for generating a data set for a sparse scene according to any one of claims 1 to 6.

10. A computer program product, characterized in that The method comprises computer instructions, wherein the computer instructions are used to cause a computer to execute the method for generating a data set for a sparse scene according to any one of claims 1 to 6.

Citation Information

Cited By

  • Express stacking detection method and device based on target detection, equipment and medium

    CN121170472A

  • Food material data generation method and device based on image processing, storage medium and electronic device

    CN121482526A

  • AI data set generation method and system based on image generation

    CN121686137A

  • Deep-sea rare biological target detection method based on stable diffusion model

    CN122116416A

  • Data set generation method and system for sparse scene

    WO2026108865A1