Scenic area garbage identification method and device based on causal characteristics, medium and product
By using asymmetric twin network and image fusion technology in the scenic spot garbage recognition system, the causal characteristics of scenic spot garbage are extracted, and the problem of insufficient garbage recognition accuracy in the scenic spot background environment is solved, and higher recognition accuracy and system reliability are achieved.
Patent Information
- Application Number
- CN202510088474.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-21
AI Technical Summary
The accuracy of garbage recognition in the scenic area background environment is insufficient in the prior art, especially due to the uneven distribution of training samples, the model misses the garbage recognition in non-featured scenarios.
The garbage recognition method of scenic spots based on causal characteristics is adopted. By obtaining the scenic spot monitoring images, adjusting the images to standard sizes, extracting garbage and background samples, image fusion generates training data, using an asymmetric twin network for model training, extracting the features of the object itself and ignoring the background information.
The model's accuracy of identifying the same garbage in different backgrounds is improved, the false positive rate is reduced, and the practical value and reliability of the model are enhanced.
Smart Images

Figure CN119942215A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data identification, and in particular to a method, device, medium and product for identifying garbage in a scenic area based on causal characteristics. Background Art
[0002] With the continuous advancement of urbanization and the continuous improvement of people's living standards, the construction and management of urban park scenic spots have received more and more attention. As an important place for urban residents to relax, entertain, and exercise, the environmental sanitation quality of park scenic spots is directly related to the citizens' sightseeing experience and the city's image.
[0003] Among related technologies, target detection algorithms in the field of computer vision are widely used in garbage identification tasks. Such algorithms use deep learning models to extract and classify features from images to achieve automatic identification of garbage. Common technical solutions use target detection algorithms such as YOLO and Faster R-CNN, combined with images collected by surveillance cameras, to detect garbage in real time, and then assign robots to clean up after detection.
[0004] However, there are deviations in the distribution of different types of garbage in the background environment of scenic spots. For example, leaves and flowers mainly appear on the road and water backgrounds, while paper scraps and fruit peels are mostly seen in the grass background. This will lead to an uneven distribution of the model's training samples. Due to the lack of relevant training samples, the model will miss some non-feature scenes of garbage, such as leaves and petals on the grass, and paper scraps and fruit peels in the road background. The garbage recognition accuracy of related technologies needs to be improved. Summary of the invention
[0005] The present application provides a method, device, medium and product for accurately identifying scenic area garbage based on causal characteristics.
[0006] In a first aspect, the present application provides a method for identifying garbage in a scenic area based on causal features, which is applied to a garbage identification device for a scenic area. The method includes: obtaining multiple scenic area monitoring images of a target park scenic area, and adjusting the multiple scenic area monitoring images to a preset standard size to obtain multiple standardized images; respectively extracting garbage areas and background areas from the multiple standardized images to obtain a garbage sample set and a background sample set; performing image fusion on each garbage area in the garbage sample set with multiple background areas in the background sample set to generate a garbage image training set; the garbage image training set includes multiple garbage image groups, each garbage image group includes multiple images of garbage in multiple backgrounds; constructing a twin network including two encoding branches, setting a gradient backpropagation blocking operation on the second encoding branch of the twin network to obtain an asymmetric twin network; selecting a garbage image pair with the same garbage but different backgrounds from the garbage image group of the garbage image training set, and respectively inputting the two images of the garbage image pair into the two encoding branches of the asymmetric twin network, performing model training with the cosine distance of the output vectors of the two encoding branches as the training standard, and determining a feature extraction model based on the first encoding branch that has completed the training; inputting the monitoring image to be tested into the feature extraction model to obtain a garbage identification result.
[0007] In the above embodiment, the scenic area garbage identification device obtains the scenic area monitoring image and extracts garbage and background samples, performs image fusion to generate training data, and adopts an asymmetric twin network for model training, so that the network can extract the characteristics of the object itself (causal characteristics) and ignore the changing background information (correlation characteristics). By preventing the gradient backpropagation of the second encoding branch, the stability of feature extraction is guaranteed, and the recognition accuracy of the model for the same garbage under different backgrounds is improved.
[0008] In combination with some embodiments of the first aspect, in some embodiments, the step of obtaining multiple scenic area monitoring images of the target park scenic area, and adjusting the multiple scenic area monitoring images to a preset standard size to obtain multiple standardized images specifically includes: obtaining multiple scenic area monitoring images of the target park scenic area, and calculating the Laplace operator variance of the scenic area monitoring images to obtain image clarity; calculating the average value of the image pixel values in the scenic area monitoring images to obtain the image brightness value; when the image clarity is greater than a preset clarity threshold and the image brightness value is within a preset brightness range, taking the corresponding scenic area monitoring image as a valid image; selecting a preset number of target images from the valid images, adjusting the target images to a preset standard size, and obtaining multiple standardized images.
[0009] In the above embodiment, the scenic area garbage identification device screens the clarity and brightness of the input image by calculating the Laplace operator variance and pixel average of the image, thereby ensuring the quality of the training data, eliminating blurred and improperly exposed images, and retaining only valid images that meet the requirements, thereby improving the accuracy of subsequent feature extraction and recognition.
[0010] In combination with some embodiments of the first aspect, in some embodiments, the step of selecting a preset number of target images from valid images, adjusting the target images to a preset standard size, and obtaining a plurality of standardized images specifically includes: selecting a preset number of target images from valid images, and obtaining image size data of each target image; when the image size data is smaller than the preset standard size, enlarging the target image using a bicubic interpolation algorithm to obtain a target enlarged image; when the image size data is larger than the preset standard size, compressing the target image using a regional mean pooling method to obtain a target reduced image; and normalizing the target image equal to the preset standard size, the target enlarged image, and the target reduced image to obtain a plurality of standardized images.
[0011] In the above embodiment, the scenic area garbage identification device uses a bicubic interpolation algorithm to enlarge small-size images, uses regional mean pooling to compress large-size images, and finally performs normalization processing to make all images meet the preset standard size requirements, which not only ensures the maximum retention of image information, but also realizes data standardization.
[0012] In combination with some embodiments of the first aspect, in some embodiments, after the steps of respectively extracting garbage areas and background areas from multiple standardized images to obtain garbage sample sets and background sample sets, the method also includes: counting the number of samples of each type of garbage samples in the garbage sample set; when the number of samples of a certain garbage sample is less than a preset number of samples, treating the garbage sample as a garbage missing sample; and completing the number of samples of the garbage missing sample to the preset number of samples.
[0013] In the above embodiment, the scenic area garbage identification device counts the number of various types of garbage samples, identifies the categories with insufficient sample numbers, and performs data enhancement processing such as rotation, flipping and brightness adjustment on these missing samples, thereby effectively solving the problem of uneven distribution of training data and improving the model's ability to identify various types of garbage.
[0014] In combination with some embodiments of the first aspect, in some embodiments, each garbage area in the garbage sample set is image-fused with multiple background areas in the background sample set to generate a garbage image training set, which specifically includes: obtaining historical monitoring images of the target park scenic area in different seasons as seasonal characteristic images; extracting the same background area with different image features in each seasonal characteristic image to obtain a seasonal characteristic background set; and each garbage area in the garbage sample set is image-fused with multiple background areas in the seasonal characteristic background set to generate a garbage image training set.
[0015] In the above embodiment, the scenic area garbage identification device constructs a seasonal feature background set based on historical monitoring images of different seasons, extracts seasonal background features, realizes the fusion of garbage samples with different seasonal backgrounds, and enhances the adaptability of the model to garbage identification in different seasonal environments.
[0016] In combination with some embodiments of the first aspect, in some embodiments, each garbage area in the garbage sample set is respectively image-fused with multiple background areas in the seasonal characteristic background set to generate a garbage image training set, which specifically includes: calculating the main color tone, brightness distribution and texture characteristics of each background area in the seasonal characteristic background set to obtain seasonal characteristic parameters; performing color space conversion and parameter adjustment on the garbage areas in the garbage sample set according to the seasonal characteristic parameters to obtain seasonally adaptive garbage samples; extracting contour features of the seasonally adaptive garbage samples, and smoothing the contour edges to obtain edge-optimized garbage samples; performing gradient fusion on the edge-optimized garbage samples and the background areas in the seasonal characteristic background set to generate training images; and combining multiple training images with the same garbage but different seasons or different backgrounds into a garbage image group to generate a garbage image training set.
[0017] In the above embodiment, the scenic area garbage identification device extracts the main color tone, brightness distribution and texture features of the background area, performs color space conversion and parameter adjustment on the garbage samples, and performs contour optimization and gradient fusion to generate more natural training images, thereby improving the authenticity and diversity of the training data.
[0018] In combination with some embodiments of the first aspect, in some embodiments, after the step of inputting the monitoring image to be tested into the feature extraction model to obtain the garbage identification result, the method also includes: obtaining the user's annotation information on the garbage identification result, and counting the differences between the annotation information and the garbage identification result to obtain the model performance evaluation index; the model performance evaluation index includes the recognition accuracy, the missed detection rate and the false detection rate; according to the model performance evaluation index, determining the scene type whose performance is lower than the preset threshold, and obtaining the scene to be optimized; from the garbage image training set, screening the training samples with the same characteristics as the scene to be optimized, and constructing the scene-oriented training set; using the scene-oriented training set to perform incremental training on the feature extraction model to obtain the feature extraction optimization model; inputting the monitoring image to be tested into the feature extraction optimization model to obtain the garbage identification optimization result.
[0019] In the above embodiment, the scenic area garbage identification device evaluates the model performance based on the user annotation information, and constructs a targeted training set for scenes with poor performance for incremental training, which can continuously improve the recognition effect of the model in various actual scenes and ensure the practicality of the system.
[0020] In a second aspect, an embodiment of the present application provides a scenic area garbage identification device, which includes: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to enable the scenic area garbage identification device to perform the method described in the first aspect and any possible implementation method of the first aspect.
[0021] In a third aspect, an embodiment of the present application provides a computer program product comprising instructions. When the above-mentioned computer program product is run on a scenic area garbage identification device, the above-mentioned scenic area garbage identification device executes the method described in the first aspect and any possible implementation method of the first aspect.
[0022] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, comprising instructions. When the instructions are executed on a scenic area garbage identification device, the scenic area garbage identification device executes the method described in the first aspect and any possible implementation method of the first aspect.
[0023] It can be understood that the scenic area garbage identification device provided in the second aspect, the computer program product provided in the third aspect, and the computer storage medium provided in the fourth aspect are all used to execute the method provided in the embodiment of the present application. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method, and will not be repeated here.
[0024] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: 1. Image standardization and sample extraction are used to establish basic training data, and then image fusion technology is used to construct a group of garbage images with different backgrounds. An asymmetric twin network structure is adopted to maintain the stability of feature extraction by preventing the gradient feedback of one encoding branch, while allowing the other branch to continuously optimize to extract the essential characteristics of the garbage. This effectively solves the problem of dependence between model recognition and background information and insufficient generalization ability in related technologies, and realizes the extraction of the characteristics of the object itself (causal characteristics) while ignoring the changing background information (correlation characteristics), ensuring the accurate identification of the same garbage in different scenarios and improving the practical value and reliability of the model.
[0025] 2. Due to the use of statistical analysis of sample quantity, various data enhancement processes such as rotation, flipping and brightness adjustment are performed on the missing garbage samples until the preset sample quantity requirement is reached. This effectively solves the problems of insufficient training and recognition bias caused by the scarcity of samples of certain garbage categories in related technologies, achieves a balanced distribution of training data, and improves the model's recognition accuracy for various types of garbage.
[0026] 3. By collecting historical monitoring images of different seasons, extracting the characteristic background of each season, and realizing the fusion of garbage samples with diverse seasonal backgrounds, the problem of poor adaptability of models in related technologies to seasonal environmental changes is effectively solved, thereby achieving stable recognition effects under different seasonal backgrounds, enabling the system to adapt to changes in various environmental conditions in the scenic area throughout the year, and improving the practicality and reliability of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 It is a flow chart of a method for identifying scenic area garbage based on causal characteristics in an embodiment of the present application; Figure 2 (a) is a schematic diagram of the architecture of model training based on an asymmetric twin network in an embodiment of the present application; Figure 2 (b) is a schematic diagram of a scenario in which a feature extraction model is used to identify garbage in an embodiment of the present application; Figure 3 is another flow chart of the method for identifying scenic area garbage based on causal features in an embodiment of the present application; Figure 4 It is a schematic diagram of the structure of a physical device of a scenic area garbage identification device in an embodiment of the present application. DETAILED DESCRIPTION
[0028] The terms used in the following embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to be used as limitations to the present application. As used in the specification of the present application, the singular expressions "one", "a kind of", "above", "the" and "this" are intended to also include plural expressions, unless there is a clear indication to the contrary in the context. It should also be understood that the term "and / or" used in the present application refers to any or all possible combinations comprising one or more of the listed items.
[0029] In the following, the terms "first" and "second" are used for descriptive purposes only and are not to be understood as suggesting or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features, and in the description of the embodiments of the present application, unless otherwise specified, "plurality" means two or more.
[0030] For ease of understanding, the application scenarios of the embodiments of the present application are introduced below.
[0031] The park area of a large scenic spot receives more than 8,000 visitors every day. More than 50 high-definition surveillance cameras are deployed in the scenic park to monitor each area in real time. However, it faces severe challenges in garbage disposal in daily operations. For example, snack packaging bags and beverage bottles in the picnic area are often discarded at will; cigarette butts on the fitness trail are hidden in the grass and difficult to find; paper scraps in the rest area are mixed with the ground texture and are difficult to identify; dead branches and fallen leaves in the green belt blend into the environment.
[0032] In the related art, garbage identification can be performed by using a YOLO-based target detection algorithm to achieve automatic detection of garbage in scenic spots. The following describes a scenario in which a method for identifying garbage in scenic spots based on causal features in the related art is used.
[0033] A scenic spot has introduced an intelligent garbage identification system based on the YOLO algorithm to detect garbage under a 512×512 resolution monitoring screen. However, the system has obvious shortcomings in practical applications. For example, when packaging bags or beverage bottles appear on the ground with complex textures, the system's recognition accuracy is only 65%; in areas blocked by trees, due to uneven lighting, the detection rate of small garbage such as cigarette butts is less than 50%; garbage in green belts is often confused with vegetation features, and the missed detection rate is as high as 40%. In addition, due to the lack of recognition of the essential characteristics of garbage, the system often misjudges ground textures as paper scraps and vegetation as garbage, resulting in missed detections and false detections.
[0034] The feature extraction method based on the asymmetric twin network in the embodiment of the present application is used to learn the essential features of the garbage, and accurately identify the garbage in a complex background, which not only improves the recognition accuracy, but also reduces the false alarm rate. The following introduces the scenario of using the scenic area garbage identification method based on causal features in the present application.
[0035] The scenic area garbage identification device using this solution first collected 20,000 real environment background pictures from a wetland park, and collected sample pictures of 8 common garbage including snack packaging bags, beverage bottles, cigarette butts, etc. The scenic area garbage identification device uses an asymmetric twin network architecture and extracts features based on the ResNet18 model. After 500 rounds of training, it successfully built an intelligent system that can accurately identify garbage. In actual application, the new system can complete the recognition of the input monitoring image in only 3 seconds. Compared with the traditional YOLO algorithm, the garbage recognition miss rate in complex backgrounds is reduced by about 10%: the recognition accuracy in areas with complex ground textures is increased to 85%; the small garbage detection rate in shaded areas reaches 75%; the garbage recognition accuracy in green belts exceeds 80%.
[0036] It can be seen that the asymmetric twin network feature extraction method in the embodiment of the present application can not only realize accurate garbage identification, but also effectively solve the interference problem caused by complex background, thereby realizing stable and reliable garbage detection.
[0037] For ease of understanding, the following describes the process of the method provided by this implementation in combination with the above scenario. Figure 1 , which is a flow chart of a method for identifying scenic area garbage based on causal features in an embodiment of the present application.
[0038] S101, obtaining a plurality of scenic area monitoring images of a target park scenic area, and adjusting the plurality of scenic area monitoring images to a preset standard size to obtain a plurality of standardized images.
[0039] Among them, the target park scenic area represents the specific park area where garbage identification needs to be carried out; the scenic area monitoring image refers to the image data collected in real time by the monitoring cameras fixed in the park; the preset standard size is used to represent the image size that has been uniformly standardized, usually set to 512×512 or 1024×1024 pixels; the standardized image refers to the standardized image data after resizing and normalization.
[0040] When the scenic area garbage identification device starts to build a training data set, it first needs to obtain basic image data. Specifically, the device collects real-time images through the surveillance camera network distributed in the scenic area, performs quality inspection and resolution inspection on the collected original images, and selects the corresponding adjustment method according to the actual size when it is found that the image size does not meet the preset standard. Finally, all images are uniformly adjusted to the preset standard size, laying the foundation for the subsequent garbage area and background area extraction.
[0041] S102 , extracting garbage areas and background areas from a plurality of standardized images respectively to obtain a garbage sample set and a background sample set.
[0042] Among them, the garbage area refers to the various object areas contained in the image that need to be cleaned, such as leaves, beverage bottles, cigarette butts, etc.; the background area refers to the environmental scene area other than the garbage area, such as grass, water surface, road surface, etc.; the garbage sample set refers to the set of all extracted garbage areas; the background sample set refers to the set of all extracted background areas.
[0043] After obtaining the standardized image, the scenic area garbage identification device needs to separate and extract the image content. Specifically, the device first performs semantic segmentation on the standardized image, identifies the boundaries of the garbage area and the background area in the image, and then extracts and labels these areas respectively. The garbage areas are sorted by category to form a garbage sample set, and the background areas are sorted by scene type to form a background sample set, providing basic data for subsequent image fusion.
[0044] S103 , performing image fusion on each garbage area in the garbage sample set and multiple background areas in the background sample set to generate a garbage image training set.
[0045] The garbage image training set refers to a data set obtained after fusion for model training. The garbage image training set includes multiple garbage image groups, and each garbage image group includes multiple images of garbage in multiple backgrounds.
[0046] After obtaining the garbage sample set and the background sample set, the scenic area garbage identification device needs to construct training data. Specifically, the device naturally merges each garbage sample with multiple different types of background samples to form an image group containing the same garbage but different backgrounds. All image groups together constitute the final training data set to ensure that the model can learn garbage features that are unrelated to the background.
[0047] In some embodiments, image fusion can be achieved in a variety of ways: optionally, first adjust the color space of the garbage area to make it coordinated with the color tone of the target background, then perform edge feathering to achieve smooth transition, and finally use the Poisson fusion algorithm to seamlessly embed the garbage area into the background; optionally, first match the lighting conditions of the garbage area, then use the gradient domain fusion method to achieve regional splicing, and finally adjust the fusion effect through the global optimization algorithm. It is understandable that other image processing algorithms can also be used to achieve natural fusion effects, which are not limited here.
[0048] After the images are fused, the content of each image is composed of "garbage + background", such as leaves in the grass, cigarette boxes on the water, flowers on the snow, beverage bottles on the ground, etc. P={pi} represents the garbage to be identified, including leaves, cigarette boxes, flowers, beverage bottles, etc., and G={gi} represents all possible backgrounds, including grass, water, snow, ground, road surface, etc. When constructing the training set, a garbage object pi (such as leaves) to be identified is fused with each background to obtain a set of samples Spi={sk}={leaves in the grass, leaves on the water, leaves on the snow, leaves on the ground, etc.}. Each garbage object to be identified will generate such a set of samples, which are merged together to form the garbage image training set S=∪Spi.
[0049] In addition, one piece of garbage here contains multiple forms, such as the complete form, incomplete form and folded form of a cigarette box, from which more garbage images can be derived, such as Spi={{sk1}, {sk2}, {sk3}, ...}={{complete cigarette box in the grass, complete cigarette box on the water, complete cigarette box on the ground...}, {incomplete cigarette box in the grass, incomplete cigarette box on the water, incomplete cigarette box on the ground...}, {folded cigarette box in the grass, folded cigarette box on the water, folded cigarette box on the ground...}...}.
[0050] S104. Construct a twin network including two coding branches, set a gradient backpropagation blocking operation for the second coding branch of the twin network, and obtain an asymmetric twin network.
[0051] Among them, the Siamese Network is a special neural network architecture that contains two or more identical sub-networks that share the same parameters and weights. They are mainly used for comparative learning and comparing the similarity of two inputs. For example, two identical Resnet18 networks constitute a Siamese network. The encoding branch is a sub-network in the Siamese network, responsible for encoding (converting) the input image into a feature vector. For example, an encoding branch can be a Resnet18 structure that converts a 512×512 image into a feature vector. In deep learning, the model updates parameters through backpropagation, and Stop-gradient prevents the gradient from passing through certain layers or branches during backpropagation. An asymmetric Siamese network is a twin network structure with two branches with different training strategies.
[0052] After obtaining the training data set, the scenic area garbage identification device needs to build a feature extraction model. Specifically, the device first builds a twin network containing two identical encoding network branches, and then sets a gradient blocking operation on the second branch so that its parameters remain unchanged during the training process, and only updates the parameters of the first branch, thereby forming an asymmetric training structure, which helps to learn to extract the characteristics of the object itself (causal characteristics) while ignoring the changing background information (correlation characteristics), and then extract stable feature representations. That is, let one branch remain stable as a feature extractor, and the other branch learns to extract key features through gradient updates, preventing training instability caused by simultaneous changes in the two branches, and helping the network converge to a better feature representation.
[0053] Correspondingly, the ultimate goal of model training is to extract the features of pi from sk. To this end, the self-supervised method and the decoupled representation learning method are combined, and a contrastive learning framework based on an asymmetric twin network is adopted to extract feature information.
[0054] In some embodiments, the construction of an asymmetric twin network can be achieved in a variety of ways: optionally, using the ResNet series model as the basic network of the encoding branch, building a twin structure by sharing weights, and then using PyTorch's detach() function to achieve gradient blocking; optionally, using Vision Transformer as a feature extractor, replicating and building a dual-branch structure, and achieving asymmetry of parameter updates through the stop_gradient operation. It is understandable that other deep learning architectures can also be used to achieve the construction of feature extraction networks, which are not limited here.
[0055] S105. From the junk image group in the junk image training set, select a junk image pair with the same junk but different backgrounds, input the two images of the junk image pair into the two encoding branches of the asymmetric twin network respectively, perform model training with minimizing the cosine distance of the output vectors of the two encoding branches as the training standard, and determine the feature extraction model based on the first encoding branch that has completed the training.
[0056] Among them, the garbage image pair refers to two images containing the same garbage but different backgrounds; the cosine distance refers to a metric to measure the similarity of two vectors.
[0057] After the network structure is built, the scenic area garbage identification device needs to be trained. Specifically, the device selects image pairs of the same garbage in different backgrounds from the training set and inputs them into the two branches of the asymmetric twin network respectively. The network is trained by minimizing the cosine distance between the feature vectors output by the two branches, so that it can extract the essential features of garbage that are independent of the background. Finally, the first branch that has been trained is selected, that is, the sub-network that is used for parameter update is used as the feature extraction model.
[0058] Correspondingly, during the training process, by performing a stop-gradient operation on a path, the contrastive learning method can better learn the effective features of the sample. The model is trained in the above manner for each Spi in the set S until convergence. After the training, the obtained encoding network has the ability to extract objects (i.e., causal features) from the input samples.
[0059] For details, please refer to Figure 2 , Figure 2 middle Figure 2 (a) is a schematic diagram of the architecture for model training based on an asymmetric twin network; Figure 2 In (a), the training architecture of the asymmetric Siamese network extracts causal features of objects by inputting image pairs containing the same object but different backgrounds.
[0060] The scenic area garbage recognition device first fuses the object to be identified (beverage bottle) with the background i (grassland) and background j (open space) respectively to generate two training samples. These two samples are simultaneously input into two encoding networks with exactly the same structure.
[0061] Although the two encoding networks have the same structure, they adopt an asymmetric training strategy. The right branch sets a gradient block (stop-grad, i.e. stop-gradient) to keep its parameters unchanged during training, while the left branch updates its parameters normally. This design enables the left branch to stably learn the essential features of the object, because the branch makes its output features as similar as possible to the output features of the right branch (achieved by minimizing the cosine distance between the two).
[0062] Since the two input samples differ only in the background, the training process will cause the network to gradually ignore the influence of the background and focus on extracting the features of the object itself. The left branch also adds an additional prediction layer for the final classification task to generate a feature extraction model.
[0063] S106: Input the monitored image to be tested into the feature extraction model to obtain a garbage recognition result.
[0064] The monitored image to be tested refers to a real-time monitored image that needs to be garbage identified; the garbage identification result refers to the model's judgment result on whether there is garbage in the input image.
[0065] After completing the model training, the scenic area garbage identification device can be put into practical use. Specifically, the device receives the real-time images collected by the scenic area monitoring camera, inputs them into the trained feature extraction model, extracts and analyzes the image features, determines whether there is garbage that needs to be cleaned up in the image, and outputs the identification results to provide a basis for subsequent cleaning work.
[0066] For details, please refer to Figure 2 , Figure 2 middle Figure 2 (b) is a schematic diagram of the scenario of garbage identification using feature extraction model; Figure 2 In (b), the system receives a scene image containing potential garbage objects (a beverage bottle on the grass) as input and sends it to the trained encoding network (i.e., the left branch network after training) for processing.
[0067] The encoding network extracts key features of objects in the image and outputs recognition results based on these features. Possible recognition results include multiple predefined garbage categories (such as packaging bags, beverage bottles, cigarette butts, etc.) or a judgment of "no garbage".
[0068] Since the model has learned to ignore background interference and focus on extracting the essential features of objects during the training phase, it can maintain stable recognition performance in various complex background environments. For example, for this scene image (a beverage bottle on the grass), the model will prioritize the essential features of possible garbage objects such as shape and texture, and will not be disturbed by background features such as the green color or texture of the grass, thereby making accurate recognition judgments.
[0069] It should be noted that the scenic area garbage identification device has been tested in a railway park in a certain city. When identifying various types of garbage in different environments, it has been significantly improved compared with ordinary target recognition and garbage detection algorithms. The details are as follows: 1. Environment preparation: The operating environment of the "Scenic Area Garbage Identification Device Based on Causal Feature Extraction" is: Intel i9-10900X processor, NVIDIA GeForce RTX 3090 GPU, 128G memory, deployment of Python 3.6 and Pytorch1.4.0.
[0070] 2. Training data: To adapt to the actual environmental characteristics of a certain city's park scenic area, we sampled 20,000 background images from real environments such as a railway park and a wetland park, and each background image was adjusted to 512 x 512. We collected 80 images of 8 common types of garbage, including snack packaging bags, beverage bottles, cigarette butts, cigarette boxes, paper scraps, branches, leaves, and stones, and each image was adjusted to 24x24 to construct a training data set.
[0071] 3. Model training and packaging. Resnet18 is used as the basic architecture of the encoding network to form an asymmetric twin network. It is trained according to the above method for 500 epochs to generate a parameterized encoding network, which is then packaged to form a "scenic area garbage identification device based on causal feature extraction".
[0072] 4. Garbage identification: The single-frame image captured by the surveillance camera of the park scenic area is resized and input into the device, and the output is whether there is garbage in the image. One identification takes about 3 seconds.
[0073] 5. Recognition effect: Compared with traditional target detection models (such as YOLO), in a real park scenic environment, the device's missed detection rate for garbage recognition in various environmental backgrounds is reduced by about 10%.
[0074] In the above embodiment, the scenic area garbage identification device realizes feature extraction of garbage in different scenes based on image fusion technology. In practical applications, the identification strategy can also be dynamically adjusted according to the characteristics of different scenes, such as identifying garbage based on scenes in different seasons, to further improve system performance. The following is a supplement to the scene of this embodiment.
[0075] The garbage identification device in the scenic area was upgraded to seasonal adaptability six months later. By collecting data on environmental changes in four seasons, the system established a seasonal feature background library: when the sea of flowers blooms in spring, the system can accurately distinguish between garbage and petals; when the grass and trees are lush in summer, it can maintain a high recognition rate for garbage hidden in the vegetation; when the leaves fall in autumn, it can accurately distinguish between garbage and fallen leaves; when snow accumulates in winter, it can still identify white garbage.
[0076] After combining the above scenarios, the following is a more detailed description of the process of the method provided by this implementation. Figure 3 , is another flow chart of the method for identifying scenic area garbage based on causal features in an embodiment of the present application.
[0077] S301, obtaining multiple scenic area monitoring images of a target park scenic area, and calculating the Laplace operator variance of the scenic area monitoring images to obtain image clarity.
[0078] The Laplace operator variance represents a mathematical indicator for evaluating the clarity of an image; and image clarity refers to a quantitative value of the clarity of the edges and details of an image.
[0079] When the scenic area garbage identification device is screening images, it first needs to evaluate the image quality. Specifically, the device collects the original image through the scenic area monitoring system, and then calculates the second-order differential Laplace operator for each image, and calculates the variance value of the operator response. The larger the value, the clearer the image edges and details, thereby obtaining an objective image clarity evaluation result.
[0080] S302: Calculate the average value of the image pixel values in the scenic area monitoring image to obtain the image brightness value.
[0081] Among them, the image pixel value represents the brightness or color information of each pixel in the image; the image brightness value refers to the quantitative index of the overall brightness level.
[0082] When evaluating image quality, the scenic area garbage identification device needs to calculate the overall brightness of the image. Specifically, the device converts the color image into a grayscale image, counts the brightness values of all pixels, calculates the arithmetic mean of these values, and obtains a quantitative index reflecting the overall brightness level of the image.
[0083] In some embodiments, the brightness value calculation can be implemented in a variety of ways: optionally, using RGB to YUV color space conversion, directly extracting the average value of the Y channel as the brightness index; optionally, using the V channel value of the HSV color space, combined with the statistical characteristics of the image histogram to calculate the weighted average brightness. It is understandable that other brightness calculation methods can also be used to obtain the image brightness value, which is not limited here.
[0084] S303: When the image clarity is greater than a preset clarity threshold and the image brightness value is within a preset brightness range, the corresponding scenic area monitoring image is taken as a valid image.
[0085] Among them, the preset clarity threshold represents the lowest acceptable value of image clarity; the preset brightness range refers to the reasonable value range of image brightness; and the valid image represents an image that meets the quality requirements and can be used for subsequent processing.
[0086] After obtaining the image quality evaluation index, the scenic area garbage identification device needs to screen qualified images. Specifically, the device compares the clarity value of each image with the preset clarity threshold, and checks whether the image brightness value falls within the preset reasonable range. Only images that meet both conditions will be marked as valid images, ensuring that the images processed subsequently have sufficient quality.
[0087] In some embodiments, image screening can be achieved in a variety of ways: optionally, the clarity threshold is set to 80% of the statistical mean of the Laplace operator variance, the brightness interval is [40, 220], and a double condition filtering is performed to obtain a valid image; optionally, an adaptive threshold strategy is used to dynamically adjust the screening criteria according to the historical image quality distribution to achieve more flexible quality control. It is understandable that other image quality control methods can also be used to achieve effective image screening, which is not limited here.
[0088] S304: Select a preset number of target images from the valid images, adjust the target images to a preset standard size, and obtain a plurality of standardized images.
[0089] Among them, the preset number indicates the number of target images that need to be selected; the target image refers to the image selected from the valid images for subsequent processing; and the preset standard size indicates a unified image processing size specification.
[0090] After obtaining a valid image set, the scenic area garbage identification device needs to perform normalization processing. Specifically, the device selects a predetermined number of images from the valid image set as target images according to a certain strategy, and then resizes these images and uniformly converts them into a preset standard size to provide standardized input data for subsequent feature extraction and analysis.
[0091] In some embodiments, the scenic area garbage identification device will enlarge or reduce the target image based on a preset standard size, that is, select a preset number of target images from the valid images and obtain the image size data of each target image; when the image size data is smaller than the preset standard size, use the bicubic interpolation algorithm to enlarge the target image to obtain the target enlarged image; when the image size data is larger than the preset standard size, use the regional mean pooling method to compress the target image to obtain the target reduced image; normalize the target image equal to the preset standard size, the target enlarged image and the target reduced image to obtain multiple standardized images.
[0092] Among them, the preset standard size represents a unified image processing specification; the bicubic interpolation algorithm is an image enlargement method; regional mean pooling is a way of image compression; and normalization processing refers to the operation of mapping pixel values to a standard range.
[0093] When processing target images, the garbage identification device in scenic spots needs to standardize the size. Specifically, the device first obtains the size information of each target image, and then uses the bicubic interpolation algorithm to enlarge the image that is too small, and uses regional mean pooling to compress the image that is too large, and finally normalizes the pixel values of all processed images to obtain standardized images with uniform specifications. Among them, the selection of the preset standard size needs to balance the computational efficiency and image detail preservation, and can usually be set to common sizes such as 224×224 or 256×256. The bicubic interpolation algorithm achieves smooth image enlargement through weighted calculation of 16 adjacent pixels, and the regional mean pooling achieves image compression by calculating the average value of the local area. The normalization process maps the pixel value to the [0, 1] interval to ensure the consistency of data distribution.
[0094] S305 , extracting garbage areas and background areas from a plurality of standardized images respectively to obtain a garbage sample set and a background sample set.
[0095] Referring to step S102, the scenic area garbage identification device obtains a garbage sample set and a background sample set.
[0096] S306: Count the number of garbage samples of each type in the garbage sample set.
[0097] Among them, type refers to different kinds of garbage, such as plastic bottles, paper scraps, leaves, etc.; sample quantity refers to the statistical value of the number of samples of each type of garbage.
[0098] After obtaining the garbage sample set, the scenic area garbage identification device needs to perform sample analysis. Specifically, the device classifies and counts the samples in the garbage sample set, calculates the number of each type of garbage samples, and provides a decision-making basis for subsequent sample balancing and data enhancement.
[0099] In some embodiments, sample statistics can be implemented in a variety of ways: optionally, a garbage type dictionary is established, and category mapping and counting are performed by traversing the sample set to generate a sample distribution statistics report; optionally, a database management system is used to record sample information, and the number of samples in each category is counted through SQL query. It is understandable that other data analysis methods can also be used to implement sample statistics, which are not limited here.
[0100] S307: When the sample quantity of a garbage sample is less than the preset sample quantity, the garbage sample is regarded as a garbage missing sample.
[0101] Among them, the preset number of samples represents the minimum number standard that each type of garbage samples needs to meet; garbage missing samples refer to garbage category samples with insufficient sample quantity.
[0102] After completing the sample statistics, the scenic area garbage identification device needs to identify the category with insufficient samples. Specifically, the device compares the number of samples of each type of garbage with the preset minimum number of samples. When the number of samples of a certain type of garbage is lower than the preset value, it is marked as a garbage missing sample.
[0103] S308, completing the number of samples of garbage missing samples to the preset number of samples.
[0104] Among them, sample completion will be applied to data enhancement processing, which refers to the process of generating new samples through image transformation, including image rotation, flipping and brightness adjustment; image rotation refers to the angle transformation around the center point; flipping includes horizontal and vertical mirroring; brightness adjustment refers to changing the overall brightness of the image.
[0105] After identifying the missing garbage samples, the scenic area garbage identification device needs to expand the samples. Specifically, the device applies multiple image transformation operations to each missing garbage sample, including rotation at different angles, horizontal and vertical flipping, and brightness adjustment to different degrees, to generate new samples until the number of samples reaches the preset requirement.
[0106] In some embodiments, data enhancement can be achieved in a variety of ways: optionally, the image is randomly rotated within the range of [-30°, 30°], a random horizontal flip is performed, and the brightness range is adjusted to be between [0.8, 1.2] times the original value; optionally, a variety of enhancement methods are combined, such as adding Gaussian noise, adjusting contrast, cropping, etc., to generate more diverse samples. It is understandable that other image processing methods can also be used to achieve data enhancement, which is not limited here.
[0107] S309: Acquire historical monitoring images of the target park scenic area in different seasons as seasonal feature images.
[0108] Among them, the seasonal feature images represent historical monitoring images that reflect the characteristics of different seasons, including four periods: spring, summer, autumn and winter.
[0109] When expanding background samples, the garbage identification device in scenic spots needs to consider seasonal changes. Specifically, the device selects images that can clearly reflect the characteristics of different seasons in spring, summer, autumn and winter from historical monitoring data as seasonal feature images, providing a data source for subsequent background feature extraction.
[0110] In some embodiments, seasonal feature image acquisition can be achieved in a variety of ways: optionally, historical images are classified by season according to timestamps, and representative scene images are selected from each season; optionally, cluster analysis is performed based on the color features and texture features of the image to extract typical scene images of different seasons. It is understandable that other image analysis methods can also be used to achieve seasonal feature image acquisition, which is not limited here.
[0111] S310, extracting the same background area of different image features in each seasonal characteristic image to obtain a seasonal characteristic background set.
[0112] Among them, image features refer to the visual features of the background in different seasons; and the seasonal feature background set refers to a set of background samples containing different seasonal features.
[0113] After obtaining the seasonal characteristic images, the scenic area garbage identification device needs to extract the seasonal background. Specifically, the device extracts the background areas of the same position or similar type from the characteristic images of each season, analyzes the characteristic changes in different seasons, and organizes them into a background sample set containing seasonal change characteristics.
[0114] In some embodiments, seasonal feature background extraction can be implemented in a variety of ways: optionally, using image registration technology to align images of different seasons, extracting background samples of the same area, and forming a seasonal change sequence; optionally, based on scene classification results, extracting background samples of similar scenes in different seasons, and building a seasonal feature background library. It is understandable that other computer vision methods can also be used to implement seasonal feature background extraction, which is not limited here.
[0115] S311, performing image fusion on each garbage area in the garbage sample set and multiple background areas in the seasonal characteristic background set to generate a garbage image training set.
[0116] After obtaining the seasonal characteristic background set, the scenic area garbage identification device needs to generate training data. Specifically, the device naturally merges each garbage area in the garbage sample set with multiple backgrounds in the seasonal characteristic background set to generate image groups containing the same garbage but with different seasonal backgrounds. All image groups together constitute the final training data set.
[0117] In some embodiments, the scenic area garbage identification device will also adjust the garbage samples, that is, calculate the main color tone, brightness distribution and texture characteristics of each background area in the seasonal characteristic background set to obtain seasonal characteristic parameters; according to the seasonal characteristic parameters, perform color space conversion and parameter adjustment on the garbage area in the garbage sample set to obtain seasonal adaptive garbage samples; extract the contour features of the seasonal adaptive garbage samples, and smooth the contour edges to obtain edge-optimized garbage samples; gradually fuse the edge-optimized garbage samples with the background areas in the seasonal characteristic background set to generate training images; combine multiple training images with the same garbage but in different seasons or different backgrounds into a garbage image group to generate a garbage image training set.
[0118] Among them, the main color tone refers to the main color distribution of the image; the seasonal characteristic parameters include visual features such as color, brightness and texture; and gradient fusion refers to the image synthesis method to achieve smooth transition.
[0119] When generating training data, the garbage identification device in scenic spots needs to ensure the naturalness of the samples. Specifically, the device first analyzes the visual feature parameters of the background, adjusts the color and brightness features of the garbage samples accordingly, then optimizes the edges of the garbage samples, and finally uses gradient fusion technology to naturally embed the processed garbage samples into the backgrounds of different seasons to form a realistic training data set. Among them, the extraction of seasonal feature parameters can use methods such as color histograms and grayscale co-occurrence matrices, the adjustment of garbage samples needs to consider color balance and lighting consistency, edge optimization can use smoothing algorithms such as Gaussian blur, and gradient fusion can be achieved using alpha blending or Poisson editing.
[0120] S312. Construct a twin network including two coding branches, set a gradient backpropagation blocking operation for the second coding branch of the twin network, and obtain an asymmetric twin network.
[0121] Referring to step S104, the scenic area garbage identification device will construct an asymmetric twin network.
[0122] S313. From the junk image group in the junk image training set, select a junk image pair with the same junk but different backgrounds, input the two images of the junk image pair into the two encoding branches of the asymmetric twin network respectively, perform model training with minimizing the cosine distance of the output vectors of the two encoding branches as the training standard, and determine the feature extraction model based on the first encoding branch that has completed the training.
[0123] Referring to step S105, the scenic area garbage identification device will be trained to obtain a feature extraction model.
[0124] S314: Input the monitored image to be tested into the feature extraction model to obtain the garbage recognition result.
[0125] Referring to step S106 , the scenic area garbage identification device will perform garbage identification based on the feature extraction model.
[0126] In some embodiments, the scenic area garbage identification device will perform targeted optimization of the model based on the identification results, that is, obtain the user's annotation information on the garbage identification results, count the differences between the annotation information and the garbage identification results, and obtain model performance evaluation indicators; model performance evaluation indicators include recognition accuracy, missed detection rate and false detection rate; according to the model performance evaluation indicators, determine the scene type with performance lower than the preset threshold, and obtain the scene to be optimized; from the garbage image training set, screen training samples with the same characteristics as the scene to be optimized, and construct a scene-oriented training set; use the scene-oriented training set to perform incremental training on the feature extraction model to obtain a feature extraction optimization model; input the monitoring image to be tested into the feature extraction optimization model to obtain a garbage identification optimization result.
[0127] Among them, labeling information refers to the manually confirmed garbage location information; performance evaluation indicators are used to quantify the recognition effect of the model; incremental training refers to targeted optimization based on the original model.
[0128] In the actual application process of scenic area garbage identification device, it is necessary to continuously optimize the model performance. Specifically, the device collects user feedback annotation information, obtains performance indicators through comparative analysis with model recognition results, identifies the scene types that need to be optimized, and then specifically constructs training sample sets for incremental training of the model, and finally obtains an optimized model with better performance. Among them, the calculation of performance evaluation indicators is based on the confusion matrix, including standard indicators such as accuracy and recall rate. The construction of scene-oriented training sets needs to consider the representativeness and balance of samples. Incremental training requires the use of appropriate learning rates and training strategies to ensure that the model can maintain its original performance while improving the recognition effect of specific scenes.
[0129] In addition, after the garbage identification device was running for a period of time, the analysis of the operation data revealed that the recognition effect varied in different time periods and areas. For example, in the fitness area from 7 to 9 in the morning, due to the dim light and dense crowds, the recognition accuracy of mineral water bottles was only 75%; in the dining area from 11 to 14 noon, due to the complex types of garbage and severe obstruction, the missed detection rate reached 30%; and the viewing platform from 16 to 18 o'clock was affected by backlight, and the false detection rate exceeded 20%.
[0130] In order to optimize the recognition effect in these specific scenarios, the system will first collect garbage labeling information that has not been recognized by the system through the mobile phone app of the target user, the cleaning staff, including data such as time, location, and garbage type, and record the location information of the system's false alarms. Based on this garbage labeling information, the system constructs a confusion matrix and calculates the performance indicators of each area in each time period. When it is found that the recognition performance of the fitness area, dining area, and viewing platform is lower than the preset threshold of 80%, additional training samples are collected in these scenarios. Subsequently, the system uses a small learning rate of 0.001 to perform incremental training on the model, improving the recognition effect of specific scenarios while maintaining the original recognition ability.
[0131] Specifically, the scenic area garbage identification device has established a complete model optimization closed-loop system, which conducts scene-oriented performance evaluation through user feedback, and conducts targeted training after identifying the optimization target. In the performance evaluation stage, the system uses the confusion matrix to calculate various indicators: the accuracy rate is equal to the number of correctly identified samples divided by the total number of samples predicted to be garbage, the missed detection rate is equal to the number of unidentified garbage samples divided by the total number of actual garbage samples, and the false positive rate is equal to the number of samples incorrectly identified as garbage divided by the total number of samples predicted to be garbage. When performing incremental training, the system uses the gradient descent method to update the model parameters, but uses a smaller learning rate to ensure that the model can be fine-tuned to a better state while maintaining the original performance.
[0132] For example, suppose that in a restaurant area during a certain period of time, the system predicts a total of 100 suspected garbage locations, of which 70 are correct, 30 are false alarms, and 30 actual garbage locations are not identified. Then the accuracy rate in this scenario is 70% (70 / 100), the missed detection rate is 30% (30 / 100), and the false detection rate is 30% (30 / 100). Based on these indicators, the system determines that the scenario needs to be optimized, and then specifically collects training samples containing various occlusion situations and complex backgrounds, and improves the performance of the model in this specific scenario through incremental training.
[0133] In the embodiment of the present application, due to the use of a feature extraction method based on an asymmetric twin network, combined with a seasonal background sample library for data enhancement and model training, it is possible to effectively extract garbage features that are independent of the background, and to achieve accurate identification of garbage in different seasons and scenarios. This solution effectively solves the problem of low recognition accuracy and high false alarm rate of traditional computer vision methods in complex backgrounds and seasonal changes, and thus achieves all-weather, all-scenario intelligent garbage recognition. Through a dynamically updated sample library and incremental training mechanism, the system can continuously optimize recognition performance, adapt to new types of garbage and environmental changes, and improve the efficiency and quality of scenic area environmental management.
[0134] The following describes the scenic area garbage identification device in the embodiment of the present invention from the perspective of hardware processing. Figure 4 , which is a schematic diagram of the structure of a physical device of a scenic area garbage identification device in an embodiment of the present application.
[0135] It should be noted that Figure 4 The structure of the scenic area garbage identification device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.
[0136] like Figure 4 As shown, the scenic area garbage identification device includes a central processing unit (CPU) 401, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 402 or the program loaded from the storage part 408 to the random access memory (RAM) 403, such as executing the method described in the above embodiment. In RAM 403, various programs and data required for system operation are also stored. CPU 401, ROM 402 and RAM 403 are connected to each other through bus 404. Input / output (I / O) interface 405 is also connected to bus 404.
[0137] The following components are connected to the I / O interface 405: an input section 406 including an audio input device, a button switch, etc.; an output section 407 including a liquid crystal display (LCD) and an audio output device, an indicator light, etc.; a storage section 408 including a hard disk, etc.; and a communication section 409 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to the I / O interface 405 as needed. A removable medium 411, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 410 as needed so that a computer program read therefrom is installed into the storage section 408 as needed.
[0138] In particular, according to an embodiment of the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through the communication part 409, and / or installed from a removable medium 411. When the computer program is executed by the central processing unit (CPU) 401, various functions defined in the present invention are performed.
[0139] It should be noted that specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, apparatus, or device.
[0140] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present invention. Each box in the flowchart or block diagram may represent a module, a program segment, or a part of a code, and the above-mentioned module, program segment, or a part of a code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the box may also occur in an order different from that marked in the accompanying drawings.
[0141] Specifically, the scenic area garbage identification device of this embodiment includes a processor and a memory, and a computer program is stored in the memory. When the computer program is executed by the processor, the scenic area garbage identification method based on causal features provided in the above embodiment is implemented.
[0142] As another aspect, the present invention further provides a computer-readable storage medium, which may be included in the scenic area garbage identification device described in the above embodiment; or may exist independently without being assembled into the scenic area garbage identification device. The above storage medium carries one or more computer programs, and when the above one or more computer programs are executed by a processor of the scenic area garbage identification device, the scenic area garbage identification device implements the scenic area garbage identification method based on causal features provided in the above embodiment.
[0143] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.
[0144] As used in the above embodiments, the term "when..." may be interpreted to mean "if..." or "after..." or "in response to determining..." or "in response to detecting...", depending on the context. Similarly, the phrases "upon determining..." or "if (the stated condition or event) is detected" may be interpreted to mean "if determining..." or "in response to determining..." or "upon detecting (the stated condition or event)" or "in response to detecting (the stated condition or event)", depending on the context.
[0145] Those skilled in the art can understand that to implement all or part of the processes in the above-mentioned embodiments, the processes can be completed by computer programs to instruct related hardware, and the programs can be stored in computer-readable storage media. When the programs are executed, they can include the processes of the above-mentioned method embodiments. The aforementioned storage media include: ROM or random access memory RAM, magnetic disk or optical disk and other media that can store program codes.
Claims
1. A method for identifying garbage in scenic spots based on causal characteristics, characterized in that: Applied to a garbage identification device in a scenic area, the method comprises: Acquire multiple scenic area monitoring images of a target park scenic area, and adjust the multiple scenic area monitoring images to a preset standard size to obtain multiple standardized images; Extracting garbage areas and background areas from the plurality of standardized images respectively to obtain a garbage sample set and a background sample set; Perform image fusion on each garbage area in the garbage sample set and multiple background areas in the background sample set to generate a garbage image training set; the garbage image training set includes multiple garbage image groups, each of which includes multiple images of garbage in multiple backgrounds; Constructing a twin network including two encoding branches, setting a gradient backpropagation blocking operation on the second encoding branch of the twin network to obtain an asymmetric twin network; From the junk image group of the junk image training set, a junk image pair with the same junk but different backgrounds is selected, and two images of the junk image pair are respectively input into two encoding branches of the asymmetric twin network, and model training is performed with minimizing the cosine distance of output vectors of the two encoding branches as the training standard, and a feature extraction model is determined based on the first encoding branch after training; The monitored image to be tested is input into the feature extraction model to obtain a garbage recognition result.
2. The method according to claim 1, characterized in that The step of acquiring a plurality of scenic area monitoring images of the target park scenic area, and adjusting the plurality of scenic area monitoring images to a preset standard size to obtain a plurality of standardized images specifically includes: Acquire multiple scenic area monitoring images of the target park scenic area, and calculate the Laplace operator variance of the scenic area monitoring images to obtain image clarity; Calculate the average value of the pixel values in the scenic area monitoring image to obtain the image brightness value; When the image clarity is greater than a preset clarity threshold and the image brightness value is within a preset brightness range, the corresponding scenic area monitoring image is taken as a valid image; A preset number of target images are selected from the valid images, and the target images are adjusted to a preset standard size to obtain a plurality of standardized images.
3. The method according to claim 2, characterized in that The step of selecting a preset number of target images from the valid images, adjusting the target images to a preset standard size, and obtaining a plurality of standardized images specifically includes: Selecting a preset number of target images from the valid images, and obtaining image size data of each of the target images; When the image size data is smaller than a preset standard size, a bicubic interpolation algorithm is used to enlarge the target image to obtain a target enlarged image; When the image size data is larger than the preset standard size, compressing the target image using a regional mean pooling method to obtain a target reduced image; The target image equal to the preset standard size, the target enlarged image and the target reduced image are normalized to obtain a plurality of standardized images.
4. The method according to claim 1, characterized in that After the step of respectively extracting garbage areas and background areas from the plurality of standardized images to obtain a garbage sample set and a background sample set, the method further comprises: Counting the number of garbage samples of each type in the garbage sample set; When the sample quantity of a certain garbage sample is less than the preset sample quantity, the garbage sample is regarded as a garbage missing sample; The sample quantity of the garbage missing samples is supplemented to the preset sample quantity.
5. The method according to claim 1, characterized in that The step of fusing each garbage area in the garbage sample set with a plurality of background areas in the background sample set to generate a garbage image training set specifically includes: Acquire historical monitoring images of the target park scenic area in different seasons as seasonal characteristic images; Extracting the same background area of different image features in each of the seasonal characteristic images to obtain a seasonal characteristic background set; Each garbage area in the garbage sample set is image-fused with a plurality of background areas in the seasonal characteristic background set to generate a garbage image training set.
6. The method according to claim 5, characterized in that The step of fusing each garbage area in the garbage sample set with a plurality of background areas in the seasonal characteristic background set to generate a garbage image training set specifically includes: Calculate the main color tone, brightness distribution and texture characteristics of each background area in the seasonal characteristic background set to obtain seasonal characteristic parameters; According to the seasonal characteristic parameters, color space conversion and parameter adjustment are performed on the garbage area in the garbage sample set to obtain seasonally adaptive garbage samples; Extracting the contour features of the seasonally adaptive garbage sample and smoothing the contour edge to obtain an edge-optimized garbage sample; Gradually merge the edge-optimized garbage samples with the background area in the seasonal characteristic background concentration to generate a training image; A plurality of the training images having the same garbage but in different seasons or different backgrounds are combined into a garbage image group to generate a garbage image training set.
7. The method according to claim 1, characterized in that After the step of inputting the monitored image to be tested into the feature extraction model to obtain the garbage identification result, the method further includes: Obtaining user's annotation information on the spam identification result, counting the difference between the annotation information and the spam identification result, and obtaining a model performance evaluation index; the model performance evaluation index includes recognition accuracy, missed detection rate and false detection rate; According to the model performance evaluation index, determine the scene type whose performance is lower than a preset threshold, and obtain the scene to be optimized; From the junk image training set, select training samples having the same characteristics as the scene to be optimized to construct a scene-oriented training set; Using the scene-oriented training set to incrementally train the feature extraction model to obtain a feature extraction optimization model; The monitored image to be tested is input into the feature extraction optimization model to obtain a garbage recognition optimization result.
8. A scenic area garbage identification device, characterized in that: The scenic area garbage identification device includes: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to enable the scenic area garbage identification device to execute the method described in any one of claims 1-7.
9. A computer-readable storage medium comprising instructions, characterized in that: When the instruction is executed on a scenic area garbage identification device, the scenic area garbage identification device is caused to execute the method as described in any one of claims 1-7.
10. A computer program product, characterized in that When the computer program product runs on a scenic area garbage identification device, the scenic area garbage identification device is enabled to execute the method as described in any one of claims 1-7.
Citation Information
Patent Citations
Garbage type detection and identification method and device based on deep learning
CN114863255A
Picture generation method and device, computer equipment and storage medium
CN116597252A
Unmanned aerial vehicle remote sensing garbage identification method, medium and system
CN118314484A
Cited By
Intelligent classification and identification method and system for urban solid waste
CN120823452A
Garbage detection training data generation method and device based on image fusion and medium
CN121280850A