Spam detection method, computer device and medium

By combining image acquisition equipment with a trained recognition model to identify litter in the environment and using the distance between objects and people to determine litter, the accuracy and efficiency of litter detection in the environment are solved, realizing intelligent litter detection and timely cleanup.

CN114937208BActive Publication Date: 2026-01-27BOE TECHNOLOGY GROUP CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210685202.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-17
Publication Date
2026-01-27
Estimated Expiration
2042-06-17

AI Technical Summary

Technical Problem

Existing technologies are ineffective at detecting waste in the environment, especially in non-specific scenarios.

Method used

An image acquisition device is used to acquire environmental images, and a first recognition model that has been trained is used to identify objects of a preset category. A second recognition model that has been trained is used to identify people. The distance between the object and the person is used to determine whether it is garbage, and a preset distance threshold is used to determine whether it is garbage.

Benefits of technology

It enables accurate and efficient detection of waste in the environment, reduces false alarm rates, provides a basis for timely waste removal, and enhances environmental beautification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114937208B_ABST
    Figure CN114937208B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a garbage detection method, a computer device and a medium. In a specific embodiment, the method comprises: acquiring an environment image collected by an image collection device; inputting the environment image into a trained first recognition model and a trained second recognition model respectively, so as to recognize objects belonging to a preset category in the environment image through the trained first recognition model, and recognize a person in the environment image through the trained second recognition model; and determining garbage in the objects belonging to the preset category in the environment image according to a distance between the objects belonging to the preset category and the person in the environment image. The embodiment can accurately and efficiently intelligently detect garbage in the environment, and provides a reliable basis for timely cleaning of the garbage to realize beautification of the environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer technology. More specifically, it relates to a waste detection method, computer equipment, and media. Background Technology

[0002] Currently, waste detection solutions are typically designed for specific scenarios such as trash cans and waste stations, and cannot be applied to waste detection in the environment. Summary of the Invention

[0003] The purpose of this invention is to provide a waste detection method, computer equipment, and medium to solve at least one of the problems existing in the prior art.

[0004] To achieve the above objectives, the present invention adopts the following technical solution:

[0005] The first aspect of this invention provides a waste detection method, comprising:

[0006] Acquire environmental images captured by the image acquisition device;

[0007] The environmental image is input into a trained first recognition model and a trained second recognition model, respectively, so that the trained first recognition model identifies objects belonging to a preset category in the environmental image, and the trained second recognition model identifies people in the environmental image; and

[0008] Based on the distance between objects belonging to a preset category and people in the environmental image, the garbage in the objects belonging to the preset category in the environmental image is determined.

[0009] Optionally, determining the presence of litter among objects belonging to a preset category in the environmental image based on the distance between them and people includes:

[0010] Determine whether the distance between an object belonging to a preset category in the environmental image and the person closest to that object is greater than a preset distance threshold. If not, then the object is determined to be trash.

[0011] Optionally, the preset distance threshold is set to 3 to 7 times the width of the person closest to the object in the environmental image.

[0012] Optionally,

[0013] The environmental images acquired by the image acquisition device include multiple frames of environmental images acquired by the image acquisition device.

[0014] The step of inputting the environmental images into a trained first recognition model and a trained second recognition model respectively, so as to identify objects belonging to a preset category in the environmental images through the trained first recognition model and to identify people in the environmental images through the trained second recognition model, includes: inputting the multiple frames of environmental images into the trained first recognition model and the trained second recognition model respectively, so as to identify objects belonging to a preset category in the multiple frames of environmental images through the trained first recognition model and to identify people in the multiple frames of environmental images through the trained second recognition model;

[0015] The step of determining the trash in the environmental image based on the distance between objects belonging to a preset category and people in the environmental image includes:

[0016] In the multi-frame environmental images, it is determined whether the distance between the object belonging to the same preset category and the person closest to the object in each frame is greater than a preset distance threshold. If not, the object is determined to be garbage.

[0017] Optionally, the preset distance threshold is set to 3 to 7 times the width of the person closest to the object in the environmental image.

[0018] Optionally, after determining the trash in the objects belonging to the preset category in the environmental image based on the distance between the objects and people in the environmental image, the method further includes: issuing a prompt message based on the determined trash in the objects belonging to the preset category in the environmental image.

[0019] Optionally, the prompt information includes the geographic location information of the environmental image, as well as a reminder of the presence of litter and / or an environmental image marked with litter.

[0020] Optionally, before inputting the environmental image into the trained first recognition model and the trained second recognition model respectively, the method further includes:

[0021] A first recognition model is trained using images containing objects belonging to a preset category and images containing objects belonging to other categories different from the preset category.

[0022] Optionally, after training a first recognition model using images containing objects belonging to a preset category and images containing objects belonging to other categories different from the preset category, the method further includes:

[0023] In response to the adjustment operation, the preset category is updated; and

[0024] A first recognition model is obtained by retraining using images containing objects belonging to a preset category and images containing objects belonging to other categories different from the preset category.

[0025] A second aspect of the present invention provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the garbage detection method provided in the first aspect of the present invention.

[0026] A third aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the garbage detection method provided in the first aspect of the present invention.

[0027] The beneficial effects of this invention are as follows:

[0028] The technical solution described in this invention can accurately, efficiently, and intelligently detect garbage in the environment, providing a reliable foundation for timely garbage removal and environmental beautification. Attached Figure Description

[0029] The specific embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.

[0030] Figure 1 An exemplary system architecture diagram is shown, in which an embodiment of the present invention can be applied.

[0031] Figure 2 This diagram illustrates a flow chart of the waste detection method provided in an embodiment of the present invention.

[0032] Figure 3 This diagram illustrates another flow chart of the waste detection method provided in an embodiment of the present invention.

[0033] Figure 4 A schematic diagram of a waste detection system provided in an embodiment of the present invention is shown.

[0034] Figure 5 A schematic diagram of the structure of a computer system implementing the server provided in the embodiments of the present invention is shown. Detailed Implementation

[0035] To more clearly illustrate the present invention, the following description, in conjunction with embodiments and accompanying drawings, further explains the invention. Similar components in the drawings are indicated by the same reference numerals. Those skilled in the art should understand that the specific description below is illustrative rather than restrictive and should not be construed as limiting the scope of protection of the present invention.

[0036] This invention provides a waste detection method, comprising the following steps:

[0037] Acquire environmental images captured by the image acquisition device;

[0038] The environmental image is input into a trained first recognition model and a trained second recognition model, respectively, so that the trained first recognition model identifies objects belonging to a preset category in the environmental image, and the trained second recognition model identifies people in the environmental image; and

[0039] Based on the distance between objects belonging to a preset category and people in the environmental image, the garbage in the objects belonging to the preset category in the environmental image is determined.

[0040] The waste detection method provided in this embodiment of the invention inputs collected environmental images into two trained models to identify people and objects belonging to preset categories. Then, it determines whether an object belonging to a preset category is discarded waste based on the distance between the object and the person. This method can accurately and efficiently detect waste in the environment, providing a reliable foundation for timely waste cleanup and environmental beautification. The first trained recognition model identifies objects belonging to preset categories in the environmental images. Defining the categories of objects that may be waste accelerates the convergence of the first recognition model and improves its accuracy. Furthermore, this embodiment provides a judgment logic that determines whether an object belonging to a preset category is discarded waste based on the distance between the object and the person, which can reduce the false alarm rate and improve the accuracy of waste detection.

[0041] The garbage detection method provided in this embodiment can be implemented by a computer device with data processing capabilities. Specifically, the computer device can be a computer with data processing capabilities, including a personal computer (PC), a minicomputer, or a mainframe, or a server or server cluster with data processing capabilities. This embodiment does not limit the specific computer device to this type.

[0042] To facilitate understanding of the technical solution in this embodiment, the following is combined with... Figure 1 The method provided in this embodiment will be described in a practical scenario. See [link to relevant documentation]. Figure 1The scenario includes a training server 101, a recognition server 102, and a judgment server 103. In this embodiment, the training server 101 uses images containing objects belonging to a preset category and images containing objects belonging to other categories different from the preset category as first training samples to train a first recognition model for recognizing objects belonging to the preset category in environmental images, thereby obtaining a trained or pre-trained first recognition model. It then uses images containing people and images containing trees, vehicles, and other objects as second training samples to train a second recognition model for recognizing people in environmental images, thereby obtaining a trained or pre-trained second recognition model. Subsequently, the recognition server 102 can acquire environmental images, such as urban environmental images, collected by image acquisition devices 111 (e.g., road cameras installed at intersections, surveillance cameras installed at shopping mall entrances), and then use the first and second recognition models trained by the training server 101 to recognize objects and people belonging to the preset category in the environmental images, respectively. Subsequently, the judgment server 103 can determine the garbage based on the recognition results output by the recognition server 102 and the distance between objects and people belonging to the preset category in the environmental image, and obtain the garbage judgment result or garbage detection result, that is, obtain the garbage information in the environment. Based on the garbage detection result, it can issue prompt information such as cleaning prompts, which can be received and viewed by sanitation workers through terminal devices 121 such as smartphones.

[0043] It is important to note that Figure 1 The training server 101, recognition server 102, and decision server 103 in the configuration can, in practical applications, be three independent servers or a single server integrating model training, recognition, and decision functions. When they are independent servers, they can communicate with each other via a network, which can include various connection types, such as wired, wireless communication links, or fiber optic cables. The recognition server 102 and the image acquisition device 111, and the decision server 103 and the terminal device 121, can communicate via networks, which can also include various connection types, such as... Figure 1 The wireless communication links shown are examples of this.

[0044] Next, from the perspective of a processing device with data processing capabilities, the waste detection method provided in this embodiment will be described.

[0045] One embodiment of the present invention provides a waste detection method, such as... Figure 2 As shown, the process includes steps S210-S250, where step S210 belongs to the training phase, and steps S220 and thereafter belong to the detection phase. A detailed explanation follows.

[0046] like Figure 2 As shown, the waste detection method provided in this embodiment includes the following steps:

[0047] S210. Train a first recognition model for recognizing objects belonging to a preset category in an image and a second recognition model for recognizing people in an image.

[0048] In one possible implementation, the first recognition model trained to identify objects belonging to a preset category in an image further includes:

[0049] A first recognition model is obtained by training images containing objects belonging to a preset category and images containing objects belonging to other categories different from the preset category. That is, a trained or pre-trained first recognition model is obtained.

[0050] In a specific example, the first recognition model can employ CNN-based image instance segmentation algorithms such as Mask-RCNN (Mask-Regions with CNN features), SSD (Single Shot MultiBox Detector), or YOLO (You Only Look Once), or classification deep neural network algorithms such as ResNet50 (residual network) or DenseNet (classification network). For instance, the first recognition model can use the YOLOv5 model. YOLO (You Only Look Once) is a deep neural network-based object recognition and localization model. The YOLOv5 model has good recognition performance and fast inference speed, making it suitable for online deployment. Furthermore, YOLOv5 considers multi-scale target scenarios in its network structure and anchor box design, effectively addressing the recognition of smaller targets (objects or people belonging to a predefined category). When training the YOLOv5 model, it uses a data loader to pass and augment each batch of training data (the data loader performs three types of data augmentation: scaling, color space adjustment, and mosaic enhancement). This process significantly broadens the training data, greatly improving the model's generalization ability. For example, the YOLOv5 model includes a backbone network, a fusion network, and a bounding box prediction layer. The backbone network, which can be a convolutional neural network, performs multi-level upsampling of the image to obtain images of different fine-grained resolutions. The fine-grainedness of the upsampled feature map corresponding to each level is greater than that of the previous level, and the image features of each upsampled feature map are extracted. The fusion network includes a series of network layers that mix and combine image features. The fusion network can be a Feature Pyramid Network (FPN) or a Path Aggregation Network (PANet) to fuse image features and pass the fused image features to the bounding box prediction layer. The bounding box prediction layer is used to generate the bounding box corresponding to the target to be identified, and to segment the target to be identified from the image to be identified based on the bounding box, thus obtaining the image of the target to be identified.

[0051] In a specific example, the first step is to set preset categories, which define the types of objects that might be classified as trash in subsequent assessments. If the number of preset categories is small and the negative sample richness is insufficient, more targets that distinguish the target trash types can be added, i.e., more categories can be added as positive samples for training. In this way, defining the types and range of objects that may be trash can make the inter-class differences between target types obvious, which facilitates training; adding more other types that are different from the target types to train the first recognition model can make the target category feature distribution clearer and reduce false detections.

[0052] Specifically:

[0053] When setting preset categories and defining waste types, common household waste can be categorized as: trash bottles, trash bags, boxed waste, and paper scraps. These categories are characterized by clear intra-category features and large inter-category distances. For example, trash bags are mostly irregularly shaped plastic objects, trash bottles are mostly cylindrical, boxed waste is mostly square, and paper scraps are mostly thin, light-colored, and small in area. The primary recognition model can learn to extract features such as shape and color through training and perform identification. By clearly defining the range of waste types, model convergence can be accelerated, and recognition accuracy improved.

[0054] If, as described above, there are four types of trash—garbage bottles, garbage bags, boxed trash, and paper scraps—and one background type, then the feature distribution is as follows: the first recognition model learns features for all four target categories (positive samples) and learns some features for negative samples. However, because negative samples typically lack richness and all negative samples are learned as a single category, their features are not easily converged (due to the limited similarity between target categories). This might lead to the first recognition model not being able to identify an unknown object as a negative sample and not as a specific type of trash, thus compromising the accuracy of the first recognition model and hindering its practical application. To address this issue, additional categories, such as clothing and shoes, can be added to the existing trash categories. After adding these new categories, both the existing and new categories will converge based on their respective positive sample features. Simultaneously, the number of target categories corresponding to the background class decreases, which is beneficial for convergence and improves the performance of the first recognition model, achieving the required accuracy.

[0055] In a specific example, the second recognition model trained to identify people in an image further includes:

[0056] Using images containing people and images containing other objects that are distinct from people, such as trees and vehicles, as second training samples, a second recognition model for identifying people in environmental images is trained, resulting in a trained or pre-trained second recognition model.

[0057] In a specific example, similar to the first recognition model type, the second recognition model can also use the YOLOv5 model.

[0058] S220: Acquire environmental images captured by the image acquisition device.

[0059] In a specific example, the image acquisition devices, such as road cameras installed at intersections or surveillance cameras installed at the entrance of shopping malls, can be one or more. If there are multiple devices, the environmental images acquired by each image acquisition device will be acquired separately, and subsequent processes will be performed on the acquired multiple environmental images. The device can be configured to acquire the environmental images and the corresponding geographical location information of the image acquisition devices. The geographical location information is such as XX Street XX Intersection, XX Shopping Mall West Gate, XX Park North Gate, etc.

[0060] In a specific example, environmental images captured by an image acquisition device can be displayed on a monitoring screen. For instance, the geographical location information of the corresponding image acquisition device can also be displayed next to the corresponding environmental image.

[0061] In a specific example, to ensure the accuracy of subsequent identification and judgment, the environmental images can be specified. For example, 1080P environmental images from park cameras can be directly processed in subsequent processes. If the environmental images are from high-definition road cameras, such as 2K images, the environmental images can be first sliced ​​or cut into sub-images, and then the subsequent processes can be carried out on the sub-images respectively.

[0062] S230. The environmental image is input into the trained first recognition model and the trained second recognition model respectively, so as to identify objects belonging to a preset category in the environmental image through the trained first recognition model, and to identify people in the environmental image through the trained second recognition model.

[0063] The trained first recognition model identifies objects in the environmental image that belong to a preset category. This allows for optimization of the object recognition process by defining the categories of objects that may be garbage, thereby accelerating the convergence of the first recognition model and improving its recognition accuracy.

[0064] S240. Based on the distance between objects belonging to a preset category and people in the environmental image, determine the garbage in the objects belonging to the preset category in the environmental image.

[0065] In one possible implementation, step S230 further includes:

[0066] Determine whether the distance between an object belonging to a preset category in the environmental image and the person closest to that object is greater than a preset distance threshold. If not, then the object is determined to be trash.

[0067] For example, for each object belonging to a preset category in an environmental image, the distance between the object's bounding box and the bounding boxes of all identified people can be traversed (calculated one by one). It is then determined whether the minimum distance is greater than a preset distance threshold. If not, the object is determined to be unattended and thus classified as garbage. Conversely, if the distance is, the object is determined to be under someone's care and thus not classified as garbage.

[0068] In one possible implementation, the preset distance threshold is set to 3 to 7 times the width of the person closest to the object in the environmental image. For example, the preset distance threshold is set to 5 times the width of the person closest to the object in the environmental image. The width of the person in the environmental image can be calculated from the pixel width of the person's bounding box in the environmental image, satisfying the principle of objects appearing larger when closer and smaller when farther away. Therefore, this adaptive distance threshold setting method can improve the accuracy of the judgment.

[0069] In one possible implementation,

[0070] Step S220 further includes: acquiring multiple frames of environmental images captured by the image acquisition device;

[0071] Step S230 further includes: inputting the multi-frame environmental images into a trained first recognition model and a trained second recognition model respectively, so as to identify objects belonging to a preset category in the multi-frame environmental images through the trained first recognition model, and to identify people in the multi-frame environmental images through the trained second recognition model; and

[0072] Step S240 further includes: determining whether the distance between the object belonging to the same preset category and the person closest to the object in each frame of the multi-frame environmental images is greater than a preset distance threshold; if not, the object is determined to be garbage. Similar to the aforementioned implementation, in this implementation, the preset distance threshold is, for example, set to 3 to 7 times the width of the person closest to the object in the environmental image.

[0073] Therefore, by combining time and space information, it is possible to determine whether objects belonging to a preset category are being watched, thus avoiding misjudgments such as misclassifying an object belonging to a preset category that is not being watched as being watched and not considered garbage if a pedestrian happens to be passing by in a frame of an environmental image. This further clarifies the act of discarding and improves the accuracy of the judgment.

[0074] In one possible implementation, the duration of the multi-frame environmental images is between 3 and 10 seconds. For example, the duration of the multi-frame environmental images is 5 seconds. This duration is the time threshold, and the judgment logic is that only if a person is present within the distance threshold for 5 consecutive seconds is an object belonging to a preset category determined to be trash.

[0075] In one possible implementation, the multi-frame environmental images are consecutive multi-frame environmental images. This ensures the accuracy of the determination. Furthermore, considering factors such as data processing volume, when the image acquisition device has a high acquisition frequency, the multi-frame environmental images can also be set to a preset interval. For example, if acquisition occurs once per second, the multi-frame environmental images can be consecutive; if acquisition occurs three times per second, the multi-frame environmental images can be set to a preset interval of two.

[0076] In a specific example, during the process of image acquisition by the image acquisition device in real time or at fixed intervals, steps S220-S240 are executed cyclically to continuously perform garbage detection.

[0077] In one possible implementation, during the cyclic execution of steps S220-S240 of the garbage detection method provided in this embodiment, i.e., the detection phase after the training phase, the garbage detection method provided in this embodiment further includes:

[0078] In response to the adjustment operation, the preset category is updated; and

[0079] A first recognition model is obtained by retraining using images containing objects belonging to a preset category and images containing objects belonging to other categories different from the preset category.

[0080] In this way, the first recognition model can be retrained by adjusting the preset categories according to the actual situation, so as to flexibly adjust the categories of objects identified. For example, a new category of "discarded masks" can be added, so that environmental images containing discarded masks can be used as positive samples in the original first training samples for retraining, so that the first recognition model can learn the ability to identify discarded masks from environmental images.

[0081] In one possible implementation, such as Figure 2 As shown, the waste detection method provided in this embodiment also includes:

[0082] S250. Based on the garbage in the environmental image obtained from the determination that belongs to a preset category, issue a prompt message.

[0083] In one possible implementation, the prompt information includes the geographic location information of the environmental image, as well as a reminder of the presence of litter and / or an environmental image marked with litter.

[0084] Therefore, receiving prompts, such as from smartphones, can output prompts through display and / or voice broadcast, allowing sanitation workers to know the type, quantity, and location of the garbage so they can take timely action.

[0085] In a specific example, step S220 can acquire the geographical location information of the camera that acquired the environmental image while acquiring the environmental image, and use this information as the geographical location information of the environmental image.

[0086] In a specific example, combining the above implementation methods, such as Figure 3 As shown, a typical process of the waste detection method provided in this embodiment is as follows:

[0087] After acquiring environmental image frames from the image acquisition device, the system performs person recognition and object recognition belonging to preset categories, and tracks them separately using a tracker. This includes summarizing and updating person bounding boxes, determining the tracking ID of objects, and then checking whether an object has been flagged based on its tracking ID. If no flag has been given, the system checks the distance between the object and the person based on a distance threshold and times the event to determine if the object is being monitored within a set time. If someone is monitoring the object within the set time, the suspected "garbage" status is cleared and the timer is reset to zero; otherwise, the object is considered "garbage," the flagging status is updated, and a flagging message is issued. The tracker can utilize target tracking algorithms such as Meanshift, Sort, particle filter-based motion estimation, contour-based tracking, and AI-based target tracking (MDNet, TCNN, SiamFC, GOTURN, etc.). For example, a tracker based on the Sort algorithm can be used to track all object detection boxes and person detection boxes in real time.

[0088] like Figure 4 As shown, another embodiment of the present invention provides a waste detection system, including a server;

[0089] The servers include:

[0090] The acquisition module is used to acquire environmental images captured by the image acquisition device;

[0091] The recognition module is used to input the environmental image into a trained first recognition model and a trained second recognition model, respectively, so as to identify objects belonging to a preset category in the environmental image through the trained first recognition model and to identify people in the environmental image through the trained second recognition model.

[0092] The determination module is used to determine the garbage in the environmental image based on the distance between objects belonging to a preset category and people in the environmental image.

[0093] In one possible implementation, the server also includes a display module for displaying the acquired environmental image.

[0094] In one possible implementation, such as Figure 4 As shown, the waste detection system provided in this embodiment also includes at least one terminal device;

[0095] The server also includes a publishing module, which is used to publish a prompt message based on the garbage in the environmental image that belongs to a preset category;

[0096] The terminal device is used to receive and output the prompt information.

[0097] It should be noted that the principle and workflow of the waste detection system provided in this embodiment are similar to the waste detection method described above. The relevant parts can be referred to the above description and will not be repeated here.

[0098] like Figure 5 As shown, a computer system suitable for implementing the server in the waste detection system provided in the above embodiments includes a central processing module (CPU), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) or a program loaded from a storage portion into a random access memory (RAM). Various programs and data required for the operation of the computer system are also stored in the RAM. The CPU, ROM, and RAM are connected via a bus. An input / output (I / O) interface is also connected to the bus.

[0099] The following components are connected to the I / O interface: input sections including keyboards, mice, etc.; output sections including liquid crystal displays (LCDs) and speakers, etc.; storage sections including hard disks, etc.; and communication sections including network interface cards such as LAN cards and modems. The communication sections perform communication processing via networks such as the Internet. Drives are also connected to the I / O interface as needed. Removable media, such as disks, optical disks, magneto-optical disks, semiconductor memories, etc., are installed on the drive as needed so that computer programs read from them can be installed into the storage section as required.

[0100] Specifically, according to this embodiment, the process described in the flowchart above can be implemented as a computer software program. For example, this embodiment includes a computer program product comprising a computer program tangibly embodied on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication component, and / or installed from a removable medium.

[0101] The flowcharts and schematic diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of the system, method, and computer program product of this embodiment. In this regard, each block in the flowchart or schematic diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the schematic diagram and / or flowchart, and combinations of blocks in the schematic diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0102] The modules described in this embodiment can be implemented in software or hardware. These modules can also be housed in a processor; for example, it can be described as: a processor including an acquisition module, an identification module, and a determination module. The names of these modules do not necessarily limit the functionality of the module itself. For example, the determination module can also be described as a "judgment module" or a "detection module."

[0103] On the other hand, this embodiment also provides a non-volatile computer storage medium. This non-volatile computer storage medium can be the non-volatile computer storage medium included in the above-described device, or it can be a separate non-volatile computer storage medium not installed in the terminal. The non-volatile computer storage medium stores one or more programs. When these programs are executed by a device, the device: acquires an environmental image captured by an image acquisition device; inputs the environmental image into a trained first recognition model and a trained second recognition model, respectively, to identify objects belonging to a preset category in the environmental image using the trained first recognition model and to identify people in the environmental image using the trained second recognition model; and determines whether the objects belonging to the preset category in the environmental image are garbage based on the distance between the objects and people in the environmental image.

[0104] In the description of this invention, it should be noted that the terms "upper," "lower," etc., indicating the orientation or positional relationship are based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Unless otherwise expressly specified and limited, the terms "installed," "connected," and "linked" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication between two elements. For those skilled in the art, the specific meaning of the above terms in this invention can be understood according to the specific circumstances.

[0105] It should also be noted that in the description of this invention, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0106] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not intended to limit the implementation of the present invention. For those skilled in the art, other variations or modifications can be made based on the above description. It is impossible to exhaustively list all the implementation methods here. All obvious variations or modifications derived from the technical solutions of the present invention are still within the protection scope of the present invention.

Claims

1. A waste detection method, characterized in that, include: Acquire multiple frames of environmental images captured by the image acquisition device; The multi-frame environmental images are respectively input into a trained first recognition model and a trained second recognition model, so that the trained first recognition model can identify objects belonging to a preset category in the multi-frame environmental images, and the trained second recognition model can identify people in the multi-frame environmental images. as well as Based on the distance between objects belonging to a preset category and people in the multi-frame environmental images, the garbage in the objects belonging to the preset category in the multi-frame environmental images is determined; The step of determining the trash in the multi-frame environmental images based on the distance between objects belonging to a preset category and people in the multi-frame environmental images includes: In the multi-frame environmental images, it is determined whether the distance between the object belonging to the same preset category and the person closest to the object in each frame is greater than a preset distance threshold. If not, the object is determined to be garbage. While acquiring multiple frames of environmental images captured by the image acquisition device, the geographical location information of the image acquisition device that acquired the environmental images is also acquired, and used as the geographical location information of the environmental images; After determining that the objects in the multi-frame environmental images belong to a preset category of garbage, the method further includes: issuing a prompt message based on the determined garbage in the multi-frame environmental images belonging to the preset category, the prompt message including the geographical location information of the environmental image, as well as a reminder message that garbage exists and an environmental image marked with garbage; Before inputting the multi-frame environmental images into the trained first recognition model and the trained second recognition model respectively, the method further includes: A first recognition model is trained using images containing objects belonging to a preset category and images containing objects belonging to other categories different from the preset category.

2. The method according to claim 1, characterized in that, The preset distance threshold is set to 3 to 7 times the width of the person closest to the object in the environmental image.

3. The method according to claim 1, characterized in that, After training a first recognition model using images containing objects belonging to a preset category and images containing objects belonging to other categories different from the preset category, the method further includes: In response to the adjustment operation, the preset category is updated; and A first recognition model is obtained by retraining using images containing objects belonging to a preset category and images containing objects belonging to other categories different from the preset category.

4. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method as described in any one of claims 1-3.

5. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-3.

Citation Information

Patent Citations

  • Recognition method and recognition device for random garbage throwing behavior and readable storage medium

    CN112115846A

  • Environmental health monitoring method and device based on computer vision

    CN113420730A

  • Mobile vehicle-mounted intelligent junk information management method and system

    CN113705638A