Target Object Detection Method, Device, Readable Storage Medium and Electronic Device
By detecting the target object of auxiliary materials in construction or decoration sites, determining the quantity of auxiliary materials and generating push information using a pre-trained detection model, the problem of inefficient auxiliary materials management in the existing technology is solved, real-time automatic monitoring and efficient management are achieved.
Patent Information
- Application Number
- CN202110438624.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-22
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-04-22
AI Technical Summary
In construction, decoration and other scenarios, it is difficult for the existing technology to accurately monitor and manage the remaining quantity, used quantity and available days of auxiliary materials, resulting in ineffective management.
By obtaining the site image taken on the site where the target object is stored, the image is input into the pre-trained target object detection model, the number of target objects is determined, and information for push is generated.
Real-time automatic monitoring of the number of target objects is achieved, reducing dependence on people and improving management efficiency.
Smart Images

Figure CN113111826B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and in particular, to a method and apparatus for detecting a target object, a computer-readable storage medium, and an electronic device. Background Art
[0002] Currently, when performing operations such as construction, maintenance, and construction on some sites, some necessary items are required, and the remaining quantity, used quantity, available days, etc. of these items often need to be estimated more accurately.
[0003] For example, in the scenario of an indoor decoration construction site, the support of decoration auxiliary materials is required, and the identification of the types of auxiliary materials and the effective monitoring of the quantity are crucial for construction site management. Common decoration auxiliary materials on construction sites usually include: waterproof coating, tile adhesive, plaster for plastering, water-resistant putty, wall primer, floor primer, quick-drying adhesive powder, etc. After the auxiliary materials enter the construction site, they are usually stacked in a designated area. Currently, the management of auxiliary materials in decoration construction sites mostly relies on foremen to go to the site for inspection or on monitoring cameras to check the usage of auxiliary materials at the current construction site. Summary of the Invention
[0004] Embodiments of the present disclosure provide a method and apparatus for detecting a target object, a computer-readable storage medium, and an electronic device.
[0005] Embodiments of the present disclosure provide a method for detecting a target object. The method includes: obtaining a site image of a site storing a target object; inputting the site image into a pre-trained target object detection model to obtain at least one target object individual detection information and target object category information corresponding to at least one target object individual detection information respectively; determining the current quantity of the target object corresponding to the same target object category information based on at least one target object individual detection information; and generating a push message for pushing to a target user terminal based on the quantity.
[0006] In some embodiments, generating a push message for pushing to a target user terminal based on the quantity includes: determining the consumption amount per unit time of the target object corresponding to the same target object category information; determining the available time of the target object corresponding to the same target object category information based on the consumption amount per unit time and the current quantity of the corresponding target object; and generating a push message including the available time.
[0007] In some embodiments, generating a push message for pushing to a target user terminal based on the quantity includes: determining whether the current quantity of the target object corresponding to the same target object category information meets a preset replenishment condition; and if the preset replenishment condition is met, generating a push message for reminding the target user that the target object corresponding to the same target object category information needs to be replenished.
[0008] In some embodiments, the target object is production materials; after generating a push message for reminding the target user that the target object corresponding to the same target object category information needs to be replenished if a preset replenishment condition is met, the method further includes: receiving feedback information from the target user's terminal; if the feedback information indicates that the production materials corresponding to the feedback information are not replenished, determining the remaining construction period of the application link based on the application link of the production materials corresponding to the feedback information and the quantity of the production materials corresponding to the feedback information, and pushing it.
[0009] In some embodiments, the target object detection model includes a first sub-model, a second sub-model, and a third sub-model; inputting the site image into the pre-trained target object detection model to obtain at least one target object individual detection information and the target object category information corresponding to at least one target object individual detection information respectively, including: inputting the site image into the first sub-model to obtain region mask information representing the stacking area of the target object; based on the region mask information, extracting the target object region image from the site image; inputting the target object region image into the second sub-model to obtain at least one target object stacking pile mask information and the target object category information corresponding to at least one target object stacking pile mask information respectively; based on at least one target object stacking pile mask information, extracting at least one target object stacking pile image from the target object region image; inputting at least one target object stacking pile image into the third sub-model respectively to obtain at least one target object individual mask information as the target object individual detection information, and obtaining the target object category information corresponding to at least one target object individual mask information respectively.
[0010] In some embodiments, based on at least one target object individual detection information, determining the current quantity of the target object corresponding to the same target object category information includes: based on at least one target object stacking pile mask information, determining at least one first target object stacking pile mask information representing the target object stacking pile located on the outer layer and at least one second target object stacking pile mask information representing the target object stacking pile located on the inner layer; based on at least one first target object stacking pile mask information, at least one second target object stacking pile mask information, and at least one target object individual mask information, determining the first quantity of the target object located on the outer layer and the second quantity of the target object located on the inner layer corresponding to each target object category information respectively; based on the first quantity and the second quantity, determining the current quantity of the target object corresponding to the same target object category information.
[0011] In some embodiments, the first sub-model is trained based on the following steps: obtaining a training sample set, wherein the training samples in the training sample set include sample site images and corresponding annotation region mask information for characterizing the stacking regions of target objects in the sample site images; using the sample site images included in the training samples in the training sample set as the input of a preset first initial model, and using the annotation region mask information corresponding to the input sample site images as the expected output of the first initial model to train the first sub-model.
[0012] In some embodiments, the second sub-model is trained based on the following steps: obtaining sample target object region images intercepted from the sample site images and corresponding annotation target object stacking mask information and annotation target object category information for characterizing the positions of at least one target object stacking; using the sample target object region images as the input of a preset second initial model, and using the annotation target object stacking mask information and annotation target object category information corresponding to the input sample target object region images as the expected output of the second initial model to train the second sub-model.
[0013] In some embodiments, the annotation target object stacking mask information includes first annotation target object stacking mask information for characterizing that the target object stacking is located in the outer layer and second annotation target object stacking mask information for characterizing that the target object stacking is located in the inner layer; using the annotation target object stacking mask information and annotation target object category information corresponding to the input sample target object region images as the expected output of the second initial model to train the second sub-model includes: using the first annotation target object stacking mask information, the second annotation target object stacking mask information, and the annotation target object category information corresponding to the input sample target object region images as the expected output of the second initial model to train the second sub-model.
[0014] In some embodiments, the third sub-model is trained based on the following steps: obtaining at least one sample target object stacking image intercepted from the sample target object region images and corresponding annotation target object individual mask information and annotation target object category information for characterizing the positions of the individuals of the target objects respectively; using the sample target object stacking images as the input of a preset third initial model, and using the annotation target object individual mask information and annotation target object category information corresponding to the input sample target object stacking images as the expected output of the third initial model to train the third sub-model.
[0015] According to another aspect of the embodiments of the present disclosure, a target object detection device is provided. The device includes: an acquisition module configured to acquire a site image of a site where a target object is stored; a detection module configured to input the site image into a pre-trained target object detection model to obtain at least one target object individual detection information and target object category information corresponding to the at least one target object individual detection information respectively; a determination module configured to determine the current quantity of the target object corresponding to the same target object category information based on the at least one target object individual detection information; and a generation module configured to generate a push message for pushing to a target user terminal based on the quantity.
[0016] In some embodiments, the generation module includes: a first determination unit configured to determine the consumption per unit time of the target object corresponding to the same target object category information; a second determination unit configured to determine the available time of the target object corresponding to the same target object category information based on the consumption per unit time and the current quantity of the corresponding target object; and a first generation unit configured to generate a push message including the available time.
[0017] In some embodiments, the generation module includes: a third determination unit configured to determine whether the current quantity of the target object corresponding to the same target object category information meets a preset replenishment condition; and a second generation unit configured to, if the preset replenishment condition is met, generate a push message for reminding the target user that the target object corresponding to the same target object category information needs to be replenished.
[0018] In some embodiments, the target object is production materials; the device further includes: a reception module configured to receive feedback information from the target user terminal; and a push module configured to, if the feedback information indicates not to replenish the production materials corresponding to the feedback information, determine the remaining construction period of the application link based on the application link of the production materials corresponding to the feedback information and the quantity of the production materials corresponding to the feedback information and push it.
[0019] In some embodiments, the target object detection model includes a first sub-model, a second sub-model, and a third sub-model; the detection module includes: a first detection unit configured to input a site image into the first sub-model to obtain region mask information representing the stacking area of the target object; a first extraction unit configured to extract a target object region image from the site image based on the region mask information; a second detection unit configured to input the target object region image into the second sub-model to obtain at least one target object stacking mask information and target object category information corresponding to the at least one target object stacking mask information respectively; a second extraction unit configured to extract at least one target object stacking image from the target object region image based on the at least one target object stacking mask information; a third detection unit configured to input the at least one target object stacking image into the third sub-model respectively to obtain at least one target object individual mask information as target object individual detection information, and to obtain target object category information corresponding to the at least one target object individual mask information respectively.
[0020] In some embodiments, the determination module includes: a fourth determination unit configured to determine at least one first target object stacking mask information representing the target object stacking located on the outer layer and at least one second target object stacking mask information representing the target object stacking located on the inner layer based on the at least one target object stacking mask information; a fifth determination unit configured to determine a first quantity of the target objects located on the outer layer and a second quantity of the target objects located on the inner layer corresponding to each target object category information based on the at least one first target object stacking mask information, the at least one second target object stacking mask information, and the at least one target object individual mask information; a sixth determination unit configured to determine the current quantity of the target objects corresponding to the same target object category information based on the first quantity and the second quantity.
[0021] In some embodiments, the first sub-model is trained based on the following steps: obtaining a training sample set, wherein the training samples in the training sample set include sample site images and corresponding labeled region mask information for representing the stacking areas of the target objects in the sample site images; using the sample site images included in the training samples in the training sample set as the input of a preset first initial model, and using the corresponding labeled region mask information of the input sample site images as the expected output of the first initial model to train and obtain the first sub-model.
[0022] In some embodiments, the second sub-model is trained based on the following steps: obtaining an image of a sample target object region cropped from a sample site image, and corresponding labeled target object stack mask information and labeled target object category information for characterizing the positions where at least one target object stack is placed; using the image of the sample target object region as the input of a preset second initial model, and using the corresponding labeled target object stack mask information and labeled target object category information of the input sample target object region image as the expected output of the second initial model, to train the second sub-model.
[0023] In some embodiments, the labeled target object stack mask information includes first labeled target object stack mask information for characterizing that the target object stack is located on the outer layer and second labeled target object stack mask information for characterizing that the target object stack is located on the inner layer; using the corresponding labeled target object stack mask information and labeled target object category information of the input sample target object region image as the expected output of the second initial model to train the second sub-model, includes: using the first labeled target object stack mask information, the second labeled target object stack mask information, and the labeled target object category information corresponding to the input sample target object region image as the expected output of the second initial model, to train the second sub-model.
[0024] In some embodiments, the third sub-model is trained based on the following steps: obtaining at least one image of a sample target object stack cropped from the sample target object region image, and corresponding labeled target object individual mask information and labeled target object category information for characterizing the positions where the individuals of the target object are located respectively for each of the at least one image of the sample target object stack; using the image of the sample target object stack as the input of a preset third initial model, and using the corresponding labeled target object individual mask information and labeled target object category information of the input sample target object stack image as the expected output of the third initial model, to train the third sub-model.
[0025] According to another aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium storing a computer program for executing the above-mentioned target object detection method.
[0026] According to another aspect of the embodiments of the present disclosure, there is provided an electronic device including: a processor; a memory for storing executable instructions of the processor; the processor for reading the executable instructions from the memory and executing the instructions to implement the above-mentioned target object detection method.
[0027] Based on the object detection method, device, computer-readable storage medium, and electronic device provided in the above embodiments of the present disclosure, by taking images of the site where the target object is stored, inputting the site images into the target object detection model, obtaining the individual detection information of the target object and the corresponding target object category information, then determining the current quantity of the target objects corresponding to the same target object category information based on the individual detection information of the target object, and finally generating a push message according to the quantity. It realizes the real-time automatic monitoring of the quantity of various target objects by using an electronic device, enables the user to timely know the usage situation of the target object, thereby reducing the dependence on people when monitoring the quantity of the target object and improving the efficiency of monitoring the quantity of the target object.
[0028] The technical solution of the present disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] By describing the embodiments of the present disclosure in more detail in conjunction with the accompanying drawings, the above and other objects, features, and advantages of the present disclosure will become more obvious. The accompanying drawings are used to provide a further understanding of the embodiments of the present disclosure, and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the present disclosure, and do not constitute a limitation to the present disclosure. In the accompanying drawings, the same reference numerals generally represent the same components or steps.
[0030] Figure 1 It is a system diagram applicable to the present disclosure.
[0031] Figure 2 It is a flowchart of the object detection method provided by an exemplary embodiment of the present disclosure.
[0032] Figure 3 It is a flowchart of the object detection method provided by another exemplary embodiment of the present disclosure.
[0033] Figure 4 It is a flowchart of the object detection method provided by another exemplary embodiment of the present disclosure.
[0034] Figure 5 It is a flowchart of the object detection method provided by another exemplary embodiment of the present disclosure.
[0035] Figure 6 It is a schematic diagram of extracting the object region image in the embodiment of the present disclosure.
[0036] Figure 7 It is a schematic diagram of extracting the object stacking image in the embodiment of the present disclosure.
[0037] Figure 8 It is a schematic diagram of determining the individual mask information of the object in the embodiment of the present disclosure.
[0038] Figure 9 It is a schematic structural diagram of an object detection model according to an embodiment of the present disclosure.
[0039] Figure 10 It is a schematic flowchart of an object detection method provided by another exemplary embodiment of the present disclosure.
[0040] Figure 11 It is a schematic structural diagram of an object detection device provided by an exemplary embodiment of the present disclosure.
[0041] Figure 12 It is a schematic structural diagram of an object detection device provided by another exemplary embodiment of the present disclosure.
[0042] Figure 13 It is a structural diagram of an electronic device provided by an exemplary embodiment of the present disclosure. Detailed Embodiments
[0043] Next, exemplary embodiments according to the present disclosure will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all of the embodiments of the present disclosure. It should be understood that the present disclosure is not limited by the exemplary embodiments described herein.
[0044] It should be noted that: Unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions, and numerical values set forth in these embodiments do not limit the scope of the present disclosure.
[0045] Those skilled in the art can understand that terms such as "first" and "second" in the embodiments of the present disclosure are only used to distinguish different steps, devices, or modules, etc., and neither represent any specific technical meaning nor indicate an inevitable logical order between them.
[0046] It should also be understood that in the embodiments of the present disclosure, "a plurality" may refer to two or more, and "at least one" may refer to one, two, or more.
[0047] It should also be understood that for any component, data, or structure mentioned in the embodiments of the present disclosure, unless otherwise clearly defined or given a contrary indication in the context, it can generally be understood as one or more.
[0048] In addition, the term " / and" in the present disclosure is only a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A / and B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in the present disclosure generally represents an "or" relationship between the associated objects before and after.
[0049] It should also be understood that the descriptions of the various embodiments in this disclosure emphasize the differences between the various embodiments, and the similarities or resemblances between them can be referred to each other. For the sake of brevity, they will not be elaborated one by one.
[0050] At the same time, it should be understood that, for the convenience of description, the dimensions of the various parts shown in the drawings are not drawn according to the actual proportional relationship.
[0051] The following description of at least one exemplary embodiment is actually only illustrative and in no way restricts the present disclosure and its application or use.
[0052] Techniques, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the said techniques, methods, and devices should be regarded as part of the specification.
[0053] It should be noted that: like reference numerals and letters denote like items in the following drawings, and thus, once an item is defined in one drawing, it does not need to be further discussed in subsequent drawings.
[0054] The embodiments of the present disclosure can be applied to electronic devices such as terminal devices, computer systems, servers, etc., which can operate together with many other general or special computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments, and / or configurations suitable for use with electronic devices such as terminal devices, computer systems, servers, etc. include, but are not limited to: personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network personal computers, small computer systems, large computer systems, and distributed cloud computing technology environments including any of the above systems, and so on.
[0055] Terminal devices, computer systems, servers and other electronic devices can be described in the general context of computer system-executable instructions (such as program modules) executed by a computer system. Generally, program modules may include routines, programs, object programs, components, logics, data structures, etc., which perform specific tasks or implement specific abstract data types. The computer system / server can be implemented in a distributed cloud computing environment where tasks are executed by remote processing devices linked through a communication network. In a distributed cloud computing environment, program modules can be located on local or remote computing system storage media including storage devices.
[0056] Overview of the Application
[0057] In scenarios such as construction, building, and decoration of some sites, the management of target objects is mostly carried out by relevant personnel going to the site for on-site inspection or by viewing the current usage of target objects through surveillance cameras. For example, in the scenario of an indoor decoration construction site, a foreman generally manages multiple construction sites at the same time, resulting in situations such as untimely supervision and inspection of auxiliary materials and missed inspections. Moreover, this kind of on-line or off-line inspection and statistics by human eyes is also very time-consuming. Therefore, the existing perception and statistics efficiency of auxiliary materials is very low, and there is also a situation where data resources cannot be effectively utilized.
[0058] Exemplary System
[0059] Figure 1 Exemplary system architecture 100 of a target object detection method or a target object detection device to which embodiments of the present disclosure can be applied is shown.
[0060] As Figure 1 shown, system architecture 100 may include a terminal device 101, a network 102, a server 103, and a camera 104. The network 102 is used to provide a medium for a communication link between the terminal device 101 and the server 103. The network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0061] A user may use the terminal device 101 to interact with the server 103 through the network 102 to receive or send messages, etc. Various communication client applications may be installed on the terminal device 101, such as monitoring applications, web browser applications, instant messaging tools, etc.
[0062] The terminal device 101 may be various electronic devices, including but not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), vehicle terminals (such as vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc.
[0063] The camera 104 is used to photograph the site where the target object is stored, obtain a site image, and send the site image to the terminal device 101 or the server 102 through a wired or wireless transmission method.
[0064] The server 103 may be a server that provides various services, such as a background image processing server for detecting target objects in the images uploaded by the terminal device 101 or the camera 104. The background image processing server may process the received images, obtain processing results (such as the number and category of target objects), and feedback the processing results to the terminal device 101.
[0065] It should be noted that the target object detection method provided by the embodiments of the present disclosure can be executed by the server 103 or the terminal device 101. Correspondingly, the target object detection device can be set in the server 103 or the terminal device 101.
[0066] It should be understood that Figure 1 the numbers of the terminal devices, networks, servers, and cameras in
[0067] Exemplary Method
[0068] Figure 2 is a schematic flow chart of the target object detection method provided by an exemplary embodiment of the present disclosure. This embodiment can be applied to an electronic device (such as Figure 1 the terminal device 101 or the server 103 shown), as Figure 2 shown, the method includes the following steps:
[0069] Step 201, obtain a site image of a site where a target object is stored.
[0070] In this embodiment, the electronic device can obtain a site image of a site where a target object is stored. Among them, the site image can be an image taken by the camera 104 shown Figure 1 of various sites. The target object can be a pre-specified specific category of object. As an example, the site can be an indoor decoration construction site, and the target object can be auxiliary materials for decoration, such as waterproof coating, tile adhesive, plaster for plastering, water-resistant putty, wall primer, floor primer, quick stick powder, etc.
[0071] Step 202, input the site image into a pre-trained target object detection model to obtain at least one target object individual detection information and target object category information corresponding to at least one target object individual detection information respectively.
[0072] In this embodiment, the electronic device can input the site image into a pre-trained target object detection model to obtain at least one target object individual detection information and target object category information corresponding to at least one target object individual detection information respectively. Among them, the target object individual detection information is used to characterize the position of the individual of the target object in the site image. Based on the target object individual detection information, a mark covering the target object image can be generated in the site image. The target object category information is used to characterize the category of the target object, and the form of the target object category information can include but is not limited to forms such as numbers, texts, symbols, etc.
[0073] The target object detection model is used to represent the corresponding relationship between the site image and the individual detection information and category information of the target object. Usually, the target object detection model can parse the input site image to obtain image feature information, which can be used to represent features such as line shapes, textures, and colors in the site image. Then, based on the image feature information, the area where the target object is located in the site image is determined and classified to obtain the individual detection information and category information of the target object. The target object detection model can be trained using machine learning methods based on a preset training sample set. Specifically, the electronic device can use machine learning methods, take the sample site images included in the training samples in the training sample set as input, and take the individual detection information and category information of the target object corresponding to the input sample site image as the expected output, and train the initial model (such as including the resnet network, yolact++ network, classifier, etc.). For each training input site image, the actual output can be obtained. Among them, the actual output is the data actually output by the initial model, which is used to represent the target object and its category. Then, the electronic device can use the gradient descent method and the backpropagation method to adjust the parameters of the initial model based on the actual output and the expected output, take the model obtained after each parameter adjustment as the initial model for the next training, and end the training when the preset training end condition is met (such as the number of training times reaches the preset number, or the loss value obtained based on the preset loss function converges), so as to train the target object detection model.
[0074] Step 203: Based on at least one piece of individual detection information of the target object, determine the current number of target objects corresponding to the same target object category information.
[0075] In this embodiment, the electronic device can determine the current number of target objects corresponding to the same target object category information based on at least one piece of individual detection information of the target object. Specifically, the electronic device can determine the number of stacks of the target object placement according to the arrangement rule of the above at least one piece of individual detection information of the target object, then determine the number of target objects included in each stack, and further calculate the current number of target objects corresponding to the same target object category information.
[0076] Step 204: Generate push information for pushing to the target user terminal based on the quantity.
[0077] In this embodiment, the electronic device can generate push information for pushing to the target user terminal based on the current number of target objects corresponding to the same target object category information. Among them, the target user terminal can be a preset terminal for receiving the above push information. For example, it can be a mobile terminal used by the foreman in charge of site construction.
[0078] In some alternative implementation manners, such as Figure 3 shown, step 204 may be performed as follows:
[0079] Step 2041, determine the consumption amount per unit time of the target objects corresponding to the same target object category information.
[0080] Wherein, the unit time may be any set time period, such as one day, one week, etc. The consumption amount per unit time may be determined according to the historical quantity of each type of target object. For example, for a certain type of target object, the remaining quantity the day before yesterday and the remaining quantity yesterday may be extracted, and the difference between the two is the consumption amount per unit time. Or, according to the remaining quantity every day in the past N days, the average of the consumption amounts every day may be taken to obtain the consumption amount per unit time.
[0081] Step 2042, based on the consumption amount per unit time and the current quantity of the corresponding target object, determine the available time of the target objects corresponding to the same target object category information.
[0082] For example, divide the current quantity of a certain type of target object by the consumption amount per unit time to obtain the available time of this type of target object.
[0083] Step 2043, generate a push message including the available time.
[0084] In this implementation manner, by generating the available time of the target objects corresponding to the same target object category information and pushing the available time to the target user, the target user can timely know the remaining available time of a certain type of target object, reminding the target user to reasonably arrange the usage of the target object, without manual calculation, thereby improving the efficiency of monitoring the usage of the target object.
[0085] In some alternative implementation manners, such as Figure 4 shown, step 204 may further include the following steps:
[0086] Step 2044, determine whether the current quantity of the target objects corresponding to the same target object category information meets a preset replenishment condition.
[0087] Wherein, the target objects corresponding to the same target object category information correspond to a preset replenishment condition. The preset replenishment condition is a condition indicating that the target object needs to be replenished. As an example, a corresponding quantity threshold may be set for each type of target object, so that it can be determined whether the quantity of each type of target object is less than or equal to the corresponding preset quantity threshold. If it is less than or equal to the preset quantity threshold, it is determined that the preset replenishment condition is met. Or, using the consumption amount per unit time determined by the above implementation manner, it is determined that the current quantity of a certain type of target object is less than or equal to its corresponding consumption amount per unit time, or less than or equal to the product of the consumption amount per unit time and a preset multiple, and it is determined that the preset replenishment condition is met.
[0088] Step 2045, if the preset condition to be supplemented is met, generate a push message for reminding the target user that the target object corresponding to the same target object category information needs to be supplemented.
[0089] In this implementation, by generating a corresponding push message when the quantity of a certain type of target object is insufficient, the target user can timely learn from the push message that the target object should be supplemented currently, without the need for manual inspection of each site, improving the efficiency of monitoring the usage of the target object.
[0090] In some optional implementation manners, the target object is production materials; after step 2045, the electronic device may further perform the following steps:
[0091] First, receive the feedback message of the target user terminal.
[0092] Among them, the feedback message is the information generated by the feedback operation of the target user on the push message after the target user terminal receives the push message sent in the above step 2045. For example, if the target user is a foreman at a construction site and receives a push message that a certain production material needs to be supplemented, and confirms that the production material does not need to be supplemented because the construction has been completed or is about to be completed, the feedback message is generated by clicking a button, sending text, etc., and the target user terminal sends the feedback message to the above electronic device.
[0093] Then, if the feedback message indicates not to supplement the production material corresponding to the feedback message, determine the remaining construction period of the application link according to the application link of the production material corresponding to the feedback message and the quantity of the production material corresponding to the feedback message, and push it.
[0094] Among them, the push object of the remaining construction period may include the above target user terminal, or other terminal devices with a pre-established corresponding relationship. For example, in the scenario of indoor decoration, the push object may be the terminal of the owner, the terminal of the worker, the terminal of the salesman, etc.
[0095] As an example, a certain production material is floor glue, and the application link of this production material is laying the floor. If the feedback message indicates that the construction has been completed and the floor glue does not need to be supplemented, it is determined that the remaining construction period is 0, and the remaining construction period is pushed to the above push object; if the feedback message indicates that the construction is not completed and the floor glue does not need to be supplemented, the remaining construction period is determined according to the current quantity of the floor glue and the consumption per unit time corresponding to the floor glue, and the remaining construction period is pushed to the above push object.
[0096] According to the feedback information received from the target user terminal, when the feedback information indicates that production materials are not replenished, the remaining construction period of the application link of the production materials can be pushed, enabling relevant personnel to obtain the usage situation of the production materials and the corresponding construction period status in a timely manner, thereby further improving the efficiency of the management of production materials.
[0097] The method provided by the above-mentioned embodiment of the present disclosure obtains individual detection information of the target object and corresponding target object category information by taking an image of the site where the target object is stored and inputting the site image into the target object detection model, and then determines the current quantity of the target object corresponding to the same target object category information based on the individual detection information of the target object. Finally, a push message is generated according to the quantity. It realizes the real-time automatic monitoring of the quantity of various target objects by using an electronic device, enables the user to know the usage situation of the target object in a timely manner, thereby reducing the dependence on people when monitoring the quantity of the target object and improving the efficiency of monitoring the quantity of the target object.
[0098] Further referring to Figure 5 , a flowchart showing another embodiment of the target object detection method is shown. In this embodiment, the target object detection model includes a first sub-model, a second sub-model, and a third sub-model. As Figure 5 shown, based on the above-mentioned Figure 2 shown embodiment, step 202 may include the following steps:
[0099] Step 2021, input the site image into the first sub-model to obtain region mask information representing the stacking area of the target object.
[0100] Among them, the first sub-model is used to represent the correspondence between the site image and the region mask information. According to the region mask information, a mark covering all target objects can be generated on the site image. As Figure 6 shown, the site image 601 includes an image of multiple target objects stacked together, and the region mask information 6011 covers the target objects.
[0101] Step 2022, extract the target object region image from the site image based on the region mask information.
[0102] Specifically, the electronic device can intercept the target object region image from the site image according to the region covered by the region mask information. As Figure 6 shown, the smallest circumscribed rectangle including the region covered by the region mask information is extracted from the site image 601 to obtain the target object region image 602.
[0103] Step 2023: Input the target object area image into the second sub-model to obtain at least one target object stacking mask information and the corresponding target object category information for each of the at least one target object stacking mask information.
[0104] Among them, the second sub-model is used to represent the correspondence between the target object area image and the target object stacking mask information and the target object category information. According to the target object stacking mask information, marks covering each target object stacking can be generated on the target object area image. As Figure 7 shown, the target object area image 701 includes images of multiple target object stackings. The target object stacking mask information 7011 - 7014 respectively cover the target object stackings located on the outer layer, and the target object stacking mask information 7015 - 7016 respectively cover the target object stackings located on the inner layer. The target object category information corresponding to the target object stacking mask information 7011 - 7013 indicates that the category of the target object is Category One (such as gypsum powder), and the target object category information corresponding to the target object stacking mask information 7014 - 7015 indicates that the category of the target object is Category Two (such as tile adhesive).
[0105] Step 2024: Based on at least one target object stacking mask information, extract at least one target object stacking image from the target object area image.
[0106] Specifically, the electronic device can intercept at least one target object stacking image from the target object area image according to the area covered by the target object stacking mask information. As Figure 7 shown, the minimum bounding rectangles containing the areas covered by the area mask information 7011 - 7014 are respectively extracted from the target object area image 701 to obtain the target object stacking images 702 - 705.
[0107] Step 2025: Input at least one target object stacking image into the third sub-model respectively to obtain at least one target object individual mask information as the target object individual detection information, and obtain the corresponding target object category information for each of the at least one target object individual mask information.
[0108] Among them, the third sub-model is used to represent the correspondence between the target object stacking image and the target object individual mask information and the target object category information. According to the target object individual mask information, marks covering each target object can be generated on the target object stacking image. As Figure 8 shown, one of the target object stacking images 801 includes images of multiple target objects. The target object individual mask information 8011 - 8014 respectively cover each target object, and the corresponding target object category information indicates that the category of the target object is Category One (such as gypsum powder).
[0109] In this embodiment, the structures of the above-mentioned first sub-model, second sub-model, and third sub-model can be the same. For example, the structures of the first sub-model, second sub-model, and third sub-model can be configured using the yolact++ model.
[0110] As an example, the model structure can be as Figure 9 shown. The model includes a residual network (such as ResNet101) 901. The image is input into the residual network 901, and after obtaining the feature map, it is input into the Feature Pyramid network 902 to obtain a new feature map. This feature map is sent to two branches respectively. One branch is the protonet network 903, which is a 4-layer fully connected network used to generate feature information based on the entire image with a size of S*S*K, where S*S is the size of the feature map and K is the number of channels of the feature map, which is a hyperparameter. The other branch is first sent to the convolutional layer 904 to obtain a new feature map, and then class prediction (performed by 905), score prediction for each mask information (performed by 906), coordinate box regression for bbox (bounding box) instances (performed by 907), and protonet coefficient prediction (performed by 908) are respectively performed on this feature map. The non-maximum suppression (NMS) module 911 in the figure arranges the candidate boxes in descending order of confidence. Starting from the candidate box with the highest confidence, it calculates the overlap degree between all the remaining candidate boxes and this candidate box. When the overlap degree is greater than 0.7, it deletes the candidate box. Then, it selects the candidate box with the highest confidence from the remaining candidate boxes and repeats the above operation until all candidate boxes are traversed, obtaining the final bbox.
[0111] The mask information for each instance is obtained by multiplying the predicted coefficient by the feature map output by the protonet. The number and position information of each instance are obtained by setting thresholds for the scores of the bbox and mask information. Specifically, the scores of the bbox and mask information range from [0, 1], and the larger the score value, the higher the accuracy of the bbox and mask information. By setting a threshold, bbox and mask information with scores less than the threshold can be deleted (performed by 909).
[0112] Finally, cropping (performed by 910) is performed on the mask information obtained by the protonet to obtain the mask information 912 for each instance.
[0113] The above Figure 5For the method provided in the corresponding embodiment, by setting a first sub-model, a second sub-model, and a third sub-model in the target object detection model, the site image, the target object area image, and the target object stacking image can be gradually detected respectively. Each sub-model has corresponding detection pertinence, which helps to improve the accuracy of detecting various target objects.
[0114] In some alternative implementation manners, based on the above Figure 5 corresponding embodiment, as Figure 10 shown, step 203 can be executed as follows:
[0115] Step 2031: Based on at least one target object stacking mask information, determine at least one first target object stacking mask information representing the target object stacking located on the outer layer and at least one second target object stacking mask information representing the target object stacking located on the inner layer.
[0116] Among them, the target object stacking located on the outer layer is the target object stacking that is completely displayed in the scene image. The target object stacking located on the inner layer is the target object stacking that is partially displayed in the scene image. As Figure 7 shown, the target object stacking mask information 7011 - 7014 is the first target object stacking mask information, and the target object stacking mask information 7015 - 7016 is the second target object stacking mask information.
[0117] Step 2032: Based on at least one first target object stacking mask information, at least one second target object stacking mask information, and at least one target object individual mask information, determine the first quantity of the target objects located on the outer layer and the second quantity of the target objects located on the inner layer respectively corresponding to each target object category information.
[0118] Among them, for a certain target object category, the first quantity can be obtained by counting the target object individual mask information included in the first target object stacking mask information, and the second quantity can be obtained according to the number of layers of the individual mask information included in the first target object stacking mask information. For example, the second quantity = the number of layers × the number of the second target object stacking mask information.
[0119] Step 2033: Based on the first quantity and the second quantity, determine the current quantity of the target objects corresponding to the same target object category information.
[0120] Specifically, the first quantity and the second quantity corresponding to each target object category information can be added together to obtain the current quantity of the target objects corresponding to the same target object category information.
[0121] In this implementation manner, by dividing the stacking mask information of the target objects into inner and outer layers and respectively counting the number of target objects in the outer layer and the number of target objects in the inner layer, the number of occluded target objects can be estimated, and the accuracy of determining the number of target objects can be improved.
[0122] In some alternative implementation manners, the above first sub-model is trained based on the following steps:
[0123] First, obtain a training sample set.
[0124] Among them, the training samples in the training sample set include sample site images and corresponding annotation region mask information for characterizing the stacking regions of the target objects in the sample site images.
[0125] Then, use the sample site images included in the training samples in the training sample set as the input of a preset first initial model, and use the corresponding annotation region mask information of the input sample site images as the expected output of the first initial model to train the first sub-model.
[0126] Specifically, the electronic device can use machine learning methods, use the sample site images included in the training samples in the training sample set as the input, and use the corresponding annotation region mask information of the input sample site images as the expected output to train the first initial model (such as Figure 9 the model structure shown). For each training input site image, an actual output can be obtained. Among them, the actual output is the data actually output by the first initial model and is used to characterize the region where the target object is located. Then, the electronic device can adopt the gradient descent method and the backpropagation method, adjust the parameters of the first initial model based on the actual output and the expected output, use the model obtained after each parameter adjustment as the initial model for the next training, and end the training when the preset training end condition is met (such as the number of training times reaches the preset number of times, and the loss value obtained based on the preset loss function converges), so as to train the first sub-model.
[0127] In some alternative implementation manners, the second sub-model is trained based on the following steps:
[0128] First, obtain the sample target object region images intercepted from the sample site images and the corresponding annotation target object stacking mask information and annotation target object category information for characterizing the positions of at least one target object stacking.
[0129] Then, use the sample target object region images as the input of a preset second initial model, and use the corresponding annotation target object stacking mask information and annotation target object category information of the input sample target object region images as the expected output of the second initial model to train the second sub-model.
[0130] Among them, the structure of the second initial model can be the same as that of the first initial model, and the training method is basically the same as that of the first sub-model described above, which will not be elaborated here.
[0131] In some optional implementation manners, the above-mentioned annotated target object stacking mask information includes first annotated target object stacking mask information representing that the target object stacking is located on the outer layer and second annotated target object stacking mask information representing that the target object stacking is located on the inner layer.
[0132] The electronic device can use the first annotated target object stacking mask information, the second annotated target object stacking mask information, and the annotated target object category information corresponding to the input sample target object region image as the expected output of the second initial model, and train to obtain the second sub-model.
[0133] The trained second sub-model can detect the first target object stacking mask information representing that the target object is located on the outer layer of the stacking and the corresponding target object category information, as well as the second target object stacking mask information representing that the target object is located on the inner layer of the stacking and the corresponding target object category information.
[0134] In this implementation manner, by separately annotating the target object stacking on the inner layer and the target object stacking on the outer layer, the trained second sub-model can detect the target object stacking on the outer layer and the inner layer, thereby improving the comprehensiveness and accuracy of detecting the target object stacking.
[0135] In some optional implementation manners, the third sub-model is trained based on the following steps:
[0136] First, obtain at least one sample target object stacking image intercepted from the sample target object region image and the annotated target object individual mask information and the annotated target object category information corresponding to the at least one sample target object stacking image respectively, which are used to represent the location of the individual of the target object.
[0137] Then, use the sample target object stacking image as the input of a preset third initial model, and use the annotated target object individual mask information and the annotated target object category information corresponding to the input sample target object stacking image as the expected output of the third initial model, and train to obtain the third sub-model.
[0138] Among them, the structure of the third initial model can be the same as that of the first initial model and the second initial model, and the training method is basically the same as that of the first sub-model and the second initial model described above, which will not be elaborated here.
[0139] The method for training the first sub-model, the second sub-model, and the third sub-model can make the trained first sub-model, second sub-model, and third sub-model accurately segment the images of the target object by obtaining a large number of sample site images as training samples and respectively annotating the sample site images, the sample target object region images, and the sample target object stacking images, thereby improving the accuracy of determining the number of target objects from the site images.
[0140] Exemplary Device
[0141] Figure 11 FIG. is a schematic structural diagram of a target object detection device provided by an exemplary embodiment of the present disclosure. This embodiment can be applied to an electronic device, such as Figure 11 As shown, the target object detection device includes: an acquisition module 1101, configured to acquire a site image of a site where a target object is stored; a detection module 1102, configured to input the site image into a pre-trained target object detection model to obtain at least one target object individual detection information and target object category information corresponding to the at least one target object individual detection information respectively; a determination module 1103, configured to determine the current number of target objects corresponding to the same target object category information based on the at least one target object individual detection information; a generation module 1104, configured to generate a push message for pushing to a target user terminal based on the number.
[0142] In this embodiment, the acquisition module 1101 can acquire a site image of a site where a target object is stored. Among them, the site image can be an image taken by Figure 1 As shown, the camera 104 takes images of various sites. The target object can be a specific category of object specified in advance. As an example, the site can be an indoor decoration construction site, and the target object can be auxiliary materials for decoration, such as waterproof coating, tile adhesive, plaster for painting, water-resistant putty, wall primer, floor primer, quick-drying powder, etc.
[0143] In this embodiment, the detection module 1102 can input the site image into a pre-trained target object detection model to obtain at least one target object individual detection information and target object category information corresponding to the at least one target object individual detection information respectively. Among them, the target object individual detection information is used to represent the position of the individual of the target object in the site image. Based on the target object individual detection information, a mark covering the target object image can be generated in the site image. The target object category information is used to represent the category of the target object, and the form of the target object category information can include but is not limited to forms such as numbers, texts, symbols, etc.
[0144] In this embodiment, the determining module 1103 may determine the current quantity of the target objects corresponding to the same target object category information based on at least one target object individual detection information. Specifically, the determining module 1103 may determine the number of stacks of the target objects placed according to the arrangement rule of the above at least one target object individual detection information, then determine the quantity of the target objects included in each stack, and further calculate the current quantity of the target objects corresponding to the same target object category information.
[0145] In this embodiment, the generating module 1104 may generate push information for pushing to the target user terminal based on the current quantity of the target objects corresponding to the same target object category information. The target user terminal may be a preset terminal for receiving the above push information. For example, it may be a mobile terminal used by a foreman in charge of site construction.
[0146] Refer to Figure 12 , Figure 12 which is a schematic structural diagram of a target object detection device provided by another exemplary embodiment of the present disclosure.
[0147] In some optional implementation manners, the generating module 1104 includes: a first determining unit 11041 for determining the consumption quantity per unit time of the target objects corresponding to the same target object category information; a second determining unit 11042 for determining the available time of the target objects corresponding to the same target object category information based on the consumption quantity per unit time and the current quantity of the corresponding target objects; a first generating unit 11043 for generating push information including the available time.
[0148] In some optional implementation manners, the generating module 1104 includes: a third determining unit 11044 for determining whether the current quantity of the target objects corresponding to the same target object category information meets a preset replenishment condition; a second generating unit 11045 for generating push information for reminding the target user that the target objects corresponding to the same target object category information need to be replenished if the preset replenishment condition is met.
[0149] In some optional implementation manners, the target object is production materials; the device further includes: a receiving module 1105 for receiving feedback information from the target user terminal; a pushing module 1106 for determining the remaining construction period of the application link and pushing it if the feedback information indicates not to replenish the production materials corresponding to the feedback information based on the application link of the production materials corresponding to the feedback information and the quantity of the production materials corresponding to the feedback information.
[0150] In some alternative implementation manners, the target object detection model includes a first sub-model, a second sub-model, and a third sub-model; the detection module 1102 includes: a first detection unit 11021, configured to input a site image into the first sub-model to obtain region mask information representing the stacking area of the target object; a first extraction unit 11022, configured to extract a target object region image from the site image based on the region mask information; a second detection unit 11023, configured to input the target object region image into the second sub-model to obtain at least one target object stacking mask information and target object category information corresponding to the at least one target object stacking mask information respectively; a second extraction unit 11024, configured to extract at least one target object stacking image from the target object region image based on the at least one target object stacking mask information; a third detection unit 11025, configured to input the at least one target object stacking image into the third sub-model respectively to obtain at least one target object individual mask information as target object individual detection information, and obtain target object category information corresponding to the at least one target object individual mask information respectively.
[0151] In some alternative implementation manners, the determination module 1103 includes: a fourth determination unit 11031, configured to determine at least one first target object stacking mask information representing the target object stacking located on the outer layer and at least one second target object stacking mask information representing the target object stacking located on the inner layer based on the at least one target object stacking mask information; a fifth determination unit 11032, configured to determine a first quantity of the target objects located on the outer layer and a second quantity of the target objects located on the inner layer corresponding to each target object category information based on the at least one first target object stacking mask information, the at least one second target object stacking mask information, and the at least one target object individual mask information; a sixth determination unit 11033, configured to determine the current quantity of the target objects corresponding to the same target object category information based on the first quantity and the second quantity.
[0152] In some alternative implementation manners, the first sub-model is trained through the following steps: obtaining a training sample set, where the training samples in the training sample set include sample site images and corresponding annotation region mask information for representing the stacking areas of the target objects in the sample site images; using the sample site images included in the training samples in the training sample set as the input of a preset first initial model, and using the corresponding annotation region mask information of the input sample site images as the expected output of the first initial model, and training to obtain the first sub-model.
[0153] In some alternative implementations, the second sub-model is trained based on the following steps: Obtain a sample target object region image cropped from a sample site image, and corresponding annotation target object stacking mask information and annotation target object category information for characterizing the positions where at least one stack of target objects is placed; Use the sample target object region image as the input of a preset second initial model, and use the corresponding annotation target object stacking mask information and annotation target object category information of the input sample target object region image as the expected output of the second initial model to train the second sub-model.
[0154] In some alternative implementations, the annotation target object stacking mask information includes first annotation target object stacking mask information for characterizing that the target object stack is located on the outer layer and second annotation target object stacking mask information for characterizing that the target object stack is located on the inner layer; Using the corresponding annotation target object stacking mask information and annotation target object category information of the input sample target object region image as the expected output of the second initial model to train the second sub-model includes: Using the first annotation target object stacking mask information, second annotation target object stacking mask information, and annotation target object category information corresponding to the input sample target object region image as the expected output of the second initial model to train the second sub-model.
[0155] In some alternative implementations, the third sub-model is trained based on the following steps: Obtain at least one sample target object stack image cropped from the sample target object region image, and corresponding annotation target object individual mask information and annotation target object category information for characterizing the positions where the individuals of the target objects are located; Use the sample target object stack image as the input of a preset third initial model, and use the corresponding annotation target object individual mask information and annotation target object category information of the input sample target object stack image as the expected output of the third initial model to train the third sub-model.
[0156] The target object detection device provided in the above embodiments of the present disclosure captures an image of the site where the target object is stored, inputs the site image into the target object detection model to obtain target object individual mask information and corresponding target object category information, then determines the current quantity of the target objects corresponding to the same target object category information based on the target object individual mask information, and finally generates a push message according to the quantity. It realizes the real-time automatic monitoring of the quantities of various target objects by using an electronic device, enables the user to timely know the usage situation of the target objects, thereby reducing the dependence on people when monitoring the quantity of the target objects and improving the efficiency of monitoring the quantity of the target objects.
[0157] Exemplary Electronic Device
[0158] Next, with reference to Figure 13 an electronic device according to an embodiment of the present disclosure will be described. The electronic device may be any one or both of the terminal device 101 and the server 103 as shown in Figure 1 FIG. 1, or a stand-alone device independent of them. The stand-alone device may communicate with the terminal device 101 and the server 103 to receive the input signals collected from them.
[0159] Figure 13 FIG. 2 is a block diagram of an electronic device according to an embodiment of the present disclosure.
[0160] As Figure 13 shown in FIG. 3, the electronic device 1300 includes one or more processors 1301 and a memory 1302.
[0161] The processor 1301 may be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 1300 to perform desired functions.
[0162] The memory 1302 may include one or more computer program products. The computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage media. The processor 1301 may run the program instructions to implement the object detection method of the various embodiments of the present disclosure above and / or other desired functions. Various contents such as site images and mask information may also be stored in the computer-readable storage media.
[0163] In one example, the electronic device 1300 may further include: an input device 1303 and an output device 1304. These components are interconnected through a bus system and / or other forms of connection mechanisms (not shown).
[0164] For example, when the electronic device is the terminal device 101 or the server 103, the input device 1303 may be a camera, a mouse, or a keyboard device for inputting images. When the electronic device is a stand-alone device, the input device 1303 may be a communication network connector for receiving the input images from the terminal device 101 and the server 103.
[0165] The output device 1304 can output various types of information to the outside, including the generated push information. The output device 1304 can include, for example, a display, a speaker, a printer, a communication network, and remote output devices connected thereto, and the like.
[0166] Of course, for simplicity, Figure 13 only some of the components related to the present disclosure in the electronic device 1300 are shown, and components such as a bus, an input / output interface, and the like are omitted. In addition, according to specific application scenarios, the electronic device 1300 may further include any other appropriate components.
[0167] Exemplary Computer Program Product and Computer Readable Storage Medium
[0168] In addition to the above methods and devices, an embodiment of the present disclosure may also be a computer program product, which includes computer program instructions that, when executed by a processor, cause the processor to execute the steps in the target object detection method according to various embodiments of the present disclosure described in the "Exemplary Method" section of this specification.
[0169] The computer program product can be written in any combination of one or more programming languages for programming code to perform the operations of the embodiments of the present disclosure. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0170] Furthermore, an embodiment of the present disclosure may also be a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the processor is caused to execute the steps in the target object detection method according to various embodiments of the present disclosure described in the "Exemplary Method" section of this specification.
[0171] The computer-readable storage medium may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may include, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0172] The basic principles of the present disclosure have been described in conjunction with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present disclosure are only examples and not limitations. It cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present disclosure. Additionally, the specific details disclosed above are only for illustrative and facilitating understanding purposes and are not limitations. The above details do not limit the present disclosure to necessarily adopt the above specific details for implementation.
[0173] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the embodiments, reference may be made to each other. For system embodiments, since they basically correspond to method embodiments, the description is relatively simple. For related parts, reference may be made to the partial description of the method embodiments.
[0174] The block diagrams of the devices, apparatuses, equipment, and systems involved in the present disclosure are only illustrative examples and do not intend to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including", "comprising", "having", etc. are open-ended terms meaning "including but not limited to" and can be used interchangeably with each other. The word "or" and "and" used herein refer to the word "and / or" and can be used interchangeably with each other unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to" and can be used interchangeably with each other.
[0175] The methods and apparatuses of the present disclosure may be implemented in many ways. For example, the methods and apparatuses of the present disclosure may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of steps for the methods is for illustration only, and the steps of the methods of the present disclosure are not limited to the specific order described above, unless otherwise specifically stated. In addition, in some embodiments, the present disclosure may also be implemented as a program recorded in a recording medium, and these programs include machine-readable instructions for implementing the methods according to the present disclosure. Therefore, the present disclosure also covers a recording medium storing a program for executing the methods according to the present disclosure.
[0176] It should also be noted that in the apparatuses, devices, and methods of the present disclosure, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of the present disclosure.
[0177] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the aspects shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0178] The above description has been presented for purposes of illustration and description. In addition, this description is not intended to limit the embodiments of the present disclosure to the form disclosed herein. Although several example aspects and embodiments have been discussed above, those skilled in the art will recognize some of their variations, modifications, alterations, additions, and subcombinations.
Claims
1. A method for detecting a target object, characterized in that, Including: Obtain a site image of a site storing a target object; Input the site image into a pre-trained target object detection model to obtain at least one target object individual detection information and target object category information corresponding to the at least one target object individual detection information respectively; Based on the at least one target object individual detection information, determine the current quantity of the target object corresponding to the same target object category information; Generate push information for pushing to a target user terminal based on the quantity; Wherein, the target object detection model includes a first sub-model, a second sub-model and a third sub-model; The step of inputting the site image into the pre-trained target object detection model to obtain at least one target object individual detection information and target object category information corresponding to the at least one target object individual detection information respectively includes: Input the site image into the first sub-model to obtain region mask information representing the stacking area of the target object; Based on the region mask information, extract a target object region image from the site image; Input the target object region image into the second sub-model to obtain at least one target object stacking mask information and target object category information corresponding to the at least one target object stacking mask information respectively; Based on the at least one target object stacking mask information, extract at least one target object stacking image from the target object region image; Input the at least one target object stacking image into the third sub-model respectively to obtain at least one target object individual mask information as target object individual detection information, and obtain target object category information corresponding to the at least one target object individual mask information respectively.
2. The method according to claim 1, characterized in that, Wherein, The step of generating push information for pushing to a target user terminal based on the quantity includes: Determine the consumption amount per unit time of the target object corresponding to the same target object category information; based on the consumption amount per unit time and the current quantity of the corresponding target object, determine the available time of the target object corresponding to the same target object category information; generate push information including the available time.
3. The method according to claim 1 or 2, characterized in that Wherein, The step of generating push information for pushing to a target user terminal based on the quantity includes: Determine whether the current quantity of the target object corresponding to the same target object category information meets a preset replenishment condition; if the preset replenishment condition is met, generate push information for reminding the target user that the target object corresponding to the same target object category information needs to be replenished.
4. The method according to claim 3, wherein Wherein, The target object is production materials; After generating the push information for reminding the target user that the target object corresponding to the same target object category information needs to be replenished if the preset replenishment condition is met, the method further includes: Receive feedback information from the target user terminal; If the feedback information indicates not to replenish the production materials corresponding to the feedback information, determine the remaining construction period of the application link according to the application link of the production materials corresponding to the feedback information and the quantity of the production materials corresponding to the feedback information and push it.
5. The method according to claim 1, wherein Wherein, Determining the current quantity of the target objects corresponding to the same target object category information based on the at least one target object individual detection information includes: Based on the at least one target object stacking pile mask information, determining at least one first target object stacking pile mask information representing the target object stacking piles located on the outer layer and at least one second target object stacking pile mask information representing the target object stacking piles located on the inner layer; Based on the at least one first target object stacking pile mask information, the at least one second target object stacking pile mask information, and the at least one target object individual mask information, determining the first quantity of the target objects located on the outer layer and the second quantity of the target objects located on the inner layer respectively corresponding to each target object category information; Based on the first quantity and the second quantity, determining the current quantity of the target objects corresponding to the same target object category information.
6. The method according to claim 1, wherein Wherein, The first sub-model is trained based on the following steps: Obtaining a training sample set, wherein the training samples in the training sample set include sample site images and corresponding annotation area mask information for characterizing the stacking areas of the target objects in the sample site images; Taking the sample site images included in the training samples in the training sample set as the input of a preset first initial model, and taking the annotation area mask information corresponding to the input sample site images as the expected output of the first initial model, and training to obtain the first sub-model.
7. The method according to claim 6, wherein Wherein, The second sub-model is trained based on the following steps: Obtaining a sample target object region image intercepted from the sample site image and corresponding annotation target object stacking pile mask information and annotation target object category information for characterizing the locations of at least one target object stacking pile; Taking the sample target object region image as the input of a preset second initial model, and taking the annotation target object stacking pile mask information and annotation target object category information corresponding to the input sample target object region image as the expected output of the second initial model, and training to obtain the second sub-model.
8. The method according to claim 7, wherein Wherein, The annotation target object stacking pile mask information includes first annotation target object stacking pile mask information representing that the target object stacking piles are located on the outer layer and second annotation target object stacking pile mask information representing that the target object stacking piles are located on the inner layer; The step of taking the annotation target object stacking pile mask information and annotation target object category information corresponding to the input sample target object region image as the expected output of the second initial model and training to obtain the second sub-model includes: Taking the first annotation target object stacking pile mask information, the second annotation target object stacking pile mask information, and the annotation target object category information corresponding to the input sample target object region image as the expected output of the second initial model, and training to obtain the second sub-model.
9. The method according to claim 7 or 8, characterized in that, Wherein, The third sub-model is trained based on the following steps: Obtain at least one sample target object stacking image intercepted from the sample target object area image, and the annotation target object individual mask information and annotation target object category information corresponding to the at least one sample target object stacking image, which are used to characterize the location of the individuals of the target object. Use the sample target object stacking image as the input of a preset third initial model, and use the annotation target object individual mask information and annotation target object category information corresponding to the input sample target object stacking image as the expected output of the third initial model, and train to obtain the third sub-model.
10. An object detection device, characterized in that, It includes: An acquisition module, configured to acquire a site image of a site storing target objects. A detection module, configured to input the site image into a pre-trained target object detection model to obtain at least one target object individual detection information and the target object category information corresponding to the at least one target object individual detection information. A determination module, configured to determine the current quantity of target objects corresponding to the same target object category information based on the at least one target object individual detection information. A generation module, configured to generate push information for pushing to a target user terminal based on the quantity. Wherein, the target object detection model includes a first sub-model, a second sub-model and a third sub-model. The step of inputting the site image into a pre-trained target object detection model to obtain at least one target object individual detection information and the target object category information corresponding to the at least one target object individual detection information includes: Input the site image into the first sub-model to obtain region mask information characterizing the stacking region of the target object. Based on the region mask information, extract the target object area image from the site image. Input the target object area image into the second sub-model to obtain at least one target object stacking mask information and the target object category information corresponding to the at least one target object stacking mask information. Based on the at least one target object stacking mask information, extract at least one target object stacking image from the target object area image. Input the at least one target object stacking image into the third sub-model respectively to obtain at least one target object individual mask information as the target object individual detection information, and obtain the target object category information corresponding to the at least one target object individual mask information respectively.
11. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, and the computer program is used to execute the method according to any one of claims 1-9 above.
12. An electronic device, characterized in that, The electronic device includes: A processor; A memory for storing executable instructions of the processor; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the method according to any one of claims 1-9 above.
Citation Information
Patent Citations
Real-time detection method and device for different types of entity objects in construction site image
CN108052881A