Target inventory method and related apparatus
By performing interval frame extraction and inter-frame motion correction on image data captured by moving cameras, and combining multiple matching strategies, a low-cost and efficient target inventory is achieved, solving the problem of low efficiency and accuracy in the inventory of moving targets in existing technologies. It is particularly widely used in the inventory of agricultural biological assets.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ALIBABA CLOUD COMPUTING CO LTD
- Filing Date
- 2023-06-05
- Publication Date
- 2026-05-22
AI Technical Summary
In existing technologies, still images cannot be used for inventory checks when the target is in motion. Although images captured by fixed cameras can be used, they are costly to deploy and have low efficiency and accuracy in target inventory checks.
By acquiring image data captured by a moving camera, performing interval frame extraction, determining the detection box position of the target object, correcting the detection box position using inter-frame motion vectors, and combining multiple matching strategies for target matching and tracking, automated inventory can be achieved.
It improves the accuracy and efficiency of target inventory, reduces deployment costs, and expands the applicable scenarios, especially in the inventory of agricultural biological assets.
Smart Images

Figure CN117011330B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a target inventory method and related equipment. Background Technology
[0002] This section is intended to provide background or context for the embodiments of the invention set forth in the claims. It should not be construed as an admission that the description herein is prior art.
[0003] Agricultural biological asset inventory is a key service scenario in smart agriculture. Related technologies primarily use still images or images captured by fixed cameras as data input for target inventory. However, still images are unsuitable for inventorying moving objects, and while images captured by fixed cameras are applicable, they require deploying a large number of cameras around the target area, resulting in high deployment costs. How to achieve fast, accurate, and low-cost target inventory is a pressing technical problem that needs to be solved. Summary of the Invention
[0004] The target inventory method and related equipment provided in this invention at least solve the problem of how to use image data captured by a moving camera to perform automated target inventory, reduce the deployment cost of target inventory, and improve the accuracy and efficiency of target inventory.
[0005] To address the aforementioned problems, one aspect of this invention provides a target inventory method, comprising:
[0006] The system acquires image data captured by a moving camera targeting the target area, performs interval frame extraction on the image data, and obtains the target frame image; wherein the target frame image includes two frames before and after the interval node.
[0007] The target frame image is processed for target detection to determine the detection box position of the target object. The detection box position of the target object is used as the mask area. The inter-frame motion of the camera is estimated based on the non-mask area of the target frame image to determine the inter-frame motion vector of the camera. The detection box position of the target object is corrected based on the inter-frame motion vector.
[0008] The target object is matched and tracked based on the position of the detection box after correction, and the target inventory is performed based on the target matching and tracking results to obtain the target inventory value.
[0009] Furthermore, the steps for target detection processing of the target frame image include:
[0010] The target frame image is classified according to the image image classification type;
[0011] The target frame image is divided into at least one image region according to the classification type indicated by the image classification result. The target detection calculation parameters corresponding to the image regions in the target frame image are configured according to the classification type. The target frame image is then processed for target detection based on the target detection calculation parameters to determine the detection box position of the target object.
[0012] Furthermore, after classifying the target frame image according to the image image classification type, the method also includes:
[0013] Determine whether the image classification result indicates an image anomaly;
[0014] If the image classification result indicates that the image is abnormal, the target frame image with the abnormal image is marked as an invalid frame image, and the target frame image is updated.
[0015] Furthermore, the step of performing target matching and tracking on the target object based on the corrected detection box position, and then conducting target inventory based on the target matching and tracking results, includes:
[0016] The target object in the target frame image is matched and tracked according to the matching strategy and the position of the detection box after correction, and the target is counted according to the target matching and tracking results. The matching strategy includes one or more of the following: the cross-union matching strategy of the detection box of the target object, the center distance matching strategy of the detection box of the target object, and the feature vector matching strategy of the target object.
[0017] Furthermore, after performing target detection processing on the target frame image, the detection value of the target object is determined; after the step of performing target inventory based on the target matching and tracking results to obtain the target inventory value, the method also includes:
[0018] The target object for review is determined based on the detection threshold and the detection value of the target object.
[0019] Perform target matching and tracking verification on the target objects to be reviewed, and update the target inventory values based on the verification results.
[0020] Furthermore, the steps of performing interval frame extraction on the image data to obtain the target frame image include:
[0021] The image data is deframed to obtain the corresponding frame image.
[0022] The frame image is divided according to the interval node, and the two frames before and after the interval node are taken as the target frame image.
[0023] Furthermore, the steps for classifying the target frame image according to the image image classification type include:
[0024] Configure the image classification type, and configure the number of input channels of the image status classifier to be consistent with the number of images included in the target frame image, and the number of output channels to be consistent with the number of classification types;
[0025] The target frame image obtained by interval frame extraction is processed into grayscale and then input into the image image state classifier to classify the target frame image according to the image image classification type.
[0026] To address the aforementioned problems, another aspect of the present invention provides a target inventory device, comprising:
[0027] The acquisition module is used to acquire image data captured by a moving camera targeting the target area where the target object is located, and to perform interval frame extraction processing on the image data to obtain the target frame image; wherein, the target frame image includes two frames before and after the interval node;
[0028] The processing module is used to perform target detection processing on the target frame image, determine the detection box position of the target object, use the detection box position of the target object as the mask area, perform inter-frame motion estimation of the camera based on the non-mask area of the target frame image, determine the inter-frame motion vector of the camera, and correct the detection box position of the target object based on the inter-frame motion vector.
[0029] The inventory module is used to perform target matching and tracking on the target object based on the position of the detection box after correction, and to perform target inventory based on the target matching and tracking results to obtain the target inventory value.
[0030] To address the aforementioned problems, another aspect of the present invention provides an electronic device, including: a processor and a memory storing a program, wherein the program includes instructions that, when executed by the processor, cause the processor to perform any of the aforementioned target inventory methods.
[0031] To address the aforementioned problems, another aspect of the present invention provides a non-transitory machine-readable medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute any of the aforementioned target inventory methods.
[0032] The beneficial effects of this invention are as follows: using image data captured by a moving camera as the data input source, the detection box position of the target object is determined through target detection processing, and the detection box position is corrected according to the inter-frame motion vector of the camera. Then, the target object is matched and tracked according to the corrected detection box position, and the target inventory value is determined according to the target matching and tracking result. This realizes automated inventory of target objects using image data captured by a moving camera as the data input source, which improves the efficiency and accuracy of target inventory, reduces the number of cameras required, reduces deployment costs, and expands the applicable scenarios of target inventory.
[0033] Details of one or more embodiments of the present invention are set forth in the following drawings and description, so that other features, objects and advantages of the invention will be more readily understood. Attached Figure Description
[0034] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other embodiments can be obtained based on these drawings without creative effort.
[0035] Figure 1 This is a flowchart illustrating a target inventory method provided in one embodiment of the present invention.
[0036] Figure 2 This is a flowchart illustrating a target inventory method provided in another embodiment of the present invention.
[0037] Figure 3 This is a schematic diagram of the frame of a target inventory device provided in one embodiment of the present invention.
[0038] Figure 4 This is a schematic diagram of the structure of the electronic device in this embodiment. Detailed Implementation
[0039] Embodiments of this embodiment will now be described in more detail with reference to the accompanying drawings. While some embodiments of this embodiment are shown in the drawings, it should be understood that this embodiment can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this embodiment. It should be understood that the accompanying drawings and embodiments are for illustrative purposes only and are not intended to limit the scope of protection of this embodiment.
[0040] Regarding how to achieve automated inventory of target objects, the relevant technologies include: First, using still images to count the number of target objects by detecting the number of target objects in still images. The disadvantage of this is that it is difficult to apply to the inventory of target objects in motion. Second, using image data captured by fixed cameras as the data input source. The disadvantage of this is that a series of cameras need to be set up around the area where the target objects are located, which requires high deployment costs and has low inventory efficiency and accuracy. Third, using image data captured by mobile devices as the data input source. The disadvantage of this is that the superposition of the movement of the mobile devices and the movement of the target objects, as well as the changing scene captured by the images, leads to low accuracy in the target inventory.
[0041] To address the aforementioned problems, embodiments of the present invention provide a target inventory method. For example... Figure 1 As shown, the target inventory method mainly includes:
[0042] Step S101: Acquire image data captured by a moving camera targeting the target area where the target object is located, and perform interval frame extraction processing on the image data to obtain the target frame image; wherein, the target frame image includes two frames before and after the interval node.
[0043] The target object provided in the embodiments of the present invention can be an object in a static state or an object in a moving state, such as biological materials such as pigs, cattle, and sheep in indoor or outdoor farms.
[0044] By performing interval frame extraction on image data captured by a moving camera, a target frame image is obtained, consisting of two frames before and after the interval node. This helps in the subsequent calculation of the inter-frame motion vector between the two frames included in the target frame image. The position of the detection box of the target object is then corrected based on this inter-frame motion vector. In other words, the inter-frame motion vector of the camera is used to compensate for the image shift caused by camera movement, which improves the accuracy of subsequent target matching and tracking, thereby improving the accuracy of target inventory. This enables automated inventory of agricultural biological data by using image data captured by a moving camera as the data input source, expanding the applicable scenarios of target inventory. At the same time, it also reduces the number of cameras required to deploy in related technologies that use image data captured by fixed cameras as the data input source, thus reducing deployment costs.
[0045] For example, according to an embodiment of the present invention, the step of performing interval frame extraction processing on image data to obtain a target frame image includes: performing frame de-framing processing on the image data to obtain a frame image corresponding to the image data; dividing the frame image according to the interval node, and taking the two frame images before and after the interval node as the target frame image in sequence.
[0046] Specifically, according to one embodiment of the present invention, the acquired image data captured by the moving camera is first subjected to full-frame rate frame decoding to obtain all frame images, which are then saved according to the shooting sequence. Then, all frame images are divided according to a preset interval node or interval step value, and the two frames before and after the interval node or the division node are taken as the target frame image. The setting of the interval node or interval step value can be adjusted according to the actual situation. For example, if the camera's movement speed or the target object's movement speed is fast, the interval nodes can be appropriately increased or the interval step value shortened; if the camera's movement speed or the target object's movement speed is slow, the interval nodes can be appropriately decreased or the interval step value extended.
[0047] Step S102: Perform target detection processing on the target frame image, determine the detection box position of the target object, use the detection box position of the target object as the mask area, perform inter-frame motion estimation of the camera based on the non-mask area of the target frame image, determine the inter-frame motion vector of the camera, and correct the detection box position of the target object based on the inter-frame motion vector.
[0048] After performing target detection processing on the target frame image and determining the detection box position of the target object, since the target object itself may be in motion, the detection box position of the target object is used as a mask area, and the non-masked area in the target frame image is used as a reference to perform inter-frame motion estimation of the camera, determine the inter-frame motion vector of the camera, and then correct the detection box position of the target object based on the inter-frame motion vector of the camera. This setup compensates for the errors caused by the camera's own motion due to using image data captured by a moving camera as the data input source, ensuring the accuracy of target inventory based on the target matching and tracking results.
[0049] For example, according to an embodiment of the present invention, the above-mentioned steps for performing target detection processing on the target frame image include: classifying the target frame image according to the image image classification type; dividing the target frame image into at least one image region according to the classification type indicated by the image image classification result; configuring the target detection calculation parameters corresponding to the image regions in the target frame image according to the classification type; and performing target detection processing on the target frame image according to the target detection calculation parameters to determine the detection box position of the target object.
[0050] Because the image data captured by a moving camera includes not only the camera's own motion but also the motion of the target object, changes in the shooting scene and shooting angle, etc., different image types may appear in the image, such as dark images, crowded images, and images where the target object is moving at a high speed. If the same set of target detection calculation parameters is used for these different image types during target frame image processing, it will lead to significant errors in the target detection results, resulting in a large deviation in the final target count value. Therefore, by combining the above settings with image classification types to classify the target frame image, corresponding target detection calculation parameters can be configured for the image classification type corresponding to the current target frame image. When the current target frame image corresponds to multiple classification types, multiple sub-parameters in the target detection calculation parameters can be configured simultaneously. Furthermore, the current target frame image can be further divided into image regions according to classification types to determine the classification type of different image regions in the target frame image. Based on the classification type of different image regions, corresponding target detection calculation parameters can be configured, further improving the accuracy of target detection.
[0051] Optionally, according to an embodiment of the present invention, after the step of classifying the target frame image according to the image image classification type, the above method further includes: determining whether the image image classification result indicates an image image abnormality; if the image image classification result indicates an image image abnormality, marking the target frame image with the image image abnormality as an invalid frame image, and updating the target frame image.
[0052] In actual processing, since the camera is in motion, there may be instances in the captured image data where no target object appears in a certain frame, or the image is a distant view that cannot effectively detect the target object. Such images are classified as image abnormalities. By setting the above, target frames with abnormal images are marked as invalid frames, which reduces the number of target frames to be processed for subsequent target detection, reduces the amount of target detection computation, improves target detection efficiency, and thus improves target inventory efficiency.
[0053] Furthermore, according to an embodiment of the present invention, the step of classifying the target frame image according to the image image classification type includes: configuring the image image classification type, and configuring the number of input channels of the image image state classifier to be consistent with the number of images included in the target frame image, and the number of output channels to be consistent with the number of classification types; performing grayscale processing on the target frame image obtained by interval frame extraction processing, and inputting it into the image image state classifier, so as to classify the target frame image according to the image image classification type.
[0054] Step S103: Perform target matching and tracking on the target object according to the position of the detection box after correction, and perform target inventory based on the target matching and tracking results to obtain the target inventory value.
[0055] In the aforementioned steps, the position of the detection box of the target object is corrected by the inter-frame motion vector of the camera. After compensating for the error caused by the camera's own motion on the target matching and tracking, the target object is directly matched and tracked based on the corrected detection box position. The target inventory is then performed based on the target matching and tracking results, which can quickly and accurately obtain the target inventory value.
[0056] Specifically, according to an embodiment of the present invention, the step of performing target matching and tracking on the target object based on the corrected detection box position, and then performing target inventory based on the target matching and tracking result, includes: performing target matching and tracking on the target object in the target frame image based on the matching strategy and the corrected detection box position, and performing target inventory based on the target matching and tracking result; wherein, the matching strategy includes one or more of the following: the target object's detection box intersection-union ratio matching strategy, the target object's detection box center distance matching strategy, and the target object's feature vector matching strategy.
[0057] By adding the above settings to the traditional target tracking algorithm, multiple matching strategies are added. Specifically, one or more of the following strategies are used to match and track the corrected detection boxes: the intersection-union-ratio (IU) matching strategy, the center distance matching strategy, and the feature vector matching strategy. This effectively improves the accuracy of matching and tracking target objects in motion, thereby improving the accuracy of target inventory.
[0058] For example, according to an embodiment of the present invention, after the target frame image is processed by target detection, the detection value of the target object is determined; after the step of performing target inventory based on the target matching and tracking results to obtain the target inventory value, the method further includes: determining the verification target object based on the detection threshold and the detection value of the target object; performing target matching and tracking verification on the verification target object, and updating the target inventory value based on the verification results.
[0059] In the target detection process, not only is the detection bounding box position of the target object determined, but the detection value of the target object is also determined based on the target detection parameters. The detection value of the target object is compared with the detection threshold. Target objects with detection values lower than the detection threshold are set as verification targets for target matching and tracking verification. The target inventory value is updated based on the verification results. Through the above settings, a multi-level matching method is realized. That is, based on the detection threshold and the detection value of the target object, target objects with low detection values during the target detection process are not directly discarded, but are verified a second time to determine the final target inventory value. Through the above settings, the situation of target object matching loss due to poor image quality is avoided, and the target inventory accuracy is further improved.
[0060] The present invention provides the following method: acquiring image data captured by a moving camera targeting a target area, performing interval frame extraction on the image data to obtain a target frame image; wherein the target frame image includes two frames before and after the interval node; performing target detection processing on the target frame image to determine the detection box position of the target object; using the detection box position of the target object as a mask area, performing inter-frame motion estimation of the camera based on the non-mask area of the target frame image to determine the inter-frame motion vector of the camera; and correcting the detection box position of the target object based on the inter-frame motion vector; and performing target matching and tracking on the target object based on the corrected detection box position. This technical solution involves performing target inventory based on target matching and tracking results to obtain target inventory values. Because the inter-frame motion vectors of the camera correct the detection box position of the target object, the error caused by the camera's own motion is compensated. Based on the corrected detection box position of the target object, target matching and tracking are performed, and then target inventory is conducted based on the target matching and tracking results. This allows for automated inventory of target objects using image data captured by a moving camera as the data input source, significantly improving the efficiency and accuracy of target inventory, reducing the number of cameras required, lowering deployment costs, and expanding the applicable scenarios for target inventory.
[0061] This invention also provides a target inventory method applied to the inventory of agricultural biological data. For example... Figure 2 As shown, the target inventory method mainly includes:
[0062] Step S201: Obtain image data captured by a moving camera targeting the target area where the target object is located; perform frame de-framing on the image data to obtain the frame image corresponding to the image data; divide the frame image according to the interval node, and take the two frames before and after the interval node as the target frame image.
[0063] Specifically, according to an embodiment of the present invention, the acquired image data captured by a moving camera is first subjected to full-frame-rate frame decoding, and all the obtained frame images are then saved in the order of capture. Then, all frame images are divided according to a preset interval node or interval step value, and two frames before and after the interval node or the division node are taken as target frame images. The setting of the interval node or interval step value can be adjusted according to actual conditions. It is understood that, depending on the different settings of the interval node or interval step value, multiple sets of target frame images can ultimately be obtained based on the image data captured by the moving camera, and each set of frame images includes two frames before and after the interval node.
[0064] By performing interval frame extraction on image data captured by a moving camera, a target frame image is obtained, consisting of two frames before and after the interval node. This helps in the subsequent calculation of the inter-frame motion vector between the two frames included in the target frame image. The position of the detection box of the target object is then corrected based on this inter-frame motion vector. In other words, the inter-frame motion vector of the camera is used to compensate for the image shift caused by camera movement, which improves the accuracy of subsequent target matching and tracking, thereby improving the accuracy of target inventory. This enables automated inventory of agricultural biological data by using image data captured by a moving camera as the data input source, expanding the applicable scenarios of target inventory. At the same time, it also reduces the number of cameras required to deploy in related technologies that use image data captured by fixed cameras as the data input source, thus reducing deployment costs.
[0065] Step S202: Configure the image frame classification type, and configure the number of input channels of the image frame state classifier to be consistent with the number of images included in the target frame image, and the number of output channels to be consistent with the number of classification types; perform grayscale processing on the target frame image obtained by interval frame extraction processing, and input it into the image frame state classifier to classify the target frame image according to the image frame classification type.
[0066] Considering that image data captured by a moving camera includes not only the camera's own motion but also the motion of the target object, as well as changes in the shooting scene and shooting angle, the corresponding frame images of the image data include multiple image types. Among them, the image classification types provided in this embodiment mainly include four categories: abnormal image, dark image, crowded target object in the image, and fast-moving target object in the image. It should be noted that the above classification configuration should not be construed as a limitation of this invention, and classification types can be added or reduced according to the actual shooting situation.
[0067] The image state classifier provided in this embodiment of the invention can employ common classifiers, including but not limited to ResNet. Based on the actual situation of this embodiment, the number of input channels of the image state classifier is configured to be consistent with the number of images included in the target frame image; the number of output channels is configured to be consistent with the number of classification types. In a specific implementation of this embodiment, the number of input channels of the classifier can be set to 2, corresponding to the two frames before and after the interval node included in the target frame image; the number of output channels can be configured to 4, corresponding to four different image state classification types.
[0068] For example, this embodiment of the invention also provides an image classification method, which can be used in the step of classifying target frame images in the target inventory method provided in the above embodiments. Specifically, after acquiring image data captured by a moving camera, multiple frame images are obtained through frame de-framing. Then, some frame images are labeled and classified as a validation set, and the remaining frame images are used as a training set. During the labeling and classification process, it is required to ensure that the classification types are independent of each other, that is, multiple classification types can be labeled in a single frame image. For example, in a frame image of a dark scene, there may be a situation where the target object is crowded in the image. The frame images in the validation set and the training set are grayscale processed to remove classification interference caused by color information while retaining the relative information between frames. Then, the frame images in the grayscale-processed training set are input into the image classifier through the input channel for classification training. The classification results are then combined with the frame images in the grayscale-processed validation set and input into the classifier for verification, thereby optimizing the training of the image classifier.
[0069] During the training and optimization of the image classifier, to improve the accuracy and efficiency of subsequent target detection, the loss function can be dynamically adjusted based on the classification score of image anomalies within the image classification types. Specifically, the emphasis on image anomalies in the image classifier can be adjusted by regulating the parameter γ. The formula for the loss function is as follows:
[0070] L = L1 + (1 - P1) γ *(L2+L3+L4)
[0071] Where L represents the overall loss function; L1 represents the loss function corresponding to the image anomaly type; P1 represents the probability value that the current frame image is an image anomaly type; γ is an adjustment parameter used to adjust the emphasis of the image classifier on the image anomaly classification type; L2 represents the loss function corresponding to the dark image type; L3 represents the loss function corresponding to the fast-moving target object type in the image type; and L4 represents the loss function corresponding to the crowded target object type in the image type.
[0072] Furthermore, according to an embodiment of the present invention, after the step of classifying the target frame image according to the image image classification type, the above method further includes:
[0073] Determine whether the image classification result indicates an image anomaly;
[0074] If the image classification result indicates that the image is abnormal, the target frame image with the abnormal image is marked as an invalid frame image, and the target frame image is updated.
[0075] By setting the above parameters, target frames with abnormal image quality are marked as invalid frames, which reduces the number of target frames to be processed for subsequent target detection, reduces the computational load of target detection, improves target detection efficiency, and thus improves target inventory efficiency.
[0076] Step S203: Divide the target frame image into at least one image region according to the classification type indicated by the image classification result, configure the target detection calculation parameters corresponding to the image regions in the target frame image according to the classification type, and perform target detection processing on the target frame image according to the target detection calculation parameters to determine the detection box position of the target object.
[0077] Since the target detection processing involves processing image data captured by a moving camera to obtain target frame images, using the same set of target detection calculation parameters for different image types during the target frame image detection process can lead to significant errors in the target detection results, resulting in substantial deviations in the final target detection values. Therefore, the target frame image is divided into multiple image regions based on classification type, and corresponding target detection calculation parameters are configured for different image regions according to their classification types. Then, target detection processing is performed on the target frame image based on these different target detection calculation parameters to determine the detection box position and detection value of the target object. Through this setting, combining image classification types to classify the target frame image and determine the classification type of each image region, and configuring corresponding target detection calculation parameters based on the classification type of each image region, the accuracy of target detection is effectively improved.
[0078] For example, according to embodiments of the present invention, if the classification type is a dark image, the detection threshold in the target detection calculation parameters can be reduced to improve the sensitivity of target detection; if the classification type is that the target object in the image moves quickly, the detection frame rate in the target detection calculation parameters can be increased to offset the detection loss problem caused by the fast movement of the target object; if the classification type is that the target object in the image is crowded, the detection box intersection-union threshold in the target detection calculation parameters can be increased to avoid the missed detection problem caused by the overlap of target objects. If an image region corresponds to multiple classification types, multiple sub-parameters in the target calculation parameters can also be adjusted simultaneously. The specific type and value of the adjusted sub-parameters can be set according to the actual situation.
[0079] Step S204: Using the detection box position of the target object as the mask area, perform inter-frame motion estimation of the camera based on the non-mask area of the target frame image, determine the inter-frame motion vector of the camera, and correct the detection box position of the target object based on the inter-frame motion vector.
[0080] After performing target detection processing on the target frame image and determining the detection box position of the target object, since the target object itself may be in motion, the detection box position of the target object is used as a mask area, and the non-masked area in the target frame image is used as a reference to perform inter-frame motion estimation of the camera, determine the inter-frame motion vector of the camera, and then correct the detection box position of the target object based on the inter-frame motion vector of the camera. This setup compensates for the errors caused by the camera's own motion due to using image data captured by a moving camera as the data input source, ensuring the accuracy of target inventory based on the target matching and tracking results.
[0081] In the process of estimating inter-frame motion of a camera and determining its inter-frame motion vectors, motion estimation algorithms such as optical flow, pixel recursion, and block matching can be used.
[0082] Step S205: Perform target matching and tracking on the target object in the target frame image according to the matching strategy and the position of the detection box after correction, and perform target inventory based on the target matching and tracking results; wherein, the matching strategy includes one or more of the following: the cross-union matching strategy of the detection box of the target object, the center distance matching strategy of the detection box of the target object, and the feature vector matching strategy of the target object.
[0083] By adding the above settings to the traditional target tracking algorithm, multiple matching strategies are added. Specifically, one or more of the following strategies are used to match and track the corrected detection boxes: the intersection-union-ratio (IU) matching strategy, the center distance matching strategy, and the feature vector matching strategy. This effectively improves the accuracy of matching and tracking target objects in motion, thereby improving the accuracy of target inventory.
[0084] Step S206: Determine the target object to be reviewed based on the detection threshold and the detection value of the target object; perform target matching tracking and verification on the target object to be reviewed, and update the target inventory value based on the verification results.
[0085] In the target detection process, not only is the detection bounding box position of the target object determined, but the detection value of the target object is also determined based on the target detection parameters. The detection value of the target object is compared with the detection threshold. Target objects with detection values lower than the detection threshold are set as verification targets for target matching and tracking verification. The target inventory value is updated based on the verification results. Through the above settings, a multi-level matching method is realized. That is, based on the detection threshold and the detection value of the target object, target objects with low detection values during the target detection process are not directly discarded, but are verified a second time to determine the final target inventory value. Through the above settings, the situation of target object matching loss due to poor image quality is avoided, and the target inventory accuracy is further improved.
[0086] The target inventory method provided in this embodiment of the invention compensates for the error of target matching and tracking caused by the camera's own motion by correcting the detection box position of the target object through the inter-frame motion vector of the camera; it also improves the accuracy of target detection by configuring corresponding target detection calculation parameters based on different classification types of the image; and it performs target matching and tracking based on the corrected detection box position of the target object, and then performs target inventory based on the target matching and tracking results. This method can realize the use of image data captured by a moving camera as the data input source to perform automated inventory of target objects, significantly improving the efficiency and accuracy of target inventory, reducing the number of cameras required, reducing deployment costs, and expanding the applicable scenarios of target inventory.
[0087] Based on the target inventory method provided in the embodiments of the present invention, the embodiments of the present invention also provide a target inventory device, such as... Figure 3 As shown, the target inventory device 300 includes:
[0088] The acquisition module 301 is used to acquire image data captured by a moving camera targeting the target area where the target object is located, and to perform interval frame extraction processing on the image data to obtain the target frame image; wherein, the target frame image includes two frames before and after the interval node.
[0089] The target object provided in the embodiments of the present invention can be an object in a static state or an object in a moving state, such as biological materials such as pigs, cattle, and sheep in indoor or outdoor farms.
[0090] By performing interval frame extraction on image data captured by a moving camera, a target frame image is obtained, consisting of two frames before and after the interval node. This helps in the subsequent calculation of the inter-frame motion vector between the two frames included in the target frame image. The position of the detection box of the target object is then corrected based on this inter-frame motion vector. In other words, the inter-frame motion vector of the camera is used to compensate for the image shift caused by camera movement, which improves the accuracy of subsequent target matching and tracking, thereby improving the accuracy of target inventory. This enables automated inventory of agricultural biological data by using image data captured by a moving camera as the data input source, expanding the applicable scenarios of target inventory. At the same time, it also reduces the number of cameras required to deploy in related technologies that use image data captured by fixed cameras as the data input source, thus reducing deployment costs.
[0091] For example, according to an embodiment of the present invention, the acquisition module 301 is further configured to: perform frame de-framing processing on the image data to obtain the frame image corresponding to the image data; divide the frame image according to the interval node, and take the two frame images before and after the interval node as the target frame image in sequence.
[0092] Specifically, according to one embodiment of the present invention, the acquired image data captured by the moving camera is first subjected to full-frame rate frame decoding to obtain all frame images, which are then saved according to the shooting sequence. Then, all frame images are divided according to a preset interval node or interval step value, and the two frames before and after the interval node or the division node are taken as the target frame image. The setting of the interval node or interval step value can be adjusted according to the actual situation. For example, if the camera's movement speed or the target object's movement speed is fast, the interval nodes can be appropriately increased or the interval step value shortened; if the camera's movement speed or the target object's movement speed is slow, the interval nodes can be appropriately decreased or the interval step value extended.
[0093] The processing module 302 is used to perform target detection processing on the target frame image, determine the detection box position of the target object, use the detection box position of the target object as the mask area, perform inter-frame motion estimation of the camera based on the non-mask area of the target frame image, determine the inter-frame motion vector of the camera, and perform correction processing on the detection box position of the target object based on the inter-frame motion vector.
[0094] After performing target detection processing on the target frame image and determining the detection box position of the target object, since the target object itself may be in motion, the detection box position of the target object is used as a mask area, and the non-masked area in the target frame image is used as a reference to perform inter-frame motion estimation of the camera, determine the inter-frame motion vector of the camera, and then correct the detection box position of the target object based on the inter-frame motion vector of the camera. This setup compensates for the errors caused by the camera's own motion due to using image data captured by a moving camera as the data input source, ensuring the accuracy of target inventory based on the target matching and tracking results.
[0095] For example, according to an embodiment of the present invention, the processing module 302 is further configured to: classify the target frame image according to the image image classification type; divide the target frame image into at least one image region according to the classification type indicated by the image image classification result; configure the target detection calculation parameters corresponding to the image regions in the target frame image according to the classification type; and perform target detection processing on the target frame image according to the target detection calculation parameters to determine the detection box position of the target object.
[0096] Because the image data captured by a moving camera includes not only the camera's own motion but also the movement of the target object, changes in the shooting scene and shooting angle, etc., different image types may appear in the image, such as dark images, crowded images, and images where the target object is moving at a high speed. If the same set of target detection calculation parameters is used for these different image types during target frame image processing, it will lead to significant errors in the target detection results, resulting in a large deviation in the final target count value. Therefore, by combining the above settings with image classification types to classify the target frame image, and further refining the classification type of each image region in the target frame image, and configuring corresponding target detection calculation parameters based on the classification type of each image region, the accuracy of target detection is effectively improved.
[0097] Optionally, according to an embodiment of the present invention, the target inventory device 300 further includes a target frame image update module. After the step of classifying the target frame image according to the image image classification type, the target frame image update module is used to: determine whether the image image classification result indicates an image image abnormality; if the image image classification result indicates an image image abnormality, mark the target frame image with the image image abnormality as an invalid frame image, and update the target frame image.
[0098] In actual processing, since the camera is in motion, there may be instances in the captured image data where no target object appears in a certain frame, or the image is a distant view that cannot effectively detect the target object. Such images are classified as image abnormalities. By setting the above, target frames with abnormal images are marked as invalid frames, which reduces the number of target frames that need to be processed for target detection, reduces the amount of target detection computation, improves target detection efficiency, and thus improves target inventory efficiency.
[0099] Furthermore, according to an embodiment of the present invention, the processing module 302 is further configured to: configure the image frame classification type, and configure the number of input channels of the image frame state classifier to be consistent with the number of images included in the target frame image, and the number of output channels to be consistent with the number of classification types; perform grayscale processing on the target frame image obtained by interval frame extraction processing, and input it into the image frame state classifier, so as to classify the target frame image according to the image frame classification type.
[0100] The inventory module 303 is used to perform target matching and tracking on the target object based on the position of the detection box after correction, and to perform target inventory based on the target matching and tracking results to obtain the target inventory value.
[0101] During the execution of the aforementioned module, the position of the detection box of the target object is corrected by the inter-frame motion vector of the camera. After compensating for the error caused by the camera's own motion on the target matching and tracking, the target object is directly matched and tracked based on the corrected detection box position. The target inventory is then performed based on the target matching and tracking results, which can quickly and accurately obtain the target inventory value.
[0102] Specifically, according to an embodiment of the present invention, the inventory module 303 is further configured to: perform target matching and tracking on the target object in the target frame image according to the matching strategy and the position of the detection box after correction, and perform target inventory according to the target matching and tracking result; wherein, the matching strategy includes one or more of the following: the detection box intersection-union ratio matching strategy of the target object, the detection box center distance matching strategy of the target object, and the feature vector matching strategy of the target object.
[0103] By adding the above settings to the traditional target tracking algorithm, multiple matching strategies are added. Specifically, one or more of the following strategies are used to match and track the corrected detection boxes: the intersection-union-ratio (IU) matching strategy, the center distance matching strategy, and the feature vector matching strategy. This effectively improves the accuracy of matching and tracking target objects in motion, thereby improving the accuracy of target inventory.
[0104] For example, according to an embodiment of the present invention, after the target frame image is processed by target detection, the detection value of the target object is also determined; the target inventory device 300 further includes a target inventory value update module. After the step of performing target inventory based on the target matching and tracking results to obtain the target inventory value, the target inventory value update module is used to: determine the target object to be reviewed based on the detection threshold and the detection value of the target object; perform target matching and tracking verification on the target object to be reviewed, and update the target inventory value based on the verification results.
[0105] In the target detection process, not only is the detection bounding box position of the target object determined, but the detection value of the target object is also determined based on the target detection parameters. The detection value of the target object is compared with the detection threshold. Target objects with detection values lower than the detection threshold are set as verification targets for target matching and tracking verification. The target inventory value is updated based on the verification results. Through the above settings, a multi-level matching method is realized. That is, based on the detection threshold and the detection value of the target object, target objects with low detection values during the target detection process are not directly discarded, but are verified a second time to determine the final target inventory value. Through the above settings, the situation of target object matching loss due to poor image quality is avoided, and the target inventory accuracy is further improved.
[0106] The target inventory device provided in this embodiment of the invention corrects the detection box position of the target object by using the inter-frame motion vector of the camera, thus compensating for the error influence of the camera's own motion on target matching and tracking. Based on the corrected detection box position of the target object, it performs target matching and tracking, and then performs target inventory based on the target matching and tracking results. This enables the use of image data captured by a moving camera as the data input source for automated inventory of target objects, significantly improving the efficiency and accuracy of target inventory, reducing the number of cameras required, lowering deployment costs, and expanding the applicable scenarios for target inventory.
[0107] This invention also provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program executable by the at least one processor, which, when executed by the at least one processor, causes the electronic device to perform the method of this invention.
[0108] The present invention also provides a non-transitory machine-readable medium storing a computer program, wherein the computer program, when executed by a computer's processor, is used to cause the computer to perform the method of the present invention.
[0109] This invention also provides a computer program product, including a computer program, wherein the computer program, when executed by a computer's processor, is used to cause the computer to perform the method of this invention.
[0110] refer to Figure 4 The present invention will now describe a structural block diagram of an electronic device that can serve as a server or client in embodiments of the present invention, which is an example of a hardware device that can be applied to various aspects of the present invention. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0111] like Figure 4 As shown, the electronic device includes a computing unit 401, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 402 or a computer program loaded from a storage unit 408 into a random access memory (RAM) 403. The RAM 403 may also store various programs and data required for the operation of the electronic device. The computing unit 401, ROM 402, and RAM 403 are interconnected via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.
[0112] Multiple components in the electronic device are connected to I / O interface 405, including: input unit 406, output unit 407, storage unit 408, and communication unit 409. Input unit 406 can be any type of device capable of inputting information into the electronic device. Input unit 406 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of the electronic device. Output unit 407 can be any type of device capable of presenting information and may include, but is not limited to, a display, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 408 may include, but is not limited to, disks and optical discs. Communication unit 409 allows the electronic device to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers, and / or chipsets, such as Bluetooth devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.
[0113] The computing unit 401 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, CPUs, graphics processing units (GPUs), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any suitable processors, controllers, microcontrollers, etc. The computing unit 401 performs the various methods and processes described above. For example, in some embodiments, the method embodiments of the present invention can be implemented as a computer program tangibly contained in a machine-readable medium, such as storage unit 408. In some embodiments, part or all of the computer program can be delivered via ROM.
[0114] 402 and / or communication unit 409 are loaded and / or installed on the electronic device. In some embodiments, computing unit 401 may be configured to perform the methods described above by any other suitable means (e.g., by means of firmware).
[0115] Computer programs for implementing the methods of embodiments of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0116] In the context of embodiments of the present invention, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable signal medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0117] It should be noted that the term "comprising" and its variations used in the embodiments of the present invention are open-ended, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The modifications of "one" and "multiple" mentioned in the embodiments of the present invention are illustrative and not restrictive. Those skilled in the art should understand that, unless explicitly indicated otherwise in the context, they should be understood as "one or more".
[0118] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the embodiments of the present invention are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0119] The steps described in the method embodiments provided by this invention can be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of protection of this invention is not limited in this respect.
[0120] The term "embodiment" in this specification refers to a specific feature, structure, or characteristic described in connection with an embodiment that may be included in at least one embodiment of the invention. The appearance of this phrase in various places in the specification does not necessarily imply the same embodiment, nor does it imply independence or alternativeity from other embodiments. The various embodiments in this specification are described in a related manner, with reference to each other for similar or identical parts. In particular, for apparatus, device, and system embodiments, since they are substantially similar to method embodiments, the description is relatively simple, and relevant details are referred to in the description of the method embodiments.
[0121] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of patent protection. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the appended claims.
Claims
1. A target inventory method, comprising: The image data captured by a moving camera targeting the target area is obtained, and the image data is subjected to interval frame extraction to obtain the target frame image; wherein, the target frame image includes two frames before and after the interval node; The target frame image is classified according to the image image classification type; the target frame image is divided into at least one image region according to the classification type indicated by the image image classification result; target detection calculation parameters corresponding to the image regions in the target frame image are configured according to the classification type, and target detection processing is performed on the target frame image according to the target detection calculation parameters to determine the detection box position of the target object; using the detection box position of the target object as the mask area, the inter-frame motion estimation of the camera is performed based on the non-mask area of the target frame image to determine the inter-frame motion vector of the camera, and the detection box position of the target object is corrected according to the inter-frame motion vector; The target object is matched and tracked based on the position of the detection box after correction, and the target inventory is performed based on the target matching and tracking results to obtain the target inventory value.
2. The method according to claim 1, wherein, After the step of classifying the target frame image according to the image image classification type, the method further includes: Determine whether the image classification result indicates an image anomaly; If the image classification result indicates that the image is abnormal, the target frame image with the abnormal image is marked as an invalid frame image, and the target frame image is updated.
3. The method according to claim 1, wherein, The step of performing target matching and tracking on the target object based on the corrected detection box position, and then performing target inventory based on the target matching and tracking results, includes: The matching strategy is based on one or more of the frame center distance matching strategy and the feature vector matching strategy of the target object.
4. The method according to claim 3, wherein, After performing target detection processing on the target frame image, the detection value of the target object is also determined; After the step of performing target inventory based on the target matching and tracking results to obtain the target inventory value, the method further includes: The target object for review is determined based on the detection threshold and the detection value of the target object; The target object to be reviewed is subjected to target matching and tracking verification, and the target inventory value is updated based on the verification results.
5. The method according to claim 1, wherein, The step of performing interval frame extraction processing on the image data to obtain the target frame image includes: The image data is deframed to obtain the frame image corresponding to the image data; The frame image is divided according to the interval node, and the two frame images before and after the interval node are taken as the target frame image.
6. The method according to claim 1, wherein, The step of classifying the target frame image according to the image image classification type includes: Configure the image frame classification type, and configure the number of input channels of the image frame state classifier to be consistent with the number of images included in the target frame image, and the number of output channels to be consistent with the number of classification types; The target frame image obtained by the interval frame extraction process is processed into grayscale and then input into the image image state classifier to classify the target frame image according to the image image classification type.
7. A target inventory device, comprising: The acquisition module is used to acquire image data captured by a moving camera targeting the target area where the target object is located, and to perform interval frame extraction processing on the image data to obtain a target frame image; wherein, the target frame image includes two frames before and after the interval node; The processing module is used to classify the target frame image according to the image image classification type; divide the target frame image into at least one image region according to the classification type indicated by the image image classification result; configure the target detection calculation parameters corresponding to the image region in the target frame image according to the classification type; perform target detection processing on the target frame image according to the target detection calculation parameters to determine the detection box position of the target object; use the detection box position of the target object as a mask area; perform inter-frame motion estimation of the camera according to the non-mask area of the target frame image to determine the inter-frame motion vector of the camera; and perform correction processing on the detection box position of the target object according to the inter-frame motion vector. The inventory module is used to perform target matching and tracking on the target object based on the position of the detection box after correction, and to perform target inventory based on the target matching and tracking results to obtain the target inventory value.
8. An electronic device, comprising: A processor, and a memory storing a program, wherein the program includes instructions that, when executed by the processor, cause the processor to perform the method according to any one of claims 1-6.
9. A non-transitory machine-readable medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-6.