Image processing method and device, medium and computer equipment

By comparing the similarity of the target object features in the current image and historical image, and combining timing information, we can determine whether the operation sequence meets the preset order, which solves the problem of poor monitoring of timing operation specifications in the prior art, and improves the accuracy and reliability of monitoring.

CN120071232APending Publication Date: 2025-05-30SHENGDOUSHI SHANGHAI SCI & TECH DEV CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311632334.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-30
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing monitoring methods have poor monitoring effects on operating specifications with timing, and it is difficult to accurately determine whether the operating sequence meets the preset order.

Method used

By acquiring the first feature of the target object in the current image and the second feature of the target object in the historical image, the similarity comparison is performed. If the similarity is greater than the preset threshold, based on the timing information and features, it is determined whether the operation sequence conforms to the preset order.

Benefits of technology

It effectively improves the monitoring accuracy and reliability of operating specifications with timing, and can accurately determine whether the operating sequence meets the preset order.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071232A_ABST
    Figure CN120071232A_ABST
Patent Text Reader

Abstract

The invention discloses an image processing method and device, a medium and computer equipment. The method comprises the following steps: acquiring a current image obtained by performing image acquisition on an operation environment where a target object is located; acquiring a first feature of a target object in the current image; performing similarity comparison on the first feature and a second feature of a target object in a historical image obtained by performing image acquisition on the operation environment; and if the similarity between the first feature and the second feature is greater than a preset similarity threshold, based on the time sequence information of the current image, the time sequence information of the historical image, the first feature and the second feature, judging whether the operation sequence of the target object in the operation environment accords with a preset sequence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer vision technology, and in particular, to an image processing method, apparatus, medium, and computer device. Background Art

[0002] With the continuous development of deep learning technology and machine vision technology, deep learning and machine vision technology are widely used to monitor the standardization of the operation process of target objects. However, the existing monitoring methods have poor monitoring effects on some operation specifications with temporality. Summary of the Invention

[0003] In a first aspect, an embodiment of the present disclosure provides an image processing method, the method including: obtaining a current image obtained by collecting an image of an operation environment where a target object is located; obtaining a first feature of the target object in the current image; comparing the similarity between the first feature and a second feature of the target object in a historical image obtained by collecting an image of the operation environment; if the similarity between the first feature and the second feature is greater than a preset first similarity threshold, determining whether the operation sequence of the target object in the operation environment conforms to a preset sequence based on the temporal information of the current image, the temporal information of the historical image, the first feature, and the second feature.

[0004] In some embodiments, the obtaining the first feature of the target object in the current image includes: detecting the target object in the current image; if the detection is successful, extracting the pixel region where the target object in the current image is located; performing feature extraction on the pixel region to obtain the first feature of the target object in the current image.

[0005] In some embodiments, the if the detection is successful, extracting the pixel region where the target object in the current image is located includes: if the detection is successful and the target object in the current image is not occluded, extracting the pixel region where the target object in the current image is located; if it is detected that the target object in the current image is occluded, filtering the current image.

[0006] In some embodiments, the second feature is a feature in a cache; after comparing the similarity between the first feature and the second feature of the target object in a historical image obtained by collecting an image of the operation environment, the method further includes: replacing the feature in the cache with the first feature; after obtaining the next image of the current image, taking the next image as the current image and returning to the step of obtaining the first feature of the target object in the current image.

[0007] In some embodiments, the method further includes: after obtaining the current image, if no new image is obtained within a preset time period, outputting the first feature and / or the first image, and initializing the features in the cache.

[0008] In some embodiments, the method further includes: if the similarity between the first feature and the second feature is less than or equal to a preset second similarity threshold, outputting the features in the cache and / or the images corresponding to the features in the cache; the second similarity threshold is less than or equal to the first similarity threshold.

[0009] In some embodiments, the second similarity threshold is less than the first similarity threshold; the method further includes: if the similarity between the first feature and the second feature is greater than the second similarity threshold and less than or equal to the first similarity threshold, matching the added objects of the target object in the current image with the added objects of the target object in the historical image; if the matching is successful, based on the timing information of the current image, the timing information of the historical image, the first feature, and the second feature, determining whether the operation sequence of the target object in the operation environment conforms to a preset sequence; if the matching fails, outputting the features in the cache and / or the images corresponding to the features in the cache.

[0010] In some embodiments, the method further includes: if the number of objects with the same category in the added objects of the target object in the current image and the added objects of the target object in the historical image is greater than a preset number, determining that the added objects of the target object in the current image and the added objects of the target object in the historical image are successfully matched.

[0011] In some embodiments, the method further includes: obtaining a quality detection result of the target object in the operation environment based on the output features and / or images.

[0012] In some embodiments, the replacing the features in the cache with the first feature includes: if the number of target objects in the current image is greater than the number of target objects in the historical image, adding at least one initial feature to the cache; the number of added initial features is determined based on the difference in the number of target objects in the current image and the number of target objects in the historical image; matching the first features of each target object in the current image with the cached second features; replacing the features in the cache with the first features that match the features, and replacing the initial features with the first features that do not match the features in the cache successfully.

[0013] In some embodiments, the method further includes: if it is determined that the operation sequence of the target object in the operation environment does not conform to the preset sequence, output a prompt message.

[0014] In a second aspect, an image processing apparatus provided by an embodiment of the present disclosure includes: a first acquisition module, configured to acquire a current image obtained by performing image acquisition on an operation environment where a target object is located; a second acquisition module, configured to acquire a first feature of the target object in the current image; a comparison module, configured to compare the similarity between the first feature and a second feature of the target object in a historical image obtained by performing image acquisition on the operation environment; a determination module, configured to, if the similarity between the first feature and the second feature is greater than a preset first similarity threshold, determine whether the operation sequence of the target object in the operation environment conforms to the preset sequence based on the timing information of the current image, the timing information of the historical image, the first feature, and the second feature.

[0015] In some embodiments, the second acquisition module is specifically configured to: detect the target object in the current image; if the detection is successful, extract the pixel region where the target object in the current image is located; perform feature extraction on the pixel region to obtain the first feature of the target object in the current image.

[0016] In some embodiments, the second acquisition module is specifically configured to: if the detection is successful and the target object in the current image is not occluded, extract the pixel region where the target object in the current image is located; if it is detected that the target object in the current image is occluded, filter the current image.

[0017] In some embodiments, the second feature is a feature in the cache; the apparatus further includes: a replacement module, configured to replace the feature in the cache with the first feature; after acquiring the next image of the current image, use the next image as the current image and return to execute the function of the second acquisition module.

[0018] In some embodiments, the apparatus further includes: an initialization module, configured to, after acquiring the current image, if no new image is acquired within a preset time period, output the first feature and / or the first image, and perform initialization processing on the feature in the cache.

[0019] In some embodiments, the apparatus further includes: a first output module, configured to, if the similarity between the first feature and the second feature is less than or equal to a preset second similarity threshold, output the feature in the cache and / or the image corresponding to the feature in the cache; the second similarity threshold is less than or equal to the first similarity threshold.

[0020] In some embodiments, the second similarity threshold is less than the first similarity threshold; the apparatus further includes: a matching module, configured to match the added object of the target object in the current image and the added object of the target object in the historical image if the similarity between the first feature and the second feature is greater than the second similarity threshold and less than or equal to the first similarity threshold; a second determination module, configured to determine whether the operation sequence of the target object in the operation environment conforms to a preset sequence based on the timing information of the current image, the timing information of the historical image, the first feature, and the second feature if the matching is successful; a third output module, configured to output the feature in the cache and / or the image corresponding to the feature in the cache if the matching fails.

[0021] In some embodiments, the apparatus further includes: a determination module, configured to determine that the matching between the added object of the target object in the current image and the added object of the target object in the historical image is successful if the number of objects with the same category in the added object of the target object in the current image and the added object of the target object in the historical image is greater than a preset number.

[0022] In some embodiments, the apparatus further includes: a quality detection module, configured to obtain a quality detection result of the target object in the operation environment based on the output feature and / or image.

[0023] In some embodiments, the apparatus further includes: a second output module, configured to output a prompt message if it is determined that the operation sequence of the target object in the operation environment does not conform to the preset sequence.

[0024] In a third aspect, an embodiment of the present disclosure provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the method described in any embodiment of the present disclosure is implemented.

[0025] In a fourth aspect, an embodiment of the present disclosure provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the method described in any embodiment of the present disclosure is implemented.

[0026] In an embodiment of the present disclosure, a similarity comparison is made between a first feature of a target object in a current image and a second feature of the target object in a historical image. When the similarity between the first feature and the second feature is greater than a preset first similarity threshold, it can be considered that the target object in the current image and the target object in the historical image are the same object. Therefore, based on the timing information of the current image, the timing information of the historical image, the first feature of the target object in the current image, and the second feature of the target object in the historical image, it can be determined whether the operation sequence of the target object in the operation environment conforms to a preset sequence, thereby effectively monitoring the operation specifications with timing, and improving the accuracy and reliability of the monitoring results.

[0027] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The accompanying drawings herein are incorporated into the specification and constitute a part of this disclosure. These drawings illustrate embodiments consistent with the present disclosure and, together with the specification, are used to explain the technical solutions of the present disclosure.

[0029] Figure 1A is a schematic diagram of an application scenario of an embodiment of the present disclosure.

[0030] Figure 1B is a schematic diagram of an image processing process in the related art.

[0031] Figure 2 is a flowchart of an image processing method according to an embodiment of the present disclosure.

[0032] Figure 3 is a general flowchart of an embodiment of the present disclosure.

[0033] Figure 4 is a flowchart of an image deduplication process according to an embodiment of the present disclosure.

[0034] Figure 5A and Figure 5B are schematic diagrams of different numbers of target objects according to an embodiment of the present disclosure.

[0035] Figure 6 is a schematic diagram of an added object of a target object according to an embodiment of the present disclosure.

[0036] Figure 7 is a block diagram of an image processing apparatus according to an embodiment of the present disclosure.

[0037] Figure 8 is a schematic diagram of a computer device according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0038] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0039] The terms used in the present disclosure are for the purpose of describing particular embodiments only and are not intended to limit the present disclosure. The singular forms "a", "the", and "said" used in the present disclosure and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. Additionally, the term "at least one" as used herein represents any one of a plurality or any combination of at least two of a plurality.

[0040] It should be understood that although the terms first, second, third, etc. may be used in the present disclosure to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of the present disclosure, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "upon" or "in response to determining".

[0041] In order to enable those skilled in the art of the present technology to better understand the technical solutions in the embodiments of the present disclosure and to make the above-mentioned objects, features, and advantages of the embodiments of the present disclosure more apparent and understandable, the technical solutions in the embodiments of the present disclosure will be further described in detail below with reference to the accompanying drawings.

[0042] During the process of operating a target object, it is necessary to monitor the operation process of the target object in a standardized manner. For example, in the catering industry, the target object may be food, the operation process of the target object may be the food production process or the food delivery process, and the standardized monitoring may be to monitor whether the addition order of various ingredients required for the food, the production duration of the food, the production tools, etc. comply with the pre-specified specifications. Some operation specifications often have temporality. Take Figure 1ATaking the application scenario shown as an example, this application scenario includes an operating environment, which includes an operating table D, a steak Z (i.e., the target object) placed on the operating table D, a chef U standing beside the operating table D, and a camera 11 installed on the operating table D. The chef U can make the steak Z on the operating table D according to a certain production sequence. For example, the chef U can first sprinkle seasonings (such as salt, pepper, etc.) on the steak Z, then add side dishes (such as potatoes, lettuce) to the steak Z, and then sprinkle cheese on the side dishes. When conducting normative monitoring, it is necessary to monitor whether the chef U makes the steak Z according to the above production sequence. It can be understood that Figure 1A The application scenario shown is only for illustrative purposes and is not used to limit the present disclosure. For example, in other application scenarios, the target object can be pizza or other ingredients, or other types other than ingredients, and the operating environment may only include the target object and its producer (such as the chef U in the above embodiment), without including the operating table D.

[0043] The image processing method in the related art is as Figure 1B shown. First, a batch of images to be detected are received, then duplicate images are removed through a deduplication algorithm, then the target object required is detected through a detection model, and finally downstream tasks are completed through a subsequent model. When conducting normative monitoring, it is only possible to infer whether an operation has been performed from the result of target detection to judge the normativity of the operation process. Therefore, the monitoring effect for some operation specifications with time sequence is poor.

[0044] Based on this, an embodiment of the present disclosure provides an image processing method. Refer to Figure 2 , the image processing method includes:

[0045] Step S1: Obtain a current image obtained by collecting an image of the operating environment where the target object is located;

[0046] Step S2: Obtain the first feature of the target object in the current image;

[0047] Step S3: Compare the similarity between the first feature and the second feature of the target object in the historical image obtained by collecting an image of the operating environment;

[0048] Step S4: If the similarity between the first feature and the second feature is greater than a preset first similarity threshold, based on the time sequence information of the current image, the time sequence information of the historical image, the first feature, and the second feature, judge whether the operation sequence of the target object in the operating environment conforms to the preset sequence.

[0049] Next, in combination with the application scenario shown in FIG. 1 and taking the target object as the steak Z as an example, the solution of the embodiment of the present disclosure will be illustrated by examples.

[0050] In step S1, the camera 11 can collect images of the operation environment at a certain frame rate to obtain at least one image. Among them, the current image can be an image collected in real time, and the historical image can be an image collected before the current image. For example, assuming that the image collected in real time is the (i + 1)-th image, then the (i + 1)-th image is the current image, and the 1st image to the i-th image are all historical images, where i is a positive integer. The current image and the historical image can be distinguished based on the timing information of the images. Among them, the timing information of the current image is later than the timing information of the historical image. The timing information can be represented by the timestamp of the image.

[0051] The current image may include a complete target object. Since the target object may be occluded, the current image may also not include the target object, or the target object in the current image is incomplete. In addition, the target object in the operation environment may change. For example, between time T1 and time T2, the target object in the operation environment is steak 1, and between time T2 and time T3, the operation on steak 1 is completed, steak 1 is taken away, and a new steak 2 is placed in the operation environment for operation. Thus, between time T2 and time T3, the target object in the operation environment changes to steak 2. The characteristics of different target objects are often different. Therefore, it is possible to determine whether the target object in the operation environment has changed based on the characteristics of the target object.

[0052] In step S2, the first feature of the target object in the current image can be obtained. Specifically, the target object in the current image can be detected to obtain the pixel region where the target object in the current image is located. The target object in the current image can be detected by a pre-trained target detection model. In some embodiments, the target detection model can output the information of the target object in the current image, and based on the information of the target object output by the target detection model, the pixel region where the target object in the current image is located can be determined. The target detection model includes but is not limited to the Fast R-CNN model, the YOLO model, the SSD model, etc. A sample image including the target object can be obtained, and the information of the target object in the sample image can be pre-annotated, and a target detection model can be trained based on the sample image and the annotated information of the target object. Among them, the annotated information of the target object includes the specific position X-Y and size W-H of the target object in the sample image (where X-Y represents the X-axis coordinate and Y-axis coordinate of the whole target object or a component on the target object in the image pixel coordinate system, and W-H represents the width and height of the whole target object or a component on the target object in the image pixel coordinate system). Further, there may be other objects in the operation environment (such as chef U and the production tool of the target object), and the categories of each object in the sample image can be further annotated, so as to facilitate subsequent determination of whether the target object is occluded.

[0053] After obtaining the pixel region where the target object is located in the current image, feature extraction can be performed on this pixel region to obtain the first feature of the target object in the current image. Feature extraction can be implemented through models such as convolutional neural networks and recurrent neural networks. In some embodiments, the pixel region where the target object is located can be cropped from the current image, and feature extraction is performed on the cropped pixel region. In some scenarios, the resolution of the current image is large, but the size of the target object is small, so the pixel region occupied by the target object in the entire image is small, and most of the features in the current image are features unrelated to the target object. If feature extraction is directly performed on the entire current image, the background region other than the target object in the current image may cause certain interference to subsequent feature comparison. In this embodiment, by cropping the pixel region where the target object is located from the current image and only performing feature extraction on the cropped pixel region, the interference of the background region on subsequent feature comparison can be reduced, thereby improving the accuracy of feature comparison.

[0054] Since the target object may be occluded, in some cases, it may not be possible to successfully detect the target object from the current image. Therefore, it is possible to first determine whether the target object in the current image can be detected successfully based on the detection result of the target object. If the detection is successful, the pixel region where the target object is located in the current image is extracted. If the detection fails, the current image can be directly filtered out, thereby reducing unnecessary processing of the current image subsequently and reducing the consumption of processing resources.

[0055] In some embodiments, it is also possible to detect whether the target object in the current image is occluded. If the target object in the current image is detected successfully and the target object in the current image is not occluded, the pixel region where the target object is located in the current image is extracted. If it is detected that the target object in the current image is occluded, the current image is filtered. Specifically, by detecting the current image, the positions of each object included in the current image can be obtained. If there are other objects other than the target object in the current image and the positions of the other objects are within the range of the detection frame of the target object, it is determined that the target object in the current image is occluded. If the current image only includes the target object, or although there are other objects other than the target object in the current image, but the positions of the other objects are outside the range of the detection frame of the target object, it is determined that the target object in the current image is not occluded. Whether the target object is occluded can also be detected through a detection model.

[0056] In this way, a high-quality current image of the target object when it is not occluded can be obtained. When the target object is not occluded, the first features of the target object can be obtained more comprehensively, thereby improving the accuracy and reliability of subsequent feature comparison. In addition, in some business scenarios, it may be necessary to perform some business processes based on the image or features of the target object. For example, display the image of the target object, or perform quality detection on the target object based on its features. Therefore, by obtaining the current image of the target object when it is not occluded, the accuracy and reliability of subsequent business processes can be improved.

[0057] In step S3, the second features of the target object in the historical image obtained by image acquisition of the operating environment can be obtained. Among them, the historical image can be any image collected historically. Optionally, the historical image can be the previous image collected before the current image. For example, assuming that the current image is the i-th image collected by the camera 11, the historical image can be the (i - 1)-th image collected by the camera 11. Since the previous image is the latest image collected before the current image, the features of the target object in the previous image can best reflect the latest features of the target object. Therefore, comparing the similarity between the first features of the target object in the current image and the second features of the target object in the previous image can best reflect the feature change of the target object in the current image. The target object in the historical image can be detected, the pixel region where the target object in the historical image is located can be extracted based on the detection result, and feature extraction can be performed on this pixel region to obtain the second features of the target object in the historical image. The method for obtaining the second features can refer to the method for obtaining the first features, which will not be elaborated in detail here.

[0058] After obtaining the second features, the first features obtained in step S2 can be compared with the second features obtained in this step for similarity. Both the second features in this step and the first features in step S2 can be feature vectors. On this basis, the similarity between the two feature vectors can be represented by the distance between the feature vectors (for example, cosine distance).

[0059] In some embodiments, the second feature can be cached. When comparing the first feature and the second feature, the first feature can be compared with the features in the cache. Specifically, for each feature of the target object obtained in an image, the feature of the target object in this image can be used to replace the existing feature in the cache. For example, the feature in the cache can be empty initially, or a feature vector of all 0s or all 1s, and the above vector is called an initialization feature. When the first image is obtained, the feature of the target object in the first image can be compared with the initialization feature in the cache, and after the similarity comparison, the feature in the cache can be replaced with the feature of the target object in the first image. When the second image is obtained, the feature of the target object in the second image can be compared with the feature in the cache (which is the feature of the target object in the first image at this time), and after the similarity comparison, the feature in the cache can be replaced with the feature of the target object in the second image. And so on. Therefore, after comparing the first feature and the second feature, the feature in the cache can be replaced with the first feature. In this way, after the next image of the current image is obtained, the next image can be used as the new current image, the feature in the cache (which is the first feature at this time) can be used as the new second feature, and return to step S2.

[0060] In some embodiments, there may be a situation where the number of target objects in the historical image is inconsistent with the number of target objects in the current image. Taking the target object as a steak for example, assume that the historical image includes two plates of steaks. At a certain moment, one plate of steak may be taken away, so that the current image only includes one plate of steak. Or, at a certain moment, a new plate of steak may be added, so that the current image includes three plates of steaks.

[0061] If the number of target objects in the current image is greater than the number of target objects in the historical image, at least one initialization feature can be added to the cache. Among them, the number of added initialization features is determined based on the difference in the number of target objects between the current image and the historical image. Assume that the number of target objects in the historical image is N1, and the number of target objects in the current image is N2 (N1 is greater than N2), then the number of added initialization features can be N1 - N2. As Figure 5A shown, the number of steaks in the historical image is 1, and the number of steaks in the current image is 2, then it can be determined that the number of added initialization features is 1. The initialization feature can be an empty feature, or a feature vector of all 0s or all 1s.

[0062] The first features of each target object in the current image can be matched with the cached second features, the features in the cache can be replaced with the first features that match the features, and the initial features can be replaced with the first features that do not successfully match the features in the cache. Suppose there are two target objects in the historical image and three target objects in the current image. Then, each target object in the historical image can be respectively feature-matched with the three target objects in the current image. Continuing to refer to Figure 5A , if the similarity between the features of a certain target object (such as steak 1) in the historical image and the features of a certain target object (such as steak 2) in the current image is greater than a preset first similarity threshold, it is determined that these two features match successfully, and thus it can be determined that steak 1 and steak 2 are the same object. Therefore, the features of steak 1 in the cache can be replaced with the features of steak 2. If the features of a certain target object (such as steak 1) in the historical image do not successfully match the features of any target object in the current image, it can be determined that the target object (such as steak 3) in the current image is a newly added target object, and thus the features of steak 3 can be used to replace the initial features in the cache. After the above processing, the current features in the cache are the features of steak 2 and the features of steak 3. By the above method, the feature replacement error caused by replacing the features of the original target object with the features of the newly added target object during the feature replacement process is avoided.

[0063] If the number of target objects in the current image is less than the number of target objects in the historical image, the first features of each target object in the current image can be matched with the cached second features. If the features in the cache match the first features successfully, the features in the cache can be replaced with the first features that match the features. If the features in the cache do not match the first features successfully, the features in the cache and / or the image corresponding to the features in the cache can be output, and the features can be deleted from the cache.

[0064] Refer to Figure 5B , suppose there are two target objects (such as steak 1 and steak 2) in the historical image and one target object (such as steak 3) in the current image. The features of steak 1 can be matched with the features of steak 3, and the features of steak 2 can be matched with the features of steak 3. If the features of steak 1 match the features of steak 3 successfully while the features of steak 2 do not match the features of steak 3 successfully, it means that steak 2 has been taken away when the current image is captured, and steak 1 and steak 3 are the same plate of steak. Therefore, the features of steak 1 in the cache can be replaced with the features of steak 3, and the features of steak 2 and / or the image of steak 2 in the cache can also be output, and the features of steak 2 can be deleted from the cache.

[0065] In some embodiments, after obtaining the current image, if no new image is obtained within a preset time period (e.g., 30 s), it indicates that the current image is the last image. According to the foregoing processing method, the first feature of the target object in the current image will replace the second feature and be cached. However, since no new image is obtained, the first feature will be stored in the cache all the time and will not be processed. To solve this problem, after obtaining the current image, if no new image is obtained within the preset time period, the first image and / or the first feature can be output, and the features in the cache can be initialized. Specifically, the first image and / or the first feature can be output to the service processing unit for service processing. The service processing includes, but is not limited to, displaying the first image, performing quality detection on the target object based on the first image and / or the first feature, etc.

[0066] In step S4, if the similarity between the first feature and the second feature is greater than a preset first similarity threshold, it indicates that the target object in the current image and the target object in the historical image are the same object. Therefore, it can be determined whether the operation sequence of the target object in the operation environment conforms to the preset sequence, that is, it is determined whether the operation performed on the target object in the current image and the operation performed on the target object in the historical image conform to the preset sequence. If the similarity between the first feature and the second feature is less than or equal to the second similarity threshold, it indicates that the target object in the current image and the target object in the historical image are not the same object, and thus there is no need to perform a timing judgment on the operations of the target object in the current image and the target object in the historical image.

[0067] The process of determining whether the target object in the current image and the target object in the historical image are the same object by comparing the first feature and the second feature can also be referred to as deduplication processing of the image including the target object. In some business scenarios, usually only one image of the same target object and / or the features of the target object in the above-mentioned one image are required for service processing. Therefore, through deduplication processing, one image can be screened out from multiple images including the same target object, which facilitates the service processing unit to perform service processing. The deduplication processing can be implemented by using a pre-trained deduplication model. The deduplication model can be, for example, the Fast ReID model. Sample images including the target object and the label information of the sample images can be obtained, where the label information is used to characterize whether the target objects in multiple sample images are the same object. For example, the label information of the sample images including the same target object can be the same, while the label information of the sample images including different target objects can be different.

[0068] In some embodiments, if the similarity between the first feature and the second feature is less than or equal to the second similarity threshold, it indicates that the target object in the current image is not the same object as the target object in the historical image, that is, the target object in the operating environment is no longer the target object in the historical image. Therefore, the features in the cache and / or the images corresponding to the features in the cache can be output. Specifically, the features in the cache and / or the images corresponding to the features in the cache can be output to the service processing unit for service processing. Specifically, the features in the cache and / or the images corresponding to the features in the cache can be output to the service processing unit for service processing. The service processing includes, but is not limited to, displaying the images corresponding to the features in the cache, performing quality detection on the target object based on the features in the cache and / or the images corresponding to the features in the cache, etc. Further, the service processing unit can record the quality detection result and / or output a prompt message when the quality detection result is lower than the preset quality detection result.

[0069] In some embodiments, the first similarity threshold may be equal to the second similarity threshold. In other embodiments, the first similarity threshold may be greater than the second similarity threshold. When the similarity between the first feature and the second feature is greater than the larger similarity threshold (i.e., the first similarity threshold), it can be determined that the target object in the current image is the same object as the target object in the historical image; when the similarity between the first feature and the second feature is less than the smaller similarity threshold (i.e., the second similarity threshold), it can be determined that the target object in the current image is not the same object as the target object in the historical image. In this way, the accuracy of judging the same target object can be improved.

[0070] When the similarity between the first feature and the second feature is greater than the second similarity threshold and less than or equal to the first similarity threshold, the added object in the current image (hereinafter referred to as the first added object) and the added object in the historical image (hereinafter referred to as the second added object) can be matched to further determine whether the target object in the current image is the same object as the target object in the historical image. Among them, the added object can be an object added to or around the target object. As Figure 6 shown, the target object is a steak, and the added objects can be side dishes, such as pineapple, broccoli, potato, carrot, etc. If the first added object and the second added object match successfully, it can be determined that the target object in the current image is the same object as the target object in the historical image, and thus, based on the timing information of the current image, the timing information of the historical image, the first feature and the second feature, it can be determined whether the operation sequence of the target object in the operating environment conforms to the preset sequence. If the first added object and the second added object do not match successfully, it can be determined that the target object in the current image is not the same object as the target object in the historical image, and thus, the features in the cache and / or the images corresponding to the features in the cache can be output.

[0071] Among them, the added objects can include one or more categories. The categories of the first added object and the second added object can be matched. If the number of objects with the same category among the first added object and the second added object is greater than or equal to a preset number, it is determined that the first added object and the second added object are successfully matched; otherwise, it is determined that the first added object and the second added object are failed to be matched. The preset number can be a fixed value, for example, numerical values such as 1, 2, or 3, or it can be the product of the target number and the preset proportional coefficient, where the target number is the larger of the number of the first added object and the number of the second added object. Suppose the number of the first added object is 3 and the number of the second added object is 4, then the target number is 4. Suppose the preset proportional coefficient is 50%, then the preset number is 2. Therefore, if the number of objects with the same category among the first added object and the second added object is greater than 2, it is determined that the first added object and the second added object are successfully matched.

[0072] Still taking Figure 6 as an example, suppose the added objects in the historical image include pineapples, and the added objects in the current image include pineapples and broccoli, then there is 1 kind of added object with the same category in the current image and the historical image. Suppose the preset number is also 1, then it meets the condition that the number of objects with the same category among the first added object and the second added object is greater than or equal to the preset number, so it can be determined that the first added object and the second added object are successfully matched. By matching the added objects, it can be more accurately judged whether the target objects in different images are the same object, improving the judgment accuracy.

[0073] In other examples, it can also be determined whether the added objects in different images are matched based on the difference in the number of added objects. For example, if the difference in the number between the first added object in the current image and the second added object in the historical image is less than the preset number difference, it is determined that the first added object and the second added object are successfully matched; otherwise, it is determined that the first added object and the second added object are failed to be matched. Further, if there are multiple categories of the first added object and / or multiple categories of the second added object, it can be determined whether the difference in the number between the first added object of each category and the second added object of the corresponding category is less than the preset number difference. If the difference in the number between the first added object of each category and the second added object of the corresponding category is less than the preset number difference, it is determined that the first added object and the second added object are successfully matched. If the difference in the number between the first added object of any category and the second added object of the corresponding category is greater than or equal to the preset number difference, it is determined that the first added object and the second added object are failed to be matched. In practical applications, the first added object and the second added object can also be matched based on other conditions, which will not be listed one by one here.

[0074] The basis for timing judgment includes the timing information of the current image, the timing information of the historical image, the first feature, and the second feature. Among them, the timing information of the current image can be the time when the current image is acquired, which can be determined by the timestamp of the current image. Similarly, the timing information of the historical image can be the time when the historical image is acquired, which can be determined by the timestamp of the historical image. Since performing different operations on the target object will cause different changes in the features of the target object, the first feature can reflect the operations that have been performed on the target object when the current image is acquired, and the second feature can reflect the operations that have been performed on the target object when the historical image is acquired. Therefore, based on the timing information of the current image, the timing information of the historical image, the first feature, and the second feature, the operation sequence of the target object can be reflected.

[0075] For example, the object (referred to as the added object) on the target object that has been added to the current image can be determined based on the first feature. If the added object determined based on the first feature and added to the target object in the current image includes the target added object, it means that in the current image, the operation of adding the target added object to the target object has been performed. Taking the target object as steak Z as an example, the added object can include, but is not limited to, side dishes, sauces, etc. Another example is that the target object in the current image can be classified based on the first feature to obtain the category of the target object, and different categories correspond to different operation stages of the target object. Still taking the target object as steak Z as an example, the category of steak Z can include the pre-cut category and the post-cut category. Among them, the pre-cut category means that steak Z has not been cut into small pieces, and the post-cut category means that steak Z has been cut into small pieces. If the operation stage indicated by the category of the target object is the target operation stage, it means that in the current image, the operation corresponding to the target operation stage has been performed.

[0076] The standard operation sequence of each operation performed on the target object can be determined in advance. For example, for steak Z, its standard operation sequence is to add side dishes first and then add sauces. If the operation sequence of the target object is consistent with the standard operation sequence, it is determined that the operation sequence of the target object in the operation environment conforms to the preset sequence; otherwise, it is determined that the operation sequence of the target object in the operation environment does not conform to the preset sequence.

[0077] When it is determined that the operation sequence for the target object in the operating environment does not conform to the preset sequence, a prompt message can be output to prompt the operator. Among them, the prompt message can include at least one of, but not limited to, sound prompt information and visual prompt information (for example, text prompt information, light prompt information). The prompt message can include the identification information (for example, the name of the operation) of the operation that does not conform to the preset sequence. Suppose the operation of adding side dishes and the operation of adding sauce do not conform to the preset sequence, then the prompt message can be "The operation of adding side dishes was not performed before the operation of adding sauce", where "adding side dishes" and "adding sauce" are both the names of the corresponding operations. For the sake of simplicity, the prompt message can also only include general prompt information such as "operation violation".

[0078] Figure 3 and Figure 4 respectively show the overall flowchart of the embodiments of the present disclosure and the flowchart of the deduplication process. Taking the target object as steak Z as an example, referring to Figure 3 , n (n is a positive integer) images including steaks (referred to as steak images) can be input, and the n steak images are input into the target detection model for detection. Through detection, it can be determined which of the n steak images have the steak occluded and which do not, and the images with the steak occluded are filtered out. For the images where the steak is not occluded, the pixel region where the steak is located (referred to as the steak region) can be extracted according to the detection results, and the steak regions in the n steak images are input into the deduplication model for deduplication processing, so as to filter out the images of the same plate of steak. After filtering, one steak image or the feature of this steak image can be obtained, and this steak image or the feature of this steak image is output to other models for processing. For example, quality detection of the steak can be performed based on this steak image or the feature of this steak image, and at the same time, the image of this steak and the quality detection result can be output. For multiple images of the same plate of steak, normative monitoring can be performed based on the timing information of these multiple images. If an operation that does not conform to the specification is found, a prompt message can be output or a warning can be issued.

[0079] Figure 4The process of deduplication processing for each image is shown. First, the object detection result of the image and the steak area in the image are input into the deduplication model. If the deduplication model determines that the steak in the image belongs to the same steak as the steak in the historical image (i.e., the image is a duplicate of the historical image), then the features of the steak in the image are used to replace the cached features (the image and its features do not need to be output backward), and based on the timing information of the image and the timing information of the historical image, it is determined whether there are any irregular operations. If so, a prompt is issued. If the deduplication model determines that the steak in the image does not belong to the same steak as the steak in the historical image (i.e., the image is not a duplicate of the historical image), then the features of the steak in the image are used to replace the cached features, and the results of the current stage (historical image and / or its features) are output backward. The output image and / or features can be used by the service processing unit for service processing.

[0080] In the case where the installation environment of the camera 11 in the embodiments of the present disclosure is complex or the shooting conditions are not ideal, unnecessary images and duplicate images can be effectively filtered, which provides guarantee for the accuracy of subsequent service processing; moreover, the embodiments of the present disclosure introduce timing logic judgment, so that the model can not only supervise the quality of the target object, but also judge whether some operations are standard during the deduplication process.

[0081] See Figure 7 , the embodiments of the present disclosure also provide an image processing apparatus, and the apparatus includes:

[0082] A first acquisition module 110, configured to acquire a current image obtained by performing image acquisition on an operation environment where a target object is located;

[0083] A second acquisition module 120, configured to acquire first features of the target object in the current image;

[0084] A comparison module 130, configured to compare the similarity between the first features and second features of the target object in a historical image obtained by performing image acquisition on the operation environment;

[0085] A judgment module 140, configured to, if the similarity between the first features and the second features is greater than a preset first similarity threshold, based on the timing information of the current image, the timing information of the historical image, the first features, and the second features, judge whether the operation sequence of the target object in the operation environment conforms to a preset sequence.

[0086] In some embodiments, the functions or modules included in the apparatus provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments, and the specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0087] An embodiment of the present disclosure also provides a computer device, which at least includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the method described in any of the foregoing embodiments is implemented.

[0088] Figure 8 FIG. shows a more specific schematic diagram of the hardware structure of a computing device provided by an embodiment of the present disclosure. The device may include: a processor 21, a memory 22, an input / output interface 23, a communication interface 24, and a bus 25. Among them, the processor 21, the memory 22, the input / output interface 23, and the communication interface 24 are communicatively connected to each other inside the device through the bus 25.

[0089] The processor 21 can be implemented in a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present disclosure. The processor 21 may further include a graphics card, and the graphics card may be an Nvidia titan X graphics card or a 1080Ti graphics card, etc.

[0090] The memory 22 can be implemented in the form of a read-only memory (ROM), a random access memory (RAM), a static storage device, a dynamic storage device, etc. The memory 22 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of the present disclosure through software or firmware, the relevant program codes are stored in the memory 22 and are called and executed by the processor 21.

[0091] The input / output interface 23 is used to connect to an input / output module to implement information input and output. The input / output module can be configured as a component in the device (not shown in the figure) or can be externally connected to the device to provide corresponding functions. Among them, the input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.

[0092] The communication interface 24 is used to connect to a communication module (not shown in the figure) to implement communication interaction between this device and other devices. Among them, the communication module can communicate through a wired method (such as USB, network cable, etc.) or can communicate through a wireless method (such as a mobile network, WIFI, Bluetooth, etc.).

[0093] The bus 25 includes a path for transmitting information between various components of the device, such as the processor 21, the memory 22, the input / output interface 23, and the communication interface 24.

[0094] It should be noted that although the above device only shows the processor 21, the memory 22, the input / output interface 23, the communication interface 24, and the bus 25, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solution of the embodiments of the present disclosure, and does not necessarily include all the components shown in the figure.

[0095] The embodiments of the present disclosure also provide a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the method described in any one of the foregoing embodiments.

[0096] Computer-readable media includes both permanent and non-permanent, removable and non-removable media and can be implemented by any method or technology for storing information. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.

[0097] From the description of the above embodiments, those skilled in the art can clearly understand that the embodiments of the present disclosure can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the embodiments of the present disclosure, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present disclosure.

[0098] The systems, devices, modules or units illustrated in the above embodiments may be specifically implemented by a computer device or entity, or by a product with certain functions. A typical implementation device is a computer, and the specific form of the computer may be a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email transceiver device, a game console, a tablet computer, a wearable device, or a combination of any several of these devices.

[0099] Each embodiment in this disclosure is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, reference can be made to the partial description of the method embodiments. The device embodiments described above are only illustrative. The modules described as separate components may or may not be physically separated. When implementing the solutions of the embodiments of this disclosure, the functions of the various modules may be implemented in one or more software and / or hardware. It is also possible to select some or all of the modules according to actual needs to achieve the purpose of the solutions of this embodiment. A person of ordinary skill in the art can understand and implement it without creative efforts.

[0100] The above are only the specific implementation manners of the embodiments of this disclosure. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the embodiments of this disclosure, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the embodiments of this disclosure.

Claims

1. An image processing method, the method comprises: acquiring a current image obtained by performing image acquisition on an operation environment where a target object is located; acquiring a first feature of the target object in the current image; performing a similarity comparison between the first feature and a second feature of the target object in a historical image obtained by performing image acquisition on the operation environment; if the similarity between the first feature and the second feature is greater than a preset first similarity threshold, based on the timing information of the current image, the timing information of the historical image, the first feature, and the second feature, determining whether the operation sequence of the target object in the operation environment conforms to a preset sequence.

2. The method according to claim 1, wherein acquiring the first feature of the target object in the current image comprises: detecting the target object in the current image; if the detection is successful, extracting the pixel region where the target object in the current image is located; performing feature extraction on the pixel region to obtain the first feature of the target object in the current image.

3. The method according to claim 2, wherein if the detection is successful, extracting the pixel region where the target object in the current image is located comprises: if the detection is successful and the target object in the current image is not occluded, extracting the pixel region where the target object in the current image is located; if it is detected that the target object in the current image is occluded, filtering the current image.

4. The method according to claim 1, wherein the second feature is a feature in a cache; after performing a similarity comparison between the first feature and the second feature of the target object in a historical image obtained by performing image acquisition on the operation environment, the method further comprises: replacing the feature in the cache with the first feature; after acquiring the next image of the current image, taking the next image as the current image, and returning to the step of acquiring the first feature of the target object in the current image.

5. The method according to claim 4, the method further comprises: after acquiring the current image, if no new image is acquired within a preset time period, outputting the first feature and / or the first image, and performing an initialization process on the feature in the cache.

6. The method according to claim 4, the method further comprises: if the similarity between the first feature and the second feature is less than or equal to a preset second similarity threshold, outputting the feature in the cache and / or the image corresponding to the feature in the cache; the second similarity threshold is less than or equal to the first similarity threshold.

7. The method according to claim 6, wherein the second similarity threshold is less than the first similarity threshold; the method further comprises: if the similarity between the first feature and the second feature is greater than the second similarity threshold and less than or equal to the first similarity threshold, matching the added object of the target object in the current image and the added object of the target object in the historical image; If the matching is successful, based on the timing information of the current image, the timing information of the historical image, the first feature, and the second feature, determine whether the operation sequence of the target object in the operation environment conforms to a preset sequence; If the matching fails, output the features in the cache and / or the images corresponding to the features in the cache.

8. The method according to claim 7, the method further comprises: If the number of objects with the same category in the added objects of the target object in the current image and the added objects of the target object in the historical image is greater than or equal to a preset number, determine that the added objects of the target object in the current image and the added objects of the target object in the historical image match successfully.

9. The method according to claim 6, the method further comprises: Obtain a quality detection result of the target object in the operation environment based on the output features and / or images.

10. The method according to claim 4, replacing the features in the cache with the first feature, comprises: If the number of target objects in the current image is greater than the number of target objects in the historical image, add at least one initial feature to the cache; The number of added initial features is determined based on the difference in the number of target objects in the current image and the number of target objects in the historical image; Match the first feature of each target object in the current image with the cached second feature; Replace the features in the cache with the first feature that matches the feature, and replace the initial feature with the first feature that fails to match the feature in the cache successfully.

11. The method according to claim 4, replacing the features in the cache with the first feature, comprises: If the number of target objects in the current image is less than the number of target objects in the historical image, match the first feature of each target object in the current image with the cached second feature; If the feature in the cache matches the first feature successfully, replace the feature in the cache with the first feature that matches the feature; If the feature in the cache fails to match the first feature, output the features in the cache and / or the images corresponding to the features in the cache, and delete the feature from the cache.

12. The method according to claim 1, the method further comprises: If it is determined that the operation sequence of the target object in the operation environment does not conform to the preset sequence, output a prompt message.

13. An image processing device, the device comprises: A first acquisition module, configured to acquire a current image obtained by performing image acquisition on an operation environment where a target object is located; A second acquisition module, configured to acquire a first feature of the target object in the current image; A comparison module, configured to perform a similarity comparison between the first feature and a second feature of the target object in a historical image obtained by performing image acquisition on the operation environment; A judgment module, configured to, if the similarity between the first feature and the second feature is greater than a preset first similarity threshold, judge whether the operation sequence of the target object in the operation environment conforms to a preset sequence based on the timing information of the current image, the timing information of the historical image, the first feature, and the second feature.

14. A computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the method according to any one of claims 1 to 12 is implemented.

15. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the method according to any one of claims 1 to 12 is implemented.