Target object recognition method and device, and vending equipment

By acquiring images from vending machines and determining parallel detection boxes, and cropping target sub-images to identify product attribute information, the problem of low efficiency and high cost in existing technologies is solved, achieving efficient and accurate product identification.

CN115170791BActive Publication Date: 2026-02-17BOE TECHNOLOGY GROUP CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210885986.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-26
Publication Date
2026-02-17
Estimated Expiration
2042-07-26

AI Technical Summary

Technical Problem

Existing technologies for identifying product attribute information are inefficient, provide a poor user experience, and are costly.

Method used

By using a camera to acquire images in the vending machine, ensuring that the axis of the detection frame is parallel to the axis of the target object, cropping the target sub-image, and using an object recognition model to identify attribute information, the need for an image acquisition component for alignment is avoided.

Benefits of technology

It improves the efficiency and accuracy of attribute information acquisition and identification, simplifies user operations, and reduces costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115170791B_ABST
    Figure CN115170791B_ABST
Patent Text Reader

Abstract

The application discloses a target object recognition method and device and a vending equipment, and relates to the technical field of image processing. The vending equipment can acquire a photographed image and recognize attribute information of a target object based on the photographed image. Since the bar code of the target object does not need to be aligned with the image acquisition component of the automatic settlement equipment for scanning by the automatic settlement equipment to recognize the attribute information of the commodity, on the one hand, the attribute information acquisition efficiency is improved, and on the other hand, the operation of the user is simplified, and the user experience is improved. Moreover, since the axis of the detection frame determined by the vending equipment in the photographed image is parallel to the axis of the target object, the invalid background information in the acquired target sub-image can be ensured to be less. Therefore, the attribute information acquisition efficiency can be further improved, and the attribute information recognition accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing, and particularly relates to a target object recognition method and device and a vending equipment. BACKGROUND

[0002] In the shopping process, a user can align the barcode of a commodity to be purchased with an image acquisition component of an automatic settlement equipment. The automatic settlement equipment can scan the barcode and identify attribute information (such as a price) of the commodity. Then, the automatic settlement equipment can display a payment code based on the identified attribute information, so as to be scanned by a mobile terminal of the user and paid.

[0003] However, the above method for identifying attribute information of a commodity is low in efficiency. SUMMARY

[0004] The present application provides a target object recognition method and device and a vending equipment, and can solve the problem of low efficiency of a method for identifying attribute information of a commodity in the related art. The technical solution is as follows.

[0005] In one aspect, a target object recognition method is provided, and is applied to a vending equipment, wherein the vending equipment comprises a camera; and the method comprises the following steps.

[0006] obtaining a photographed image photographed by the camera;

[0007] determining a detection frame comprising a target object in the photographed image, at least one side of the detection frame being parallel to an axis of the target object;

[0008] obtaining a target sub-image in the detection frame;

[0009] identifying attribute information of the target object based on the target sub-image.

[0010] In another aspect, a target object recognition device is provided, and is configured in a vending equipment, wherein the vending equipment comprises a camera; and the device comprises the following modules.

[0011] a first obtaining module, configured to obtain a photographed image photographed by the camera;

[0012] a determining module, configured to determine a detection frame comprising a target object in the photographed image, at least one side of the detection frame being parallel to an axis of the target object;

[0013] a second obtaining module, configured to obtain a target sub-image in the detection frame;

[0014] an identifying module, configured to identify attribute information of the target object based on the target sub-image.

[0015] In yet another aspect, a vending device is provided, which includes a camera, a processor and a memory having instructions stored therein, the instructions being loaded and executed by the processor to implement the target object recognition method as described in the above aspects.

[0016] In yet another aspect, a computer readable storage medium is provided, which has instructions stored therein, the instructions being loaded and executed by a processor to implement the target object recognition method as described in the above aspects.

[0017] In yet another aspect, a computer program product is provided, which includes computer instructions, the instructions being loaded and executed by a processor to implement the target object recognition method as described in the above aspects.

[0018] The technical solutions provided by the present application have at least the following beneficial effects:

[0019] The present application provides a target object recognition method, device and vending device, which can obtain a photographed image and recognize attribute information of a target object based on the photographed image. Since the bar code of the target object does not need to be aligned with the image acquisition component of the automatic settlement device for scanning by the automatic settlement device to recognize the attribute information of the commodity, on the one hand, the attribute information acquisition efficiency is improved, and on the other hand, the user's operation is simplified, and the user experience is improved.

[0020] Moreover, since the axis of the detection frame determined by the vending device in the photographed image is parallel to the axis of the target object, the invalid background information in the obtained target sub-image can be ensured to be less. Therefore, the attribute information acquisition efficiency can be further improved, and the attribute information recognition accuracy can be improved. BRIEF DESCRIPTION OF DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0022] Figure 1 is a flowchart of a target object recognition method provided by an embodiment of the present application;

[0023] Figure 2 is a flowchart of another target object recognition method provided by an embodiment of the present application;

[0024] Figure 3 is a schematic diagram of obtaining a second image based on a first image provided by an embodiment of the present application;

[0025] Figure 4 FIG. 4 is a schematic diagram of obtaining a fourth image and a fifth image based on a first image according to an embodiment of the present application;

[0026] Figure 5 FIG. 5 is a schematic diagram of obtaining a first image according to an embodiment of the present application;

[0027] Figure 6 FIG. 6 is a schematic diagram of a rectangular detection frame according to the related art;

[0028] Figure 7 FIG. 7 is a schematic diagram of a rectangular detection frame according to an embodiment of the present application;

[0029] Figure 8 FIG. 8 is a schematic diagram of a first coordinate system and a second coordinate system according to an embodiment of the present application;

[0030] Figure 9 FIG. 9 is a structural schematic diagram of a target object recognition device according to an embodiment of the present application;

[0031] Figure 10 FIG. 10 is a structural schematic diagram of another target object recognition device according to an embodiment of the present application;

[0032] Figure 11 FIG. 11 is a structural schematic diagram of a vending device according to an embodiment of the present application. DETAILED DESCRIPTION

[0033] In order to make the purpose, technical scheme and advantages of the present application more clear, the embodiments of the present application will be further described in detail below with reference to the drawings.

[0034] The present application provides a target object recognition method, which is applied to a vending device (for example, a vending device without a staff), and the vending device comprises a camera. Optionally, the vending device can be a refrigerator, a beverage cabinet or the like. Referring to FIG. 1, the method comprises the following steps. Figure 1

[0035] Step 101: obtaining a shooting image shot by the camera.

[0036] In the embodiments of the present application, the vending device comprises a device body and a door body connected to the device body. The vending device can control the camera to shoot an image after detecting that the door body is in an open state. Correspondingly, the vending device can obtain the shooting image shot by the camera.

[0037] ​Alternatively, the device body of the vending device is in a groove shape, and the vending device can further include a shelf arranged in the device body. The shelf is used to place an object (for example, a commodity). During use of the vending device, if the vending device detects that the pressure applied to the shelf decreases, the camera can be controlled to capture an image. Correspondingly, the vending device can obtain the captured image captured by the camera.

[0038] Step 102, determining a detection box including the target object in the captured image.

[0039] The axis of the detection box is parallel to the axis of the target object. The detection box can be an axisymmetric figure, and the axis of the detection box can be an axis extending in the length direction of the detection box. Alternatively, the detection box can be a rectangular detection box, or an oval detection box, etc. If the detection box is a rectangular detection box, the length direction refers to the extension direction of the long side of the detection box, and if the detection box is an oval detection box, the length direction refers to the extension direction of the long side of the minimum circumscribed rectangle of the detection box.

[0040] If the target object is an axisymmetric object, the axis of the target object is a symmetry axis extending in the length direction (also referred to as the height direction) of the target object. If the target object is not an axisymmetric object, the axis of the target object can be a straight line passing through the center point of the target object and extending in the length direction of the target object. The length direction of the target object can be parallel to the extension direction of the long side of the minimum circumscribed rectangle of the target object.

[0041] It can be understood that the axis of the detection box being parallel to the axis of the target object means that the axis of the detection box is substantially parallel to the axis of the target object. That is, when the included angle between the axis of the detection box and the axis of the target object is less than an included angle threshold, it can be considered that the axis of the detection box is parallel to the axis of the target object. The included angle threshold can be pre-stored by the vending device. For example, it can be 5°.

[0042] Step 103, obtaining a target sub-image in the detection box.

[0043] In the embodiments of the present application, the vending device can directly crop the captured image according to the detection box to obtain the target sub-image in the detection box. Alternatively, the detection box is a rectangular detection box. The vending device can first determine the conversion relationship between the first coordinate system and the second coordinate system of the captured image. Then, the vending device can convert the pixels in the detection box to the second coordinate system based on the conversion relationship, thereby obtaining the target sub-image.

[0044] The two coordinate axes of the first coordinate system are parallel to two edges (for example, two edges perpendicular to each other) of the rectangular detection frame, and the two coordinate axes of the second coordinate system are parallel to a pixel row direction and a pixel column direction of the captured image. Two edges of the target sub-image are parallel to the two coordinate axes of the second coordinate system.

[0045] In step 104, attribute information of the target object is identified based on the target sub-image.

[0046] In the embodiment of the present application, the object recognition model is pre-stored in the vending device. The vending device can input the target sub-image into the object recognition model to obtain attribute information of the target object output by the object recognition model.

[0047] To sum up, the embodiment of the present application provides a target object recognition method. The vending device can capture an image and identify attribute information of a target object based on the captured image. Since the bar code of the target object does not need to be aligned with the image acquisition component of the automatic settlement device for scanning by the automatic settlement device to identify attribute information of the commodity, on the one hand, the efficiency of obtaining attribute information is improved, and on the other hand, the operation of the user is simplified, and the user experience is improved.

[0048] Moreover, since the axis of the detection frame determined by the vending device in the captured image is parallel to the axis of the target object, the amount of invalid background information in the target sub-image obtained can be ensured to be small. Therefore, the efficiency of obtaining attribute information can be further improved, and the recognition accuracy of attribute information can be improved.

[0049] The embodiment of the present application takes a rectangular detection frame as an example to exemplarily describe the target object recognition method provided by the embodiment of the present application. The method can be applied to a vending device, and the vending device includes a camera. Referring to Figure 2 The method can include the following steps.

[0050] In step 201, a plurality of sample data are obtained.

[0051] Each sample data includes a sample image of a target object and sample information. The sample information is used to indicate the attribute of the target object. The attribute of the target object can include the price of the target object.

[0052] In the embodiments of the present application, the sample images in the plurality of sample data can include: first images obtained by photographing the target object from different photographing angles. Further, the sample images in the plurality of sample data can also include: a plurality of second images. Each second image can be obtained by adding a third image of an object to one first image. The object can be an object for picking up the target object, for example, can be a human hand. In this way, it can be ensured that a larger number of sample data is obtained, and the diversity of the obtained sample data is higher, so as to ensure that the reliability of the target recognition model obtained by training is higher. Each first image and each second image are sample images in one sample data.

[0053] For the implementation mode in which the sample images in the plurality of sample data include: a plurality of first images and a plurality of second images, the process in which the vending device obtains the sample images in the plurality of sample data can include:

[0054] The vending device obtains a plurality of first images obtained by photographing the target object from different photographing angles. Then, for each first image, the vending device adds a third image of an object in the first image to obtain a second image, thereby obtaining the sample images in the plurality of sample data. The third image can cover part of the sub-image of the target object, and the position of the third image in the first image can be random.

[0055] It can be understood that the sample images in the plurality of sample data can also include: a plurality of fourth images and a plurality of fifth images. The background image of each fourth image is different from the background image of each first image, and the background image of each fifth image is different from the background image of each second image. In this way, it can be further ensured that a larger number of sample data is obtained, and the diversity is higher.

[0056] For example, for each first image, the vending device can also update the background image in the first image to obtain a fourth image. Further, the vending device can add the third image of the object in the fourth image to obtain a fifth image.

[0057] For example, for each first image A, the vending device can add a third image a of an object to the first image A to obtain a second image B. And, referring to Figure 3 , for each first image A, the vending device can add a third image a of an object to the first image A to obtain a second image B. And, referring to Figure 4 , the vending device can perform image segmentation processing on the first image A to extract a sub-image b of the target object from the first image A, and add the sub-image b to a first background image to obtain a fourth image C. The first background image is different from the second background image in the first image A. Further, the vending device can add the third image a of the object to the fourth image C to obtain a fifth image D. From Figure 3It can be seen that the target object is a hand image of a human body.

[0058] Optionally, the plurality of first images can be collected by the image collection device. For example, refer to Figure 5 The target object 01 can be located on the rotating table 02 (also referred to as a gimbal), and the image collection device can be fixedly arranged relative to the rotating table. Then, the rotating table can be rotated, and correspondingly, the front view of the target object within the collection range of the image collection device can change, so that the image collection device can capture the target object from different shooting angles, thereby obtaining a plurality of first objects.

[0059] According to the above description, the vending device can obtain the data enhanced data set by processing the plurality of first images. The data enhanced data set can include: the plurality of first images, the plurality of second images, the plurality of fourth images, and the plurality of fifth images.

[0060] It can be understood that the sample images in the plurality of sample data can also include images collected by the camera during use of the sample vending device. In this way, the reliability of the object recognition model obtained by training can be further ensured, and in turn the accuracy of the attribute information of the target object obtained by recognition can be ensured to be high.

[0061] Step 202, model training is performed on the plurality of sample data to obtain an object recognition model.

[0062] After the vending device obtains the plurality of sample data, the vending device can perform model training on the plurality of sample data to obtain an object recognition model.

[0063] Step 203, a shooting image captured by the camera is obtained.

[0064] In the embodiment of the present application, the vending device includes a device body and a door body connected to the device body. The vending device can control the camera to capture an image after detecting that the door body is in an open state. Correspondingly, the vending device can obtain the shooting image captured by the camera.

[0065] Alternatively, the device body of the vending device is in a groove shape, and the vending device further includes a shelf arranged in the device body. The shelf is used to place objects. During use of the vending device, if the vending device detects that the pressure applied to the shelf decreases, the camera can be controlled to capture an image. Correspondingly, the vending device can obtain the shooting image captured by the camera.

[0066] Step 204, a rectangular detection box including the target object is determined in the shooting image.

[0067] The axis of the rectangular detection frame is parallel to the axis of the target object. An extension direction of the axis of the rectangular detection frame is parallel to an extension direction of a long side of the rectangular detection frame.

[0068] If the target object is an axisymmetric object, the axis of the target object is a symmetry axis extending along a length direction of the target object. If the target object is not an axisymmetric object, the axis of the target object can be a straight line passing through a center point of the target object and extending along a length direction of the target object. The length direction of the target object can be parallel to an extension direction of a long side of a minimum circumscribed rectangle of the target object. Correspondingly, the rectangular detection frame can be the minimum circumscribed rectangle of the target object.

[0069] It can be understood that the axis of the rectangular detection frame being parallel to the axis of the target object means that the axis of the rectangular detection frame is substantially parallel to the axis of the target object. That is, when an included angle between the axis of the rectangular detection frame and the axis of the target object is less than an included angle threshold, it can be considered that the axis of the rectangular detection frame is parallel to the axis of the target object. The included angle threshold can be pre-stored by the vending device. For example, the included angle threshold can be 5°.

[0070] Since the axis of the rectangular detection frame determined by the vending device in the photographed image is parallel to the axis of the target object, it can be ensured that there is less invalid background information in the obtained target sub-image. Therefore, the computing complexity of the vending device can be reduced, so as to improve the acquisition efficiency of the attribute information of the target object, and the influence of the background information on recognition can be effectively reduced, so as to improve the recognition accuracy of the attribute information.

[0071] In the embodiments of the present application, the vending device can input the photographed image into the target detection model to obtain a detection result output by the target detection model. The detection result can be used to indicate the position of the detection frame in the photographed image.

[0072] For example, the detection result can include the position of each of the four vertices of the rectangular detection frame in the photographed image. Alternatively, the detection result can include the position of the center point of the rectangular detection frame in the photographed image, the length of each side of the rectangular detection frame, and the included angle between one side of the rectangular detection frame and the pixel row direction or the pixel column direction of the photographed image.

[0073] The position of each of the four vertices and the center point in the captured image can refer to the coordinates of each point in an image coordinate system in which the captured image is located. The image coordinate system can refer to a coordinate system established with a certain vertex (for example, the top-left vertex) of the captured image as the origin, the pixel row direction of the captured image as the extension direction of the first coordinate axis, and the pixel column direction of the captured image as the extension direction of the second coordinate axis. The first coordinate axis can be one of the horizontal axis and the vertical axis, and the second coordinate axis can be the other of the horizontal axis and the vertical axis. For example, the first coordinate axis is the horizontal axis, and the second coordinate axis is the vertical axis.

[0074] In the embodiment of the present application, the target detection model can be obtained by pre-training the target detection model based on a plurality of reference data by a vending equipment. Each reference data can include a reference image and a sub-image of a target object in the reference image. Optionally, the target detection model can be obtained by training a you only look once (YOLO) network. For example, the target detection model can be obtained by training a YOLO version 5 (YOLOV5) network.

[0075] The YOLO algorithm is a convolutional neural network used for target detection. The convolutional neural network has strong learning ability and efficient feature expression ability, and has great advantages in computer vision tasks such as image segmentation, target detection, object recognition, and target tracking.

[0076] In the example, Figure 6 is a schematic diagram of a detection frame in a reference image in the related art. Figure 7 is a schematic diagram of a rectangular detection frame in a reference image provided by an embodiment of the present application. Compared with Figure 6 and Figure 7 It can be seen that, Figure 6 the axis line c of the detection frame 03 in is not parallel to the axis line d of the target object 00, while Figure 7 the axis line e of the detection frame 04 in is parallel to the axis line d of the target object 00.

[0077] In addition, Figure 6 the area of the background image included in the sub-image in the detection frame 03 is larger than the area of the background image included in the sub-image in the detection frame 04. That is, Figure 7 the invalid background information included in the sub-image in the detection frame 03 is more than the invalid background information included in the sub-image in the detection frame 04. Figure 6 Figure 7

[0078] Step 205, obtaining a target sub-image in the rectangular detection frame.

[0079] ​​In this sub-image, the two sides are parallel to the pixel row direction and pixel column direction of the captured image, respectively. The two sides of the sub-image are perpendicular to each other.

[0080] In this application embodiment, there are multiple ways for the vending machine to acquire the target sub-image within the rectangular detection box. This application embodiment takes the following three optional implementation methods as examples to illustrate the process of the vending machine acquiring the target sub-image within the rectangular detection box.

[0081] In a first alternative implementation, the vending device can directly crop the captured object based on the rectangular detection box to obtain the target sub-image. For example, the vending device can determine the minimum bounding rectangle of the rectangular detection box and crop the captured image with the boundary of the minimum bounding rectangle as the boundary line to obtain the target image, which includes the target sub-image.

[0082] In a second alternative implementation, the vending device can determine a first perspective transformation matrix between the first and second coordinate systems based on the positions of each vertex of the rectangular detection box in the captured image. Then, the vending device can determine the target sub-image from the captured image based on the first perspective transformation matrix and the positions of each vertex in the captured image.

[0083] Among them, see Figure 8 The two coordinate axes of the first coordinate system (X1O1Y1) are parallel to the two perpendicular sides of the rectangular detection box, and the two coordinate axes of the second coordinate system (X1O1Y1) are parallel to the pixel row direction and pixel column direction of the captured image, respectively.

[0084] In this embodiment, the vending device can determine the position of each vertex of the rectangular detection box in the first coordinate system based on the position of each vertex of the rectangular detection box in the captured image and the transformation relationship between the first coordinate system and the image coordinate system where the captured image is located. Then, the vending device can obtain a first perspective transformation matrix based on the positions of each vertex in the first coordinate system and the positions of each vertex in the second coordinate system.

[0085] The positions of each vertex in the second coordinate system can be pre-determined by the vending machine based on the size of the rectangular detection box (i.e., the target sub-image). The size of the rectangular detection box can be determined by the vending machine based on the positions of each vertex of the rectangular detection box in the captured image.

[0086] For example, see Figure 8, the four vertices of the rectangular detection frame 04 are sequentially S1 to S4. The length of one side of the rectangular detection frame is the distance d1 between the vertex S1 and the vertex S2, and the length of the other side of the rectangular detection frame is the distance d2 between the vertex S1 and the vertex S4. The one side and the other side are perpendicular to each other. It can be determined that the positions of the four vertices in the second coordinate system are sequentially (0, 0), (d1, 0), (d1, d2) and (0, d2) determined by the vending equipment.

[0087] It can be understood that the vending equipment can call the function “getPerspectiveTransform()” in the open source computer vision database (OpenCV), and input the positions of the vertices of the target sub-image in the captured image and the positions of the vertices in the second coordinate system as input parameters of the function “getPerspectiveTransform()”, to obtain the first perspective transformation matrix output by the function “getPerspectiveTransform()”.

[0088] And the vending equipment can call the function “warpAffiine()” in the OpenCV, and input the first perspective transformation matrix, the positions of the vertices in the captured image, and the captured image as input parameters of the function “warpAffiine()”, to obtain the target sub-image output by the function “warpAffiine()”.

[0089] It can be understood that after the vending equipment obtains the position of the rectangular detection frame in the captured image, the vending equipment can detect whether each side of the rectangular detection frame is parallel to the pixel row direction of the captured image. If the vending equipment determines that any side of the rectangular detection frame is not parallel to the pixel row direction, the vending equipment can determine the first perspective transformation matrix between the first coordinate system and the second coordinate system based on the positions of the vertices of the rectangular detection frame in the captured image, and then obtain the target sub-image from the captured image based on the first perspective transformation matrix. In this way, the processing resources of the vending equipment can be saved.

[0090] If the vending equipment determines that two sides of the rectangular detection frame are parallel to the pixel row direction of the captured image, the vending equipment can directly extract the target sub-image in the rectangular detection frame.

[0091] In a third optional implementation manner, the vending equipment can obtain an included angle between one side of the rectangular detection frame and the pixel row direction or the pixel column direction of the captured image. For example, the vending equipment can obtain an included angle between one side of the rectangular detection frame parallel to the axis of the target object and the pixel row direction of the captured image.

[0092] Then, the vending device can determine a second perspective transformation matrix between the first coordinate system and the second coordinate system based on the included angle and the position of the center point of the rectangular detection frame in the captured image, and then determine the target sub-image from the captured image based on the second perspective transformation matrix, the size of the rectangular detection frame, and the position of the center point of the rectangular detection frame in the captured image.

[0093] The size of the rectangular detection frame includes the lengths of two mutually perpendicular sides. The two coordinate axes of the first coordinate system are respectively parallel to the two sides of the rectangular detection frame, and the two coordinate axes of the second coordinate system are respectively parallel to the pixel row direction and the pixel column direction of the captured image. The second coordinate system is obtained by rotating the first coordinate system with the coordinate origin of the first coordinate as the center of rotation. That is, the coordinate origin of the first coordinate system is the same as the coordinate origin of the second coordinate system. Optionally, the coordinate origin can be the center point of the rectangular detection frame.

[0094] In the embodiments of the present application, the vending device can first map the captured image into the second coordinate system based on the second perspective transformation matrix. Then, the vending device can call the function "getRectSubPix()" in OpenCV and take the size of the rectangular detection frame, the position of the center point of the rectangular detection frame in the captured image, and the captured image as input parameters of the function "getRectSubPix()" to obtain the target sub-image output by the function.

[0095] For example, the vending device can execute the instruction getRectSubPix(image, size(image.cols / 2, image.rows / 2), Point2f center, Output dst, int patchType = -1) to obtain the target sub-image in the rectangular detection frame. Wherein, image is the captured image, image.cols / 2 is the width of the captured image, image.rows / 2 is the height of the captured image, Point2f center is the position of the center point of the rectangular detection frame in the captured image. Output dst is the output image, and int patchType = -1 represents the depth of the output image, which is -1 by default.

[0096] In the embodiments of the present application, after the vending device obtains the included angle between one side of the rectangular detection frame and the pixel row direction or the pixel column direction of the captured image, the vending device can compare the included angle with a target value. If the vending device determines that the included angle is less than or equal to the target value, the vending device can directly determine the second perspective transformation matrix according to the included angle and the position of the center point of the rectangular detection frame in the captured image. The target value can be pre-stored by the vending device, for example, the target value can be 45 degrees (°).

[0097] If the vending device determines that the included angle is greater than the target value, the vending device can first determine a difference between the included angle and the target value. Then, the vending device can determine the second perspective transformation matrix based on the difference and a position of the center point of the rectangular detection frame in the captured image.

[0098] It can be understood that the rectangular detection frame is a long rectangular detection frame. For a case where an included angle between one side of the rectangular detection frame and a pixel row direction or a pixel column direction of the captured image is greater than a target value, when the vending device calls the function “getRectSubPix()” in OpenCV to obtain the target sub-image, the width and the height of the rectangular detection frame need to be interchanged. In this way, the integrity of the obtained target sub-image can be ensured.

[0099] Optionally, after the vending device obtains the included angle between one side of the rectangular detection frame and the pixel row or the pixel column, the vending device can compare the included angle with an angle threshold. If the vending device determines that the included angle is greater than the angle threshold, the vending device can determine the second perspective transformation matrix between the first coordinate system and the second coordinate system based on the included angle and a position of the center point of the rectangular detection frame in the captured image. In this way, the processing resources of the vending device can be saved. The angle threshold can be pre-stored by the vending device. For example, the angle threshold can be 0°.

[0100] It can be understood that if the vending device obtains, through step 204, the detection result output by the target detection model to include a position of each of the four vertices of the rectangular detection frame in the captured image, the second optional implementation manner is used to obtain the target sub-image from the captured image.

[0101] If the vending device obtains, through step 204, the detection result output by the target detection model to include a position of a center point of the rectangular detection frame in the captured image, a length of each side of the rectangular detection frame, and an included angle between one side of the rectangular detection frame and a pixel row direction or a pixel column direction of the captured image, the third optional implementation manner is used to obtain the target sub-image from the captured image.

[0102] In step 206, the target sub-image is input to the object recognition model to obtain attribute information of the target object output by the object recognition model.

[0103] In the embodiments of the present application, the vending device can input the target sub-image to the object recognition model trained through steps 201 and 202 to obtain attribute information of the target object output by the object recognition model. The attribute information can include price information of the target object and an identifier of the target object. The price information is used to indicate the price of the target object. The identifier can include a code of the target object. For example, the identifier can include the code and the name of the target object.

[0104] It can be understood that the object recognition model can first filter out the background information in the target sub-image, and then recognize the attribute information of the target object based on the target sub-image after filtering out the background information. In this way, the recognition efficiency and accuracy of the attribute information of the target object can be improved.

[0105] Optionally, the object recognition model can perform image segmentation processing on the target sub-image to filter out the background information in the target sub-image.

[0106] In the embodiments of the present application, the attribute information of the target object includes price information of the target object. After the vending device recognizes the attribute information of the target object, the vending device can generate a payment code based on the price information of the target object. The payment code can be scanned by the mobile terminal for payment.

[0107] Optionally, the payment code can be a payment two-dimensional code.

[0108] In related technologies, when shopping, the user needs to go to the cashier to scan the bar code of the product by the cashier using the settlement device, or go to the automatic settlement device to align the bar code of the product with the image acquisition component of the automatic settlement device, so that the settlement device recognizes the attribute information (such as price) of the product. Then, the settlement device can display a payment code based on the recognized attribute information, so that the user can pay by using a mobile terminal. However, this method of recognizing the attribute information of the product is low in efficiency and poor in user experience.

[0109] In order to improve the recognition efficiency of the attribute information of the product, related technologies can set a radio frequency identification (RFID) tag on the product, so that the settlement device can recognize the attribute information of the product based on the RFID tag. However, this method of recognizing the attribute information of the product is high in cost.

[0110] However, by using the method provided in the embodiments of the present application, the vending device can obtain the photographed image and recognize the attribute information of the target object based on the photographed image. Since the bar code of the target object does not need to be aligned with the image acquisition component of the settlement device for the settlement device to scan and recognize the attribute information of the product, on the one hand, the efficiency of obtaining the attribute information is improved, and on the other hand, the operation of the user is simplified, and the user experience is improved. Moreover, since the RFID tag does not need to be set on the target object, the cost of recognizing the attribute information of the target object can be reduced.

[0111] It can be understood that the order of the steps of the target object recognition method provided in the embodiments of the present application can be adjusted appropriately, and the steps can be increased or decreased as appropriate. For example, steps 201 and 202 can be deleted as appropriate. Any person skilled in the art can easily think of changes within the scope of the technology disclosed in the present application, which should be covered within the protection scope of the present application, and thus will not be described again.

[0112] In summary, the embodiments of the present application provide a target object recognition method, and the vending equipment can obtain a photographed image and recognize attribute information of the target object based on the photographed image. Since the bar code of the target object does not need to be aligned with the image acquisition component of the automatic settlement equipment for scanning by the automatic settlement equipment to recognize the attribute information of the commodity, on the one hand, the attribute information acquisition efficiency is improved, and on the other hand, the user operation is simplified, and the user experience is improved.

[0113] In addition, since the axis of the detection box determined by the vending equipment in the photographed image is parallel to the axis of the target object, the invalid background information in the obtained target sub-image can be ensured to be less. Therefore, the attribute information acquisition efficiency can be further improved, and the attribute information recognition accuracy can be improved.

[0114] The embodiments of the present application provide a target object recognition device, which is configured in the vending equipment. The vending equipment includes a camera. Referring to Figure 9 The device 300 can include:

[0115] The first acquisition module 301 is configured to obtain a photographed image photographed by the camera.

[0116] The determination module 302 is configured to determine a detection box including the target object in the photographed image, and an axis of the detection box is parallel to an axis of the target object.

[0117] The second acquisition module 303 is configured to obtain a target sub-image in the detection box.

[0118] The recognition module 304 is configured to recognize attribute information of the target object based on the target sub-image.

[0119] Optionally, the determination module 302 can be configured to:

[0120] input the photographed image into a target detection model to obtain a detection result output by the target detection model, and the detection result is used to indicate a position of the detection box in the photographed image.

[0121] Optionally, the detection box is a rectangular detection box. The second acquisition module 303 can be configured to:

[0122] determine a first perspective transformation matrix between the first coordinate system and a second coordinate system based on the positions of the vertices of the detection frame in the captured image, wherein two coordinate axes of the first coordinate system are parallel to two edges of the detection frame respectively, and two coordinate axes of the second coordinate system are parallel to a pixel row direction and a pixel column direction of the captured image respectively;

[0123] determine the target sub-image from the captured image based on the first perspective transformation matrix and the positions of the vertices of the detection frame in the captured image, wherein two edges of the target sub-image are parallel to the pixel row direction and the pixel column direction of the captured image respectively.

[0124] Optionally, the second obtaining module 303 can be configured to:

[0125] If neither of the edges of the detection frame is parallel to the pixel row direction of the captured image, determine the first perspective transformation matrix between the first coordinate system and the second coordinate system based on the positions of the vertices of the detection frame in the captured image.

[0126] Optionally, the detection frame is a rectangular detection frame. The second obtaining module 303 can be configured to:

[0127] obtain an included angle between one edge of the detection frame and the pixel row direction or the pixel column direction of the captured image;

[0128] determine a second perspective transformation matrix between the first coordinate system and the second coordinate system based on the included angle and the position of the center point of the detection frame in the captured image, wherein two coordinate axes of the first coordinate system are parallel to two edges of the detection frame respectively, and two coordinate axes of the second coordinate system are parallel to the pixel row direction and the pixel column direction of the captured image respectively;

[0129] determine the target sub-image from the captured image based on the second perspective transformation matrix, the size of the detection frame, and the position of the center point of the detection frame in the captured image, wherein two edges of the target sub-image are parallel to the pixel row direction and the pixel column direction of the captured image respectively.

[0130] Optionally, the second obtaining module 303 can be configured to:

[0131] If the included angle is greater than an angle threshold, determine the second perspective transformation matrix between the first coordinate system and the second coordinate system based on the included angle and the position of the center point of the detection frame in the captured image.

[0132] Optionally, the identification module 304 can be configured to: filter out background information in the target sub-image; and identify attribute information of the target object based on the target sub-image after the background information is filtered out.

[0133] Optionally, the identification module 304 can be configured to:

[0134] Input the target sub-image into the object recognition model to obtain attribute information of the target object.

[0135] Figure 10 is a structural schematic diagram of another target object recognition device provided by an embodiment of the present application. Referring to Figure 10 The device 300 further includes:

[0136] The third acquisition module 305 is configured to acquire a plurality of sample data, each sample data including a sample image of a target object and sample information, the sample information being used to indicate an attribute of the target object.

[0137] The training module 306 is configured to perform model training on the plurality of sample data to obtain an object recognition model.

[0138] Optionally, the third acquisition module 305 can be configured to:

[0139] acquire a plurality of first images of the target object captured from different shooting angles;

[0140] For each first image, add a third image of a target object in the first image to obtain a second image, the target object being an object used to take the target object;

[0141] Each first image and each second image is a sample image.

[0142] Optionally, the attribute information of the target object includes price information of the target object. Please continue to refer to Figure 10 The device 300 can further include:

[0143] The generation module 307 is configured to generate a payment code based on the price information of the target object, the payment code being used for scanning by a mobile terminal.

[0144] To sum up, the embodiments of the present application provide a target object recognition device, which can acquire a captured image and can recognize attribute information of a target object based on the captured image. Since the bar code of the target object does not need to be aligned with the image acquisition component of the automatic settlement device for scanning by the automatic settlement device to recognize the attribute information of the goods, on the one hand, the acquisition efficiency of the attribute information is improved, and on the other hand, the operation of the user can be simplified, and the user experience is improved.

[0145] Moreover, since the axis of the detection frame determined by the vending device in the captured image is parallel to the axis of the target object, the invalid background information in the target sub-image acquired can be ensured to be less. Therefore, the acquisition efficiency of the attribute information can be further improved, and the recognition accuracy of the attribute information can be improved.

[0146] It can be understood that the target object recognition apparatus provided by the above embodiments is only exemplified by the division of the above functional modules, and in actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.

[0147] The embodiments of the present application also provide a vending device, which can include the target object recognition apparatus provided by the above embodiments.

[0148] As shown in the above Figure 11 , the vending device can include a camera 401, a processor 402, and a memory 403, the memory 403 stores instructions, the instructions are loaded and executed by the processor 402 to implement the target object recognition method provided by the above method embodiments.

[0149] The embodiments of the present application also provide a computer readable storage medium, which stores instructions, the instructions are loaded and executed by the processor to implement the target object recognition method provided by the above method embodiments, for example Figure 1 or Figure 2 the method shown in the above.

[0150] The embodiments of the present application also provide a computer program product or computer program, which includes computer instructions, the computer instructions are loaded and executed by the processor to implement the target object recognition method provided by the above method embodiments, for example Figure 1 or Figure 2 the method shown in the above.

[0151] It can be understood that the term "at least one" in the present application means one or more, and the meaning of "multiple" is two or more.

[0152] In this paper, "and / or" means that there can be three relationships, for example, A and / or B, which can mean: A exists alone, A and B exist together, and B exists alone. The character " / " generally represents that the associated objects before and after are in an "or" relationship.

[0153] A person of ordinary skill in the art can understand that all or part of the steps of the above embodiments can be completed by hardware, or by program to instruct related hardware, and the program can be stored in a computer readable storage medium, and the above mentioned storage medium can be a read-only memory, a magnetic disk or an optical disk, etc.

[0154] It should be noted that the information (including but not limited to user equipment information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.) and signals involved in the present application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions. For example, the photographed images involved in the present application are obtained under sufficient authorization.

[0155] The above only illustrates the exemplary embodiments of the present application, and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A method of identifying a target object, characterized by, The application is applied to a vending device, the vending device comprising a camera; the method comprising: obtaining a shooting image shot by the camera; determining a detection box comprising a target object in the shooting image, an axis of the detection box being parallel to an axis of the target object, the detection box being a rectangular detection box; obtaining a target sub-image in the detection box; recognizing attribute information of the target object based on the target sub-image; wherein the obtaining of the target sub-image in the detection box comprises: determining a first perspective transformation matrix between a first coordinate system and a second coordinate system based on positions of each vertex of the detection box in the shooting image, wherein two coordinate axes of the first coordinate system are parallel to two edges of the detection box respectively, and two coordinate axes of the second coordinate system are parallel to a pixel row direction and a pixel column direction of the shooting image; determining the target sub-image from the shooting image based on the first perspective transformation matrix and the positions of the vertices in the shooting image, two edges of the target sub-image being parallel to the pixel row direction and the pixel column direction of the shooting image respectively; or obtaining an included angle between an edge of the detection box and the pixel row direction or the pixel column direction of the shooting image; determining a second perspective transformation matrix between the first coordinate system and the second coordinate system based on the included angle and a position of a center point of the detection box in the shooting image; determining the target sub-image from the shooting image based on the second perspective transformation matrix, a size of the detection box, and the position of the center point in the shooting image.

2. The method of claim 1, wherein, The determining of the detection box comprising the target object in the shooting image comprises: inputting the shooting image into a target detection model to obtain a detection result output by the target detection model, the detection result being used to indicate a position of the detection box in the shooting image.

3. The method of claim 1, wherein, The determining of the first perspective transformation matrix between the first coordinate system and the second coordinate system based on the positions of the vertices of the detection box in the shooting image comprises: if neither of the edges of the detection box is parallel to the pixel row direction of the shooting image, determining the first perspective transformation matrix between the first coordinate system and the second coordinate system based on the positions of the vertices of the detection box in the shooting image.

4. The method of claim 1, wherein, The determining of the second perspective transformation matrix between the first coordinate system and the second coordinate system based on the included angle and the position of the center point of the detection box in the shooting image comprises: if the included angle is greater than an angle threshold, determining the second perspective transformation matrix between the first coordinate system and the second coordinate system based on the included angle and the position of the center point of the detection box in the shooting image.

5. The method according to any one of claims 1 to 4, characterized in that, The recognizing of the attribute information of the target object based on the target sub-image comprises: filtering background information in the target sub-image; recognizing the attribute information of the target object based on the target sub-image after the filtering of the background information.

6. The method according to any one of claims 1 to 4, characterized in that, The recognizing of the attribute information of the target object based on the target sub-image comprises: Input the target sub-image into an object recognition model to obtain attribute information of the target object.

7. The method of claim 6, wherein, Before the inputting the target sub-image into the object recognition model to obtain the attribute information of the target object, the method further comprises: obtaining a plurality of sample data, each of the sample data comprising: a sample image of the target object and sample information, the sample information being used to indicate the attribute of the target object; training the object recognition model based on the plurality of sample data.

8. The method of claim 7, wherein, The sample image in the plurality of sample data comprises: obtaining a plurality of first images obtained by shooting the target object from different shooting angles; for each of the first images, adding a third image of a target object in the first image to obtain a second image, the target object being an object used to take the target object; each of the first images and each of the second images is a sample image.

9. The method according to any one of claims 1 to 4, characterized in that, The attribute information of the target object comprises price information of the target object; after the identifying the attribute information of the target object based on the target sub-image, the method further comprises: generating a payment code based on the price information of the target object, the payment code being used for scanning by a mobile terminal.

10. An apparatus for identifying a target object, characterized by comprising: The device is configured in a vending device, and the vending device comprises a camera; the device comprises: a first obtaining module configured to obtain a shooting image shot by the camera; a determining module configured to determine a detection frame comprising a target object in the shooting image, an axis of the detection frame being parallel to an axis of the target object, the detection frame being a rectangular detection frame; a second obtaining module configured to obtain a target sub-image in the detection frame; an identifying module configured to identify attribute information of the target object based on the target sub-image; The second obtaining module is configured to determine a first perspective transformation matrix between a first coordinate system and a second coordinate system based on positions of each vertex of the detection frame in the shooting image, wherein two coordinate axes of the first coordinate system are parallel to two edges of the detection frame, and two coordinate axes of the second coordinate system are parallel to a pixel row direction and a pixel column direction of the shooting image; determine the target sub-image from the shooting image based on the first perspective transformation matrix and the positions of the vertices in the shooting image, two edges of the target sub-image being parallel to the pixel row direction and the pixel column direction of the shooting image; or, obtain an included angle between an edge of the detection frame and the pixel row direction or the pixel column direction of the shooting image; determine a second perspective transformation matrix between the first coordinate system and the second coordinate system based on the included angle and a position of a center point of the detection frame in the shooting image; and determine the target sub-image from the shooting image based on the second perspective transformation matrix, a size of the detection frame, and the position of the center point in the shooting image.

11. A vending apparatus characterized by comprising: The vending device comprises a camera, a processor and a memory, the memory storing instructions which are loaded and executed by the processor to implement the target object identification method according to any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that, The storage medium stores instructions which are loaded and executed by the processor to implement the target object identification method according to any one of claims 1 to 9.

13. A computer program product, characterised in that, The computer program product comprises computer instructions which are loaded and executed by the processor to implement the target object identification method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Target object identification method and device, computer equipment and storage medium

    CN113449606A