Artificial intelligence-based target recognition method, device, and storage medium

By collecting shelf images through robots and utilizing target detection and recognition models, the problem of inefficient shelf merchandise counting in large supermarkets has been solved, and automated and efficient merchandise identification has been achieved.

CN115840417BActive Publication Date: 2025-09-30ECOVACS COMML ROBOTICS CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202011011330.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-09-23
Publication Date
2025-09-30
Estimated Expiration
2040-09-23

AI Technical Summary

Technical Problem

In large supermarkets, the distribution statistics of shelf merchandise require human participation, which is inefficient and lacks accuracy, and existing technology is difficult to effectively improve.

Method used

The robot collects shelf images while moving along the shelves, and uses the target detection model and recognition model to use the rotation angle information to mark and identify the targets, and extract local images for automatic recognition.

Benefits of technology

The accuracy and efficiency of target recognition are improved, thereby improving the efficiency of product statistics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115840417B_ABST
    Figure CN115840417B_ABST
Patent Text Reader

Abstract

The embodiment of the present application provides a target recognition method, device and storage medium based on artificial intelligence. In the embodiment of the present application, shelf images collected by the robot while moving along the shelf can be obtained. In the target detection stage of the shelf image, the shelf image is input into the target detection model to obtain the spatial information of the detection frame for target annotation of the shelf image. In the target recognition stage, the local image corresponding to the detection frame is extracted from the shelf image based on the spatial information including the rotation angle; and the local image is input into the target recognition model to identify the target object in the local image, thereby realizing automatic recognition of the target object, helping to improve the efficiency of target recognition, and further helping to improve the efficiency of subsequent commodity statistics based on the target recognition results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to an artificial intelligence-based target recognition method, device, and storage medium. Background Art

[0002] In large supermarkets, counting the inventory of goods on shelves usually requires human participation. Due to the large area of ​​supermarkets and the coverage of thousands of types of goods, counting the categories and quantities of goods is time-consuming, labor-intensive and inefficient. Summary of the Invention

[0003] Various aspects of the present application provide an artificial intelligence-based target recognition method, device, and storage medium to improve the accuracy of target recognition.

[0004] An embodiment of the present application provides an artificial intelligence-based target recognition method, including: obtaining a shelf image collected by a robot while moving along a shelf; inputting the shelf image into a target detection model to obtain spatial information of a first detection frame for target annotation of the shelf image; extracting a local image corresponding to the first detection frame from the shelf image based on the spatial information of the first detection frame; and inputting the local image into a target recognition model to identify a target object contained in the local image.

[0005] The embodiment of the present application further provides a robot, comprising: a mechanical body; a camera, a memory, and a processor mounted on the mechanical body; the memory is used to store a computer program;

[0006] The camera is used to capture shelf images when the robot moves along the shelf;

[0007] The processor is coupled to the memory and is configured to execute the computer program for: inputting the shelf image into a target detection model to obtain spatial information of a first detection frame for target labeling the shelf image; extracting a partial image corresponding to the first detection frame from the shelf image based on the spatial information of the first detection frame; and inputting the rotation angle into a target recognition model to identify a target object contained in the partial image.

[0008] The embodiment of the present application further provides a computer device, comprising: a memory and a processor; the memory is used to store a computer program;

[0009] The processor is coupled to the memory and is configured to execute the computer program to perform the steps in the above-mentioned artificial intelligence-based target recognition method.

[0010] An embodiment of the present application also provides a computer-readable storage medium storing computer instructions. When the computer instructions are executed by one or more processors, the one or more processors are caused to execute the steps in the above-mentioned artificial intelligence-based target recognition method.

[0011] In an embodiment of the present application, shelf images captured by a robot as it moves along a shelf can be obtained. During the target detection phase of the shelf image, the shelf image is input into a target detection model to obtain spatial information of a detection frame for target annotation of the shelf image. During the target recognition phase, a partial image corresponding to the detection frame is extracted from the shelf image based on the spatial information including the rotation angle. This partial image is then input into a target recognition model to identify the target object in the partial image. This achieves automatic recognition of the target object, helps improve target recognition efficiency, and further helps improve the efficiency of subsequent commodity statistics based on the target recognition results. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0013] Figure 1a A hardware structure diagram of the robot provided in the embodiment of the present application;

[0014] Figure 1b A schematic diagram of a scenario in which a robot collects shelf images according to an embodiment of the present application;

[0015] Figure 1c A schematic diagram of a shelf image captured by a robot according to an embodiment of the present application;

[0016] Figure 1d A schematic diagram of target detection effect provided in an embodiment of the present application;

[0017] Figure 1e A schematic diagram of the structure of a multi-angle image acquisition system provided in an embodiment of the present application;

[0018] Figure 2a and Figure 2b A flowchart of an artificial intelligence-based target recognition method provided in an embodiment of the present application;

[0019] Figure 3 A schematic diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0020] To make the purpose, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the specific embodiments of this application and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0021] To address the technical issue of low target recognition accuracy in existing products, in some embodiments of the present application, shelf images captured by a robot as it moves along a shelf can be obtained. During the target detection phase of the shelf image, the shelf image is input into a target detection model to obtain spatial information of a detection frame for target annotation of the shelf image. During the target recognition phase, a local image corresponding to the detection frame is extracted from the shelf image based on the spatial information including the rotation angle. This local image is then input into a target recognition model to identify the target object in the local image. This achieves automatic recognition of the target object, helps improve target recognition efficiency, and further helps improve the efficiency of subsequent commodity statistics based on the target recognition results.

[0022] The following describes in detail the technical solutions provided by various embodiments of the present application in conjunction with the accompanying drawings.

[0023] It should be noted that the same reference numerals denote the same objects in the following drawings and embodiments, and therefore, once an object is defined in one drawing or embodiment, it does not need to be further discussed in the subsequent drawings and embodiments.

[0024] Figure 1a This is a hardware structure diagram of a robot provided by an exemplary embodiment of the present application. Figure 1a As shown, the robot 100 includes a mechanical body 101 , on which a processor 102 and a memory 103 for storing computer instructions are disposed. In addition, the mechanical body 101 is also provided with a camera 104 .

[0025] It is worth noting that the number of processors 102 and memories 103 can be one or more. "More" means two or more. In this embodiment, the processor 102 and memory 103 can be disposed inside the mechanical body 101 or on the surface of the mechanical body 101. The camera 104 is disposed on the surface of the mechanical body.

[0026] The mechanical body 101 is the actuator of the robot 100, which can execute the operation specified by the processor 102 in a certain environment. The mechanical body 101 reflects the appearance of the robot 100 to a certain extent. In this embodiment, the appearance of the robot 100 is not limited. For example, the robot 100 can be Figure 1b The robot 100 is a humanoid robot, and the mechanical body 101 may include but is not limited to: the robot's head, hands, wrists, arms, waist, base and other mechanical structures. In addition, the robot 100 may also be a non-humanoid robot, and the mechanical body 101 mainly refers to the body of the robot 100.

[0027] It is worth noting that the mechanical body 101 also houses some of the basic components of the robot 100, such as a drive assembly, an odometer, a power supply assembly, an audio assembly, and the like. Optionally, the drive assembly may include drive wheels, a drive motor, universal wheels, and the like. The basic components and their composition may vary between robots 100 , and the examples listed in this application are merely some examples.

[0028] The memory 103 is primarily used to store one or more computer instructions that are executable by the processor 102, causing the processor 102 to control the robot 100 to implement corresponding functions, perform corresponding actions, or complete tasks. In addition to storing computer instructions, the memory 103 can also be configured to store various other data to support operations on the robot 100. Examples of such data include instructions for any application or method operating on the robot 100, an environmental map corresponding to the environment in which the robot 100 is located, and the like. The environmental map can be one or more pre-stored maps corresponding to the entire environment, or it can be a partial map that is currently being constructed.

[0029] The memory 103 may be implemented by any type of volatile or non-volatile memory device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0030] Processor 102 can be considered the control system of robot 100 and is configured to execute computer instructions stored in memory 103 to control robot 100 to perform corresponding functions, actions, or tasks. In this embodiment, robot 100 can move autonomously and complete certain tasks based on autonomous movement. For example, in shopping environments such as supermarkets and shopping malls, robot 100 can perform inventory on shelves. In another example, in some warehouse sorting scenarios, sorting robots can sort goods.

[0031] In this embodiment, whether the robot 100 is counting goods on the shelf or sorting goods, it is not necessary to detect and identify the goods. In this embodiment, in order to realize the detection and identification of goods by the robot 100, Figure 1bAs shown, processor 102 can control robot 100 to move along the shelves and, during the movement of robot 100 along the shelves, control camera 104 to capture shelf images. Camera 104 captures real-time shelf images of the location where robot 100 moves along the shelves. In this embodiment, the shelf images include images of products displayed on the shelves.

[0032] Furthermore, considering that the quality of shelf images will affect the subsequent product detection and recognition effects, the existing technology usually uses handheld devices or fixed cameras to collect shelf images. The shooting angle of handheld devices is not fixed, and the collected images are prone to tilt. Fixed cameras are generally fixed at high positions with a large viewing angle but are far away from the products, and cannot obtain high-definition images, which will reduce the accuracy of product recognition.

[0033] To address the aforementioned issues, the processor 102 in the robot 100 can control the camera 104 to maintain a stable relative position between its capture angle of view and the merchandise on the shelf. Alternatively, the processor 102 can control the camera 104 to focus directly on the merchandise on the shelf. Furthermore, the processor 102 can control the robot 100 to move parallel to the shelf and control the camera 104 to capture images of the shelf during movement. Keeping the robot 100's movement parallel to the shelf helps ensure that the camera 104's capture angle remains stable, improving the quality of the shelf images captured by the camera 104 and, in turn, improving the accuracy of subsequent detection and recognition of merchandise in the shelf images.

[0034] Optionally, the processor 102 can determine the position distribution of the shelves in the environmental map based on a known environmental map; and based on the position distribution of the shelves in the environmental map, plan a moving path parallel to the shelves for the robot 100, and control the robot 100 to move along the moving path parallel to the shelves.

[0035] Furthermore, the robot 100 can automatically adjust the distance between it and the shelf to ensure that the camera 104 can capture images of the entire shelf. Optionally, the robot 100 can automatically adjust the height of the camera 104 and the distance between the robot 100 and the shelf to ensure that the camera 104 can capture images of the entire shelf.

[0036] Accordingly, the processor 102 can also obtain the shelf image captured by the camera 104 and input the shelf image into the object detection model to obtain the spatial information of the detection frame used to annotate the shelf image. In practical applications, when performing object detection on an image, a rectangular detection frame is typically used to annotate the target object contained in the image. The spatial information of the rectangular detection frame specifically refers to the spatial information of the rectangular detection frame on the image to be processed, which can reflect the spatial distribution of the target object in the image. The spatial information of the rectangular detection frame includes: the center position and size of the detection frame. Optionally, the center position of the detection frame can be represented by the center coordinates of the rectangular detection frame, and the size of the detection frame can be represented by the width and height of the rectangular detection frame. Accordingly, the spatial information of the detection frame can be expressed as (x, y, w, h). Where (x, y) represents the center coordinates of the rectangular detection frame, i.e., the coordinates of the center of the rectangular detection frame in the image to be detected, and w and h represent the width and height of the rectangular detection frame, respectively. Alternatively, the spatial information of the rectangular detection frame includes: the coordinates of the vertices of the rectangular detection frame. The center coordinates and vertex coordinates of the rectangular detection frame are both coordinates on the image to be processed.

[0037] Furthermore, the processor 102 can extract a local image corresponding to the target detection frame from the shelf image based on the spatial information of the target detection frame; and input the local image into the target recognition model to identify the target object contained in the local image, thereby realizing automatic recognition of the target object, helping to improve the efficiency of target recognition, and further helping to improve the efficiency of subsequent commodity statistics based on the target recognition results.

[0038] In the embodiment of the present application, it is considered that the placement of items on the shelf may be tilted. For example, in shopping scenes such as shopping malls and supermarkets, customers may put the items back to a position that is inconsistent with the initial placement position when purchasing goods, and the goods may be tilted. Figure 1b and Figure 1c The pencil case shown in . Figure 1c is a schematic diagram of a frame of shelf image collected by the robot 100 during its movement along the shelf. Figure 1d As shown in the left figure, for tilted target objects, the local image extracted based on the spatial information of the rectangular detection box without angle information often contains excessive background noise, affecting the accuracy of subsequent target recognition based on the local image.

[0039] To address the above issues, this embodiment introduces rotation angle detection into the target detection model. The spatial information of the target detection frame output by the target detection model includes the rotation angle. The rotation angle included in the spatial information of the target detection frame corresponds to the rotation angle of the horizontal or vertical detection frame described above. The target detection frame is a rectangular frame, and its spatial information may include: the center coordinates, dimensions, and rotation angle of the target detection frame, which can be expressed as (x, y, w, h, θ); where θ represents the rotation angle of the target detection frame. Alternatively, the spatial information of the target detection frame may also include: the vertex coordinates and rotation angle of the target detection frame.

[0040] Based on the above-mentioned target detection model, the processor 102 can input the acquired shelf image into the target detection model to obtain the spatial information of the target detection frame for target annotation of the shelf image; the spatial information of the target detection frame includes: the center coordinates, size and rotation angle of the target detection frame, or the spatial information of the target detection frame includes: the vertex coordinates and rotation angle of the target detection frame. Further, the processor 102 can extract the local image corresponding to the target detection frame from the shelf image based on the spatial information including the rotation angle; the local image also has a rotation angle. The rotation angle is consistent with the tilt angle of the target object in the local image, and the extracted local image is as follows Figure 1d As shown in the right figure, the background noise of the target object is reduced, which helps to improve the accuracy of subsequent target recognition based on the local image. Figure 1d The dotted box shown in the right figure is the target detection box.

[0041] Optionally, the processor 102 may determine the position of the image marked with the target detection frame in the shelf image based on the spatial information including the rotation angle. The specific calculation formula is as follows:

[0042] (1)

[0043] In formula (1), (x, y) is the center coordinate of the target detection frame, that is, the coordinate of the center of the target detection frame in the shelf image. i=1,2. ( )and( ) represent the positions of the two vertices of the diagonal line of the target detection frame when it is at a positive angle. ( )and( ) represent the rotation angle of the target detection frame The positions of the two vertices of the diagonal line afterwards.

[0044] Furthermore, since the partial image has a rotation angle, in this embodiment, in order to achieve target recognition for the partial image with a rotation angle, a target recognition model that supports multi-angle target recognition can also be provided. Accordingly, the processor 102 can input the partial image with a rotation angle into the target recognition model that supports multi-angle target recognition to identify the target object contained in the partial image.

[0045] The robot provided in this embodiment can move along the shelf and collect shelf images while moving along the shelf. In the target detection stage of the shelf image, the shelf image is input into the target detection model to obtain spatial information including the rotation angle of the detection frame for target annotation of the shelf image. Therefore, extracting the local image corresponding to the detection frame with the rotation angle can reduce the background noise of the target object contained in the extracted local image, which helps to improve the accuracy of subsequent recognition of the target object contained in the local image; in the target recognition stage, based on the spatial information including the rotation angle, the local image corresponding to the detection frame is extracted from the shelf image, and the local image also has the above-mentioned rotation angle; and the local image with the rotation angle is input into the target recognition model that supports multi-angle target recognition to identify the target object in the local image. The target object in the local image can be identified without returning the local image, which helps to improve the efficiency of target recognition.

[0046] It is worth noting that, in the embodiment of the present application, before inputting the shelf image into the target detection model, the processor 102 may also perform model training on the target detection model. Optionally, the processor 102 may obtain multi-angle images and multi-distance images of a sample shelf on which a variety of commodities are placed as a sample image set. Among them, the multi-angle image of the sample shelf refers to an image obtained by capturing the sample shelf using multiple shooting angles. Optionally, the multi-angle image includes: a top-view image, a horizontal image, and a bottom-up image. The multi-distance image refers to an image obtained by capturing the sample at different shooting distances.

[0047] Furthermore, the processor 102 may also use the spatial information of the detection frame including the rotation angle to target each sample image in the sample image set, and obtain the spatial information of the detection frame for target annotation of each sample image, and the spatial information includes the rotation angle. In the embodiment of the present application, for the convenience of description and distinction, the detection frame for target annotation of the sample image is defined as the reference detection frame; and the target detection frame for annotating the shelf image collected by the robot 100 is defined as the first detection frame. The spatial information of the first detection frame and the reference detection frame both include the rotation angle information. Optionally, manual annotation can also be used to target each sample image in the sample image set to obtain the spatial information of the reference detection frame, and the spatial information includes the rotation angle.

[0048] Furthermore, the processor 102 may use the sample image set to perform model training on the target detection model to obtain a target detection model. The initial model for model training of the target detection model is referred to as the initial detection model. The initial detection model has the same model architecture as the target detection model finally obtained by the model training, that is, the parameters of the model are the same. The model training in this embodiment mainly refers to: using the sample image set to train the parameters of the target detection model to minimize the loss function. That is, with the minimization of the loss function as the training goal, the model is trained using the sample image set to obtain the initial target detection model. The loss function can be determined based on the spatial information of the detection frame containing the rotation angle obtained by the model training, and the spatial information of the reference detection frame for target annotation of the sample image set before model training. For the convenience of description and distinction, the detection frame containing the rotation angle obtained by the model training is defined as the second detection frame. Optionally, the loss function can be expressed as:

[0049] (2)

[0050] In the loss function (2), represents the center loss, that is, the loss between the center coordinates of the second detection box obtained by model training and the center coordinates of the reference detection box; represents the scale loss, that is, the loss between the width and height of the second detection box obtained by model training and the width and height of the reference detection box; represents the offset loss, i.e., the offset loss of the center coordinates of the second detection box obtained by model training compared with the center coordinates of the reference detection box; Represents the rotation angle loss, that is, the loss of the rotation angle of the second detection frame obtained by model training compared with the rotation angle of the reference detection frame. 、 and They represent the weights of scale loss, offset loss, and rotation angle loss respectively, and can be flexibly set according to actual conditions.

[0051] Among them, the rotation angle loss It can be expressed as:

[0052] (3)

[0053] Among them, K is the number of positive samples in the sample image set, and are the angle of the reference detection box and the rotation angle of the model training output, respectively.

[0054] The rotation angle prediction branch in the aforementioned object detection model can be introduced into the known object detection model architecture. The known object detection model architecture may be, but is not limited to, a Single Shot Detector (SSD) model, a YOLO (You only look once) series model, or a CenterNet model. For example, for a CenterNet model, the detection head of the CenterNet model can include three branches: the center coordinates of the second detection box, the center coordinate offset, and the scale prediction. These branches share the parameters of the Hourglass-104 model, and two convolutional layers are cascaded in the head to respectively implement the aforementioned rotation angle prediction. The Hourglass-104 model uses a large number of convolutional and deconvolutional layers to fuse multi-scale features. The final output feature map size is 1 / 4 of the input size, which helps reduce computational complexity.

[0055] Among them, the above-mentioned target detection model training stage can be completed before the robot 100 leaves the factory, or after leaving the factory, the user of the robot 100 provides a sample image set and starts the target detection model training related computer instructions to complete the target detection model training.

[0056] After the target detection model is trained, the processor 102 can use the target detection model to detect targets in the input shelf image and obtain spatial information of a first detection frame for target annotation of the input shelf image, where the spatial information includes a rotation angle. The rotation angle of the first detection frame is consistent with the rotation angle of the target object annotated by the first detection frame. Furthermore, based on the spatial information including the rotation angle, a local image corresponding to the first detection frame can be extracted from the shelf image, where the local image also has a rotation angle. Furthermore, the local image with the rotation angle can be input into a target recognition model that supports multi-angle target recognition to identify the target object contained in the local image.

[0057] In this embodiment, before the partial image with the rotation angle is input into the target recognition model that supports multi-angle target recognition, the target recognition model can also be trained so that the target recognition model can recognize targets from multiple angles. Optionally, the target recognition model can be trained using the ReID method. The specific implementation method is as follows:

[0058] In this embodiment, the robot 100 can obtain multi-angle images of a variety of sample commodities. Herein, multiple refers to 2 or more kinds. Preferably, the sample commodity is a single commodity. The multi-angle images of the sample commodity may include: a top-view image, a horizontal image, and an upward-view image. Herein, the top-view angle and the upward-view angle can be flexibly set. Optionally, the top-view angle and the upward-view angle can be 45°, etc. A variety of sample commodities include: the same commodity under different brands, different kinds of commodities under the same brand, and different specifications of the same commodity under the same brand, etc. Each sample commodity corresponds to a set of multi-angle images. Herein, multi-angle refers to multiple angles. Multiple refers to 2 or more, and the specific number of angles can be flexibly set.

[0059] In the embodiment of the present application, the specific implementation method of the robot 100 for acquiring multi-angle images of a plurality of sample commodities is not limited. In some embodiments, the robot 100 can flexibly adjust the acquisition angle of the camera 104 and use a different acquisition angle to shoot each sample commodity, thereby obtaining multi-angle images of a plurality of sample commodities. In other embodiments, other devices or systems can also acquire multi-angle images of a plurality of sample commodities and provide the acquired multi-angle images to the robot 100. For example, Figure 1e As shown, the multi-angle image acquisition system may include: a fixed camera and a rotatable tray S1. The installation position of the fixed camera is as follows: Figure 1e As shown. Figure 1e The multi-angle image acquisition system shown includes three cameras (i.e., cameras 1-3) with a 45-degree downward view, a 0-degree horizontal view, and a 45-degree upward view. Sample product A can be placed on a rotating tray, and each time the rotating tray rotates a set angle, a set of images is captured, including images captured by the camera with a 45-degree downward view, images captured by the camera with a 0-degree horizontal view, and images captured by the camera with a 45-degree upward view. For example, a set of images is captured every 6° rotation of the rotating tray, so that a total of 180 images are captured for each sample product A. Optionally, each frame of image captured by the multi-angle image acquisition system can be cropped to retain only the sample product portion, and the cropped image is used as the above-mentioned multi-angle image.

[0060] In practice, products from different brands vary significantly in appearance, making them easy to distinguish. These products can be of the same type, such as Brand A cola and Brand B cola, or they can be of different types, such as Brand A Sprite and Brand B cola. Different types of products from the same brand also vary significantly in appearance, making them easy to distinguish. For example, Brand A Sprite and Coke; Brand B potato chips and bread. However, different sizes of the same product under the same brand can be very similar in appearance, making identification more difficult. For example, Brand A 500ml cola and Brand B 250ml cola.

[0061] In an embodiment of the present application, in order to improve the accuracy of target recognition, the target recognition model can be trained by distinguishing sample products with different granularity to perform coarse and fine recognition model training. In this embodiment, for the convenience of description and distinction, the coarse recognition model is defined as the first target recognition sub-model, and the fine recognition model is defined as the second target recognition sub-model. Among them, the first target recognition sub-model and the second target recognition sub-model both support multi-angle target recognition. The fineness of the image feature vector extracted by the second target recognition sub-model is greater than that of the image feature vector extracted by the first target recognition sub-model. Optionally, the first target recognition sub-model can distinguish and identify sample products of different brands and different types of products under the same brand; the second target recognition sub-model can distinguish and identify different specifications of the same type of products under the same brand.

[0062] Furthermore, the processor 102 may train the first object recognition sub-model and the second object recognition sub-model separately. Optionally, when training the first object recognition sub-model, the processor 102 may obtain images of the same type and same brand of sample products from multi-angle images of multiple sample products as positive image pairs (referred to as the first positive image pair); obtain images of sample products of different brands or images of sample products of different types within the same brand as negative image pairs (referred to as the first negative image pair); and form a triplet (referred to as the first triplet) of positive and negative image pairs containing the same anchor image. That is, for an anchor image in a first triplet, one of the other two image frames is an image of a sample product of the same brand and type as the sample product contained in the anchor image. This image and the anchor image constitute the first positive image pair; the other image is an image of a sample product of a different brand than the sample product contained in the anchor image, or an image of a sample product of a different type within the same brand as the sample product contained in the anchor image. This image and the anchor image constitute the first negative image pair.

[0063] Optionally, the sample products contained in each frame of the multi-angle image of the sample products may be pre-labeled with a primary product identifier, where the primary product identifier may be a brand name + product name, such as Brand A Coke, Brand A Sprite, etc. In this way, the processor 102 may determine, based on the primary product identifiers of the sample products contained in each frame of the multi-angle image of the multiple sample products, images of the same type of sample products of the same brand in the multi-angle image of the sample products; and images of sample products of different brands or images of different types of sample products of the same brand in the multi-angle image of the sample products.

[0064] Optionally, the sample products contained in each frame of the multi-angle images of the plurality of sample products may be pre-labeled with secondary product identifications, where the secondary product identifications may be brand name + product name + specifications, such as Brand A Coke 500ml, Brand A Coke 250ml, etc. Thus, in the following embodiments, during the training of the second object recognition sub-model, the processor 102 may determine, based on the secondary product identifications of the sample products contained in each frame of the multi-angle images of the plurality of sample products, images of sample products of the same brand, type, and specification in the multi-angle images of the sample products; and images of sample products of the same brand, type, and different specifications in the multi-angle images of the sample products.

[0065] In this embodiment, training the first target recognition sub-model primarily involves training the feature extraction layer within the first target recognition sub-model. To distinguish it from the feature extraction layer within the second target recognition sub-model described below, the feature extraction layer within the first target recognition sub-model is defined as the first feature extraction layer, and the feature extraction layer within the second target recognition sub-model described below is defined as the second feature extraction layer. The image feature vectors extracted by the trained second feature extraction layer are more refined than those extracted by the trained first feature extraction layer.

[0066] The initial network model used to train the first feature extraction layer is referred to as the first initial network model. The first initial network model has the same model architecture as the first feature extraction layer ultimately obtained from model training, i.e., the model parameters are the same. Model training in this embodiment primarily refers to training the parameters of the first initial network model using triples to minimize the loss function.

[0067] For the ReID method, the feature extraction layer can be connected to the classifier, and the classifier determines the product identification contained in each sample image in the first triplet based on the image feature vector output by the feature extraction layer. In this embodiment, the product identification is information that uniquely identifies a product. For example, the product identification can be a product ID, or the brand, product name and specifications of the product. Accordingly, the loss function for training the first feature extraction layer can be a joint loss function composed of a triplet loss function, a center loss function and a product category loss function. For the convenience of description and distinction, the above-mentioned loss function for training the target detection model is defined as the first loss function; the loss function for training the first feature extraction layer is defined as the second loss function; and the loss function for training the second target recognition sub-model below is defined as the third loss function.

[0068] Among them, the triple loss function and the center loss function in the second loss function can be determined based on the image feature vector output by the first initial network model in each training; the product category loss in the second loss function can be determined based on the differences in the product categories contained in the triples obtained in each training and the actual differences in the product categories contained in the triples.

[0069] Furthermore, the first training model can be trained using triplets of positive and negative image pairs, with minimizing the second loss function as the training objective, to obtain a first feature extraction layer. The first training model includes a first initial network model and a classifier for training the first feature extraction layer. Specifically, the first initial network model when minimizing the second loss function is the first feature extraction layer. Of course, the parameters of the first initial network model when minimizing the second loss function have changed compared to the parameters of the first initial network model during initial training.

[0070] Furthermore, the processor 102 may input the multi-angle images of the aforementioned multiple sample products into the trained first feature extraction layer to obtain image feature vectors corresponding to the multi-angle images as product feature vectors in a product feature vector set (referred to as the first product feature vector set). Based on the product feature vectors in the first product feature vector set, the processor 102 constructs a first product feature vector set. Furthermore, the processor 102 establishes a correspondence between each product feature vector in the first product feature vector set and the product identifier. Optionally, the dimension of the product feature vectors in the first product feature vector set may be 256, but is not limited thereto. The product identifier in the correspondence between the first product feature vector set and the product identifier may be the aforementioned secondary product identifier. Because the negative sample image pairs in the triplets used to train the first feature extraction layer are images of sample products of different brands or different types of sample products of the same brand, the image feature vectors of the negative sample image pairs differ significantly. Therefore, the image feature vectors in the first product feature vector set have a coarse granularity. Furthermore, because the sample images used to train the first feature extraction layer are multi-angle images of a variety of sample products, the trained first feature extraction layer can support feature extraction from multi-angle images, and the image feature vectors in the first product feature vector set are image feature vectors from multi-angle images. Therefore, the trained first object recognition sub-model can support multi-angle object recognition.

[0071] In an embodiment of the present application, the second target recognition sub-model can also be trained. Optionally, when training the second target recognition sub-model, the processor 102 can obtain images of sample products of the same brand, type, and specification from multi-angle images of multiple sample products as positive sample image pairs; and obtain images of sample products of the same brand, type, and specifications but different from each other as negative sample image pairs, and form a triplet (recorded as a second triplet) with the positive sample image pairs and the negative sample image pairs containing the same anchor image. That is, for an anchor image in a second triplet, one of the other two frames of images is an image of a sample product of the same brand, type, and specification as the sample product contained in the anchor image, and this frame of image and the anchor image constitute a positive sample image pair; the other frame of image is an image of a sample product of the same brand, type, and specifications but different from each other as the sample product contained in the anchor image, and this frame of image and the anchor image constitute a negative sample image pair.

[0072] In this embodiment, training the second target recognition sub-model primarily involves training the feature extraction layer (referred to as the second feature extraction layer) in the second target recognition sub-model. The initial network model used to train the second feature extraction layer is referred to as the second initial network model, where the second initial network model can be the trained first feature extraction layer described above. The second initial network model has the same model architecture as the second feature extraction layer ultimately obtained through model training, i.e., the model parameters are the same. Model training in this embodiment primarily refers to training the parameters of the second initial network model using the second triplet to minimize the loss function.

[0073] For ReID methods, the feature extraction layer can be coupled with a classifier, which determines the product identity of each sample image in the second triplet based on the image feature vectors output by the feature extraction layer. Accordingly, the third loss function used to train the second feature extraction layer can be a joint loss function composed of a triplet loss function, a center loss function, and a product category loss function.

[0074] Among them, the triplet loss function and the center loss function in the third loss function can be determined based on the image feature vector output by the first feature extraction layer in each training; the product category loss in the third loss function can be determined based on the differences in the product categories contained in the triples obtained in each training and the actual differences in the product categories contained in the triples.

[0075] Furthermore, the second training model can be trained using a second triplet consisting of a positive sample image pair and a negative sample image pair with the minimization of the third loss function as the training objective to obtain a second feature extraction layer. The second training model includes: a second initial network model and a classifier. Optionally, the second initial network model can be the first feature extraction layer that has been trained as described above. Specifically, the second initial network model when the third loss function is minimized is the second feature extraction layer. Of course, the parameters of the second feature extraction layer when the third function is minimized have changed compared to those in the second initial network model during initial training.

[0076] Furthermore, the processor 102 may input the multi-angle images of the aforementioned multiple sample products into the trained second feature extraction layer to obtain image feature vectors corresponding to the multi-angle images as product feature vectors in a product feature vector set (referred to as the second product feature vector set). Based on the product feature vectors in the second product feature vector set, a second product feature vector set is constructed. A correspondence between each product feature vector in the second product feature vector set and the product identifier is established. Optionally, the dimension of the product feature vectors in the second product feature vector set may be 256, but is not limited thereto. The product identifier in the correspondence between the second product feature vector set and the product identifier may be the aforementioned secondary product identifier. Because the negative sample image pairs in the second triplet used to train the second feature extraction layer are images of sample products of the same brand and type but of different sizes, the image feature vectors in the negative sample image pairs have relatively small differences. Therefore, the image feature vectors in the second product feature vector set have finer granularity, i.e., the image feature vectors in the second product feature vector set have greater refinement than the image feature vectors in the first product feature vector set. Furthermore, because the sample images used to train the second feature extraction layer are multi-angle images of a variety of sample products, the trained second feature extraction layer can support feature extraction from multi-angle images, and the image feature vectors in the second product feature vector set are image feature vectors from multi-angle images. Therefore, the trained second object recognition sub-model can support multi-angle object recognition.

[0077] Based on the above-mentioned first target recognition sub-model and second target recognition sub-model, during target recognition, the processor 102 can input the local image with a rotation angle into the first target recognition sub-model to identify the target object contained in the local image, and if the first target recognition sub-model cannot identify the target object contained in the local image, the local image with the selected angle will be input into the second target recognition sub-model again, and the second target recognition sub-model will identify the target object contained in the local image.

[0078] Optionally, the first object recognition sub-model may include a first feature extraction layer that supports feature extraction from multi-angle images and a first object recognition layer that supports multi-angle object recognition. When identifying a target object contained in a partial image, the processor 102 may input the rotated partial image into the first feature extraction layer to obtain an image feature vector for the partial image (denoted as a first image feature vector). The processor 102 then inputs the first image feature vector into the first object recognition layer to identify the target object contained in the partial image.

[0079] Furthermore, in the first target recognition layer, the similarity between the first image feature vector of the partial image and the product feature vectors in the first product feature vector set can be calculated. Optionally, the cosine distance between the first image feature vector of the partial image and each product feature vector in the first product feature vector set can be calculated, where the shorter the cosine distance, the greater the similarity. M product feature vectors are selected from the first product feature vector set in descending order of similarity to the first image feature vector. Then, based on the correspondence between the product feature vectors and product identifiers in the first product feature vector set, the product identifiers corresponding to the M product feature vectors are determined. Furthermore, if the number Q of the M product feature vectors that has the same product identifier as the product identifier corresponding to the product feature vector with the greatest similarity is ≥ N, the product identifier corresponding to the product feature vector with the greatest similarity is used as the identifier of the target object contained in the partial image. In this embodiment, M and N are both integers, with M ≥ 2 and 1 ≤ N ≤ (M - 1). In this embodiment, the specific values ​​of M and N are not limited; alternatively, M = 5 and N = 3 are possible, but are not limited to these values.

[0080] Accordingly, if Q < N, it is determined that the first object recognition sub-model cannot recognize the target object contained in the partial image. Further, the processor 102 inputs the partial image with the selected angle into the second object recognition sub-model, and the second object recognition sub-model recognizes the target object contained in the partial image.

[0081] Optionally, the second object recognition sub-model includes: a second feature extraction layer that supports feature extraction of multi-angle images and a second object recognition layer that supports multi-angle object recognition, wherein the image feature vectors extracted by the second feature extraction layer have a higher degree of refinement than the image feature vectors extracted by the first feature extraction layer.

[0082] Correspondingly, when Q is less than N, the processor 102 inputs the above-mentioned local image with the rotation angle into the second feature extraction layer to obtain a second image feature vector of the local image; and inputs the second image feature vector into the second target recognition layer to identify the target object contained in the local image.

[0083] Optionally, in the second object recognition layer, the correspondence between product feature vectors and product identifiers in the second product feature vector set can be determined. The first target product feature vectors corresponding to the product identifiers corresponding to the M product feature vectors can be obtained from the second product feature vector set. The similarity between the second image feature vector and the first target product feature vector can also be calculated. Alternatively, the cosine distance between the second image feature vector and the first target product feature vector of the partial image can be calculated, where a shorter cosine distance indicates a greater similarity.

[0084] Furthermore, if there is a second target product feature vector in the first target product feature vector whose similarity with the second image feature vector is greater than or equal to the set similarity threshold, the identity of the target object contained in the partial image is determined based on the correspondence between the product feature vector and the product identification in the second product feature vector set and the second target product feature vector.

[0085] Optionally, a product feature vector having the greatest similarity with the second image feature vector can be determined from the second target product feature vector; and based on the correspondence between the product feature vectors and product identifications in the second product feature vector set, the product identification corresponding to the product feature vector having the greatest similarity with the second image feature vector is determined as the identification of the target object contained in the partial image.

[0086] Alternatively, in the second object recognition layer, the similarity between the second image feature vector of the partial image and the product feature vectors in the second product feature vector set can be calculated; wherein the product feature vectors in the second product feature vector set have a higher degree of refinement than the product feature vectors in the first product feature vector set. Alternatively, the cosine distance between the second image feature vector of the partial image and each product feature vector in the first product feature vector set can be calculated, wherein the shorter the cosine distance, the greater the similarity.

[0087] Furthermore, if there is a target product feature vector in the second product feature vector set whose similarity with the second image feature vector is greater than or equal to a set similarity threshold, the identity of the target object contained in the partial image is determined based on the correspondence between the product feature vector and the product identification in the second product feature vector set and the target product feature vector.

[0088] Optionally, a product feature vector having the greatest similarity with the second image feature vector can be determined from the target product feature vector; and based on the correspondence between the product feature vectors and product identifications in the second product feature vector set, the product identification corresponding to the product feature vector having the greatest similarity with the second image feature vector is determined as the identification of the target object contained in the partial image.

[0089] The above-described object recognition process involves capturing any frame of shelf imagery as the robot 100 moves along the shelf. In practical applications, the robot 100 may capture multiple frames of shelf imagery, which may overlap to some extent. In this embodiment, the processor 102 may also extract feature points from the multiple frames of shelf imagery. Feature points are local expressions of shelf image features and can reflect the local specificity of the shelf image.

[0090] Optionally, the processor 102 further performs spot detection or corner detection on the multiple frames of shelf images. If the processor 102 performs spot detection on the shelf images, the LOG method, DOH method, SIFI algorithm, or SURF algorithm may be used to obtain feature points in the shelf images. If the processor 102 performs corner detection on the shelf images, the Harris algorithm or FAST algorithm may be used to obtain feature points in the shelf images.

[0091] Furthermore, the processor 102 may perform deduplication processing on the multiple frames of shelf images based on their feature points. Optionally, for any first shelf image and second shelf image captured adjacently in time in the multiple frames of shelf images, the similarity between the feature points of the first shelf image and the second shelf image is calculated; and based on the similarity between the feature points of the first shelf image and the second shelf image, a perspective transformation matrix between the first shelf image and the second shelf image is calculated; then, based on the perspective transformation matrix, an affine transformation is performed on the first shelf image and the second shelf image to determine the overlapping area of ​​the first shelf image and the second shelf image; and the overlapping area of ​​the first shelf image and the second shelf image is deduplicated, thereby achieving deduplication of the first shelf image and the second shelf image.

[0092] Furthermore, the processor 102 may perform deduplication processing on the target objects contained in the multiple frames of shelf images according to the result of deduplication processing on the multiple frames of shelf images, thereby achieving target detection and recognition on the entire shelf.

[0093] It's worth noting that the object recognition method provided in the above embodiments can be executed not only by a robot but also by another computer device with which the robot communicates. In this case, the robot can provide the acquired shelf image to the other computer device, which then performs object recognition on the shelf image. The specific implementation of object recognition on the shelf image by the other computer device can be found in the aforementioned discussion of object recognition performed by the robot and will not be further elaborated here.

[0094] In addition to the above-mentioned robot, some exemplary embodiments of the present application also provide a target recognition method, which is described in detail below with reference to the accompanying drawings.

[0095] Figure 2aThe following is a flow chart of a target recognition method based on artificial intelligence provided in an embodiment of the present application. Figure 2a As shown, the method includes:

[0096] 20a. Acquire shelf images collected by the robot while it moves along the shelf.

[0097] 20b. Input the shelf image into the target detection model to obtain spatial information of a first detection frame for target labeling of the shelf image.

[0098] 20c. Extract a partial image corresponding to the first detection frame from the shelf image based on the spatial information of the first detection frame.

[0099] 20d. Input the local image into the target recognition model to identify the target object contained in the local image.

[0100] The target recognition method provided in this embodiment can be executed by an autonomous mobile robot or by other computer devices that communicate with the autonomous mobile robot. Regardless of the device that executes the target recognition method, in step 20a, the shelf image collected by the robot while it moves along the shelf can be obtained. The shelf image includes images of goods placed on the shelf. For the case where the execution subject is a robot, an optional implementation method of step 20a is: control the robot to move along the shelf, and control the camera on the robot to collect shelf images during the robot's movement along the shelf. For the case where the execution subject is other devices that communicate with the robot, another optional implementation method of step 20a is: receive the shelf image collected by the robot while it moves along the shelf.

[0101] Furthermore, considering that the quality of shelf images will affect the subsequent product detection and recognition effects, the existing technology usually uses handheld devices or fixed cameras to collect shelf images. The shooting angle of handheld devices is not fixed, and the collected images are prone to tilt. Fixed cameras are generally fixed at high positions with a large viewing angle but are far away from the products, and cannot obtain high-definition images, which will reduce the accuracy of product recognition.

[0102] To address the above issues, in step 20a, the camera's capture angle of view on the robot can be controlled to maintain a stable relative position to the merchandise on the shelf. Optionally, the camera's capture angle of view can be controlled to face the merchandise on the shelf. Furthermore, the robot can be controlled to move parallel to the shelf, and the camera can be controlled to capture images of the shelf during the robot's movement. Keeping the robot's movement parallel to the shelf helps ensure that the camera's capture angle remains stable, improving the quality of the shelf images captured by the camera and, in turn, increasing the accuracy of subsequent detection and recognition of merchandise in the shelf images.

[0103] Optionally, the position distribution of the shelves in the environmental map can be determined based on a known environmental map; and based on the position distribution of the shelves in the environmental map, a moving path parallel to the shelves can be planned for the robot, and the robot can be controlled to move along the moving path parallel to the shelves.

[0104] Furthermore, the distance between the robot and the shelf can be automatically adjusted to ensure that the camera can capture images of the entire shelf. Optionally, the robot can automatically adjust the height of the camera and the distance between the robot and the shelf to ensure that the camera can capture images of the entire shelf.

[0105] Furthermore, in step 20b, the shelf image can be input into the target detection model to obtain the spatial information of the first detection frame for target annotation of the shelf image. For the description of the spatial information of the first detection frame, please refer to the relevant content of the above-mentioned robot embodiment, which will not be repeated here. Then, in step 20c, based on the spatial information of the first detection frame, the local image corresponding to the first detection frame can be extracted from the shelf image; and in step 20d, the local image is input into the target recognition model to identify the target object contained in the local image, thereby realizing automatic recognition of the target object, helping to improve the efficiency of target recognition, and further helping to improve the efficiency of subsequent commodity statistics based on the target recognition results.

[0106] Furthermore, consider that items on shelves may be tilted. For example, in shopping scenarios like malls and supermarkets, customers may place items back in a different location than where they were originally placed, causing the items to appear tilted. For tilted objects, the local image extracted based on the spatial information of the rectangular detection frame, which lacks angular information, often contains excessive background noise, affecting the accuracy of subsequent object recognition based on the local image.

[0107] In order to solve the above problems, the embodiment of the present application also provides another target recognition method based on artificial intelligence. Figure 2b As shown, the method includes:

[0108] 201. Obtain shelf images collected by the robot while it moves along the shelf.

[0109] 202. Input the shelf image into the target detection model to obtain spatial information of a first detection frame for target annotation of the shelf image; wherein the spatial information of the first detection frame includes: the center coordinates, size and rotation angle of the first detection frame, or the vertex coordinates and rotation angle of the first detection frame.

[0110] 203. Extract a partial image corresponding to the first detection frame from the shelf image according to the spatial information including the rotation angle; the partial image has a rotation angle.

[0111] 204. Input the local image with the rotation angle into a target recognition model supporting multi-angle target recognition to recognize the target object contained in the local image.

[0112] In this embodiment, to improve the accuracy of target recognition, the target detection model introduces rotation angle detection. The spatial information of the target detection box output by the target detection model includes the rotation angle. The description of the rotation angle can be found in the relevant content of the above embodiment and will not be repeated here.

[0113] Based on the aforementioned object detection model, in step 202, the acquired shelf image can be input into the object detection model to obtain spatial information for a target detection frame used to annotate the shelf image. This spatial information includes a rotation angle. Furthermore, in step 203, based on the spatial information including the rotation angle, a partial image corresponding to the target detection frame can be extracted from the shelf image. This partial image also has a rotation angle. This rotation angle aligns with the tilt angle of the target object in the partial image, reducing background noise of the target object and improving the accuracy of subsequent target recognition based on the partial image.

[0114] Furthermore, since the partial image has a rotation angle, in this embodiment, in order to achieve target recognition for the partial image with a rotation angle, a target recognition model that supports multi-angle target recognition can also be provided. Accordingly, in step 204, the partial image with a rotation angle can be input into the target recognition model that supports multi-angle target recognition to identify the target object contained in the partial image.

[0115] In this embodiment, shelf images captured by the robot as it moves along the shelf can be obtained. During the target detection phase of the shelf image, the shelf image is input into a target detection model to obtain spatial information, including the rotation angle, of the detection frame used to annotate the shelf image. Therefore, extracting the local image corresponding to the detection frame with the rotation angle can reduce background noise of the target object contained in the extracted local image, helping to improve the accuracy of subsequent recognition of the target object contained in the local image. During the target recognition phase, based on the spatial information including the rotation angle, a local image corresponding to the detection frame is extracted from the shelf image, and this local image also has the aforementioned rotation angle. The local image with the rotation angle is then input into a target recognition model that supports multi-angle target recognition to identify the target object in the local image. This allows the target object in the local image to be identified without having to correct the local image, helping to improve target recognition efficiency.

[0116] It is worth noting that, in the embodiment of the present application, before the shelf image is input into the target detection model, the target detection model can also be trained. Optionally, multi-angle images of a sample shelf with a variety of commodities can be obtained as a sample image set. It is also possible to use the detection frame spatial information containing the rotation angle to target each sample image in the sample image set, and obtain the spatial information of the detection frame for target marking each sample image, and these spatial information include the rotation angle. In the embodiment of the present application, for the convenience of description and distinction, the detection frame for target marking the sample image is defined as the reference detection frame; and the target detection frame for marking the shelf image collected by the robot is defined as the first detection frame. The spatial information of the first detection frame and the reference detection frame both include the rotation angle information. Optionally, it is also possible to use manual marking to target each sample image in the sample image set, and obtain the spatial information of the reference detection frame, and the spatial information includes the rotation angle.

[0117] Furthermore, multi-angle images and multi-distance images of sample shelves with a variety of commodities can be obtained as the first sample image set. Particularly, for the description of multi-angle images and multi-distance images, please refer to the relevant content of the above embodiment, which will not be repeated here. Further, the minimization of the first loss function can be used as the training goal, and the first sample image set can be used for model training to obtain a target detection model; wherein, the first loss function is determined based on the spatial information of the second detection frame containing the rotation angle obtained by model training, and the spatial information of the reference detection frame for target annotation of the first sample image set before model training; the spatial information of the reference detection frame includes the rotation angle. Particularly, for the first loss function and the specific training process of the target detection model, please refer to the relevant content of the above robot embodiment, which will not be repeated here.

[0118] After the target detection model is trained, it can be used to detect targets in the input shelf image, obtaining spatial information of a first detection frame for target annotation of the input shelf image. This spatial information includes a rotation angle. The rotation angle of the first detection frame is consistent with the rotation angle of the target object annotated by the first detection frame. Furthermore, based on the spatial information including the rotation angle, a partial image corresponding to the first detection frame can be extracted from the shelf image. This partial image also has a rotation angle. Furthermore, the partial image with the rotation angle can be input into a target recognition model that supports multi-angle target recognition to identify the target object contained in the partial image.

[0119] In this embodiment, before the partial image at the selected angle is input into the target recognition model supporting multi-angle target recognition, the target recognition model may be trained to enable the target recognition model to recognize targets from multiple angles. Optionally, a ReID method may be used to train the target recognition model. The specific implementation method can be found in the relevant content of the above embodiment and will not be repeated here.

[0120] In this embodiment, the multi-angle images of the aforementioned multiple sample products can be input into a trained first feature extraction layer to obtain image feature vectors corresponding to the multi-angle images, which are used as product feature vectors in a product feature vector set (referred to as a first product feature vector set). Based on the product feature vectors in the first product feature vector set, a first product feature vector set is constructed. A corresponding relationship between each product feature vector in the first product feature vector set and the product identifier is established. Correspondingly, the multi-angle images of the aforementioned multiple sample products can also be input into a trained second feature extraction layer to obtain image feature vectors corresponding to the multi-angle images, which are used as product feature vectors in a product feature vector set (referred to as a second product feature vector set). Based on the product feature vectors in the second product feature vector set, a second product feature vector set is constructed. A corresponding relationship between each product feature vector in the second product feature vector set and the product identifier is established. Based on the above-mentioned first target recognition sub-model and second target recognition sub-model, since the fineness of the image feature vector extracted by the second target recognition sub-model is greater than that of the image feature vector extracted by the first target recognition sub-model, the local image with a rotation angle can be first input into the first target recognition sub-model to identify the target object contained in the local image; and when the first target recognition sub-model cannot identify the target object contained in the local image, the local image with a rotation angle is input into the second target recognition sub-model to identify the target object contained in the local image.

[0121] In this embodiment, the first target recognition sub-model may include: a first feature extraction layer that supports feature extraction from multi-angle images, and a first target recognition layer that supports multi-angle target recognition. Accordingly, based on the first target recognition sub-model, when identifying a target object contained in a partial image, the partial image with a rotation angle may be input into the first feature extraction layer to obtain an image feature vector for the partial image (denoted as a first image feature vector); the first image feature vector is then input into the first target recognition layer to identify the target object contained in the partial image.

[0122] Optionally, in the first target recognition layer, the similarity between the image feature vector of the partial image and the product feature vectors in the first product feature vector set can be calculated. M product feature vectors are selected from the first product feature vector set in descending order of similarity to the first image feature vector. Then, based on the correspondence between the product feature vectors and product identifiers in the first product feature vector set, the product identifiers corresponding to the M product feature vectors are determined. Furthermore, if the number Q of the M product feature vectors that has the same product identifier as the product identifier corresponding to the product feature vector with the greatest similarity is ≥ N, the product identifier corresponding to the product feature vector with the greatest similarity is used as the identifier of the target object contained in the partial image. In this embodiment, M and N are both integers, with M ≥ 2 and 1 ≤ N ≤ (M - 1). In this embodiment, the specific values ​​of M and N are not limited; M = 5 and N = 3 are optional, but are not limited to these values.

[0123] Correspondingly, if Q<N, it is determined that the first target sub-model cannot recognize the target object contained in the partial image. Further, the partial image with a rotation angle can be input into the second target recognition sub-model to recognize the target object contained in the partial image.

[0124] The second target recognition sub-model may include: a second feature extraction layer that supports feature extraction from multi-angle images and a second target recognition layer that supports multi-angle target recognition. Accordingly, if Q < N, the rotated partial image may be input into the second feature extraction layer to obtain a second image feature vector for the partial image; the second image feature vector is then input into the second target recognition layer to identify the target object contained in the partial image.

[0125] Optionally, in the second target recognition layer, based on the correspondence between the product feature vectors and the product identification in the second product feature vector set, the first target product feature vectors corresponding to the product identifications corresponding to the M product feature vectors are obtained from the second product feature vector set; and the similarity between the second image feature vector and the first target product feature vector is calculated; if there is a second target product feature vector in the first target product feature vector, and the similarity between the second image feature vector and the second target product feature vector is greater than or equal to the set similarity threshold, then based on the correspondence between the product feature vector and the product identification in the second product feature vector set and the second target product feature vector, the identification of the target object contained in the local image is determined.

[0126] Optionally, a product feature vector having the greatest similarity with the second image feature vector is determined from the second target product feature vector; and based on the correspondence between the product feature vectors and product identifications in the second product feature vector set, the product identification corresponding to the product feature vector having the greatest similarity with the second image feature vector is determined as the identification of the target object contained in the partial image.

[0127] For other implementations of the second target recognition layer for identifying the target object, please refer to the relevant content of the above-mentioned robot embodiment, which will not be repeated here.

[0128] The shelf image used in the target recognition process is any frame of shelf imagery captured by the robot as it moves along the shelf. In actual applications, the robot may capture multiple frames of shelf imagery as it moves along the shelf, and these multiple frames may overlap to some extent. In this embodiment, feature points from these multiple frames of shelf imagery may also be extracted. Feature points are local expressions of shelf image features and can reflect the local specificity of the shelf image. The specific implementation method for extracting feature points from multiple frames of shelf imagery can be found in the relevant content of the aforementioned robot embodiment and will not be further elaborated here.

[0129] Furthermore, deduplication processing can be performed on the multiple frames of shelf images based on their feature points. Optionally, for any first shelf image and second shelf image captured adjacently in time in the multiple frames of shelf images, the similarity between the feature points of the first shelf image and the second shelf image is calculated; and based on the similarity between the feature points of the first shelf image and the second shelf image, a perspective transformation matrix between the first shelf image and the second shelf image is calculated; then, based on the perspective transformation matrix, an affine transformation is performed on the first shelf image and the second shelf image to determine the overlapping area of ​​the first shelf image and the second shelf image; and the overlapping area of ​​the first shelf image and the second shelf image is deduplicated, thereby achieving deduplication of the first shelf image and the second shelf image.

[0130] Furthermore, based on the result of deduplication processing on the multiple frames of shelf images, deduplication processing can be performed on the target objects contained in the multiple frames of shelf images, thereby achieving target detection and recognition of the entire shelf.

[0131] It should be noted that the execution entity of each step of the method provided in the above embodiment can be the same device, or the method can be executed by different devices. For example, the execution entity of steps 201 and 202 can be device A; for another example, the execution entity of step 201 can be device A, and the execution entity of step 202 can be device B; and so on.

[0132] In addition, some of the processes described in the above embodiments and the accompanying drawings include multiple operations that appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order in which they appear in this document or may be executed in parallel. The sequence numbers of the operations, such as 201, 202, etc., are merely used to distinguish between different operations and do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel.

[0133] Accordingly, an embodiment of the present application further provides a computer-readable storage medium storing computer instructions, which, when executed by one or more processors, causes the one or more processors to execute the steps in the above-mentioned target recognition method.

[0134] Figure 3 A schematic diagram of the structure of a computer device provided in an embodiment of the present application. The computer device may be a terminal device such as a smartphone or a computer, or a server device. The server device may be a single server device, a cloud-based server array, or a virtual machine (VM) running in a cloud-based server array. Furthermore, the server device may also refer to other computing devices with corresponding service capabilities, such as a terminal device such as a computer (running a service program).

[0135] like Figure 3 As shown, the computer device includes: a memory 30a and a processor 30b. The memory 30a is used to store a computer program. The processor 30b is coupled to the memory 30a and is used to execute the computer program to: obtain shelf images captured by the robot while moving along the shelf; input the shelf images into a target detection model to obtain spatial information of a first detection frame for target annotation of the shelf images; extract a partial image corresponding to the first detection frame from the shelf image based on the spatial information of the first detection frame; and input the partial image into a target recognition model to identify the target object contained in the partial image.

[0136] In some embodiments, the spatial information of the first detection frame includes a rotation angle. The spatial information of the first detection frame may also include the center coordinates and dimensions of the first detection frame; or the vertex coordinates of the first detection frame. Accordingly, when identifying a target object contained in a partial image, the processor 30b is specifically configured to input the partial image with the rotation angle into a target recognition model that supports multi-angle target recognition to identify the target object contained in the partial image.

[0137] In some embodiments, the computer device further includes a communication component 30c. Optionally, when the processor 30b obtains the shelf image captured by the robot while moving along the shelf, it is specifically configured to: receive the shelf image captured by the robot while moving in a direction parallel to the shelf through the communication component 30c.

[0138] In some embodiments, the processor 30b is further used to: before inputting the shelf image into the target detection model, obtain multi-angle images of a sample shelf with a variety of commodities as a first sample image set; with minimization of the first loss function as the training goal, use the first sample image set to perform model training to obtain a target detection model; wherein the first loss function is determined based on the spatial information of the second detection frame including the rotation angle obtained during model training, and the spatial information of the reference detection frame for target annotation of the first sample image set before model training; the spatial information of the reference detection frame includes the rotation angle.

[0139] Optionally, the processor 30b is further configured to: perform object annotation on the first sample image set using the detection frame spatial information including the rotation angle, so as to obtain the reference detection frame spatial information including the rotation angle.

[0140] In other embodiments, the target recognition model includes: a first target recognition sub-model supporting multi-target recognition and a second target recognition sub-model supporting multi-target recognition; wherein the image feature vector extracted by the second target recognition sub-model has a higher degree of refinement than the image feature vector extracted by the first target recognition sub-model. Accordingly, when identifying a target object contained in a partial image, the processor 30b is specifically configured to: input the partial image with a rotation angle into the first target recognition sub-model to identify the target object contained in the partial image; and if the first target recognition sub-model cannot identify the target object contained in the partial image, input the partial image with a rotation angle into the second target recognition sub-model to identify the target object contained in the partial image.

[0141] Optionally, the first target sub-model includes: a first feature extraction layer that supports feature extraction from multi-angle images; and a first target recognition layer that supports multi-angle target recognition. Accordingly, when identifying a target object contained in a partial image, processor 30b is specifically configured to: input the rotated partial image into the first feature extraction layer to obtain a first image feature vector for the partial image; and input the first image feature vector into the first target recognition sub-model to identify the target object contained in the partial image.

[0142] Furthermore, when the processor 30b identifies the target object contained in the partial image, it is specifically used to: calculate the similarity between the first image feature vector and the product feature vector in the first product feature vector set in the first target recognition layer; select M product feature vectors from the first product feature vector set in descending order of similarity with the first image feature vector; determine the product identifications corresponding to the M product feature vectors based on the correspondence between the product feature vectors and the product identifications in the first product feature vector set; if Q≥N, then use the product identification corresponding to the product feature vector with the greatest similarity as the identification of the target object contained in the partial image; wherein Q is the number of product feature vectors in the M product feature vectors that have the same product identification as the product feature vector with the greatest similarity; M≥2, 1≤N≤(M-1), and M and N are integers.

[0143] Optionally, the second target recognition sub-model includes: a second feature extraction layer that supports feature extraction from multi-angle images and a second target recognition layer that supports multi-angle target recognition; wherein the image feature vectors extracted by the second feature extraction layer have a higher degree of refinement than the image feature vectors extracted by the first feature extraction layer. Accordingly, the processor 30b is further configured to: if Q < N, input the partial image with a rotation angle into the second feature extraction layer to obtain a second image feature vector for the partial image; the second image feature vector has a higher degree of refinement than the first image feature vector; and input the second image feature vector into the second target recognition sub-model to identify the target object contained in the partial image.

[0144] Furthermore, when identifying the target object contained in the partial image, the processor 30b is specifically configured to:

[0145] In the second target recognition layer, based on the correspondence between the product feature vectors and the product identifiers in the second product feature vector set, the first target product feature vectors corresponding to the product identifiers corresponding to the M product feature vectors are obtained from the second product feature vector set;

[0146] Calculate the similarity between the second image feature vector and the first target product feature vector; wherein the fineness of the product feature vectors in the second product feature vector set is greater than that of the product feature vectors in the first product feature vector set; if there is a second target product feature vector in the first target product feature vector whose similarity with the second image feature vector is greater than or equal to a set similarity threshold, determine the identity of the target object contained in the partial image based on the correspondence between the product feature vectors and the product identity in the second product feature vector set and the second target product feature vector.

[0147] Furthermore, when determining the identification of the target object contained in the partial image, the processor 30b is specifically used to: determine, from the second target product feature vector, a product feature vector having the greatest similarity with the second image feature vector; and based on the correspondence between the product feature vectors and product identifications in the second product feature vector set, determine the product identification corresponding to the product feature vector having the greatest similarity with the second image feature vector as the identification of the target object contained in the partial image.

[0148] Optionally, the processor 30b is also used to: obtain multi-angle images of multiple sample products before inputting the local image with a rotation angle into the first feature extraction layer; obtain images of the same type of sample products belonging to the same brand from the multi-angle images of the sample products as the first positive sample image pair; and obtain images of sample products belonging to different brands or images of different types of sample products under the same brand as the first negative sample image pair; with minimization of the second loss function as the training goal, use the first triplet consisting of the first positive sample image pair and the first negative sample image pair to train the first training model to obtain the first feature extraction layer; wherein the first training model includes: a first initial network model and a classifier for training the first feature extraction layer; the second loss function is composed of a first triplet loss function, a first center loss function and a first product category loss; wherein the first triplet loss function and the first center loss function are determined based on the image feature vector output by the first initial network model each time training; the first product category loss is determined based on the difference in product categories contained in the first triplet obtained in each training, and the actual difference in product categories contained in the first triplet.

[0149] Optionally, the multi-angle image of each sample product includes: a top view image, a level view image, and a bottom view image of the sample product.

[0150] Optionally, the processor 30b is also used to: input the multi-angle images of the sample product into the trained first feature extraction layer to obtain image feature vectors corresponding to the multi-angle images of the sample product, as the product feature vectors in the first product feature vector set; construct the first product feature vector set based on the product feature vectors in the first product feature vector set; and establish a correspondence between each product feature vector in the first product feature vector set and the product identification.

[0151] Optionally, the processor 30b is also used to: before inputting the second image feature vector into the second target recognition sub-model, obtain images of sample products of the same type and specification under the same brand from multi-angle images as a second positive sample image pair, and obtain images of sample products of the same type and specification under the same brand but different specifications as a second negative sample image pair; with minimization of the third loss function as the training goal, train the second training model using the triplet consisting of the second positive sample image pair and the second negative sample image pair to obtain a second feature extraction layer; wherein the second training model includes: a first feature extraction layer and a classifier; the third loss function is composed of a second triplet loss function, a second center loss function and a second product category loss; wherein the second triplet loss function and the second center loss function are determined based on the image feature vector output by the second initial network model each time of training; the second product category loss is determined based on the difference in product categories contained in the second positive sample image pair and the second negative sample image pair output by the classifier each time of training, and the actual difference in product categories contained in the second positive sample image pair and the second negative sample image pair.

[0152] Correspondingly, the processor 30b is also used to: input multi-angle images of multiple sample products into the trained second feature extraction layer to obtain image feature vectors corresponding to the multi-angle images, respectively, as product feature vectors in the second product feature vector set; construct a second product feature vector set based on the product feature vectors in the second product feature vector set; and establish a correspondence between each product feature vector in the second product feature vector set and the product identification.

[0153] In some other embodiments, the number of shelf images is multiple frames; the processor 30b is further used to: extract feature points of the multiple frames of shelf images; deduplicate the multiple frames of shelf images based on the feature points of the multiple frames of shelf images; and deduplicate the target objects contained in the multiple frames of shelf images based on the results of the deduplication processing of the multiple frames of shelf images.

[0154] Furthermore, when the processor 30b performs deduplication processing on multiple frames of shelf images, it is specifically used to: calculate the similarity between the feature points of the first shelf image and the second shelf image that are adjacent in any acquisition time in the multiple frames of shelf images; calculate the perspective transformation matrix between the first shelf image and the second shelf image based on the similarity between the feature points of the first shelf image and the second shelf image; perform affine transformation on the first shelf image and the second shelf image based on the perspective transformation matrix to determine the overlapping area of ​​the first shelf image and the second shelf image; and perform deduplication processing on the overlapping area of ​​the first shelf image and the second shelf image.

[0155] In some optional embodiments, such as Figure 3As shown, the computer device may further include optional components such as a power supply component 30d, a display component 30e, and an audio component 30f. Figure 3 Only some components are shown schematically, which does not mean that the computer equipment must include Figure 3 The components shown do not necessarily mean that the computer equipment can only include Figure 3 Components shown.

[0156] The computer device provided in this embodiment can acquire shelf images captured by a robot as it moves along a shelf. During the target detection phase of the shelf image, the shelf image is input into a target detection model to obtain spatial information of a detection frame for labeling the target in the shelf image. During the target recognition phase, a partial image corresponding to the detection frame is extracted from the shelf image based on the spatial information including the rotation angle. This partial image is then input into a target recognition model to identify the target object in the partial image. This enables automatic recognition of the target object, helps improve target recognition efficiency, and in turn, helps improve the efficiency of subsequent product statistics based on the target recognition results.

[0157] In the embodiments of the present application, the memory is used to store computer programs and can be configured to store various other data to support operations on the device in which it is located. The processor can execute the computer program stored in the memory to implement the corresponding control logic. For details on the implementation of the memory, please refer to the relevant content of the above embodiments and will not be repeated here.

[0158] In the embodiments of the present application, the processor can be any hardware processing device that can execute the logic of the above method. Optionally, the processor can be a central processing unit (CPU), a graphics processing unit (GPU), or a microcontroller unit (MCU); it can also be a programmable device such as a field-programmable gate array (FPGA), a programmable array logic device (PAL), a general array logic device (GAL), a complex programmable logic device (CPLD); or it can be an advanced reduced instruction set (RISC) processor (Advanced RISC Machines, ARM) or a system on chip (System on Chip, SOC), etc., but is not limited to these.

[0159] In an embodiment of the present application, the communication component is configured to facilitate wired or wireless communication between the device in which it is located and other devices. The device in which the communication component is located can access a wireless network based on a communication standard, such as WiFi, 2G or 3G, 4G, 5G or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component can also be implemented based on near field communication (NFC) technology, radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology or other technologies.

[0160] In an embodiment of the present application, the display assembly may include a liquid crystal display (LCD) and a touch panel (TP). If the display assembly includes a touch panel, the display assembly may be implemented as a touch screen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touches, slides, and gestures on the touch panel. The touch sensors may not only sense the boundaries of a touch or slide action, but also detect the duration and pressure associated with the touch or slide action.

[0161] In embodiments of the present application, a power supply assembly is configured to provide power to various components of the device in which it is located. The power supply assembly may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device in which the power supply assembly is located.

[0162] In an embodiment of the present application, the audio component may be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC), and when the device where the audio component is located is in an operating mode, such as call mode, recording mode, and voice recognition mode, the microphone is configured to receive external audio signals. The received audio signal may be further stored in a memory or sent via a communication component. In some embodiments, the audio component further includes a speaker for outputting audio signals. For example, for a device with a language interaction function, voice interaction with the user may be achieved through the audio component.

[0163] It should be noted that the descriptions such as "first" and "second" in this article are used to distinguish different messages, devices, modules, etc., and do not represent the order of precedence, nor do they limit "first" and "second" to different types.

[0164] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0165] This application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate a device for implementing the functions specified in one or more processes in the flowchart and / or one or more blocks in the block diagram.

[0166] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0167] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0168] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0169] Memory may include non-permanent storage in a computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0170] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can be implemented using any method or technology to store information. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change RAM (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.

[0171] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0172] The foregoing is merely an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.

Claims

1. A target recognition method based on artificial intelligence, characterized in that: include: Acquire shelf images collected by the robot while it moves along the shelf; Inputting the shelf image into an object detection model to obtain spatial information of a first detection frame for object annotation of the shelf image, wherein the spatial information of the first detection frame includes: the center coordinates, size, and rotation angle of the first detection frame; or the vertex coordinates and rotation angle of the first detection frame; extracting a local image corresponding to the first detection frame from the shelf image according to the spatial information of the first detection frame; Inputting the partial image with the rotation angle into a target recognition model supporting multi-angle target recognition to recognize the target object contained in the partial image; The target recognition model includes: a first target recognition sub-model supporting multi-angle target recognition and a second target recognition sub-model supporting multi-angle target recognition; wherein the fineness of the image feature vector extracted by the second target recognition sub-model is greater than the fineness of the image feature vector extracted by the first target recognition sub-model; The step of inputting the partial image having the rotation angle into a target recognition model supporting multi-angle target recognition to recognize the target object contained in the partial image includes: Inputting the partial image with the rotation angle into the first target recognition sub-model to recognize the target object contained in the partial image; In the case that the first target recognition sub-model cannot recognize the target object contained in the partial image, the partial image with the rotation angle is input into the second target recognition sub-model to recognize the target object contained in the partial image.

2. The method according to claim 1, characterized in that The first target recognition sub-model includes: a first feature extraction layer supporting feature extraction of multi-angle images and a first target recognition layer supporting multi-angle target recognition; Inputting the partial image with the rotation angle into the first target recognition sub-model to recognize the target object contained in the partial image includes: Inputting the partial image with the rotation angle into the first feature extraction layer to obtain a first image feature vector of the partial image; The first image feature vector is input into the first target recognition layer to recognize the target object contained in the partial image.

3. The method according to claim 2, characterized in that Inputting the first image feature vector into the first target recognition layer to identify the target object contained in the partial image includes: In the first object recognition layer, calculating the similarity between the first image feature vector and the product feature vectors in the first product feature vector set; Selecting M product feature vectors from the first product feature vector set in descending order of similarity with the first image feature vector; Determining the product identifiers corresponding to the M product feature vectors based on the correspondence between the product feature vectors and product identifiers in the first product feature vector set; If Q ≥ N, the product identifier corresponding to the product feature vector with the greatest similarity is used as the identifier of the target object contained in the partial image; Wherein, Q is the number of product feature vectors among the M product feature vectors having the same product identifier as the product feature vector corresponding to the product feature vector with the greatest similarity; M ≥ 2, 1 ≤ N ≤ (M-1), and M and N are integers.

4. The method according to claim 3, characterized in that The second target recognition sub-model includes: a second feature extraction layer supporting feature extraction of multi-angle images and a second target recognition layer supporting multi-angle target recognition; wherein the fineness of the image feature vector extracted by the second feature extraction layer is greater than the fineness of the image feature vector extracted by the first feature extraction layer; The method further comprises: If Q<N, inputting the partial image with the rotation angle into the second feature extraction layer to obtain a second image feature vector of the partial image; the fineness of the second image feature vector is greater than the fineness of the first image feature vector; The second image feature vector is input into the second target recognition layer to identify the target object contained in the partial image.

5. The method according to claim 4, characterized in that Inputting the second image feature vector into the second target recognition layer to identify the target object contained in the partial image includes: In the second target recognition layer, based on the correspondence between the product feature vectors and the product identifiers in the second product feature vector set, the first target product feature vectors corresponding to the product identifiers corresponding to the M product feature vectors are obtained from the second product feature vector set; Calculating the similarity between the second image feature vector and the first target product feature vector; wherein the fineness of the product feature vectors in the second product feature vector set is greater than the fineness of the product feature vectors in the first product feature vector set; If a second target product feature vector exists in the first target product feature vector and the similarity between the second target product feature vector and the second image feature vector is greater than or equal to a set similarity threshold, the identity of the target object contained in the partial image is determined based on the correspondence between the product feature vectors and the product identity in the second product feature vector set and the second target product feature vector.

6. The method according to claim 5, characterized in that The determining the identifier of the target object contained in the partial image according to the correspondence between the product feature vectors and the product identifiers in the second product feature vector set and the second target product feature vector includes: Determining, from the second target product feature vectors, a product feature vector having the greatest similarity to the second image feature vector; According to the correspondence between the product feature vectors and product identifiers in the second product feature vector set, the product identifier corresponding to the product feature vector having the greatest similarity to the second image feature vector is determined as the identifier of the target object contained in the partial image.

7. The method according to claim 5, characterized in that Before inputting the partial image with the rotation angle into the first feature extraction layer, the method further includes: Obtain multi-angle images of various sample products; Inputting the multi-angle images of the sample product into the trained first feature extraction layer to obtain image feature vectors corresponding to the multi-angle images of the sample product, as product feature vectors in the first product feature vector set; constructing the first product feature vector set based on the product feature vectors in the first product feature vector set; A correspondence between each product feature vector in the first product feature vector set and a product identifier is established.

8. The method according to claim 7, characterized in that Before inputting the second image feature vector into the second object recognition sub-model, the method further includes: Inputting the multi-angle images of the sample product into the trained second feature extraction layer to obtain image feature vectors corresponding to the multi-angle images of the sample product, as product feature vectors in the second product feature vector set; constructing the second product feature vector set based on the product feature vectors in the second product feature vector set; A correspondence between each product feature vector in the second product feature vector set and the product identifier is established.

9. The method according to claim 7, characterized in that The multi-angle images of each sample product include a top-view image, a horizontal-view image, and a bottom-view image of the sample product.

10. The method according to any one of claims 1 to 9, characterized in that The step of acquiring shelf images collected by the robot while it moves along the shelf includes: The robot is controlled to move in a direction parallel to the shelf; and the robot is controlled to collect the shelf image during the movement.

11. The method according to any one of claims 1 to 9, characterized in that The number of the shelf images is multiple frames; the method further includes: Extract feature points of multiple frames of shelf images; performing deduplication processing on the multiple frames of shelf images according to the feature points of the multiple frames of shelf images; According to the result of deduplication processing on the multiple frames of shelf images, deduplication processing is performed on the target objects included in the multiple frames of shelf images.

12. A robot, characterized in that: include: A mechanical body; a camera, a memory, and a processor are installed on the mechanical body; The memory is used to store computer programs; The camera is used to capture shelf images when the robot moves along the shelf; The processor is coupled to the memory and is configured to execute the computer program to: input the shelf image into an object detection model to obtain spatial information of a first detection frame for object annotation of the shelf image, wherein the spatial information of the first detection frame includes: the center coordinates, size, and rotation angle of the first detection frame; or the vertex coordinates and rotation angle of the first detection frame; and extract a partial image corresponding to the first detection frame from the shelf image based on the spatial information of the first detection frame; Inputting the partial image with the rotation angle into a target recognition model supporting multi-angle target recognition to recognize the target object contained in the partial image; The target recognition model includes: a first target recognition sub-model supporting multi-angle target recognition and a second target recognition sub-model supporting multi-angle target recognition; wherein the fineness of the image feature vector extracted by the second target recognition sub-model is greater than the fineness of the image feature vector extracted by the first target recognition sub-model; The step of inputting the partial image having the rotation angle into a target recognition model supporting multi-angle target recognition to recognize the target object contained in the partial image includes: Inputting the partial image with the rotation angle into the first target recognition sub-model to recognize the target object contained in the partial image; In the case that the first target recognition sub-model cannot recognize the target object contained in the partial image, the partial image with the rotation angle is input into the second target recognition sub-model to recognize the target object contained in the partial image.

13. A computer device, characterized in that: include: memory and processor; The memory is used to store computer programs; The processor is coupled to the memory and configured to execute the computer program to perform the steps of the method according to any one of claims 1 to 11.

14. A computer-readable storage medium storing computer instructions, characterized in that: When the computer instructions are executed by one or more processors, the one or more processors are caused to perform the steps of the method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Commodity and tag automatic correlation method, system and equipment and storage medium

    CN108416403A

  • Vehicle image accurate retrieval method based on big data

    CN108491797A

  • Smart equipment and commodity counting method, device and equipment

    CN108846449A

  • System and method for product identification

    US20160155011A1