Method for counting the number of targets in a region, method for training a region recognition model, and device
By training the area identification model to automatically identify the target scene area and count the number of objects, the problem of inefficient boundaries of manual configuration area is solved, and efficient and accurate counting of the target object number is achieved to adapt to scene changes.
Patent Information
- Application Number
- CN202210454214.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-27
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-04-27
AI Technical Summary
In the prior art, the target object count statistics method requires manual configuration of regional boundaries, resulting in inefficiency and labor cost, and the existing methods cannot be automatically corrected when the area changes, resulting in statistical errors.
By obtaining the image to be identified, the trained area recognition model automatically recognizes the area of the target scene, divides the image blocks, and determines whether the target object is in the area by calculating the repetition ratio of the image blocks and the trajectory information, and counts its number.
It realizes that areas are not divided manually, improves the efficiency and accuracy of area identification in the target scenario, can adapt to scene changes, reduce computing resources, and improves the accuracy of the number of target objects statistics.
Smart Images

Figure CN114842416B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular, to a method for counting the number of targets in a region, a method for training a region recognition model, and an apparatus therefor. Background Art
[0002] Methods for counting the number of target objects (such as people, vehicles, etc.) in public places can be applied to multiple fields such as video surveillance, smart cities, and public safety, and have high practical value. At present, a large number of video surveillance devices are widely used in public places such as stations, museums, squares, banks, and supermarkets. Using the image information already collected by the video surveillance devices, effectively monitoring and analyzing the number of target objects in different regions of public places (such as the fresh food area of a supermarket, the ticket office of a station, the sidewalk at an intersection, etc.) is an indispensable part in the construction and management of public places.
[0003] However, in the method for counting the number of target objects, generally, the boundaries of each region in a scene are manually configured first, and then the number of recognized objects in each region is counted. In this way, it will consume a lot of manpower and affect the counting efficiency of the number of target objects. Summary of the Invention
[0004] The embodiments of this application provide a method for counting the number of targets in a region, a method for training a region recognition model, and an apparatus therefor, which can improve the accuracy of counting the number of people in a region.
[0005] In a first aspect, the embodiments of this application provide a method for counting the number of targets in a region. The specific method includes: obtaining an image to be recognized, where the image to be recognized is an image captured of a target scene; inputting the image to be recognized into a region recognition model to obtain feature information of the target scene, where the feature information is used to indicate one or more target regions of the target scene; dividing the image to be recognized into multiple image blocks; determining whether a first target object in the image to be recognized is in a target region according to the obtained multiple image blocks; and determining the number of target objects in the target region.
[0006] Based on the technical solution provided by this application, at least the following beneficial effects can be produced: Based on the trained region recognition model, this method first inputs the image to be recognized of the target scene into the region recognition model to automatically recognize one or more target regions of the target scene. In this way, there is no need to manually divide each region, saving a lot of manpower and improving the recognition efficiency of one or more regions in the target scene. Further, this method can sequentially recognize whether each target object in the image to be recognized is in a region of the target scene, so that it can accurately determine the region where each target object is located, and then count the number of target objects in each region, improving the accuracy of counting the number of target objects in the region.
[0007] In a possible implementation, determining whether a first target object in an image to be recognized is within a target area based on multiple image blocks includes: determining a repetition ratio of the number of identical image blocks among the image blocks included in the first target object in the image to be recognized and the image blocks included in the target area to the number of image blocks included in the first target object; and when the obtained repetition ratio is greater than or equal to a preset threshold, determining that the first target object is within the target area.
[0008] It can be understood that in some images to be recognized, if the first target object is within the target area, then all the image blocks included in the first target object should belong to the image blocks included in the above-mentioned target area. However, in some other images to be recognized, when the first target object is within the target area, some of the image blocks included in the first target object belong to the image blocks included in the above-mentioned target area, and some other image blocks included in the first target object may not belong to the image blocks included in the above-mentioned target area, but the proportion of the number of image blocks belonging to the image blocks included in the above-mentioned target area to the total number of image blocks included in the first target object is usually greater than or equal to the preset threshold. Therefore, in the above implementation, the method can determine whether the first target object is within the target area by calculating the above-mentioned repetition ratio.
[0009] In another possible implementation, before determining the repetition ratio of the number of identical image blocks among the image blocks included in the first target object in the image to be recognized and the image blocks included in the target area to the number of image blocks included in the first target object, the above method further includes: determining the image blocks included in the first target object and the image blocks included in the target area in the image to be recognized.
[0010] In yet another possible implementation, determining the image blocks included in the first target object and the image blocks included in the target area in the image to be recognized specifically includes: for any one of the multiple image blocks, when the center point of any one image block is within the recognition range of the first target object, determining that the first target object in the image to be recognized includes any one image block; and / or when the center point of any one image block is within the recognition range of the target area, determining that the target area in the image to be recognized includes any one image block.
[0011] In yet another possible implementation, the method further includes: when the above-mentioned repetition ratio is less than the preset threshold, determining that the first target object is not within the target area.
[0012] In yet another possible implementation, the above-mentioned feature information further includes one or more boundary lines of the target scene, and the method further includes: obtaining trajectory information of a first target object in a preset time period, where the preset time period includes a first moment when the first target object enters the target area; determining a connection line between a trajectory point at the first moment and a trajectory point at a second moment in the trajectory information of the preset time period, the second moment being before the first moment and belonging to the preset time period; if the connection line has an intersection with one or more boundary lines, determining that the first target object is not within the target area.
[0013] It can be understood that in the image to be recognized, if there are transparent and non-directly crossable obstacles (such as glass walls, transparent shelves, etc.) in the target scene, and the first target object (such as a person) is near the glass wall, whether the person is inside or outside the glass wall, the image information of the person is included in the image to be recognized. However, one or more areas in the target scene are usually divided according to the intersection line of the obstacle and the ground (i.e., the above-mentioned interface), so the inside and outside of the glass wall belong to different areas. Therefore, when the above-mentioned repetition ratio is greater than or equal to the preset threshold, the statistical device can determine that the person is near the glass wall. Further, the statistical device can determine whether the connection line between the trajectory point at the first moment and the trajectory point at the second moment of the person has an intersection with one or more dividing lines of the target scene according to the trajectory information of the person. Since the glass wall cannot be directly crossed, if there is an intersection, it is determined that the person is not within the wall area.
[0014] In a second aspect, an embodiment of the present application further provides a method for training a region recognition model, the method including: obtaining a training sample set, where the training sample set includes one or more images labeled with feature information of a target scene, and the feature information includes one or more regions and one or more boundary lines of the target scene; training an initial model according to the obtained training sample set to obtain a region recognition model, where the region recognition model is used for region recognition of an image of the target scene.
[0015] It can be understood that for a target scene that requires region recognition, the training device uses one or more images labeled with feature information of the target scene as a training sample set to train an initial model, so as to obtain a region recognition model that can be used for region recognition of the target scene. In this way, using the trained region recognition model, automatic recognition of one or more regions of the target scene can be achieved.
[0016] In a possible implementation, the above-mentioned obtaining of the training sample set specifically includes: obtaining multiple images of the target scene; for the first image taken at the first shooting angle among the multiple images, converting the first image into an image with a top-down shooting angle, where the first shooting angle is a horizontal shooting or an upward shooting; then obtaining the feature information labeled on the image with a top-down shooting angle; converting the first image labeled with the feature information back into the first image taken at the first shooting angle; thus, determining the first image with the first perspective shooting angle and labeled with the feature information as the image in the training sample set.
[0017] It can be understood that in the embodiments of the present application, the shooting angle of the image to be recognized may include an upward shooting, a horizontal shooting, or a top-down shooting. Generally, for a target scene (such as a restaurant), usually one or more areas are divided according to the specific item placement positions (such as the table and chair placement positions in the dining area) and wall positions in the scene, that is, usually one or more areas are divided according to the floor plan layout of the target scene. Therefore, in the image to be recognized with the first shooting angle, the actual placement positions of the items in the target area may not be captured, and it is not convenient to label one or more areas of the target scene on this image. Thus, in this implementation, the image can be transformed, converting the first image into a first image with a top-down shooting angle, and then obtaining the labeled one or more areas to obtain more accurate feature information of the target scene. Furthermore, determining the first image with the first perspective shooting angle and labeled with the feature information as the image in the training sample set can improve the recognition accuracy of the obtained area recognition model.
[0018] In another possible implementation, the above method further includes: when the area included in the target scene changes, obtaining one or more images labeled with the feature information of the changed target scene; retraining the area recognition model according to the one or more images labeled with the feature information of the changed target scene to obtain a retrained area recognition model.
[0019] In a third aspect, the present application provides a device for counting the number of targets within an area. The device includes: a transceiver module for obtaining an image to be recognized, where the image to be recognized is an image taken of the target scene; an identification module for inputting the image to be recognized into the area recognition model to obtain the feature information of the target scene, where the feature information is used to indicate one or more target areas of the target scene; a processing module for dividing the image to be recognized into multiple image blocks; the processing module is further used to determine whether the first target object in the image to be recognized is within the target area according to the multiple image blocks; and determine the number of target objects within the target area.
[0020] In a possible implementation, the above-mentioned processing module is specifically configured to: determine the repetition ratio of the number of identical image patches between the image patches included in the first target object and the image patches included in the target area in the image to be recognized, and the number of image patches included in the first target object; when the repetition ratio is greater than or equal to a preset threshold, determine that the first target object is within the target area.
[0021] In another possible implementation, the above-mentioned processing module is further configured to determine the image patches included in the first target object and the image patches included in the target area in the image to be recognized.
[0022] In yet another possible implementation, the above-mentioned processing module is specifically configured to: for any one of a plurality of image patches, when the center point of any one of the image patches is within the recognition range of the first target object, determine that the first target object in the image to be recognized includes any one of the image patches; and / or, when the center point of any one of the image patches is within the recognition range of the target area, determine that the target area in the image to be recognized includes any one of the image patches.
[0023] In yet another possible implementation, the above-mentioned processing module is further configured to determine that the first target object is not within the target area when the repetition ratio is less than the preset threshold.
[0024] In yet another possible implementation, the above-mentioned feature information further includes one or more boundary lines of the target scene, and the transceiver module is further configured to obtain the trajectory information of the first target object in a preset time period, the preset time period includes the first moment when the first target object enters the target area, and the processing module is further configured to determine the connection line between the trajectory point at the first moment and the trajectory point at the second moment in the trajectory information of the preset time period, the second moment is before the first moment, and the second moment belongs to the preset time period; and is used to determine that the first target object is not within the target area when the connection line intersects with one or more boundary lines.
[0025] In a fourth aspect, an embodiment of the present application provides a training device for a region recognition model, and the device includes: a transceiver module, configured to obtain a training sample set, the training sample set includes one or more images labeled with feature information of a target scene, where the feature information includes one or more regions and one or more boundary lines of the target scene; a training module, configured to train an initial model according to the training sample set to obtain a region recognition model, and the region recognition model is used to perform region recognition on an image of the target scene.
[0026] In a possible implementation, the device further includes a processing module. The transceiver module is specifically configured to: obtain multiple images of a target scene; for a first image captured at a first shooting angle among the multiple images, the processing module is configured to convert the first image into a first image with a shooting angle of top-down shooting, where the first shooting angle is horizontal shooting or upward shooting; the transceiver module is further specifically configured to obtain feature information annotated on the image with a shooting angle of top-down shooting; the processing module is further configured to convert the first image annotated with the feature information into a first image captured at the first shooting angle; the processing module is further configured to determine the first image with a shooting angle of the first perspective and annotated with the feature information as an image in the training sample set.
[0027] In another possible implementation, the transceiver module is further configured to obtain one or more images annotated with the feature information of the changed target scene when the area included in the target scene changes; the training module is further configured to retrain the area recognition model according to the one or more images annotated with the feature information of the changed target scene to obtain a retrained area recognition model.
[0028] In a fifth aspect, the present application provides an electronic device, which includes: one or more processors; one or more memories; wherein, the one or more memories are used to store computer program code, and the computer program code includes computer instructions. When the one or more processors execute the computer instructions, the electronic device executes the method provided in the first aspect as described above, or the method provided in the second aspect as described above.
[0029] In a sixth aspect, the present application provides a chip system, which is applied to an electronic device; the chip system includes one or more interface circuits and one or more processors. The interface circuits and the processors are interconnected by lines; the interface circuits are configured to receive signals from the memory of the electronic device and send the signals to the processors, and the signals include the computer instructions stored in the memory. When the processors execute the computer instructions, the electronic device executes the method provided in the first aspect as described above, or the method provided in the second aspect as described above.
[0030] In a seventh aspect, the present application provides a computer-readable storage medium, which stores computer instructions. When the computer instructions are run on an electronic device, the electronic device executes the method provided in the first aspect as described above, or the method provided in the second aspect as described above.
[0031] In an eighth aspect, the present application provides a computer program product, which includes computer instructions. When the computer instructions are run on an electronic device, the electronic device executes the method provided in the first aspect as described above, or the method provided in the second aspect as described above.
[0032] For the specific descriptions of the third to eighth aspects and their various implementation manners in this application, reference may be made to the detailed descriptions in the first aspect or the second aspect and their various implementation manners; moreover, for the beneficial effects of the third to eighth aspects and their various implementation manners, reference may be made to the analysis of the beneficial effects in the first aspect or the second aspect and their various implementation manners, which will not be elaborated here. Description of the Drawings
[0033] Figure 1 It is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of this application;
[0034] Figure 2 It is a flowchart of a method for counting the number of targets in a region provided by an embodiment of this application;
[0035] Figure 3 It is a schematic diagram of a target scenario provided by an embodiment of this application;
[0036] Figure 4 It is a schematic diagram of trajectory information provided by an embodiment of this application;
[0037] Figure 5 It is a flowchart of another method for counting the number of targets in a region provided by an embodiment of this application;
[0038] Figure 6 It is a schematic diagram of another target scenario provided by an embodiment of this application;
[0039] Figure 7 It is a flowchart of a method for training a region recognition model provided by an embodiment of this application;
[0040] Figure 8 It is a schematic diagram of the composition of a statistical device provided by an embodiment of this application;
[0041] Figure 9 It is a schematic diagram of the composition of a training device provided by an embodiment of this application. Detailed Embodiments
[0042] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Apparently, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.
[0043] In the description of this application, unless otherwise specified, " / " means "or". For example, A / B can mean A or B. "And / or" in this text is merely a correlative relationship describing related objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, "at least one" means one or more, and "a plurality" means two or more. The terms "first", "second", etc. do not limit the quantity and execution order, and the terms "first", "second", etc. do not necessarily limit to being different.
[0044] It should be noted that in this application, words such as "exemplary" or "for example" are used to give examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in this application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.
[0045] Currently, the application of the method for counting the number of targets in a region is also becoming more and more extensive. For example, it can count the number of people in public places such as stations, museums, squares, banks, and supermarkets. Taking a supermarket as an example, when arranging the display of supermarket goods, specific adjustments can be made to the goods according to the statistical results of the number of people in different regions, and the shelf layout can also be adjusted according to the statistical results of the number of people in different regions.
[0046] Currently, for counting the number of people in each region, population heat data can be obtained periodically, and the population heat data can be matched with the dot matrix data table pre-constructed according to the unit region to obtain the unit region to which each piece of population heat data belongs. On the one hand, the population heat data in this method is determined according to the positioning information of the client provided by the location based service (LBS) network service provider. However, the number of clients in a region cannot fully identify the number of people in that region, and there may still be people without a client in that region. Moreover, the error of satellite positioning is relatively large and is not suitable for counting the number of people in areas with a relatively small area such as the fresh food area of a supermarket or the order-taking area of a coffee shop. On the other hand, with the adjustment or change of the layout in the scene, the size of a region may also change, but this method cannot automatically correct the boundary of the unit region in the pre-constructed dot matrix data table, which may lead to incorrect statistics of the number of people in the region.
[0047] In view of this, an embodiment of the present application provides a method for counting the number of regional targets. The specific method includes: obtaining an image to be recognized, where the image to be recognized is an image captured of a target scene; inputting the image to be recognized into a regional recognition model to obtain feature information of the target scene; the feature information is used to indicate one or more target regions of the target scene; dividing the image to be recognized into multiple image blocks; determining whether a first target object in the image to be recognized is within the target region according to the obtained multiple image blocks; and determining the number of target objects in the target region. This method can automatically recognize one or more target regions of a target scene based on a trained regional recognition model. In this way, the efficiency of recognizing one or more regions in the target scene can be improved.
[0048] An embodiment of the present application also provides a device for counting target objects within a region. The device for counting target objects within a region can be used to execute the above method for counting the number of people in a region. Optionally, the device for counting target objects within a region can be an electronic device with data processing capabilities, or a functional module in the electronic device, and this is not limited. For example, the electronic device can be a server, which can be a single server, or alternatively, a server cluster composed of multiple servers. For another example, the electronic device can be a mobile phone, a tablet computer, a desktop computer, a laptop computer, a handheld computer, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, and terminal devices such as a cellular phone, a personal digital assistant (PDA), an augmented reality (AR) / virtual reality (VR) device, etc. For another example, the electronic device can also be a video recording device, a video surveillance device, etc., which can be used to capture an image to be recognized of a target scene. The present disclosure does not impose special restrictions on the specific form of the electronic device.
[0049] Optionally, taking the example of counting the number of people in a supermarket area, a trained regional recognition model applicable to the supermarket can be pre-stored in the device for counting target objects within a region. A user can input an image to be recognized of the supermarket into the device for counting target objects within a region, and the device for counting target objects within a region can recognize one or more regions of the supermarket through the regional recognition model and further determine the number of people in each region.
[0050] Next, taking the device for counting target objects within a region as an electronic device as an example, as Figure 1 shown, Figure 1 a hardware structure of an electronic device 100 is shown.
[0051] As Figure 1As shown, the electronic device 100 includes a processor 110, a communication line 120, and a communication interface 130.
[0052] Optionally, the electronic device 100 may further include a memory 140. Among them, the processor 110, the memory 140, and the communication interface 130 may be connected through the communication line 120.
[0053] Among them, the processor 110 may be a central processing unit (CPU), a general-purpose processor, a network processor (NP), a digital signal processor (DSP), a microprocessor, a microcontroller, a programmable logic device (PLD), or any combination thereof. The processor 101 may also be any other device with processing functions, such as a circuit, a device, or a software module, without limitation.
[0054] In one example, the processor 110 may include one or more CPUs, such as Figure 1 CPU0 and CPU1 in
[0055] As an optional implementation, the electronic device 100 includes multiple processors. For example, in addition to the processor 110, it may further include a processor 170. The communication line 120 is used to transmit information between the components included in the electronic device 100.
[0056] The communication interface 130 is used to communicate with other devices or other communication networks. The other communication network may be an Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc. The communication interface 130 may be a module, a circuit, a transceiver, or any device capable of implementing communication.
[0057] The memory 140 is used to store instructions. Among them, the instructions may be computer programs.
[0058] Among them, the memory 140 can be a read-only memory (ROM) or other types of static storage devices that can store static information and / or instructions. It can also be a random access memory (RAM) or other types of dynamic storage devices that can store information and / or instructions. It can also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, without limitation.
[0059] It should be noted that the memory 140 can exist independently of the processor 110 or be integrated with the processor 110. The memory 140 can be used to store instructions, program codes, or some data, etc. The memory 140 can be located inside the electronic device 100 or outside the electronic device 100, without limitation.
[0060] The processor 110 is configured to execute the instructions stored in the memory 140 to implement the communication method provided in the following embodiments of the present application. For example, when the electronic device 100 is a terminal or a chip in the terminal, the processor 110 can execute the instructions stored in the memory 140 to implement the steps performed by the sending end in the following embodiments of the present application.
[0061] As an optional implementation, the electronic device 100 further includes an output device 150 and an input device 160. Among them, the output device 150 can be a display screen, a speaker, or other devices that can output the data of the electronic device 100 to the user. The input device 160 can be a keyboard, a mouse, a microphone, or a joystick, or other devices that can input data to the electronic device 100.
[0062] It should be noted that Figure 1 the structure shown in Figure 1 does not constitute a limitation on the computing device. In addition to
[0063] In addition, for the above-mentioned area recognition model, the embodiments of the present application further provide a training device for the area recognition model (hereinafter simply referred to as the training device for simplicity). This training device can be used to train the area recognition model and provide a trained area recognition model applicable to a specific scenario (such as supermarket A). Moreover, this training device can also be an electronic device module with data processing capabilities or a functional module in the electronic device, and no limitation is made thereto. Optionally, the hardware structure of the training device can also be as Figure 1 shown.
[0064] In some embodiments, the above-mentioned target object statistics device in the area and the above-mentioned training device can be integrated into one device; alternatively, the above-mentioned target object statistics device in the area and the above-mentioned training device can be two independent devices.
[0065] Next, the embodiments provided by the present application will be specifically introduced with reference to the accompanying drawings of the specification.
[0066] The embodiments of the present application provide a method for counting the number of people in an area, as Figure 2 shown. This method is applied to a target object statistics device in the area with the hardware structure as Figure 1 shown (hereinafter simply referred to as the statistics device for simplicity), and specifically includes the following steps:
[0067] S101. The statistics device acquires an image to be recognized.
[0068] Among them, the above-mentioned image to be recognized is an image taken of the target scene. The target scene is the scene where the number of target objects in the area needs to be counted. Optionally, the target scene can be a public place such as supermarket A, station B, coffee shop C, etc.
[0069] Therefore, the above-mentioned image to be recognized is one or more images of the target scene. Moreover, to improve the accuracy of counting the number of people in the area, the above-mentioned image to be recognized can include images of multiple shooting angles of the target scene. In an actual scene, an area can be determined by the specific layout in the scene such as the placement position of the shelves, the setting position of the cashier desk, and the placement position of the tables and chairs. Therefore, the above-mentioned image to be recognized is an image that can show the overall decoration and layout of the specific scene.
[0070] In addition, in the embodiments of the present application, the above-mentioned shooting angles include upward shooting, horizontal shooting, and downward shooting. Among them, upward shooting means that the shooting device shoots from a low position upward, and during the shooting process, the height of the shooting device is lower than the object to be photographed. Horizontal shooting means that the shooting device and the object to be photographed are at the same horizontal line. Downward shooting means that the shooting device shoots from a high position downward, and during the shooting process, the height of the shooting device is higher than the object to be photographed.
[0071] It should be noted that, under normal circumstances, for a target scene (such as a restaurant), one or more areas are usually divided according to the specific item placement positions (such as the placement positions of tables and chairs in the dining area) and wall positions in the scene, that is, one or more areas are usually divided according to the floor plan of the target scene. Therefore, in the to-be-recognized image taken from a flat or upward shooting angle, the actual placement positions of items in the target area may not be captured. Then, the to-be-recognized image taken from a downward shooting angle can clearly show the overall decoration and layout of the target scene, that is, the to-be-recognized image taken from a downward shooting angle can clearly show the area division of the target scene.
[0072] S102. The statistical device inputs the to-be-recognized image into the area recognition model to obtain the feature information of the target scene.
[0073] Among them, the above area recognition model is a pre-trained recognition model that can be used to recognize one or more areas included in the target scene.
[0074] In addition, the above feature information includes one or more target areas of the above target scene. Optionally, the above feature information may also include one or more dividing lines in the above target scene.
[0075] Among them, the above boundary line refers to the intersection line of an obstacle that cannot be directly crossed and the ground. Exemplarily, the obstacle may be a wall, a cabinet, a shelf, etc. in the target scene.
[0076] It should be noted that one or more areas of the target scene can be determined according to the specific layout in the scene such as the wall position, the placement position of the shelf, the setting position of the cashier desk, the placement position of the tables and chairs, etc. Thus, the above boundary line can also be used as the dividing line of one or more areas.
[0077] In addition, during the process of recognizing the number of people in the target area, if there are transparent obstacles that cannot be directly crossed (such as glass walls, transparent shelves, etc.) in the target scene, and the first target object (such as a person) is near the glass wall, whether the person is inside or outside the glass wall, the image information of the person is included in the to-be-recognized image. However, the inside and outside of the glass wall may belong to different areas, or the outside of the glass wall may not belong to the target scene. Therefore, the above feature information may also include one or more dividing lines in the above target scene. In this way, when the target scene has transparent obstacles that cannot be directly crossed, it is convenient to more accurately determine the area where the target object is located according to the dividing line. Exemplarily, as Figure 3 shown, Figure 3 is a to-be-recognized image of coffee shop X including feature information. Among them, the feature information of this image includes Figure 3Regions 11, 12, 13, 14, 15, and 16 therein. The feature information of the image includes Figure 3 Boundary lines 21, 22, and 23 therein.
[0078] In one implementation, the statistical device can identify the feature information of the target scene at a first preset frequency. Furthermore, based on the identified feature information, the number of regional target objects in the target scene can be counted multiple times.
[0079] Among them, the above-mentioned first preset frequency can be once a week, once a month, etc.
[0080] It should be noted that when the overall decoration and layout of the target scene remain unchanged, the regional division of the target scene usually does not change either. Therefore, the statistical device can obtain the image to be recognized of the target scene at the first preset frequency to reduce the frequency of regional recognition and save computing resources.
[0081] In another implementation, the statistical device can obtain the image to be recognized of the target scene each time the number of regional target objects in the target scene needs to be counted, and then identify the feature information of the target scene based on the image to be recognized.
[0082] It should be noted that when the statistical device obtains the image to be recognized of the target scene each time the number of regional target objects needs to be counted, relatively accurate feature information can be immediately recognized.
[0083] S103. The statistical device divides the image to be recognized into multiple image blocks.
[0084] Specifically, the statistical device can evenly divide the image to be recognized into multiple image blocks. Among them, an image block can be square, rectangular, or other possible shapes, and this application does not limit this.
[0085] Exemplarily, the above-mentioned image blocks can be obtained by performing a mesh segmentation on an image in the form of X rows × Y columns, and then X × Y image blocks of the complete image can be obtained. An image block can include one or more pixel points of the image to be recognized. Exemplarily, the statistical device can perform a mesh segmentation on the image according to the resolution or size of the obtained image. Exemplarily, the statistical device can divide an image to be recognized into 1984 image blocks according to 32 rows × 62 columns.
[0086] In addition, the above-mentioned pixel refers to the smallest unit in the image, and the pixel point can be regarded as an indivisible unit or element in the entire image. The complete image is composed of multiple pixel points, and the colors and positions of all the pixel points of the image can determine the overall pattern and size presented by the image.
[0087] S104. The statistical device determines whether a first target object in the image to be recognized is within a target area based on a plurality of image blocks.
[0088] Wherein, the above-mentioned target area is any area in the target scene. The above-mentioned first target object can be any target object in the target scene, or the above-mentioned first target object can be any target object in the image to be recognized that has the same image blocks as the target area in terms of their positions. Or, the above-mentioned first target object can be a target object with trajectory information close to the target area.
[0089] Exemplarily, taking the first target object as a person, when the person appears within the field of view of a monitoring device installed in the target scene, the monitoring device can track the person and record the position of the person at a second preset frequency. For example, the trajectory information of the person within a period of time can be as Figure 4 shown in (a) of, which is the person images of the person at multiple moments within that period of time. Among them, the person image 31 can be the trajectory information of the person at one moment. Also for example, the trajectory information of the person within a period of time can be as Figure 4 shown in (b) of, which is multiple trajectory points of the person at multiple moments within that period of time. Among them, the trajectory point 32 can be the trajectory information of the person at one moment within a preset time period. Optionally, the trajectory point 32 can be the centroid position of the person or the foot position of the person, etc. at that moment. Again for example, the trajectory information of the person within a period of time can be multiple images obtained by the monitoring device when photographing the person at multiple moments within that period of time, that is, the multiple images can be multiple above-mentioned images to be recognized.
[0090] In some embodiments, the statistical device can determine whether the first target object is within the target area based on the trajectory information of the first target object at the current moment.
[0091] Optionally, taking the trajectory information shown in (b) of Figure 4 as an example, the statistical device can determine whether the trajectory point 32 is within the target area based on one or more target areas of the recognized target scene.
[0092] Optionally, as shown in Figure 5 , the statistical device can perform the following steps S1041 to S1043 to determine whether the first target object is within the target area based on the multiple image blocks in the image to be recognized:
[0093] S1041. The statistical device determines the repetition ratio of the number of identical image blocks between the image blocks included in the first target object in the image to be recognized and the image blocks included in the target area to the number of image blocks included in the first target object.
[0094] In some embodiments, the statistical device may first determine the image blocks included in the first target object in the image to be recognized and the image blocks included in the target area, and then calculate the above-mentioned repetition ratio.
[0095] It should be understood that for any one of the multiple image blocks in the image to be recognized, the image block may be an image block included in the first target object, or the image block may be an image block included in the target area, or the image block is both an image block included in the first target object and an image block included in the target area. Or, the image block neither belongs to the image blocks included in the first target object nor belongs to the image blocks included in the target area
[0096] Optionally, for any one of the multiple image blocks in the image to be recognized, when the center point of any one image block is within the recognition range of the first target object, the statistical device may determine that the first target object in the image to be recognized includes any one image block; and / or when the center point of any one image block is within the recognition range of the target area, it is determined that the target area in the image to be recognized includes any one image block.
[0097] Exemplarily, as Figure 6 shown, taking the first target object as the person 40 as an example, the recognition range of the above-mentioned first target object may be the range included in the person recognition frame 41 of the person 40 in the image to be recognized, and all the image blocks with the center point in the person recognition frame 41 are the image blocks included in the above-mentioned first target object. The recognition range of the above-mentioned target area may be the range included in the area recognition frame 17 in the image to be recognized, and all the image blocks with the center point in the area recognition frame 17 are the image blocks included in the above-mentioned target area.
[0098] Furthermore, the statistical device may determine the number of image blocks included in the first target object in the image to be recognized and the number of identical image blocks among the image blocks included in the first target object and the image blocks included in the target area in the image to be recognized according to the image blocks included in the first target object and the image blocks included in the target area in the image to be recognized.
[0099] In addition, the above-mentioned repetition ratio is the ratio between the number of identical image blocks among the image blocks included in the first target object and the image blocks included in the target area in the image to be recognized and the number of image blocks included in the first target object in the image to be recognized.
[0100] S1042. When the repetition ratio is greater than or equal to the preset threshold, the statistical device determines that the first target object is within the target area.
[0101] Wherein, the above-mentioned preset threshold may be 50% or other reasonable ratio values.
[0102] It can be understood that in the target image to be recognized, for example, in the image to be recognized with a top-down shooting angle, if the first target object is within the target area, all the image blocks included in the first target object should belong to the image blocks included in the above-mentioned target area. However, in some other images to be recognized, for example, in the image to be recognized with a flat shooting angle, when the first target object is within the target area, some of the image blocks included in the first target object belong to the image blocks included in the above-mentioned target area, and some other image blocks included in the first target object may not belong to the image blocks included in the above-mentioned target area, but the proportion of the number of image blocks belonging to the image blocks included in the above-mentioned target area to the total number of image blocks included in the first target object is usually greater than or equal to a preset threshold. Therefore, the statistical device can determine whether the first target object is within the target area by calculating the above-mentioned repetition ratio.
[0103] Among them, the above-mentioned top-down shooting is a type of top-down shooting. During the shooting process with a top-down shooting angle, the height of the camera is higher than the height of the object to be photographed, and the shooting direction of the camera is perpendicular to the horizontal plane.
[0104] In some embodiments, when the above-mentioned repetition ratio is greater than or equal to the preset threshold, the statistical device can obtain the trajectory information of the first target object in a preset time period, determine the connection line between the trajectory point at the first moment and the trajectory point at the second moment in the trajectory information of the preset time period. If the above-mentioned connection line has no intersection with one or more boundary lines, it is determined that the first target object is within the target area.
[0105] Among them, the above-mentioned preset time period includes the first moment when the first target object enters the target area. And the duration of the above-mentioned preset time period can be a relatively short duration such as 10 seconds or 5 seconds. Exemplarily, the above-mentioned preset time period can be a time period with the end moment being the first moment, or a time period with the middle moment being the first moment.
[0106] And the above-mentioned second moment is before the first moment and belongs to the above-mentioned preset time period.
[0107] It can be understood that taking the first target object as a person as an example, since under normal circumstances, a person cannot directly pass through a solid wall, the movement trajectory of a person has no intersection with one or more boundary lines. Therefore, if the above-mentioned connection line has an intersection with one or more boundary lines, it indicates that the first target object is not within the target area.
[0108] It should be noted that in the image to be recognized, if there are transparent and non-directly crossable obstacles (such as glass walls, transparent shelves, etc.) in the target scene, and the first target object (such as a person) is near the glass wall, whether the person is inside or outside the glass wall, the image information of the person is included in the image to be recognized. However, one or more regions in the target scene are usually divided according to the intersection line of the obstacle and the ground (i.e., the above-mentioned interface), so the inside and outside of the glass wall belong to different regions. Therefore, when the above-mentioned repetition ratio is greater than or equal to the preset threshold, the statistical device can determine that the person is near the glass wall. Further, the statistical device can determine whether there are intersections between the line connecting the trajectory points at the first moment and the trajectory points at the second moment and one or more boundary lines of the target scene according to the trajectory information of the person. Since the glass wall cannot be directly crossed, if there are intersections, it is determined that the person is not in the walled area. For example, if the outer wall of restaurant E is a glass wall, the inside of the glass wall is the dining area of restaurant E, and the outside of the glass wall belongs to the outside of restaurant E. However, in the image to be recognized, the repetition ratio of the image information of the first target object outside the glass wall and the image information of the dining area of restaurant E may be greater than the preset threshold.
[0109] Thus, when the above-mentioned repetition ratio is greater than or equal to the preset threshold, the statistical device can further determine whether there are intersections between the line connecting the trajectory points at the first moment and the trajectory points at the second moment and one or more boundary lines of the target scene, and use this to determine whether the first target object is in the target area.
[0110] S1043. When the repetition ratio is less than the preset threshold, the statistical device determines that the first target object is not in the target area.
[0111] S105. The statistical device determines the number of target objects in the target area. Optionally, for all target objects in the target scene, the statistical device can sequentially determine the area where each target object is located, so as to obtain the number of target objects included in the target area in the image.
[0112] Based on the technical solution provided by the present application, at least the following beneficial effects can be produced: Based on the trained area recognition model, the method first inputs the image to be recognized of the target scene into the area recognition model, and automatically recognizes one or more target areas of the target scene. In this way, there is no need to manually divide each area, saving a large amount of manpower and improving the efficiency of recognizing one or more areas in the target scene. Further, the method can sequentially recognize whether each target object in the image to be recognized is in an area of the target scene, so as to accurately determine the area where each target object is located, and then count the number of target objects in each area, improving the accuracy of counting the number of target objects in the area.
[0113] In some embodiments, the embodiments of the present application further provide a method for training a region recognition model, which is applied to the above training device, as Figure 7 shown. The method includes:
[0114] S201. The training device obtains a training sample set.
[0115] Among them, the above training sample set includes multiple images of a target scene.
[0116] Optionally, the above target scene is a scene that needs to perform region recognition or a same-type scene with a layout similar to the scene that needs to perform region recognition. For example, if the scene that needs to perform region recognition is supermarket A, then the above target scene can be supermarket A and the same-type supermarket A1 with a layout similar to supermarket A.
[0117] It should be noted that in actual implementation, supermarket A and supermarket A1 can be chain supermarkets of the same brand opened at different locations. In this way, the layouts of supermarket A and supermarket A1 are similar, and the region divisions of supermarket A and supermarket A1 are also similar. Therefore, in order to collect a large number of samples to train the region recognition model, the training sample set can also include multiple images of supermarket A1.
[0118] Optionally, the multiple images of the above target scene can be images from multiple perspectives of the target scene.
[0119] Among them, the images in the above training sample set are labeled with target scene feature information. The feature information includes one or more regions of the target scene. Exemplarily, one or more regions labeled on each image can be as Figure 3 shown by region 12, region 13, region 14, region 15, and region 16 in.
[0120] And, the above feature information further includes one or more boundary lines of the target scene. The boundary line refers to the intersection line of an obstacle that cannot be directly crossed and the ground. Exemplarily, one or more boundary lines labeled on each image can be as Figure 3 shown by demarcation line 21, demarcation line 22, and demarcation line 23 in.
[0121] In some embodiments, the training device can directly obtain multiple images labeled with target scene feature information and determine the obtained multiple images as the above training sample set.
[0122] Alternatively, the training device can first obtain multiple images of the target scene without labeled feature information, then display the obtained multiple images of the target scene to the user, and then receive the multiple images of the target scene labeled with feature information by the user to obtain the training sample set.
[0123] Optionally, when the training device obtains multiple images of a target scene without feature information marked thereon, the training device may display the multiple images of the target scene without feature information marked thereon to the user through a built-in or external display device, and obtain the feature information marked by the user on each image through a built-in or external input device such as a mouse, a keyboard, or a touch device on the display device, so as to obtain a training sample set.
[0124] Alternatively, the training device may send multiple images of the target scene without feature information marked thereon to the user's terminal device. Thus, the user's terminal device may obtain the feature information marked by the user on each image, and send multiple images of the target scene with feature information marked thereon to the training device. Correspondingly, the training device may receive the multiple images and obtain a training sample set.
[0125] In some embodiments, when the training device obtains multiple images of a target scene without feature information marked thereon, the training device may determine a first image taken at a first shooting angle among the multiple images, and convert the first image into an image with a shooting angle of top-down shooting, where the first shooting angle is a flat shooting or an upward shooting.
[0126] Wherein, the shooting angle may refer to the relevant description in step S101 above, and will not be elaborated herein.
[0127] Optionally, the training device may convert the image at the first shooting angle into an image with a top-down shooting angle by means of projective transformation.
[0128] Among them, projective transformation, also known as perspective transformation (Perspective Transformation), refers to a transformation that projects an image from its original plane to a new plane (such as a horizontal plane) by using the condition that the perspective center, the image point, and the target point are collinear, and rotating the projection plane (perspective plane) around the trace line (perspective axis) by a certain angle according to the perspective rotation law.
[0129] It should be noted that the top-down image can intuitively display the horizontal layout of the target scene. In this way, it is convenient for the user to determine the boundary of a region of the target scene on the horizontal plane and the boundary line of the target scene. In addition, the user can more accurately judge the demarcation of a region of the target scene on the horizontal plane and the position of the above boundary line. Therefore, compared with the image at the first shooting angle, based on the top-down image, one or more regions and one or more boundary lines of the target scene marked by the user are more accurate. Thus, the training device trains the region recognition model accordingly, and a region recognition model with higher recognition accuracy can be obtained.
[0130] Further, the training device can obtain the feature information annotated on the top-down image of the shooting angle, and convert the top-down image with the annotated feature information into an image with the first shooting angle. Thus, the training device can determine the first image with the first perspective and annotated with the feature information as the image in the training sample set.
[0131] S202. The training device trains the initial model according to the above training sample set to obtain a region recognition model.
[0132] Among them, the above region recognition model is used to perform region recognition on the image of the target scene.
[0133] In some embodiments, the process of training the region recognition model to be trained according to the training samples includes:
[0134] S1. The training device obtains the initial model to be trained and the training sample set.
[0135] S2. The training device inputs the first image in the training sample set into the initial model to be trained to obtain the feature information of the first image.
[0136] S3. The training device compares the feature information of the first image output by the initial model with the feature information annotated on the first image to determine the image loss value.
[0137] S4. The training device adjusts the model parameters of the initial mode according to the above loss value.
[0138] S5. The training device uses another image in the training sample set as the new first image and repeats S1 - S5 until the model converges.
[0139] Among them, this another image can be the above first image, or any image in the training sample set other than the above first image, and there is no limitation on this.
[0140] Among them, the training device can determine whether the model converges based on the loss value output each time during the initial model process, or determine whether the mode converges based on the number of training times. There is no limitation on this in the embodiments of the present application.
[0141] For example, when the loss value is less than the loss value threshold, the training device can determine that the model converges.
[0142] For another example, when the number of training times exceeds the number threshold, the training device can determine that the model converges.
[0143] In this way, the training device can determine the converged model as the region recognition model that has completed training.
[0144] In some embodiments, the method for training the above-mentioned region recognition model further includes: the training device obtains test sample images annotated with target scene feature information, inputs the test sample images into the region recognition model, and tests the above-mentioned region recognition model. Exemplarily, the process of the training device testing the above-mentioned region recognition model may refer to the above steps S1 to S5, which will not be elaborated here.
[0145] In some embodiments, when the regions included in the target scene change, the training device may obtain one or more images annotated with the changed target scene feature information; and re-train the region recognition model according to the one or more images annotated with the changed target scene feature information, so as to obtain a re-trained region recognition model.
[0146] Among them, the process of the training device re-training the region recognition model is as shown in the above steps S1 to S6.
[0147] It should be noted that as the layout in the target scene is adjusted or changed, the feature information of the target scene may also change. At this time, the training device may obtain multiple images of the target scene and the corrected feature information on each image, and re-train the region recognition model. It should be understood that if the boundaries of one or more regions in the target scene in the feature information change with the adjustment or change of the layout in the target scene, and the region recognition model is not corrected in time, it will lead to a large error in the region recognition result obtained by using the region recognition model, thereby affecting the accuracy of the above-mentioned region population statistics.
[0148] Among them, the corrected feature information annotated on each image of the above-mentioned target scene may be re-entered by the user.
[0149] Based on the above embodiments, for a target scene that requires region recognition, the training device uses one or more images annotated with the feature information of the target scene as a training sample set to train the initial model, so as to obtain a region recognition model that can be used for region recognition of the target scene. In this way, using the trained region recognition model, automatic recognition of one or more regions of the target scene can be achieved.
[0150] It can be seen that the above mainly introduces the solution provided by the embodiments of the present application from the perspective of methods. To implement the above functions, the embodiments of the present application provide the corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should easily realize that, in combination with the modules and algorithm steps of each example described in the embodiments disclosed herein, the embodiments of the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0151] The embodiments of the present application can divide the functional modules of the above statistical device or training device according to the above method examples. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. Optionally, the division of modules in the embodiments of the present application is illustrative, only a logical function division, and there can be other division methods in actual implementation.
[0152] As Figure 8 shown, the embodiments of the present application also provide a structural schematic diagram of a statistical device. The statistical device 200 may include: a transceiver module 201, an identification module 202, and a processing module 203.
[0153] Among them, the transceiver module 201 is used to obtain the image to be recognized, and the image to be recognized is an image obtained by photographing a target scene.
[0154] The identification module 202 is used to input the image to be recognized into the region recognition model to obtain the feature information of the target scene; the feature information is used to indicate one or more target regions of the target scene.
[0155] The processing module 203 is used to divide the image to be recognized into multiple image blocks; the processing module 203 is further used to determine whether the first target object in the image to be recognized is in the target region according to the multiple image blocks; and determine the number of target objects in the target region.
[0156] In a possible implementation manner, the above processing module 203 is specifically used to: determine the repetition ratio of the number of identical image blocks in the image blocks included in the first target object in the image to be recognized and the image blocks included in the target region to the number of image blocks included in the first target object; when the repetition ratio is greater than or equal to a preset threshold, determine that the first target object is in the target region.
[0157] In another possible implementation, the above-mentioned processing module 203 is further configured to determine the image blocks included in the first target object in the image to be recognized and the image blocks included in the target area.
[0158] In yet another possible implementation, the above-mentioned processing module 203 is specifically configured to: for any one of the multiple image blocks, when the center point of any one of the image blocks is within the recognition range of the first target object, determine that the first target object in the image to be recognized includes any one of the image blocks; and / or, when the center point of any one of the image blocks is within the recognition range of the target area, determine that the target area in the image to be recognized includes any one of the image blocks.
[0159] In yet another possible implementation, the above-mentioned processing module 203 is further configured to determine that the first target object is not within the target area when the repetition ratio is less than a preset threshold.
[0160] In yet another possible implementation, the above-mentioned feature information further includes one or more boundary lines of the target scene, and the transceiver module 201 is further configured to obtain the trajectory information of the first target object in a preset time period, the preset time period includes the first moment when the first target object enters the target area, and the processing module 203 is further configured to determine the connection line between the trajectory point at the first moment and the trajectory point at the second moment in the trajectory information of the preset time period, the second moment is before the first moment and belongs to the preset time period; and is configured to determine that the first target object is not within the target area when the connection line has an intersection with one or more boundary lines.
[0161] For the specific descriptions of the above optional methods, reference may be made to the foregoing method embodiments, which will not be elaborated here. In addition, for any of the above-described explanations of the statistical device 200 and the descriptions of the beneficial effects, reference may be made to the corresponding method embodiments above, which will not be elaborated.
[0162] As an example, in combination with Figure 1 , the functions implemented by the message processing module 203 in the statistical device 200 can be executed by the processor 110 or the processor 170 in Figure 1 and the program code in the memory 140 in Figure 1 . The functions implemented by the transceiver module 101 can be implemented by the communication line 120 in Figure 1 , and of course, it is not limited to this.
[0163] As Figure 9 shown, a structural schematic diagram of a training device 300 provided by an embodiment of the present application is also shown. The device 300 may include: a transceiver module 301 and a training module 302. Optionally, the device 300 may further include a processing module 303.
[0164] Among them, the transceiver module 301 is configured to obtain a training sample set, which includes one or more images labeled with feature information of the target scene. The feature information includes one or more regions and one or more boundary lines of the target scene.
[0165] The training module 302 is configured to train an initial model according to the training sample set to obtain a region recognition model, which is used to perform region recognition on an image of the target scene.
[0166] In a possible implementation manner, the above device further includes a processing module 303. The transceiver module 301 is specifically configured to: obtain multiple images of the target scene. For the first image captured at the first shooting angle among the multiple images, the processing module 303 is configured to convert the first image into a first image with a top-down shooting angle, where the first shooting angle is a flat shooting angle or a bottom-up shooting angle; the transceiver module is further specifically configured to obtain the feature information labeled on the image with a top-down shooting angle. The processing module 303 is further configured to convert the first image labeled with the feature information into a first image captured at the first shooting angle. The processing module 303 is further configured to determine the first image captured at the first viewing angle and labeled with the feature information as an image in the training sample set. In another possible implementation manner, the transceiver module 301 is further configured to, when the region included in the target scene changes, obtain one or more images labeled with the feature information of the changed target scene. The training module 302 is further configured to retrain the region recognition model according to the one or more images labeled with the feature information of the changed target scene to obtain a retrained region recognition model.
[0167] For the specific description of the above optional manner, reference may be made to the foregoing method embodiments, which will not be elaborated here. In addition, the explanations and descriptions of the beneficial effects of any of the above provided training devices 300 may refer to the corresponding method embodiments above and will not be elaborated.
[0168] As an example, in combination with Figure 1 , the functions implemented by the processing module 303 in the training device can be executed by the processor 110 or the processor 170 in Figure 1 and the program code in the memory 140 in Figure 1 . The functions implemented by the transceiver module 301 can be implemented by the communication line 120 in Figure 1 , but of course it is not limited to this.
[0169] Those skilled in the art should easily realize that, for the units and algorithm steps of each example described in combination with the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described function for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0170] It should be noted that Figure 8 or Figure 9 The division of modules in [] is illustrative, merely a logical function division, and there may be other division methods in actual implementation. For example, two or more functions can also be integrated into one processing module. The above integrated module can be implemented in the form of hardware or in the form of a software function module.
[0171] The embodiments of the present application also provide a computer-readable storage medium, including computer-executable instructions, which when running on a computer, cause the computer to execute any one of the methods provided in the above embodiments. For example, Figure 2 One or more features of S101 to S105 in [] can be borne by one or more computer-executable instructions stored in the computer-readable storage medium.
[0172] The embodiments of the present application also provide a computer program product containing computer-executable instructions, which when running on a computer, cause the computer to execute any one of the methods provided in the above embodiments.
[0173] The embodiments of the present application also provide a chip, including: a processor and an interface, the processor is coupled to a memory through the interface, and when the processor executes the computer program or computer-executable instructions in the memory, any one of the methods provided in the above embodiments is executed.
[0174] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using a software program, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer execution instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more media integrated therein. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.
[0175] As described above, the above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for counting the number of targets in a region, characterized in that, The method includes: Obtain an image to be recognized, where the image to be recognized is an image captured from a target scene; Input the image to be recognized into a region recognition model to obtain feature information of the target scene; the feature information is used to indicate one or more target regions of the target scene; the feature information further includes one or more boundary lines of the target scene, and the boundary line refers to the intersection line between an obstacle that cannot be directly crossed and the ground; Divide the image to be recognized into multiple image blocks; Determine the repetition ratio of the number of identical image blocks between the image blocks included in the first target object in the image to be recognized and the image blocks included in the target region, and the number of image blocks included in the first target object; When the repetition ratio is greater than or equal to a preset threshold and there is no intersection between the trajectory connection line corresponding to the first target object and the one or more boundary lines, determine that the first target object is within the target region; the trajectory connection line is the connection line between the trajectory point at the first moment and the trajectory point at the second moment in the trajectory information of the first target object during a preset period; the preset period includes the first moment when the first target object enters the target region; the second moment is before the first moment and belongs to the preset period; When there is an intersection between the trajectory connection line corresponding to the first target object and the one or more boundary lines, determine that the first target object is not within the target region; Determine the number of target objects within the target region.
2. The method according to claim 1, wherein The method further includes: When the repetition ratio is less than the preset threshold, determine that the first target object is not within the target region.
3. The method according to claim 1, characterized in that, Before determining the repetition ratio of the number of identical image blocks between the image blocks included in the first target object in the image to be recognized and the image blocks included in the target region, and the number of image blocks included in the first target object, the method further includes: For any one of the multiple image blocks, when the center point of the any one image block is within the recognition range of the first target object, determine that the first target object in the image to be recognized includes the any one image block; and / or, When the center point of the any one image block is within the recognition range of the target region, determine that the target region in the image to be recognized includes the any one image block.
4. The method according to claim 1, characterized in that, The method further includes: Obtain a training sample set, where the training sample set includes one or more images annotated with the feature information of the target scene, and the feature information includes one or more regions and one or more boundary lines of the target scene; Train an initial model according to the training sample set to obtain a region recognition model, and the region recognition model is used to perform region recognition on the image of the target scene.
5. The method according to claim 4, characterized in that, The obtaining of the training sample set includes: Obtain multiple images of the target scene; For the first image captured at a first shooting angle among the multiple images, convert the first image into an image with a top-down shooting angle, where the first shooting angle is a flat shooting angle or an upward shooting angle; Obtain the feature information annotated on the image with a top-down shooting angle; Convert the first image marked with the feature information into the first image captured at the first shooting angle; Determine the first image with the shooting angle being the first perspective and marked with the feature information as an image in the training sample set.
6. The method according to claim 4, wherein The method further includes: When the area included in the target scene changes, obtain one or more images marked with the feature information of the changed target scene; Retrain the area recognition model according to one or more images marked with the feature information of the changed target scene to obtain a retrained area recognition model.
7. A target quantity statistical device within a region, characterized in that, The device includes: A transceiver module, configured to obtain an image to be recognized, where the image to be recognized is an image captured of a target scene; An identification module, configured to input the image to be recognized into an area recognition model to obtain the feature information of the target scene; the feature information is used to indicate one or more target areas of the target scene; the feature information further includes one or more boundary lines of the target scene, and the boundary line refers to the intersection line between an obstacle that cannot be directly crossed and the ground; A processing module, configured to divide the image to be recognized into a plurality of image blocks; The processing module is further configured to determine the repetition ratio of the number of identical image blocks between the image blocks included in the first target object in the image to be recognized and the image blocks included in the target area, and the number of image blocks included in the first target object; when the repetition ratio is greater than or equal to a preset threshold and there is no intersection between the trajectory connection line corresponding to the first target object and the one or more boundary lines, determine that the first target object is within the target area; the trajectory connection line is the connection line between the trajectory point at the first moment and the trajectory point at the second moment in the trajectory information of the first target object in a preset time period; the preset time period includes the first moment when the first target object enters the target area; the second moment is before the first moment and belongs to the preset time period; when there is an intersection between the trajectory connection line corresponding to the first target object and the one or more boundary lines, determine that the first target object is not within the target area; and determine the number of target objects within the target area; The processing module is further configured to determine that the first target object is not within the target area when the repetition ratio is less than the preset threshold; The processing module is specifically configured to: For any one of the plurality of image blocks, when the center point of the any one image block is within the recognition range of the first target object, determine that the first target object in the image to be recognized includes the any one image block; and / or, When the center point of the any one image block is within the recognition range of the target area, determine that the target area in the image to be recognized includes the any one image block.
8. The device according to claim 7, characterized in that The device further includes: A transceiver module, configured to obtain a training sample set, where the training sample set includes one or more images marked with the feature information of the target scene, and the feature information includes one or more areas and one or more boundary lines of the target scene; A training module, configured to train an initial model according to the training sample set to obtain a region recognition model, where the region recognition model is used to perform region recognition on an image of the target scene; The apparatus further includes a processing module, and the transceiver module is specifically configured to obtain multiple images of the target scene; For a first image captured at a first shooting angle among the multiple images, the processing module is configured to convert the first image into an image with a shooting angle of top-down shooting, where the first shooting angle is horizontal shooting or upward shooting; The transceiver module is further specifically configured to obtain feature information annotated on the image with a shooting angle of top-down shooting; The processing module is further configured to convert the first image annotated with the feature information into the first image captured at the first shooting angle; The processing module is further configured to determine the first image with a shooting angle of a first perspective and annotated with the feature information as an image in the training sample set; The transceiver module is further configured to, when a region included in the target scene changes, obtain one or more images annotated with feature information of the changed target scene; The training module is further configured to retrain the region recognition model according to the one or more images annotated with feature information of the changed target scene to obtain a retrained region recognition model.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and when the computer instructions run, the computer is caused to execute the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Method and device for determining sample image, equipment and storage medium
CN110298838A
Region identification method and device, equipment and medium
CN111241881A
Method for detecting number of people in target place, recommendation method, detection system and medium.
CN112036345A