Image recognition method and device
By performing frame processing and area recognition on the video, combining the recognition results of multiple images and object constraints, the problem of low image recognition accuracy in the prior art is solved, and higher recognition accuracy is achieved.
Patent Information
- Application Number
- CN202311868630.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2043-12-29
AI Technical Summary
The existing image recognition model has the problem of low recognition accuracy in the recognition image, especially when the recognition degree reaches the preset similarity, it is prone to recognition errors.
By performing frame-based processing on the target video, multiple images to be identified are obtained, and the area division model and image recognition model are used to identify the area to be identified. Combined with the set number of image recognition results and object constraint conditions, it is determined whether the target object is identified.
The probability of error recognition is reduced, the accuracy of image recognition is improved, and the accuracy of recognition results is ensured by analyzing multiple images.
Smart Images

Figure CN117830896B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image recognition technology, and in particular to an image recognition method. The present application also relates to an image recognition device, a computing device, and a computer-readable storage medium. Background Art
[0002] With the development of Internet technology, the application of image recognition in practical applications has become increasingly extensive. Image recognition aims to enable computer systems to understand and process image and video data like humans.
[0003] In existing technologies, image recognition is a technology that uses computers to process, analyze, and understand images to identify various patterns of objects and targets. The recognition process includes image preprocessing, image segmentation, feature extraction, and judgment matching. Simply put, image recognition is how computers understand content like humans. Image recognition technology not only enables faster access to information through search, but also creates new ways to interact with the external world, and even makes the external world more intelligent.
[0004] However, most of the image recognition models currently on the market only recognize one image. If the recognition degree in the image reaches a preset similarity, the recognition method that determines that the target object has been recognized may have recognition errors, and the recognition accuracy is not high.
[0005] Therefore, there is an urgent need for an image recognition method that can improve the accuracy of image recognition. Summary of the Invention
[0006] In view of this, the embodiments of the present application provide an image recognition method to solve the technical defects existing in the prior art. The embodiments of the present application also provide an image recognition device, a computing device, and a computer-readable storage medium.
[0007] According to a first aspect of an embodiment of the present application, there is provided an image recognition method, comprising:
[0008] Perform frame processing on the target video to obtain multiple images to be identified;
[0009] Performing image recognition on a target image to be recognized to obtain a recognition result, wherein the target image to be recognized is any one of the multiple images to be recognized, and the recognition result is used to indicate whether the target image to be recognized includes a target object;
[0010] According to the recognition results of a set number of images to be recognized, it is determined whether the object constraint condition is met. If so, it is determined that the target object is recognized, wherein the object constraint condition is used to restrict the number of images including the target object in the images to be recognized.
[0011] In one embodiment of the present disclosure, performing image recognition on the target image to be recognized to obtain a recognition result includes:
[0012] Acquiring region parameters, and determining a region to be identified in the target image to be identified using a region partitioning model based on the region parameters;
[0013] The image recognition model is used to perform image recognition on the area to be recognized to obtain the recognition result.
[0014] In one embodiment of the present disclosure, determining whether the object constraint condition is met based on the recognition results of a set number of images to be recognized includes:
[0015] Determining the number of tracking images among the set number of images to be identified, wherein the tracking images are images to be identified that include the target object;
[0016] The ratio of the number of images to the set number is determined, and if the ratio is greater than a ratio threshold, it is determined that the set number of images to be identified meet the object constraint condition.
[0017] In one embodiment of the present disclosure, after determining that the target object is identified, the method includes:
[0018] Get the tracking area range;
[0019] determining an object position of a target object in each tracking image;
[0020] Determining an object to be tracked in each tracking image according to the object position and the tracking area range, wherein the object to be tracked is a target object that meets the tracking area range;
[0021] Tracking information is obtained according to the object position of the object to be tracked.
[0022] In one embodiment of the present disclosure, obtaining the tracking area range includes:
[0023] Obtaining a virtual reference line and an offset preset for the image to be recognized;
[0024] The tracking area range is determined according to the virtual reference line and the offset.
[0025] In one embodiment of the present disclosure, the tracking information includes an object motion trajectory and movement parameters; and obtaining the tracking information according to the object position of the object to be tracked includes:
[0026] generating a motion trajectory of the target object according to the object position of the object to be tracked;
[0027] The movement parameters of the target object are calculated according to the motion trajectory.
[0028] In one embodiment of the present disclosure, after determining and identifying the target object, the method further includes:
[0029] determining an object type of the identified target object;
[0030] The target objects are counted according to the object type.
[0031] According to a second aspect of an embodiment of the present application, there is provided an image recognition device, comprising:
[0032] A processing module is configured to perform frame processing on the target video to obtain a plurality of images to be identified;
[0033] a recognition module configured to perform image recognition on a target image to be recognized, and obtain a recognition result, wherein the target image to be recognized is any one of the multiple images to be recognized, and the recognition result is used to indicate whether the target image to be recognized includes a target object;
[0034] The judgment module is configured to determine whether the object constraint condition is met based on the recognition results of a set number of images to be recognized, and if so, determine that the target object is recognized, wherein the object constraint condition is used to constrain the number of images that include the target object in the images to be recognized.
[0035] According to a third aspect of an embodiment of the present application, a computing device is provided, including:
[0036] memory and processor;
[0037] The memory is used to store computer-executable instructions, and the processor implements the steps of the image recognition method when executing the computer-executable instructions.
[0038] According to a fourth aspect of an embodiment of the present application, a computer-readable storage medium is provided, which stores computer-executable instructions, and when the instructions are executed by a processor, the steps of the image recognition method are implemented.
[0039] According to a fifth aspect of the embodiments of the present application, a chip is provided, which stores a computer program, and when the computer program is executed by the chip, the steps of the image recognition method are implemented.
[0040] The image recognition method provided in the present application determines whether a target object is recognized by performing frame processing on a video to obtain multiple images to be recognized. This is different from the method of only recognizing one image and determining whether a target object exists by the recognized similarity. It reduces the probability of false recognition and improves the recognition accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 This is a flowchart of an image recognition method provided by an embodiment of the present application;
[0042] Figure 2 is a schematic diagram of a virtual reference line provided in an embodiment of the present application;
[0043] Figure 3 This is a processing flow chart of an image recognition method for vehicle identification provided by an embodiment of the present application;
[0044] Figure 4 This is a schematic structural diagram of an image recognition device provided in one embodiment of the present application;
[0045] Figure 5 This is a structural block diagram of a computing device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0046] The following description sets forth many specific details to facilitate a thorough understanding of the present application. However, the present application can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the scope of the present application. Therefore, the present application is not limited to the specific implementations disclosed below.
[0047] The terms used in one or more embodiments of the present application are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of the present application. The singular forms "a", "the" and "the" used in one or more embodiments of the present application and the appended claims are also intended to include plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of the present application refers to and includes any or all possible combinations of one or more associated listed items.
[0048] It should be understood that although the terms "first," "second," and the like may be used to describe various information in one or more embodiments of the present application, such information should not be limited to these terms. These terms are merely used to distinguish information of the same type from one another. For example, "first" may also be referred to as "second," and similarly, "second" may also be referred to as "first," without departing from the scope of one or more embodiments of the present application.
[0049] First, the terms involved in one or more embodiments of the present invention are explained.
[0050] Model convergence refers to the process by which the loss function gradually decreases and stabilizes during model training. Typically, the most commonly used metric for determining model convergence is the change in the loss function. The loss function is a function of the model's weights and bias parameters; smaller values indicate a better model fit to the data. Therefore, monitoring changes in the loss function can reflect the model's fit. In addition to the loss function, other metrics can also be used to determine model convergence, such as accuracy and F1 score, which vary depending on the specific problem.
[0051] This application provides an image recognition method, an image recognition apparatus, a computing device, and a computer-readable storage medium, which are described in detail in the following embodiments.
[0052] Figure 1 A flowchart of an image recognition method provided according to an embodiment of the present application is shown, which specifically includes the following steps 102-106.
[0053] Step 102: Frame-segment the target video to obtain a plurality of images to be identified.
[0054] Among them, the target video specifically refers to any video that needs to be recognized. The target video can be one or more, which can be a video recorded in advance or a real-time video transmitted back by a surveillance camera; the image to be recognized specifically refers to the image obtained by frame processing of the target video.
[0055] In a specific implementation, during the frame processing of the target video, all frames can be extracted, frames can be extracted at time intervals, or frames can be extracted by the number of frames. For example, if a video has a total of 100 frames and 24 frames per second, all 100 frames can be extracted, or only the 1st, 25th, 49th, 73rd, etc. can be extracted, or one frame can be extracted every 10 frames. Specifically, video frames can be extracted using methods and tools such as FFmpeg, OpenCV, Libav, Python's MoviePy library, GStreamer, MATLAB, VLC media player, etc. The video frames extracted from the target video are the images to be recognized.
[0056] In order to ensure the efficiency of image recognition, after the image to be recognized is extracted, it can be stored in the same storage space.
[0057] Step 104: performing image recognition on the target image to be recognized to obtain a recognition result, wherein the target image to be recognized is any one of the multiple images to be recognized, and the recognition result is used to indicate whether the target image to be recognized includes a target object.
[0058] Among them, the target image to be identified is the image to be identified for image recognition, and the target object can be any object. Any element appearing in the image to be identified can be the target object, such as: cars, people, plants, etc. can all be target objects. Of course, a certain object or element can also be specified as the target object. For example: if only a car needs to be identified, the car is determined as the target object, and other objects in the image to be identified will not be identified. It is easy to understand that there may be one or more target objects in an image to be identified.
[0059] In a specific implementation, each of the multiple images to be recognized obtained by the frame segmentation process is used as a target image to be recognized for image recognition. When performing image recognition, the image recognition can be performed using a trained image recognition model or other image recognition tools.
[0060] In order to ensure the efficiency of image recognition and reduce the image recognition time, OpenCV is used to narrow the recognition range of the image so that the image recognition model or image recognition tool only needs to recognize the image within the recognition range of the image to be recognized. In this embodiment, this is specifically achieved through the following steps.
[0061] The performing image recognition on the target image to be recognized to obtain a recognition result includes:
[0062] Acquiring region parameters, and determining a region to be identified in the target image to be identified using a region partitioning model based on the region parameters;
[0063] The image recognition model is used to perform image recognition on the area to be recognized to obtain the recognition result.
[0064] Among them, the region parameters specifically refer to the parameters of the divided region, which can be the parameters of the height and width of the divided region, or the size of the region; the region parameters can include one or more groups, for example: a height parameter of 5cm and a width parameter of 6cm is a group of region parameters, a height parameter of 4cm and a width parameter of 6cm is a group of region parameters, and a height parameter of 8cm and a width parameter of 5cm is a group of region parameters. When dividing the region, the division can be performed according to any of these three groups of region parameters. Considering that the size of the target object in the image to be identified is related to the distance between the target object and the photographed object, and the sizes of different target objects are also different, the number of groups of region parameters can be appropriately determined; the region division model specifically refers to a model for dividing the target image to be identified into regions. The algorithm of the region division model can be based on OpenCV. The region to be identified specifically refers to the region that may contain the target object identified by the region division model; the image recognition model specifically refers to a model for identifying the target object. Models such as the YOLO model, the Faster-RCNN model, and the Mask-RCNN model can all be used as image recognition models.
[0065] In a specific implementation, multiple sets of region parameters are pre-set, and the region parameters and the target image to be identified are input into a region partition model. The region partition model outputs the parameters of the region to be identified in the target image to be identified. The parameters of the region to be identified and the target image to be identified are then input into an image recognition model, which outputs a recognition result. Of course, if the target object is identified, the image recognition model will output a recognition result including the location of the target object in the target image to be identified. Of course, in addition to the location of the target object in the target image to be identified, the recognition result may also include the location of other objects in the target image to be identified, the object type of each object in the target image to be identified, the confidence level (similarity), etc. It should be noted that the region partition model and the image recognition model can be two independent models or two sub-models within a single model.
[0066] In addition, in order to ensure the accuracy of image recognition, the model needs to be trained before using it for image recognition. Here we take the YOLO model as an example.
[0067] Step 1: Collect and organize training samples. A certain number of images, which may or may not contain the target object, should be the majority of the total number of images. These images serve as training samples. The number of images in the training sample should be as large as possible. The more images there are, the more accurate the recognition results of the trained YOLO model will be.
[0068] Step 2: Preprocess the training samples. Label each image in the collected training samples, label the target objects in the images, and assign names to the labeled target objects. The names can be output as recognition results. The names can be in English or Chinese. After labeling is completed, a training sample set is generated.
[0069] Step 3: Train the YOLO model. Existing YOLO models are divided into versions s, l, m, and x. S, l, m, and x represent small, large, medium, and xlarge, respectively, reflecting the size and complexity of the model. In actual training, you can choose based on your needs. For example, the small (s) version of the YOLO model is the smallest model, requires low computing resources, and can run on resource-constrained devices. If the device's computing resources are too limited, you can use this version of the YOLO model. If computing resources are sufficient, you can choose the extra-large (x) version of the YOLO model. The extra-large (x) version is the largest and most complex model in the YOLO model. It is designed to provide the best performance without considering the use of computing resources. Therefore, this model requires a lot of computing resources and may not run in real time on some devices. However, if computing resources are sufficient, this model can provide the best performance. The large (l) and medium (m) versions of the YOLO model strike a balance between size and performance. Therefore, if computing resources are sufficient, the large (l) or medium (m) versions of the YOLO model can be selected. During training, the training sample set generated after annotation is input into the selected YOLO model. Note that if training samples are annotated in Chinese, the encoding format recognized by the model can be added before training to train specifically for Chinese. During training, the number of training rounds can be pre-set. If the model converges before training is complete, training can be terminated early. If the model has not converged after the set number of training rounds, further training can be performed until the model converges. For example, if a training round requires 400 rounds but the final model has not converged, 100 additional rounds can be added. Conversely, if a training round requires 500 rounds, it can be stopped at 300 rounds and a trained model can still be generated after stopping. It should be noted that during training, the training sample set can be divided into a validation set and a comparative test set. During the training process, testing and comparison can be performed while training, and finally a comparative result can be generated.
[0070] In summary, since not all areas in the image to be identified need to be identified, the recognition range is narrowed by the region division model, so that the image recognition model only needs to identify the area to be identified, rather than the entire target image to be identified, which reduces the recognition time and improves the recognition efficiency.
[0071] Step 106: Determine whether the object constraint condition is met based on the recognition results of a set number of images to be recognized. If so, determine that the target object is recognized, wherein the object constraint condition is used to restrict the number of images including the target object in the images to be recognized.
[0072] The set number is set by those skilled in the art according to their needs, such as 2, 10, or 100, which is not limited in this application. The object constraint condition refers to the condition for determining the recognition of the target object, for example, whether the recognition result is a preset number threshold for the number of images to be recognized that contain the target object, and if it is reached, it is determined that the object constraint condition is met, or whether the percentage of images to be recognized that contain the target object in the set number of images to be recognized meets a preset ratio threshold, and if it is reached, it is determined that the object constraint condition is met, which is not limited in this application.
[0073] In specific implementation, after image recognition, the image to be recognized with the recognition result can be placed in a pre-set image pool. The number of images to be recognized stored in the image pool is set. Every time the recognition result of an image to be recognized is obtained, it is placed in the image pool. When the image pool is full, it is determined that the set number of images to be recognized is obtained. The multiple images to be recognized in the image pool are calculated to determine whether the images to be recognized in the image pool recognize the target object.
[0074] Specifically, in this embodiment, whether the target object is identified is determined by determining whether the percentage of images to be identified that include the target object among a set number of images to be identified meets a preset ratio threshold to determine whether the object constraint condition is met. The specific steps are as follows.
[0075] The step of determining whether the object constraint condition is met based on the recognition results of a set number of images to be recognized includes:
[0076] Determining the number of tracking images among the set number of images to be identified, wherein the tracking images are images to be identified that include the target object;
[0077] The ratio of the number of images to the set number is determined, and if the ratio is greater than a ratio threshold, it is determined that the set number of images to be identified meet the object constraint condition.
[0078] The ratio threshold is pre-set by those skilled in the art, and may be 80% or 90%, which is not limited in this application.
[0079] In specific implementation, the ratio of the number of images to the set number can be calculated by the following formula:
[0080] Number of traced images / set number*100%=ratio
[0081] The ratio is compared with a ratio threshold. If the ratio is greater than the ratio threshold, the set number of images to be identified are determined to meet the object constraint conditions, and the target object is determined to be identified. If the ratio is less than the ratio threshold, the current image to be identified is discarded, and all current images to be identified are judged as not having identified the target object.
[0082] In addition, to prevent missed recognition of the target object, sliding window recognition can be performed. For example, a video has 1000 frames, and the target object only appears between frames 325-425. The number is set to 100 frames (100 images to be recognized), and the ratio threshold is 80%. If recognition is performed every 100 frames, when frames 300-400 are recognized, the obtained ratio is 76%, which is less than the ratio threshold, and it is determined that the target object is not recognized. When frames 400-500 are recognized, the obtained ratio is 24%, which is less than the ratio threshold, and it is also determined that the target object is not recognized. Obviously, an obvious error has occurred. Therefore, recognition can be performed every few frames, that is, 300-400 frames, 320-420 frames, 440-540 frames, ... can be recognized. Recognition using this method can reduce the probability of false recognition to a certain extent and improve the recognition accuracy to a certain extent.
[0083] In summary, the image recognition method disclosed in this embodiment determines whether the target object is recognized by obtaining multiple images to be recognized by frame processing of the video. This is different from the method of only recognizing one image and determining whether the target object exists by the recognized similarity. It reduces the probability of false recognition and improves the recognition accuracy.
[0084] After confirming that the target object has been identified, in order to analyze the identified target object, a tracking area can be set during the analysis. When the target object appears within the tracking area, it is analyzed. The reason for setting the tracking area is to exclude some objects that accidentally appear in the image to be identified. Specifically, in this embodiment, the target object is analyzed through the following steps.
[0085] After the target object is determined and identified, the method includes:
[0086] Get the tracking area range;
[0087] determining an object position of a target object in each tracking image;
[0088] Determining an object to be tracked in each tracking image according to the object position and the tracking area range, wherein the object to be tracked is a target object that meets the tracking area range;
[0089] Tracking information is obtained according to the object position of the object to be tracked.
[0090] Among them, the tracking area range specifically refers to the range of tracking the target object. If the target object does not appear within the tracking area range, no processing is performed on it. The tracking area range is set by technical personnel in this field according to needs; the tracking image is the image to be identified containing the target object, and the tracking information is the movement information of the target object.
[0091] In a specific implementation, if the target object is within the tracking area, the target object is determined to be the object to be tracked, and tracking information is determined according to the object position of the object to be tracked in the plurality of images to be identified.
[0092] Specifically, the tracking area range is preset. The diagonal coordinates of the tracking area range can be set, and a rectangular range can be determined as the tracking area range according to the diagonal coordinates. For example, if the coordinates of the upper right corner are set to (20, 15) and the coordinates of the lower left corner are set to (0, 5), then the range formed by connecting (0, 5), (0, 15), (20, 15), and (20, 5) in sequence is determined as the tracking area range. Of course, a virtual reference line and the offset corresponding to the virtual reference line can also be set to determine the tracking area range.
[0093] In this embodiment, the tracking area range is determined by using a virtual reference line and an offset corresponding to the virtual reference line. The specific implementation steps are as follows.
[0094] Obtaining the tracking area range includes:
[0095] Obtaining a virtual reference line and an offset preset for the image to be recognized;
[0096] The tracking area range is determined according to the virtual reference line and the offset.
[0097] Among them, the virtual reference line is a custom virtual line of any direction and length in the target video. The direction and length of the virtual line and other parameters can be customized. The offset is the farthest offset distance from the target object to the virtual reference line. The offset can also be customized. The offset can be set in both directions, that is, on both sides of the virtual reference line. For example: Figure 2 As shown in FIG, a virtual reference line is set horizontally in the middle of the height direction of the target video, the offset is set to 10, and the area 10 above and below the virtual reference line is the tracking area range.
[0098] Specifically, the motion trajectory and movement parameters of the object can be determined according to the object position of the object to be tracked in each tracking image. In this embodiment, this is achieved through the following steps.
[0099] The tracking information includes an object motion trajectory and movement parameters; and obtaining the tracking information according to the object position of the object to be tracked includes:
[0100] generating a motion trajectory of the target object according to the object position of the object to be tracked;
[0101] The movement parameters of the target object are calculated according to the motion trajectory.
[0102] Specifically, the motion trajectory refers to the motion path of the target object within the tracking area. An object position of the object to be tracked is generated in each tracking image. The position coordinates of each object position in multiple tracking images are recorded and connected in sequence to obtain the motion trajectory of the target object. After obtaining the motion trajectory, the movement parameters can be determined based on the motion trajectory. Specifically, the movement parameters can refer to at least one of the movement direction, movement speed, movement time, and movement distance. The movement direction of the target object is determined by determining two points in the motion trajectory and the distance between the two points and the virtual reference line. The movement speed and movement distance of the target object are determined by the size ratio of the scene recorded in the motion trajectory and the actual scene in the target video. At the same time, the movement time of the target object from one point to another can be calculated based on the number of frames of the object to be tracked in the image to be identified and the motion trajectory.
[0103] Considering that after the target object is identified, the traffic of the target object and other conditions can be analyzed, and the target object can also be counted, this is specifically achieved through the following steps.
[0104] After the target object is determined and identified, the method further includes:
[0105] determining an object type of the identified target object;
[0106] The target objects are counted according to the object type.
[0107] The object type specifically refers to the type of the object, for example, a fire truck, a police car, and other object types are all cars.
[0108] In a specific implementation, multiple images to be identified are identified, and each image to be identified may contain multiple target objects. The target objects are counted according to their respective classifications determined in the images to be identified, or the number can be calculated based on a predetermined number of images to be identified and the data of the images containing the target objects.
[0109] In specific implementation, the image recognition method can be deployed in any camera, and the deployment can be completed by specifying a GPU server, which is simple to operate and has good adaptability.
[0110] The following combined Figure 3 Taking the application of the image recognition method provided in this application to a vehicle as an example, the image recognition method is further explained. Figure 3 A processing flow chart of an image recognition method for vehicle identification provided in an embodiment of the present application is shown, which specifically includes the following steps:
[0111] Step 302: Frame-segment the target video to obtain a plurality of images to be identified.
[0112] Step 304: Acquire region parameters, and determine the region to be identified in the target image to be identified using a region partition model based on the region parameters.
[0113] Step 306: Perform image recognition on the area to be recognized using an image recognition model to obtain the recognition result.
[0114] Step 308: Determine the number of tracking images among the set number of images to be identified, wherein the tracking images are images to be identified that include the target object;
[0115] Step 310: Determine the ratio of the number of images to the set number. If the ratio is greater than a ratio threshold, determine that the set number of images to be identified meet the object constraint condition. If so, determine that the vehicle is identified.
[0116] Step 312: Obtain a virtual reference line and an offset preset for the image to be recognized.
[0117] Step 314: Determine the tracking area range according to the virtual reference line and the offset.
[0118] Step 316: Obtain the tracking area range.
[0119] Step 318: Determine the object position of the vehicle in each tracking image.
[0120] Step 320: Determine the vehicle to be tracked in each tracking image according to the object position and the tracking area range, wherein the vehicle to be tracked is a vehicle that meets the tracking area range.
[0121] Step 322: Generate a motion trajectory of the vehicle according to the object position of the vehicle to be tracked.
[0122] Step 324: Calculate the movement parameters of the vehicle according to the motion trajectory.
[0123] Corresponding to the above method embodiment, the present application also provides an image recognition device embodiment, Figure 4 FIG. 1 shows a schematic diagram of the structure of an image recognition device provided by an embodiment of the present application. Figure 4 As shown, the device includes:
[0124] The processing module 402 is configured to perform frame processing on the target video to obtain a plurality of images to be identified;
[0125] The recognition module 404 is configured to perform image recognition on a target image to be recognized, and obtain a recognition result, wherein the target image to be recognized is any one of the multiple images to be recognized, and the recognition result is used to indicate whether the target image to be recognized includes a target object;
[0126] The judgment module 406 is configured to determine whether the object constraint condition is met based on the recognition results of a set number of images to be recognized. If so, it is determined that the target object is recognized, wherein the object constraint condition is used to constrain the number of images including the target object in the images to be recognized.
[0127] The recognition module 404 is further configured to obtain region parameters, and determine the region to be recognized in the target image to be recognized using a region division model based on the region parameters; perform image recognition on the region to be recognized using an image recognition model to obtain the recognition result.
[0128] The judgment module 406 is further configured to determine the number of tracking images in the set number of images to be identified, wherein the tracking images are images to be identified that include the target object; determine the ratio of the number of images to the set number, and if the ratio is greater than a ratio threshold, determine that the set number of images to be identified meet the object constraint condition.
[0129] The device also includes:
[0130] A range acquisition module, configured to acquire a tracking area range;
[0131] a position determination module configured to determine an object position of a target object in each tracking image;
[0132] an object determination module configured to determine an object to be tracked in each tracking image according to the object position and the tracking area range, wherein the object to be tracked is a target object that meets the tracking area range;
[0133] The information acquisition module is configured to obtain tracking information according to the object position of the object to be tracked.
[0134] The range acquisition module 408 is further configured to acquire a virtual reference line and an offset preset for the image to be recognized; and determine the tracking area range according to the virtual reference line and the offset.
[0135] The tracking information includes an object motion trajectory and movement parameters; the information acquisition module 414 is further configured to generate a motion trajectory of the target object according to the object position of the object to be tracked; and calculate the movement parameters of the target object according to the motion trajectory.
[0136] The counting module is configured to determine the object type of the identified target object; and count the target object according to the object type.
[0137] The above is a schematic scheme of an image recognition device of this embodiment. It should be noted that the technical solution of the image recognition device and the technical solution of the above-mentioned image recognition method belong to the same concept. For details not described in detail in the technical solution of the image recognition device, please refer to the description of the technical solution of the above-mentioned image recognition method. In addition, the various components in the device embodiment should be understood as functional modules that must be established to implement each step of the program flow or each step of the method, and each functional module is not an actual functional division or separation definition. The device claim defined by such a group of functional modules should be understood as a functional module architecture that mainly implements the solution through the computer program recorded in the specification, and should not be understood as a physical device that mainly implements the solution through hardware.
[0138] Figure 5 The block diagram shows a structure of a computing device 500 according to an embodiment of the present application. The components of the computing device 500 include, but are not limited to, a memory 510 and a processor 520. The processor 520 is connected to the memory 510 via a bus 530, and a database 550 is used to store data.
[0139] The computing device 500 also includes an access device 540 that enables the computing device 500 to communicate via one or more networks 560. Examples of such networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 540 may include one or more of any type of network interface (e.g., a network interface card (NIC)) whether wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, a near field communication (NFC) interface, and the like.
[0140] In one embodiment of the present application, the above components of the computing device 500 and Figure 5 Other components not shown in the figure may also be connected to each other, for example, via a bus. Figure 5 The computing device structure block diagram shown is for illustrative purposes only and is not intended to limit the scope of the present application. Those skilled in the art may add or replace other components as needed.
[0141] Computing device 500 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, personal digital assistant, laptop computer, notebook computer, netbook computer, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or personal computer (PC). Computing device 500 may also be a mobile or stationary server.
[0142] The processor 520 is configured to execute computer executable instructions of the image recognition method.
[0143] The above is a schematic diagram of a computing device according to this embodiment. It should be noted that the technical solution of the computing device and the technical solution of the above-mentioned image recognition method are based on the same concept. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the above-mentioned image recognition method.
[0144] An embodiment of the present application further provides a computer-readable storage medium storing computer instructions, which are used in an image recognition method when executed by a processor.
[0145] The above is a schematic diagram of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solution of the above-mentioned image recognition method are based on the same concept. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solution of the above-mentioned image recognition method.
[0146] The computer instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium. It should be noted that the content contained in the computer-readable medium may be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.
[0147] An embodiment of the present application further provides a chip storing a computer program, which implements the steps of the image recognition method when executed by the chip.
[0148] It should be noted that for the aforementioned method embodiments, for ease of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.
[0149] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0150] The preferred embodiments of the present application disclosed above are intended only to help illustrate the present application. The optional embodiments do not describe all details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made based on the content of this application. This application selects and describes these embodiments in detail in order to better explain the principles and practical applications of this application, so that those skilled in the art can better understand and utilize this application. This application is limited only by the claims and their full scope and equivalents.
Claims
1. An image recognition method, characterized in that: include: Performing frame processing on the target video to obtain multiple images to be identified, wherein the frame processing uses a sliding window method to continuously extract video frames. The sliding window method divides the frame sequence based on a preset window length and a sliding step size, and the frame sequences of adjacent windows partially overlap; performing image recognition on a target image to be recognized to obtain a recognition result, wherein the target image to be recognized is any one of the multiple images to be recognized, and the recognition result is used to indicate whether the target image to be recognized includes a target object. The performing image recognition on the target image to be recognized to obtain a recognition result comprises: determining the number of groups of region parameters and each group of region parameters, wherein the number of groups of region parameters is determined based on the distance between the target object and the photographed object and the size of the target object; determining, according to the region parameters, a region to be recognized in the target image to be recognized obtained by the frame processing using a region partitioning model, wherein the size of the region to be recognized corresponds to the region parameters; performing image recognition on the region to be recognized using an image recognition model to obtain the recognition result; Determining whether an object constraint condition is met based on the recognition results of a set number of images to be recognized, and if so, determining that the target object is recognized, wherein the object constraint condition is used to restrict the number of images that include the target object in the images to be recognized; determining whether the object constraint condition is met based on the recognition results of the set number of images to be recognized includes: Determining the number of tracking images among the set number of images to be identified, wherein the tracking images are images to be identified that include the target object; The ratio of the number of images to the set number is determined, and if the ratio is greater than a ratio threshold, it is determined that the set number of images to be identified meet the object constraint condition.
2. The image recognition method according to claim 1, wherein: After the target object is determined and identified, the method includes: Get the tracking area range; determining an object position of a target object in each tracking image; Determining an object to be tracked in each tracking image according to the object position and the tracking area range, wherein the object to be tracked is a target object that meets the tracking area range; Tracking information is obtained according to the object position of the object to be tracked.
3. The image recognition method according to claim 2, wherein: Obtaining the tracking area range includes: Obtaining a virtual reference line and an offset preset for the image to be recognized; The tracking area range is determined according to the virtual reference line and the offset.
4. The image recognition method according to claim 2, wherein: The tracking information includes an object motion trajectory and movement parameters; and obtaining the tracking information according to the object position of the object to be tracked includes: generating a motion trajectory of the target object according to the object position of the object to be tracked; The movement parameters of the target object are calculated according to the motion trajectory.
5. The image recognition method according to claim 1, wherein: After the target object is determined and identified, the method further includes: determining an object type of the identified target object; The target objects are counted according to the object type.
6. An image recognition device, characterized in that: include: a processing module configured to perform frame processing on the target video to obtain a plurality of images to be recognized, wherein the frame processing adopts a sliding window method to continuously extract video frames, the sliding window method divides the frame sequence based on a preset window length and a sliding step size, and the frame sequences of adjacent windows partially overlap; a recognition module configured to perform image recognition on a target image to be recognized and obtain a recognition result, wherein the target image to be recognized is any one of the multiple images to be recognized, and the recognition result is used to indicate whether the target image to be recognized includes a target object, and performing image recognition on the target image to be recognized and obtaining the recognition result comprises: determining the number of groups of region parameters and each group of region parameters, wherein the number of groups of region parameters is determined based on the distance between the target object and the photographed object and the size of the target object; obtaining multiple groups of region parameters, and determining, based on the region parameters, a region to be recognized in the target image to be recognized obtained by the frame processing using a region partitioning model, wherein the size of the region to be recognized corresponds to the region parameters; and performing image recognition on the region to be recognized using an image recognition model to obtain the recognition result; The judgment module is configured to determine whether the object constraint condition is met based on the recognition results of a set number of images to be recognized, and if so, determine that the target object is recognized, wherein the object constraint condition is used to constrain the number of images that include the target object in the images to be recognized; the determination of whether the object constraint condition is met based on the recognition results of the set number of images to be recognized includes: determining the number of tracking images in the set number of images to be recognized, wherein the tracking images are images to be recognized that include the target object; determining the ratio of the number of images to the set number, and if the ratio is greater than a ratio threshold, determining that the set number of images to be recognized meet the object constraint condition.
7. A computing device, characterized in that include: memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the steps of the image recognition method according to any one of claims 1 to 5.
8. A computer-readable storage medium storing computer instructions, characterized in that: When the instruction is executed by a processor, the steps of the image recognition method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Target object identification method and device, computer equipment and storage medium
CN113449606A
Vehicle classification counting method and system based on multi-target detection and tracking
CN114973169A
Video data processing method and device, electronic equipment and storage medium
CN115170867A