Image recognition system and method based on visual handle

By using visual handles with resolution and frequency below the threshold for image recognition, combined with pre-trained models and position prediction, the problem of slow image recognition in complex scenes is solved, and computing resources are optimized and recognition speed is improved.

CN120807968APending Publication Date: 2025-10-17SHENYANG UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510804188.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing technologies have slow image recognition speeds and high computing resource consumption in complex scenarios, making it difficult to meet real-time recognition requirements. This is especially true on embedded devices and mobile terminals where hardware resources are limited, resulting in poor multi-target recognition results.

Method used

Image recognition is performed using visual handles with a resolution lower than a preset threshold and a frequency lower than a preset threshold. Image recognition is performed through a pre-trained model, combining position prediction and feature extraction to reduce the amount of computational data and improve recognition speed.

Benefits of technology

Through low-resolution and low-frequency visual handle recognition, computing resource consumption is reduced, image recognition speed is improved, computational complexity is reduced, and recognition effect is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120807968A_ABST
    Figure CN120807968A_ABST
Patent Text Reader

Abstract

The invention provides an image recognition system and method based on a visual handle. The system comprises a visual sensor and a data processing module. The visual sensor is used for collecting a visual handle under a target task; the visual handle is an image of which the resolution is lower than a preset resolution threshold and the image frequency is lower than a preset frequency threshold; the visual sensor comprises at least one imaging camera; the data processing module is used for performing image recognition based on the visual handle to obtain an initial recognition result; and in response to the condition that the initial recognition result meets a preset condition, determining the initial recognition result as an image recognition result of the target task. According to the invention, the structure and parameters of the visual model can be simplified, the robustness of the visual model can be improved, the computing power demand can be reduced, the image recognition speed can be improved, and target recognition can be completed in a complex scene.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of machine vision, and in particular to an image recognition system and method based on a visual handle. BACKGROUND

[0002] The image recognition method adopted by the related art mainly uses a camera to collect clear image data, and then uses a complex visual model and a large amount of computing resources, which is difficult to meet the real-time recognition requirement of a target in a complex scene. For example, in the autonomous walking robot technology, a large amount of data captured by a high-definition camera is usually used to analyze and recognize an object and understand the behavior of the object. However, due to the large amount of data of the high-definition camera, and the fact that the high-definition camera is easily affected by uncertain factors such as light, shielding, and angle change, the processing speed is slow, which can easily lead to delay and recognition error. On embedded devices and mobile terminals, the hardware resources are limited, and the bottleneck problem of image processing speed is more prominent. In particular, when the machine vision technology is used in a complex scene for multi-target recognition, it is difficult to effectively distinguish all the targets or understand the content in the image. SUMMARY

[0003] The embodiments of the present application provide an image recognition system and method based on a visual handle, which can improve the image recognition speed.

[0004] The technical solution of the embodiments of the present application is as follows:

[0005] The embodiments of the present application provide an image recognition method based on a visual handle, which comprises: acquiring a visual handle under a target task; the visual handle is an image with a resolution lower than a preset resolution threshold and an image frequency lower than a preset frequency threshold; performing image recognition based on the visual handle to obtain an initial recognition result; and in response to the initial recognition result meeting a preset condition, determining the initial recognition result as an image recognition result of the target task.

[0006] The embodiments of the present application provide an image recognition system based on a visual handle, which comprises: a visual sensor and a data processing module; the visual sensor is used to acquire a visual handle under a target task; the visual handle is an image with a resolution lower than a preset resolution threshold and an image frequency lower than a preset frequency threshold; the visual sensor comprises at least one imaging camera; the data processing module is used to perform image recognition based on the visual handle to obtain an initial recognition result; and in response to the initial recognition result meeting a preset condition, determining the initial recognition result as an image recognition result of the target task.

[0007] The embodiment of the present application provides a kind of image recognition device based on visual handle, comprising: visual handle module, for obtaining the visual handle under target task;The visual handle is the image with resolution lower than preset resolution threshold and image frequency lower than preset frequency threshold;Image recognition module, for carrying out image recognition based on the visual handle, obtains initial identification result;Result determination module, for responding to the initial identification result meets preset condition, determines the initial identification result as the image recognition result of the target task.

[0008] In the above scheme, the image recognition module is further configured to perform image recognition on the visual handle by using a pre-trained model to obtain an initial identification result, wherein the pre-trained model is obtained by training a set of visual handles, and the set of visual handles includes at least one sample visual handle.

[0009] In the above scheme, the visual handle module is further configured to determine at least one target non-focus imaging camera in a working state and a distance value between a lens module and an image sensor in the at least one target non-focus imaging camera in the electronic device based on a task type of the target task, the preset frequency threshold, and the preset resolution threshold, perform image acquisition on the target task by using the at least one target non-focus imaging camera based on the distance value to obtain the visual handle, or determine at least one target focus imaging camera in a working state and an aperture value of the at least one target focus imaging camera in the electronic device based on the task type of the target task, the preset frequency threshold, and the preset resolution threshold, perform image acquisition on the target task by using the at least one target focus imaging camera based on the aperture value to obtain the visual handle, or determine a target focus imaging camera in a working state in the electronic device, perform image acquisition on the target task by using the target focus imaging camera to obtain a target image, and perform convolution blur and down-sampling at least once on the target image based on the task type of the target task, the preset frequency threshold, and the preset resolution threshold to obtain the visual handle.

[0010] In the above scheme, the result determination module is further configured to obtain first image data of the target task in response to the initial identification result not meeting the preset condition, perform image recognition based on the visual handle and the first image data to obtain an image recognition result of the target task.

[0011] In the scheme, the result determination module is further configured to determine a position prediction model corresponding to a task type of the target task; perform position prediction on the visual handle through the position prediction model to obtain at least one position information of the visual handle; perform image registration on the first image data according to the visual handle to obtain a pixel position mapping relationship between the visual handle and the first image data; perform image cropping on the first image data according to the at least one position information and the pixel position mapping relationship to obtain at least one first local image; and perform image recognition based on the visual handle and the at least one first local image to obtain an image recognition result of the target task.

[0012] In the scheme, the result determination module is further configured to respectively perform feature extraction on the visual handle and the at least one first local image to correspondingly obtain a first feature vector and at least one second feature vector; the second feature vector includes a texture feature vector; perform vector fusion on the first feature vector and the at least one second feature vector to obtain a fused feature vector; and perform image recognition based on the fused feature vector to obtain an image recognition result of the target task.

[0013] In the scheme, the result determination module is further configured to perform target region cropping on the first image data to obtain a target region image of the first image data; and perform image recognition based on the visual handle, the at least one first local image, and the target region image to obtain an image recognition result of the target task.

[0014] In the scheme, the result determination module is further configured to perform image cropping on the first image data according to preset position information to obtain a second local image; generate position information of at least one random position in the first image data; perform image cropping on the first image data according to the position information of the at least one random position to obtain a third local image; and determine at least one of the second local image and the third local image as the target region image.

[0015] In the scheme, the result determination module is further configured to, in response to the initial recognition result not satisfying the preset condition, determine at least one target sensor in a working state from the electronic device based on a task type of the target task; perform data collection on the target task through the at least one target sensor to obtain at least one target data; and perform image recognition based on the visual handle and the at least one target data to obtain an image recognition result of the target task.

[0016] In the scheme, the result determination module is further configured to, in response to the initial recognition result not satisfying the preset condition, acquire first image data of the target task, and determine at least one target sensor in a working state from the electronic device based on a task type of the target task; acquire data of the target task through the at least one target sensor to obtain at least one target data; and perform image recognition based on the visual handle, the first image data, and the at least one target data to obtain an image recognition result of the target task.

[0017] In the scheme, the result determination module is further configured to determine a position prediction model corresponding to the task type; perform position prediction on the visual handle through the position prediction model to obtain at least one position information of the visual handle; perform image cropping on the first image data according to the at least one position information to obtain at least one first local image; and perform image recognition on the visual handle and the at least one first local image based on the at least one target data to obtain an image recognition result of the target task.

[0018] The electronic device provided in the embodiments of the present application comprises: a memory configured to store computer executable instructions or computer programs; and a processor configured to execute the computer executable instructions or computer programs stored in the memory to implement the image recognition method based on a visual handle provided in the embodiments of the present application.

[0019] The computer readable storage medium provided in the embodiments of the present application stores computer programs or computer executable instructions, and is configured to be executed by a processor to implement the image recognition method based on a visual handle provided in the embodiments of the present application.

[0020] The computer program product provided in the embodiments of the present application comprises computer programs or computer executable instructions, and is configured to be executed by a processor to implement the image recognition method based on a visual handle provided in the embodiments of the present application.

[0021] The embodiments of the present application have the following beneficial effects:

[0022] First, a visual handle under a target task is acquired, the visual handle being an image with a resolution lower than a preset resolution threshold and an image frequency lower than a preset frequency threshold, the low resolution can effectively reduce the data amount of the image, the low frequency can effectively reduce the interference of high-frequency noise, remove the detail redundancy, retain the key part in the image, eliminate the influence of light, view angle, target or scene change over time, and simplify the structure and parameters of the visual model. Then, image recognition is performed based on the visual handle to obtain an initial recognition result, and when the initial recognition result meets a preset condition, the initial recognition result is determined as an image recognition result of the target task. In this way, the image recognition can be performed through the visual handle with low resolution and low frequency, the calculation data amount can be reduced, the calculation resource consumption can be reduced, and the image recognition speed can be improved. BRIEF DESCRIPTION OF DRAWINGS

[0023] Figure 1 is an optional structural schematic diagram of an image recognition system based on a visual handle provided by an embodiment of the present application;

[0024] Figure 2 is an optional flow schematic diagram of an image recognition method based on a visual handle provided by an embodiment of the present application;

[0025] Figure 3 is a schematic diagram of an image provided by an embodiment of the present application;

[0026] Figure 4 is another optional flow schematic diagram of an image recognition method based on a visual handle provided by an embodiment of the present application;

[0027] Figure 5 is an optional structural schematic diagram of a non-focus imaging camera provided by an embodiment of the present application;

[0028] Figure 6 is an optional principle schematic diagram of a non-focus imaging camera provided by an embodiment of the present application;

[0029] Figure 7 is an implementation flowchart of obtaining an image recognition result provided by an embodiment of the present application;

[0030] Figure 8 is a principle schematic diagram of obtaining a first local image provided by an embodiment of the present application;

[0031] Figure 9 is another optional structural schematic diagram of an image recognition system based on a visual handle provided by an embodiment of the present application;

[0032] Figure 10 is an optional scene schematic diagram of an image recognition method based on a visual handle provided by an embodiment of the present application;

[0033] Figure 11 is another optional scene schematic diagram of the image recognition method based on a visual handle provided by the embodiment of the application;

[0034] Figure 12 is a structural block diagram of an image recognition device based on a visual handle provided by the embodiment of the application;

[0035] Figure 13 is a structural schematic diagram of an electronic device provided by the embodiment of the application. DETAILED DESCRIPTION

[0036] In order to make the purpose, technical scheme and advantages of the application clearer, the application will be further described in detail below with reference to the drawings. The described embodiments should not be regarded as limiting the application. All other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the application.

[0037] In the following description, "some embodiments" are described, which describe a subset of all possible embodiments, but it can be understood that "some embodiments" can be the same subset or different subset of all possible embodiments, and can be combined with each other without conflict.

[0038] In the following description, the terms "first\second\third" are only to distinguish similar objects, and do not represent a specific order of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence as allowed, so that the embodiments of the application described herein can be implemented in an order other than that illustrated or described herein.

[0039] In the embodiments of the application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works with other related parts to achieve a predetermined target, and can be implemented entirely or partially by using software, hardware (such as processing circuit or memory) or combination thereof. Similarly, one processor (or multiple processors or memory) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an integral module or unit that includes the function of the module or unit.

[0040] Unless otherwise defined, all technical and scientific terms used in the embodiments of the application have the same meanings as those commonly understood by those skilled in the art. The terms used in the embodiments of the application are only for the purpose of describing the embodiments of the application, and are not intended to limit the application.

[0041] The related data collection and processing in the embodiments of the present application should strictly comply with the requirements of relevant laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of authorization of laws and regulations and the personal information subject.

[0042] Before the embodiments of the present application are further described in detail, the terms and phrases involved in the embodiments of the present application are explained, which are applicable to the following explanations.

[0043] 1) In response to: for indicating the conditions or states on which the operations performed depend, when the dependent conditions or states are met, one or more operations performed can be real-time or have a set delay; in the absence of special instructions, there is no restriction on the execution sequence of multiple operations performed.

[0044] In order to better understand the image recognition system and method based on visual handle provided by the embodiments of the present application, the image recognition system and method based on visual handle in the related art will be described first.

[0045] The visual system uses a high-definition camera to collect images, and then performs subsequent image processing. However, the key information and local details in high-resolution images are often intertwined. The target or scene is easily affected by uncertainty factors such as light, occlusion, angle change, etc., resulting in large processing data volume, complex model, slow processing speed, and recognition errors in complex scenes.

[0046] In view of the problems in the related art, the embodiments of the present application provide an image recognition system and method based on visual handle. First, the visual handle under the target task is obtained. The visual handle is an image with a resolution lower than a preset resolution threshold and an image frequency lower than a preset frequency threshold. Low resolution can effectively reduce the data volume of the image, and low frequency can split the target and scene, use scene prediction to identify the object, and remove the detail redundancy of the identified object. Then, image recognition is performed based on the visual handle to obtain an initial recognition result. When the initial recognition result meets a preset condition, the initial recognition result is determined as the image recognition result of the target task. Therefore, image recognition can be performed through a low-resolution and low-frequency visual handle, which can reduce the calculation data volume, eliminate visual redundancy, reduce the consumption of computing resources, and improve the image recognition speed.

[0047] The embodiments of the present application provide an image recognition system based on visual handle, which is described with reference to Figure 1 , Figure 1is an optional structural schematic diagram of an image recognition system based on a visual handle provided by the embodiment of the present application, the system comprises a visual sensor and a data processing module (not shown in the figure); the visual sensor can be used to collect a visual handle under a target task; here, the visual handle is an image with a resolution lower than a preset resolution threshold and an image frequency lower than a preset frequency threshold; the visual sensor comprises at least one imaging camera; the data processing module can be used to perform image recognition based on the visual handle to obtain an initial recognition result; and in response to the initial recognition result satisfying a preset condition, the initial recognition result is determined as an image recognition result of the target task.

[0048] In the embodiment of the present application, the visual sensor can be a sensor formed by the combination of the non-focus imaging camera and the focus imaging camera. For example, as shown in the figure, the visual sensor can be composed of the non-focus imaging camera 101, the non-focus imaging camera 102, the non-focus imaging camera 103, the focus imaging camera 104, the focus imaging camera 105 and the focus imaging camera 106. It should be noted that the number of imaging cameras in the visual sensor is not limited in the embodiment of the present application, and the arrangement manner is not limited, and these imaging cameras can be arranged in a line or in a matrix. The imaging camera can be a color image sensor or a black-and-white image sensor. The visual sensor can be installed on a fixing device 107, the fixing device 107 can be connected to a base 109 through a rotating shaft 108, the fixing device 107 can rotate around the x1 axis of the rotating shaft 108, and the rotating shaft 108 can rotate around the z2 axis of the base 109. Figure 1

[0049] In some embodiments, the visual sensor can comprise a non-focus imaging camera; the non-focus imaging camera can be used to collect an image with a resolution lower than a preset resolution threshold and an image frequency lower than a preset frequency threshold.

[0050] In some embodiments, the visual sensor can comprise a non-focus imaging camera and a focus imaging camera; the non-focus imaging camera can be used to collect a visual handle with a resolution lower than a preset resolution threshold and an image frequency lower than a preset frequency threshold; and the focus imaging camera can be used to collect first image data of the target task; the image resolution corresponding to the first image data is greater than the preset resolution, and the image frequency is greater than the preset frequency threshold.

[0051] Correspondingly, the data processing module can also be used to, in response to the initial recognition result not satisfying the preset condition, acquire first image data of the target task collected by the focus imaging camera; and perform image recognition based on the visual handle and the first image data to obtain an image recognition result of the target task.

[0052] ​In some embodiments, the system can further include at least one target sensor; the target sensor can be used to collect target data of the target task. The target sensor can be any one or more of the light intensity sensor 110, the electronic compass 111, the gyroscope 112, the real-time clock 113, and the satellite positioning module 114. It should be noted that the number of target sensors is not limited in the embodiments of the present application.

[0053] Correspondingly, the data processing module can be further configured to, in response to the initial recognition result not satisfying the preset condition, acquire at least one target data of the target task collected by at least one target sensor; and perform image recognition based on the visual handle and the at least one target data to obtain an image recognition result of the target task.

[0054] In some embodiments, the visual sensor can include a non-focus imaging camera and a focus imaging camera; the system can further include at least one target sensor; the non-focus imaging camera can be used to collect an image with a resolution lower than a preset resolution threshold and an image frequency lower than a preset frequency threshold to obtain the visual handle; the focus imaging camera can be used to collect first image data of the target task; the first image data corresponds to an image with a resolution greater than the preset resolution and an image frequency greater than the preset frequency threshold; and the target sensor can be used to collect target data of the target task.

[0055] Correspondingly, the data processing module can be further configured to, in response to the initial recognition result not satisfying the preset condition, acquire the first image data collected by the focus imaging camera and at least one target data of the target task collected by at least one target sensor; and perform image recognition based on the visual handle, the first image data, and the at least one target data to obtain an image recognition result of the target task.

[0056] In some embodiments, the visual sensor can include a focus imaging camera; the system can further include a downsampling module and a convolution blur module; the focus imaging camera can be used to collect first image data of the target task; the first image data corresponds to an image with a resolution greater than a preset resolution and an image frequency greater than a preset frequency threshold; the convolution blur module can be used to perform at least one convolution blur on the first image data to obtain a convolution blur result; and the downsampling module can be used to perform at least one downsampling on the convolution blur result to obtain the visual handle.

[0057] The image recognition method based on the visual handle provided in the embodiments of the present application can be applied to electronic devices such as notebook computers, tablet computers, desktop computers, etc. The embodiments of the present application do not make any limitation on the specific type of electronic devices.

[0058] The image recognition method based on the visual handle provided in the embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0059] The image recognition method based on the visual handle provided by the embodiments of the present application can be executed by an electronic device, which can be a server or a terminal, i.e., the image recognition method based on the visual handle provided by the embodiments of the present application can be executed by a server, a terminal, or an interaction between the server and the terminal.

[0060] Figure 2 is an optional flowchart of the image recognition method based on the visual handle provided by the embodiments of the present application, which can be applied to an electronic device. In the following, the electronic device is taken as a server for example. As shown in Figure 2 , the method comprises the following steps S101 to S103:

[0061] Step S101, acquiring a visual handle under a target task.

[0062] Here, the visual handle is an image with a resolution lower than a preset resolution threshold and an image frequency lower than a preset frequency threshold.

[0063] In the embodiments of the present application, the target task can be any task related to image processing, such as pedestrian detection, obstacle detection, expression detection, traffic sign detection, animal detection, product defect detection, gesture detection, license plate detection, etc. The server can acquire the corresponding visual handle (VisualHandle) for different target tasks, which is an image with a resolution lower than a preset resolution threshold and an image frequency lower than a preset frequency threshold. The resolution threshold and the frequency threshold can be pre-set according to actual conditions. Referring to Figure 3 Figure 3 is a schematic diagram of an image provided by the embodiments of the present application, wherein, Figure 3 the (a) part in is an image of the letter A acquired by a high-resolution camera, Figure 3 the (b) part in is a visual handle, which can be acquired by a low-resolution camera in an out-of-focus state. As can be seen, the visual handle removes the influence of details such as the small ball and background text in the (a) part in Figure 3 but still clearly reflects the (a) part in Figure 3 ​The combination relationship A between the elements in the (b) part image in the (a) part image. The visual handle can be an abstract of all the contents in the image obtained by optical imaging technology or image processing method. The visual handle can be a breakthrough point for grasping and understanding all the contents in the image. The role of the visual handle mainly reflects two aspects: one is that the visual handle itself can realize the detection and recognition of the target or the understanding of the relationship between objects, and the other is that it can cooperate with the detailed information to know the internal state of the object in the real world. The visual handle contains fewer pixels than high-definition images. The visual handle can aggregate some detailed contents into larger contents and ignore the influence of details in the image. There are many ways to obtain the visual handle, including obtaining from network resources, obtaining from storage devices, obtaining through image generation technology, real-time acquisition through imaging cameras, and obtaining from public image data sets, etc.

[0064] In step S102, image recognition is performed based on the visual handle to obtain an initial recognition result.

[0065] In the embodiments of the present application, the server can obtain an initial recognition result by performing image recognition on the visual handle. After obtaining the visual handle, the server can perform preprocessing, such as size adjustment, normalization, and data enhancement. The server can load a pre-trained image model, and the image model is used to infer the visual handle. The image model can be a convolutional neural network model. The preprocessed visual handle is input into the image model to obtain the output of the image model, i.e., the initial recognition result. For example, for a classification task, the output of the image model is usually a probability distribution, and the corresponding class label can be obtained according to the highest probability value output by the image model. For a recognition task, after the image model recognizes the visual handle, the output result can be the position of one or more key objects in the visual task.

[0066] In step S103, the initial recognition result is determined as the image recognition result of the target task in response to the initial recognition result satisfying a preset condition.

[0067] In the embodiments of the present application, a preset condition can be set in advance. When the initial recognition result satisfies the preset condition, the initial recognition result can be determined as the image recognition result of the target task. For example, for a classification task, the initial recognition result can include a class label and a corresponding probability value. The preset condition can be set according to the probability value, for example, the preset condition is set to be greater than 85%. For example, if the initial recognition result for the visual handle is cat-90%, then cat-90% can be taken as the image recognition result of the target task.

[0068] The embodiment of the present application can obtain a visual handle under a target task. The visual handle is an image with a resolution lower than a preset resolution threshold and an image frequency lower than a preset frequency threshold. The low resolution can effectively reduce the data amount of the image, and the low frequency can effectively reduce the interference of high-frequency noise. Then, image recognition is performed based on the visual handle to obtain an initial recognition result. When the initial recognition result meets a preset condition, the initial recognition result is determined as an image recognition result of the target task. In this way, image recognition can be performed through a visual handle with low resolution and low frequency, the amount of calculation data can be reduced, the consumption of calculation resources can be reduced, and the image recognition speed can be improved.

[0069] Figure 4 Another optional flowchart of the image recognition method based on the visual handle provided by the embodiment of the present application is shown in FIG. 2. As shown in FIG. 2, the method includes the following steps S201 to S216. Figure 4

[0070] In step S201, the terminal receives an image recognition operation input by a user.

[0071] The image recognition operation includes a selection operation or an input operation. The selection operation is used to select a task to be executed, or the input operation is used to input a task identifier of a target task to be executed.

[0072] In step S202, the terminal encapsulates the task identifier of the task to be executed into an image recognition request.

[0073] The image recognition request is used to request the server to perform image recognition on the task to be executed.

[0074] In step S203, the terminal sends the image recognition request to the server.

[0075] In the embodiment of the present application, the terminal sends the image recognition request to the server to request the server to perform image recognition on the task to be executed.

[0076] In step S204, the server acquires a visual handle under the target task in response to the image recognition request.

[0077] ​In some embodiments, the "obtaining a visual handle under the target task" in step S204 can be implemented in the following manners: based on the task type of the target task, the preset frequency threshold and the preset resolution threshold, determining at least one target non-focus imaging camera in a working state and a distance value between a lens module and an image sensor in the at least one target non-focus imaging camera from the electronic device; then, based on the distance value, performing image acquisition on the target task by the at least one target non-focus imaging camera to obtain the visual handle; or, based on the task type of the target task, the preset frequency threshold and the preset resolution threshold, determining at least one target focus imaging camera in a working state and an aperture value of the at least one target focus imaging camera from the electronic device; then, based on the aperture value, performing image acquisition on the target task by the at least one target focus imaging camera to obtain the visual handle; or, determining a target focus imaging camera in a working state from the electronic device; then performing image acquisition on the target task by the target focus imaging camera to obtain a target image; finally, based on the task type of the target task, the preset frequency threshold and the preset resolution threshold, performing at least one convolution blur and down-sampling on the target image to obtain the visual handle.

[0078] In the embodiments of the present application, the server can determine at least one target non-focus imaging camera in a working state and a distance value between a lens module and an image sensor in the at least one target non-focus imaging camera from the electronic device according to the task type of the target task, the preset frequency threshold and the preset resolution threshold, and then adjust the distance between the lens module and the image sensor according to the distance value, and perform image acquisition on the target task by the target non-focus imaging camera with the adjusted distance between the lens module and the image sensor to obtain the visual handle. The number of non-focus imaging cameras corresponding to different task types can be preset, for example, one target non-focus imaging camera can be selected for a video monitoring task type, and multiple target non-focus imaging cameras can be selected for an automatic driving task type. The server determines the focus value of the non-focus imaging camera according to the preset frequency threshold and the preset resolution threshold. The distance value between the lens module and the image sensor is determined according to the focus value, and then the current direction in the coil is determined according to the distance value. Referring to Figure 5 , Figure 5 is an optional structural schematic diagram of the non-focus imaging camera provided by the embodiments of the present application. The magnetic field can be generated by the current in the coil 503, and different directions of the current will make the coil 503 generate magnetic forces in different directions. The same or opposite magnetic force of the permanent magnet is generated by the coil to push the lens module 501 to do reciprocating motion. In this way, the moving direction of the lens module 501 can be determined according to the distance value between the lens module 501 and the image sensor 502, and then the current direction in the coil 503 is determined. Referring to Figure 6 , Figure 6is an optional schematic diagram of a non-focus imaging camera provided by an embodiment of the present application. The lens module 601 is responsible for focusing light, and the light focused by the lens module 601 can fall on the image sensor 602. The image sensor 602 can convert the light signal into an electrical signal to form a digital image. The processor 603 receives data from the image sensor and performs further processing. The micro control unit 604 is responsible for coordinating the work of various components, receiving information from the processor, and controlling the motor drive and other components as needed. The motor drive 605 receives instructions from the micro control unit 604 to control the operation of the motor. The motor drive 605 is responsible for adjusting the position of the lens module 601 to change the focal length (i.e. the distance between the lens module and the image sensor). The motor 606 is a component that performs mechanical movement, and moves the lens module 601 according to the instructions of the motor drive to achieve automatic focusing. The optical linear encoder 607 is used to accurately measure the position of the lens module 601, and provides feedback to the micro control unit 604 so that the micro control unit 604 knows the current focal length, and the micro control unit 604 adjusts the motor according to the current focal length to achieve accurate focusing. Finally, when the distance between the lens module 501 and the image sensor 502 reaches the distance value, the target non-focus imaging camera can be used to collect images of the target task to obtain a visual handle.

[0079] The server can also determine the number of target focus imaging cameras according to the task type of the target task, select target focus imaging cameras in a working state from the electronic device according to the number, and then determine the aperture value of the target focus imaging camera according to the preset frequency threshold and the preset resolution threshold. The server can control the aperture through the motor, control the amount of light entering through the different opening and closing sizes of the aperture, and achieve the effect of extracting the outline of the visual blur, that is, a small aperture can obtain a blurred effect. The server can control the opening and closing size of the aperture through the motor according to the aperture value, and collect images of the target task through the target focus imaging camera to obtain a visual handle.

[0080] The server can also determine the target focus imaging camera in a working state from the electronic device; collect images of the target task through the target focus imaging camera to obtain a target image. Then, according to the task type of the target task, the preset frequency threshold and the preset resolution threshold, the target image is subjected to convolution blur processing and down-sampling processing to obtain a visual handle. For example, the target image is subjected to Gaussian pyramid processing until the resolution of the processed image is lower than the preset frequency threshold and the frequency of the processed image is lower than the preset resolution threshold. For example, during the navigation walking process of the robot, a non-focus imaging camera can be installed on the robot device to collect a visual handle in real time.

[0081] It should be noted that the visual handle is directly obtained by the non-focus imaging camera and the focus imaging camera focus method, and the acquisition efficiency is higher than the image pyramid method.

[0082] In the above manner, the visual handle can be obtained in multiple ways, and the adaptability of the system to different scenes is improved.

[0083] In step S205, the server performs image recognition based on the visual handle to obtain an initial recognition result.

[0084] In some embodiments, step S205 can be implemented in the following manner: performing image recognition based on the visual handle by a pre-trained model to obtain an initial recognition result. Here, the pre-trained model is obtained by training a visual handle set; the visual handle set includes at least one sample visual handle.

[0085] In the embodiments of the present application, a model can be pre-trained, and the model can be obtained by training a visual handle set. The visual handle set includes at least one sample visual handle. The sample visual handle can be obtained by the implementation manner in step S204, which will not be described here. The server can perform image recognition on the visual handle by the pre-trained model to obtain an initial recognition result. It should be noted that the visual handle obtained in step S204 can be saved as a sample visual handle in the server.

[0086] In the above manner, the image recognition can be performed by the pre-trained model, and the initial recognition result can be quickly obtained.

[0087] In step S206, the server determines the initial recognition result as the image recognition result of the target task in response to the initial recognition result satisfying a preset condition.

[0088] It should be noted that step S206 is the same as step S103 described above, and the implementation details of step S206 will not be described here. For example, the preset condition is set to a probability value greater than 85%. The initial recognition result of the visual handle collected during the navigation walking process of the robot is cat-90%, and then cat-90% can be used as the image recognition result of the target task.

[0089] In step S207, the server obtains first image data of a target task in response to the initial recognition result not satisfying the preset condition.

[0090] In the embodiments of the present application, the first image data is an image with a resolution higher than or equal to a preset resolution threshold and an image frequency higher than or equal to a preset frequency threshold. The first image data can be acquired in real time by a focus imaging camera or obtained from a public data set. When the initial recognition result does not meet the preset condition, the server can obtain the first image data of the target task. For example, the preset condition is set as a probability value greater than 85%. The initial recognition result of the visual handle collected during the robot navigation walking process is cat-70%, and then the robot can further include a focus imaging camera, and the server can call the focus imaging camera to collect the first image data.

[0091] In step S208, the server performs image recognition based on the visual handle and the first image data to obtain an image recognition result of the target task.

[0092] In some embodiments, referring to Figure 7 , Figure 7 It is shown that step S208 can be implemented by the following steps S2081 to S2085:

[0093] In step S2081, a position prediction model corresponding to the task type of the target task is determined.

[0094] In the embodiments of the present application, different position prediction models can be set in advance for different task types, for example, a face recognition task corresponds to a face recognition position prediction model, and a target detection task corresponds to a target detection position prediction model. The position prediction model can be a pre-trained deep learning model. Then the server can determine the corresponding position prediction model according to the task type of the target task.

[0095] In step S2082, the position prediction model is used to perform position prediction on the visual handle to obtain at least one position information of the visual handle.

[0096] In the embodiments of the present application, the server can obtain at least one position information of the visual handle by performing position prediction on the visual handle through the position prediction model. The position information can be used to represent the position of the local image in the visual handle. For example, a coordinate system can be used, and a rectangular box can be used to represent the size of the local image and the position of the local image in the visual handle. The coordinates are in pixels, and the boundary box is defined by two opposite coordinate points in the format (x1, y1, x2, y2), where (x1, y1) is the coordinate of the upper left corner of the rectangle, and (x2, y2) is the coordinate of the lower right corner of the rectangle. For example, the size of the visual handle is 640*480, and the determined position information can be represented as (120, 80, 220, 180).

[0097] In step S2083, the first image data is image-registered according to the visual handle, to obtain a pixel position mapping relationship between the visual handle and the first image data.

[0098] In the embodiments of the present application, the server can image-register the first image data according to the visual handle, so that the first image data can be aligned with the visual handle, so that the visual handle and the first image data have a consistent reference coordinate system in space, thereby obtaining the pixel position mapping relationship between the visual handle and the first image data. For example, first, feature points are extracted from the visual handle and the first image data, respectively. These feature points can be corner points, edge points, etc. The feature points in the two images can be matched by a feature point extraction algorithm. After the feature point matching, a series of feature point pairs are obtained. The feature point pairs can be used to estimate the geometric transformation between the two images. Since the resolutions of the visual handle and the first image data are different, the geometric transformation can include scaling, rotation, and translation, etc. A transformation model such as an affine transformation or a homography transformation can be fitted according to these transformations. Through the transformation model, the corresponding position of each pixel point in the visual handle in the first image data can be calculated, that is, the pixel position mapping relationship between the visual handle and the first image data, which can be represented by a transformation matrix.

[0099] In step S2084, the first image data is image-cropped according to the at least one position information and the pixel position mapping relationship, to obtain at least one first local image.

[0100] In the embodiments of the present application, after obtaining the pixel position mapping relationship between the visual handle and the first image data, the corresponding position information of the at least one position information in the first image data can be determined through the pixel position mapping relationship, and the first image data is cropped according to the corresponding position information, to obtain at least one first local image.

[0101] In step S2085, image recognition is performed based on the visual handle and the at least one first local image, to obtain an image recognition result of the target task.

[0102] In some embodiments, step S2085 can be implemented in the following manner: feature extraction is performed on the visual handle and the at least one first local image respectively, to obtain a first feature vector and at least one second feature vector; the second feature vector includes a texture feature vector; vector fusion is performed on the first feature vector and the at least one second feature vector, to obtain a fused feature vector; and image recognition is performed based on the fused feature vector, to obtain an image recognition result of the target task.

[0103] In the embodiments of the present application, the server can perform feature extraction on the visual handle and the at least one first local image respectively, and obtain a first feature vector and at least one second feature vector, the first feature vector can be used to represent shape features and color features, etc., and the second feature vector can include a texture feature vector. Then, the first feature vector and the at least one second feature vector are fused to obtain a fused feature vector. Finally, the server performs image recognition on the fused feature vector through a pre-trained image model to obtain an image recognition result of the target task. For example, through the visual handle, it can be recognized that the shape of the object in the image is a square shape, and the color is brown. In combination with the first local image, it can be recognized that the texture of the object is wood grain. Therefore, the pre-trained image model can quickly infer that the object is wood.

[0104] In the above manner, image recognition can be accurately performed through a small amount of image data, the amount of calculation data can be reduced, the consumption of computing resources can be reduced, and the speed of image recognition can be improved.

[0105] In some embodiments, step S2085 can also be implemented in the following manner: performing target region cropping on the first image data to obtain a target region image of the first image data; and performing image recognition based on the visual handle, the at least one first local image, and the target region image to obtain an image recognition result of the target task.

[0106] In the embodiments of the present application, the target region can be a region corresponding to preset position information or a region corresponding to position information of a random position. The preset position information and the position information of the random position can also be obtained when step S2082 is performed. The first image data can be cropped according to the preset position information to obtain a second local image, which can be implemented in the manner of steps S2082 to S2083. Alternatively, the server can generate at least one position information of a random position in the first image data, and then crop the first image data according to the at least one position information of the random position to obtain a third local image, which can be implemented in the manner of steps S2082 to S2083. Finally, at least one of the second local image and the third local image is determined as the target region image.

[0107] Referring to Figure 8 , Figure 8is a schematic diagram of a principle for obtaining the first local image provided in the embodiment of the present application. According to the image registration of the first image data (not shown in the figure) by the visual handle 804, the pixel position mapping relationship between the visual handle 804 and the first image data can be obtained. Then, the first image data can be image cropped according to the obtained prediction position information 801 (i.e., at least one position information), the center position information 802 (i.e., the preset position information), the position information 803 of the random position, and the pixel position mapping relationship, to obtain the corresponding local images including at least one first local image 805, a second local image 806, and a third local image 807. Then, the feature extraction is performed on the visual handle 804 and the local images (including at least one first local image 805, a second local image 806, and a third local image 807) respectively, to obtain the global feature 809 (i.e., a first feature vector) and the local feature 808 (i.e., at least one second feature vector) correspondingly.

[0108] It should be noted that, in the embodiment of the present application, when the image recognition is performed by using the first image data, the local image of the first image data can be used. Therefore, when the image resolution of the first image data is high, the image size of the local image of the first image data can still be relatively small. Therefore, the image resolution of the first image data can not be limited by the computer computing power and the like.

[0109] In the above manner, the determination manner of the target region can be flexibly selected, the adaptability of image processing is enhanced, the diversity of image data is increased, the accuracy of the image recognition result can be improved, the amount of calculation data can be reduced, the consumption of computing resources can be reduced, and the image recognition speed can be improved.

[0110] In step S209, the server determines at least one target sensor in a working state from the electronic device based on the task type of the target task in response to the initial recognition result not satisfying the preset condition.

[0111] In the embodiment of the present application, the target sensor corresponding to the task type can be preset, for example, when the task type is indoor robot navigation, the corresponding target sensor can be a gyroscope. When the initial recognition result does not satisfy the preset condition, the server can also determine at least one target sensor in a working state from the electronic device according to the task type of the target task.

[0112] In step S210, the server collects data for the target task by using the at least one target sensor, to obtain at least one target data.

[0113] In the embodiment of the present application, continuing to refer to Figure 1If the target sensor is the light intensity sensor 110, the server can read the data of the light intensity sensor to obtain the light level of the current environment. If the target sensor is the electronic compass 111, the server can read the output of the electronic compass to obtain the orientation. If the target sensor is the gyroscope 112, the server can collect the data of the gyroscope to analyze the rotation dynamics of the electronic device. If the target sensor is the real-time clock 113, the server can read the data of the real-time clock to obtain the current date and time. If the target sensor is the satellite positioning module 114, the server can receive the data of the satellite positioning module to obtain the accurate position of the electronic device.

[0114] In step S211, the server performs image recognition based on the visual handle and the at least one target data to obtain an image recognition result of the target task.

[0115] In the embodiments of the present application, if the target data is the light intensity, in the context of robot navigation transferring from indoor to outdoor, the visual handle can identify that there is no obstacle around the current environment, and the robot can move to the direction with stronger light intensity (i.e., the light intensity of outdoor is greater than that of indoor). If the target data is the current orientation of the electronic compass, in the context of the navigation of the sweeping robot sweeping in the indoor environment, the visual handle can identify that there is an obstacle in the south direction and an obstacle in the north direction in the current environment, and the sweeping robot can determine that the obstacle in the south direction is a sofa and the obstacle in the north direction is a dining table (i.e., the sofa is in the south direction of the room and the dining table is in the north direction of the room in the indoor environment) according to the orientation. If the target data is the current date and time, in the automatic driving scenario, the visual handle can identify that there is an obstacle in front, the obstacle has a large volume, and the color is brown, and combined with the current time being winter, it can be inferred that the obstacle is a tree. Correspondingly, if the obstacle has a large volume and the color is green, combined with the current time being summer, it can also be inferred that the obstacle is a tree. If the target data is the positioning data, in the automatic driving scenario, the scene is in a campus, the visual handle can identify that there is an obstacle in front, the obstacle has a large volume, the positioning data can be obtained, and the specific fixed buildings (e.g., dormitory, library, canteen, etc.) can be determined according to the positioning data, and it can be inferred that the obstacle is a library (generally, the positioning data of the building does not change).

[0116] In some embodiments, if the image recognition based on the visual handle and the at least one target data cannot obtain the image recognition result of the target task that meets the preset condition, the visual sensor can be adjusted based on the target sensor to reacquire the visual handle. And the new visual handle is used to perform image recognition again. For example, the visual sensor can be rotated by using the gyroscope sensor to adjust the viewing angle of the visual sensor to reacquire the visual handle.

[0117] Through steps S209 to S211, when the initial recognition result does not meet the preset conditions, the server can automatically determine and use at least one target sensor in a working state from the electronic device to collect data, obtain key target data, and then combine the visual handle and target data to accurately perform image recognition, which can reduce the amount of calculated data, reduce computing resource consumption, and improve image recognition speed.

[0118] In step S212 , in response to the initial recognition result not meeting the preset condition, the server obtains first image data of the target task, and determines at least one target sensor in a working state from the electronic device based on the task type of the target task.

[0119] It should be noted that step S212 is the same as the above-mentioned step S207 and step S209, and the implementation details of step S212 are not repeated in this embodiment of the application.

[0120] In step S213 , the server collects data on the target task through at least one target sensor to obtain at least one target data.

[0121] It should be noted that step S213 is the same as the above-mentioned step S210, and the implementation details of step S213 are not repeated in this embodiment of the application.

[0122] In step S214 , the server performs image recognition based on the visual handle, the first image data, and at least one target data to obtain an image recognition result of the target task.

[0123] In some embodiments, step S214 can also be implemented in the following ways: determining a position prediction model corresponding to the task type; performing position prediction on the visual handle through the position prediction model to obtain at least one position information of the visual handle; performing image cropping on the first image data according to the at least one position information to obtain at least one first partial image; performing image recognition on the visual handle and the at least one first partial image based on at least one target data to obtain the image recognition result of the target task.

[0124] In the embodiments of the present application, the server can first determine a position prediction model corresponding to the task type, then perform position prediction on the visual handle through the position prediction model to obtain at least one position information of the visual handle, then perform image cropping on the first image data according to the at least one position information to obtain at least one first local image, and finally perform image recognition on the visual handle and the at least one first local image based on the at least one target data to obtain an image recognition result of the target task. If the target data is the light intensity, in the context of robot navigation transferring from indoor to outdoor, the current surrounding can be recognized as having no obstacle through the visual handle, and the texture information of the surrounding environment can be obtained through the first local image, and the robot can move to the direction where the light intensity is stronger and the environmental texture information includes the texture information of natural objects (i.e., the light intensity outdoors is greater than that indoors). If the target data is the direction of the current electronic device, in the context of the robot cleaner navigating in the indoor cleaning, it can be recognized through the visual handle that there is an obstacle in the south and north directions in the current environment, and the texture information of the two obstacles can be obtained through the first local image, such as the texture information of the obstacle in the south direction including leather texture information, and the texture information of the obstacle in the north direction including wood grain texture information, and the robot cleaner can determine that the obstacle in the south direction is a sofa and the obstacle in the north direction is a dining table according to the direction and the texture information (i.e., the sofa is in the south direction of the room and the dining table is in the north direction of the room in the indoor environment). If the target data is the current date and time, in the automatic driving scene, it can be recognized through the visual handle that there is an obstacle in front, the obstacle has a large volume, and the color is brown, and the texture information of the obstacle including wood grain texture information can be obtained through the first local image, and combined with the current time being winter, it can be inferred that the obstacle is a tree. Correspondingly, if the obstacle has a large volume and the color is green, the texture information of the obstacle including wood grain texture information can be obtained through the first local image, and combined with the current time being summer, it can also be inferred that the obstacle is a tree.

[0125] In the above manner, the first local image containing detailed information can be cropped from the first image data, and then the visual handle and the first local image are subjected to image recognition in combination with the at least one target data, so as to obtain an accurate target task image recognition result, thereby reducing the amount of calculation data, reducing the consumption of computing resources, and improving the image recognition speed.

[0126] Through steps S212 to S214, when the initial recognition result does not meet the preset condition, the server can obtain the first image data, determine at least one target sensor in a working state from the electronic device, and collect data by using the at least one target sensor to obtain key target data, and then accurately perform image recognition in combination with the visual handle, the first image data and the target data, thereby reducing the amount of calculation data, reducing the consumption of computing resources, and improving the image recognition speed.

[0127] S215, the server sends the image recognition result to the terminal.

[0128] S216, the terminal displays the image recognition result.

[0129] In the following, an exemplary application of the embodiments of the present application in a practical application scenario will be described.

[0130] In the following, an example of two imaging cameras included in an electronic device will be described. The non-focus imaging camera focuses on extracting overall information, and the focus imaging camera extracts details, and the image information acquired by the two can be fused to realize image recognition.

[0131] Referring to Figure 9 , Figure 9 is another optional structure diagram of the image recognition system based on visual handle provided by the embodiments of the present application. The non-focus imaging camera 901 can be mounted on the fixing device 905, the fixing device 905 can be connected to the second rotating shaft 909 through the first rotating shaft 903, the fixing device 905 can rotate around the x3 axis of the first rotating shaft 903, and the first rotating shaft 903 can rotate around the z5 axis of the second rotating shaft 909. The second rotating shaft 909 can be mounted with the base 910, and the second rotating shaft 909 can rotate around the z6 axis of the base 910. The focus imaging camera 902 can be mounted on the fixing device 906, the fixing device 906 can be connected to the second rotating shaft 909 through the first rotating shaft 904, the fixing device 906 can rotate around the x4 axis of the first rotating shaft 904, and the first rotating shaft 904 can rotate around the z5 axis of the second rotating shaft 909. The second rotating shaft 909 can be mounted with the base 910, and the second rotating shaft 909 can rotate around the z6 axis of the base 910. The sensors 907 and 908 can be gyroscopes.

[0132] The embodiments of the present application can be applied to various cases, for example, one case can refer to Figure 10 , Figure 10 is an optional scene diagram of the image recognition method based on visual handle provided by the embodiments of the present application. As shown in Figure 10 , Figure 10 may be the layout of a large event room (which can hold a dance party or other gatherings). The dots in the figure represent the crowd gathered in the event room. The important application direction of the image recognition system based on visual handle mentioned by the embodiments of the present application can be towards self-walking robots. The self-walking robot can participate in the human gathering and help organize the human gathering. The self-walking robot can monitor the personnel and articles in the room by accessing the data of the visual sensor (i.e. visual sensor) installed in the center of the room top. Figure 10 part (a) in the figure is a layout of the event room formed by the flow of personnel, Figure 10Part (b) shows another layout of the activity room due to the flow of people. Figure 10 Part (a) and Figure 10 As shown in part (b) of the figure, the scene will form a complex layout due to the flow of people and the changes in the placement of objects. When a self-propelled robot locates key parts in a room (for example, looking for doors, windows, emergency exits, or projection screens), if a conventional camera is used to capture images of the indoor situation, it must traverse all the grids in the image each time before finally locating the door, emergency exit, or projection screen. Figure 10 Part (c) of the

[0133] Using the visual sensor 1001 provided in the embodiment of the present application, Figure 10 Part (d) in the image is the visual handle. In particular, after using the visual handle, the visual handle has a powerful detail removal capability. It can directly remove the details of movable objects such as people, or some static objects from the image, which greatly simplifies the layout of the room. Figure 10 As shown in part (d) of the figure, based on the visual handles, a self-propelled robot can predict the location of key objects in the room, such as doors, emergency exits, and projection screens, through simple training. Furthermore, by leveraging the visual handles and the focus imaging camera, the robot can find and locate other objects in the room. This can assist with the entry of people at the start of a party. In an emergency, people can be organized to escape through the main door or emergency exit. Both visual handles and local details can be captured simultaneously using multiple cameras with different resolutions, which can provide additional benefits. For example, after aligning the visual handles with the image captured by the focus imaging camera, the position of the projection screen can be predicted using the visual handles. The area surrounding the projection screen can then be cropped directly from the image captured by the focus imaging camera. If the focus imaging camera (e.g., with a resolution of 8 megapixels) captures the area as not only the projection screen but also a book on the ground within the predicted frame, the title cannot be recognized due to insufficient resolution. At this time, another focus imaging camera (for example, with a resolution of 12 million pixels) can be called to capture images, and the captured images are aligned with the visual handles, from which the prediction area is cropped for recognition.

[0134] The target sensor in the embodiment of the present application can assist in obtaining simpler visual handles to improve the predictive ability of the visual handles. In the above scenario, if the self-propelled robot does not directly use the visual sensor data at the top of the room for positioning. Instead, it uses the data collected from the visual sensor from the robot's perspective (close to the human perspective) for positioning. Even if the visual handle is obtained from the robot's perspective, a large amount of personnel information may remain in the obtained visual handle. If the self-propelled robot can adjust the robot's perspective with the assistance of the gyroscope on the visual sensor and turn to collect the structural information at the top of the room, the information of mobile personnel on the ground can be removed, and then the position of the windows and doors of the room can be further determined by the light sensor (for example, the window is in a high-brightness area and the door is in a low-brightness area), and the electronic compass can also participate in the determination of the north-south direction. Both the light sensor data and the electronic compass data can be used as constraints to assist the visual handle in predicting the key areas in the room.

[0135] For example, another case can be seen Figure 11 , Figure 11 This is another optional scene diagram of the image recognition method based on visual handles provided by the embodiment of the present application. For example, in an outdoor scene, an autonomous vehicle 1101 is driving on the road. If conventional pure visual technology is used, the autonomous vehicle will obtain a bird's-eye view 1102. The autonomous vehicle needs to process all pixels of the image in each small box in turn before obtaining the key prediction box 1104 (i.e., the key area in front) required by the autonomous vehicle. Through the embodiment of the present application, the detailed information on the road surface can be blurred to obtain Figure 11 The visual handle 1103 shown in part (c) of the image is shown. Based on the visual handle 1103, the pre-trained visual handle dataset can be used to obtain the key prediction frame in front of the vehicle in the case of following the vehicle. The prediction frame is then aligned with the image captured by the focus imaging camera to obtain the details that the autonomous vehicle needs to pay attention to. In addition, cameras with different resolutions can obtain details at different resolutions according to different needs, thereby identifying lane lines, vehicle outlines, license plates, metal textures on the vehicle surface, and subtle changes in vehicle taillights. In addition, the current time can be obtained in conjunction with the real-time clock in the target sensor. If it is winter, it will assist the autonomous vehicle in determining that the white object on the ground may be snow. When the lane line is covered with snow, the forward driving position is directly predicted based on the predicted position of the visual handle, breaking away from the limitations of the ground lane line.

[0136] In summary, first, the embodiment of the present application takes visual handle (blurred image) as the main means of image recognition method. Compared with the multi-scale method (i.e. the method of taking clear image as the center and blurred image as the auxiliary in image recognition), the embodiment of the present application establishes a data set by using the blurred image corresponding to the visual handle, performs image recognition by the visual handle, or predicts the details in the image by the visual handle, and the visual handle has good resistance to light and angle transformation. Secondly, the embodiment of the present application can directly obtain the visual handle by using the non-focus method, and the acquisition efficiency will be greatly improved compared with the image pyramid method. Finally, when high-definition image needs to be used, the obtained high-definition image can not be limited by the image resolution, and the image resolution of the high-definition image far exceeds the limitation of the image resolution (5-8 million pixels) of the automatic driving system.

[0137] The image recognition method based on the visual handle according to the above embodiment, Figure 12 A structure block diagram of an image recognition device based on a visual handle provided by the embodiment of the present application is shown, the image recognition device based on the visual handle 1200 can be a device in an electronic device (for example, a server), the image recognition device based on the visual handle 1200 can be realized in a software manner, which can be software in the form of programs and plug-ins, and includes the following software modules: a visual handle module 1201, an image recognition module 1202 and a result determination module 1203. These modules are logical, and therefore can be combined or further split according to the realized functions.

[0138] The visual handle module 1201 is configured to obtain a visual handle under a target task; the visual handle is an image with a resolution lower than a preset resolution threshold and an image frequency lower than a preset frequency threshold; the image recognition module 1202 is configured to perform image recognition based on the visual handle to obtain an initial recognition result; and the result determination module 1203 is configured to determine the initial recognition result as an image recognition result of the target task in response to the initial recognition result meeting a preset condition.

[0139] In some embodiments, the image recognition module 1202 is further configured to perform image recognition on the visual handle by a pre-trained model to obtain an initial recognition result; wherein the pre-trained model is obtained by training a visual handle set; and the visual handle set includes at least one sample visual handle.

[0140] In some embodiments, the visual handle module 1201 is further configured to: determine, based on the task type of the target task, the preset frequency threshold, and the preset resolution threshold, at least one target non-focus imaging camera in an active state and a distance value between a lens module and an image sensor in the at least one target non-focus imaging camera from the electronic device; perform image collection on the target task by the at least one target non-focus imaging camera based on the distance value to obtain the visual handle; or determine, based on the task type of the target task, the preset frequency threshold, and the preset resolution threshold, at least one target focus imaging camera in an active state and an aperture value of the at least one target focus imaging camera from the electronic device; perform image collection on the target task by the at least one target focus imaging camera based on the aperture value to obtain the visual handle; or determine, from the electronic device, a target focus imaging camera in an active state; perform image collection on the target task by the target focus imaging camera to obtain a target image; and perform at least one convolution blur and down-sampling on the target image based on the task type of the target task, the preset frequency threshold, and the preset resolution threshold to obtain the visual handle.

[0141] In some embodiments, the result determination module 1203 is further configured to, in response to the initial recognition result not satisfying the preset condition, acquire first image data of the target task; and perform image recognition based on the visual handle and the first image data to obtain an image recognition result of the target task.

[0142] In some embodiments, the result determination module 1203 is further configured to determine a position prediction model corresponding to the task type of the target task; perform position prediction on the visual handle by the position prediction model to obtain at least one position information of the visual handle; perform image registration on the first image data according to the visual handle to obtain a pixel position mapping relationship between the visual handle and the first image data; perform image cropping on the first image data according to the at least one position information and the pixel position mapping relationship to obtain at least one first local image; and perform image recognition based on the visual handle and the at least one first local image to obtain an image recognition result of the target task.

[0143] In some embodiments, the result determination module 1203 is further configured to perform feature extraction on the visual handle and the at least one first partial image respectively, to obtain a first feature vector and at least one second feature vector; the second feature vector includes a texture feature vector; perform vector fusion on the first feature vector and the at least one second feature vector to obtain a fused feature vector; and perform image recognition based on the fused feature vector to obtain the image recognition result of the target task.

[0144] In some embodiments, the result determination module 1203 is further configured to perform target region cropping on the first image data to obtain a target region image of the first image data; and perform image recognition based on the visual handle, the at least one first partial image, and the target region image to obtain the image recognition result of the target task.

[0145] In some embodiments, the result determination module 1203 is further configured to perform image cropping on the first image data according to preset position information to obtain a second partial image; generate position information of at least one random position in the first image data; perform image cropping on the first image data according to the position information of the at least one random position to obtain a third partial image; and determine at least one of the second partial image and the third partial image as the target region image.

[0146] In some embodiments, the result determination module 1203 is further configured to, in response to the initial recognition result not satisfying the preset condition, determine at least one target sensor in a working state from the electronic device based on a task type of the target task; perform data collection on the target task through the at least one target sensor to obtain at least one target data; and perform image recognition based on the visual handle and the at least one target data to obtain the image recognition result of the target task.

[0147] In some embodiments, the result determination module 1203 is further configured to, in response to the initial recognition result not satisfying the preset condition, obtain first image data of the target task, and determine at least one target sensor in a working state from the electronic device based on a task type of the target task; perform data collection on the target task through the at least one target sensor to obtain at least one target data; and perform image recognition based on the visual handle, the first image data, and the at least one target data to obtain the image recognition result of the target task.

[0148] In some embodiments, the result determination module 1203 is also used to determine a position prediction model corresponding to the task type; perform position prediction on the visual handle through the position prediction model to obtain at least one position information of the visual handle; perform image cropping on the first image data according to the at least one position information to obtain at least one first partial image; perform image recognition on the visual handle and the at least one first partial image based on at least one target data to obtain the image recognition result of the target task.

[0149] It should be noted that the description of the device embodiment of the present application is similar to the description of the method embodiment described above, and has similar beneficial effects as the method embodiment, so it will not be repeated. For technical details not disclosed in the device embodiment, please refer to the description of the method embodiment of the present application for understanding.

[0150] An embodiment of the present application provides an electronic device, Figure 13 Schematic diagram of the structure of the electronic device provided in the embodiment of the present application. Figure 13 As shown, the electronic device 130 includes: at least one processor 131 ( Figure 13 Only one is shown), a memory 132 and computer executable instructions 133 stored in the memory 132 and executable on at least one processor 131, when the processor 131 executes the computer executable instructions 133, the steps of any of the above-mentioned embodiments of the image recognition method based on visual handles are implemented.

[0151] The electronic device may include but is not limited to a processor 131 and a memory 132. It will be understood by those skilled in the art that Figure 13 This is merely an example of the electronic device 130 and does not constitute a limitation on the electronic device 130 . The electronic device 130 may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, it may also include input and output devices, network access devices, etc.

[0152] The processor 131 may be a central processing unit (CPU), or other general-purpose processors, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPG), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.

[0153] The memory 132 can be an internal storage unit of the electronic device 130, such as a hard disk or a memory of the electronic device 130 in some embodiments. The memory 132 can also be an external storage device of the electronic device 130, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, or the like equipped on the electronic device 130 in other embodiments. Further, the memory 132 can include both an internal storage unit and an external storage device of the electronic device 130. The memory 132 is used to store an operating system, an application program, a boot loader, data, and other programs, such as program codes of a computer program, and the like. The memory 132 can also be used to temporarily store data that has been output or will be output.

[0154] The computer program product includes a computer program or computer executable instructions stored in a computer readable storage medium. The processor of the electronic device reads the computer executable instructions from the computer readable storage medium, and the processor executes the computer executable instructions to cause the electronic device to perform the image recognition method based on a visual handle according to the embodiments of the present application.

[0155] The computer readable storage medium stores computer executable instructions or a computer program. When the computer executable instructions or the computer program are executed by the processor, the processor performs the image recognition method based on a visual handle according to the embodiments of the present application, for example, the image recognition method based on a visual handle as shown in the above. Figure 2

[0156] In some embodiments, the computer readable storage medium can be a RAM, a ROM, a flash memory, a magnetic surface memory, an optical disc, or a CD-ROM, and the like. The computer readable storage medium can also be various devices including one or any combination of the above storage devices.

[0157] In some embodiments, the computer executable instructions can be in the form of a program, software, software module, script, or code, written in any form of programming language (including a compiled or interpreted language, or a declarative or procedural language), and can be deployed in any form, including being deployed as a standalone program or being deployed as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0158] ​By way of example, computer readable media can include computer- storage media and communication media. Computer-storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. Computer-storage media includes, but is not limited to, RAM, ROM, EEPROM, solid state drives (SSDs), flash memory, phase-change (PC) memory, optical disks (e.g., CDs, DVDs), magnetic disks (e.g., diskettes, hard disk drives), magnetic tapes, magnetic strips, and other storage media. Computer readable media also includes any wireless media such as term-endal, cellular links, microwave links or other wireless communication channels.

[0159] By way of example, computer readable media can include computer- storage media and communication media. Computer-storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. Computer-storage media includes, but is not limited to, RAM, ROM, EEPROM, solid state drives (SSDs), flash memory, phase-change (PC) memory, optical disks (e.g., CDs, DVDs), magnetic disks (e.g., diskettes, hard disk drives), magnetic tapes, magnetic strips, and other storage media. Computer readable media also includes any wireless media such as term-endal, cellular links, microwave links or other wireless communication channels.

[0160] The above description is implemented only as an example of the present application, and is not intended to limit the protection scope of the present application. Any modification, equivalent replacement and improvement within the spirit and scope of the present application are included in the protection scope of the present application.

Claims

1. An image recognition system based on visual handles, characterized in that: The system includes: a visual sensor and a data processing module; The visual sensor is used to collect visual handles under the target task; the visual handle is an image with a resolution lower than a preset resolution threshold and an image frequency lower than a preset frequency threshold; the visual sensor includes at least one imaging camera; The data processing module is used to perform image recognition based on the visual handle to obtain an initial recognition result; and in response to the initial recognition result meeting a preset condition, determine the initial recognition result as the image recognition result of the target task.

2. The system according to claim 1, wherein: The visual sensor includes a non-focus imaging camera; The non-focus imaging camera is used to collect images with a resolution lower than a preset resolution threshold and an image frequency lower than a preset frequency threshold; the non-focus imaging camera includes a black and white camera or an RGB color camera, and the horizontal viewing angle of the non-focus imaging camera is 0-200 degrees, the vertical viewing angle is 0-160 degrees, and the focal length is 0-200 mm.

3. The system according to claim 1, wherein: The visual sensor includes a non-focus imaging camera and a focus imaging camera; The non-focus imaging camera is used to collect images with a resolution lower than a preset resolution threshold and an image frequency lower than a preset frequency threshold to obtain the visual handle; The focus imaging camera is used to collect first image data of the target task; the image resolution corresponding to the first image data is greater than the preset resolution, and the image frequency is greater than the preset frequency threshold; Correspondingly, the data processing module is also used to obtain the first image data of the target task captured by the focus imaging camera in response to the initial recognition result not meeting the preset conditions; and perform image recognition based on the visual handle and the first image data to obtain the image recognition result of the target task.

4. The system according to claim 1, wherein: The system further includes at least one target sensor; The target sensor is used to collect target data under the target task; the target data includes at least one of time data and space data; Correspondingly, the data processing module is also used to obtain at least one target data of the target task collected by the at least one target sensor in response to the initial recognition result not meeting the preset condition; and perform image recognition based on the visual handle and the at least one target data to obtain the image recognition result of the target task.

5. The system according to claim 1, wherein: The visual sensor includes a non-focus imaging camera and a focus imaging camera; the system also includes at least one target sensor; The non-focus imaging camera is used to collect images with a resolution lower than a preset resolution threshold and an image frequency lower than a preset frequency threshold to obtain the visual handle; The focus imaging camera is used to collect first image data of the target task; The image resolution corresponding to the first image data is greater than the preset resolution, and the image frequency is greater than the preset frequency threshold; The target sensor is used to collect target data under the target task; the target data includes at least one of time data and space data; Correspondingly, the data processing module is further configured to, in response to the initial recognition result not satisfying the preset condition, acquire the first image data acquired by the focus imaging camera and the at least one target data of the target task acquired by the at least one target sensor; Image recognition is performed based on the visual handle, the first image data, and the at least one target data to obtain an image recognition result of the target task.

6. The system according to claim 1, wherein: The visual sensor includes a focus imaging camera; the system also includes a downsampling module and a convolution blur module; The focus imaging camera is used to collect first image data of the target task; the image resolution corresponding to the first image data is greater than the preset resolution, and the image frequency is greater than the preset frequency threshold; The convolution blur module is used to perform convolution blur on the first image data at least once to obtain a convolution blur result; The downsampling module is used to downsample the convolution blur result at least once to obtain the visual handle.

7. A method for image recognition based on visual handles, characterized in that: The method comprises: Obtaining a visual handle for a target task; the visual handle is an image having a resolution lower than a preset resolution threshold and an image frequency lower than a preset frequency threshold; Performing image recognition based on the visual handle to obtain an initial recognition result; In response to the initial recognition result satisfying a preset condition, the initial recognition result is determined as the image recognition result of the target task.

8. The method according to claim 7, characterized in that The performing image recognition based on the visual handle to obtain an initial recognition result includes: Performing image recognition on the visual handle using a pre-trained model to obtain an initial recognition result; The pre-trained model is obtained by training a visual handle set; the visual handle set includes at least one sample visual handle.

9. The method according to claim 7, characterized in that The step of obtaining the visual handle under the target task includes: Determining, from the electronic device, at least one target non-focus imaging camera in operation and a distance value between a lens module and an image sensor in the at least one target non-focus imaging camera based on the task type of the target task, the preset frequency threshold, and the preset resolution threshold; Based on the distance value, acquiring an image of the target task by using the at least one target non-focus imaging camera to obtain the visual handle; or, Determining, from the electronic device, at least one target focus imaging camera in operation and an aperture value of the at least one target focus imaging camera based on the task type of the target task, the preset frequency threshold, and the preset resolution threshold; Based on the aperture value, acquiring an image of the target task by using the at least one target focus imaging camera to obtain the visual handle; or, Determining from the electronic device a target focus imaging camera in an operational state; Capturing the target task image by the target focus imaging camera to obtain a target image; Based on the task type of the target task, the preset frequency threshold, and the preset resolution threshold, convolution blurring and downsampling are performed on the target image at least once to obtain the visual handle.

10. The method according to claim 7, characterized in that The method further comprises: In response to the initial recognition result not satisfying the preset condition, acquiring first image data of the target task; Image recognition is performed based on the visual handle and the first image data to obtain an image recognition result of the target task.

11. The method according to claim 10, characterized in that The performing image recognition based on the visual handle and the first image data to obtain the image recognition result of the target task includes: Determining a location prediction model corresponding to the task type of the target task; Performing position prediction on the visual handle using the position prediction model to obtain at least one position information of the visual handle; performing image registration on the first image data according to the visual handle to obtain a pixel position mapping relationship between the visual handle and the first image data; Performing image cropping on the first image data according to the at least one position information and the pixel position mapping relationship to obtain the at least one first partial image; Image recognition is performed based on the visual handle and the at least one first partial image to obtain an image recognition result of the target task.

12. The method according to claim 11, characterized in that The performing image recognition based on the visual handle and the at least one first partial image to obtain the image recognition result of the target task includes: Performing feature extraction on the visual handle and the at least one first partial image respectively to obtain a first feature vector and at least one second feature vector correspondingly; the second feature vector includes a texture feature vector; Performing vector fusion on the first eigenvector and the at least one second eigenvector to obtain a fused eigenvector; Image recognition is performed based on the fused feature vector to obtain an image recognition result of the target task.

13. The method according to claim 7, characterized in that The method further comprises: In response to the initial recognition result not satisfying the preset condition, determining at least one target sensor in an operating state from the electronic device based on a task type of the target task; Collecting data for the target task through the at least one target sensor to obtain the at least one target data; Image recognition is performed based on the visual handle and the at least one target data to obtain an image recognition result of the target task.

14. The method according to claim 7, wherein: The method further comprises: In response to the initial recognition result not satisfying the preset condition, acquiring first image data of the target task, and determining at least one target sensor in an operating state from the electronic device based on a task type of the target task; Collecting data for the target task through the at least one target sensor to obtain the at least one target data; Image recognition is performed based on the visual handle, the first image data, and the at least one target data to obtain an image recognition result of the target task.