Object recognition method, device, server and medium
By performing image quality evaluation and object attribute recognition on the server side, the problems of low accuracy of object recognition and high hardware cost in the prior art are solved, and efficient and accurate object recognition and the effect of reducing hardware cost is achieved.
Patent Information
- Application Number
- CN202011081284.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-10-10
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2040-10-10
AI Technical Summary
The prior art is difficult to improve identification accuracy in object recognition, and it has a high dependence on terminal hardware configuration, resulting in high hardware costs.
By receiving the video stream data uploaded by the terminal, the image quality is evaluated, the color image frame with better image quality is selected, and the target object is identified at attributes. This method is executed on the server side to reduce dependence on terminal hardware.
It effectively improves the accuracy of the attribute information of object recognition, reduces the cost of terminal hardware, and improves recognition efficiency and user experience.
Smart Images

Figure CN113518217B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of Internet technologies, specifically to the field of image processing technologies, and more particularly to an object recognition method, an object recognition device, a server, and a computer storage medium. Background Art
[0002] With the popularization of the concept of Artificial Intelligence (AI), object recognition has become a hot research topic. The so-called object recognition refers to a technology that uses a computer to process, analyze, and understand an image to identify the attribute information of the target object in the image. The target object here can be: a human face, a gesture, and other organisms other than humans (such as cats and dogs); currently, how to better perform object recognition has become a research hotspot. Summary of the Invention
[0003] Embodiments of the present invention provide an object recognition method, device, server, and medium, which can better perform object recognition and improve recognition accuracy.
[0004] On the one hand, embodiments of the present invention provide an object recognition method, the method comprising:
[0005] Receiving video stream data about a target object uploaded by a terminal, the video stream data at least comprising: a color map stream composed of multiple color image frames about the target object;
[0006] Performing image quality evaluation on each color image frame in the color map stream to obtain an image quality score for each color image frame;
[0007] Based on the image quality scores of each color image frame, selecting a target color image frame from the color map stream; the target color image frame refers to the color image frame with the highest image quality score, or the color image frame with an image quality score greater than a quality score threshold;
[0008] Performing attribute recognition on the target object according to the target color image frame to obtain the attribute information of the target object.
[0009] On the other hand, embodiments of the present invention provide an object recognition device, the device comprising:
[0010] A communication unit, receiving video stream data about a target object uploaded by a terminal, the video stream data at least comprising: a color map stream composed of multiple color image frames about the target object;
[0011] A processing unit, configured to perform image quality evaluation on each color image frame in the color map stream to obtain an image quality score for each color image frame;
[0012] The processing unit is further configured to select a target color image frame from the color map stream based on the image quality scores of the color image frames; the target color image frame refers to the color image frame with the highest image quality score, or the color image frame with an image quality score greater than the quality score threshold.
[0013] The processing unit is further configured to perform attribute recognition on the target object according to the target color image frame to obtain attribute information of the target object.
[0014] In one implementation, when the processing unit is configured to perform image quality assessment on each color image frame in the color map stream to obtain the image quality scores of the color image frames, it is specifically configured to:
[0015] Obtain a quality assessment rule, and perform image quality assessment on each color image frame in the color map stream according to the quality assessment rule to obtain the image quality scores of the color image frames.
[0016] Wherein, the quality assessment rule includes multiple evaluation dimensions, an initial weight value corresponding to each evaluation dimension, and a scoring algorithm under each evaluation dimension; the multiple evaluation dimensions include: image sharpness dimension, illumination quality dimension, resolution dimension, and object integrity dimension.
[0017] In another implementation, when the processing unit is configured to perform image quality assessment on each color image frame in the color map stream according to the quality assessment rule to obtain the image quality scores of the color image frames, it may be specifically configured to:
[0018] Perform quality detection on any color image frame in the color data stream under each evaluation dimension to obtain the quality detection results of the any color image frame under each evaluation dimension.
[0019] Determine the quality scores of the any color image frame under each evaluation dimension according to the quality detection results of the any color image frame under each evaluation dimension and the scoring algorithm under each evaluation dimension.
[0020] Determine the target weight values corresponding to the evaluation dimensions according to the quality detection results of the any color image frame under each evaluation dimension and the initial weight values corresponding to the evaluation dimensions.
[0021] Perform weighted summation on the quality scores of the any color image frame under each evaluation dimension by using the target weight values corresponding to the evaluation dimensions to obtain the image quality score of the any color image frame.
[0022] In another implementation, if there is an occluder in any of the color image frames, resulting in an incomplete target object in any of the color image frames, the quality detection result of any of the color image frames in the object integrity dimension includes: the occlusion attribute of the occluder; the occlusion attribute includes at least one of the following: occlusion area and occlusion position;
[0023] The method for determining the target weight value corresponding to the object integrity dimension is as follows: determining a weight adjustment factor for the object integrity dimension according to the occlusion attribute; and adjusting the initial weight value corresponding to the object integrity dimension by using the weight adjustment factor to obtain the target weight value;
[0024] Wherein, if the occlusion attribute includes the occlusion area, the weight adjustment factor is: a first adjustment factor calculated according to the occlusion area and the image size of any of the color image frames; if the occlusion attribute includes the occlusion position, the weight adjustment factor is: a second adjustment factor corresponding to the occlusion position; if the occlusion attribute includes the occlusion area and the occlusion position, the weight adjustment factor is: an adjustment factor obtained by integrating the weights of the first adjustment factor and the second adjustment factor.
[0025] In another implementation, when the processing unit is used to identify the attributes of the target object according to the target color image frame to obtain the attribute information of the target object, it may specifically be used for:
[0026] Identifying the attributes of the target object according to the target color image frame;
[0027] If the identification is successful, the identified attribute information is used as the attribute information of the target object;
[0028] If the identification fails, a target color image frame is reselected from the remaining unselected color image frames in the color map stream, and the attributes of the target object are identified according to the reselected target color image frame to obtain the attribute information of the target object.
[0029] In another implementation, when the processing unit is used to, if the identification fails, reselect a target color image frame from the remaining unselected color image frames in the color map stream, it may specifically be used for:
[0030] If the identification fails, detect image anomaly problems in the target color image frame;
[0031] If it is detected that the target color image frame has a target image anomaly problem, adjust the quality assessment rule according to the target image anomaly problem;
[0032] Perform image quality assessment on each remaining color image frame in the color map stream that has not been selected according to the adjusted quality assessment rules to update the image quality scores of each remaining color image frame;
[0033] Based on the updated image quality scores of each remaining color image frame, reselect new target color image frames from each remaining color image frame.
[0034] In another implementation manner, when the processing unit is used to adjust the quality assessment rules according to the target image anomaly problem, it may specifically be used for:
[0035] Obtain a correspondence table between image anomaly problems and evaluation dimensions, where the correspondence table includes multiple image anomaly problems and the evaluation dimensions corresponding to each image anomaly problem;
[0036] Query the target evaluation dimension corresponding to the target image anomaly problem from the correspondence table;
[0037] If the quality assessment rules include the target evaluation dimension, increase the initial weight value corresponding to the target evaluation dimension in the quality assessment rules;
[0038] If the quality assessment rules do not include the target evaluation dimension, add the target evaluation dimension, the initial weight value corresponding to the target evaluation dimension, and the scoring algorithm under the target evaluation dimension to the quality assessment rules.
[0039] In another implementation manner, the video stream data further includes: an infrared map stream composed of multiple infrared image frames of the target object, and a depth map stream composed of multiple depth image frames of the target object; correspondingly, when the processing unit is used to perform attribute recognition on the target object according to the target color image frame, it may specifically be used for:
[0040] Select a target infrared image frame corresponding to the target color image frame from the infrared map stream, and select a target depth image frame corresponding to the target color image frame from the depth map stream;
[0041] Perform attribute recognition on the target object according to the target color image frame, the target infrared image frame, and the target depth image frame.
[0042] In another implementation manner, when the processing unit is used to perform attribute recognition on the target object according to the target color image frame, the target infrared image frame, and the target depth image frame, it may specifically be used for:
[0043] Perform attribute recognition on the target object according to the target color image frame and the target infrared image frame to obtain first attribute information;
[0044] Perform attribute recognition on the target object according to the target color image frame and the target depth image frame to obtain second attribute information;
[0045] If the first attribute information and the second attribute information match, determine that the recognition is successful; otherwise, determine that the recognition fails.
[0046] In another implementation manner, the processing unit can also be used to:
[0047] Return the attribute information of the target object to the terminal, so that the terminal performs service processing according to the attribute information of the target object;
[0048] Wherein, the target object is a target face, and the attribute information of the target object includes the identity information corresponding to the target face; the service processing includes: face attendance processing, face payment processing, or face image classification processing.
[0049] On the other hand, an embodiment of the present invention provides a server, the server includes an input interface and an output interface, and the server further includes:
[0050] A processor, adapted to implement one or more instructions; and,
[0051] A computer storage medium, the computer storage medium stores one or more instructions, and the one or more instructions are adapted to be loaded and executed by the processor to perform the following steps:
[0052] Receive video stream data of a target object uploaded by a terminal, the video stream data at least includes: a color map stream composed of multiple color image frames of the target object;
[0053] Perform image quality evaluation on each color image frame in the color map stream to obtain the image quality score of each color image frame;
[0054] Based on the image quality scores of each color image frame, select a target color image frame from the color map stream; the target color image frame refers to the color image frame with the highest image quality score, or the color image frame whose image quality score is greater than the quality score threshold;
[0055] Perform attribute recognition on the target object according to the target color image frame to obtain the attribute information of the target object.
[0056] In another aspect, an embodiment of the present invention provides a computer storage medium storing one or more instructions, which are adapted to be loaded and executed by a processor to perform the following steps:
[0057] Receive video stream data of a target object uploaded by a terminal, where the video stream data at least includes: a color map stream composed of multiple color image frames of the target object;
[0058] Perform image quality assessment on each color image frame in the color map stream to obtain the image quality score of each color image frame;
[0059] Based on the image quality scores of the color image frames, select target color image frames from the color map stream; the target color image frames refer to the color image frames with the highest image quality score or the color image frames with an image quality score greater than a quality score threshold;
[0060] Perform attribute recognition on the target object according to the target color image frame to obtain the attribute information of the target object.
[0061] The server in the embodiment of the present invention can receive the video stream data of the target object uploaded by the terminal, and perform image quality assessment on each color image frame in the color map stream included in the video stream data to obtain the image quality score of each color image frame. Then, based on the image quality scores of the color image frames, target color image frames with better image quality can be selected to perform attribute recognition on the target object, which can effectively improve the accuracy of the recognized attribute information. And since all the operations involved in the entire object recognition process (such as image selection operation, attribute recognition operation) are executed in the server, the dependence on the hardware configuration of the terminal can be effectively eliminated, thereby effectively reducing the hardware cost of the terminal. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0063] Figure 1a is a system architecture diagram of an object recognition system provided by an embodiment of the present invention;
[0064] Figure 1b is a schematic principle diagram of an object recognition solution provided by an embodiment of the present invention;
[0065] Figure 2 is a schematic flowchart of an object recognition method provided by an embodiment of the present invention;
[0066] Figure 3 It is a schematic flowchart of an object recognition method provided by another embodiment of the present invention;
[0067] Figure 4a It is a schematic diagram of writing image identifiers for each image frame provided by an embodiment of the present invention;
[0068] Figure 4b It is a schematic diagram of selecting a target color image frame, a target depth image frame, and a target infrared image frame provided by an embodiment of the present invention;
[0069] Figure 5a It is an application scenario diagram of an object recognition method provided by an embodiment of the present invention;
[0070] Figure 5b It is an application scenario diagram of an object recognition method provided by an embodiment of the present invention;
[0071] Figure 5c It is an application scenario diagram of an object recognition method provided by an embodiment of the present invention;
[0072] Figure 6 It is a schematic structural diagram of an object recognition device provided by an embodiment of the present invention;
[0073] Figure 7 It is a schematic structural diagram of a server provided by an embodiment of the present invention. Detailed implementation manners
[0074] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention.
[0075] In the embodiments of the present invention, an object recognition system is involved. Refer to Figure 1aAs shown, the object recognition system may at least include: a terminal 11 and a server 12; and the terminal 11 and the server 12 may be directly or indirectly connected through wired or wireless communication means, without limitation thereto. Among them, the terminal 11 refers to any device that can call a traditional camera component or a 3D camera component to collect images of a target object; the target object here refers to any object to be subjected to attribute recognition, such as a target face, a target gesture, and other organisms other than humans (such as cats, dogs), etc. The traditional camera component refers to a camera component used to collect color image frames (i.e., RGB image frames), while the 3D camera component refers to a camera component that adds relevant software and hardware for live detection on the basis of the traditional camera component. For example, on the basis of including the traditional camera component, the 3D camera component may further include: a depth camera for collecting depth image frames, and an infrared camera for collecting infrared image frames, etc. Among them, an RGB image frame refers to an image in which the RGB values of each point in the captured scene are used as pixel values; that is, an RGB image frame refers to an image with RGB values, such as a color face image, a color gesture image, etc. A depth image frame, also known as a distance image frame, refers to an image in which the distance of each point in the captured scene relative to the depth camera is used as a pixel value; that is, a depth image frame refers to an image with depth information, such as a face image with depth information, a gesture image with depth information, etc. An infrared image frame refers to an image formed by an infrared camera collecting the radiation of each point in the captured scene in the infrared band; that is, an infrared image frame refers to an image with infrared information, such as a face image with infrared information, a gesture image with infrared information, etc. Specifically, the terminal 11 may include, but is not limited to: a smart phone, a tablet computer, a laptop computer, a desktop computer, a face-swiping payment device (a device that automatically deducts electronic resources from a resource account associated with the identity information after obtaining the identity information through face recognition), an automatic cashier device, and so on.
[0076] Server 12 refers to a service device that can provide multiple services such as streaming media services, image optimization services, object recognition services, etc.; the streaming media service here refers to: a service based on the image frames collected and uploaded by the streaming media receiving terminal 11. The so-called streaming media refers to a media form that transmits audio, video, and multimedia files in a streaming manner over the network. The image optimization service refers to: a service that selects image frames with better image quality from the image frames received through the streaming media service, such as a service that selects the optimal frame; the so-called optimal frame refers to the image frame with the highest image quality among the received image frames. The object recognition service refers to: a service that identifies the attribute information of the target object in the image frame. The attribute information of the target object mentioned here refers to the information that can be used to characterize some features of the target object. For example, if the target object is a target face, the attribute information can be the identity information corresponding to the target face; another example is that if the target object is a target gesture, the attribute information can be the gesture information corresponding to the target gesture (such as the shape of the gesture, the position information of each node in the hand, etc.). Specifically, the server 12 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms, etc.
[0077] Based on the above object recognition system, an object recognition solution is proposed in an embodiment of the present invention to better perform object recognition. Refer to Figure 1b As shown, the general principle of this object recognition solution is as follows: When there is an object recognition requirement, the terminal can call the 3D camera component to collect images of the target object to be recognized; and use the streaming media method to upload each collected image frame to the server in the form of a video stream for object recognition. The so-called streaming media method refers to a transmission method in which the already collected image frames are synchronously transmitted during the process of collecting new image frames. Correspondingly, the server can receive the video stream data uploaded by the terminal through the streaming media service. Then, the image optimization service can use the image optimization logic to optimize the video stream data to select image frames with better quality from the video stream data; then, the object recognition service can perform attribute recognition on the target object based on the selected image frames with better quality. After successful recognition, the attribute information about the target object can also be returned to the terminal so that the terminal can perform relevant business operations based on this attribute information. It should be noted that Figure 1b This is only an example to illustrate the general principle of the object recognition solution and does not limit it. For example, Figure 1bThe principle of the shown solution is to upload each captured image frame in the form of a video stream by using the streaming media method; in other embodiments, after capturing each image frame, the terminal can also make each image frame into video stream data and then upload each frame image in the video stream data to the server uniformly; in this case, the server can directly receive the video stream data, and so on.
[0078] It can be seen from this that the object recognition solution proposed in the embodiments of the present invention has the following beneficial effects: ① By selecting image frames with better image quality for attribute recognition, the accuracy of the recognized attribute information can be effectively improved. ② Since operations such as image optimization and attribute recognition are all executed by the server, on the one hand, the dependence on the hardware configuration of the terminal can be eliminated, and the hardware cost of the terminal can be reduced; on the other hand, when the preferred image frames cannot be recognized, it is convenient for the server to directly select new image frames from the video stream data for recognition without waiting for the terminal to re-capture and upload new image frames, which can effectively improve the recognition efficiency and thus improve the user experience.
[0079] Based on the above description, an object recognition method is proposed in the embodiments of the present invention, and this object recognition method can be executed by the server in the above-mentioned object recognition system. Please refer to Figure 2 , and this object recognition method may include the following steps S201-S204:
[0080] S201, receive the video stream of the target object uploaded by the terminal.
[0081] In a specific implementation, the terminal can detect an object recognition trigger event regarding the target object in real time; this object recognition trigger event may include but is not limited to: an event of detecting a language password indicating object recognition, an event of detecting a trigger operation (such as a click operation, a press operation) on the recognition trigger component in the user interface displayed by the terminal, and so on. If the terminal detects this object recognition trigger event, it can call the 3D camera component to collect images of the target object to obtain video stream data of the target object and upload the video stream data to the server. Correspondingly, the server can receive the video stream data of the target object uploaded by the terminal. Among them, this video stream data at least includes: a color stream composed of multiple color image frames of the target object; and this color stream can be uploaded by using the streaming media method or can be uploaded uniformly after each color image frame is captured, and there is no limitation on this.
[0082] Optionally, the video stream data may further include: an infrared image stream composed of multiple infrared image frames of the target object, a depth image stream composed of multiple depth image frames of the target object, and so on; so that the subsequent server can refer to the infrared image frames in the infrared image stream and the depth image frames in the depth image stream to perform attribute recognition on the target object, thereby improving the accuracy of the attribute information. And similar to the color image stream, the aforementioned infrared image stream can be uploaded in a streaming media manner or can be uploaded uniformly after each infrared image frame is collected, without any restrictions on this. The aforementioned depth image stream can be uploaded in a streaming media manner or can be uploaded uniformly after each depth image frame is collected, without any restrictions on this. Moreover, the upload methods used for these three types of data, namely the color image stream, the infrared image stream, and the depth image stream, can be the same or different; for example, the color image stream can be uploaded in a streaming media manner, while the infrared image stream and the depth image stream can be uploaded in a unified upload manner. Uploading the color image stream in a streaming media manner enables the server to perform image quality assessment on the received color image frames in parallel during the process of receiving new color image frames, so as to improve the quality assessment efficiency; while using the unified upload manner for the infrared image stream and the depth image stream can avoid the infrared image stream and the depth image stream competing with the color image stream for network transmission resources, and can effectively improve the transmission efficiency of the color image stream.
[0083] S202. Perform image quality assessment on each color image frame in the color image stream to obtain the image quality scores of each color image frame.
[0084] In a specific implementation, if the color image stream is uploaded uniformly after each color image frame is collected by the terminal, the server can perform image quality assessment on each color image frame in the color image stream after receiving each color image frame in the color image stream to obtain the image quality scores of each color image frame. In another specific implementation, if the color image stream is uploaded in a streaming media manner, the server can execute two processes in parallel to implement the reception processing and image quality assessment processing of the color image stream; specifically, the first process can be called to receive each color image frame in the color image stream, and the second process can be called to perform image quality assessment on each color image frame in the color image stream to obtain the image quality scores of each color image frame. That is to say, in this specific implementation, each color image frame in the color image stream is received through the first process, and each color image frame in the color image stream is subjected to image quality assessment through the second process; where the first process and the second process are executed in parallel. Such a processing method can enable the server to receive new color image frames in the color image stream uploaded by the terminal while concurrently performing image quality assessment on the received color image frames, thereby effectively improving the efficiency of image quality assessment.
[0085] Specifically, in the specific implementation process of the server for evaluating the image quality of each color image frame in the color map stream to obtain the image quality score of each color image frame, the quality evaluation rule can be obtained first. Among them, the quality evaluation rule can include multiple evaluation dimensions, the initial weight value corresponding to each evaluation dimension, and the scoring algorithm under each evaluation dimension; the multiple evaluation dimensions here can include: image sharpness dimension, illumination quality dimension, resolution dimension, and object integrity dimension; the initial weight value corresponding to each evaluation dimension and the scoring algorithm under each evaluation dimension can both be set according to empirical values or business requirements. It should be noted that the embodiments of the present invention only exemplarily list multiple evaluation dimensions, not an exhaustive list; for example, when the target object is a human face or the face of other organisms other than humans, the quality evaluation rule can also include the facial symmetry dimension, the initial weight value corresponding to the facial symmetry dimension, and the scoring algorithm under the facial symmetry dimension, and so on. Then, the image quality of each color image frame in the color map stream can be evaluated according to the quality evaluation rule to obtain the image quality score of each color image frame. Since the implementation method of evaluating the image quality of each color image frame to obtain the corresponding image quality score is the same, the principle of image quality evaluation is described by taking any color image frame in the color map stream as an example; specifically, the quality score of any color image frame under each evaluation dimension can be determined respectively according to the scoring algorithm under each evaluation dimension; then, the quality score of any color image frame under each evaluation dimension and the initial weight value corresponding to each evaluation dimension are weighted and summed to obtain the image quality score of any color image frame.
[0086] S203. Based on the image quality scores of each color image frame, select a target color image frame from the color map stream.
[0087] Since the image quality scores of each color image frame can be used to reflect the image quality of each color image frame, the higher the image quality score of a color image frame, the better its image quality; therefore, when the server executes step S203, it can select a color image frame with better image quality from the color map stream as the target color image frame based on the image quality scores of each color image frame. Based on this, step S203 can specifically include the following implementation methods:
[0088] In one implementation method, the color image frame with the highest image quality score can be selected from the color map stream as the target image frame according to the image quality scores of each color image frame.
[0089] In another implementation, a preset quality score threshold can be obtained. Secondly, at least one candidate color image frame with an image quality score greater than the quality score threshold can be filtered out from the color map stream according to the image quality scores of the respective color image frames in the color map stream. Then, a target color image frame can be selected from the at least one candidate color image frame. Specifically, the server can directly select any color image frame from the at least one candidate color image frame as the target color image frame. Alternatively, each evaluation dimension can also have a corresponding priority; then the server can first determine the benchmark evaluation dimension with the highest priority from multiple evaluation dimensions according to the priorities of the respective evaluation dimensions; then, according to the quality scores of the respective candidate color image frames under the benchmark evaluation dimension, the candidate color map frame with the highest quality score under the benchmark evaluation dimension can be selected as the target color image frame. For example, assume that the color map stream includes a total of 7 color image frames (i.e., color image frames 1-7), and their image quality scores are: 75 points, 80 points, 95 points, 98 points, 99 points, 85 points; and the quality score threshold is 90 points, then 3 candidate color image frames can be obtained: color image frame 3, color image frame 4, and color image frame 5. And the quality scores of these 3 candidate color image frames under the benchmark evaluation dimension (such as the object integrity dimension) are: 60 points, 58 points, 59 points; then, it can be seen that the candidate color map frame with the highest quality score under the benchmark evaluation dimension is color image frame 3. Therefore, color image frame 3 can be selected as the target image frame.
[0090] S204. Perform attribute recognition on the target object according to the target color image frame to obtain the attribute information of the target object.
[0091] After the target color image frame is selected through step S203, the attribute of the target object can be recognized according to the target color image frame. If the recognition is successful, the recognized attribute information can be used as the attribute information of the target object; if the recognition fails, a target color image frame can be reselected from the remaining unselected color image frames in the color map stream, and the attribute of the target object can be recognized according to the reselected target color image frame to obtain the attribute information of the target object.
[0092] Among them, a specific implementation manner of performing attribute recognition on a target object according to a target color image frame may be: directly performing attribute recognition on the target object according to the target color image frame. Further, if the video stream also includes: an infrared image stream composed of multiple infrared image frames of the target object, and a depth image stream composed of multiple depth image frames of the target object; then another specific implementation manner of performing attribute recognition on the target object according to the target color image frame may be: selecting a target infrared image frame corresponding to the target color image frame from the infrared image stream, and selecting a target depth image frame corresponding to the target color image frame from the depth image stream; performing attribute recognition on the target object according to the target color image frame, the target infrared image frame, and the target depth image frame.
[0093] The server in the embodiment of the present invention can receive the video stream data of the target object uploaded by the terminal, and perform image quality evaluation on each color image frame in the color image stream included in the video stream data to obtain the image quality score of each color image frame. Then, based on the image quality scores of each color image frame, a target color image frame with better image quality can be selected to perform attribute recognition on the target object, which can effectively improve the accuracy of the recognized attribute information. And since all the operations involved in the entire object recognition process (such as image selection operation, attribute recognition operation) are executed in the server, the dependence on the hardware configuration of the terminal can be effectively eliminated, thereby effectively reducing the hardware cost of the terminal.
[0094] Based on the above Figure 2 related description of the object recognition method embodiment shown, the embodiment of the present invention also proposes another more specific object recognition method, which can be executed by the server in the object recognition system mentioned above. And the embodiment of the present invention mainly takes the video stream data including a color image stream, an infrared image stream, and a depth image stream, and combines these three image streams to perform attribute recognition on the target object as an example for description. Please refer to Figure 3 , the object recognition method may include the following steps S301 - S307:
[0095] S301, receiving the video stream data of the target object uploaded by the terminal.
[0096] In a specific implementation, the terminal can call the 3D camera component to collect images of the target object, so as to obtain a color map stream, an infrared map stream, and a depth map stream of the target object. For the color map stream, infrared map stream, and depth map stream generated by the 3D camera component, the terminal can write a unique image ID (i.e., image identifier) into the frame data of each image frame generated at the same moment in the three streams of images (i.e., the color image frame, infrared image frame, and depth image frame generated at the same moment), so that after the server selects the target color image frame, it can obtain the corresponding target infrared image frame and target depth image frame through this image ID. Specifically, the terminal can first call the ID generator to generate a unique image ID for the image frames generated by the three streams of images at the same moment, and then write the image ID into the image data of the corresponding image frame in the form of a watermark or other form, such as Figure 4a shown. Among them, the aforementioned image ID can be a UUID (Universally Unique Identifier), or a timestamp used to represent the generation time of the image frame, and so on. After the terminal writes the corresponding image ID into the image data of each image frame in the three streams of images, it can upload the three streams of images as the video stream data of the target object to the server. Correspondingly, the server can receive the video stream of the target object uploaded by the terminal.
[0097] S302, perform image quality evaluation on each color image frame in the color map stream to obtain the image quality score of each color image frame.
[0098] In the specific implementation process, the server can perform image quality evaluation on each color image frame in the color map stream according to the quality evaluation rules to obtain the image quality score of each color image frame. As can be seen from the foregoing, the quality evaluation rules can include multiple evaluation dimensions, the initial weight value corresponding to each evaluation dimension, and the scoring algorithm under each evaluation dimension; the multiple evaluation dimensions mentioned here include: image clarity dimension, lighting quality dimension, resolution dimension, and object integrity dimension. Then correspondingly, for any color image frame in the color map stream, the specific implementation method for the server to perform image quality evaluation on the any color image frame to obtain the image quality score of the any color image frame can be as follows:
[0099] The server can first perform quality detection on any color image frame in the color data stream under each evaluation dimension to obtain the quality detection results of any color image frame under each evaluation dimension. Among them, the quality detection result of any color image frame in the image sharpness dimension may include the gray values of each pixel point in any color image frame; the quality detection result of any color image frame in the illumination quality dimension may include the available range of the gray intensity of each pixel in any color image; the quality detection result of any color image frame in the resolution dimension may include the resolution of each pixel point in any color image frame; the quality detection result of any color image frame in the object integrity dimension may include the integrity of the target object in any color image frame. Further, if the target object is incomplete, it may be caused by the target object being blocked by an occluder (such as a mask, sunglasses, other obstacles, etc.), or it may be caused by the target object not being entirely within the shooting range of the 3D camera component. Based on this, if the target object in any color image frame is incomplete due to the existence of an occluder in any color image frame, the quality detection result of any color image frame in the object integrity dimension may further include: the occlusion attribute of the occluder; the occlusion attribute may include at least one of the following: occlusion area and occlusion position.
[0100] Secondly, the server can determine the quality scores of any color image frame under each evaluation dimension according to the quality detection results of any color image frame under each evaluation dimension and the scoring algorithms under each evaluation dimension. Among them, the scoring algorithms under each evaluation dimension are as follows:
[0101] The scoring algorithm in the image sharpness dimension may include, but is not limited to: Brenner gradient method (an algorithm that obtains the quality score by calculating the gray difference between two adjacent pixel points), Tenengrad gradient method (an algorithm that determines the quality score by using the Sobel operator to calculate the gradients in the horizontal and vertical directions respectively) energy gradient function, and so on. Further, if the target object is a human face, the scoring algorithm in the image sharpness dimension can also be: first find the facial feature points from the color image frame to be evaluated, and construct a mask (mask) according to the facial feature points, and there are no background pixels on the mask; then the average Laplacian operator response of the mask can be used as the quality score of the color image frame in the image sharpness dimension.
[0102] The algorithm principle of the scoring algorithm in the illumination quality dimension is roughly as follows: evaluate the quality score of the color image frame in the illumination quality dimension by determining the length of the available range of the gray intensity; specifically, 5% of the darkest pixel and the brightest pixel can be removed, and then the quality score of the color image frame in the illumination quality dimension can be determined according to the available range of the gray intensity of the remaining pixels.
[0103] The scoring algorithm in the resolution dimension may include, but is not limited to: a linear function based on the size of the bounding box in terms of resolution; specifically, using \(x\) to represent the resolution of any pixel point, and \(s\) to represent the quality score of the color image frame in the resolution dimension; correspondingly, the function formula of this linear function can be seen as shown in the following formula:
[0104]
[0105] The general principle of the scoring algorithm in the object integrity dimension is as follows: Obtain a scoring table between integrity and quality score, where this scoring table includes multiple corresponding relationships between integrity and quality score; Query the quality score corresponding to the integrity of the target object in the color image frame from the scoring table, and take the queried quality score as the quality score of the color image frame in the object integrity dimension. Or, obtain the default quality score corresponding to the object integrity dimension, calculate the product between the integrity of the target object in the color image frame and the default quality score, and take the calculated product as the quality score of the color image frame in the object integrity dimension, and so on. It should be noted that the quality scores of the color image frames mentioned in the embodiments of the present invention in the object integrity dimension are all values less than or equal to 0.
[0106] Then, the server can determine the target weight value corresponding to each evaluation dimension according to the quality detection results of any color image frame in each evaluation dimension and the initial weight value corresponding to each evaluation dimension. In one implementation manner, the initial weight value corresponding to each evaluation dimension can be directly used as the target weight value corresponding to each evaluation dimension. In another implementation manner, the dynamic adjustment factor corresponding to each evaluation dimension can be determined respectively according to the quality detection results of any color image frame in each evaluation dimension, and the initial weight value corresponding to each evaluation dimension can be dynamically adjusted by using the dynamic adjustment factor corresponding to each evaluation dimension to obtain the target weight value corresponding to each evaluation dimension. Taking the object integrity dimension as an example, the determination method of the target weight value corresponding to this object integrity dimension can be as follows: First, the weight adjustment factor regarding the object integrity dimension (i.e., the dynamic adjustment factor corresponding to the object integrity dimension) can be determined according to the occlusion attribute; and the initial weight value corresponding to the object integrity dimension is adjusted by using the weight adjustment factor to obtain the target weight value.
[0107] Among them, if the occlusion attribute includes the occlusion area, the weight adjustment factor can be: the first adjustment factor calculated according to the occlusion area and the image size of any color image frame; specifically, the first adjustment factor can be the ratio between the occlusion area and the image size, or the product of the ratio between the occlusion area and the image size and the default adjustment factor, and so on. If the occlusion attribute includes the occlusion position, the weight adjustment factor is: the second adjustment factor corresponding to the occlusion position; specifically, the second adjustment factor is proportional to the importance degree of the occlusion position, that is, the more important the occlusion position is, the larger the second adjustment factor is. If the occlusion attribute includes the occlusion area and the occlusion position, the weight adjustment factor is: the adjustment factor obtained by integrating the weights of the first adjustment factor and the second adjustment factor; specifically, the weight integration here can include but is not limited to: summation operations (such as weighted summation, direct summation, etc.), product operations (such as weighted product, direct product, etc.), and so on.
[0108] Finally, the server can perform a weighted sum on the quality scores of any color image frame in each evaluation dimension using the target weight values corresponding to each evaluation dimension to obtain the image quality score of any color image frame. By iterating the above steps, the server can implement the image quality evaluation of each color image frame in the color image stream and obtain the image quality scores of each color image frame.
[0109] S303, select a target color image frame from the color image stream based on the image quality scores of each color image frame.
[0110] Among them, the target color image frame refers to the color image frame with the highest image quality score, or the color image frame with an image quality score greater than the quality score threshold; further, when the target color image frame refers to the color image frame with an image quality score greater than the quality score threshold, it can specifically refer to any color image frame with an image quality score greater than the quality score threshold, or specifically refer to the color image frame with an image quality score greater than the quality score threshold and the highest quality score in the benchmark evaluation dimension, and so on. It should be noted that the specific implementation manner of step S303 can refer to the relevant description of step S203 in the above-mentioned invention embodiments and will not be elaborated here.
[0111] S304, perform attribute recognition on the target object according to the target color image frame.
[0112] In the specific implementation process, the target infrared image frame corresponding to the target color image frame can be selected from the infrared image stream first, and the target depth image frame corresponding to the target color image frame can be selected from the depth image stream. In the specific implementation: as described above, the image ID is written in the image data of each image frame. Therefore, the server can first determine the image ID of the target color image frame; then select the infrared image frame with this image ID from the infrared image stream as the target infrared image frame, and select the depth image frame with this image ID from the depth image stream as the target depth image frame. Taking the target color image frame as the color image frame with the highest image quality score as an example, assume the target color image frame is Figure 4b the color image frame shown by the black frame in Figure 4b as shown.
[0113] Then, the attribute recognition of the target object can be performed according to the target color image frame, the target infrared image frame and the target depth image frame. In the specific implementation, the attribute recognition of the target object can be performed according to the target color image frame and the target infrared image frame to obtain the first attribute information. Specifically, the living body detection of the target object can be performed according to the target infrared image frame first; if the target object passes the living body detection, the attribute recognition model can be called to perform the attribute recognition of the target object according to the target color image frame to obtain the first attribute information. In addition, the attribute recognition of the target object can be performed according to the target color image frame and the target depth image frame to obtain the second attribute information. Specifically, the image reconstruction can be performed according to the target color image frame and the target depth image frame to obtain the three-dimensional image frame of the target object; then the attribute recognition of the target object can be performed according to this three-dimensional image frame to obtain the second attribute information. After obtaining the first attribute information and the second attribute information, it can be judged whether the first attribute information and the second attribute information match. If the first attribute information and the second attribute information match, it can be determined that the recognition is successful; in this case, the server can jump to execute step S305. Otherwise, it can be determined that the recognition fails; in this case, the server can jump to execute steps S306 - S307. Among them, the first attribute information and the second attribute information match means that: the first attribute information and the second attribute information are the same, or the difference between the first attribute information and the second attribute information is less than the difference threshold.
[0114] S305, if the recognition is successful, the recognized attribute information is used as the attribute information of the target object; the recognized attribute information mentioned here is the first attribute information or the second attribute information.
[0115] S306, if the recognition fails, the target color image frame is reselected from the remaining unselected color image frames in the color image stream.
[0116] In one embodiment, if the recognition fails, the server can directly reselect the target color image frame from the remaining color image frames not selected in the color map stream according to the image quality scores of the remaining color image frames not selected in the color map stream; in this embodiment, the reselected target color image frame refers to: the remaining color image frame with the highest image quality score among the remaining color image frames not selected in the color map stream, or the remaining color image frame with an image quality score greater than the quality score threshold among the remaining color image frames not selected in the color map stream.
[0117] In another embodiment, as can be seen from the foregoing, the recognition failure is caused by the mismatch between the first attribute information and the second attribute information. The mismatch between the first attribute information and the second attribute information may be due to image abnormality problems in the target color image frame (such as low clarity, incomplete target object, backlight in the target color image frame, etc.), resulting in too large a difference between the first attribute information and the second attribute information; it may also be due to the fact that there are no image abnormality problems in the target color image frame, but misrecognition occurs during the recognition process, resulting in too large a difference between the first attribute information and the second attribute information. Based on this, to improve the subsequent recognition success rate; if the recognition fails, the server can detect image abnormality problems in the target color image frame; if it is detected that the target color image frame has target image abnormality problems, the quality assessment rule can be adjusted according to the target image abnormality problems. Then, the image quality of the remaining color image frames not selected in the color map stream can be evaluated according to the adjusted quality assessment rule to update the image quality scores of the remaining color image frames; and based on the updated image quality scores of the remaining color image frames, a new target color image frame can be reselected from the remaining color image frames. By adjusting the quality assessment rule according to the target image abnormality problem, it is possible to focus on the image abnormality problem existing in the target color image frame when reselecting the target color image frame, so as to avoid the situation that the reselected target color image frame has the same image abnormality problem, thereby improving the accuracy of the reselected target color image frame and further improving the subsequent recognition effect.
[0118] Among them, the specific implementation of adjusting the quality assessment rule according to the target image abnormality problem can be: obtaining a correspondence table between the image abnormality problem and the evaluation dimension, which includes multiple image abnormality problems and the evaluation dimensions corresponding to each image abnormality problem. Secondly, the target evaluation dimension corresponding to the target image abnormality problem can be queried from the correspondence table. If the quality assessment rule includes the target evaluation dimension, the initial weight value corresponding to the target evaluation dimension can be increased in the quality assessment rule; if the quality assessment rule does not include the target evaluation dimension, the target evaluation dimension, the initial weight value corresponding to the target evaluation dimension, and the scoring algorithm under the target evaluation dimension can be added to the quality assessment rule.
[0119] It should be noted that if no target image abnormality is detected in the target color image frame, there is no need to adjust the quality assessment rule, nor to re-evaluate the image quality of each remaining color image frame not selected in the color image stream. In this case, the target color image frame can be directly reselected from each remaining color image frame not selected in the color image stream according to the image quality scores of each remaining color image frame not selected in the color image stream.
[0120] S307. Perform attribute recognition on the target object according to the reselected target color image frame to obtain the attribute information of the target object.
[0121] In specific implementation, the target infrared image frame corresponding to the reselected target color image frame can be reselected from the infrared image stream, and the target depth image frame corresponding to the reselected target color image frame can be reselected from the depth image stream. Then, attribute recognition can be performed on the target object according to the reselected target color image frame, the reselected target infrared image frame, and the reselected target depth image frame. It should be understood that the specific implementation manner of this step can refer to the relevant description of step S304 above and will not be elaborated here.
[0122] After the server obtains the attribute information of the target object through the above steps S301 - S307, it can also return the attribute information of the target object to the terminal, so that the terminal can perform business processing according to the attribute information of the target object. Among them, if the target object is a target face, the attribute information of the target object may include the identity information corresponding to the target face. Correspondingly, the business processing may include: face attendance processing, face payment processing, or face image classification processing, etc.; among them, face payment processing refers to: the processing of deducting a certain amount of electronic resources from the resource account associated with the identity information corresponding to the target face. If the target object is a gesture, the attribute information of the target object may include the gesture information corresponding to the target gesture (such as the shape of the gesture, the position information of each node in the hand, etc.). Correspondingly, the business processing may include: gesture photographing processing, gesture image classification processing, etc.
[0123] In an embodiment of the present invention, the server can receive a video stream of a target object uploaded by a terminal, and perform image quality evaluation on each color image frame in the color stream included in the video stream to obtain the image quality score of each color image frame. Then, based on the image quality scores of each color image frame, target color image frames with better image quality can be selected to perform attribute recognition on the target object, which can effectively improve the accuracy of the recognized attribute information. Moreover, since all operations involved in the entire object recognition process (such as image selection operation, attribute recognition operation) are executed in the server, the dependence on the hardware configuration of the terminal can be effectively eliminated, thereby effectively reducing the hardware cost of the terminal.
[0124] In practical applications, the object recognition method shown above can be applied to different application scenarios according to business requirements; for example, face recognition-based face payment scenarios, gesture recognition-based photo-taking scenarios, and so on. The so-called face recognition refers to the technology of obtaining the identity information corresponding to a face by using face multimedia information (such as an image frame containing a face); and the face recognition-based face payment scenario means: collecting an image frame containing the user's face for identity recognition, and automatically deducting the corresponding electronic resources from the resource account associated with the recognized identity information according to the item resources purchased by the user; a scenario where payment can be realized without the user operating their own mobile terminal (such as a mobile phone, tablet). The following takes applying this object recognition method to a face recognition-based face payment scenario as an example to elaborate on the specific application process of this object recognition method: Figures 2 - 3 In a specific application, when a user needs to pay the target electronic resources corresponding to a certain target item resource due to purchasing the target item resource, the user can align the target face with the 3D camera component of the terminal (such as a self-checkout device), as
[0125] shown. At this time, the terminal can call the relevant interfaces of the 3D camera component to collect face data of the target face, and obtain video stream data including an RGB stream (i.e., a color stream), an infrared stream, and a depth stream. After collecting the image frame of the target face, the terminal can display the collected image frame on the user interface. Then, the terminal can perform an uplink operation on the obtained video stream data by calling the relevant interfaces of the streaming media service of the server to upload the video stream data to the server, as Figure 5a shown. It should be noted that each image frame in the video stream data has a corresponding image ID. Figure 5b shown.
[0126] Correspondingly, the server side can define a quality score threshold through the image optimization service. Only when the image quality score of the RGB frame exceeds this quality score threshold will it be selected and sent to the object recognition service of the server. Then, after receiving the video stream data uploaded by the terminal, the image optimization service can perform image optimization processing on the RGB image stream. Specifically, the image quality of each RGB frame (i.e., color image frame) in the RGB image stream can be evaluated according to the image quality and face information to obtain the image quality score of each RGB frame. Secondly, according to the image quality scores of each RGB frame, the optimal frame (i.e., the RGB frame with the highest image quality score) can be selected from the color image stream as the target RGB frame, or any RGB frame with an image quality score greater than the quality score threshold can be selected from the color images as the target RGB frame. Then, the image ID of the target RGB frame can be obtained, and the target infrared image frame can be obtained from the infrared image stream and the target depth image frame can be obtained from the depth image stream according to this image ID; and the target RGB frame, the target infrared image frame, and the target depth image frame are passed to the object recognition service together. Then, the server can perform face recognition on the target face according to the target RGB frame, the target infrared image frame, and the target depth image frame through the object recognition service to obtain the identity information corresponding to the target face. When it is found that the face recognition is abnormal, the object recognition service can also notify the image optimization service to perform image optimization again to obtain the reselected target RGB frame; and obtain the corresponding infrared image frame and depth image frame according to the reselected target RGB frame to perform face recognition on the target face again until the face recognition is successful.
[0127] After the face recognition is successful, the identity information corresponding to the target face can be returned to the terminal so that the terminal can perform face payment processing according to this identity information; specifically, the terminal can obtain the target resource account associated with this identity information, and then deduct the target electronic resources corresponding to the target item resources purchased by the user from this target resource account. After successfully deducting the target electronic resources, the terminal can also output a payment success prompt in the user interface, such as Figure 5c shown.
[0128] The server in the embodiment of the present invention can receive the video image stream of the target object uploaded by the terminal, and evaluate the image quality of each color image frame in the color image stream included in the video image stream to obtain the image quality score of each color image frame. Then, based on the image quality scores of each color image frame, the target color image frame with better image quality can be selected to perform attribute recognition on the target object, which can effectively improve the accuracy of the recognized attribute information. And since all the operations involved in the entire object recognition process (such as image selection operation, attribute recognition operation) are executed in the server, the dependence on the hardware configuration of the terminal can be effectively eliminated, thereby effectively reducing the hardware cost of the terminal.
[0129] Based on the description of the embodiments of the above object recognition method, embodiments of the present invention also disclose an object recognition device, and the object recognition device may be a computer program (including program code) running in the above-mentioned server. The object recognition device can execute Figures 2 - 3 the method shown. Please refer to Figure 6 , and the object recognition device can run the following units:
[0130] A communication unit 601 that receives video stream data of a target object uploaded by a terminal, where the video stream data at least includes: a color map stream composed of multiple color image frames of the target object;
[0131] A processing unit 602 for performing image quality evaluation on each color image frame in the color map stream to obtain the image quality score of each color image frame;
[0132] The processing unit 602 is further configured to select a target color image frame from the color map stream based on the image quality scores of the color image frames; the target color image frame refers to the color image frame with the highest image quality score or the color image frame with an image quality score greater than a quality score threshold;
[0133] The processing unit 602 is further configured to perform attribute recognition on the target object according to the target color image frame to obtain attribute information of the target object.
[0134] In one implementation, the color map stream is uploaded in a streaming media manner, and the streaming media manner refers to a transmission manner in which the already captured image frames are synchronously transmitted during the process of capturing new image frames;
[0135] wherein, each color image frame in the color map stream is received by a first process, and each color image frame in the color map stream is subjected to image quality evaluation by a second process; the first process and the second process are executed in parallel.
[0136] In another implementation, when the processing unit 602 is used to perform image quality evaluation on each color image frame in the color map stream to obtain the image quality scores of the color image frames, it is specifically configured to:
[0137] Obtain a quality evaluation rule, and perform image quality evaluation on each color image frame in the color map stream according to the quality evaluation rule to obtain the image quality scores of the color image frames;
[0138] wherein, the quality evaluation rule includes multiple evaluation dimensions, an initial weight value corresponding to each evaluation dimension, and a scoring algorithm under each evaluation dimension; the multiple evaluation dimensions include: an image clarity dimension, an illumination quality dimension, a resolution dimension, and an object integrity dimension.
[0139] In another embodiment, when the processing unit 602 is configured to perform image quality assessment on each color image frame in the color map stream according to a quality assessment rule to obtain the image quality scores of the color image frames, it may specifically be configured to:
[0140] Perform quality detection on any color image frame in the color data stream under each evaluation dimension to obtain the quality detection results of the any color image frame under each evaluation dimension;
[0141] Determine the quality scores of the any color image frame under each evaluation dimension according to the quality detection results of the any color image frame under each evaluation dimension and the scoring algorithms under each evaluation dimension;
[0142] Determine the target weight values corresponding to the evaluation dimensions according to the quality detection results of the any color image frame under each evaluation dimension and the initial weight values corresponding to the evaluation dimensions;
[0143] Perform weighted summation on the quality scores of the any color image frame under each evaluation dimension by using the target weight values corresponding to the evaluation dimensions to obtain the image quality score of the any color image frame.
[0144] In another embodiment, if there is an occluder in the any color image frame, resulting in the target object in the any color image frame being incomplete, the quality detection result of the any color image frame under the object integrity dimension includes: the occlusion attribute of the occluder; the occlusion attribute includes at least one of the following: occlusion area and occlusion position;
[0145] The method for determining the target weight value corresponding to the object integrity dimension is as follows: Determine a weight adjustment factor for the object integrity dimension according to the occlusion attribute; and adjust the initial weight value corresponding to the object integrity dimension by using the weight adjustment factor to obtain the target weight value;
[0146] Wherein, if the occlusion attribute includes the occlusion area, the weight adjustment factor is: a first adjustment factor calculated according to the occlusion area and the image size of the any color image frame; if the occlusion attribute includes the occlusion position, the weight adjustment factor is: a second adjustment factor corresponding to the occlusion position; if the occlusion attribute includes the occlusion area and the occlusion position, the weight adjustment factor is: an adjustment factor obtained by integrating the weights of the first adjustment factor and the second adjustment factor.
[0147] In another implementation, when the processing unit 602 is used to perform attribute recognition on the target object according to the target color image frame to obtain the attribute information of the target object, it may specifically be used for:
[0148] Perform attribute recognition on the target object according to the target color image frame;
[0149] If the recognition is successful, use the recognized attribute information as the attribute information of the target object;
[0150] If the recognition fails, reselect a target color image frame from the remaining unselected color image frames in the color map stream, and perform attribute recognition on the target object according to the reselected target color image frame to obtain the attribute information of the target object.
[0151] In another implementation, when the processing unit 602 is used to, if the recognition fails, reselect a target color image frame from the remaining unselected color image frames in the color map stream, it may specifically be used for:
[0152] If the recognition fails, detect image anomaly problems in the target color image frame;
[0153] If it is detected that the target color image frame has a target image anomaly problem, adjust the quality assessment rule according to the target image anomaly problem;
[0154] Perform image quality assessment on the remaining unselected color image frames in the color map stream according to the adjusted quality assessment rule to update the image quality scores of the remaining color image frames;
[0155] Based on the updated image quality scores of the remaining color image frames, reselect a new target color image frame from the remaining color image frames.
[0156] In another implementation, when the processing unit 602 is used to adjust the quality assessment rule according to the target image anomaly problem, it may specifically be used for:
[0157] Obtain a correspondence table between image anomaly problems and evaluation dimensions, where the correspondence table includes multiple image anomaly problems and the evaluation dimensions corresponding to each image anomaly problem;
[0158] Query the target evaluation dimension corresponding to the target image anomaly problem from the correspondence table;
[0159] If the quality assessment rule includes the target evaluation dimension, increase the initial weight value corresponding to the target evaluation dimension in the quality assessment rule;
[0160] If the target evaluation dimension is not included in the quality evaluation rule, add the target evaluation dimension, the initial weight value corresponding to the target evaluation dimension, and the scoring algorithm under the target evaluation dimension to the quality evaluation rule.
[0161] In another implementation, the video stream data further includes: an infrared image stream composed of multiple infrared image frames of the target object, and a depth image stream composed of multiple depth image frames of the target object; correspondingly, when the processing unit 602 is used to perform attribute recognition on the target object according to the target color image frame, it can specifically be used for:
[0162] Select a target infrared image frame corresponding to the target color image frame from the infrared image stream, and select a target depth image frame corresponding to the target color image frame from the depth image stream;
[0163] Perform attribute recognition on the target object according to the target color image frame, the target infrared image frame, and the target depth image frame.
[0164] In another implementation, when the processing unit 602 is used to perform attribute recognition on the target object according to the target color image frame, the target infrared image frame, and the target depth image frame, it can specifically be used for:
[0165] Perform attribute recognition on the target object according to the target color image frame and the target infrared image frame to obtain first attribute information;
[0166] Perform attribute recognition on the target object according to the target color image frame and the target depth image frame to obtain second attribute information;
[0167] If the first attribute information and the second attribute information match, determine that the recognition is successful; otherwise, determine that the recognition fails.
[0168] In another implementation, the processing unit 602 can also be used for:
[0169] Return the attribute information of the target object to the terminal, so that the terminal performs business processing according to the attribute information of the target object;
[0170] Wherein, the target object is a target face, and the attribute information of the target object includes the identity information corresponding to the target face; the business processing includes: face attendance processing, face payment processing, or face image classification processing.
[0171] According to an embodiment of the present invention, Figure 2 or Figure 3Each step involved in the method shown can be performed by Figure 6 each unit in the object recognition device shown. For example, Figure 2 the step S201 shown in Figure 6 can be performed by the communication unit 601 shown in Figure 6 , and the steps S202 - S204 can be performed by the processing unit 602 shown in Figure 3 . Another example, Figure 6 the step S301 shown in Figure 6 can be performed by the communication unit 601 shown in , and the steps S302 - S307 can be performed by the processing unit 602 shown in
[0172] , and so on. Figure 6
[0173] According to another embodiment of the present invention, Figure 2 or Figure 3 each step involved in the corresponding method shown, by running a computer program (including program code) on a general computing device such as a computer including processing elements and storage elements such as a central processing unit (CPU), a random access storage medium (RAM), and a read-only storage medium (ROM), an object recognition device as shown in Figure 6 can be constructed, and the object recognition method of the embodiment of the present invention can be implemented. The computer program can be recorded on, for example, a computer-readable recording medium, loaded into the above computing device through the computer-readable recording medium, and run therein.
[0174] In the embodiments of the present invention, the server can receive the video stream of the target object uploaded by the terminal, and perform image quality evaluation on each color image frame in the color stream included in the video stream to obtain the image quality scores of each color image frame. Then, based on the image quality scores of each color image frame, the target color image frame with better image quality can be selected to perform attribute recognition on the target object, so as to effectively improve the accuracy of the recognized attribute information. Moreover, since all the operations involved in the entire object recognition process (such as image selection operation, attribute recognition operation) are executed in the server, the dependence on the hardware configuration of the terminal can be effectively eliminated, thereby effectively reducing the hardware cost of the terminal.
[0175] Based on the descriptions of the above method embodiments and apparatus embodiments, the embodiments of the present invention also provide a server. Please refer to Figure 7 , the server at least includes a processor 701, an input interface 702, an output interface 703, and a computer storage medium 704. Among them, the processor 701, the input interface 702, the output interface 703, and the computer storage medium 704 in the server can be connected through a bus or other means.
[0176] The computer storage medium 704 can be stored in the memory of the server. The computer storage medium 704 is used to store a computer program, and the computer program includes program instructions. The processor 701 is used to execute the program instructions stored in the computer storage medium 704. The processor 701 (or CPU (Central Processing Unit)) is the computing core and control core of the server, which is suitable for implementing one or more instructions, specifically suitable for loading and executing one or more instructions to implement the corresponding method flow or corresponding function. In one embodiment, the processor 701 in the embodiments of the present invention can be used to perform a series of object recognition processes, specifically including: receiving the video stream data of the target object uploaded by the terminal, and the video stream data at least includes: a color stream composed of multiple color image frames of the target object; performing image quality evaluation on each color image frame in the color stream to obtain the image quality scores of each color image frame; based on the image quality scores of each color image frame, selecting a target color image frame from the color stream; the target color image frame refers to the color image frame with the highest image quality score, or the color image frame with an image quality score greater than the quality score threshold; performing attribute recognition on the target object according to the target color image frame to obtain the attribute information of the target object, and so on.
[0177] An embodiment of the present invention further provides a computer storage medium (Memory). The computer storage medium is a memory device in the server and is used to store programs and data. It can be understood that the computer storage medium here can include both the built-in storage medium in the server and, of course, the extended storage medium supported by the server. The computer storage medium provides a storage space, and this storage space stores the operating system of the server. And, in this storage space, one or more instructions suitable for being loaded and executed by the processor 701 are also stored. These instructions can be one or more computer programs (including program codes). It should be noted that the computer storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory; optionally, it can also be at least one computer storage medium located far from the aforementioned processor.
[0178] In one embodiment, one or more instructions stored in the computer storage medium can be loaded and executed by the processor 701 to implement the corresponding steps of the method in the above-mentioned embodiment of the object recognition method; in a specific implementation, one or more instructions in the computer storage medium are loaded and executed by the processor 701 as follows:
[0179] Receive the video stream data about the target object uploaded by the terminal. The video stream data at least includes: a color map stream composed of multiple frame color image frames about the target object;
[0180] Perform image quality evaluation on each color image frame in the color map stream to obtain the image quality score of each color image frame;
[0181] Based on the image quality scores of each color image frame, select a target color image frame from the color map stream; the target color image frame refers to the color image frame with the highest image quality score or the color image frame whose image quality score is greater than the quality score threshold;
[0182] Perform attribute recognition on the target object according to the target color image frame to obtain the attribute information of the target object.
[0183] In one implementation manner, the color map stream is uploaded in a streaming media manner. The streaming media manner refers to a transmission manner in which the already captured image frames are synchronously transmitted during the process of capturing new image frames;
[0184] Among them, each color image frame in the color map stream is received through a first process, and each color image frame in the color map stream is subjected to image quality evaluation through a second process; the first process and the second process are executed in parallel.
[0185] In another implementation, when performing image quality assessment on each color image frame in the color map stream to obtain the image quality scores of the color image frames, the one or more instructions can be loaded and specifically executed by the processor 701:
[0186] Obtain a quality assessment rule, and perform image quality assessment on each color image frame in the color map stream according to the quality assessment rule to obtain the image quality scores of the color image frames;
[0187] Among them, the quality assessment rule includes multiple assessment dimensions, the initial weight value corresponding to each assessment dimension, and the scoring algorithm under each assessment dimension; the multiple assessment dimensions include: image sharpness dimension, lighting quality dimension, resolution dimension, and object integrity dimension.
[0188] In another implementation, when performing image quality assessment on each color image frame in the color map stream according to the quality assessment rule to obtain the image quality scores of the color image frames, the one or more instructions can be loaded and specifically executed by the processor 701:
[0189] Perform quality detection on any color image frame in the color data stream under each assessment dimension to obtain the quality detection results of the any color image frame under each assessment dimension;
[0190] Determine the quality scores of the any color image frame under each assessment dimension according to the quality detection results of the any color image frame under each assessment dimension and the scoring algorithm under each assessment dimension;
[0191] Determine the target weight values corresponding to the assessment dimensions according to the quality detection results of the any color image frame under each assessment dimension and the initial weight values corresponding to the assessment dimensions;
[0192] Perform weighted summation on the quality scores of the any color image frame under each assessment dimension by using the target weight values corresponding to the assessment dimensions to obtain the image quality score of the any color image frame.
[0193] In another implementation, if there is an occluder in the any color image frame, resulting in the target object in the any color image frame being incomplete, the quality detection result of the any color image frame in the object integrity dimension includes: the occlusion attribute of the occluder; the occlusion attribute includes at least one of the following: occlusion area and occlusion position;
[0194] The method for determining the target weight value corresponding to the object integrity dimension is as follows: determining a weight adjustment factor for the object integrity dimension according to the occlusion attribute; and adjusting the initial weight value corresponding to the object integrity dimension by using the weight adjustment factor to obtain the target weight value;
[0195] Wherein, if the occlusion attribute includes the occlusion area, the weight adjustment factor is: a first adjustment factor calculated according to the occlusion area and the image size of any color image frame; if the occlusion attribute includes the occlusion position, the weight adjustment factor is: a second adjustment factor corresponding to the occlusion position; if the occlusion attribute includes the occlusion area and the occlusion position, the weight adjustment factor is: an adjustment factor obtained by integrating the weights of the first adjustment factor and the second adjustment factor.
[0196] In another implementation manner, when performing attribute recognition on the target object according to the target color image frame to obtain the attribute information of the target object, the one or more instructions can be loaded and specifically executed by the processor 701:
[0197] Performing attribute recognition on the target object according to the target color image frame;
[0198] If the recognition is successful, the recognized attribute information is used as the attribute information of the target object;
[0199] If the recognition fails, a target color image frame is reselected from the remaining unselected color image frames in the color map stream, and attribute recognition is performed on the target object according to the reselected target color image frame to obtain the attribute information of the target object.
[0200] In another implementation manner, when, if the recognition fails, a target color image frame is reselected from the remaining unselected color image frames in the color map stream, the one or more instructions can be loaded and specifically executed by the processor 701:
[0201] If the recognition fails, detecting an image anomaly problem with the target color image frame;
[0202] If it is detected that the target color image frame has a target image anomaly problem, adjusting the quality assessment rule according to the target image anomaly problem;
[0203] Performing image quality assessment on the remaining unselected color image frames in the color map stream according to the adjusted quality assessment rule to update the image quality scores of the remaining color image frames;
[0204] Based on the updated image quality scores of the remaining color image frames, a new target color image frame is reselected from the remaining color image frames.
[0205] In another implementation, when adjusting the quality assessment rule according to the target image anomaly problem, the one or more instructions can be loaded and specifically executed by the processor 701:
[0206] Obtain a correspondence table between the image anomaly problems and the evaluation dimensions, where the correspondence table includes multiple image anomaly problems and the evaluation dimensions corresponding to each image anomaly problem;
[0207] Query the target evaluation dimension corresponding to the target image anomaly problem from the correspondence table;
[0208] If the target evaluation dimension is included in the quality assessment rule, increase the initial weight value corresponding to the target evaluation dimension in the quality assessment rule;
[0209] If the target evaluation dimension is not included in the quality assessment rule, add the target evaluation dimension, the initial weight value corresponding to the target evaluation dimension, and the scoring algorithm under the target evaluation dimension to the quality assessment rule.
[0210] In another implementation, the video stream data further includes: an infrared image stream composed of multiple infrared image frames of the target object, and a depth image stream composed of multiple depth image frames of the target object; correspondingly, when performing attribute recognition on the target object according to the target color image frame, the one or more instructions can be loaded and specifically executed by the processor 701:
[0211] Select a target infrared image frame corresponding to the target color image frame from the infrared image stream, and select a target depth image frame corresponding to the target color image frame from the depth image stream;
[0212] Perform attribute recognition on the target object according to the target color image frame, the target infrared image frame, and the target depth image frame.
[0213] In another implementation, when performing attribute recognition on the target object according to the target color image frame, the target infrared image frame, and the target depth image frame, the one or more instructions can be loaded and specifically executed by the processor 701:
[0214] Perform attribute recognition on the target object according to the target color image frame and the target infrared image frame to obtain first attribute information;
[0215] Perform attribute recognition on the target object according to the target color image frame and the target depth image frame to obtain second attribute information;
[0216] If the first attribute information and the second attribute information match, it is determined that the recognition is successful; otherwise, it is determined that the recognition fails.
[0217] In another implementation manner, the one or more instructions may also be loaded and specifically executed by the processor 701:
[0218] Return the attribute information of the target object to the terminal, so that the terminal performs service processing according to the attribute information of the target object;
[0219] Wherein, the target object is a target face, and the attribute information of the target object includes the identity information corresponding to the target face; the service processing includes: face attendance processing, face payment processing, or face image classification processing.
[0220] In the embodiment of the present invention, the server can receive the video stream of the target object uploaded by the terminal, and perform image quality evaluation on each color image frame in the color stream included in the video stream to obtain the image quality score of each color image frame. Then, based on the image quality scores of each color image frame, a target color image frame with better image quality can be selected to perform attribute recognition on the target object, which can effectively improve the accuracy of the recognized attribute information. And since all the operations involved in the entire object recognition process (such as image selection operation, attribute recognition operation) are executed in the server, the dependence on the hardware configuration of the terminal can be effectively eliminated, thereby effectively reducing the hardware cost of the terminal.
[0221] It should be noted that, according to one aspect of the present application, a computer program product or a computer program is also provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above Figure 2 or Figure 3 The methods provided in various alternative ways in the object recognition method embodiments shown.
[0222] Moreover, it should be understood that the above-disclosed are only the preferred embodiments of the present invention. Of course, the scope of the rights of the present invention cannot be limited by this. Therefore, equivalent changes made according to the claims of the present invention still fall within the scope covered by the present invention.
Claims
1. An object recognition method, characterized in that, it includes: Receiving video stream data about a target object uploaded by a terminal, where the video stream data at least includes: a color map stream composed of multiple color image frames about the target object, an infrared map stream composed of multiple infrared image frames about the target object, and a depth map stream composed of multiple depth image frames about the target object; wherein, the color image frame, the infrared image frame, and the depth image frame generated at the same moment contain the same image identifier, and the image frames generated at different moments include different image identifiers; Performing image quality assessment on each color image frame in the color map stream to obtain the image quality score of each color image frame; Based on the image quality scores of each color image frame, selecting a target color image frame from the color map stream; the target color image frame refers to the color image frame with the highest image quality score, or the color image frame with an image quality score greater than the quality score threshold; Performing attribute recognition on the target object according to the target color image frame to obtain the attribute information of the target object; Among them, the method of performing attribute recognition on the target object according to the target color image frame includes: based on the image identifier of the target color image frame, selecting a target infrared image frame corresponding to the target color image frame from the infrared map stream, and selecting a target depth image frame corresponding to the target color image frame from the depth map stream; performing attribute recognition on the target object according to the target color image frame and the target infrared image frame to obtain first attribute information; performing attribute recognition on the target object according to the target color image frame and the target depth image frame to obtain second attribute information; if the first attribute information and the second attribute information match, it is determined that the recognition is successful; otherwise, it is determined that the recognition fails.
2. The method according to claim 1, characterized in that, the color map stream is uploaded in a streaming media manner, and the streaming media manner refers to a transmission manner in which the already collected image frames are synchronously transmitted during the process of collecting new image frames; wherein, each color image frame in the color map stream is received by a first process, and each color image frame in the color map stream is subjected to image quality assessment by a second process; the first process and the second process are executed in parallel.
3. The method according to claim 1 or 2, characterized in that, performing image quality assessment on each color image frame in the color map stream to obtain the image quality score of each color image frame includes: Obtaining a quality assessment rule, and performing image quality assessment on each color image frame in the color map stream according to the quality assessment rule to obtain the image quality score of each color image frame; wherein, the quality assessment rule includes multiple evaluation dimensions, an initial weight value corresponding to each evaluation dimension, and a scoring algorithm under each evaluation dimension; the multiple evaluation dimensions include: image clarity dimension, lighting quality dimension, resolution dimension, and object integrity dimension.
4. The method according to claim 3, characterized in that, Performing image quality assessment on each color image frame in the color map stream according to the quality assessment rules to obtain the image quality scores of the respective color image frames, including: Performing quality detection on any color image frame in the color map stream under each evaluation dimension to obtain the quality detection results of the any color image frame under each evaluation dimension; Determining the quality scores of the any color image frame under each evaluation dimension according to the quality detection results of the any color image frame under each evaluation dimension and the scoring algorithms under each evaluation dimension; Determining the target weight values corresponding to the respective evaluation dimensions according to the quality detection results of the any color image frame under each evaluation dimension and the initial weight values corresponding to the respective evaluation dimensions; Performing weighted summation on the quality scores of the any color image frame under each evaluation dimension by using the target weight values corresponding to the respective evaluation dimensions to obtain the image quality score of the any color image frame.
5. The method according to claim 4, wherein, if there is an occluder in the any color image frame, resulting in the target object in the any color image frame being incomplete, then the quality detection result of the any color image frame under the object integrity dimension includes: the occlusion attribute of the occluder; the occlusion attribute includes at least one of the following: occlusion area and occlusion position; The method for determining the target weight value corresponding to the object integrity dimension is as follows: determining a weight adjustment factor for the object integrity dimension according to the occlusion attribute; and adjusting the initial weight value corresponding to the object integrity dimension by using the weight adjustment factor to obtain the target weight value; wherein, if the occlusion attribute includes the occlusion area, the weight adjustment factor is: a first adjustment factor calculated according to the occlusion area and the image size of the any color image frame; if the occlusion attribute includes the occlusion position, the weight adjustment factor is: a second adjustment factor corresponding to the occlusion position; if the occlusion attribute includes the occlusion area and the occlusion position, the weight adjustment factor is: an adjustment factor obtained by integrating the weights of the first adjustment factor and the second adjustment factor.
6. The method according to claim 3, wherein, the performing attribute recognition on the target object according to the target color image frame to obtain the attribute information of the target object includes: Performing attribute recognition on the target object according to the target color image frame; If the recognition is successful, using the recognized attribute information as the attribute information of the target object; If the recognition fails, reselecting a target color image frame from the remaining unselected color image frames in the color map stream, and performing attribute recognition on the target object according to the reselected target color image frame to obtain the attribute information of the target object.
7. The method according to claim 6, wherein, the reselecting a target color image frame from the remaining unselected color image frames in the color map stream if the recognition fails includes: If the recognition fails, perform image anomaly problem detection on the target color image frame; If it is detected that the target color image frame has a target image anomaly problem, adjust the quality assessment rule according to the target image anomaly problem; Perform image quality assessment on each remaining color image frame in the color image stream that has not been selected according to the adjusted quality assessment rule to update the image quality scores of each remaining color image frame; Based on the updated image quality scores of each remaining color image frame, re-select a new target color image frame from each remaining color image frame.
8. The method according to claim 7, wherein, the adjusting the quality assessment rule according to the target image anomaly problem includes: Obtain a correspondence table between image anomaly problems and evaluation dimensions, where the correspondence table includes multiple image anomaly problems and the evaluation dimensions corresponding to each image anomaly problem; Query the target evaluation dimension corresponding to the target image anomaly problem from the correspondence table; If the quality assessment rule includes the target evaluation dimension, increase the initial weight value corresponding to the target evaluation dimension in the quality assessment rule; If the quality assessment rule does not include the target evaluation dimension, add the target evaluation dimension, the initial weight value corresponding to the target evaluation dimension, and the scoring algorithm under the target evaluation dimension to the quality assessment rule.
9. The method according to claim 1, wherein, the method further includes: Return the attribute information of the target object to the terminal, so that the terminal performs business processing according to the attribute information of the target object; wherein, the target object is a target face, and the attribute information of the target object includes the identity information corresponding to the target face; the business processing includes: face attendance processing, face payment processing, or face image classification processing.
10. An object recognition device, wherein, it includes: A communication unit that receives video stream data of a target object uploaded by a terminal, where the video stream data at least includes: a color image stream composed of multiple color image frames of the target object, an infrared image stream composed of multiple infrared image frames of the target object, and a depth image stream composed of multiple depth image frames of the target object; wherein, the color image frame, infrared image frame, and depth image frame generated at the same moment contain the same image identifier, and the image frames generated at different moments include different image identifiers; A processing unit for performing image quality assessment on each color image frame in the color image stream to obtain the image quality scores of each color image frame; The processing unit is further configured to select a target color image frame from the color image stream based on the image quality scores of each color image frame; the target color image frame refers to the color image frame with the highest image quality score, or the color image frame with an image quality score greater than the quality score threshold; The processing unit is further configured to perform attribute recognition on the target object according to the target color image frame to obtain the attribute information of the target object; Among them, the method for performing attribute recognition on the target object according to the target color image frame includes: based on the image identifier of the target color image frame, selecting a target infrared image frame corresponding to the target color image frame from the infrared image stream, and selecting a target depth image frame corresponding to the target color image frame from the depth image stream; performing attribute recognition on the target object according to the target color image frame and the target infrared image frame to obtain first attribute information; performing attribute recognition on the target object according to the target color image frame and the target depth image frame to obtain second attribute information; if the first attribute information and the second attribute information match, it is determined that the recognition is successful; otherwise, it is determined that the recognition fails.
11. A server, comprising an input interface and an output interface, Characterized in that, It further includes: A processor, adapted to implement one or more instructions; And, A computer storage medium, the computer storage medium stores one or more instructions, and the one or more instructions are adapted to be loaded and executed by the processor to perform the object recognition method according to any one of claims 1-10.
12. A computer storage medium, Characterized in that, The computer storage medium stores one or more instructions, and the one or more instructions are adapted to be loaded and executed by a processor to perform the object recognition method according to any one of claims 1-10.
13. A computer program product, Characterized in that, The computer program product includes computer instructions, and the computer instructions are executed by a processor to implement the object recognition method according to any one of claims 1-10.
Citation Information
Patent Citations
Interaction method and interaction device based on lip language
CN106774856A
Method and device for detecting face image
CN110276277A