Image recognition method, attendance checking device, storage medium and computer program product
By integrating tongue coating analysis and facial recognition into attendance equipment, the problem of single function of attendance equipment is solved, the integration of identity authentication and health monitoring is achieved, the attendance efficiency is improved and the health management application is expanded.
Patent Information
- Application Number
- CN202510682732.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-09-12
AI Technical Summary
Existing attendance equipment has a single function and cannot effectively integrate user health data with attendance records, resulting in data fragmentation and inefficiency, and cannot achieve continuous tracking of user health status.
By acquiring facial images containing tongue coating, facial recognition is performed to obtain identity information, and health status information is obtained based on tongue coating analysis. Combined with a dual-spectrum camera and deep learning algorithm, facial recognition and tongue coating analysis functions are integrated to achieve identity authentication and health monitoring.
It realizes the simultaneous acquisition of user identity and health status information in the attendance device, improves attendance efficiency, expands health management application scenarios, and reduces the upgrade cost of health management.
Smart Images

Figure CN120636007A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of big data technology, and in particular to an image recognition method, attendance equipment, storage medium and computer program product. Background Art
[0002] With the development of science and technology, it has become quite common to use attendance equipment to record whether people arrive on time. However, existing attendance equipment has the problem of single function. Summary of the Invention
[0003] To solve related technical problems, embodiments of the present application provide an image recognition method, attendance equipment, storage medium, and computer program product.
[0004] The technical solution of the embodiment of the present application is implemented as follows:
[0005] The present invention provides an image recognition method for use in an attendance device. The method includes:
[0006] Acquire a first image; the first image represents a face image including a tongue coating, or different images in the first image include a face image or a tongue coating image;
[0007] Performing facial recognition on the first image to obtain first information, where the first information indicates identity information of the user;
[0008] Based on the tongue coating contained in the first image, second information is obtained, where the second information indicates the user's health status information.
[0009] In the above solution, the second information includes tongue coating category information; and obtaining the second information based on the tongue coating included in the first image includes:
[0010] determining a first image from the first image, where the first image represents an image of the first image at least including tongue coating;
[0011] Marking or extracting a region containing tongue coating in the first image to obtain a second image;
[0012] The second image is processed by the first model to obtain tongue coating category information corresponding to the second image; the first model is used to process the image containing the tongue coating and output the tongue coating category information.
[0013] In the above solution, the first model includes an input layer, a first layer, and an output layer; and processing the second image by the first model to obtain tongue coating category information corresponding to the second image includes:
[0014] Inputting the second image into the input layer and performing a first operation to obtain a third image; the first operation includes one or more of the following: normalization, cropping, scaling, and image enhancement;
[0015] Input the third image into the first layer to perform tongue coating category probability prediction, and obtain the probability of each tongue coating category corresponding to the third image;
[0016] The probabilities of the tongue coating categories corresponding to the third image are input into the output layer for category conversion to obtain tongue coating category information corresponding to the third image. The tongue coating category information corresponding to the third image is used to determine the tongue coating category information corresponding to the second image.
[0017] In the above solution, the first layer includes a convolutional layer, a first residual layer, a pooling layer, a second residual layer, and a fully connected layer; inputting the third image into the first layer to perform tongue coating category probability prediction to obtain the probability of each tongue coating category corresponding to the third image includes:
[0018] Inputting the third image into the convolution layer to perform feature extraction to obtain first feature information of the third image;
[0019] Inputting the first feature information of the third image into the first residual layer for residual processing to obtain second feature information of the third image;
[0020] Inputting the first feature information of the third image into the pooling layer for feature fusion to obtain third feature information of the third image;
[0021] Inputting the third feature information and the second feature information of the third image into the second residual layer for residual processing to obtain fourth feature information of the third image;
[0022] The fourth feature information of the third image is input into the fully connected layer for linear transformation to obtain the probability of each tongue coating category corresponding to the third image.
[0023] In the above solution, the convolution layer includes a first convolution kernel and / or a second convolution kernel, and the first feature information of the third image includes local feature information and / or global feature information of the third image; inputting the third image into the convolution layer for feature extraction to obtain the first feature information of the third image includes:
[0024] performing feature extraction on the third image using the first convolution kernel to obtain local feature information of the third image; and / or
[0025] Feature extraction is performed on the third image using the second convolution kernel to obtain global feature information of the third image, and the size of the second convolution kernel is larger than the size of the first convolution kernel.
[0026] In the above solution, the first residual layer includes a first residual block, a second residual block, and a third residual block, and the second feature information of the third image includes the fifth feature information, the sixth feature information, the seventh feature information, or the eighth feature information of the third image; and inputting the first feature information of the third image into the first residual layer for residual processing to obtain the second feature information of the third image includes:
[0027] performing residual processing on the first feature information of the third image using the first residual block to obtain fifth feature information of the third image; or
[0028] In a case where the third image has multiple textures and / or colors interwoven, performing residual processing on the first feature information of the third image using the first residual block and the second residual block to obtain sixth feature information of the third image; or
[0029] In the case where there are flaws and / or defects in the third image, performing residual processing on the first feature information of the third image using the first residual block and the third residual block to obtain seventh feature information of the third image; or
[0030] When there are multiple textures and / or colors interwoven in the third image, and there are flaws and / or defects in the third image, residual processing is performed on the first feature information of the third image through the first residual block, the second residual block and the third residual block to obtain the eighth feature information of the third image.
[0031] In the above solution, obtaining the first image includes:
[0032] Determining or adjusting photography parameters based on the brightness of the shooting environment, the photography parameters including one or more of fill light brightness, sensitivity, and shutter speed;
[0033] Image acquisition is performed based on the photographic parameters to obtain the first image.
[0034] The present application also provides an image recognition device, comprising:
[0035] an acquisition module, configured to acquire a first image; the first image represents a face image including a tongue coating, or different images in the first image include a face image or a tongue coating image;
[0036] a recognition module, configured to perform face recognition on the first image to obtain first information, where the first information indicates identity information of a user;
[0037] The information acquisition module is used to obtain second information based on the tongue coating contained in the first image, where the second information indicates the user's health status information.
[0038] The embodiment of the present application also provides an attendance device, comprising: a processor and a communication interface; wherein,
[0039] The processor is used to obtain a first image; the first image represents a facial image containing tongue coating, or different images in the first image include facial images or tongue coating images; face recognition is performed on the first image to obtain first information, and the first information indicates the user's identity information; based on the tongue coating contained in the first image, second information is obtained, and the second information indicates the user's health status information.
[0040] The embodiment of the present application further provides an attendance device, comprising a processor and a memory for storing a computer program that can be run on the processor.
[0041] The processor is configured to execute the steps of any of the above methods when running the computer program.
[0042] An embodiment of the present application further provides a storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above methods are implemented.
[0043] An embodiment of the present application further provides a computer program product, comprising a computer program, which implements the steps of any of the above methods when executed by a processor.
[0044] In the image recognition method, attendance device, storage medium, and computer program product provided in the embodiments of the present application, a first image is obtained; facial recognition is performed on the first image to obtain first information, the first information indicating the user's identity information; based on the tongue coating contained in the first image, second information is obtained, the second information indicating the user's health status information. The above scheme can obtain both the user's identity information and the user's health status information through the first image, integrate the functions of facial recognition and tongue coating analysis, and integrate the attendance and health monitoring functions, thereby improving attendance efficiency while expanding the application scenarios of health management. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 This is a flowchart of an image recognition method according to an embodiment of the present application;
[0046] Figure 2 This is a schematic diagram of a full-face image of a user sticking out his tongue in an embodiment of the present application;
[0047] Figure 3 This is a schematic image diagram of the tongue coating of a user according to an embodiment of the present application;
[0048] Figure 4 This is a schematic structural diagram of the first model of the embodiment of the present application;
[0049] Figure 5 This is a schematic diagram of the structure of the image recognition device according to an embodiment of the present application;
[0050] Figure 6 This is a structural diagram of the attendance equipment in an embodiment of the present application. DETAILED DESCRIPTION
[0051] Traditional attendance devices only implement identity authentication and have the following problems:
[0052] User health data and attendance records are managed separately, resulting in data fragmentation;
[0053] User health monitoring requires separate equipment, and data fusion between the two devices relies on manual operation, resulting in low efficiency.
[0054] It is impossible to continuously track the user's health status and lack dynamic monitoring of the user's health status.
[0055] Based on this, in various embodiments of the present application, a first image is acquired; facial recognition is performed on the first image to obtain first information, which indicates the user's identity information; and second information is obtained based on the tongue coating contained in the first image, which indicates the user's health status information. The above solution can obtain both the user's identity information and the user's health status information through the first image, integrating the functions of facial recognition and tongue coating analysis, and integrating the attendance and health monitoring functions, thereby improving attendance efficiency while expanding the application scenarios of health management.
[0056] The present application will be described in further detail below with reference to the accompanying drawings and embodiments.
[0057] The embodiment of the present application provides an image recognition method, which is applied to an attendance device. The attendance device may have a display screen or display function, or may be connected to a display screen. Figure 1 Shown, including:
[0058] Step 101: Acquire a first image.
[0059] The first image representation includes a face image containing tongue coating, or the different images in the first image include a face image or a tongue coating image.
[0060] Here, the first image can be acquired using a dual-spectrum camera or other device that supports image acquisition. The dual-spectrum camera can integrate a visible light camera and a near-infrared camera. For example, a visible light camera can be used to acquire facial images for facial recognition, while a near-infrared camera can be used to acquire tongue coating images to capture tongue coating details. The specific embodiments are not limited thereto.
[0061] In actual application, the visible light camera can use a white light-emitting diode (LED) light source that is close to natural light, and the color temperature can be set between 5000 Kelvin (K) and 6500K to more realistically restore the color and texture of the tongue coating; multiple LED light sources can also be used to make the light evenly illuminate the tongue coating, reduce shadows and reflections, and ensure the clarity of the tongue coating image.
[0062] It is worth noting that the device for acquiring the first image can be placed on a rotatable bracket, or can be integrated with the rotatable bracket so that the device can be adjusted through the rotatable bracket and / or the user can be guided to shoot by voice, so that the user's tongue coating in the captured first image is naturally stretched and can be accurately focused.
[0063] The first image may include an image and / or a video. A full-face image of the user sticking out the tongue may be obtained by a visible light camera, for example Figure 2 As shown, at this time, the first image representation includes a face image of a person with tongue coating. The image of the tongue coating of the user can be obtained by the near infrared camera in the dual-spectrum camera, for example Figure 3 As shown, a full-face image of the user without sticking out the tongue is obtained through the visible light camera in the dual-spectrum camera. At this time, the different images in the first image include a face image or a tongue coating image.
[0064] In one embodiment, acquiring the first image includes:
[0065] Determining or adjusting photography parameters based on the brightness of the shooting environment, the photography parameters including one or more of fill light brightness, sensitivity, and shutter speed;
[0066] Image acquisition is performed based on the photographic parameters to obtain the first image.
[0067] Here, the brightness of the shooting environment can be monitored using a light sensor built into the device used to acquire the first image. The first image can be acquired by determining or adjusting photographic parameters of an image sensor built into the device used to acquire the first image. The image sensor can be high-resolution to capture subtle texture and color variations of the tongue coating, facilitating subsequent analysis of the tongue coating image. The image sensor can be set to RGB (Red, Green, Blue) color mode to accurately capture color information and color details of the tongue coating.
[0068] Determine or adjust photography parameters based on the brightness of the shooting environment, which may include one or more of the following:
[0069] When the brightness of the shooting environment is less than or equal to a first threshold, increasing the fill light brightness and / or improving the ISO (International Standards Organization);
[0070] When the brightness of the shooting environment is greater than the second threshold, reducing the fill light brightness and / or adjusting the ISO to be less than or equal to the third threshold;
[0071] It is determined that the shutter speed is greater than or equal to a fourth threshold.
[0072] The first threshold, the second threshold, the third threshold, and the fourth threshold can be set based on actual needs. The first threshold and the second threshold can be the same or different. The second threshold is greater than or equal to the first threshold.
[0073] It is worth noting that the shutter speed may not be determined or adjusted based on the brightness of the shooting environment.
[0074] In this embodiment, determining or adjusting the photographic parameters can improve the quality of the first image and make the quality of the first image stable; determining or adjusting the fill light brightness can avoid overbrightness or overdarkness; determining or adjusting the sensitivity can reduce image noise and increase image quality; using a faster shutter speed can reduce image blur caused by slight movement of the device.
[0075] Step 102: Perform face recognition on the first image to obtain first information.
[0076] The first information indicates identity information of the user.
[0077] Here, performing face recognition on the first image to obtain the first information may include:
[0078] Determining a fourth image from the first image, the fourth image representing an image in the first image that at least includes a human face;
[0079] Marking or extracting a region containing a face in the fourth image to obtain a fifth image;
[0080] The fifth image is processed by the second model to obtain the first information corresponding to the first image; the second model is used to process the image containing the face and output the identity information corresponding to the face.
[0081] One fourth image corresponds to one fifth image.
[0082] The area containing the face is marked or extracted in the fourth image. A deep learning algorithm that supports face region recognition can be used to perform face region recognition on the fourth image, outputting bounding box coordinates of the face. The fourth image is then cropped based on the bounding box coordinates of the face to extract the area containing the face in the fourth image. It should be understood that the cropping of the fourth image based on the bounding box coordinates of the face can also include normalizing the cropped image to a fixed resolution size to obtain the fifth image.
[0083] Performing face recognition on the first image to obtain first information may also include:
[0084] Determining a fourth image from the first image, the fourth image representing an image in the first image that at least includes a human face;
[0085] The fourth image is processed by the second model to obtain the first information corresponding to the first image; the second model is used to process the image containing the face and output the identity information corresponding to the face.
[0086] The second model can be a model that supports face recognition. For example, the second model includes a ranking neural network (RankNet) and key point detection. Key point detection and the ranking neural network are combined to improve recognition robustness. The ranking neural network can be understood as a RankNet neural network. Variants or improved RankNet neural networks all belong to ranking neural networks.
[0087] Processing the fourth image using the second model to obtain first information corresponding to the first image may include: sorting the quality of multiple fourth images in the first image using a sorting neural network, performing feature recognition on the quality-sorted fourth images using key point detection, and obtaining the first information corresponding to the first image when the feature recognition results of the multiple fourth images are the same; and reacquiring the first image or determining the most feature recognition results as the first information corresponding to the first image when the feature recognition results of the multiple fourth images are different. Sorting the quality of multiple fourth images in the first image using the sorting neural network may be performed by sorting all fourth images in the first image using the sorting neural network, or by sorting a set number of fourth images in the first image using the sorting neural network.
[0088] The implementation process of processing the fifth image through the second model to obtain the first information corresponding to the first image can refer to the description content of processing the fourth image through the second model to obtain the first information corresponding to the first image, and will not be repeated here.
[0089] Step 103: Obtain second information based on the tongue coating contained in the first image.
[0090] The second information indicates health status information of the user.
[0091] Here, based on the tongue coating included in the first image, the category information of the tongue coating can be identified, and each category information of the tongue coating corresponds to a health status information.
[0092] It should be understood that the first information and / or the second information may be displayed on the attendance device, or may be output or transmitted to other terminals or devices; the first information and the second information may be associated and stored in the backend, or may be encrypted and transmitted to the cloud. If the second information indicates that the user's health status is abnormal, a prompt message may be output by the attendance device, or a prompt message may be sent to the terminal of the user corresponding to the first information, to alert the user.
[0093] It is worth noting that the first information and the second information are stored in association. Based on the first information and the second information stored in association, a correlation model between the user's health status and attendance pattern can be established according to the time series to warn the user of sub-health risks.
[0094] In this embodiment, through the first image, both the user's identity information and the user's health status information can be obtained, the functions of face recognition and tongue coating analysis are integrated, and the attendance and health monitoring functions are integrated, which improves attendance efficiency while expanding the application scenarios of health management.
[0095] In one embodiment, the second information includes tongue coating category information; and obtaining the second information based on the tongue coating included in the first image includes:
[0096] determining a first image from the first image, where the first image represents an image of the first image at least including tongue coating;
[0097] Marking or extracting a region containing tongue coating in the first image to obtain a second image;
[0098] The second image is processed by the first model to obtain tongue coating category information corresponding to the second image; the first model is used to process the image containing the tongue coating and output the tongue coating category information.
[0099] Here, the first image determined from the first image may include one or more images. In the case where the first image includes multiple images, one first image corresponds to one second image, and the tongue coating category information that the first image ultimately corresponds to is determined based on the tongue coating category information corresponding to the multiple second images. Determining the tongue coating category information that the first image ultimately corresponds to based on the tongue coating category information corresponding to the multiple second images may include: when the tongue coating category information corresponding to the multiple second images is the same, determining the tongue coating category information corresponding to any second image as the tongue coating category information that the first image ultimately corresponds to; or, when the tongue coating category information corresponding to a set number of second images is the same, determining any tongue coating category information corresponding to a set number of second images as the tongue coating category information that the first image ultimately corresponds to, without specific limitation.
[0100] Marking or extracting the area containing the tongue coating in the first image to obtain the second image may include: identifying and segmenting the area containing the tongue coating in the first image through a Mask R-CNN (Mask Region-based Convolutional Neural Network) or other image segmentation algorithms that support tongue coating recognition to obtain a second image containing only the tongue coating area; or, identifying the tongue coating area of the first image through a deep learning algorithm that supports tongue coating area recognition, and marking the identified tongue coating area on the first image to obtain the second image. For example, by using other image segmentation algorithms that support tongue coating recognition, invalid areas such as teeth and / or lips in the first image are automatically cropped to obtain the second image, so as to improve the accuracy of tongue coating feature extraction. For example, the identified tongue coating area on the first image is marked by a mask, and the mask includes binary information indicating whether each pixel on the first image belongs to the tongue coating area.
[0101] Before marking or extracting the area containing the tongue coating in the first image, the first image may be converted from the RGB color space to the Hue-Saturation-Value (HSV) color space to better distinguish the color of the tongue coating.
[0102] In practical applications, Mask R-CNN is used to identify and segment the tongue coating region in the first image, extracting features from the first image and classifying them at the pixel level. During Mask R-CNN training, convolutional layers can be added to Mask R-CNN to capture local features in the first image. Pooling layers can also be added to Mask R-CNN to reduce the resolution of the first image, reducing computational effort while retaining important feature information. A cross-entropy loss function can also be added to Mask R-CNN to improve the accuracy of tongue coating region segmentation.
[0103] It is understandable that the first model may be a trained model. For example, the first model includes but is not limited to Residual Network 50 (ResNet-50). Before training the first model, sample images are obtained; the sample images include tongue coating images of people of different ages, genders, and races under different lighting conditions and / or different shooting angles; the sample images are divided into a training set, a validation set, and a test set. Among them, the training set is used for parameter learning of the first model; the validation set is used to adjust the hyperparameters of the first model and monitor the training process of the first model to prevent overfitting; the test set is used to evaluate the final performance of the first model.
[0104] The training process of the first model may include: inputting the sample images of the training set and the validation set into the first model to obtain the tongue coating category information corresponding to each sample image; calculating the loss value based on the tongue coating category information of each sample image and the tongue coating category information calibrated by the corresponding sample image; updating the parameters of the first model based on the loss value and the back propagation algorithm until the end training condition of the first model is reached. The end training condition of the first model represents that the updating of the parameters of the first model can be ended, or the training of the first model is completed. The end training condition of the first model is, for example, that the accuracy of the first model is greater than the set value, and / or, the parameter update of the first model reaches a set number of times, and / or, the calculated loss value is less than the set loss value. The specific end training condition of the first model is not limited. The trained model can be evaluated using the test set, and the first model can be further optimized according to the evaluation results to finally obtain the trained first model.
[0105] Tongue coating category information includes but is not limited to one or more of the following: normal tongue coating, damp-heat tongue coating, and blood stasis tongue coating.
[0106] It is worth noting that the second information may also include health recommendations and / or analysis conclusions. When outputting tongue coating category information, the health recommendations and / or analysis conclusions corresponding to the tongue coating category information may be determined based on a set association relationship. For example, the set association relationship may represent the association between the tongue coating category information and cases in a traditional Chinese medicine database.
[0107] In this embodiment, a method for obtaining the second information based on the tongue coating contained in the first image is clearly defined, and the efficiency of obtaining the second information is improved. The first image is determined from the first image, and the area containing the tongue coating is marked or extracted in the first image to obtain the second image. This makes the tongue coating area in the second image clearer, and can improve the accuracy of the first model in processing the second image. The efficiency of processing the second image by the first model is also higher.
[0108] In one embodiment, the first model includes an input layer, a first layer, and an output layer; and processing the second image using the first model to obtain tongue coating category information corresponding to the second image includes:
[0109] Inputting the second image into the input layer and performing a first operation to obtain a third image; the first operation includes one or more of the following: normalization, cropping, scaling, and image enhancement;
[0110] Input the third image into the first layer to perform tongue coating category probability prediction, and obtain the probability of each tongue coating category corresponding to the third image;
[0111] The probabilities of the tongue coating categories corresponding to the third image are input into the output layer for category conversion to obtain tongue coating category information corresponding to the third image. The tongue coating category information corresponding to the third image is used to determine the tongue coating category information corresponding to the second image.
[0112] Here, the second image is input to the input layer and normalized, for example, by normalizing the pixel values of the second image to the range [0, 1] or [-1, 1]. The second image is input to the input layer and cropped and scaled to a uniform size, for example, by adjusting the size of the second image to 224×224 pixels. Image enhancement may include one or more of the following: random rotation, flipping, brightness adjustment, and contrast adjustment.
[0113] It should be understood that a second image may correspond to one or more third images. In the case where a second image corresponds to multiple third images, the tongue coating category information corresponding to the second image is determined based on the tongue coating category information corresponding to each third image in the multiple third images corresponding to the second image. Determining the tongue coating category information corresponding to the second image based on the tongue coating category information corresponding to each third image in the multiple third images corresponding to the second image may include: when the tongue coating category information corresponding to each third image in the multiple third images corresponding to the second image is the same, determining the tongue coating category information corresponding to any third image corresponding to the second image as the tongue coating category information corresponding to the second image; or, when the tongue coating category information corresponding to a set number of third images corresponding to the second image is the same, determining the tongue coating category information corresponding to any one of the set number of third images as the tongue coating category information corresponding to the second image, without specific limitation.
[0114] Inputting the probabilities of each tongue coating category corresponding to the third image into the output layer for category conversion to obtain tongue coating category information corresponding to the third image can include: the output layer uses a normalized exponential (Softmax) function to perform category conversion on the probabilities of each tongue coating category corresponding to the third image to obtain tongue coating category information corresponding to the third image.
[0115] In this embodiment, clarifying that the first model includes an input layer, a first layer, and an output layer, and clarifying the content of processing the second image by each layer of the first model can improve the processing efficiency of the first model on the second image.
[0116] In one embodiment, the first layer includes a convolutional layer, a first residual layer, a pooling layer, a second residual layer, and a fully connected layer; inputting the third image into the first layer to perform tongue coating category probability prediction to obtain the probabilities of each tongue coating category corresponding to the third image includes:
[0117] Inputting the third image into the convolution layer to perform feature extraction to obtain first feature information of the third image;
[0118] Inputting the first feature information of the third image into the first residual layer for residual processing to obtain second feature information of the third image;
[0119] Inputting the first feature information of the third image into the pooling layer for feature fusion to obtain third feature information of the third image;
[0120] Inputting the third feature information and the second feature information of the third image into the second residual layer for residual processing to obtain fourth feature information of the third image;
[0121] The fourth feature information of the third image is input into the fully connected layer for linear transformation to obtain the probability of each tongue coating category corresponding to the third image.
[0122] Here, the pooling layer can be used to perform average pooling and maximum pooling. Inputting the first feature information of the third image into the pooling layer for feature fusion to obtain the third feature information of the third image may include:
[0123] Performing average pooling on the first feature information of the third image to obtain a first pooling result;
[0124] Performing maximum pooling on the first feature information of the third image to obtain a second pooling result;
[0125] The first pooling result and the second pooling result are fused to obtain third feature information of the third image.
[0126] Fusing the first pooling result and the second pooling result may include: summing the first pooling result and the second pooling result, or performing a weighted sum of the first pooling result and the second pooling result.
[0127] It's worth noting that max pooling preserves salient features in an image, while average pooling helps smooth features, reduce the effects of noise, and achieve a more stable feature representation. By setting different step sizes and pooling window sizes, the pooling layer can gradually reduce the resolution of the third image while increasing the dimensionality of features in the first and second pooling results.
[0128] It should be understood that the probabilities of each tongue coating category corresponding to the third image include the probabilities of each tongue coating category corresponding to the third image among all tongue coating categories. For example, if all tongue coating categories include normal tongue coating, damp-heat tongue coating, and blood stasis tongue coating, then the probabilities of each tongue coating category corresponding to the third image include the probabilities of normal tongue coating, damp-heat tongue coating, and blood stasis tongue coating corresponding to the third image.
[0129] In this embodiment, the structure and processing content of the first layer in the first model are clarified to help improve the processing efficiency of the first model. The third feature information and the second feature information of the third image are input into the second residual layer for residual processing, which can fuse the shallow detail information when extracting the deep features of the image and capture the comprehensive characteristics of the tongue coating.
[0130] In one embodiment, the convolution layer includes a first convolution kernel and / or a second convolution kernel, and the first feature information of the third image includes local feature information and / or global feature information of the third image; inputting the third image into the convolution layer for feature extraction to obtain the first feature information of the third image includes:
[0131] performing feature extraction on the third image using the first convolution kernel to obtain local feature information of the third image; and / or
[0132] Feature extraction is performed on the third image using the second convolution kernel to obtain global feature information of the third image, and the size of the second convolution kernel is larger than the size of the first convolution kernel.
[0133] For example, the size of the first convolution kernel may be 3×3, and the size of the second convolution kernel may be 5×5, which are not specifically limited.
[0134] In this embodiment, using a smaller first convolution kernel can more carefully extract the subtle bumps, grooves and specific morphological features on the surface of the tongue coating in the third image; using a larger second convolution kernel can more widely extract the overall color tendency of the tongue coating in the third image and the distribution of different color areas, thereby improving the accuracy of the first feature information of the third image.
[0135] In one embodiment, the first residual layer includes a first residual block, a second residual block, and a third residual block, and the second feature information of the third image includes fifth feature information, sixth feature information, seventh feature information, or eighth feature information of the third image; and inputting the first feature information of the third image into the first residual layer for performing residual processing to obtain the second feature information of the third image includes:
[0136] performing residual processing on the first feature information of the third image using the first residual block to obtain fifth feature information of the third image; or
[0137] In a case where the third image has multiple textures and / or colors interwoven, performing residual processing on the first feature information of the third image using the first residual block and the second residual block to obtain sixth feature information of the third image; or
[0138] In the case where there are flaws and / or defects in the third image, performing residual processing on the first feature information of the third image using the first residual block and the third residual block to obtain seventh feature information of the third image; or
[0139] When there are multiple textures and / or colors interwoven in the third image, and there are flaws and / or defects in the third image, residual processing is performed on the first feature information of the third image through the first residual block, the second residual block and the third residual block to obtain the eighth feature information of the third image.
[0140] Here, the second residual block is used to perform residual processing on the third image having multiple textures and / or colors interwoven, so as to extract features of the third image containing complex feature areas. The third image having multiple textures and / or colors interwoven may, for example, have multiple textures interwoven and / or uneven color distribution and / or a mixture of multiple colors.
[0141] The third residual block is used to perform residual processing on the third image having blemishes and / or defects to capture features in the image related to the specific semantic features. The third image having blemishes and / or defects may include, for example, cracks and / or bruises.
[0142] The sixth feature information of the third image includes features of the complex feature region of the third image and the fifth feature information of the third image.
[0143] The seventh feature information of the third image includes features related to the specific semantic feature in the third image and the fifth feature information of the third image.
[0144] The eighth feature information of the third image includes features of the complex feature area of the third image, features related to specific semantics in the third image, and the fifth feature information of the third image.
[0145] In this embodiment, the processing content of the first residual layer is clarified. The first residual layer includes residual blocks for complex feature areas and residual blocks related to specific semantic features, which can deepen the first model's extraction of these features and improve the comprehensiveness and accuracy of the second feature information of the third image.
[0146] The present application is described in further detail below with reference to application examples.
[0147] The image recognition method provided in the embodiment of the present application is applied to an attendance device, which may include an image acquisition module, a dual-channel feature fusion module, and an interaction module. The image acquisition module is used to obtain a first image, and may include a disinfection module such as an ultraviolet C (UV-C, Ultraviolet-C) lamp to meet hygiene requirements. The dual-channel feature fusion module is used to perform face recognition on the first image to obtain first information, and obtain second information based on the tongue coating contained in the first image. In offline mode, the first information and the second information are associated and stored locally, and in online mode, the first information and the second information can be synchronized to the cloud. The interaction module is used to guide the user to adjust his posture by voice to better obtain the first image and reduce the error rate of operation. It is also used to display the first information and the second information, and may include a 7-inch touch screen.
[0148] Taking the deployment of attendance equipment at the enterprise entrance and / or hospital triage desk as an example, the attendance equipment receives the user's wake-up voice, and prompts the user through voice "Please aim at the camera for tongue coating detection", obtains the first image, and obtains the first information and the second information; based on the first information and the second information, it generates a "Health Daily" and pushes it to the individual's mobile phone, which can give the user advice on adjusting his diet; the first information and the second information are associated and stored in a database. Human Resources (HR) can learn about the group's health trends by checking the database, which can be used to adjust work arrangements or organize health intervention activities.
[0149] The dual-channel feature fusion module is used to obtain the second information based on the tongue coating contained in the first image, including:
[0150] Determining a first image from the first image, where the first image represents an image of the first image at least including tongue coating;
[0151] Marking or extracting a region containing tongue coating in the first image to obtain a second image;
[0152] The second image is processed by the first model to obtain tongue coating category information corresponding to the second image; the first model is used to process the image containing the tongue coating and output the tongue coating category information.
[0153] Here, the first model is a ResNet-50 variant, and the structural diagram of the first model is as follows: Figure 4 shown. Figure 4 The first model includes an input layer, a convolutional layer, a first residual layer, a pooling layer, a second residual layer, a fully connected layer, and an output layer. The contents executed by each layer of the first model are described above and will not be repeated here. Figure 4 The fully connected layer outputs the probability of damp-heat tongue coating, normal tongue coating and blood stasis tongue coating. It should be understood that Figure 4 This is for illustrative purposes only and is not intended to be limiting.
[0154] The image recognition method provided in this application integrates face recognition and tongue coating analysis into the attendance equipment. Based on the brightness of the shooting environment, it determines or adjusts the photography parameters to solve the problem of cross-interference. It can also realize the closed loop of "attendance record-health monitoring-risk assessment". It adopts a modular design to be compatible with the existing attendance system, reduces the cost of health management upgrades, has the advantage of low-cost deployment, and expands the application scenarios of health management while improving attendance efficiency. It has high practicality.
[0155] In order to implement an image recognition method provided in an embodiment of the present application, an image recognition device is also provided in an embodiment of the present application, such as Figure 5 As shown, the device includes:
[0156] An acquisition module 501 is configured to acquire a first image; the first image may represent a face image including a tongue coating, or different images in the first image may include a face image or a tongue coating image;
[0157] a recognition module 502 configured to perform face recognition on the first image to obtain first information indicating user identity information;
[0158] The information acquisition module 503 is used to obtain second information based on the tongue coating contained in the first image, where the second information indicates the user's health status information.
[0159] In one embodiment, the second information includes tongue coating category information; the information obtaining module 503 is specifically configured to:
[0160] determining a first image from the first image, where the first image represents an image of the first image at least including tongue coating;
[0161] Marking or extracting a region containing tongue coating in the first image to obtain a second image;
[0162] The second image is processed by the first model to obtain tongue coating category information corresponding to the second image; the first model is used to process the image containing the tongue coating and output the tongue coating category information.
[0163] In one embodiment, the first model includes an input layer, a first layer, and an output layer; the information acquisition module 503 is specifically configured to:
[0164] Inputting the second image into the input layer and performing a first operation to obtain a third image; the first operation includes one or more of the following: normalization, cropping, scaling, and image enhancement;
[0165] Input the third image into the first layer to perform tongue coating category probability prediction, and obtain the probability of each tongue coating category corresponding to the third image;
[0166] The probabilities of the tongue coating categories corresponding to the third image are input into the output layer for category conversion to obtain tongue coating category information corresponding to the third image. The tongue coating category information corresponding to the third image is used to determine the tongue coating category information corresponding to the second image.
[0167] In one embodiment, the first layer includes a convolutional layer, a first residual layer, a pooling layer, a second residual layer, and a fully connected layer; the information acquisition module 503 is specifically configured to:
[0168] Inputting the third image into the convolution layer to perform feature extraction to obtain first feature information of the third image;
[0169] Inputting the first feature information of the third image into the first residual layer for residual processing to obtain second feature information of the third image;
[0170] Inputting the first feature information of the third image into the pooling layer for feature fusion to obtain third feature information of the third image;
[0171] Inputting the third feature information and the second feature information of the third image into the second residual layer for residual processing to obtain fourth feature information of the third image;
[0172] The fourth feature information of the third image is input into the fully connected layer for linear transformation to obtain the probability of each tongue coating category corresponding to the third image.
[0173] In one embodiment, the convolution layer includes a first convolution kernel and / or a second convolution kernel, and the first feature information of the third image includes local feature information and / or global feature information of the third image; the information acquisition module 503 is specifically configured to:
[0174] performing feature extraction on the third image using the first convolution kernel to obtain local feature information of the third image; and / or
[0175] Feature extraction is performed on the third image using the second convolution kernel to obtain global feature information of the third image, and the size of the second convolution kernel is larger than the size of the first convolution kernel.
[0176] In one embodiment, the first residual layer includes a first residual block, a second residual block, and a third residual block, and the second feature information of the third image includes fifth feature information, sixth feature information, seventh feature information, or eighth feature information of the third image; and the information obtaining module 503 is specifically configured to:
[0177] performing residual processing on the first feature information of the third image using the first residual block to obtain fifth feature information of the third image; or
[0178] In a case where the third image has multiple textures and / or colors interwoven, performing residual processing on the first feature information of the third image using the first residual block and the second residual block to obtain sixth feature information of the third image; or
[0179] In the case where there are flaws and / or defects in the third image, performing residual processing on the first feature information of the third image using the first residual block and the third residual block to obtain seventh feature information of the third image; or
[0180] When there are multiple textures and / or colors interwoven in the third image, and there are flaws and / or defects in the third image, residual processing is performed on the first feature information of the third image through the first residual block, the second residual block and the third residual block to obtain the eighth feature information of the third image.
[0181] In one embodiment, the acquisition module 501 is specifically configured to:
[0182] Determining or adjusting photography parameters based on the brightness of the shooting environment, the photography parameters including one or more of fill light brightness, sensitivity, and shutter speed;
[0183] Image acquisition is performed based on the photographic parameters to obtain the first image.
[0184] In actual application, the acquisition module 501, the recognition module 502 and the information acquisition module 503 can be implemented by a processor in the image recognition device.
[0185] It should be noted that the above embodiments provide an image recognition device for performing image recognition, using only the division of the above program modules as an example. In actual applications, the above processing can be assigned to different program modules as needed, that is, the internal structure of the device can be divided into different program modules to complete all or part of the above-described processing. In addition, the image recognition device and the image recognition method embodiments provided in the above embodiments are based on the same concept. The specific implementation process is detailed in the method embodiments and will not be repeated here.
[0186] Based on the hardware implementation of the above program modules, and in order to implement an image recognition method provided by an embodiment of the present application, an embodiment of the present application also provides an attendance device, such as Figure 6 As shown, the attendance device 600 includes:
[0187] Communication interface 601, capable of exchanging information with other network nodes;
[0188] The processor 602 is connected to the communication interface 601 to implement information exchange with other network nodes and is used to execute the method provided by one or more technical solutions when running a computer program. The computer program is stored in the memory 603.
[0189] Specifically, the processor 602 is used to obtain a first image; the first image represents a facial image containing tongue coating, or different images in the first image include facial images or tongue coating images; face recognition is performed on the first image to obtain first information, and the first information indicates the identity information of the user; based on the tongue coating contained in the first image, second information is obtained, and the second information indicates the health status information of the user.
[0190] In one embodiment, the second information includes tongue coating category information; the processor 602 is specifically configured to:
[0191] determining a first image from the first image, where the first image represents an image of the first image at least including tongue coating;
[0192] Marking or extracting a region containing tongue coating in the first image to obtain a second image;
[0193] The second image is processed by the first model to obtain tongue coating category information corresponding to the second image; the first model is used to process the image containing the tongue coating and output the tongue coating category information.
[0194] In one embodiment, the first model includes an input layer, a first layer, and an output layer; the processor 602 is specifically configured to:
[0195] Inputting the second image into the input layer and performing a first operation to obtain a third image; the first operation includes one or more of the following: normalization, cropping, scaling, and image enhancement;
[0196] Input the third image into the first layer to perform tongue coating category probability prediction, and obtain the probability of each tongue coating category corresponding to the third image;
[0197] The probabilities of the tongue coating categories corresponding to the third image are input into the output layer for category conversion to obtain tongue coating category information corresponding to the third image. The tongue coating category information corresponding to the third image is used to determine the tongue coating category information corresponding to the second image.
[0198] In one embodiment, the first layer includes a convolutional layer, a first residual layer, a pooling layer, a second residual layer, and a fully connected layer; the processor 602 is specifically configured to:
[0199] Inputting the third image into the convolution layer to perform feature extraction to obtain first feature information of the third image;
[0200] Inputting the first feature information of the third image into the first residual layer for residual processing to obtain second feature information of the third image;
[0201] Inputting the first feature information of the third image into the pooling layer for feature fusion to obtain third feature information of the third image;
[0202] Inputting the third feature information and the second feature information of the third image into the second residual layer for residual processing to obtain fourth feature information of the third image;
[0203] The fourth feature information of the third image is input into the fully connected layer for linear transformation to obtain the probability of each tongue coating category corresponding to the third image.
[0204] In one embodiment, the convolution layer includes a first convolution kernel and / or a second convolution kernel, and the first feature information of the third image includes local feature information and / or global feature information of the third image; the processor 602 is specifically configured to:
[0205] performing feature extraction on the third image using the first convolution kernel to obtain local feature information of the third image; and / or
[0206] Feature extraction is performed on the third image using the second convolution kernel to obtain global feature information of the third image, and the size of the second convolution kernel is larger than the size of the first convolution kernel.
[0207] In one embodiment, the first residual layer includes a first residual block, a second residual block, and a third residual block, and the second feature information of the third image includes fifth feature information, sixth feature information, seventh feature information, or eighth feature information of the third image; and the processor 602 is specifically configured to:
[0208] performing residual processing on the first feature information of the third image using the first residual block to obtain fifth feature information of the third image; or
[0209] In a case where the third image has multiple textures and / or colors interwoven, performing residual processing on the first feature information of the third image using the first residual block and the second residual block to obtain sixth feature information of the third image; or
[0210] In the case where there are flaws and / or defects in the third image, performing residual processing on the first feature information of the third image using the first residual block and the third residual block to obtain seventh feature information of the third image; or
[0211] When there are multiple textures and / or colors interwoven in the third image, and there are flaws and / or defects in the third image, residual processing is performed on the first feature information of the third image through the first residual block, the second residual block and the third residual block to obtain the eighth feature information of the third image.
[0212] In one embodiment, the processor 602 is specifically configured to:
[0213] Determining or adjusting photography parameters based on the brightness of the shooting environment, the photography parameters including one or more of fill light brightness, sensitivity, and shutter speed;
[0214] Image acquisition is performed based on the photographic parameters to obtain the first image.
[0215] It should be noted that the specific processing process of the processor 602 can be understood by referring to the above method.
[0216] Of course, in actual application, the various components in the attendance device 600 are coupled together through the bus system 604. It is understandable that the bus system 604 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 604 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, Figure 6 Various buses are labeled as bus system 604 .
[0217] The memory 603 in the embodiment of the present application is used to store various types of data to support the operation of the attendance device 600. Examples of such data include: any computer program used to operate on the attendance device 600.
[0218] The methods disclosed in the above embodiments of the present application can be applied to the processor 602 or implemented by the processor 602. The processor 602 may be an integrated circuit chip with signal processing capabilities. During implementation, the steps of the above methods can be completed by hardware integrated logic circuits in the processor 602 or instructions in software form. The above processor 602 may be a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 602 can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. A general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the methods disclosed in the embodiments of the present application can be directly embodied as being executed by a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium, which is located in the memory 603. The processor 602 reads the information in the memory 603 and completes the steps of the above methods in combination with its hardware.
[0219] In an exemplary embodiment, the attendance device 600 can be implemented by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers (MCUs), microprocessors, or other electronic components to perform the aforementioned method.
[0220] It is understood that the memory 603 of the embodiment of the present application can be a volatile memory or a non-volatile memory, and can also include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); the magnetic surface memory can be a magnetic disk memory or a tape memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), synchronous static random access memory (SSRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct rambus random access memory (DRRAM).The memory 603 described in the embodiments of the present application is intended to include but is not limited to these and any other suitable types of memories.
[0221] In an exemplary embodiment, the present application also provides an attendance device, including a processor and a memory for storing a computer program that can be run on the processor, wherein the processor is used to execute the steps of any of the above methods when running the computer program.
[0222] This embodiment of the present application further provides a storage medium, namely, a computer storage medium, specifically, a computer-readable storage medium, such as a memory 603 storing a computer program. The computer program can be executed by a processor 602 of the attendance device 600 to complete the steps of the aforementioned method. The computer-readable storage medium can be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, Flash Memory, magnetic surface mount storage, optical disk, or CD-ROM.
[0223] An embodiment of the present application further provides a computer program product, comprising a computer program, which implements the steps of any of the above methods when executed by a processor.
[0224] It should be noted that "first," "second," and the like are used to distinguish similar objects, and are not necessarily used to describe a specific order or precedence. The term "and / or" herein simply describes an association between related objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, and B exists alone. Furthermore, the term "one or more" herein refers to any combination of at least two of any one or more items. For example, "one or more items of A, B, and C" can represent any one or at least two or more items selected from the set consisting of A, B, and C.
[0225] In addition, the technical solutions described in the embodiments of the present application can be arbitrarily combined without conflict.
[0226] The above description is merely a preferred embodiment of the present application and is not intended to limit the scope of protection of the present application.
Claims
1. An image recognition method, characterized in that: Applied to attendance equipment, the method includes: Acquire a first image; the first image represents a face image including a tongue coating, or different images in the first image include a face image or a tongue coating image; Performing facial recognition on the first image to obtain first information, where the first information indicates identity information of the user; Based on the tongue coating contained in the first image, second information is obtained, where the second information indicates the user's health status information.
2. The method according to claim 1, characterized in that The second information includes tongue coating category information; obtaining the second information based on the tongue coating included in the first image includes: determining a first image from the first image, where the first image represents an image of the first image at least including tongue coating; Marking or extracting a region containing tongue coating in the first image to obtain a second image; The second image is processed by the first model to obtain tongue coating category information corresponding to the second image; the first model is used to process the image containing the tongue coating and output the tongue coating category information.
3. The method according to claim 2, characterized in that The first model includes an input layer, a first layer, and an output layer; and processing the second image by the first model to obtain tongue coating category information corresponding to the second image includes: Inputting the second image into the input layer and performing a first operation to obtain a third image; the first operation includes one or more of the following: normalization, cropping, scaling, and image enhancement; Input the third image into the first layer to perform tongue coating category probability prediction, and obtain the probability of each tongue coating category corresponding to the third image; The probabilities of the tongue coating categories corresponding to the third image are input into the output layer for category conversion to obtain tongue coating category information corresponding to the third image. The tongue coating category information corresponding to the third image is used to determine the tongue coating category information corresponding to the second image.
4. The method according to claim 3, characterized in that The first layer includes a convolutional layer, a first residual layer, a pooling layer, a second residual layer, and a fully connected layer; inputting the third image into the first layer to perform tongue coating category probability prediction to obtain the probability of each tongue coating category corresponding to the third image includes: Inputting the third image into the convolution layer to perform feature extraction to obtain first feature information of the third image; Inputting the first feature information of the third image into the first residual layer for residual processing to obtain second feature information of the third image; Inputting the first feature information of the third image into the pooling layer for feature fusion to obtain third feature information of the third image; Inputting the third feature information and the second feature information of the third image into the second residual layer for residual processing to obtain fourth feature information of the third image; The fourth feature information of the third image is input into the fully connected layer for linear transformation to obtain the probability of each tongue coating category corresponding to the third image.
5. The method according to claim 4, characterized in that The convolution layer includes a first convolution kernel and / or a second convolution kernel, and the first feature information of the third image includes local feature information and / or global feature information of the third image; Inputting the third image into the convolution layer for feature extraction to obtain first feature information of the third image includes: performing feature extraction on the third image using the first convolution kernel to obtain local feature information of the third image; and / or Feature extraction is performed on the third image using the second convolution kernel to obtain global feature information of the third image, and the size of the second convolution kernel is larger than the size of the first convolution kernel.
6. The method according to claim 4, characterized in that The first residual layer includes a first residual block, a second residual block, and a third residual block, and the second feature information of the third image includes fifth feature information, sixth feature information, seventh feature information, or eighth feature information of the third image; Inputting the first feature information of the third image into the first residual layer for residual processing to obtain the second feature information of the third image includes: performing residual processing on the first feature information of the third image using the first residual block to obtain fifth feature information of the third image; or In a case where the third image has multiple textures and / or colors interwoven, performing residual processing on the first feature information of the third image using the first residual block and the second residual block to obtain sixth feature information of the third image; or In the case where the third image has flaws and / or defects, performing residual processing on the first feature information of the third image using the first residual block and the third residual block to obtain seventh feature information of the third image; or When there are multiple textures and / or colors interwoven in the third image, and there are flaws and / or defects in the third image, residual processing is performed on the first feature information of the third image through the first residual block, the second residual block and the third residual block to obtain the eighth feature information of the third image.
7. The method according to any one of claims 1 to 6, characterized in that The acquiring of the first image includes: Determining or adjusting photography parameters based on the brightness of the shooting environment, the photography parameters including one or more of fill light brightness, sensitivity, and shutter speed; Image acquisition is performed based on the photographic parameters to obtain the first image.
8. An attendance device, characterized in that: comprising a processor and a memory for storing a computer program capable of being executed on the processor, Wherein, when the processor is used to run the computer program, it executes the steps of the method according to any one of claims 1 to 7.
9. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.