Intelligent make-up cabinet based on gating multi-path identification

The intelligent cosmetic cabinet based on gated multi-path recognition solves the problems of poor universality and recognition capability of cosmetics recognition in existing technologies, and realizes efficient and accurate recognition and intelligent management of cosmetic packaging, thereby improving the system's processing efficiency and user experience.

CN121312946APending Publication Date: 2026-01-13GUILIN UNIV OF ELECTRONIC TECH +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511191354.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-25
Publication Date
2026-01-13

AI Technical Summary

Technical Problem

Existing smart cosmetic cabinets lack versatility and recognition capabilities when identifying cosmetics, and cannot be compatible with unmarked cosmetics or cosmetics with missing labels, resulting in poor smart management performance.

Method used

The intelligent cosmetic cabinet adopts gated multi-path recognition. It captures cosmetic images through the image acquisition unit, and the central processing unit performs appearance category recognition and text region detection. Combining lightweight and computationally intensive models, it recognizes flat and curved packaging respectively. It uses a distortion correction model to correct the text on curved packaging. Combined with environmental perception and adjustment, it provides personalized services.

Benefits of technology

It enables efficient and accurate identification of various cosmetic packaging, improves system processing efficiency and response speed, provides an end-to-end intelligent cosmetic management experience, and enhances the system's robustness and versatility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121312946A_ABST
    Figure CN121312946A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent cosmetic cabinet based on gating multi-path identification, which is characterized in that a lightweight and efficient identification model is called for a simple plane package by a core gating classification mechanism, so that unnecessary calculation overhead is avoided; the calculation-intensive correction module is only called for the complex curved surface package, so that on-demand allocation of calculation resources is realized, and the overall processing efficiency and response speed of the system are greatly improved. In order to solve the problem of curved surface packaging, by introducing an innovative distortion correction model, a high-quality character expansion graph can be generated, geometric distortion of simple OCR recognition is eliminated fundamentally, and therefore the character recognition accuracy of curved surface packaging is improved. According to the invention, the core problem of low identification precision caused by appearance diversity of cosmetics is solved, and a complete smart home ecology can be seamlessly integrated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of smart home and artificial intelligence technology, specifically to a smart vanity cabinet based on door control multi-path recognition. Background Technology

[0002] With the improvement of living standards, consumers' demand for cosmetics (including skincare products) is increasing, and refined and intelligent management has become a new requirement. This includes not only the orderly storage of products, but also aspects such as monitoring the expiration date, suitable storage environment, and personalized usage suggestions. Traditional storage cabinets only provide basic storage functions and cannot achieve intelligent management. Although some intelligent cosmetic cabinets have attempted to introduce recognition technology, such as the "Portable Intelligent Cosmetic Cabinet Based on Image Processing" disclosed in Chinese invention patent application CN201710312192.3 and the "Method and Device for Cosmetic Storage Management, Cosmetic Cabinet" disclosed in Chinese invention patent application CN202110998452.3, they all rely on pre-attached QR codes or labels to the products. They are incompatible with unmarked cosmetics or cosmetics with missing labels, and their universality and recognition capabilities are poor, failing to achieve true intelligent management. Summary of the Invention

[0003] The present invention aims to solve the problems of poor versatility and recognition ability of existing intelligent makeup cabinets, and provides an intelligent makeup cabinet based on door control multi-path recognition.

[0004] To solve the above problems, the present invention is achieved through the following technical solution:

[0005] A smart cosmetic cabinet based on gate-controlled multi-path recognition includes a storage cabinet, a power supply unit, an image acquisition unit, a central processing unit, and a human-machine interaction unit. The power supply unit and the central processing unit are installed inside the storage cabinet. The image acquisition unit is installed at the storage entrance of the storage cabinet. The human-machine interaction unit is installed on the outer surface of the storage cabinet. The power supply unit is connected to the power supply terminals of the image acquisition unit, the central processing unit, and the human-machine interaction unit. The output terminal of the image acquisition unit is connected to the input terminal of the central processing unit, and the central processing unit is connected to the human-machine interaction unit.

[0006] When a user places cosmetics into the storage cabinet, the image acquisition unit automatically captures an image of the cosmetics and transmits the image to the central processing unit.

[0007] The central processing unit's appearance category recognition model classifies the input cosmetic images into either planar dominant or curved surface dominant types.

[0008] When the classification result is planar dominant type, the text region detection model of the central processing unit locates the text region in the cosmetic image and crops out normal text slice images; then the normal text slice images are sent to the text recognition engine of the central processing unit for content recognition to obtain cosmetic information.

[0009] When the classification result is surface-dominant, the text region detection model of the central processing unit locates the text region in the cosmetic image and crops out the distorted text slice image; then the distorted text slice image is sent to the distortion correction model of the central processing unit for correction processing to generate a normal text slice image; finally, the normal text slice image is sent to the text recognition engine of the central processing unit for content recognition to obtain cosmetic information.

[0010] The central processing unit performs structured fusion processing on the cosmetic information and then transmits it to the human-computer interaction unit for display.

[0011] In the above scheme, the distortion correction model consists of a spatial transformation network, a thin-plate spline transformation module, a depth feature extraction layer, a Transformer encoder, three upsampling layers, and four convolutional block attention modules. The spatial transformation network forms the input of the distortion correction model. The output of the spatial transformation network is connected to the input of the thin-plate spline transformation module. The output of the thin-plate spline transformation module is connected to the input of the depth feature extraction layer. The output of the depth feature extraction layer is connected to the input of the Transformer encoder. The output of the Transformer encoder is connected to the input of the first upsampling layer. The output of the first upsampling layer is connected to the input of the first convolutional attention module. The output of the first convolutional attention module is connected to the input of the second upsampling layer. The output of the second upsampling layer is connected to the input of the second convolutional attention module. The output of the second convolutional attention module is connected to the input of the third upsampling layer. The output of the third convolutional attention module is connected to the input of the fourth convolutional attention module. The output of the fourth convolutional block attention module forms the output of the distortion correction model.

[0012] In the above scheme, the Transformer encoder consists of a multi-head attention module, two additive layers, two normalization layers, and one multilayer perceptron. The input of the multi-head attention module forms the input of the Transformer encoder. The input and output of the multi-head attention module are simultaneously connected to the input of the first additive layer. The output of the first additive layer is connected to the input of the first normalization layer. The output of the first normalization layer is connected to the input of the multilayer perceptron. The input and output of the multilayer perceptron are simultaneously connected to the input of the second additive layer. The output of the second additive layer is connected to the input of the second normalization layer. The output of the second normalization layer forms the output of the Transformer encoder.

[0013] In the above scheme, the appearance category recognition model is MobileNetV3.

[0014] In the above scheme, the text region detection model is DBNet.

[0015] In the above scheme, the text recognition engine is PaddleOCR.

[0016] The above solution further includes an environmental sensing unit and an environmental control unit; the environmental sensing unit and the environmental control unit are installed inside the storage cabinet; the output end of the environmental sensing unit is connected to the input end of the central processing unit, and the output end of the central processing unit is connected to the environmental control unit.

[0017] Compared with the prior art, the present invention has the following characteristics:

[0018] 1. A perfect combination of intelligence and high efficiency: The core gating classification mechanism enables the system to teach according to individual needs. For simple planar packaging, a lightweight and efficient recognition model is used, avoiding unnecessary computational overhead; only for complex curved surface packaging, a computationally intensive correction module is used, realizing on-demand allocation of computing resources and greatly improving the overall processing efficiency and response speed of the system.

[0019] 2. High accuracy: Addressing the industry-recognized challenge of curved packaging, this invention introduces an innovative distortion correction model that generates high-quality text unfolding diagrams, fundamentally eliminating the geometric distortion of simple OCR recognition, thereby improving the accuracy of text recognition on curved packaging.

[0020] 3. Robustness and versatility: By utilizing an advanced text detection model, the system can accurately locate text of any shape and orientation. Combined with the high-performance PaddleOCR engine, the system has a high tolerance for various complex layouts, artistic fonts, and multilingual packaging, making it highly versatile.

[0021] 4. Integrated Smart Ecosystem Experience: This invention not only solves the core problem of low recognition accuracy caused by the diversity of cosmetic appearances, but also seamlessly integrates it into a complete smart home ecosystem. From high-precision recognition, intelligent storage, and environmental adaptive adjustment, to personalized APP recommendations and remote control, it provides users with an end-to-end, seamless intelligent cosmetic management experience. Attached Figure Description

[0022] Figure 1 This is a schematic diagram of a smart makeup cabinet based on gate control and multi-path recognition.

[0023] Figure 2 This is a flowchart illustrating the operation of an intelligent cosmetic cabinet based on gated multi-path recognition.

[0024] Figure 3 This is a schematic diagram of the principle of the appearance category recognition model.

[0025] Figure 4 This is a schematic diagram of the text region detection model.

[0026] Figure 5 This is a schematic diagram of the distortion correction model. Detailed Implementation

[0027] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific examples and the accompanying drawings.

[0028] A smart makeup cabinet based on gate control and multi-path recognition, such as Figure 1 As shown, the system includes a storage cabinet, a power supply unit, an image acquisition unit, a central processing unit, and a human-machine interface unit. The power supply unit is connected to the power supply terminals of the image acquisition unit, the central processing unit, and the human-machine interface unit. The output terminal of the image acquisition unit is connected to the input terminal of the central processing unit, and the central processing unit is connected to the human-machine interface unit.

[0029] The storage cabinet is a physical cabinet that includes an internal storage space for cosmetics, with an entrance for placing the cosmetics inside. A power supply unit provides stable power to the entire system. The power supply unit is installed inside the storage cabinet. In this embodiment, for fixed, large-volume storage cabinets, the power supply unit is a power source connected to mains power; for portable, small-volume storage cabinets, the power supply unit is a rechargeable battery, etc. An image acquisition unit is used to capture images of the cosmetics placed inside the storage cabinet. The image acquisition unit is installed at the storage entrance of the storage cabinet, facing the entrance. In this embodiment, the image acquisition unit is a wide-angle high-definition camera to ensure complete image capture of the items placed inside the cabinet. A central processing unit is installed inside the storage cabinet. In this embodiment, the central processing unit is a high-performance development board or embedded main controller, serving as the system core and responsible for running all algorithm models and system logic. A human-computer interaction unit is used to display information and receive user commands. The human-computer interaction unit is installed on the outer surface of the storage cabinet. In this embodiment, the human-computer interaction unit is a touch screen display, used to display real-time information about the items inside the cabinet, recognition results, temperature and humidity status, and personalized recommendations.

[0030] Based on this, the intelligent cosmetic cabinet may also include an environmental sensing unit and an environmental control unit. The environmental sensing unit and environmental control unit are installed inside the storage cabinet. The environmental sensing unit is used to sense the temperature and humidity information inside the storage cabinet. In this embodiment, the environmental sensing unit is a high-precision digital temperature and humidity sensor. The environmental control unit is used to regulate the temperature and humidity inside the storage cabinet to maintain a suitable storage environment to meet the stringent storage requirements of cosmetics (such as serums containing active ingredients). In this embodiment, the environmental sensing unit includes a semiconductor cooling chip, a small fan, and a dehumidification module. The input terminal of the environmental sensing unit is connected to the input terminal of the central processing unit, and the output terminal of the central processing unit is connected to the environmental control unit.

[0031] Cosmetic packaging comes in a wide variety of forms, but can be broadly categorized into two types: one is flat-dominated packaging (such as powder compacts and eyeshadow palettes), where the text on the surface is relatively neat; the other is curved-dominated packaging (such as lipstick tubes, serum bottles, and soft tubes), where the text on the surface is severely distorted and blurred due to perspective, uneven lighting, and geometric deformation. Applying existing Optical Character Recognition (OCR) technology indiscriminately to all packaging presents the following dilemmas: 1) The contradiction between universality and accuracy: Directly applying standard OCR technology to curved packaging results in extremely low accuracy, failing to meet application requirements; 2) The contradiction between efficiency and computing power: If complex distortion correction algorithms are used for all images, regardless of packaging form, processing flat packaging will result in significant computational waste and time delays, failing to meet the demands of a real-time, efficient user experience. To address the issues of low accuracy and efficiency of existing OCR technology when recognizing cosmetic packaging with diverse shapes (especially curved surfaces), this invention first intelligently distinguishes the physical shape of the cosmetic packaging, and then adaptively matches different recognition schemes to ensure high recognition accuracy while considering processing efficiency and energy consumption. Furthermore, it provides comprehensive intelligent management and personalized services. Figure 2 As shown, the operation process of this invention includes the following steps:

[0032] 1) When the user puts the cosmetics into the storage cabinet through the storage entrance, the image acquisition unit automatically captures the image of the cosmetics and transmits the image to the central processing unit.

[0033] 2) The central processing unit's appearance category recognition model classifies the input cosmetic images into either planar-dominated or curved-surface-dominated types. Based on the classification results, the central processing unit activates one of two parallel processing paths:

[0034] Processing Path 1: When the classification result is planar-dominated, the text region detection model of the central processing unit locates the text region in the cosmetic image and crops out normal text slice images; then, the normal text slice images are sent to the text recognition engine of the central processing unit for content recognition to obtain cosmetic information. This processing path is concise and has low computational overhead.

[0035] Processing Path Two: When the classification result is surface-dominant, the text region detection model of the central processing unit locates the text region in the cosmetic image and crops out distorted text slice images. These distorted text slice images are then fed into the distortion correction model of the central processing unit for correction processing, generating normal text slice images. Finally, these normal text slice images are fed into the text recognition engine of the central processing unit for content recognition to obtain the cosmetic information. Although this processing path involves a large amount of computation, it ensures high recognition accuracy.

[0036] 3) The central processing unit performs structured fusion processing on the cosmetic information and transmits the processed information to the human-computer interaction unit for display. The central processing unit performs structured fusion processing on the cosmetic information (such as brand, product name, model, etc.) and automatically manages key information such as the product's shelf life by combining it with the opening date initially entered by the user.

[0037] Based on this, the operation process of the present invention may further include the following steps:

[0038] 4) Environmental Data Acquisition and Control: The environmental sensing unit detects the temperature and humidity information inside the storage cabinet and transmits it back to the central processing unit. Based on the obtained cosmetic information, the central processing unit controls the environmental control unit to adjust the temperature and humidity inside the storage cabinet to maintain a suitable storage environment for the cosmetics.

[0039] 5) Personalized recommendations and remote interaction: The identified product information, user usage records, and environmental data collected by sensors are transmitted to the user's mobile phone by the central processing unit. The user's mobile phone application (APP) analyzes the user's skin type, usage habits, current weather, etc., to generate personalized skin care and makeup matching suggestions, product restocking reminders, and allows the user to remotely view the status of the cabinet and adjust the environment through the APP.

[0040] The gated multipath recognition module inside the central processing unit serves as the core recognition engine of the system, replacing the traditional single or simple recognition methods. This module receives images captured by the image acquisition unit and executes an intelligent, multipath recognition process, which mainly includes an appearance category recognition model, a text region detection model, a distortion correction model, and a character recognition engine.

[0041] The appearance category recognition model is used to classify input cosmetic images into either planar-dominated or curved-surface-dominated types. In this embodiment, the appearance category recognition model uses the lightweight convolutional neural network MobileNetV3. See also... Figure 3 The appearance category recognition model consists of two convolutional layers, two or more bottleneck modules, a global average pooling layer, and a fully connected layer. The input of the first convolutional layer forms the input to the appearance category recognition model. Two or more bottleneck modules are connected in series; the output of the first convolutional layer is connected to the input of the first bottleneck module, and the output of the last bottleneck module is connected to the input of the second convolutional layer. The output of the second convolutional layer is connected to the input of the global average pooling layer, and the output of the global average pooling layer is connected to the input of the fully connected layer. The output of the fully connected layer forms the output of the appearance category recognition model. MobileNetV3, due to its lightweight nature, is suitable for implementing fast classification on resource-constrained embedded devices, thereby optimizing the efficiency of gating decisions to achieve fast and accurate classification. The appearance category recognition model was trained using approximately 6,000 open-source, representative cosmetic packaging images, covering six common packaging appearance types (such as bottled, canned, disc-shaped, tubular, tube-shaped, and pen-shaped). Each sample was manually labeled with its dominant appearance type (i.e., planar or curved) to construct a training set with supervised information. The model then automatically extracted key appearance features (contour features, highlight distribution, and texture gradient changes) from the images using a deep neural network to achieve category (curved / planar) recognition.

[0042] A text region detection model is used to locate text regions in an image. In this embodiment, the text region detection model selected is the DBNet neural network model, which is capable of detecting irregularly shaped text. See also... Figure 4 The text region detection model consists of a pyramid feature extraction module, a multi-scale feature fusion module, a segmentation head prediction module, a differentiable binarization module, a near-vision binary map, and a post-processing module. The input to the pyramid feature extraction module forms the input to the text region detection model; the output of the pyramid feature extraction module is connected to the input of the multi-scale feature fusion module, the output of the multi-scale feature fusion module is connected to the input of the segmentation head prediction module, the output of the segmentation head prediction module is connected to the input of the differentiable binarization module, the output of the differentiable binarization module is connected to the input of the near-vision binary map, and the outputs of the differentiable binarization module and the near-vision binary map are connected to the input of the post-processing module. The output of the post-processing module forms the output of the text region detection model. DBNet can accurately locate text regions on cosmetics, and the output boundary coordinates are a set of polygon vertices describing the contour of the text region.

[0043] The distortion correction model is used to correct text images distorted by curved surfaces. It is activated only when recognizing curved surfaces and is used to flatten the distorted text image into a near-planar shape. In this embodiment, the distortion correction model is the Spatial Transformation-Attention Fusion Network STAF-Net. See also... Figure 5 The distortion correction model consists of a spatial transformation network, a thin-plate spline transformation module, a depth feature extraction layer, a Transformer encoder, three upsampling layers, and four convolutional block attention modules. The Transformer encoder comprises a multi-head attention module, two summation layers, two normalization layers, and one multilayer perceptron. The input of the multi-head attention module forms the input of the Transformer encoder. The input and output of the multi-head attention module are simultaneously connected to the input of the first summation layer. The output of the first summation layer is connected to the input of the first normalization layer. The output of the first normalization layer is connected to the input of the multilayer perceptron. The input and output of the multilayer perceptron are simultaneously connected to the input of the second summation layer. The output of the second summation layer is connected to the input of the second normalization layer. The output of the second normalization layer forms the output of the Transformer encoder. The spatial transformation network forms the input to the distortion correction model. The output of the spatial transformation network is connected to the input of the thin-plate spline transformation module. The output of the thin-plate spline transformation module is connected to the input of the deep feature extraction layer. The output of the deep feature extraction layer is connected to the input of the Transformer encoder. The output of the Transformer encoder is connected to the input of the first upsampling layer. The output of the first upsampling layer is connected to the input of the first convolutional attention module. The output of the first convolutional attention module is connected to the input of the second upsampling layer. The output of the second upsampling layer is connected to the input of the third convolutional attention module. The output of the third convolutional attention module is connected to the input of the fourth convolutional attention module. The output of the fourth convolutional attention module forms the output of the distortion correction model. STAF-Net generates a callable model path after training and detecting a large number of distorted image-planar image pairs. When a distorted text image is input again, after processing by STAF-Net, a corrected image with the text straightened and flattened will be generated.

[0044] STAF-Net is based on an encoder-decoder architecture; its unique features are: Front-end transform prediction: Before entering the backbone network, a flexible, non-rigid thin plate spline (TPS) transform parameter is predicted using a spatial transform network (STN) to perform preliminary global alignment of the input features. Global Feature Encoding: The encoder section innovatively introduces Transformer Encoding to capture the global dependencies of image features, which is crucial for understanding and reversing overall, coherent bending deformations. Attention Decoding: The decoder part reconstructs the corrected image step by step and in detail through multiple upsampling operations and in combination with the Convolutional Block Attention Module (CBAM) to ensure the clarity of key details such as text strokes.

[0045] The execution flow of STAF-Net is as follows:

[0046] Step 1: Global Deformation Prediction and Preliminary Alignment (STN+TPS)

[0047] The input receives a text slice image with geometric distortion, cropped by the text region detection module. The Spatial Transform Network (STN) does not directly modify the image but acts as a parameter predictor. It contains a localization network consisting of convolutional and fully connected layers. This localization network analyzes the entire input image and outputs a set of control point coordinates. Thin Plate Spline Transform (TPS) is a powerful non-rigid transformation model, well-suited for simulating smooth, continuous deformations in the physical world, such as the distortion caused by attaching a label to a curved bottle. The control points predicted by the STN serve as anchor points driving the TPS deformation. Through these anchor points, a smooth deformation field covering the entire image can be calculated. The purpose of this step is to perform a global, coarse-grained correction. It first macroscopically pulls the severely curved text to a roughly horizontal pose, providing a better initial state for subsequent fine-tuning.

[0048] Step 2: Deep Feature Extraction (Convolutional Layer)

[0049] The initially aligned image features are fed into one or more standard convolutional layers for deep feature extraction. This step aims to extract low- and mid-level visual features of the image, such as text edges, strokes, and corners, transforming the raw pixel information into feature maps that are easier for the network to understand.

[0050] Step 3: Global Context Encoding (Transformer Encoding)

[0051] The feature maps output from the convolutional layers are fed into a Transformer Encoding layer. Self-attention is the core of the Transformer. Unlike convolutional operations, which can only perceive information from neighboring regions, self-attention can compute the interdependencies between any two points in the feature map, regardless of their spatial distance. For a curved word, the deformation of the first letter is correlated with the deformation of the last letter. The Transformer, through its global receptive field, can capture this overall curvature pattern, rather than only seeing local distortions like traditional CNNs. This allows the network to understand the essence of the deformation from a global perspective, resulting in more coherent and consistent corrections.

[0052] Step 4: Attention-guided hierarchical decoding and reconstruction (UP + CBAM)

[0053] The deep features encoded by Transformer Encoding enter the decoder to begin reconstructing the image. This process is phased and progresses from coarse to fine. The decoder retains the upsampling layer and the convolutional block attention module. First, the features are passed to the upsampling layer. With each upsampling layer, the spatial resolution (height and width) of the feature map increases, gradually restoring it to the original image size. Then, the features are passed to the convolutional block attention module (CBAM). CBAM intelligently refines and optimizes features through two sub-modules: channel attention and spatial attention. The channel attention module analyzes each channel of the feature map and learns the importance weights of different channels. For example, in text correction tasks, feature channels related to character strokes are obviously more important than those related to background texture; channel attention automatically enhances the weight of stroke channels and suppresses the weight of background channels. The spatial attention module operates in the spatial dimension, analyzing each location on the feature map and learning the importance of different spatial regions. For example, the network learns to focus attention highly on regions containing character strokes while ignoring flat background regions. After iterative reconstruction through upsampling and convolutional block attention modules, a chain-repeating structure of UP-sampling Layer and CBAM is formed. The first upsampling reconstructs the general outline of the corrected image. Each subsequent iteration focuses on finer details based on the previous one, using the attention mechanism to make corrections, until a high-resolution, high-definition image is finally recovered.

[0054] Step 5: Final Feature Optimization and Output (CBAM + Output)

[0055] At the end of the decoder, a final CBAM module performs a final feature extraction. Finally, the output goes through a final convolutional layer to convert the optimized feature map into a three-channel (RGB) image. This output image is the final result after careful correction by the algorithm module, resulting in clear text and regular shape.

[0056] The text recognition engine is used to extract the final text information from (corrected or uncorrected) text slice images. In this embodiment, the text recognition engine is the high-performance optical character recognition engine PaddleOCR, which supports multilingual mixed recognition and ultimately accurately recognizes the text content.

[0057] It should be noted that although the embodiments described above are illustrative, they are not intended to limit the invention. Therefore, the invention is not limited to the specific embodiments described above. Any other embodiments obtained by those skilled in the art under the guidance of this invention without departing from its principles are considered to be within the protection scope of this invention.

Claims

1. A smart makeup cabinet based on gate-controlled multi-path recognition, characterized in that, It includes a storage cabinet, a power supply unit, an image acquisition unit, a central processing unit, and a human-machine interaction unit; the power supply unit and the central processing unit are installed inside the storage cabinet; the image acquisition unit is installed at the storage entrance of the storage cabinet; the human-machine interaction unit is installed on the outer surface of the storage cabinet; the power supply unit is connected to the power supply terminals of the image acquisition unit, the central processing unit, and the human-machine interaction unit; the output terminal of the image acquisition unit is connected to the input terminal of the central processing unit, and the central processing unit is connected to the human-machine interaction unit; When a user places cosmetics into the storage cabinet, the image acquisition unit automatically captures an image of the cosmetics and transmits the image to the central processing unit. The central processing unit's appearance category recognition model classifies the input cosmetic images into either planar dominant or curved surface dominant types. When the classification result is a plane-dominant type, the text region detection model of the central processing unit locates the text region in the cosmetic image and crops out the normal text slice image. The normal text slice image is then sent to the text recognition engine of the central processing unit for content recognition to obtain cosmetic information; When the classification result is surface-dominant, the text region detection model of the central processing unit locates the text region in the cosmetic image and crops out the distorted text slice image; then the distorted text slice image is sent to the distortion correction model of the central processing unit for correction processing to generate a normal text slice image; finally, the normal text slice image is sent to the text recognition engine of the central processing unit for content recognition to obtain cosmetic information. The central processing unit performs structured fusion processing on the cosmetic information and then transmits it to the human-computer interaction unit for display.

2. The intelligent cosmetic cabinet based on gate control multi-path recognition according to claim 1, characterized in that, The distortion correction model consists of a spatial transformation network, a thin-plate spline transformation module, a depth feature extraction layer, a Transformer encoder, three upsampling layers, and four convolutional block attention modules. The spatial transformation network forms the input to the distortion correction model. The output of the spatial transformation network is connected to the input of the thin-plate spline transformation module. The output of the thin-plate spline transformation module is connected to the input of the depth feature extraction layer. The output of the depth feature extraction layer is connected to the input of the Transformer encoder. The output of the Transformer encoder is connected to the input of the first upsampling layer. The output of the first upsampling layer is connected to the input of the first convolutional attention module. The output of the first convolutional attention module is connected to the input of the second upsampling layer. The output of the second upsampling layer is connected to the input of the third convolutional attention module. The output of the third convolutional attention module is connected to the input of the fourth convolutional attention module. The output of the fourth convolutional attention module forms the output of the distortion correction model.

3. The intelligent cosmetic cabinet based on gate control multi-path recognition according to claim 2, characterized in that, The Transformer encoder consists of a multi-head attention module, two summing layers, two normalization layers, and one multilayer perceptron. The input of the multi-head attention module forms the input of the Transformer encoder. The input and output of the multi-head attention module are simultaneously connected to the input of the first additive layer. The output of the first additive layer is connected to the input of the first normalization layer. The output of the first normalization layer is connected to the input of the multilayer perceptron. The input and output of the multilayer perceptron are simultaneously connected to the input of the second additive layer. The output of the second additive layer is connected to the input of the second normalization layer. The output of the second normalization layer forms the output of the Transformer encoder.

4. The intelligent cosmetic cabinet based on gate control multi-path recognition according to claim 1, characterized in that, The appearance category recognition model is MobileNetV3.

5. A smart makeup cabinet based on gate control multi-path recognition according to claim 1, characterized in that, The text region detection model is DBNet.

6. A smart makeup cabinet based on gate control multi-path recognition according to claim 1, characterized in that, The text recognition engine is PaddleOCR.

7. A smart makeup cabinet based on gate control multi-path recognition according to claim 1, characterized in that, It further includes an environmental sensing unit and an environmental control unit; the environmental sensing unit and the environmental control unit are installed inside the storage cabinet; the output end of the environmental sensing unit is connected to the input end of the central processing unit, and the output end of the central processing unit is connected to the environmental control unit.

Citation Information

Patent Citations

  • Portable intelligent cosmetic cabinet based on image processing

    CN107232798A

  • Cosmetic storage management method and device and cosmetic cabinet

    CN115732056A