A smart inventory management glasses device based on augmented reality technology
By integrating image acquisition, infrared positioning, voice interaction, and CNN technology, smart glasses devices address the shortcomings of existing devices in product recognition and management in complex environments, achieving efficient and accurate product recognition and environmental modeling, and improving user experience and device adaptability.
Patent Information
- Application Number
- CN202510447388.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-04-10
AI Technical Summary
Existing smart glasses devices lack adaptability to multi-tasking scenarios, have weak system integration capabilities, outdated interaction design, and insufficient hardware portability, making them unable to efficiently identify the type, quantity, and status of goods in complex and dynamic environments.
It integrates an image acquisition module, an infrared positioning module, a voice interaction module, and a convolutional neural network (CNN), combined with panoramic modeling capabilities, to achieve product recognition and environmental management, and supports voice control and independent operation without an external power supply.
It improves the efficiency and accuracy of product identification, enhances the visual management capabilities of the warehousing environment, increases user operating efficiency and equipment flexibility, and reduces usage costs.
Smart Images

Figure CN120494681B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of inventory management technology, and in particular relates to an intelligent inventory management glasses device based on augmented reality technology. Background Technology
[0002] Currently, augmented reality (AR) technology is increasingly being used in the field of smart glasses, with various smart glasses products being developed and applied to different industry scenarios. These products typically integrate AR technology, wireless communication modules such as 5G and UWB, cameras, and sensor modules to achieve data collection, positioning, monitoring, and remote collaboration in specific scenarios. For example, in industrial inspection, AR glasses are often used for equipment status monitoring, inspection path planning, and real-time alarms; in the logistics field, some equipment is attempting to combine sorting labels based on target object attributes to achieve item classification management; and in infrastructure monitoring, AR technology is used to acquire on-site images and combine them with virtual models to complete precise comparative analysis.
[0003] These products, with their strong scene adaptability and efficient data processing methods, improve the efficiency and security of traditional work processes. For example, CN115497189A discloses a reservoir inspection system based on 5G and UWB AR glasses. This system combines AR glasses with UWB base stations and uses a 5G communication module to map the location information of the inspection personnel with the 3D model data of the reservoir stored on the remote server, retrieve the model data of the inspection points, and complete the fusion of the virtual scene and the actual scene. It is mainly used for safety inspection and data comparison tasks in reservoirs; CN115567190A discloses a method, medium, and system for monitoring the training status of smoke-generating vehicles using AR glasses; this system collects status images of smoke-generating vehicles through AR glasses, analyzes anomalies, and generates training status diagnostic data, supporting collaborative diagnosis and communication between remote command centers and equipment, mainly used for real-time monitoring and collaboration of military training status; CN109949228A discloses an online calibration device and method for optically perceptible AR glasses, which solves the online calibration problem of virtual and real space alignment of AR glasses through the mapping of binocular vision cameras and virtual space, and is suitable for augmented reality navigation systems and related application scenarios; CN117499596A discloses a gas station inspection system and method based on smart AR glasses; this system integrates standardized inspection functions of gas equipment, uses AR glasses to receive alarm information, view equipment dynamic parameters and historical data, and cooperates with the back-end management platform to complete remote collaboration and management of inspection records, thereby improving the stability and safety of the gas industry; CN The sorting method, AR glasses, and system for multiple categories of target objects disclosed in CN115953635B solve the problem of low sorting efficiency for batch items by establishing a mapping relationship between target object attributes and sorting labels, acquiring target object images using AR glasses, and projecting sorting labels onto the display interface. The pipeline monitoring system and method based on AR and IoT disclosed in CN114910125B combine AR and IoT technologies, allowing real-time acquisition of pipeline monitoring data through air gestures and server data interaction for pipeline status analysis and fault location. The intelligent highway defect inspection system and method based on AR disclosed in CN117197412B uses AR glasses to collect defect photos, extract features, and analyze severity, providing positioning and early warning functions for intelligent monitoring and inspection of highway defects. These products utilize augmented reality technology, demonstrating strong functionality and technological innovation in inspection, monitoring, and sorting. Some products further enhance data interaction and analysis capabilities by combining 5G, IoT, and artificial intelligence technologies.
[0004] While existing products have achieved some technological innovation in specific scenarios such as inspection, monitoring, and sorting, the development and application of these technologies remain limited to the realization of single functions. The core problem lies in the lack of comprehensive adaptability to multi-task scenarios. Most current smart glasses devices are designed around single scenarios, such as path planning or data collection in inspection, or item recognition and labeling in sorting. The isolated existence of these functions makes the devices exhibit significant limitations when facing complex and ever-changing real-world environments. For example, in the warehousing and logistics field, dynamic goods are diverse, varied in quantity, and complexly stacked. Existing products typically rely on static image processing technology, which cannot accurately identify the type and condition of goods in dynamic scanning situations, nor can it count quantities or determine damage.
[0005] Furthermore, the weak system integration capabilities of existing technologies severely impact the widespread adaptability and scalability of the devices. Existing devices lag behind in interaction design, mostly relying on visual displays, resulting in cumbersome operation and failing to meet the requirements of efficiency and real-time performance. For example, the lack of voice interaction functionality prevents users from freeing their hands during actual operation, significantly limiting their work efficiency in dynamic and complex environments. Moreover, many similar products neglect the challenge of accurate recognition in complex dynamic scenarios during their technological implementation. Whether it's precise positioning of inspection points or efficient identification of product status, these devices fail to fully integrate infrared positioning technology with advanced machine learning algorithms to achieve higher recognition accuracy and processing speed. Current hardware designs tend towards modular and distributed configurations, leading to insufficient portability and requiring external power supplies or networks, which not only increases usage costs but also reduces flexibility in real-world scenarios. Therefore, this technical solution proposes an intelligent inventory management glasses device based on augmented reality technology. Summary of the Invention
[0006] This invention provides an intelligent inventory management glasses device based on augmented reality technology. By introducing a convolutional neural network (CNN) into product recognition and status analysis, the device is endowed with self-learning and dynamic recognition capabilities, enabling it to efficiently and accurately identify product type, quantity, and damage status in various complex environments. Particularly noteworthy is its dynamic scanning application in the warehousing and logistics field, providing a solution for task scenarios that traditional equipment cannot handle. To address the shortcomings of existing technologies in system integration capabilities, this invention combines infrared positioning technology with panoramic modeling functions, significantly improving the accuracy of product positioning and the visualization and management capabilities of scene data. By integrating a microcomputer into the smart glasses, this invention supports panoramic photography and modeling of the warehouse environment, and can... This invention stores and migrates data, achieving efficient integration with existing enterprise SaaS systems, thereby completely eliminating data silos and meeting the needs of modern enterprises for real-time and global data management. Regarding user interaction, the invention fully considers convenience and efficiency in actual operation, integrating a voice module to support voice command control and result broadcasting, greatly improving user efficiency and experience in dynamic and complex environments. Simultaneously, to address hardware portability and battery life issues, the invention highly integrates all core modules within the smart glasses, enabling independent operation without external power or network devices. This not only reduces the cost of using the device but also enhances its flexibility and practicality in multi-scenario applications. In summary, this invention solves the problems mentioned in the background technology.
[0007] To solve the above-mentioned technical problems, the present invention is achieved through the following technical solution:
[0008] The present invention provides an intelligent inventory management glasses device based on augmented reality technology, comprising AR glasses that integrate an image acquisition module, a panoramic modeling module, a system control and management module, an infrared positioning module, a product recognition module, a voice interaction module, a data storage and synchronization module, and a battery module.
[0009] The image acquisition module is used to capture image data of the product and its surrounding environment in real time. It acquires images through a high-resolution camera and transmits the acquired images to the product recognition module in the form of a high-speed data stream.
[0010] The infrared positioning module captures the three-dimensional position information of the goods in space by transmitting and receiving infrared signals;
[0011] The product recognition module uses a pre-trained convolutional neural network (CNN) model to analyze images, which can identify the type, quantity, and degree of damage of products. It also supports a dynamic learning mechanism. When an unrecognized product is encountered, the image is stored in the learning library, and the CNN model is updated through the background management system to improve the recognition coverage.
[0012] The data storage and synchronization module serves as a bridge connecting local storage and the back-end management system. After the product information is identified, it is stored in the local storage unit and uploaded to the back-end management platform in real time through the communication module. It also supports centralized management and long-term storage of warehouse data.
[0013] The voice interaction module enables human-computer interaction through a microphone and a speaker. The microphone is used to receive the user's voice commands and control the system to perform specific tasks. The speaker is used to broadcast the recognition results and task status information to the user in voice form.
[0014] The panoramic modeling module is responsible for 3D scanning and modeling of the warehouse environment. Using data from the image acquisition module and the infrared positioning module, it generates a virtual warehouse environment model through modeling algorithms, which is convenient for users to perform visual operations on the back-end management platform.
[0015] The system control and management module is responsible for coordinating the operating status of each module and is powered by the battery module. After the system starts up, the system control and management module initializes each hardware and algorithm module in sequence to ensure that the device enters the working state. During operation, the system control and management module schedules resources according to the progress of the task and handles abnormal situations.
[0016] Furthermore, the operation process of the AR glasses is as follows:
[0017] S1. System Startup and Initialization: When the user starts the AR glasses, the system control and management module starts first and initializes all hardware modules and software components; after all modules have successfully completed initialization, the system will enter normal working mode.
[0018] S2. Image Acquisition and Positioning: Once the system enters working mode, the image acquisition module begins to capture images of the products within the user's field of vision in real time. This module uses a high-resolution camera, which can stably capture images of the products and the background in dynamic environments. At the same time, the infrared positioning module captures the three-dimensional position information of the products by emitting and receiving infrared signals.
[0019] S3. Product Recognition and Analysis: Data acquired by the image acquisition module and the infrared positioning module is transmitted to the product recognition module in real time; the product recognition module analyzes the image data based on deep learning algorithms, namely convolutional neural networks (CNNs).
[0020] S4. Data Storage and Real-time Synchronization: The recognition results are processed through the data storage and synchronization module, which stores the results in the device's local storage unit. The data is then synchronized to the back-end management platform via the communication module. The synchronization process supports real-time performance and high efficiency, and can complete data upload and verification within seconds. The back-end management platform performs further analysis and processing based on the uploaded data to support warehouse management.
[0021] S5. Voice Broadcast and User Interaction: While storing the recognition results, the voice interaction module broadcasts the recognition information to the user in real time; the broadcast content includes the current product type, quantity, and status, as well as whether a rescan is needed; the user interacts with the device through voice commands.
[0022] S6. Panoramic Modeling and Environment Virtualization: After completing the commodity recognition task, the system calls the panoramic modeling module to model the warehouse environment. This module generates a three-dimensional model of the warehouse environment by integrating the data from the image acquisition module and the infrared positioning module.
[0023] S7. Learning Mechanism and Model Update: When the system fails to recognize a product or encounters a new product, the unrecognized product data will be stored in the learning library; the product recognition module will periodically synchronize the data in the learning library with the back-end management platform and update the weights and parameters of the CNN model through deep learning algorithms; the updated model will be automatically loaded into the device to improve the device's ability to recognize new products;
[0024] S8. Task Completion and System Shutdown: When all product recognition and modeling tasks are completed, the system notifies the user that the task has ended through the voice interaction module; the system control and management module coordinates all modules to enter standby mode and saves task logs for later query; if the user wishes to continue other tasks, the system will directly enter the new task mode without restarting.
[0025] Furthermore, step S1 specifically includes the following sub-steps:
[0026] S11. Power on and start up: After the user presses the power button, the main processor starts running and enters the initialization process;
[0027] S12. Hardware component self-test and initialization: The system detects the operating status of the image acquisition module, infrared positioning module, communication module, battery module, and voice interaction module. If any abnormalities are found, such as a module not responding or insufficient power supply, the system immediately reports the error and prompts the user to repair it.
[0028] S13. Load the pre-trained product recognition model: The main processor loads the product recognition model based on the convolutional neural network to prepare for subsequent scanning and analysis.
[0029] S14. Connect to the backend management platform: Establish a connection with the backend server through the communication module to ensure that data can be uploaded and synchronized in real time.
[0030] Furthermore, step S2 includes the following sub-steps:
[0031] S21. Start the image acquisition module to acquire images: The image acquisition module uses a camera module to start capturing images of the products in the current field of view in real time; at this time, the RGB image feature X is used. RGB Collected as initial input; X RGB It is a feature of the product image captured by the camera; it is usually a three-dimensional array representing the image's height (H), width (W), and color channels (C). RGB =3)
[0032] S22. Infrared positioning module assists in locating goods: The infrared positioning module captures the location information X of the goods. IR To reduce recognition errors caused by product obstruction or movement; infrared information is used to enhance the positioning accuracy of products in images; X IR It is the product location information obtained through infrared sensors, which is a single-channel image or heat map that shows the product's position in the infrared image;
[0033] Furthermore, step S3 includes the following sub-steps:
[0034] S31. Image input to CNN model for processing: The acquired product images are input into the convolutional neural network (CNN) model for processing through a multimodal fusion mechanism;
[0035] RGB image features X RGB and infrared image features X IR The fusion will be performed using an adaptive weighted fusion method; specifically, the fused feature map X fused The calculation process is as follows: β = 1 - α, where X RGB These are features derived from RGB images, representing the pixel information of the product within the RGB image; X IR Features derived from infrared images represent the location information of the goods; α and β are dynamically calculated weights that reflect the relative importance of RGB and infrared image information.
[0036] The fused feature map X fused Given by the following formula:
[0037] X fused =α·X RGB +β·X IR ;
[0038] S32. Judgment of Success or Failure in Product Recognition: In the process of product recognition, the determination of success or failure is based on the error between the predicted value output by the model and the actual label; specifically, a loss function is used to judge the performance of the model, and if the loss function value is lower than a set threshold, the recognition is considered successful.
[0039] Loss function L enhancedThis includes multimodal fusion feature errors and location information errors, as shown below:
[0040]
[0041] Where y i It's a real label. It is a predicted label; X ground truth These are real product image features; P(x,y) and These are the actual and predicted location information of the goods, respectively; λ1 and λ2 are regularization coefficients used to balance the influence of different error terms;
[0042] S33. Learning Process and Model Update: If recognition fails or the model's capabilities need to be expanded, the system stores images of unrecognized items in the learning library, and this data is processed through an reinforcement learning process. The system trains the CNN model using new data and updates the model parameters. After new image data is added to the learning library, incremental learning updates are performed, adjusting the model weights using the following formula:
[0043]
[0044] Where θ old These are the parameters of the current model, including weights and biases. These parameters are adjusted based on the gradient of the loss function each time the model is updated. The learning rate γ determines the step size for each parameter update.
[0045] S34. Prompt the user to perform the operation: If the learning and model update are completed, the system will prompt the user to continue scanning or re-identify the unidentified products.
[0046] Furthermore, the identification process in step S3 also includes the following:
[0047] Data preprocessing: The system normalizes the acquired image data, adjusts brightness and contrast, and eliminates noise to improve the accuracy of subsequent analysis;
[0048] Feature extraction: The CNN model extracts key feature points in the image and matches them with the training data;
[0049] Output results: Based on the model's output, the system generates a recognition report, including the type, quantity, and status of the goods; if recognition fails, the system will store the relevant data in the learning library for subsequent model updates.
[0050] Furthermore, the AR glasses perform inventory checks through the following steps:
[0051] P1. Storing inventory data: The system stores the identified product information locally and uploads it to the backend server via a 5G module to ensure data synchronization and security.
[0052] P2. Provide voice broadcast results: The voice module broadcasts the type, quantity, and damage information of the current product to the user, improving the user experience;
[0053] P3. Determine if the inventory count is complete: The system determines whether all inventory work has been completed based on the task list; if not, it prompts the user to continue; if completed, it proceeds to the modeling process.
[0054] P4. Generate panoramic modeling data: The system uses the panoramic modeling module to virtually model the warehouse environment and generate environmental data that can be used for subsequent analysis.
[0055] P5. Disconnect and end the task: After the task is completed, the system disconnects from the background, shuts down the device, and enters standby mode.
[0056] The present invention has the following advantages over the prior art:
[0057] (1) This invention significantly improves the efficiency and accuracy of product recognition by integrating a high-precision image acquisition module, an advanced infrared positioning module, and a deep learning-based product recognition module. Traditional product recognition technologies rely on static image processing and single visual input, which are easily affected by insufficient light, product occlusion, and dynamic environmental changes. This invention successfully solves these problems by combining infrared positioning and image recognition. In practical applications, even in warehouse scenarios with insufficient light or complex product stacking, the infrared module can provide accurate location information to assist the image recognition module in improving recognition accuracy.
[0058] (2) The commodity recognition module of this invention adopts a convolutional neural network (CNN), which can automatically extract the key features of commodities and match them with existing training data. This deep learning algorithm enables the system to not only efficiently recognize existing commodities, but also has the ability to learn dynamically. Commodities that are not recognized will be stored in the learning library and the model will be updated through the background management system to ensure that the system can adapt to the constantly changing commodity types in the warehousing environment. Compared with the fixed recognition model of the traditional system, the dynamic learning mechanism of this invention greatly expands the application scope of the equipment, reduces the necessity of manual intervention, and improves the intelligence level of the system.
[0059] (3) This invention achieves comprehensive virtualized management of the warehouse environment through a panoramic modeling module, significantly improving the intuitiveness and visualization capabilities of warehouse operations compared to existing technologies. In traditional warehouse management, environmental modeling typically relies on independent hardware and complex software processing, which is not only time-consuming and labor-intensive but also requires additional professional skills. In contrast, this invention achieves automated environmental modeling through deep integration of the panoramic modeling module with the image acquisition module and infrared positioning module. The system can combine the acquired warehouse images and location information to generate a high-precision three-dimensional environmental model. This model not only realistically reflects the distribution of goods but also provides real-time visualization through the backend management platform, offering users an intuitive interface. For example, users can adjust the layout of goods by dragging and dropping or quickly locate goods by querying product nodes in the model. Furthermore, this modeling technology supports the storage and backtracking of historical data, allowing users to view the warehouse status at any time, providing reliable data support for warehouse optimization and decision-making. Compared to existing environmental modeling methods that rely on independent equipment and manual operation, this invention has significant advantages in terms of speed, accuracy, and ease of use.
[0060] Of course, any product implementing this invention does not necessarily need to achieve all of the advantages described above at the same time. Attached Figure Description
[0061] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0062] Figure 1 This is a flowchart of the startup and initialization steps for the AR glasses of the present invention;
[0063] Figure 2 This is a flowchart illustrating the product scanning, recognition, and learning process of the AR glasses of this invention.
[0064] Figure 3 This is a flowchart illustrating the workflow of scanning product codes using the AR glasses of this invention.
[0065] Figure 4 This is a detailed extended description of the system operation process of the AR glasses of the present invention, in the form of a flowchart.
[0066] Figure 5 This is a diagram showing the module connection relationships of the AR glasses of the present invention;
[0067] Figure 6 This invention provides a schematic diagram of the structure of an intelligent inventory management glasses device based on augmented reality technology. Detailed Implementation
[0068] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0069] Please see Figure 1-6 As shown, this invention discloses an intelligent inventory management glasses device based on augmented reality technology, aiming to solve several key problems existing in current intelligent glasses technology. In particular, it addresses the limitations of existing products, such as limited functionality, weak system integration capabilities, insufficient dynamic environment recognition capabilities, and low user interaction efficiency, providing a comprehensive technical solution. Current intelligent glasses devices mainly focus on specific task designs, such as inspection, sorting, or monitoring, and cannot adapt to the complex needs of multi-task scenarios. This invention introduces convolutional neural networks (CNNs) into product recognition and state analysis, endowing the device with self-learning and dynamic recognition capabilities. This enables it to achieve efficient and accurate identification of product type, quantity, and damage status in various complex environments. Especially in dynamic scanning applications in the warehousing and logistics field, it provides a solution for task scenarios that traditional equipment cannot handle.
[0070] Furthermore, to address the shortcomings of existing technologies in system integration capabilities, this invention combines infrared positioning technology with panoramic modeling functionality, significantly improving the accuracy of product positioning and the visualization and management capabilities of scene data. By integrating a microcomputer into smart glasses, this invention supports panoramic photography and modeling of the warehouse environment, and can store and migrate data, achieving efficient integration with existing enterprise SaaS systems. This completely eliminates the problem of data silos and meets the needs of modern enterprises for real-time and comprehensive data management.
[0071] In terms of user interaction, this invention fully considers convenience and efficiency in actual operation, integrating a voice module to support voice command control and result broadcasting functions, greatly improving user efficiency and experience in dynamic and complex environments. Meanwhile, to address the issues of hardware portability and battery life, this invention highly integrates all core modules within the smart glasses, enabling independent operation without external power or network devices. This not only reduces the cost of using the device but also enhances its flexibility and practicality in various application scenarios.
[0072] The implementation method of AR glasses inventory management in this invention is as follows: Figure 1 , Figure 2 , Figure 3 As shown below, Figure 1 , Figure 2 and Figure 3 Detailed explanation:
[0073] like Figure 1 As shown, the startup and initialization steps for AR glasses are as follows:
[0074] 1. Power on and start the main processing module: After the user presses the power button, the main processor starts running and enters the initialization process.
[0075] 2. Hardware Component Self-Check and Initialization: The system checks the operating status of key hardware components such as the camera, infrared module, communication module, and voice module. If any abnormality is detected (such as a module not responding or insufficient power supply), the system immediately reports the error and prompts the user to fix it.
[0076] 3. Load the pre-trained product recognition model: The main processor loads the product recognition model based on the convolutional neural network to prepare for subsequent scanning and analysis.
[0077] 4. Connect to the backend management platform: Establish a connection with the backend server through the 5G communication module to ensure that data can be uploaded and synchronized in real time.
[0078] like Figure 2 As shown, the product scanning, recognition, and learning process is as follows:
[0079] 1. Camera image acquisition begins: The camera module starts capturing real-time images of the product within the current field of view. At this time, the RGB image feature X... RGB It was collected as the initial input. X RGB It is a feature of the product image captured by the camera; it is usually a three-dimensional array representing the image's height (H), width (W), and color channels (C). RGB =3).
[0080] 2. Infrared module-assisted product positioning: The infrared module captures the product's location information (X). IR This reduces recognition errors caused by product obstruction or movement. Infrared information is used to enhance the positioning accuracy of products in images. IR It is the product location information obtained through infrared sensors, usually a single-channel image or heat map, which shows the position of the product in the infrared image.
[0081] 3. Image input to CNN model for processing: The acquired product images are input into the convolutional neural network (CNN) model for processing through a multimodal fusion mechanism.
[0082] RGB image features X RGB and infrared image features X IR The fusion will be performed using an adaptive weighted fusion method. Specifically, the fused feature map X... fusedThe calculation process is as follows: β = 1 - α, where X RGB These are features derived from RGB images, representing the pixel information of the product within the RGB image. X IR These are features derived from infrared images, representing the location information of the product. α and β are dynamically calculated weights, reflecting the relative importance of RGB and infrared image information.
[0083] The fused feature map X fused Given by the following formula:
[0084] X fused =α·X RGB +β·X IR
[0085] 4. Determining success or failure:
[0086] Successful Recognition Determination: In the product recognition process, success is determined based on the error between the model's predicted value and the actual label. Specifically, we use a loss function to evaluate the model's performance. If the loss function value is below a set threshold, the recognition is considered successful.
[0087] Loss function L enhanced This includes multiple aspects such as multimodal fusion feature error and location information error, as shown below:
[0088]
[0089] Where y i It's a real label. This is a predicted label. X ground truth These are the features of a real product image. P(x,y) and These represent the actual and predicted location information of the goods, respectively. λ1 and λ2 are regularization coefficients used to balance the influence of different error terms.
[0090] 5. Learning process and model updates:
[0091] If recognition fails or the model's capabilities need to be expanded, the system stores images of unrecognized items in a learning database, and this data can be processed through reinforcement learning. The system trains the CNN model using the new data, updating the model parameters. Specifically, after new image data is added to the learning database, incremental learning updates are performed, adjusting the model weights using the following formula:
[0092]
[0093] Where θ oldThese are the parameters (weights and biases) of the current model. These parameters are adjusted based on the gradient of the loss function during each model update, and the learning rate γ determines the step size for each parameter update.
[0094] 6. Prompt the user to perform actions: If the learning and model update are complete, the system will prompt the user to continue scanning or re-identify unidentified products.
[0095] like Figure 3 As shown, the product scanning, recognition, and learning process is as follows:
[0096] 1. Storing inventory data: The system stores the identified product information locally and uploads it to the backend server via a 5G module to ensure data synchronization and security.
[0097] 2. Provide voice broadcast results: The voice module broadcasts the type, quantity, and damage information of the current product to the user, improving the user experience.
[0098] 3. Determine if inventory count is complete: The system checks the task list to determine if all inventory work is complete. If not, the user is prompted to continue; if complete, the system proceeds to the modeling process.
[0099] 4. Generate panoramic modeling data: The system uses the panoramic modeling module to virtually model the warehouse environment and generate environmental data that can be used for subsequent analysis.
[0100] 5. Disconnect and end the task: After the task is completed, the system disconnects from the background, shuts down the device, and enters standby mode.
[0101] like Figure 4 As shown, the AR glasses system operation process of this invention is divided into multiple stages, covering all steps from device startup to task completion. The modules coordinate their operation to achieve efficient and accurate product recognition, environmental modeling, and data management. The following is a detailed extended description of the system operation process:
[0102] 1. System startup and initialization
[0103] When a user activates the AR glasses, the system control and management module starts first, initializing all hardware modules and software components. The initialization process includes self-tests of hardware devices such as the image acquisition module, infrared positioning module, voice interaction module, and product recognition module, as well as functional tests of the data storage and synchronization module and the panoramic modeling module. If any module fails to respond properly, the system will display the corresponding error message to the user via the voice interaction module and suggest repairs. After all modules have successfully initialized, the system will enter normal operating mode.
[0104] 2. Image Acquisition and Positioning
[0105] Once the system enters working mode, the image acquisition module begins capturing real-time images of the products within the user's field of view. This module utilizes a high-resolution camera, enabling stable capture of both product and background images in dynamic environments. Simultaneously, the infrared positioning module captures the three-dimensional position information of the products by emitting and receiving infrared signals. Infrared positioning is particularly suitable for situations with insufficient light or partially obscured products, providing additional auxiliary information to the product recognition module and improving overall recognition accuracy.
[0106] 3. Product Identification and Analysis
[0107] Data acquired by the image acquisition module and infrared positioning module is transmitted to the product recognition module in real time. The product recognition module analyzes the image data based on deep learning algorithms (convolutional neural network, CNN). The recognition process includes the following steps:
[0108] Data preprocessing: The system normalizes the acquired image data, adjusts brightness and contrast, and eliminates noise to improve the accuracy of subsequent analysis.
[0109] Feature extraction: The CNN model extracts key feature points in the image and matches them with the training data.
[0110] Output Results: Based on the model's output, the system generates a recognition report, including the type, quantity, and status of the goods (e.g., whether they are damaged). If recognition fails, the system stores the relevant data in the learning library for subsequent model updates.
[0111] 4. Data storage and real-time synchronization
[0112] The identification results are processed through a data storage and synchronization module. This module first stores the results in the device's local storage unit, and then synchronizes the data to the backend management platform via a 5G communication module. The synchronization process supports real-time performance and high efficiency, completing data upload and verification within seconds. The backend management platform then performs further analysis and processing based on the uploaded data, providing support for warehouse management.
[0113] 5. Voice broadcasting and user interaction
[0114] While storing the recognition results, the voice interaction module broadcasts the recognition information to the user in real time. The broadcast includes the current product type, quantity, and status, as well as whether a rescan is needed. Users can also interact with the device via voice commands, such as requesting inventory progress or modifying task objectives. The design of the voice interaction module greatly reduces the complexity of user operations, making the device suitable for dynamic and high-intensity work environments.
[0115] 6. Panoramic Modeling and Environment Virtualization
[0116] After completing the product identification task, the system invokes the panoramic modeling module to model the warehouse environment. This module integrates data from the image acquisition module and the infrared positioning module to generate a 3D model of the warehouse environment. These models not only visually reflect the current product distribution but also serve as historical data for subsequent analysis and optimization. For example, users can view a virtual panoramic view of the warehouse through the backend management system and adjust the warehouse layout.
[0117] 7. Learning Mechanism and Model Update
[0118] When the system fails to recognize a product or encounters a new product, the unrecognized product data is stored in the learning library. The product recognition module periodically synchronizes the data in the learning library with the backend management platform, updating the weights and parameters of the CNN model using deep learning algorithms. The updated model is automatically loaded into the device, improving its ability to recognize new products. This dynamic learning mechanism ensures that the device can continuously adapt to new demands as the warehousing environment changes.
[0119] 8. Task completion and system shutdown
[0120] Once all product recognition and modeling tasks are completed, the system will notify the user via the voice interaction module that the task has ended. Subsequently, the system control and management module will coordinate all modules to enter standby mode and save the task logs for later retrieval. If the user wishes to continue with other tasks, the system can directly enter a new task mode without restarting.
[0121] like Figure 5 As shown, the AR smart glasses system of this invention includes an image acquisition module, an infrared positioning module, a product recognition module, a data storage and synchronization module, a voice interaction module, a panoramic modeling module, and a system control and management module. These modules, through deep integration of hardware and software, collaboratively complete the tasks of product recognition, warehouse environment modeling, and backend data interaction. The connection and positional relationships between the components are as follows: Figure 6 As shown.
[0122] The system's main hardware components include a camera, infrared sensor, microphone, speaker, and computing chip mounted on the eyeglass frame, as well as a communication module that enables high-speed data communication via a 5G network. The software integrates a convolutional neural network (CNN) algorithm for product recognition, a real-time infrared positioning algorithm, a panoramic modeling algorithm, and a backend data synchronization protocol.
[0123] The image acquisition module, the front-end module of this system, is installed in the center of the glasses and is used to capture image data of the product and its surrounding environment in real time. This module acquires images through a high-resolution camera and transmits them to the product recognition module as a high-speed data stream. The image acquisition module supports shooting in dynamic scenes and can achieve stable shooting even when the product is moving rapidly.
[0124] The infrared positioning module captures the three-dimensional position information of a product in space by emitting and receiving infrared signals. This module significantly improves recognition accuracy in low-light conditions or when the product is obscured, providing positional information for the product recognition module.
[0125] The product recognition module is the core algorithm module of the system, employing a pre-trained convolutional neural network (CNN) model to analyze images. The module can identify the type, quantity, and damage status of products, and supports a dynamic learning mechanism. When an unrecognized product is encountered, the product recognition module stores the image in a learning database and updates the CNN model through the backend management system, improving the recognition coverage.
[0126] The data storage and synchronization module serves as a bridge connecting local storage and the back-end management system. Once the product information is identified, it is stored in the local storage unit and uploaded to the back-end management platform in real time via the 5G communication module, supporting centralized management and long-term storage of warehouse data.
[0127] The voice interaction module enables human-computer interaction through a microphone and a speaker. The microphone receives the user's voice commands and controls the system to execute specific tasks; the speaker then broadcasts the recognition results, task status, and other information to the user in voice format. The voice interaction module significantly improves the system's ease of use and operational efficiency.
[0128] The panoramic modeling module is responsible for 3D scanning and modeling of the warehouse environment. This module utilizes data from the image acquisition and infrared positioning modules to generate a virtualized warehouse environment model through modeling algorithms, facilitating visualization operations by users on the backend management platform.
[0129] The system control and management module is the core management unit of the device, responsible for coordinating the operational status of each module. After system startup, the control module initializes each hardware and algorithm module sequentially to ensure the device enters a working state. During operation, the control module schedules resources according to task progress and handles abnormal situations, such as triggering learning processes or adjusting communication parameters.
[0130] The AR smart glasses system of this invention demonstrates significant technical advantages in the fields of product recognition and warehouse management. Compared with existing technologies, its beneficial effects are mainly reflected in two aspects: efficient product recognition and comprehensive environmental modeling.
[0131] First, this invention significantly improves the efficiency and accuracy of product recognition by integrating a high-precision image acquisition module, an advanced infrared positioning module, and a deep learning-based product recognition module. Traditional product recognition technologies rely heavily on static image processing and single visual input, making them susceptible to insufficient lighting, product occlusion, and dynamic environmental changes. This invention successfully solves these problems through the joint processing of infrared positioning and image recognition. In practical applications, even in warehouse scenarios with insufficient lighting or complex product stacking, the infrared module provides accurate location information, assisting the image recognition module in improving recognition accuracy. Furthermore, the product recognition module of this invention employs a convolutional neural network (CNN), which automatically extracts key product features and matches them with existing training data. This deep learning algorithm enables the system not only to efficiently recognize existing products but also to possess dynamic learning capabilities. Unrecognized products are stored in a learning library, and the model is updated through a backend management system, ensuring the system can adapt to the ever-changing product types in the warehouse environment. Compared to the fixed recognition model of traditional systems, the dynamic learning mechanism of this invention greatly expands the application scope of the equipment, reduces the need for manual intervention, and improves the system's intelligence level.
[0132] Secondly, this invention achieves comprehensive virtualized management of the warehouse environment through a panoramic modeling module, significantly improving the intuitiveness and visualization capabilities of warehouse operations compared to existing technologies. In traditional warehouse management, environmental modeling typically relies on independent hardware and complex software processes, which is not only time-consuming and labor-intensive but also requires additional specialized skills. This invention, however, achieves automated environmental modeling through deep integration of the panoramic modeling module with image acquisition and infrared positioning modules. The system combines acquired warehouse images and location information to generate a high-precision 3D environmental model. This model not only realistically reflects the distribution of goods but also provides real-time visualization through a backend management platform, offering users an intuitive interface. For example, users can adjust the product layout by dragging and dropping or quickly locate product nodes by querying them in the model. Furthermore, this modeling technology supports the storage and review of historical data, allowing users to view the warehouse status at any point in time, providing reliable data support for warehouse optimization and decision-making. Compared to existing environmental modeling methods that rely on independent equipment and manual operation, this invention offers significant advantages in speed, accuracy, and ease of use.
[0133] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. The invention is limited only by the claims and their full scope and equivalents.
Claims
1. A smart inventory management glasses device based on augmented reality technology, characterized in that, AR glasses, including those that integrate an image acquisition module, a panoramic modeling module, a system control and management module, an infrared positioning module, a product recognition module, a voice interaction module, a data storage and synchronization module, and a battery module. The image acquisition module is used to capture image data of the product and its surrounding environment in real time. It acquires images through a high-resolution camera and transmits the acquired images to the product recognition module in the form of a high-speed data stream. The infrared positioning module captures the three-dimensional position information of the goods in space by transmitting and receiving infrared signals; The product recognition module uses a pre-trained convolutional neural network (CNN) model to analyze images, which can identify the type, quantity, and degree of damage of products. It also supports a dynamic learning mechanism. When an unrecognized product is encountered, the image is stored in the learning library, and the CNN model is updated through the background management system to improve the recognition coverage. Image processing via CNN model: The acquired product images are fed into a convolutional neural network (CNN) model for processing through a multimodal fusion mechanism; RGB image features X RGB and infrared image features X IR The fusion will be performed using an adaptive weighted fusion method; specifically, the fused feature map X fused The calculation process is as follows: β = 1 - α, where X RGB These are features derived from RGB images, representing the pixel information of the product within the RGB image; X IR Features derived from infrared images represent the location information of the goods; α and β are dynamically calculated weights that reflect the relative importance of RGB and infrared image information. The fused feature map X fused Given by the following formula: X fused =α·X RGB +β·X IR ; Success or failure determination: In the product recognition process, success or failure is determined based on the error between the predicted value output by the model and the actual label; specifically, a loss function is used to judge the performance of the model, and if the loss function value is lower than a set threshold, the recognition is considered successful. Loss function L enhanced This includes multimodal fusion feature errors and location information errors, as shown below: Where y i It's a real label. It is a predicted label; X ground truth These are real product image features; P(x,y) and These are the actual and predicted location information of the goods, respectively; λ1 and λ2 are regularization coefficients used to balance the influence of different error terms; Learning Process and Model Updates: If recognition fails or the model's capabilities need to be expanded, the system stores images of unrecognized items in the learning library, and this data is processed through reinforcement learning. The system trains the CNN model using new data and updates the model parameters. After new image data is added to the learning library, incremental learning updates are performed, adjusting the model weights using the following formula: Where θ old These are the parameters of the current model, including weights and biases. These parameters are adjusted based on the gradient of the loss function each time the model is updated. The learning rate γ determines the step size for each parameter update. User prompts: If the learning and model update are complete, the system will prompt the user to continue scanning or re-identify unidentified items; The data storage and synchronization module serves as a bridge connecting local storage and the back-end management system. After the product information is identified, it is stored in the local storage unit and uploaded to the back-end management platform in real time through the communication module. It also supports centralized management and long-term storage of warehouse data. The voice interaction module enables human-computer interaction through a microphone and a speaker. The microphone is used to receive the user's voice commands and control the system to perform specific tasks. The speaker is used to broadcast the recognition results and task status information to the user in voice form. The panoramic modeling module is responsible for 3D scanning and modeling of the warehouse environment. Using data from the image acquisition module and the infrared positioning module, it generates a virtual warehouse environment model through modeling algorithms, which is convenient for users to perform visual operations on the back-end management platform. The system control and management module is responsible for coordinating the operating status of each module and is powered by the battery module. After the system starts up, the system control and management module initializes each hardware and algorithm module in sequence to ensure that the device enters the working state. During operation, the system control and management module schedules resources according to the progress of the task and handles abnormal situations.
2. The intelligent inventory management glasses device based on augmented reality technology according to claim 1, characterized in that, The operation process of the AR glasses is as follows: S1. System Startup and Initialization: When the user starts the AR glasses, the system control and management module starts first and initializes all hardware modules and software components; after all modules have successfully completed initialization, the system will enter normal working mode. S2. Image Acquisition and Positioning: Once the system enters working mode, the image acquisition module begins to capture images of the products within the user's field of vision in real time. This module uses a high-resolution camera, which can stably capture images of the products and the background in dynamic environments. At the same time, the infrared positioning module captures the three-dimensional position information of the products by emitting and receiving infrared signals. S3. Product Recognition and Analysis: Data acquired by the image acquisition module and the infrared positioning module is transmitted to the product recognition module in real time; the product recognition module analyzes the image data based on deep learning algorithms, namely convolutional neural networks (CNNs). S4. Data storage and real-time synchronization: The recognition results are processed through the data storage and synchronization module, which stores the results in the device's local storage unit and then synchronizes the data to the back-end management platform through the communication module. The synchronization process supports real-time performance and high efficiency, enabling data upload and verification to be completed within seconds; the back-end management platform performs further analysis and processing based on the uploaded data, providing support for warehouse management. S5. Voice Broadcast and User Interaction: While storing the recognition results, the voice interaction module broadcasts the recognition information to the user in real time; the broadcast content includes the current product type, quantity, and status, as well as whether a rescan is needed; the user interacts with the device through voice commands. S6. Panoramic Modeling and Environment Virtualization: After completing the commodity recognition task, the system calls the panoramic modeling module to model the warehouse environment. This module generates a three-dimensional model of the warehouse environment by integrating the data from the image acquisition module and the infrared positioning module. S7. Learning Mechanism and Model Update: When the system fails to recognize a product or encounters a new product, the unrecognized product data will be stored in the learning library; the product recognition module will periodically synchronize the data in the learning library with the back-end management platform and update the weights and parameters of the CNN model through deep learning algorithms; the updated model will be automatically loaded into the device to improve the device's ability to recognize new products; S8. Task Completion and System Shutdown: When all product recognition and modeling tasks are completed, the system notifies the user that the task has ended through the voice interaction module; the system control and management module coordinates all modules to enter standby mode and saves task logs for later query; if the user wishes to continue other tasks, the system will directly enter the new task mode without restarting.
3. The intelligent inventory management glasses device based on augmented reality technology according to claim 2, characterized in that, Step S1 specifically includes the following sub-steps: S11. Power on and start up: After the user presses the power button, the main processor starts running and enters the initialization process; S12. Hardware component self-test and initialization: The system detects the operating status of the image acquisition module, infrared positioning module, communication module, battery module, and voice interaction module. If any abnormalities are found, such as a module not responding or insufficient power supply, the system immediately reports the error and prompts the user to repair it. S13. Load the pre-trained product recognition model: The main processor loads the product recognition model based on the convolutional neural network to prepare for subsequent scanning and analysis. S14. Connect to the backend management platform: Establish a connection with the backend server through the communication module to ensure that data can be uploaded and synchronized in real time.
4. The intelligent inventory management glasses device based on augmented reality technology according to claim 2, characterized in that, S2 includes the following steps: S21. Start the image acquisition module to acquire images: The image acquisition module uses a camera module to start capturing images of the products in the current field of view in real time; at this time, the RGB image feature X is used. RGB Collected as initial input; X RGB It is a three-dimensional array representing the product image features captured by the camera, showing the image's height (H), width (W), and color channels (C). RGB =3); S22. Infrared positioning module assists in locating goods: The infrared positioning module captures the location information X of the goods. IR To reduce recognition errors caused by product obstruction or movement; infrared information is used to enhance the positioning accuracy of products in images; X IR It is the product location information obtained through infrared sensors. It is a single-channel image or heat map that shows the position of the product in the infrared image.
5. The intelligent inventory management glasses device based on augmented reality technology according to claim 2, characterized in that, The identification process in step S3 also includes the following: Data preprocessing: The system normalizes the acquired image data, adjusts brightness and contrast, and eliminates noise to improve the accuracy of subsequent analysis; Feature extraction: The CNN model extracts key feature points in the image and matches them with the training data; Output results: Based on the model's output, the system generates a recognition report, including the type, quantity, and status of the goods; if recognition fails, the system will store the relevant data in the learning library for subsequent model updates.
6. The intelligent inventory management glasses device based on augmented reality technology according to claim 1, characterized in that, The AR glasses undergo inventory checks through the following steps: P1. Storing inventory data: The system stores the identified product information locally and uploads it to the backend server via a 5G module to ensure data synchronization and security. P2. Provide voice broadcast results: The voice module broadcasts the type, quantity, and damage information of the current product to the user, improving the user experience; P3. Determine if the inventory count is complete: The system checks the task list to determine if all inventory work has been completed; if not, prompt the user to continue. If completed, proceed to the modeling process; P4. Generate panoramic modeling data: The system uses the panoramic modeling module to virtually model the warehouse environment and generate environmental data that can be used for subsequent analysis. P5. Disconnect and end the task: After the task is completed, the system disconnects from the background, shuts down the device, and enters standby mode.
Citation Information
Patent Citations
An optical perspective AR glasses online calibration device and method
CN109949228A
A pipeline monitoring system and method based on AR and IoT
CN114910125B
Reservoir inspection system based on AR glasses of 5G and UWB
CN115497189A
Method, medium and system for monitoring training state of fuming vehicle by adopting AR (Augmented Reality) glasses
CN115567190A
A sorting method for multiple categories of target objects, AR glasses and system
CN115953635B