Intelligent compact shelving system and method based on internet of things technology
By combining monitoring video and audio signals to extract features in the mobile shelving system and using a classifier to control the lighting, the problem of power waste in the mobile shelving system has been solved, achieving intelligent lighting management and energy-saving effects.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-23
- Publication Date
- 2026-03-24
AI Technical Summary
The existing mobile shelving system lacks lighting equipment, which means that the lights need to be kept on when searching for files, resulting in wasted electricity and economic losses.
By extracting features from surveillance video and audio signals, a classifier is used to control smart lighting. The tracking feature map of the mobile shelving unit and the sound association feature map are combined to perform a high-dimensional spatial unit manifold sub-dimensional hyperconvex correlation measurement, resulting in a smart lighting control classification feature map. The classifier is then used to determine whether the lighting should be turned on.
It enables intelligent lighting control, avoids power waste, improves the efficiency and convenience of file management, and provides a personalized lighting experience and energy-saving effect.
Smart Images

Figure CN118114145B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of Internet of Things (IoT) technology, and more specifically, to an intelligent mobile shelving system and method based on IoT technology. Background Technology
[0002] Intelligent mobile shelving (mobile cabinets) is a smart network mobile shelving system that integrates manual, electric, and computer control, enabling remote operation. Each shelf panel is equipped with control buttons or a touchscreen display. When an administrator needs to open any shelf, simply pressing the open button will automatically open the shelf. Most conveniently, the intelligent mobile shelving system has intelligent software installed on its computer. When files are stored, a file management database is created in the computer. During subsequent management, simply entering the file to be retrieved in the computer management interface will automatically open the corresponding mobile shelving unit.
[0003] In recent years, the archival work has made great strides, and its scale has been expanding day by day. The requirements for archival management have also become more and more stringent. However, the existing mobile shelving units do not have lighting equipment. When using them, the lights in the archive room where the mobile shelving units are stored must be kept on all the time, or the lights in the archive room must be turned on in the dark to find the documents needed. This will result in a waste of electricity and cause economic losses to the company.
[0004] Therefore, a smart mobile shelving system and method based on Internet of Things (IoT) technology is desired. Summary of the Invention
[0005] To address the aforementioned technical problems, this application is proposed. Embodiments of this application provide an intelligent mobile shelving system and method based on Internet of Things (IoT) technology. This system extracts features from monitoring video and audio signals and utilizes a classifier to control intelligent lighting, thereby improving file management efficiency and saving energy.
[0006] Accordingly, according to one aspect of this application, an intelligent mobile shelving system based on Internet of Things (IoT) technology is provided, comprising:
[0007] The mobile shelving data acquisition module is used to acquire monitoring video and audio signals from the intelligent mobile shelving.
[0008] The mobile shelving data extraction module is used to extract the mobile shelving tracking feature map from the monitoring video and the mobile shelving sound association feature map from the sound signal;
[0009] The mobile shelving data fusion module is used to perform high-dimensional spatial unit manifold sub-dimensional hyperconvex correlation measurement on the mobile shelving tracking feature map and the mobile shelving sound association feature map to obtain a smart lighting control classification feature map.
[0010] The mobile shelving data analysis module is used to pass the intelligent lighting control classification feature map through a classifier to obtain a classification result, which is used to indicate whether the lighting is turned on.
[0011] According to another aspect of this application, a method for intelligent mobile shelving based on Internet of Things (IoT) technology is provided, comprising:
[0012] Acquire surveillance video and audio signals from intelligent mobile shelving units;
[0013] Extract the mobile shelving tracking feature map from the surveillance video, and extract the mobile shelving sound association feature map from the sound signal;
[0014] The tracking feature map and the sound association feature map of the mobile shelving are subjected to high-dimensional spatial unit manifold sub-dimensional hyperconvex correlation measurement to obtain the intelligent lighting control classification feature map;
[0015] The intelligent lighting control classification feature map is passed through a classifier to obtain a classification result, which is used to indicate whether the lighting is turned on.
[0016] Compared with existing technologies, this application provides an intelligent mobile shelving system and method based on Internet of Things (IoT) technology. This system acquires monitoring video and audio signals from within the intelligent mobile shelving unit, extracts tracking feature maps and sound association feature maps from these signals, and then fuses these feature maps to obtain an intelligent lighting control classification feature map. A classifier then classifies this feature map to determine whether the lighting is on or off. This enables intelligent lighting control, avoids energy waste and economic losses, and improves the efficiency and convenience of file management. Attached Figure Description
[0017] The above and other objects, features, and advantages of this application will become more apparent from the more detailed description of the embodiments of this application in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this application and form part of the specification. They are used together with the embodiments of this application to explain this application and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.
[0018] Figure 1 This is a block diagram of an intelligent mobile shelving system based on Internet of Things (IoT) technology according to an embodiment of this application.
[0019] Figure 2 This is a block diagram of the mobile shelving data extraction module in an intelligent mobile shelving system based on Internet of Things technology according to an embodiment of this application.
[0020] Figure 3This is a block diagram of a video processing unit in an intelligent mobile shelving system based on Internet of Things technology according to an embodiment of this application.
[0021] Figure 4 This is a block diagram of the sound signal processing unit in an intelligent mobile shelving system based on Internet of Things technology according to an embodiment of this application.
[0022] Figure 5 This is a flowchart of an intelligent mobile shelving method based on Internet of Things (IoT) technology according to an embodiment of this application. Detailed Implementation
[0023] Various exemplary embodiments, features, and aspects of this application will now be described in detail with reference to the accompanying drawings. The same reference numerals in the drawings denote elements that have the same or similar functions. Although various aspects of the embodiments are shown in the drawings, they are not necessarily drawn to scale unless specifically indicated otherwise.
[0024] The term “exemplary” as used herein means “serving as an example, embodiment, or illustration.” Any embodiment illustrated herein as “exemplary” is not necessarily to be construed as superior to or better than other embodiments.
[0025] Furthermore, to better illustrate this application, numerous specific details are provided in the following detailed description. Those skilled in the art should understand that this application can be implemented without certain specific details. In some instances, methods, means, components, and circuits well-known to those skilled in the art have not been described in detail in order to highlight the main points of this application.
[0026] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.
[0027] Figure 1 The diagram illustrates a block diagram of an intelligent mobile shelving system based on Internet of Things (IoT) technology according to an embodiment of this application. Figure 1As shown, the intelligent mobile shelving system 100 based on Internet of Things technology according to an embodiment of this application includes: a mobile shelving data acquisition module 110, used to acquire monitoring video and sound signals from the intelligent mobile shelving; a mobile shelving data extraction module 120, used to extract a mobile shelving tracking feature map from the monitoring video and a mobile shelving sound association feature map from the sound signals; a mobile shelving data fusion module 130, used to perform a high-dimensional space unit manifold sub-dimensional hyperconvex correlation measurement on the mobile shelving tracking feature map and the mobile shelving sound association feature map to obtain an intelligent lighting control classification feature map; and a mobile shelving data analysis module 140, used to pass the intelligent lighting control classification feature map through a classifier to obtain a classification result, the classification result being used to indicate whether the lighting is turned on.
[0028] In this embodiment, the mobile shelving data acquisition module 110 is used to acquire monitoring video and audio signals from the intelligent mobile shelving. It should be understood that monitoring video can acquire the internal status information of the mobile shelving because video provides visual information, clearly showing the storage status of files, the operation process, etc. Monitoring video can be obtained by installing a camera inside the mobile shelving, capturing the location and quantity of files and personnel operations in real time. By analyzing and processing the video, relevant status information can be acquired. Acquiring audio signals allows monitoring the operational status of the mobile shelving because it generates specific audio signals during operation. Through sensing changes in sound and monitoring video, intelligent lighting control can be achieved. When there is no person or activity, the system can automatically turn off the lighting, thereby saving energy and reducing energy consumption. This helps reduce energy costs, carbon emissions, and is more environmentally friendly. Through the intelligent lighting system, automated lighting control can be achieved. When someone or activity is detected, the system can automatically turn on the lighting, providing the required brightness and comfort. This eliminates the need for manual operation and provides a more convenient and intelligent lighting experience. Through sensing changes in sound and monitoring video, intelligent security and safety control can be achieved. When unusual sounds or suspicious activity are detected, the system can automatically activate lighting to increase environmental visibility and provide better security. This helps prevent crime and provides a sense of security. By sensing changes in sound and monitoring video, potential problems or dangerous situations can be identified promptly. For example, when surveillance video shows someone moving near a dangerous area, the system can automatically activate lighting to alert personnel. This helps prevent losses and accidents, protecting the safety of people and property. By acquiring surveillance video and audio signals, comprehensive monitoring and management of the intelligent mobile shelving can be achieved, improving the efficiency and security of file management. Such monitoring and management can be conducted remotely, without time and space limitations. Managers can remotely access the monitoring system to observe the status of files in real time, promptly identify and address problems, thereby improving the efficiency and security of file management. Specifically, monitoring equipment such as cameras and microphones are installed inside the intelligent mobile shelving. These devices can be connected to a network or a dedicated monitoring system to transmit video and audio signals in real time.
[0029] In this embodiment, the mobile shelving data extraction module 120 is used to extract mobile shelving tracking feature maps from the monitoring video and mobile shelving sound association feature maps from the sound signals. It should be understood that by extracting mobile shelving tracking feature maps from the monitoring video, the storage status of files and personnel operations can be monitored in real time. These feature maps can provide information about the location and quantity of files and personnel activities. By analyzing the tracking feature maps, it can be determined whether personnel are active near the mobile shelving, thereby determining whether the intelligent lighting system needs to be activated or deactivated. Extracting mobile shelving sound association feature maps from the sound signals can help monitor the activity and abnormal situations of the mobile shelving. For example, by analyzing the sound signals, the opening and closing sounds of the mobile shelving doors, impact sounds, etc., can be detected. These feature maps can be used to determine whether the mobile shelving is being used or if an abnormal situation has occurred, thereby deciding whether the intelligent lighting system needs to be activated or deactivated. By extracting mobile shelving tracking feature maps and sound association feature maps, monitoring video and sound signals can be converted into a more informative data form. These feature maps can be used in conjunction with the control algorithms of the intelligent lighting system to achieve more intelligent and precise lighting control. By combining information from archival activity and sound changes, it is possible to more accurately determine whether to turn the intelligent lighting system on or off, thereby improving the efficiency and energy saving of the lighting system.
[0030] Specifically, in one embodiment of this application, Figure 2 The diagram illustrates a block diagram of a mobile shelving data extraction module in an IoT-based intelligent mobile shelving system according to an embodiment of this application. Figure 2 As shown, in the above-mentioned intelligent mobile shelving system 100 based on Internet of Things technology, the mobile shelving data extraction module 120 includes: a video processing unit 121, used to extract key frames from the monitoring video and then obtain the mobile shelving tracking feature map through target detection; and an audio signal processing unit 122, used to perform noise reduction processing on the audio signal and then obtain the mobile shelving audio association feature map through convolutional coding.
[0031] Accordingly, in a specific example of this application, the video processing unit 121 is used to extract keyframes from the surveillance video and then perform target detection to obtain the mobile shelving tracking feature map. It should be understood that surveillance videos are typically recorded in frames per second, and each frame contains a large amount of information. To reduce computation and processing time, extracting keyframes is a common strategy. Keyframes are frames in the video that contain important information or significant changes; selecting keyframes for processing can reduce computation while preserving key information. Mobile shelving refers to shelves for storing files or items. Target detection algorithms can automatically identify and locate targets in images or videos. Target detection technology can identify mobile shelving in surveillance videos and generate results containing the location and bounding boxes of the mobile shelving. By performing target detection from keyframes, the location information of the mobile shelving in each keyframe can be obtained. Integrating this location information into the feature map can generate a mobile shelving tracking feature map. The mobile shelving tracking feature map can show the distribution and changes of the mobile shelving, providing information about the location of files and personnel activities. By extracting keyframes and performing target detection, the location information of mobile shelving units can be obtained from surveillance video and converted into tracking feature maps. These feature maps can be used to analyze and determine personnel activities and file storage conditions, thereby enabling the control of intelligent lighting systems. By combining mobile shelving unit tracking feature maps with intelligent lighting control algorithms, intelligent lighting control can be achieved, improving the efficiency and energy saving of lighting systems.
[0032] further, Figure 3 The diagram illustrates a block diagram of a video processing unit in an intelligent mobile shelving system based on Internet of Things (IoT) technology, according to an embodiment of this application. Figure 3 As shown, in the mobile shelving data extraction module 120 of the intelligent mobile shelving system 100 based on the Internet of Things technology, the video processing unit 121 includes: a keyframe extraction subunit 1211, used to extract multiple mobile shelving keyframes from the monitoring video; a target detection subunit 1212, used to obtain multiple target object region of interest maps by passing the multiple mobile shelving keyframes through a target detection network based on an anchorless window; and a three-dimensional tensor unit 1213, used to arrange the multiple target object region of interest maps into a three-dimensional tensor and pass it through a three-dimensional convolutional neural network model to obtain the mobile shelving tracking feature map.
[0033] Specifically, the keyframe extraction subunit 1211 is used to extract multiple keyframes of the mobile shelving unit from the surveillance video. It should be understood that the mobile shelving unit may change position in the surveillance video. Extracting multiple keyframes can capture these changes, such as the movement, addition, or deletion of shelves. This is crucial for monitoring and managing changes in the mobile shelving unit's position. Important events may occur in the surveillance video, such as items being taken or placed on the mobile shelving unit, or people moving around the mobile shelving unit. By extracting multiple keyframes, these important events can be captured and further analyzed and processed. Surveillance videos typically contain a large number of frames; for long-term video surveillance, storing all frames can consume significant storage space. Extracting multiple keyframes reduces storage requirements while retaining key information and providing sufficient data for subsequent analysis. Target detection and analysis in surveillance videos typically require substantial computational resources and time. Extracting multiple keyframes reduces the number of frames that need to be processed, thereby improving processing efficiency. By extracting multiple keyframes of the mobile shelving unit from the surveillance video, changes in the mobile shelving unit's position and the occurrence of important events can be captured, while simultaneously reducing storage requirements and improving processing efficiency. This is very valuable for realizing intelligent monitoring and management of mobile shelving systems.
[0034] Accordingly, the keyframe extraction subunit includes: an initial frame extraction subunit for extracting an initial frame from the surveillance video; a differential image frame calculation subunit for calculating positional differences between other frames in the surveillance video and the initial frame to obtain differential image frames; and a keyframe determination subunit for determining the plurality of mobile shelving keyframes from the surveillance video based on a comparison between the sum of feature values of all pixels in the differential image frames and a predetermined threshold.
[0035] Furthermore, the second-level subunit for determining keyframes is used to: determine the third-level subunit for determining first keyframes, which is used to determine other frames corresponding to the differential feature map as first keyframes based on the sum of the feature values of all pixels in the differential image frame being greater than a predetermined threshold; and set the third-level subunit for setting initial frames, which sets the first keyframe as the initial frame.
[0036] Furthermore, the target detection subunit 1212 is used to obtain multiple Region of Interest (ROI) maps for the multiple mobile shelving keyframes by passing them through an anchor-free target detection network. It should be understood that the target detection network can detect the position and bounding box of the mobile shelving in each keyframe. This helps to identify the accurate position and size of the mobile shelving for subsequent analysis and processing. Anchor-free target detection networks can directly predict the bounding box of the target without using predefined anchor boxes. By inputting the keyframes into such a network, the Region of Interest (ROI) map of the mobile shelving in each keyframe can be obtained. These ROI maps can provide local image information of the mobile shelving, which is helpful for further tracking and analysis. The obtained ROI maps of the mobile shelving can be used for target tracking and recognition. Tracking algorithms can use features in the ROI maps to track the movement of the mobile shelving in the video sequence. Recognition algorithms can perform target recognition in the ROI maps, such as determining whether there are items on the mobile shelving or classifying items. Specifically, the target detection subunit includes: passing each of the multiple mobile shelving keyframes through multiple convolutional layers to obtain multiple shallow feature maps; and passing each of the multiple shallow feature maps through the anchor-window-based target detection network to obtain multiple target object region of interest feature maps. By inputting the multiple mobile shelving keyframes into the anchor-window-based target detection network, the region of interest map of the mobile shelving in each keyframe can be obtained, providing a foundation for subsequent target tracking and recognition. This enables more accurate analysis and processing of the mobile shelving, providing richer information and functions for intelligent monitoring and management of mobile shelving systems.
[0037] Accordingly, in a specific example of this application, the sound signal processing unit 122 is used to perform noise reduction processing on the sound signal and then obtain the sound association feature map of the mobile shelving unit through convolutional coding. It should be understood that in a monitoring scenario, the sound signal may be interfered with by environmental noise, such as background noise and wind noise. By performing noise reduction processing on the sound signal, noise components can be removed, improving the signal quality and clarity, thereby better capturing the sound information related to the mobile shelving unit. Convolutional coding is an effective feature extraction method that can capture the local features and spatial correlations of a signal through convolutional operations. By inputting the noise-reduced sound signal into a convolutional coding network, abstract features of the mobile shelving unit's sound signal can be extracted, including spectral features and temporal features. These features can reflect the correlation between the mobile shelving unit and the sound, which is helpful for subsequent analysis and processing. Through convolutional coding, the sound signal can be converted into a feature map, i.e., the sound association feature map of the mobile shelving unit. This feature map can provide information on the distribution of sound in the time and frequency domains, as well as features related to the location of the mobile shelving unit. By analyzing this feature map, we can understand the sound environment around the mobile shelving unit, such as whether there are abnormal sounds and the intensity of the sounds, thus enabling more comprehensive monitoring and management of the mobile shelving unit. By performing noise reduction processing on the sound signal and obtaining the sound association feature map of the mobile shelving unit through convolutional coding, we can extract the correlation information between the sound and the mobile shelving unit, providing more data and functions for the intelligent analysis and processing of the mobile shelving unit system.
[0038] further, Figure 4 The diagram illustrates a block diagram of the sound signal processing unit in an intelligent mobile shelving system based on Internet of Things (IoT) technology according to an embodiment of this application. Figure 4 As shown, in the mobile shelving data extraction module 120 of the aforementioned IoT-based intelligent mobile shelving system 100, the sound signal processing unit 122 includes: a noise reduction processing subunit 1221, used to perform noise reduction processing on the sound signal to obtain a noise-reduced sound signal; a sound signal feature extraction subunit 1222, used to extract the log-Mel spectrum and cochlear spectrum of the noise-reduced sound signal based on the mobile shelving; and a first convolutional coding subunit 1223, used to pass the log-Mel spectrum through a first convolutional neural network as a filter to obtain the intelligent mobile shelving sound. The system comprises: a Mel spectrogram feature vector; a second convolutional encoding subunit 1224, used to pass the cochlear spectrogram through a second convolutional neural network as a filter to obtain a cochlear spectrogram feature vector for the intelligent mobile shelving unit; a fusion vector subunit 1225, used to fuse the Mel spectrogram feature vector and the cochlear spectrogram feature vector for the intelligent mobile shelving unit to obtain a sound association feature matrix for the mobile shelving unit; and a third convolutional encoding subunit 1226, used to pass the sound association feature matrix for the mobile shelving unit through a third convolutional neural network model as a feature extractor to obtain a sound association feature map for the mobile shelving unit.
[0039] Accordingly, the noise reduction processing subunit 1221 is used to perform noise reduction processing on the sound signal to obtain a noise-reduced sound signal. It should be understood that sound signals in the real world are often accompanied by various environmental noises, such as background noise, traffic noise, and wind noise. These noises interfere with the clarity of the original sound signal, reducing its audibility and recognizability. Noise reduction processing can effectively remove these environmental noises, making the sound signal cleaner and easier to analyze. Signal-to-noise ratio (SNR) refers to the ratio of useful signal to noise signal in a sound signal. When noise is high, the SNR is low, leading to a decrease in the quality of the sound signal. Noise reduction processing can reduce the amplitude of the noise signal, improve the SNR, and make the sound signal clearer and more discernible. The noise-reduced sound signal can better reflect the characteristics and information of the original sound. This is very important for the analysis and processing of sound signals, such as in applications like speech recognition, sound classification, and speech synthesis. Noise reduction processing makes the sound signal easier for computer algorithms to parse and understand, improving the accuracy and effectiveness of subsequent processing algorithms. Noise reduction processing of audio signals can remove environmental noise, improve the signal-to-noise ratio, and enhance the quality and clarity of the audio signal. This provides a better foundation for subsequent audio analysis and processing, improving system performance and reliability.
[0040] Specifically, the sound signal feature extraction subunit 1222 is used to extract the logarithmic Mel spectrum and cochlear spectrogram of the denoised sound signal based on a compact shelving system. It should be understood that the logarithmic Mel spectrum is a commonly used method for representing sound features, converting the spectral information of a sound signal into a logarithmically scaled Mel spectrum. The Mel spectrum can better simulate the human ear's perception of sound because the human ear's perception of sound is based on the Mel scale. By extracting the logarithmic Mel spectrum of the denoised sound signal, the spectral features of the sound signal can be represented more accurately, including information such as pitch, tone, and timbre. The cochlear spectrogram is a method for representing sound features based on the human ear's perception mechanism. The cochlea in the human ear is an important organ responsible for converting sound signals into nerve impulse signals, and it performs special processing on the spectrum of sound signals. By extracting the cochlear spectrogram of the denoised sound signal, the human ear's perception mechanism can be simulated, better capturing important spectral features of the sound signal, including information such as sound intensity, pitch, and timbre. By extracting log-Melograms and cochlear spectra based on the compact shelving unit, the spectral characteristics of the denoised sound signal and the sensory mechanism simulating the human ear can be represented more accurately. These features are crucial for applications such as sound recognition, speech synthesis, and music analysis, and can improve system performance and effectiveness.
[0041] Furthermore, the first convolutional coding subunit 1223 is used to pass the log-Mel spectrogram through a first convolutional neural network acting as a filter to obtain a sound Mel spectrogram feature vector for the intelligent mobile shelving unit. It should be understood that convolutional neural networks (CNNs) are widely used in computer vision and audio processing, capable of learning and extracting features from input data. The log-Mel spectrogram is a spectral representation of a sound signal, containing important features of the sound. By using the log-Mel spectrogram as input to a CNN, the network can automatically learn and extract key features from the sound signal, such as the spectral shape of speech, and the time and frequency domain features of sound. CNNs possess locality perception, meaning that feature extraction can be performed on local regions of the input data through convolution operations. Log-Mel spectrograms typically have two dimensions: time and frequency, and convolution operations can effectively capture local features in these dimensions. This allows for better extraction of time and frequency domain information from the sound signal, including short-term changes in sound and spectral texture. CNNs have parameter sharing and dimensionality reduction characteristics, enabling the network to effectively learn and represent features of the input data. By inputting the log-Mel spectrum into the first convolutional layer of a CNN, the parameter-sharing property of convolution operations can be utilized to extract multiple different filter responses, capturing features at different frequencies and in the time domain. Simultaneously, convolution operations can be used for dimensionality reduction through pooling, decreasing the dimension of the feature vectors and improving computational efficiency. By using the log-Mel spectrum as the first convolutional neural network filter, key features of the sound signal can be effectively extracted and represented as a Mel spectrum feature vector for intelligent mobile shelving systems. These feature vectors can be used for tasks such as sound recognition and speech synthesis, improving the system's performance and accuracy.
[0042] Furthermore, the second convolutional coding subunit 1224 is used to pass the cochlear spectrogram through a second convolutional neural network acting as a filter to obtain a cochlear spectrogram feature vector for the intelligent compact shelving system. It should be understood that the cochlear spectrogram is a sound feature representation method designed based on the human ear's perception mechanism, capable of better capturing the important spectral features of sound signals. By using the cochlear spectrogram as input to the second convolutional neural network, the network can further learn and extract high-level features from the sound signal. These high-level features may include information such as the harmonic structure, pitch, and timbre of the sound, which can better represent the perceptual features of the sound signal. Convolutional neural networks have a hierarchical structure; through multiple layers of convolution and pooling operations, they can gradually learn and extract abstract representations of the input data. Using the cochlear spectrogram as input to the second convolutional layer allows the network to further learn and extract higher-level abstract features from the cochlear spectrogram. These abstract features can better represent the perceptual features of the sound signal, improving the system's ability to understand and recognize sound. Using the cochlear spectrogram as input to the second convolutional layer allows the feature information from the log-Mel spectrogram and the cochlear spectrogram to be integrated. Since log-Mel spectrograms and cochlear spectrograms reflect different aspects of sound signals, combining them can provide a more comprehensive and richer representation of sound features. Through learning in the second convolutional layer, the network can fuse and integrate these features to obtain more discriminative and expressive cochlear spectrogram feature vectors. By using the cochlear spectrogram as a filter in the second convolutional neural network, high-level features and abstract representations of sound signals can be further extracted and learned, resulting in even more expressive and discriminative cochlear spectrogram feature vectors. These feature vectors can be applied to tasks such as sound recognition and speech synthesis, improving the system's performance and effectiveness.
[0043] Specifically, the fusion vector subunit 1225 is used to fuse the sound Mel-spectrum feature vector and the cochlear spectrogram feature vector of the intelligent mobile shelving unit to obtain a sound association feature matrix. It should be understood that the sound Mel-spectrum and cochlear spectrogram of the intelligent mobile shelving unit are sound feature representations obtained based on different principles and feature extraction methods. The sound Mel-spectrum mainly reflects the spectral characteristics of the sound signal, while the cochlear spectrogram is closer to the spectral characteristics perceived by the human ear. These two features are complementary in capturing different aspects of the sound signal, and fusion can obtain more comprehensive and richer sound feature information. Fusion of different types of sound features can improve the system's robustness to different environments and sound conditions. The sound Mel-spectrum and cochlear spectrogram may have different advantages in different sound scenarios. By fusing these two features, their advantages can be fully utilized, improving the system's ability to recognize and analyze sound in various sound scenarios and conditions. Fusion of the sound Mel-spectrum and cochlear spectrogram feature vectors can enhance the expressive power of the features. These two features have different resolutions and distributions in the frequency domain, and fusion can provide more global and detailed sound feature information. By obtaining the sound association feature matrix of the mobile shelving unit, the time-domain and frequency-domain characteristics of the sound signal can be better represented, improving the accuracy and effectiveness of sound recognition and analysis. By fusing the Mel-spectral feature vector and the cochlear spectroscopic feature vector of the intelligent mobile shelving unit's sound, a sound association feature matrix can be obtained. This matrix possesses more comprehensive and richer sound feature information, enhancing the system's ability to understand and analyze sound. This fusion method can be applied to tasks such as sound recognition and speech synthesis, improving the system's performance and effectiveness.
[0044] Specifically, the third convolutional coding subunit 1226 is used to process the mobile shelving sound association feature matrix through a third convolutional neural network model, which acts as a feature extractor, to obtain a mobile shelving sound association feature map. It should be understood that the mobile shelving sound association feature matrix contains fused information from the sound Mel spectrogram and cochlear spectrogram, as well as abstract features learned by them in the convolutional neural network. By using the mobile shelving sound association feature matrix as input to the third convolutional neural network, higher-level features of the sound signal can be further extracted and learned. These higher-level features can include information such as the temporal and frequency domain structure of the sound, prosody and pitch of speech, and can better represent the semantic and emotional features of the sound signal. The mobile shelving sound association feature matrix is a two-dimensional matrix, where each element represents the distribution of sound features in the time and frequency dimensions. Through convolution and pooling operations of the third convolutional neural network model, the spatial relationships in this two-dimensional matrix can be modeled. The network can learn the interactions and dependencies between different regions, thereby better understanding and representing the temporal and spectral features of the sound signal. The third convolutional neural network model can transform the mobile shelving sound association feature matrix into a mobile shelving sound association feature map. Feature maps are a representation in convolutional neural networks (CNNs) that can better express the local and global features of input data. Through learning in the third convolutional layer, the network can extract and learn more abstract and semantic feature information from the sound association feature matrix of mobile shelving units, resulting in a more discriminative and expressive sound association feature map. By using the sound association feature matrix of mobile shelving units as a feature extractor in the third convolutional neural network model, high-level features and spatial relationships of the sound signal can be further extracted and learned, and represented as the sound association feature map of mobile shelving units. This feature map can be applied to tasks such as sound recognition, sentiment analysis, and speech synthesis, improving the system's ability to understand and generate sound.
[0045] In this embodiment, the mobile shelving data fusion module 130 is used to perform high-dimensional spatial unit manifold sub-dimensional hyperconvex correlation measurement on the mobile shelving tracking feature map and the mobile shelving sound association feature map to obtain a smart lighting control classification feature map. It should be understood that the mobile shelving tracking feature map and the mobile shelving sound association feature map originate from different perceptual modalities, namely visual and auditory. By fusing these two feature maps, visual and auditory information can be fused and integrated across modalities. This can fully utilize the advantages of different perceptual modalities and improve the smart lighting control system's perception of the environment and user behavior. The mobile shelving tracking feature map and the mobile shelving sound association feature map provide spatial and semantic features of visual and auditory information, respectively. Fusing these two feature maps can enrich the expressive power of features and provide more comprehensive and diverse feature information. This helps improve the accuracy and robustness of the smart lighting control system in classification tasks. Both the mobile shelving tracking feature map and the mobile shelving sound association feature map contain certain contextual information. By fusing these two feature maps, the contextual information of different perceptual modalities can be better integrated and utilized. For example, when an intelligent lighting control system needs to determine user behavior, fusing visual and auditory information can provide more comprehensive and accurate contextual information, thereby better understanding and classifying user behavior. By fusing the mobile shelving tracking feature map and the mobile shelving sound association feature map, a classification feature map for intelligent lighting control can be obtained. This feature map integrates multimodal information, enriches the expressive power of the features, and integrates contextual information.
[0046] Specifically, in the technical solution of this application, the mobile shelving tracking feature map and the mobile shelving sound association feature map represent information from surveillance video and sound signals, respectively. The mobile shelving tracking feature map is obtained by extracting multiple keyframes of the mobile shelving and performing target detection, and is used to represent the region of interest of the target object in the surveillance video. The mobile shelving sound association feature map, on the other hand, is obtained by extracting Mel spectrograms and cochlear spectrograms from the denoised sound signal and processing them through a convolutional neural network, and is used to represent the features of the sound signal. Since the mobile shelving tracking feature map and the mobile shelving sound association feature map originate from different data sources and processing methods, they may have differences in dimensionality and scale within high-dimensional feature space units. Specifically, the mobile shelving tracking feature map is obtained through target detection network processing and may have higher dimensionality and scale. The mobile shelving sound association feature map, however, is obtained by extracting Mel spectrograms and cochlear spectrograms and processing them through a convolutional neural network, and may have lower dimensionality and scale. When these two feature maps are fused, some technical problems may arise due to the differences in dimensionality and scale. For example, during the fusion process, feature alignment difficulties may arise due to dimensional differences, i.e., how to align two feature maps in a high-dimensional space so that they can meaningfully influence each other. Furthermore, scale differences may lead to unbalanced weight distribution during feature fusion, affecting the final classification result. These differences can also cause local structural collapse in the feature space, meaning that the structural information of some local features may be lost or become unstable during fusion. To address these issues, the technical solution in this application performs a high-dimensional space unit manifold sub-dimensional hyperconvex correlation measure on the mobile shelving tracking feature map and the mobile shelving sound association feature map to obtain a smart lighting control classification feature map. This better fuses the mobile shelving tracking feature map and the mobile shelving sound association feature map, avoiding technical problems caused by dimensional and scale differences. This improves the quality of the fused smart lighting control classification feature map and reduces the occurrence of problems such as local structural collapse or ill-fitting alignment.
[0047] Accordingly, in one embodiment of this application, the mobile shelving data fusion module is configured to: include: a position-averaging unit, configured to calculate a position-averaging feature map between the mobile shelving tracking feature map and the mobile shelving sound association feature map; calculate a position-differential feature map between the mobile shelving tracking feature map and the mobile shelving sound association feature map; a differential eccentricity feature unit, configured to calculate the differential feature maps between the mobile shelving tracking feature map and the mobile shelving sound association feature map and the position-averaging feature map respectively to obtain a first differential eccentricity feature map and a second differential eccentricity feature map; and a logarithmic difference unit, configured to calculate the... The first logarithmic difference partial center feature map and the second logarithmic difference partial center feature map are obtained by taking the base-2 logarithmic function value of the feature values at each position in the first logarithmic difference partial center feature map and the second logarithmic difference partial center feature map; the correction feature unit is used to divide the first logarithmic difference partial center feature map by the position-based summation feature map between the first logarithmic difference partial center feature map and the second logarithmic difference partial center feature map to obtain the correction feature map; the position-based weighting unit is used to calculate the position-based weighted sum between the logarithmic position-based difference feature map of the correction feature map and the position-based difference feature map to obtain the intelligent lighting control classification feature map.
[0048] In other words, considering the differences in dimension and scale between the feature manifolds of the mobile shelving tracking feature map and the mobile shelving sound association feature map in the high-dimensional feature space unit, the fused mobile shelving tracking feature map may experience technical problems such as local structural collapse or ill-fitting due to the differences in dimension and scale during the fusion of the mobile shelving tracking feature map and the mobile shelving sound association feature map.
[0049] To address the aforementioned technical problems, the technical solution of this application employs a high-dimensional spatial unit manifold sub-dimensional hyperconvex correlation measurement on the mobile shelving tracking feature map and the mobile shelving sound association feature map. This measurement uses the mean feature map of the mobile shelving tracking feature map and the mobile shelving sound association feature map as the pseudo-cluster center of the feature manifold. By constructing a hyperconvex correlation measurement function based on the feature map's feature manifold, the feature values at each position between feature maps can maintain consistency with the pseudo-cluster center of the feature manifold in their sub-dimensions. This achieves hyperconvex correlation matching of the feature map's feature manifold, effectively measuring the similarity and difference between feature maps, enhancing the hyperconvex correlation of the feature map's feature manifold, and improving the robustness and accuracy of the feature map's feature manifold.
[0050] In this embodiment, the mobile shelving data analysis module 140 is used to process the intelligent lighting control classification feature map through a classifier to obtain a classification result, which indicates whether the lighting is turned on. It should be understood that the intelligent lighting control classification feature map, after processing by the classifier, can map the input feature map to different categories or labels for classifying different lighting states. The classifier can learn patterns and features in the feature map, thereby determining whether the lighting needs to be turned on. The classification result can be directly used to control the on / off state of the lighting. For example, if the classification result indicates that the lighting needs to be turned on, the intelligent lighting system can correspondingly control the brightness, color, and other attributes of the light to meet the user's needs. Conversely, if the classification result indicates that the lighting does not need to be turned on, the system can turn off the lighting to save energy. Classifying the intelligent lighting control classification feature map through a classifier enables flexibility and personalization of lighting control. Different classification results can correspond to different lighting scenarios, such as day and night, different activity environments, etc. Based on the classification result, the intelligent lighting system can automatically adjust the lighting state to provide a personalized lighting experience that meets the user's needs and preferences. By classifying the feature maps of smart lighting control using a classifier and then using the classification results to indicate whether the lights should be turned on, automated control and personalized services for smart lighting systems can be achieved. Such systems can make intelligent decisions based on the environment and user behavior, providing a convenient, comfortable, and energy-efficient lighting experience.
[0051] Accordingly, in one embodiment of this application, the mobile shelving data analysis module is used to: process the intelligent lighting control classification feature map using the classifier with the following formula to obtain the classification result;
[0052] The formula is: O = softmax{(W c B c )|Project(F)}, where Project(F) represents projecting the intelligent lighting control classification feature map into a vector, W c Let B be the weight matrix. c represents the bias vector, softmax represents the normalization exponential function, and O represents the classification result.
[0053] In summary, the intelligent mobile shelving system and method based on IoT technology described in this application acquires monitoring video and audio signals within the intelligent mobile shelving unit, extracts tracking feature maps and sound association feature maps from them, then fuses these feature maps to obtain an intelligent lighting control classification feature map, and classifies the intelligent lighting control classification feature map using a classifier to obtain a classification result used to indicate whether the lighting is turned on. This enables intelligent lighting control, avoids power waste and economic losses, and improves the efficiency and convenience of file management.
[0054] As described above, the IoT-based intelligent mobile shelving system 100 according to the embodiments of this application can be implemented in various terminal devices, such as servers for IoT-based intelligent mobile shelving systems. In one example, the IoT-based intelligent mobile shelving system 100 can be integrated into the terminal device as a software module and / or hardware module. For example, the IoT-based intelligent mobile shelving system 100 can be a software module in the operating system of the terminal device, or it can be an application developed for the terminal device; of course, the IoT-based intelligent mobile shelving system 100 can also be one of many hardware modules of the terminal device.
[0055] Alternatively, in another example, the IoT-based smart mobile shelving system 100 and the terminal device can also be separate devices, and the IoT-based smart mobile shelving system 100 can be connected to the terminal device via wired and / or wireless networks, and transmit interactive information in accordance with an agreed data format.
[0056] Figure 5 This is a flowchart of an intelligent mobile shelving method based on Internet of Things (IoT) technology according to an embodiment of this application. Figure 5 As shown, the intelligent mobile shelving method based on Internet of Things technology according to the embodiments of this application includes the following steps: S110, acquiring monitoring video and sound signals from the intelligent mobile shelving; S120, extracting a mobile shelving tracking feature map from the monitoring video and extracting a mobile shelving sound association feature map from the sound signals; S130, performing a high-dimensional space unit manifold sub-dimensional hyperconvex correlation measurement on the mobile shelving tracking feature map and the mobile shelving sound association feature map to obtain an intelligent lighting control classification feature map; S140, passing the intelligent lighting control classification feature map through a classifier to obtain a classification result, the classification result being used to indicate whether the lighting is turned on.
[0057] Here, those skilled in the art will understand that the specific operations of each step in the above-described intelligent mobile shelving method based on Internet of Things technology have been referenced above. Figures 1 to 4The description of the IoT-based intelligent mobile shelving system is detailed here, and therefore, its repeated description will be omitted.
[0058] The basic principles of this disclosure have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in this disclosure are merely examples and not limitations, and should not be considered as essential features of each embodiment of this disclosure. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the scope of this disclosure to the necessity of employing the aforementioned specific details for implementation.
[0059] The block diagrams of devices, apparatuses, devices, and systems disclosed herein are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as “comprising,” “including,” “having,” etc., are open-ended terms meaning “including but not limited to,” and are used interchangeably with them. The terms “or” and “and” as used herein refer to the terms “and / or,” and are used interchangeably with them unless the context clearly indicates otherwise.
[0060] The word “such as” refers to the phrase “such as but not limited to”, and can be used interchangeably with it.
[0061] Additionally, as used herein, the “or” used in a list of items beginning with “at least one” indicates a separate list, such that a list of, for example, “at least one of A, B, or C” means A or B or C, or AB or AC or BC, or ABC (i.e., A and B and C). Furthermore, the word “exemplary” does not imply that the described example is preferred or better than other examples.
[0062] It should also be noted that in the systems and methods of this disclosure, the components or steps can be decomposed and / or recombined. These decompositions and / or recombinations should be considered as equivalent solutions to this disclosure.
[0063] Various changes, substitutions, and modifications can be made to the technology described herein without departing from the teachings defined by the appended claims. Furthermore, the scope of the claims of this disclosure is not limited to the specific aspects of the processes, machines, manufactures, events, means, methods, and actions described above. Currently existing or later-developed processes, machines, manufactures, events, means, methods, or actions that perform substantially the same function or achieve substantially the same result as the corresponding aspects described herein can be utilized. Therefore, the appended claims include such processes, machines, manufactures, events, means, methods, or actions within their scope.
[0064] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but rather to be carried out within the widest scope consistent with the principles and novel features disclosed herein.
[0065] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this disclosure to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations therein.
Claims
1. A smart mobile shelving system based on Internet of Things (IoT) technology, characterized in that, include: The mobile shelving data acquisition module is used to acquire monitoring video and audio signals from the intelligent mobile shelving. The mobile shelving data extraction module is used to extract the mobile shelving tracking feature map from the monitoring video and the mobile shelving sound association feature map from the sound signal; The data fusion module for the mobile shelving unit includes: The position-average unit is used to calculate the position-average feature map between the mobile shelving tracking feature map and the mobile shelving sound association feature map; Calculate the position-difference feature map between the mobile shelving tracking feature map and the mobile shelving sound association feature map; The differential off-center feature unit is used to calculate the differential feature maps between the mobile shelving tracking feature map and the mobile shelving sound association feature map and the position mean feature map, respectively, to obtain the first differential off-center feature map and the second differential off-center feature map; The logarithmic difference unit is used to calculate the base-2 logarithmic function value of the feature values at each position in the first and second logarithmic difference partial center feature maps to obtain the first and second logarithmic difference partial center feature maps. The correction feature unit is used to divide the first logarithmic difference off-center feature map by the position-sum feature map between the first logarithmic difference off-center feature map and the second logarithmic difference off-center feature map to obtain the correction feature map; The position-weighted unit is used to calculate the position-weighted sum between the logarithm of the corrected feature map and the position-differential feature map to obtain the intelligent lighting control classification feature map; The mobile shelving data analysis module is used to pass the intelligent lighting control classification feature map through a classifier to obtain a classification result, which is used to indicate whether the lighting is turned on.
2. The intelligent mobile shelving system based on Internet of Things technology according to claim 1, characterized in that, The mobile shelving data extraction module includes: The video processing unit is used to extract keyframes from the surveillance video and then obtain the tracking feature map of the mobile shelving unit through target detection; The sound signal processing unit is used to perform noise reduction processing on the sound signal and then obtain the sound association feature map of the mobile shelving unit through convolutional encoding.
3. The intelligent mobile shelving system based on Internet of Things technology according to claim 2, characterized in that, The video processing unit includes: The keyframe extraction subunit is used to extract multiple keyframes of the mobile shelving unit from the monitoring video. The target detection subunit is used to obtain multiple target object region maps by passing the multiple key frames of the mobile shelving unit through a target detection network based on an anchorless window; A three-dimensional tensor quantum unit is used to arrange the region of interest maps of the multiple target objects into a three-dimensional tensor and then pass them through a three-dimensional convolutional neural network model to obtain the tracking feature map of the compact shelving unit.
4. The intelligent mobile shelving system based on Internet of Things technology according to claim 3, characterized in that, The keyframe extraction subunit includes: Extract the initial frame secondary sub-unit, which is used to extract the initial frame from the monitoring video; A second-level sub-unit for calculating differential image frames is used to calculate the positional differences between other frames in the monitoring video and the initial frame to obtain differential image frames; A keyframe secondary subunit is determined, which is used to determine the multiple keyframes of the mobile shelving unit from the monitoring video based on a comparison between the sum of the feature values of all pixels in the differential image frame and a predetermined threshold.
5. The intelligent mobile shelving system based on Internet of Things technology according to claim 4, characterized in that, The determination of the keyframe secondary subunit is used for: The first keyframe level three sub-unit is determined, which is used to determine other frames corresponding to the differential feature map as the first keyframe based on the sum of the feature values of all pixels in the differential image frame being greater than a predetermined threshold. An initial frame is defined as a three-level sub-unit, and the first keyframe is defined as the initial frame.
6. The intelligent mobile shelving system based on Internet of Things technology according to claim 5, characterized in that, The target detection subunit includes: Each of the multiple keyframes of the compact shelving unit is passed through multiple convolutional layers to obtain multiple shallow feature maps. The multiple shallow feature maps are respectively passed through the target detection network based on anchorless windows to obtain the multiple target object region of interest feature maps.
7. The intelligent mobile shelving system based on Internet of Things technology according to claim 6, characterized in that, The sound signal processing unit includes: A noise reduction processing subunit is used to perform noise reduction processing on the sound signal to obtain a noise-reduced sound signal; The sound signal feature extraction subunit is used to extract the log-Mel spectrum and cochlear spectrum of the noise-reduced sound signal based on the compact shelving. The first convolutional coding subunit is used to pass the log-Mel spectrum through the first convolutional neural network as a filter to obtain the sound Mel spectrum feature vector of the intelligent mobile shelving. The second convolutional coding subunit is used to pass the cochlear spectrogram through a second convolutional neural network as a filter to obtain the cochlear spectrogram feature vector of the intelligent compact shelving unit. A fusion vector subunit is used to fuse the sound Mel spectrum feature vector of the intelligent mobile shelving unit and the cochlear spectrum feature vector of the intelligent mobile shelving unit to obtain the sound association feature matrix of the mobile shelving unit. The third convolutional coding subunit is used to pass the sound association feature matrix of the mobile shelving unit through the third convolutional neural network model, which acts as a feature extractor, to obtain the sound association feature map of the mobile shelving unit.
8. The intelligent mobile shelving system based on Internet of Things technology according to claim 7, characterized in that, The mobile shelving data analysis module is used to: process the intelligent lighting control classification feature map using the classifier with the following formula to obtain the classification result; Wherein, the formula is ,in This indicates that the intelligent lighting control classification feature map is projected into a vector. This is the weight matrix. This represents the bias vector. Represents the normalized exponential function, This indicates the classification result.
Citation Information
Patent Citations
Rehabilitation nursing interactive management system and method based on Internet of Things
CN117690583A