Somatosensory interactive display system based on depth camera

Through deep camera array and multimodal data analysis, the somatosensory interaction parameters are dynamically adjusted, which solves the user adaptability and accuracy problems in the existing technology, and realizes personalized and efficient somatosensory interaction display.

CN120335617BActive Publication Date: 2025-08-19HANGZHOU SICHUANG EXHIBITION TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510819537.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-08-19
Estimated Expiration
2045-06-19

AI Technical Summary

Technical Problem

In the existing somatosensory interaction technology, the fixed somatosensory interaction method cannot adapt to individual differences between different users, resulting in poor display effect and large-area data collection leads to increased difficulty and reduced accuracy of interaction.

Method used

The somatosensory interactive display system based on the depth camera is adopted, and multi-modal data such as infrared images, depth data and user habits are integrated to construct cross-dimensional analysis methods, dynamically set display parameters, and precise shooting and image processing are used to identify the user's action intentions.

Benefits of technology

It improves the interactive effect, reduces user visual fatigue, reduces interaction difficulty, increases the recognition accuracy and richness of interactive instructions, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120335617B_ABST
    Figure CN120335617B_ABST
Patent Text Reader

Abstract

The present invention discloses a somatosensory interactive display system based on a depth camera, which relates to the field of somatosensory interaction technology; the system includes a data retrieval module, a parameter analysis module, a secondary analysis module, an image processing module and a verification and display module: the data retrieval module: extracts the user's most recent N historical somatosensory interaction data from a database based on a data retrieval algorithm, and preprocesses the data; the parameter analysis module: analyzes the preprocessed data, obtains a dimming strategy, sequentially collects and analyzes the user's distance data and infrared image data, and marks the interaction area; its technical key points are: a cross-dimensional analysis method is constructed for multimodal data such as the fusion of infrared images, depth data, user habits and the user's own conditions, which can accurately grasp the user's intentions and construct somatosensory interaction display parameters exclusive to the user, thereby improving the interaction effect and greatly improving the display effect, and has good usage prospects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of somatosensory interaction technology, and in particular to a somatosensory interaction display system based on a depth camera. Background Art

[0002] A depth camera is an imaging device that can obtain the distance information between objects in the scene and the camera. Unlike traditional cameras that only record two-dimensional plane images, it constructs three-dimensional spatial information through special technical means to achieve perception of the position, shape and distance of objects in three-dimensional space.

[0003] Somatosensory interactive display is a form of display that interacts with digital content or physical devices through human movements, postures, gestures, and even expressions. It integrates sensor technology, computer vision, deep learning and other technologies, breaking the traditional "one-way viewing" model and allowing viewers to directly manipulate virtual content or trigger physical devices through body language, achieving an immersive and interactive experience.

[0004] At present, the Chinese invention patent with the existing patent application number "CN113596353B" discloses a somatosensory interactive data processing method, device and somatosensory interactive equipment, which records the following: obtaining audio data to be played and its corresponding first somatosensory image data, outputting the audio data to the sound system for sound effect processing and playback; receiving second somatosensory image data fed back in real time by the image processing system; wherein the second somatosensory image data is image data associated with the user obtained by the image processing system when playing the audio data; comparing and analyzing the second somatosensory image data according to the first somatosensory image data, and generating the user's interactive data processing results based on the comparison and analysis for display. This technical solution can provide users with intelligent somatosensory interactive feedback auxiliary functions in a timely manner in entertainment scenarios of limb movement.

[0005] However, during the implementation of the above technical solution, at least the following technical problems were found:

[0006] First, its fixed somatosensory interaction method, based on somatosensory interaction, means that children and adults share a set of setting parameters. However, different users have different preferences, so this method has different display effects when used, and its applicability is relatively low.

[0007] Secondly, the existing solution is that children and adults share a set of setting parameters, but due to the different height data, a large area needs to be photographed, and the collected pictures are large-area pictures. In order to reduce the amount of data processing, the user's small-scale movements cause little change, making it impossible to collect the instructions corresponding to the movement, making the somatosensory interaction more difficult and not meeting the needs of technical personnel in this field. For this reason, a somatosensory interaction display system based on a depth camera is provided. Summary of the Invention

[0008] (1) Technical problems solved

[0009] In response to the shortcomings of the existing technology, the present invention provides a somatosensory interactive display system based on a depth camera, which constructs a cross-dimensional analysis method for multimodal data such as infrared images, depth data, user habits, and the user's own conditions. It can accurately grasp the user's intentions and construct somatosensory interactive display parameters unique to the user, thereby greatly improving the interaction effect and the display effect, meeting market demand, having good usage prospects, and solving the problems raised in the background technology.

[0010] (2) Technical solution

[0011] To achieve the above objectives, the present invention is implemented through the following technical solutions:

[0012] The somatosensory interactive display system based on the depth camera includes a data retrieval module, a parameter analysis module, a secondary analysis module, an image processing module, and a verification and display module:

[0013] Data retrieval module: extracts the user's most recent N historical somatosensory interaction data from the database based on the data retrieval algorithm and preprocesses the data;

[0014] Parameter analysis module: Analyzes pre-processed data to obtain dimming strategies, collects and analyzes user distance data and infrared image data in sequence, and marks interactive areas;

[0015] Secondary analysis module: Analyzes the marked interaction area and dynamically sets the shooting parameters of each depth camera in the depth camera array based on the analysis results;

[0016] Image processing module: Segments the images continuously collected by the depth camera array based on the image segmentation algorithm, performs feature recognition on the segmented graphics, calculates the graphic similarity between the continuously collected images, and compares the graphic similarity with the preset similarity threshold. If the graphic similarity is less than or equal to the preset similarity threshold, no processing is performed;

[0017] If the graphic similarity is greater than the preset similarity threshold, semantic understanding technology is combined to identify the interactive instructions corresponding to the actions in the graphic;

[0018] Verification and display module: summarizes the interaction instructions, generates interaction strategies, and generates and displays results according to the interaction strategies.

[0019] Furthermore, the historical somatosensory interaction data includes age data, user vision detection data, display parameters, infrared image data and distance data. The historical somatosensory interaction data is retrieved from the database through distributed indexing technology, the user vision detection data is entered into the database by the user, the infrared image data is captured and collected by the infrared camera at the somatosensory interaction display site, and the distance data is collected by the depth camera.

[0020] Furthermore, the steps for analyzing the pre-processed data and obtaining the dimming strategy are as follows:

[0021] Feature analysis: Obtain the vision and astigmatism data from the user's vision test data, and compare the vision and astigmatism data with the standard range to obtain the vision grade and astigmatism grade;

[0022] Data mining: Use multiple linear regression functions to process the brightness, resolution, and contrast of display parameters to predict parameter requirements for the current period;

[0023] Data correction: Based on the visual acuity level, astigmatism level, and age data, the corresponding preset display parameter interval set is searched. The preset display parameter interval set includes a display brightness interval, a display resolution interval, and a display contrast interval. The intersection of the three preset display parameter interval sets is taken to form the preset display parameter interval set.

[0024] Comparison and judgment: Compare the parameter requirements predicted for the current period with the preset display parameter interval set;

[0025] If the parameter requirements for the current time period are predicted to be within the corresponding display data interval in the preset display parameter interval set, then the parameters for the current time period are extracted;

[0026] If the parameter demand for the predicted current period is outside the corresponding display data interval in the preset display parameter interval set, extract the data in the preset display parameter interval set that is closest to the parameter demand for the predicted current period;

[0027] Distance data is acquired and analyzed to determine the user's height and the distance between the user and the display device. The display screen size and center position are adjusted based on the user's height and distance. The extracted data, display screen size, and center position are aggregated to form a dimming strategy.

[0028] Furthermore, the steps of sequentially collecting and analyzing the user's distance data and infrared image data are as follows:

[0029] Distance data collection and analysis: Obtain distance data and adjust the focus of the infrared camera based on the distance data to collect infrared images;

[0030] Feature enhancement: A multi-scale retinal cortex method based on Retinex theory is used to separate the infrared image reflection component and illumination component, highlighting the user's outline and key areas of somatosensory interaction;

[0031] Noise reduction: Use non-local mean filtering algorithm to remove noise from the image, while retaining edge information;

[0032] Normalization processing: normalize the image pixel values to the [0, 1] interval to unify the data standards;

[0033] Edge extraction: Use the Canny edge detection algorithm to analyze the dynamic changes in local grayscale of the image, adjust the high and low thresholds, and extract the edge contour of the user's body;

[0034] Feature extraction: Use the deep learning-based OpenPose method to extract the human body outline and identify the somatosensory interaction area.

[0035] Furthermore, the analysis of the somatosensory interaction area of the marker is to find the location of the somatosensory interaction area of the marker, calculate the size of the area of the somatosensory interaction area of the marker, calculate the ratio of the area of each somatosensory interaction area of the marker to the area of the infrared image, and adjust the focal length data of each depth camera in the depth camera array according to the ratio and the focal length data of the infrared camera.

[0036] Furthermore, before segmenting the images continuously captured by the depth camera array, the images are first enhanced, then the enhanced images are denoised, and then the pixels of the images are normalized. When segmenting the images, the grayscale value segmentation method is used to segment the images of the user's somatosensory interaction areas.

[0037] Furthermore, the steps for image segmentation are as follows:

[0038] Feature extraction: obtain the grayscale value of each pixel in the image and construct the grayscale histogram of the image;

[0039] Threshold division: The grayscale histogram is clustered using the bimodal method to divide the grayscale value range into the interaction area interval and the background interval;

[0040] Initial marking: Compare each pixel in the image with the interaction area and the background area, and mark the pixels in the interaction area;

[0041] Matrix construction: Calculate the grayscale difference between each pixel in the image and its horizontally adjacent pixels, and construct a difference matrix based on the grayscale difference;

[0042] Optimization processing: Use the Otsu algorithm to calculate the optimal difference threshold, and compare the size of the data in the difference matrix with the optimal difference threshold;

[0043] If the data in the difference matrix are all less than or equal to the optimal difference threshold, no processing is done;

[0044] If there is data in the difference matrix that is greater than the optimal difference threshold, then mark the data that is greater than the optimal difference threshold, and find the amount of marked data of adjacent pixels. If the amount of marked data is less than or equal to the preset value, then the pixel corresponding to the data is not marked. If the amount of marked data is greater than the preset value, then mark the pixel corresponding to the data.

[0045] Graphics generation: Set the grayscale value of unmarked pixels to 0 to obtain the interactive part graphics.

[0046] Furthermore, when performing feature recognition on the segmented graphics, a detection model is used to process the graphics of the interactive parts to extract the geometric features of the interactive parts; then the difference in the geometric features of the graphics at two adjacent time points is calculated, and finally the difference is converted to calculate the graphic similarity.

[0047] Furthermore, the steps for identifying the interaction instructions corresponding to the actions in the graphics are as follows:

[0048] Construct standard feature vector: set the standard feature vector of K groups of interaction instructions;

[0049] Build a mapping library: associate standard interaction parts with interaction instructions;

[0050] Extract vector: take the geometric features of the next frame of graphics in the adjacent time as the feature vector;

[0051] Calculate distance: Calculate the distance between the feature vector and the data in the K groups of standard feature vectors, and calculate and summarize the comprehensive distance between the feature vector and the K groups of standard feature vectors;

[0052] Scoring: Normalize the K groups of comprehensive distances to obtain the K groups of proximity scores;

[0053] Decision processing: Get the maximum value among the K groups of proximity scores and compare the maximum value with the preset standard score;

[0054] If the maximum value ≥ the standard score, output the interaction instruction corresponding to the maximum value;

[0055] If the maximum value is less than the standard score, it is determined that no interactive instruction has been made.

[0056] Furthermore, the steps for summarizing the interaction instructions and generating the interaction strategy are as follows:

[0057] Determine the number of interaction instructions. If the number of interaction instructions is 0, determine that the generated interaction strategy is not to be processed.

[0058] If the interaction instruction is 1 group, the generated interaction strategy is determined to be the interaction instruction;

[0059] If there are multiple groups of interaction instructions generated, calculate the proportion of each interaction instruction in the total interaction instructions, compare the proportions, and obtain the maximum proportion. If the maximum proportion is greater than the set value, determine that the generated interaction strategy is the interaction instruction corresponding to the maximum proportion;

[0060] If the maximum proportion is ≤ the set value, the generated interaction strategy is determined to be no processing.

[0061] (3) Beneficial effects

[0062] The present invention provides a somatosensory interactive display system based on a depth camera, which has the following beneficial effects:

[0063] 1. The present invention provides a somatosensory interactive display system based on a depth camera. It constructs a cross-dimensional analysis method for multimodal data such as infrared images, depth data, user habits, and the user's own conditions. It can accurately grasp the user's intentions and construct somatosensory interactive display parameters unique to the user, thereby greatly improving the interaction effect and the display effect, meeting market demand, and has good prospects for use.

[0064] 2. The present invention provides a somatosensory interactive display system based on a depth camera, which changes the existing fixed somatosensory interaction mode, analyzes the user's situation and position, and formulates display parameters exclusive to the user. This personalized adaptive mechanism reduces user visual fatigue, improves interaction comfort, and enables users to extend the somatosensory interaction time, thereby effectively improving the display effect, and has good usage prospects.

[0065] 3. The present invention provides a somatosensory interactive display system based on a depth camera, which uses infrared cameras to locate and partition users, and then uses a depth camera array to focus on shooting the somatosensory interactive parts, and uses geometric features for comparison, which can realize the recognition of small-scale movements of the somatosensory interactive parts, thereby changing the existing method of using large-scale movements to achieve somatosensory interaction, greatly reducing the difficulty of somatosensory interaction, and greatly improving the accuracy of somatosensory interaction, with good use effect. Therefore, this solution has good application prospects.

[0066] 4. The present invention provides a somatosensory interactive display system based on a depth camera, which collects data from multiple somatosensory interactive parts of the user, allowing the user to use multiple somatosensory interactive parts to implement the operation of the same interactive instruction, which can effectively avoid the situation where some interactive instructions cannot be recognized, and accordingly can effectively increase the number of interactive instructions, making the somatosensory interaction richer, thereby greatly increasing the user's desire to explore, thereby improving the display effect, and has good application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] Figure 1This is the overall system architecture diagram of the present invention;

[0068] Figure 2 A state structure diagram is used for the system of the present invention. DETAILED DESCRIPTION

[0069] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0070] With the development of society, the application of somatosensory interactive entertainment display technology is becoming more and more common. The existing somatosensory interactive entertainment display is a high-tech display method that integrates hardware interactive equipment, somatosensory interactive system software and three-dimensional digital content. It relies on video motion capture technology, tracks human movements through 3D somatosensory cameras, separates people from the camera screen, and tracks multiple parts of the human body in real time. It can track up to two players at the same time. The audience can use natural body movements and gestures, such as waving, turning, jumping, etc., to manipulate the content on the display screen in mid-air, and realize interactive operations such as zooming in and out of pictures and videos, dragging and rotating, and controlling racing games. It also supports multiple people operating simultaneously. This display method is highly interactive, attractive and interesting, and is widely used in entertainment games, interactive exhibitions, advertising and marketing and other fields. It can bring an immersive experience to the audience, enhance brand impression and the audience's understanding and memory of the display content.

[0071] However, existing somatosensory interactive entertainment display technologies have the following limitations:

[0072] Poor applicability: Existing somatosensory interactive entertainment display technology is general-purpose and deployed in a single, stable environment. It requires users to be in a fixed position for somatosensory interaction. While the external environment is uniform, users' situations and habits are not. Using a fixed interactive display method only works for certain groups of people with similar heights and habits; for others, the display effect is poor.

[0073] Low precision: In order to allow most people to experience it, the existing somatosensory interactive entertainment display technology adopts a large-scale shooting and data collection method. Large-scale collection results in a large amount of content being captured, so the amount of data to be processed is very large. In order to enable the system to respond quickly, a method is adopted in which command recognition is only performed when there is a large difference in the pictures. This method increases the amount of movement of the user during interaction and reduces the comfort of interaction. At the same time, it may be necessary to issue multiple commands for recognition, and the effect is not good.

[0074] The present invention adopts multimodal data fusion technology to integrate information such as infrared images and depth data. Through steps such as image segmentation, feature extraction, action recognition and command mapping, it constructs a complete closed loop from data acquisition, intention understanding to interactive execution, and forms a standardized process specification from data acquisition, preprocessing, feature extraction to action recognition, command mapping and execution. It can accurately perceive user actions and needs, adaptively adjust display device parameters, provide personalized visual experience, and effectively improve the robustness, real-timeness and accuracy of the interactive system, laying a technical foundation for the field of intelligent visual interaction.

[0075] like Figure 1 As shown, the specific implementation plan is as follows:

[0076] 1. Instrument selection;

[0077] Compared with existing somatosensory interactive entertainment display technology, this solution only makes improvements in data collection and data processing. For the data collection end, the previous single camera is changed to an infrared camera and a camera matrix. For data processing, it is an improvement at the software operation level and does not involve hardware.

[0078] The data uploaded by users is uploaded directly through the APP or WeChat applet, without the need to deploy corresponding collection equipment for on-site collection. Therefore, for the implementation of the entire system, the cost of replacing the instrument is relatively low.

[0079] The focus of this solution is to study user habits and the user's own situation, and to develop display parameters that are unique to the user. These parameters will change as the user changes, so that the user experience during the interaction is better and more in line with their own habits, thereby improving the display effect.

[0080] The focus of this solution is on the acquisition of user data in the early stage and the mid-term division of human body interaction parts based on clear infrared images, the formulation of corresponding shooting parameters, and the recognition of subsequent interaction instructions based on accurate and clear shooting parameters.

[0081] 2. System solution;

[0082] like Figure 1 As shown, this system mainly analyzes the following five stages during its implementation: first, it collects the user's most recent usage data, then processes and analyzes the most recent data, formulates a preliminary strategy, and then further analyzes the current user's interaction position and situation. Combined with the current user data analysis, it sets the relevant parameters for image acquisition, and then further analyzes the acquired image data to determine whether command recognition is required. When command recognition is required, it identifies the command and selects the command for display operation.

[0083] 2.1 Data collection;

[0084] Data collection is the collection of historical data, and the collection of historical data is based on the data retrieval module.

[0085] Data retrieval module: extracts the user's most recent N historical somatosensory interaction data from the database based on the data retrieval algorithm and preprocesses the data;

[0086] The N-times data refers to the data from the past three years. If the duration exceeds three years, it will not be adopted.

[0087] The data of each use will be stored in the corresponding usage record in the database, which is convenient for subsequent retrieval and analysis.

[0088] Historical somatosensory interaction data includes age data, user vision test data, display parameters, infrared image data, and distance data. Historical somatosensory interaction data is retrieved from the database using distributed indexing technology. User vision test data is entered into the database by the user or with the assistance of relevant technical personnel on site. In actual use, it can be further optimized to add age data. Because different ages have different sensitivities to light, for example, it can further help formulate brightness parameters. For example, for children under 12 years old, the brightness requirement is between 350-450Lx and the color temperature should be between 3100-3800K; for teenagers aged 13-18, the brightness requirement is 520-680Lx and the color temperature should be between 4200-4900K; for adults aged 18-60, the brightness requirement is 710-950Lx and the color temperature should be between 4200-4900K; for the elderly over 65 years old, the brightness requirement is 1000-1400Lx and the color temperature should be between 3000-4000K.

[0089] In actual use, the infrared camera can be used in combination with the depth camera. For example, one of the depth cameras on the depth camera matrix is model OAK-D-Pro. This depth camera uses structured light ranging technology and has an infrared laser dot matrix emitter and an infrared lighting LED. Its positioning accuracy is improved to the sub-millimeter level and can be used in various scenarios such as depth measurement, image recognition and positioning.

[0090] It can also be deployed separately, and the infrared camera can be set up independently, for example, the infrared camera uses the DS-2XA3T46EF-LS network camera.

[0091] While retrieving data, distributed database retrieval technology is used to achieve parallel processing and query of data. This technology greatly improves data retrieval efficiency by distributing data across multiple nodes and using distributed indexes to quickly locate users' personalized data such as myopia degree and light brightness preferences.

[0092] This retrieval implementation builds a tree index based on each data shard, supporting fast range queries.

[0093] Preprocessing involves cleaning and transforming the data. First, the raw data is converted into a unified structured format, and then missing values, outliers, and duplicate values are processed.

[0094] The processing of missing values includes deleting records, filling with interpolation methods and filling with mean methods. The missing data are compared with the set standards according to the set rules. If the standards are not met, the records are deleted. If the standards are met, the data type is determined, and then the pre-set interpolation method or mean method is selected according to the data type to fill the missing values.

[0095] The outlier processing is to identify outliers through the 3σ principle. If a characteristic value deviates from the mean by more than 3 times the standard deviation, it is judged as an anomaly. Then the abnormal data is manually judged and corrected. The value can also be inferred based on the previous and next data.

[0096] The processing of duplicate values is to detect whether there are duplicate records. If there are duplicate records, keep the record that appears first and delete subsequent duplicates.

[0097] After data processing, the data were normalized and standardized, and the standardization was performed using Z-score standardization.

[0098] After processing, the data are aligned based on timestamps to ensure the association of multimodal data at the same time point.

[0099] The present invention provides a somatosensory interactive display system based on a depth camera. It constructs a cross-dimensional analysis method for multimodal data such as infrared images, depth data, user habits, and the user's own conditions. It can accurately grasp the user's intentions and construct somatosensory interactive display parameters unique to the user, thereby greatly improving the interaction effect and the display effect, meeting market demand, and has good usage prospects.

[0100] 2.2 Preliminary analysis;

[0101] The preliminary analysis mainly involves analyzing the collected data and formulating some display parameters. This process is based on the parameter analysis module.

[0102] Parameter analysis module: Analyzes the preprocessed data to obtain the dimming strategy, collects and analyzes the user's distance data and infrared image data in sequence, and marks the interactive area.

[0103] The steps to analyze the pre-processed data and obtain the dimming strategy are as follows:

[0104] Feature analysis: Obtain the vision and astigmatism data from the user's vision test data, and compare the vision and astigmatism data with the standard range to obtain the vision grade and astigmatism grade;

[0105] The standard interval is a pre-set interval, each interval corresponds to a level, which is set based on the opinions of people with different vision and astigmatism.

[0106] This step directly compares the visual acuity and astigmatism data with the corresponding intervals to obtain the visual acuity grade and astigmatism grade.

[0107] Data mining: Use multiple linear regression functions to process the brightness, resolution, and contrast of display parameters to predict parameter requirements for the current period;

[0108] The steps of multiple linear regression function analysis are as follows:

[0109] Integrate the data features of brightness, resolution, and contrast in historical display parameters;

[0110] By organizing historical data, we constructed brightness prediction functions, resolution prediction functions, and contrast prediction functions. For example, as people age, their demand for brightness gradually increases. The basic function is a curve in childhood and is basically linear in other subsequent periods, but the slope of the function is different. The slope of the brightness function corresponding to adults is the lowest.

[0111] The brightness, resolution and contrast of the current time are calculated in sequence.

[0112] The brightness prediction function, resolution prediction function, and contrast prediction function are all set based on the most recent data so that they can be predicted by the functions.

[0113] Data correction: Based on the visual acuity level, astigmatism level, and age data, the corresponding preset display parameter interval set is searched. The preset display parameter interval set includes a display brightness interval, a display resolution interval, and a display contrast interval. The intersection of the three preset display parameter interval sets is taken to form the preset display parameter interval set.

[0114] This step is implemented by pre-building an interval library. The doctor formulates the corresponding display brightness interval, display resolution interval and display contrast interval based on age, vision and astigmatism. When in use, the corresponding interval can be directly searched. For example, the first preset display parameter interval set is found by age, which includes the first display brightness interval. The second preset display parameter interval set is found by astigmatism level, which includes the second display brightness interval. The third preset display parameter interval set is found by vision level, which includes the third display brightness interval. The intersection of the three display brightness intervals is taken to form the preset display brightness interval. And so on, the preset display resolution interval and the preset display contrast interval are obtained. Then the preset display brightness interval, the preset display resolution interval and the preset display contrast interval are summarized to obtain the preset display parameter interval set.

[0115] Comparison and judgment: Compare the parameter requirements predicted for the current period with the preset display parameter interval set;

[0116] The parameter requirements for the current time period are predicted to be the brightness calculated by the brightness prediction function, the resolution calculated by the resolution prediction function, and the contrast calculated by the contrast prediction function.

[0117] If the parameter requirements for the current time period are within the corresponding display data interval in the preset display parameter interval set, the parameters for the current time period are extracted. This indicates that the calculated data meets the requirements and can be directly applied.

[0118] If the parameter requirements for the current time period are predicted to be outside the corresponding display data interval in the preset display parameter interval set, the data closest to the parameter requirements for the current time period in the preset display parameter interval set is extracted. In this case, the calculated data does not meet the requirements, and the closest data is used;

[0119] Distance data is acquired and analyzed to determine the user's height and the distance between the user and the display device. The display screen size and center position are adjusted based on the user's height and distance. The extracted data, display screen size, and center position are aggregated to form a dimming strategy.

[0120] When calculating the user's height, the point cloud data collected by the depth camera is calibrated through the three-dimensional spatial coordinate system, and the coordinates of the user's head and feet are extracted using the human body key point detection algorithm to calculate the user's actual height. After obtaining the user's height, combined with the distance between the user and the depth camera, the horizontal distance between the depth camera and the user can be calculated, and then the distance between the user and the display device can be calculated.

[0121] For example: Figure 2As shown, the distance between the user (top of the head) and the depth camera is L, the actual height of the user is L1, and the horizontal distance between the depth camera and the user is calculated based on the inclination angle A of the depth camera (the angle between the depth camera and the vertical direction) as L1×sinA. Combined with the distance L2 between the depth camera and the display device, the distance between the user and the display device can be calculated as L1×sinA+L2.

[0122] The center of the screen is the position where the user's eyes are facing the display device, and is set to the user's actual height minus 0.1 meters, that is, L1-0.1.

[0123] This method makes user interaction more comfortable and the display effect better.

[0124] Infrared image data is captured and collected by an infrared camera at the scene of the interactive somatosensory display, and distance data is collected by a depth camera.

[0125] The steps for sequentially collecting and analyzing the user's distance data and infrared image data are as follows:

[0126] Distance data collection and analysis: Obtain distance data and adjust the focus of the infrared camera based on the distance data to collect infrared images;

[0127] The focus data of the infrared camera is adjusted dynamically using the existing automatic focusing algorithm to ensure the best imaging state. The infrared image collected after adjustment contains the complete image of the user. This step mainly reduces the proportion of the environment in the image and reduces environmental interference, thereby reducing the difficulty of subsequent analysis.

[0128] Feature enhancement: A multi-scale retinal cortex method based on Retinex theory is used to separate the infrared image reflection component and illumination component, highlighting the user's outline and key areas of somatosensory interaction;

[0129] This step primarily eliminates interference caused by uneven ambient light and enhances the contrast between the user's body contours and key areas of somatosensory interaction (such as joints, limb extremities, and hands). By analyzing image features at different scales, this step can effectively highlight the target subject and suppress the influence of background noise.

[0130] For scenarios with low external interference (enclosed environments), this step is not required.

[0131] Noise reduction: Use non-local mean filtering algorithm to remove noise from the image, while retaining edge information;

[0132] The non-local mean filtering algorithm can remove salt and pepper noise and Gaussian noise in the image. For closed environments, no processing is required. The non-local mean filtering algorithm estimates the current pixel value by searching for neighborhood blocks similar to the current pixel in the image and taking a weighted average of the similar blocks. While removing noise, it can retain image edge information, thereby facilitating the subsequent extraction of key features.

[0133] Normalization: Normalize the image pixel values to the range [0, 1] to unify the data standards. This step is a common solution in data processing and is common knowledge, so it will not be described in detail.

[0134] Edge extraction: Using the Canny edge detection algorithm, the system calculates the gradient amplitude and direction of local grayscale changes in the image to identify potential edge pixels. It then dynamically adjusts the high and low thresholds to ensure complete extraction of the true edges of the human body while suppressing false edges.

[0135] This step can accurately outline the user's body contours and provide key basic data for subsequent analysis;

[0136] Feature extraction: Use the deep learning-based OpenPose method to extract the human body outline and identify the somatosensory interaction area.

[0137] The OpenPose method based on deep learning is an existing model architecture based on deep learning. It can automatically extract key points of the human body, then connect the key points to form a complete human outline, mark the key interactive positions (head, hands, etc.), that is, identify the somatosensory interaction area. The collection area is mainly because the position will change during interaction. Therefore, a certain range is set instead of directly selecting the position. For example, a key position area is M, and the corresponding somatosensory interaction area is set to n, 3M≤n≤10M. This method can effectively reduce the collection of environmental data, thereby reducing the amount of data processed in subsequent analysis.

[0138] The present invention provides a somatosensory interactive display system based on a depth camera, which changes the existing fixed somatosensory interaction method, analyzes the user's situation and position, and formulates display parameters exclusive to the user. This personalized adaptive mechanism reduces user visual fatigue, improves interaction comfort, and enables users to extend the somatosensory interaction time, thereby effectively improving the display effect, and has good usage prospects.

[0139] 2.3. Re-analysis

[0140] After the somatosensory interaction area is divided, it is necessary to set the parameters of the depth camera in the depth camera matrix to realize the acquisition of the somatosensory interaction area and ensure stable recognition of the instructions. This process is based on the secondary analysis module.

[0141] Secondary analysis module: Analyzes the marked interaction area and dynamically sets the shooting parameters of each depth camera in the depth camera array based on the analysis results;

[0142] The analysis of the somatosensory interaction area of the marker is to find the location of the somatosensory interaction area of the marker, calculate the size of the area of the somatosensory interaction area of the marker, calculate the ratio of the area of each somatosensory interaction area of the marker to the area of the infrared image, and adjust the focal length data of each depth camera in the depth camera array according to the ratio and the focal length data of the infrared camera.

[0143] After identifying the somatosensory interaction areas, the pixel points corresponding to each interaction area are counted, and then the area of each area is calculated using the pixel counting method.

[0144] After calculating the area of each region, divide the calculated area by the total area to obtain the area ratio.

[0145] Using the optical imaging formula, the specific formula is , where is the focal length, u is the distance between the user and the depth camera, and v is the distance from the optical center of the lens to the image sensor, which is determined by the internal structure of the camera. Combined with the current focal length data of the infrared camera, the target focal length of each camera in the depth camera array is calculated. During the shooting process, the distributed control algorithm is used to synchronize the focal lengths of each camera to avoid image stitching misalignment or field of view overlap caused by focal length differences.

[0146] This setting can accurately control the shooting area and reduce the environmental data in the captured image, thus facilitating subsequent analysis and processing.

[0147] You can also directly convert the area of each somatosensory interaction area into the actual area and then calculate the focal length.

[0148] 2.4 Image Processing

[0149] After analyzing the focal length and completing the setting of all parameters, the somatosensory interactive display can be carried out. During the somatosensory interactive display, it is necessary to collect images of each somatosensory interactive area to determine whether the user has issued an instruction. This process is based on the image processing module.

[0150] Image processing module: Segments the images continuously captured by the depth camera array based on an image segmentation algorithm, performs feature recognition on the segmented graphics, calculates the graphic similarity between continuously captured images, and compares the graphic similarity with a preset similarity threshold. If the graphic similarity is less than or equal to the preset similarity threshold, no processing is performed. In this case, the user's instinctive or unconscious actions do not require further analysis.

[0151] For example, two images are captured continuously, each image is segmented to obtain a graph, and then the graphs are compared, that is, the graphs at the same position at different time points are compared.

[0152] If the graphic similarity is greater than the preset similarity threshold, semantic understanding technology is combined to identify the interaction instructions corresponding to the actions in the graphic.

[0153] Before segmenting the images continuously collected by the depth camera array, the images are first enhanced, then the noise is reduced on the enhanced images, and then the pixels of the images are normalized. When segmenting the images, the grayscale value segmentation method is used to segment the images of the user's somatosensory interaction areas.

[0154] The enhancement, noise reduction and normalization processing in this part are the same as the principles of data cleaning and conversion during data acquisition, so we will not go into details about them.

[0155] The steps to segment an image are as follows:

[0156] Feature extraction: obtain the grayscale value of each pixel in the image and construct the grayscale histogram of the image;

[0157] The specific formula for constructing the grayscale histogram is as follows:

[0158] ;

[0159] Where, is the number of pixels with gray value g, is the coordinate in the image The pixel grayscale value, the image size is C × D, is the Kronecker function, which is used to determine whether the pixel value is equal to g and count the total number of pixels with the same gray value.

[0160] In the output grayscale histogram, the horizontal axis is the grayscale value and the vertical axis is the number of pixels.

[0161] Threshold division: The grayscale histogram is clustered using the bimodal method to divide the grayscale value range into the interaction area interval and the background interval;

[0162] The bimodal method assumes that the image consists of foreground (interaction area) and background. If there are two obvious peaks in the histogram, the valley between the two peaks is selected as the segmentation threshold T to divide the grayscale value into the interaction area interval and background interval .

[0163] In a closed area, the background and the human body are more obvious and can be directly identified. In order to enhance the effect, relevant personnel can be arranged to wear specific colors, which is more effective.

[0164] Initial marking: Compare each pixel in the image with the interaction area and the background area, and mark the pixels in the interaction area;

[0165] The interactive area is marked as 1, and the background is not marked.

[0166] Matrix construction: Calculate the grayscale difference between each pixel in the image and its horizontally adjacent pixels, and construct a difference matrix based on the grayscale difference;

[0167] The formula for calculating the grayscale difference is as follows:

[0168] ;

[0169] Where: For coordinates The pixel and its horizontal right neighbor The absolute value of the grayscale difference;

[0170] The constructed difference matrix is C×(D-1).

[0171] Optimization processing: Use the Otsu algorithm to calculate the optimal difference threshold, and compare the size of the data in the difference matrix with the optimal difference threshold;

[0172] The Otsu algorithm automatically calculates the optimal threshold by maximizing the grayscale variance between the background and foreground.

[0173] If the data in the difference matrix are all less than or equal to the optimal difference threshold, no processing is done;

[0174] If there is data in the difference matrix that is greater than the optimal difference threshold, the data greater than the optimal difference threshold is marked, and the number of marked data of adjacent pixels is found (normally 8). If the number of marked data is less than or equal to the preset value (for example, set to 5), the pixel corresponding to the data is not marked. If the number of marked data is greater than the preset value, the pixel corresponding to the data is marked.

[0175] Graphic generation: The grayscale value of unmarked pixels is set to 0 to obtain the interactive part graphic. The interactive part graphic has high contrast and clearly separates the interactive part from the background, which is convenient for subsequent further analysis.

[0176] When performing feature recognition on the segmented graphics, the detection model is used to process the interactive part graphics and extract the geometric features of the interactive parts; then the difference in the geometric features of the graphics at two adjacent time points is calculated, and finally the difference is converted to calculate the graphic similarity;

[0177] There are many detection models, and all of them are existing models. For example, for hand detection, the MediaPipe Hands model is used.

[0178] When analyzing the hand, geometric features include angle calculation between fingers, palm shape features, hand posture, hand area, etc.

[0179] Before calculating the difference between the geometric features in the graphics of two adjacent time points, all the calculated geometric features are integrated into a multi-dimensional geometric feature vector.

[0180] Then, the corresponding elements in the multidimensional geometric feature vector of the real-time collected graphics and the last collected graphics are subtracted to calculate the difference vector of the two feature vectors, and then the conversion processing and similarity mapping are performed to calculate the graphic similarity.

[0181] For example: The geometric feature vector in the current graph is , the multidimensional geometric feature vector in the last collected graph is , the conversion factor is , the calculated graph similarity is Where is the graph similarity, is the maximum conversion difference, z is the number of elements in the geometric eigenvector, The conversion factor when calculating the difference for the qth element, is the qth element in the current graphic geometric feature vector, The qth element in the geometric feature vector of the last collected graphic is obtained. The greater the graphic similarity, the higher the possibility that the user has issued an instruction, which facilitates the analysis of whether to recognize the instruction, can effectively avoid invalid recognition, and can effectively reduce the amount of data processing.

[0182] For example, if the user's finger moves slightly, the calculated graphic similarity is low, and there is no need to recognize the interaction command.

[0183] The steps to identify the interactive instructions corresponding to the actions in the graphics are as follows:

[0184] Construct standard feature vector: set the standard feature vector of K groups of interaction instructions;

[0185] When constructing the standard feature vector, each interactive instruction uses multiple standard action sample graphics, performs statistical analysis on all sample feature vectors of each standard action sample graphic, and calculates the mean vector as the standard feature vector of the group of interactive instructions.

[0186] This step is similar to the prior art, but is more detailed.

[0187] Build a mapping library: associate standard interaction parts with interaction instructions;

[0188] Create a structured mapping library, associate the constructed K groups of standard feature vectors with the corresponding interaction instructions one by one, use data structures such as hash tables and database tables to store them, maintain the mapping library regularly, and update the data in the mapping library when new interaction instructions are added to ensure the accurate correspondence between the interaction instructions and the standard feature vectors.

[0189] Extract vector: take the geometric features of the next frame of graphics in the adjacent time as the feature vector;

[0190] The next frame in the adjacent moment is the currently acquired image. The geometric features have been extracted in the previous part, so this part can be directly applied.

[0191] Calculate distance: Calculate the distance between the feature vector and the data in the K groups of standard feature vectors, and calculate and summarize the comprehensive distance between the feature vector and the K groups of standard feature vectors;

[0192] Scoring: Normalize the K groups of comprehensive distances to obtain the K groups of proximity scores.

[0193] For example: The geometric feature vector in the current graph is , the standard feature vector of a certain instruction , the conversion factor is , the calculated closeness score is Where is the graph similarity, is the maximum value of the comprehensive distance of K groups, is the minimum value of the K-group comprehensive distance, z is the number of elements in the geometric feature vector, The conversion factor when calculating the distance for the qth element, is the qth element in the current graphic geometric feature vector, It is the qth element in the standard feature vector. The larger the graph closeness score is, the closer it is to the instruction.

[0194] Since the elements are different, they cannot be directly accumulated after direct subtraction, so conversion processing is required. The conversion is combined with weighted processing, that is, the conversion coefficient is the weight multiplied by the conversion value.

[0195] Decision processing: Get the maximum value among the K groups of proximity scores and compare the maximum value with the preset standard score;

[0196] If the maximum value is greater than or equal to the standard score, it means that the similarity between the current action and a certain standard action is high enough. The standard feature vector corresponding to the maximum value is searched in the mapping library, and the interaction instruction associated with the standard feature vector is output;

[0197] If the maximum value is less than the standard score, it is determined that no interactive instruction has been made and the display operation will not be performed.

[0198] The present invention provides a somatosensory interactive display system based on a depth camera, which uses infrared cameras to locate and partition users, then uses a depth camera array to focus on shooting the somatosensory interaction parts, and uses geometric features for comparison. It can realize the recognition of small-scale movements of the somatosensory interaction parts, thereby changing the existing method of using large-scale movements to achieve somatosensory interaction, greatly reducing the difficulty of somatosensory interaction, and greatly improving the accuracy of somatosensory interaction. The use effect is good, so this solution has good application prospects.

[0199] 2.5 Verification and Display

[0200] Because this system uses multiple parts to issue interactive instructions, there is a situation where multiple interactive instructions are collected at the same time. In this case, further analysis is required, and the analysis steps are based on the verification and display module.

[0201] Verification and display module: summarizes the interaction instructions, generates interaction strategies, and generates and displays results according to the interaction strategies.

[0202] The steps to summarize the interaction instructions and generate the interaction strategy are as follows:

[0203] Determine the number of interaction instructions. If the number of interaction instructions is 0, determine that the generated interaction strategy is to do nothing and maintain the current display state.

[0204] If there is one interactive instruction, the generated interactive strategy is determined to be the interactive instruction. In this case, there is no interference and it can be executed directly.

[0205] If there are multiple groups of interaction instructions generated, calculate the proportion of each interaction instruction in the total interaction instructions, compare the proportions, and obtain the maximum proportion. If the maximum proportion is greater than the set value, determine that the generated interaction strategy is the interaction instruction corresponding to the maximum proportion;

[0206] For example, if there are three groups of interaction instructions, 1, 1, and 3 times respectively, the calculated maximum proportion is 3 / (1+1+3)=0.6, and the set value is 0.5, then the interaction strategy executed is the interaction instruction corresponding to the maximum proportion.

[0207] If the maximum proportion is ≤ the set value, the generated interaction strategy is determined to be no processing.

[0208] For example, if there are three groups of interaction instructions, 1, 2, and 2 times respectively, the calculated maximum proportion is 2 / (1+2+2)=0.4, and the set value is 0.5, then the interaction strategy is to do nothing and retain the current display status.

[0209] Through the above complete steps, the system can reasonably generate corresponding interaction strategies based on the identified interaction instructions, achieve accurate and intelligent user interaction responses, improve user experience and system interaction performance, and thus improve the display effect.

[0210] The weight coefficient is determined using the coefficient of variation method, which is a method of assigning weights to each indicator based on the degree of variation between the current value of each evaluation indicator and the target value. If the numerical difference of an indicator is large, it can clearly distinguish the evaluated objects, indicating that the indicator has rich discrimination information, and thus the indicator should be given a larger weight. On the contrary, if the numerical difference of each evaluated object on a certain indicator is small, then the ability of this indicator to distinguish the evaluation objects is weak, and thus the indicator should be given a smaller weight. This method directly uses the information contained in each indicator to obtain the weight of the indicator through calculation, and therefore is objective.

[0211] The present invention provides a somatosensory interactive display system based on a depth camera, which collects data from multiple somatosensory interactive parts of the user, allowing the user to use multiple somatosensory interactive parts to implement the operation of the same interactive instruction, effectively avoiding the situation where some interactive instructions cannot be recognized, and correspondingly effectively increasing the number of interactive instructions, making the somatosensory interaction richer, thereby greatly increasing the user's desire to explore, thereby improving the display effect, and has good application prospects.

[0212] In the application, the several formulas involved are all calculated by taking their numerical values after removing the dimensions, and the formulas are established by collecting a large amount of data and performing software simulation to obtain a formula for the most recent real situation. Some coefficients or weights in the formulas are set by technical personnel in this field according to actual conditions, so they will not be elaborated here.

[0213] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. Those skilled in the art will appreciate that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution.

[0214] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, and may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment as needed.

[0215] The above is only a specific implementation method of the present application, but the scope of protection of the present application is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed in this application, which should be covered by the scope of protection of the present application.

Claims

1. The somatosensory interactive display system based on depth camera is characterized by: The system includes: Data retrieval module: extracts the user's most recent N historical somatosensory interaction data from the database based on the data retrieval algorithm and preprocesses the data; Parameter analysis module: Analyzes pre-processed data to obtain dimming strategies, collects and analyzes user distance data and infrared image data in sequence, and marks interactive areas; Secondary analysis module: Analyzes the marked interaction area and dynamically sets the shooting parameters of each depth camera in the depth camera array based on the analysis results; Image processing module: Segments the images continuously captured by the depth camera array based on an image segmentation algorithm, performs feature recognition on the segmented graphics, calculates the graphic similarity between consecutively captured images, and compares the graphic similarity with a preset similarity threshold. If the graphic similarity is less than or equal to the preset similarity threshold, no processing is performed; otherwise, semantic understanding technology is combined to identify the interactive instructions corresponding to the actions in the graphics. Verification and display module: summarizes the interaction instructions, generates interaction strategies, and generates and displays results according to the interaction strategies.

2. The somatosensory interactive display system based on a depth camera according to claim 1, characterized in that: The historical somatosensory interaction data includes age data, user vision detection data, display parameters, infrared image data and distance data. The historical somatosensory interaction data is retrieved from the database through distributed indexing technology, the user vision detection data is entered into the database by the user, the infrared image data is captured and collected by the infrared camera at the somatosensory interaction display site, and the distance data is collected by the depth camera.

3. The somatosensory interactive display system based on a depth camera according to claim 2, characterized in that: The steps to analyze the pre-processed data and obtain the dimming strategy are as follows: Feature analysis: Obtain the vision and astigmatism data from the user's vision test data, and compare the vision and astigmatism data with the standard range to obtain the vision grade and astigmatism grade; Data mining: Use multiple linear regression functions to process the brightness, resolution, and contrast of display parameters to predict parameter requirements for the current period; Data correction: Based on the visual acuity level, astigmatism level and age data, the corresponding preset display parameter interval set is searched. The preset display parameter interval set includes display brightness interval, display resolution interval and display contrast interval; and taking the intersection of the three sets of preset display parameter interval sets to form a preset display parameter interval set; Comparison and judgment: Compare the parameter requirements predicted for the current period with the preset display parameter interval set; If the parameter requirements for the current time period are predicted to be within the corresponding display data interval in the preset display parameter interval set, then the parameters for the current time period are extracted; If the parameter demand for the predicted current period is outside the corresponding display data interval in the preset display parameter interval set, extract the data in the preset display parameter interval set that is closest to the parameter demand for the predicted current period; Distance data is acquired and analyzed to determine user height and distance, where the user distance is the distance between the user and the display device. The display screen size and center position are adjusted based on the user height and distance. The extracted data, display screen size, and center position are aggregated to form a dimming strategy.

4. The somatosensory interactive display system based on a depth camera according to claim 3, characterized in that: The steps for sequentially collecting and analyzing the user's distance data and infrared image data are as follows: Distance data collection and analysis: Obtain distance data and adjust the focus of the infrared camera based on the distance data to collect infrared images; Feature enhancement: A multi-scale retinal cortex method based on Retinex theory is used to separate the infrared image reflection component and illumination component, highlighting the user's outline and key areas of somatosensory interaction; Noise reduction: Use non-local mean filtering algorithm to remove noise from the image, while retaining edge information; Normalization processing: normalize the image pixel values to the [0, 1] interval to unify the data standards; Edge extraction: Use the Canny edge detection algorithm to analyze the dynamic changes in local grayscale of the image, adjust the high and low thresholds, and extract the edge contour of the user's body; Feature extraction: Use the deep learning-based OpenPose method to extract the human body outline and identify the somatosensory interaction area.

5. The somatosensory interactive display system based on a depth camera according to claim 4, characterized in that: The analysis of the somatosensory interaction area of the marker is to find the location of the somatosensory interaction area of the marker, calculate the size of the area of the somatosensory interaction area of the marker, calculate the ratio of the area of each somatosensory interaction area of the marker to the area of the infrared image, and adjust the focal length data of each depth camera in the depth camera array according to the ratio and the focal length data of the infrared camera.

6. The somatosensory interactive display system based on a depth camera according to claim 5, characterized in that: Before segmenting the images continuously collected by the depth camera array, the images are first enhanced, then the noise is reduced on the enhanced images, and then the pixels of the images are normalized. When segmenting the images, the grayscale value segmentation method is used to segment the images of the user's somatosensory interaction areas.

7. The somatosensory interactive display system based on a depth camera according to claim 6, characterized in that: The steps to segment an image are as follows: Feature extraction: obtain the grayscale value of each pixel in the image and construct the grayscale histogram of the image; Threshold division: The grayscale histogram is clustered using the bimodal method to divide the grayscale value range into the interaction area interval and the background interval; Initial marking: Compare each pixel in the image with the interaction area and the background area, and mark the pixels in the interaction area; Matrix construction: Calculate the grayscale difference between each pixel in the image and its horizontally adjacent pixels, and construct a difference matrix based on the grayscale difference; Optimization processing: Use the Otsu algorithm to calculate the optimal difference threshold, and compare the size of the data in the difference matrix with the optimal difference threshold; If the data in the difference matrix are all less than or equal to the optimal difference threshold, no processing is done; If there is data in the difference matrix that is greater than the optimal difference threshold, then mark the data that is greater than the optimal difference threshold, and find the amount of marked data of adjacent pixels. If the amount of marked data is less than or equal to the preset value, then the pixel corresponding to the data is not marked. If the amount of marked data is greater than the preset value, then mark the pixel corresponding to the data. Graphics generation: Set the grayscale value of unmarked pixels to 0 to obtain the interactive part graphics.

8. The somatosensory interactive display system based on a depth camera according to claim 7, characterized in that: When performing feature recognition on the segmented graphics, the detection model is used to process the interactive part graphics and extract the geometric features of the interactive parts; then the difference in the geometric features of the graphics at two adjacent time points is calculated, and finally the difference is converted to calculate the graphic similarity.

9. The somatosensory interactive display system based on a depth camera according to claim 8, characterized in that: The steps to identify the interactive instructions corresponding to the actions in the graphics are as follows: Construct standard feature vector: set the standard feature vector of K groups of interaction instructions; Build a mapping library: associate standard interaction parts with interaction instructions; Extract vector: take the geometric features of the next frame of graphics in the adjacent time as the feature vector; Calculate distance: Calculate the distance between the feature vector and the data in the K groups of standard feature vectors, and calculate and summarize the comprehensive distance between the feature vector and the K groups of standard feature vectors; Scoring: Normalize the K groups of comprehensive distances to obtain the K groups of proximity scores; Decision processing: Get the maximum value among the K groups of proximity scores and compare the maximum value with the preset standard score; If the maximum value ≥ the standard score, output the interaction instruction corresponding to the maximum value; If the maximum value is less than the standard score, it is determined that no interactive instruction has been made.

10. The somatosensory interactive display system based on a depth camera according to claim 9, characterized in that: The steps to summarize the interaction instructions and generate the interaction strategy are as follows: Determine the number of interaction instructions. If the number of interaction instructions is 0, determine that the generated interaction strategy is not to be processed. If the interaction instruction is 1 group, the generated interaction strategy is determined to be the interaction instruction; If there are multiple groups of interaction instructions generated, calculate the proportion of each interaction instruction in the total interaction instructions, compare the proportions, and obtain the maximum proportion. If the maximum proportion is greater than the set value, determine that the generated interaction strategy is the interaction instruction corresponding to the maximum proportion; If the maximum proportion is ≤ the set value, the generated interaction strategy is determined to be no processing.

Citation Information

Patent Citations

  • Somatosensory interaction data processing method, device and somatosensory interaction equipment

    CN113596353B

  • Multi-target interactive image segmentation method and device

    CN104899849A

  • Combined optimization method of visual comfort and deep sense of a stereoscopic image

    CN105959679A