A stage interaction system based on image and radar data
By combining the information acquisition module of RGB cameras and LiDAR lidar and the dual-stream network processing module, the existing stage interaction system has been solved in terms of detection range and accuracy, and the stage effect of high-precision stage interaction and virtual and real combination without additional body sensing equipment is achieved.
Patent Information
- Application Number
- CN202011609683.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-30
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2040-12-30
AI Technical Summary
The existing stage interaction system has shortcomings in detection range and accuracy, especially when the performer is facing away from the camera or facing the camera, and the somatosensory device may affect the aesthetics of the movements.
Using an information acquisition module combining RGB cameras and LiDAR lidar, the collected data is extracted and pose recognized through the dual-stream network processing module to generate a pose set and control the stage effect.
It realizes a high detection range and accuracy without additional body sensing equipment, can capture the performer's gestures, postures and expressions, and accurately control the stage effect, achieving a stage interaction that combines virtual and real.
Smart Images

Figure CN112598742B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to human-computer interaction technology, and in particular relates to a stage interaction system based on image and radar data. Background Art
[0002] As the role of stage design in enhancing the performance effect of artistic performances becomes more and more prominent, the traditional conventional stage with scenery can no longer meet the needs of performers to interact with the stage scene, so many stage interaction systems have emerged. Existing stage interaction systems usually use 3D somatosensory cameras and somatosensory devices on the performers to capture the performers' gestures, or use activated radar to detect the performers' touch points to achieve interaction.
[0003] The stage interaction system based on 3D somatosensory cameras and somatosensory devices to capture the performer's posture is usually implemented by using the somatosensory device worn on the performer to detect the range of motion of the performer's limbs and output a detection signal; after the 3D somatosensory camera detects the somatosensory signal, it transmits the information to the processing device and the control device to realize stage interaction. The advantage of this method is that it can detect information including gestures, postures, expressions, etc., and only requires the performer to wear the device, which is low cost. The disadvantage is that the detection range of the system is limited. When the performer is facing away or sideways to the 3D somatosensory camera, it is difficult for the interactive system to accurately obtain the performer's posture data, and thus it is impossible to accurately control the switching of the stage. In addition, the somatosensory device will also affect the beauty of the performer's movements to a certain extent.
[0004] The stage interaction system based on laser radar detection of performers' touch points is implemented by using a laser radar detection device to detect touch actions on the surface through the formed scanning surface, thereby locating the position information of one or more touch points, and controlling the stage switching through the position information of the touch points; the advantages are strong anti-interference ability, insensitivity to ambient light, and not restricted by the shape and boundaries of the screen; the disadvantages are small effective detection area, the effective detection range is only a semicircle with a radius of 3m, and the number of detections is limited. Even if multiple laser radars are set up for use, it is still not suitable for interactive control of medium and large stages. Secondly, it can only simply detect the touch point position, but cannot detect the performer's posture information. At the same time, there are situations such as false touches, and it is difficult to achieve precise control of the interactive switching of the stage. Summary of the invention
[0005] In view of the shortcomings of the prior art, the present invention proposes a stage interaction system based on images and radar data, which does not require performers to wear additional somatosensory equipment, and can also increase the area of the effective detection range and improve the system control accuracy.
[0006] A stage interaction system based on image and radar data comprises an information acquisition module, a processing module and a control module.
[0007] The information acquisition module uses an RGB camera and a LiDAR laser radar to respectively capture images of the stage and perform radar detection, obtains the performer information on the stage in real time, generates a radar point cloud map with the same frequency as the RGB image using the data obtained by radar detection, and then transmits the RGB image and the radar cloud point map as field data to the processing module together.
[0008] Preferably, the RGB camera and LiDAR laser radar are arranged in front of the stage.
[0009] The processing module includes a posture generation unit and a posture recognition unit. The posture generation unit generates a posture set by learning the data collected by the information collection module through a dual-stream network; the posture recognition unit finds the preset posture corresponding to the on-site posture in the posture set and sends the recognition result to the control module.
[0010] The dual-stream network includes feature extraction, feature aggregation and posture generation modules. The feature extraction module first processes the radar point cloud image separately to obtain a bird's-eye view, and then uses the VGG-16 network to extract the features of the bird's-eye view and the RGB image to obtain a feature map. The feature aggregation module selects the position of the performer in the image based on the feature information in the feature map. The posture generation module fits the anchor posture into the target area based on the position information of the performer in the image to generate a posture set.
[0011] After receiving the gesture recognition result from the processing module, the control module determines the performer's intention according to the recognition result and controls the stage to change according to the performer's gesture.
[0012] The change of the stage screen content includes the change of the virtual environment image and the background screen display content.
[0013] The present invention has the following beneficial effects:
[0014] 1. Use RGB cameras and LiDAR laser radars simultaneously to obtain the current performer's posture. The detection range is large, and the performer does not need to wear additional somatosensory devices; when the performer changes different postures, the corresponding gestures, postures, expressions and other information can be captured.
[0015] 2. The processing module uses a neural network to extract features, aggregate and process the collected on-site data to identify the performer's posture information, and then uses the control module to control the stage changes according to the performer's intentions, thereby realizing the interaction between the performer's posture and the stage effects, achieving a stage effect that combines the real and the virtual. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 Schematic diagram of the stage structure in the embodiment;
[0017] Figure 2It is the schematic diagram of the interactive system;
[0018] Figure 3 Schematic diagram of the dual-stream network structure. DETAILED DESCRIPTION
[0019] The present invention will be further explained below with reference to the accompanying drawings;
[0020] like Figure 1 As shown, an RGB camera and a LiDAR laser radar are set in front of the stage, and the collected data are sent to the control host, in which the data is processed and recognized, and the changes in stage effects are controlled.
[0021] like Figure 2 As shown, a stage interaction system based on image and radar data includes an information acquisition module, a processing module and a control module.
[0022] The information acquisition module uses an RGB camera and a LiDAR laser radar to respectively capture images of the stage and perform radar detection, obtains the performer information on the stage in real time, generates a radar point cloud map with the same frequency as the RGB image using the data obtained by radar detection, and then transmits the RGB image and the radar cloud point map as field data to the processing module together.
[0023] The processing module includes a posture generation unit and a posture recognition unit. The posture generation unit generates a posture set by learning the data collected by the information collection module through a dual-stream network. The dual-stream network includes feature extraction, feature aggregation and posture generation modules. The feature extraction module first processes the radar point cloud image separately to obtain a 6-channel bird's-eye view, and then uses two parallel VGG-16 networks to simultaneously extract the features of the bird's-eye view and the RGB image to obtain two feature maps. The feature aggregation module selects the position of the performer in the image based on the feature information in the feature map. In order to generate the performer's posture, the neural network usually needs to detect the joints of the character first and then group each joint, or first select the position where the posture needs to be generated in the input data frame as the region proposal through the region proposal algorithm, and then generate the posture in the target area. The two-stream neural network adopts the method of first detecting the position and then generating the posture. In the posture fitting process, the 5-dimensional information of each posture joint point is regressed at the same time, including the coordinates of the 2D posture and the 3D posture; the anchor box pre-generated by the two-stream neural network is projected onto the feature map view obtained by the feature extraction module, and then the RolAlign algorithm is used twice. The first RolAlign algorithm obtains the 3D target area, and the anchor posture will fit the task in this area; the second RolAlign algorithm obtains the posture details and posture score, and the recommended area obtained after cropping is the position of the performer in the image. The posture generation module fits the anchor posture into the target area according to the position information of the performer in the image to generate a posture set.
[0024] Before formal use, a large amount of field data needs to be collected and input into the two-stream network to train and optimize it and adjust the network parameters. RPN , anchor pose loss L cls , 2D pose refinement loss L 2D and 3D pose refinement loss L 3D Four losses as indicators for optimizing the two-stream network L total :
[0025] L total =L RPN +L cls +L 2D +L 3D
[0026] RPN loss L RPN Used to optimize the location selection of the target area. This part of the loss includes region regression and target classification. Region regression is to find the location of the target box in the input feature map, which is used to optimize the process of the feature aggregation module outputting a series of target areas on the feature map. Target classification, namely anchor pose classification, is to determine whether the target area box is the target object. The optimization goal of this loss is to enable the network to find the appropriate target area in multiple target areas.
[0027]
[0028] where p i represents the probability that the i-th prediction box is the foreground, is the label, when the i-th prediction box is the foreground is 1, otherwise it is 0; t i Represents the 4 position parameters of the prediction box, is the parameter of the calibration frame, N cls and N reg is the size of a small batch in one training, L cls is the anchor pose loss function, L reg is the regression loss function;
[0029] Anchor pose loss L cls The selection of anchor poses is used to optimize the selection of anchor poses, including the distinction between foreground and background and the use of similarity scores to assign the best anchor poses. The similarity calculation formula of anchor poses is:
[0030]
[0031] where a k,j represents the position of joint j of the kth anchor pose, g j represents the real annotated joint node j, J is the total number of joint nodes, and K is the total number of anchor words.
[0032] 2D pose refinement loss L 2D To optimize the final 2D anchor pose, the 2D regression increment predicted by the two-stream network is added to the anchor pose to obtain a set of final 2D pose anchor predictions P 2D :
[0033]
[0034] Where P 2D is the probability that the final predicted prediction box is the foreground, N fg is the number of foregrounds, T 2D is the true annotation corresponding to each foreground target area. i is the parameter factor, smooth_ll is the smoothing function.
[0035] 3D pose refinement loss L 3D Similar to the 2D pose refinement loss, the regression increment is added to the 3D anchor pose to obtain the final 3D pose P 3D However, since the two-stream neural network does not use 3D annotated data, the network projects the 3D posture into the 2D image space for calculation:
[0036]
[0037] N fg is the number of foregrounds, T 3D is the true annotation corresponding to the target area of each foreground, and the pr function is the projection function. 3D Projection into 2D space.
[0038] The gesture recognition unit is pre-set with the gesture information of the performer. After receiving the gesture set generated by the gesture generation unit, the gesture of the performer in the gesture set is recognized. If the recognized gesture matches the gesture in the stored preset gesture, the matched preset gesture is sent to the control unit as the recognition result.
[0039] The control module stores preset scenes corresponding to the performer's preset postures. After receiving the posture recognition results from the processing module, the control module determines the performer's intention based on the recognition results, loads the corresponding animation scene onto the stage screen, and controls the switching of the stage scene layout, including lighting, music, stage special effects, stage lifting, etc.
Claims
1. A stage interaction system based on image and radar data, characterized in that: It includes an information collection module, a processing module and a control module; The information acquisition module uses an RGB camera and a LiDAR laser radar to respectively capture images of the stage and perform radar detection, obtain information about performers on the stage in real time, generate a radar point cloud map with the same frequency as the RGB image using the data obtained by radar detection, and transmit the RGB image and the radar cloud point map as on-site data to the processing module together; The processing module includes a posture generation unit and a posture recognition unit. The posture generation unit generates a posture set by learning the data collected by the information collection module through a dual-stream network; the posture recognition unit finds the preset posture corresponding to the on-site posture in the posture set and sends the recognition result to the control module; The dual-stream network includes feature extraction, feature aggregation and posture generation modules; The feature extraction module first processes the radar point cloud image separately to obtain the bird's-eye view image, and then uses the VGG-16 network to extract the features of the bird's-eye view image and RGB image to obtain the feature map; The feature aggregation module selects the position of the performer in the image based on the feature information in the feature map; The pose generation module fits the anchor pose to the target area according to the position information of the performer in the image and generates a pose set; the RPN loss L RPN , anchor pose loss L cls , 2D pose refinement loss L 2D and 3D pose refinement loss L 3D The sum of the four losses is used as the optimization index L total Optimize the dual-stream network; where p i represents the probability that the i-th prediction box is the foreground, is the label, when the i-th prediction box is the foreground is 1, otherwise it is 0; t i Represents the 4 position parameters of the prediction box, is the parameter of the calibration frame, N cls and N reg is the size of a small batch in one training, L cls is the anchor pose loss function, L reg is the regression loss function; Anchor pose loss L cls The selection of anchor poses is used to optimize the selection of anchor poses, including the distinction between foreground and background and the use of similarity scores to assign the best anchor poses. The similarity calculation formula of anchor poses is: where a k,j represents the position of joint j of the kth anchor pose, g j represents the joint node j of the real annotation, J is the total number of joint nodes, and K is the total number of anchor words; 2D pose refinement loss L 2D To optimize the final 2D anchor pose, the 2D regression increment predicted by the two-stream network is added to the anchor pose to obtain a set of final 2D pose anchor predictions P 2D : where p 2D is the probability that the final predicted prediction box is the foreground, N fg is the number of foregrounds, T 2D The actual annotation corresponding to each foreground target area; l i is the parameter factor, smooth_ll is the smoothing function; 3D pose refinement loss L 3D Add the regression increment to the 3D anchor pose to obtain the final 3D pose P 3D : N fg is the number of foregrounds, T 3D is the true annotation corresponding to the target area of each foreground, and the pr function is the projection function. 3D Projection into 2D space; After receiving the gesture recognition result from the processing module, the control module determines the performer's intention according to the recognition result and controls the stage to change according to the performer's gesture.
2. A stage interactive system based on image and radar data as claimed in claim 1, characterized in that: The RGB camera and LiDAR laser radar are arranged in front of the stage.
3. A stage interactive system based on image and radar data as claimed in claim 1, characterized in that: The change of the stage screen content includes the change of the virtual environment image and the background screen display content.
Citation Information
Patent Citations
Radar human body posture recognition method and system based on multi-class spectrogram fusion and hierarchical learning
CN111368930A
Obstacle recognition method and device, computer equipment and storage medium
CN111797650A