Visual interaction method, device, equipment and storage medium for LED display screen
By performing lighting equalization and group posture recognition on the audience area image data, combining high dynamic range feature extraction and bidirectional visual language attention decoding on the LED display screen, using manifold distance audience behavior clustering and viewing angle difference impact processing, and optimizing the display parameters based on the potential field model and model prediction control, solving the problem that traditional LED display interactive systems cannot be dynamically adjusted, and achieving efficient audience behavior recognition and display effect optimization.
Patent Information
- Application Number
- CN202510171989.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-02-17
AI Technical Summary
Traditional LED display interaction systems cannot dynamically adjust according to the audience's real-time reaction, resulting in a large gap between the display effect and the audience's needs. Especially in the complex stadium environment, it is difficult to accurately capture the audience's group behavior and understand their interactive intentions.
By performing lighting equalization and group posture recognition on the audience area image data, combining high dynamic range feature extraction and bidirectional visual language attention decoding on the LED display screen, the manifold distance audience behavior clustering and viewing angle difference influence processing are used, and the display parameters are optimized based on the potential field model and model prediction control.
It realizes accurate identification of audience group behavior and semantic understanding of display content, accurately captures the audience's interactive intentions, dynamically adjusts the display effect, ensures the consistency of visual experience in different viewing areas, and significantly improves the interactive intelligence level of LED display screens.
Smart Images

Figure CN119625652B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of LED display screens, and in particular to a visual interaction method, device, equipment and storage medium for an LED display screen. Background Art
[0002] With the development trend of intelligent and interactive large sports venues, LED display screens, as important information display and interactive media, have a significant impact on the audience's viewing experience due to their interactive performance. Traditional LED display screen interactive systems mainly rely on preset display modes and fixed parameter configurations, and cannot be dynamically adjusted according to the audience's real-time reactions, resulting in a large gap between display effects and audience needs.
[0003] In complex sports stadium environments, the dynamics of audience group behavior, the diversity of display content, and the differences in viewing conditions pose huge challenges to the interactive control of LED display screens. Existing interactive systems lack the ability to accurately identify and understand audience group behavior, and are unable to accurately capture the audience's interactive intentions. It is also difficult to effectively optimize display parameters for viewing conditions in different areas. In addition, LED display screens face complex lighting environments and wide viewing angles in large outdoor venues. Traditional display parameter control methods are difficult to simultaneously meet the requirements of high dynamic range display and uniform display. Especially in audience interaction scenarios, how to achieve real-time optimization and precise control of display parameters, and how to ensure consistency of visual experience in different viewing areas have become key issues that need to be addressed. Summary of the invention
[0004] The present application provides a visual interaction method, device, equipment and storage medium for an LED display screen. The present invention significantly improves the interactive intelligence level of the LED display screen and effectively solves the technical difficulties of traditional methods in group behavior recognition.
[0005] The first aspect of the present application provides a visual interaction method of an LED display screen, and the visual interaction method of the LED display screen comprises:
[0006] Perform illumination balance and group posture recognition on the audience area image data to obtain audience group posture heat map data;
[0007] Perform high dynamic range feature extraction and bidirectional visual language attention decoding on the LED display screen to obtain the semantic area data of the display content;
[0008] Performing manifold distance audience behavior clustering and viewing angle difference impact processing on the audience group posture heat map data and the display content semantic area data to obtain display response trigger data;
[0009] Based on the display response trigger data, a potential field model and a model prediction control are performed on the display parameters to obtain a display parameter control instruction;
[0010] Performing interaction effect analysis based on the display parameter control instructions to obtain a display screen parameter optimization strategy;
[0011] The display screen parameter optimization strategy is subjected to interactive mode analysis and multi-region display uniformity evaluation to obtain display screen interactive configuration data.
[0012] A second aspect of the present application provides a visual interaction device for an LED display screen, the visual interaction device for an LED display screen comprising:
[0013] The recognition module is used to perform illumination balance and group posture recognition on the audience area image data to obtain audience group posture heat map data;
[0014] The extraction module is used to perform high dynamic range feature extraction and bidirectional visual language attention decoding on the LED display screen to obtain the semantic area data of the display content;
[0015] A processing module, used for performing manifold distance audience behavior clustering and viewing angle difference impact processing on the audience group posture heat map data and the display content semantic area data to obtain display response trigger data;
[0016] A control module, used to perform potential field model and model prediction control on display parameters based on the display response trigger data to obtain display parameter control instructions;
[0017] An analysis module, used to analyze the interaction effect based on the display parameter control instruction to obtain a display screen parameter optimization strategy;
[0018] The evaluation module is used to perform interactive mode analysis and multi-region display uniformity evaluation on the display screen parameter optimization strategy to obtain display screen interactive configuration data.
[0019] The third aspect of the present application provides an electronic device, comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor calls the instructions in the memory so that the electronic device executes the above-mentioned visual interaction method of the LED display screen.
[0020] The fourth aspect of the present application provides a computer-readable storage medium, in which instructions are stored. When the computer-readable storage medium is run on a computer, the computer executes the above-mentioned visual interaction method of the LED display screen.
[0021] Compared with the prior art, the present application has the following beneficial effects: by combining deep learning technology with multimodal interactive control, a complete set of visual interaction methods for LED display screens is established. This method uses a dual-branch Transformer module and a multi-layer visual feature encoding network to achieve accurate recognition of audience group behavior and semantic understanding of display content; through manifold distance audience behavior clustering and perspective difference impact processing, the audience's interactive intention is accurately captured; the display parameter optimization method based on potential field model and model predictive control realizes dynamic adjustment of display effect; through interactive effect analysis and multi-region display uniformity evaluation, the consistency of visual experience in different viewing areas is guaranteed. The present invention significantly improves the interactive intelligence level of LED display screens, effectively solves the technical difficulties of traditional methods in group behavior recognition, interactive intention understanding, display parameter optimization and visual experience uniformity, and provides a new technical solution for the intelligent interaction of LED display screens in large venues. The present invention shows good technical effects in practical applications, effectively improving the audience's interactive experience and viewing experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0023] The structures, proportions, sizes, etc. illustrated in the drawings of this specification are only used to match the contents disclosed in the specification so as to facilitate understanding and reading by persons familiar with this technology. They are not used to limit the conditions under which the present invention can be implemented, and therefore have no substantive technical significance. Any structural modification, change in proportion or adjustment of size, without affecting the effects and purposes that can be achieved by the present invention, should still fall within the scope of the technical contents disclosed by the present invention.
[0024] Figure 1 It is a flow chart of a visual interaction method of an LED display screen provided by an embodiment of the present invention;
[0025] Figure 2 It is a schematic block diagram of the structure of a visual interaction device of an LED display screen provided by an embodiment of the present invention;
[0026] Figure 3 It is a schematic block diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0027] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0028] The flowcharts shown in the accompanying drawings are only examples and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may also be decomposed, combined or partially merged, so the actual execution order may change according to actual conditions.
[0029] It should also be understood that the terms used in this application specification are only for the purpose of describing specific embodiments and are not intended to limit the application. As used in this application specification and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.
[0030] It should be further understood that the term "and / or" used in the specification and appended claims of this application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations. Figure 1 , an embodiment of the visual interaction method of the LED display screen in the embodiment of the present application includes:
[0031] Step 100: Perform illumination balance and group posture recognition on the audience area image data to obtain audience group posture heat map data;
[0032] It is understandable that the execution subject of the present application can be a visual interaction device of an LED display screen, or a terminal or a server, which is not limited here. The present application embodiment is described by taking a server as the execution subject as an example.
[0033] Specifically, multi-scale Gaussian filtering is performed on the image data of the audience area, and the image is blurred by Gaussian kernels of different scales to generate multi-layer image pyramid data. Local contrast calculation is performed based on the multi-layer image pyramid data to obtain regional light intensity distribution data. Highlighting the local brightness difference of the image helps to identify details under different lighting conditions, and in complex scenes, it can eliminate the interference caused by uneven lighting. Nonlinear mapping and histogram equalization are performed on the regional light intensity distribution data. Nonlinear mapping enhances the details of the dark and bright parts of the image by adjusting the brightness and contrast of the image, while histogram equalization optimizes the image quality by balancing the overall brightness distribution of the image, ensuring that the image after illumination equalization can obtain a good display effect under various lighting environments. Feature extraction is performed on the illumination-equalized image data to extract basic feature data. The basic feature data is feature enhanced through the multi-layer encoding layer of the dual-branch Transformer module. The dual-branch Transformer module adopts a structure that combines the Transformer layer with the hole convolution layer. The Transformer layer can capture the global contextual relationship in the image and strengthen the correlation between features, while the atrous convolution layer can effectively improve the efficiency of feature extraction by expanding the receptive field. After feature enhancement, enhanced feature data is obtained. Key point detection and spatial coordinate mapping are performed based on the enhanced feature data. The joints in the image are accurately located by the deep learning model, and the key point data of the audience posture obtained include head position data, torso direction data and arm movement data, reflecting the audience's posture, movement and position relative to the display screen. The posture key point data is subjected to radial distortion correction and perspective transformation. The radial distortion correction is to correct the image deviation caused by lens distortion, and the perspective transformation is used to correct the spatial distortion caused by the perspective problem, so that the posture data obtained is more accurate and close to the actual spatial distribution. The spatial density distribution matrix is constructed based on the corrected posture data, and the posture heat map data of the audience group is calculated by kernel density estimation. The kernel density calculation can obtain the audience density at each position in the space by estimating the density of the audience distribution in the space, reflecting the spatial distribution of the audience group. At the same time, combined with posture data, the audience's activity level is calculated. Through this information, the audience group posture heat map data finally obtained can accurately reflect the audience's distribution density and interactive activity.
[0034] Step 200: extract high dynamic range features and perform bidirectional visual language attention decoding on the LED display screen to obtain semantic area data of display content;
[0035] Specifically, the LED display screen is converted from RGB to HSV color space. Each pixel in the RGB color space is represented by the intensity values of the three channels of red, green and blue, while the HSV color space describes the color through the three parameters of hue, saturation and brightness, which is more suitable for color processing and analysis. In the HSV color space, the brightness (brightness) component of the image is separated from the color (hue and saturation), which can better adjust the illumination and contrast. The image is decomposed at multiple scales using wavelet transform to obtain a high dynamic range feature image. Through wavelet transform, the image is decomposed into low-frequency, medium-frequency and high-frequency components. Each part carries different spatial information. The low-frequency component represents the basic outline and main structure of the image, the medium-frequency component reflects the regional characteristics, and the high-frequency component contains the detailed information in the image. The high dynamic range feature image is then input into a three-layer cascade convolutional network for feature hierarchical extraction, and features at different levels in the image are gradually extracted. The first layer uses 64 3×3 convolution kernels and ReLU activation functions to extract basic texture features. The ReLU activation function helps capture edge and texture information in the image. The second layer uses 128 5×5 convolution kernels and parameterized ReLU (PReLU) activation function. PReLU can adaptively adjust the slope of the negative semi-axis, so as to effectively process the regional structural features in the image and enhance the learning ability of the network. The third layer uses 256 7×7 convolution kernels and LeakyReLU activation function with a negative slope of 0.2 to extract global semantic features, which can capture information at a larger image range and semantic level. After the extraction of these convolution layers, three-level convolution features are obtained, which contain visual information at different levels of the image. For the three-level convolution features, spatial attention mechanism and channel attention mechanism are used to process them so as to give different attention weights to different features. The spatial attention mechanism adjusts each region in the feature map according to the importance of spatial position, while the channel attention mechanism focuses on the weight distribution of different channel features. Through these two attention mechanisms, the most representative and valuable features in the image are focused on to obtain feature attention weight data. The weighted feature data is input into the bidirectional visual language attention module to enhance the association between visual information and semantic information. In the bidirectional visual-language attention module, the visual-to-language attention submodule calculates the mapping relationship between visual features and language semantics through an 8-head attention mechanism and a 1024-dimensional query-key-value matrix. Through the multi-head attention mechanism, the system can learn different visual semantic relationships in parallel on multiple subspaces. The language-to-visual attention submodule fuses semantic information through a cross-attention mechanism based on cosine similarity and a 512-dimensional context vector to capture the cross-modal connection between language features and visual features. After processing by these two submodules, the obtained bidirectional attention output feature map integrates multiple information from vision and language. The bidirectional attention output feature map is reconstructed through a four-layer deconvolution network.The function of the deconvolution network is to gradually restore the spatial resolution of the feature map and reconstruct the feature map in the process. Each layer of the deconvolution uses different numbers of channels and convolution kernels to restore the details and semantic information of the image layer by layer. The first layer uses a 4×4 deconvolution kernel with 256 channels and a stride of 2, the second layer uses a 4×4 deconvolution kernel with 128 channels and a stride of 2, the third layer uses a 4×4 deconvolution kernel with 64 channels and a stride of 2, and the last layer uses a 4×4 deconvolution kernel with 32 channels and a stride of 2. Through these deconvolution layers, the high-resolution feature map of the image is gradually restored. Multi-scale feature fusion is performed on the reconstructed feature map to effectively merge features from different scales and improve the feature expression capability. The fused feature map is upsampled 16 times and its resolution is restored from 32×32 to the original 512×512 resolution using the bicubic interpolation kernel function to better restore the details and semantic structure of the image and obtain pixel-level feature map data. Semantic segmentation is performed based on the pixel-level feature map data to obtain the semantic area data of the displayed content. The process of semantic segmentation classifies each pixel in the image into different region categories, obtains region category labels and interactive attribute annotations, and reflects the function and attribute information of each region in the displayed content.
[0036] Step 300, performing manifold distance audience behavior clustering and perspective difference impact processing on audience group posture heat map data and display content semantic area data to obtain display response trigger data;
[0037] It should be noted that the feature extraction of the audience group posture heat map data is performed to obtain the group feature vector, which includes the audience's spatial distribution characteristics and the group's activity characteristics. The spatial distribution characteristics mainly describe the audience's position distribution in space, including information such as aggregation degree and density distribution, while the activity characteristics reflect the activeness of the audience's behavior, such as the audience's line of sight concentration and action frequency. The semantic area data of the displayed content is analyzed for regional attributes to obtain the regional feature vector. The regional feature vector includes regional category labels and interactive attribute tags. The regional category labels indicate the functional categories of different areas on the display screen, such as advertising areas, information display areas, interactive areas, etc., while the interactive attribute tags record the interactive attributes of the area, such as whether touch, voice control, gesture recognition, etc. are supported. Regional attribute analysis helps to understand the response type of each area during the interaction process and its relationship with the audience's behavior. Based on the calculation of the group feature vector and the regional feature vector, similarity measurement is performed to obtain the group-region association matrix, which reflects the degree of association between the audience group and each area of the display screen, and indicates how the content of different areas attracts the attention of different groups. By calculating the manifold distance of the group-region association matrix, the audience behavior clustering data is obtained. Manifold distance calculation is a nonlinear dimensionality reduction method that maps complex behavior data to a low-dimensional space for cluster analysis. The audience behavior clustering data obtained includes the behavior type of the group (such as concentrated viewing, dispersed viewing, interaction, etc.) and the activity score of the group. The activity score helps to determine the current participation of the audience group, and then provides a basis for interactive response. Perspective deviation calculation is performed based on the audience seat layout diagram. Consider the perspective differences caused by different audiences due to their different positions, including horizontal perspective deviation and vertical perspective deviation. Perspective differences will directly affect the audience's perception of the displayed content, especially in the case of large display screens or multi-area displays, the audience will see different pictures due to different seat angles. The calculation of perspective deviation accurately compensates for this difference to ensure that all audiences can get similar visual experience at different perspectives. The audience behavior clustering data is compensated according to the perspective difference data to obtain perspective correction data. Perspective correction is to adjust the audience's behavior pattern by compensating for the perspective deviation to make it consistent with the actual visual perception. The perspective correction data is feature fused with the regional feature vector to obtain interactive feature data. Response decision analysis is performed based on the interactive feature data to determine the response mode of the display screen. The display screen should respond based on the audience's behavior and interaction characteristics. The response includes an interaction trigger signal and response type information. The interaction trigger signal indicates that a specific behavior of the audience (such as gesture, touch, eye focus, etc.) triggers a response of the display screen, while the response type information specifically describes how the display screen should give feedback to the audience (such as changes in display content, brightness adjustment, area highlighting, etc.).
[0038] Step 400: Based on the display response trigger data, the display parameters are subjected to potential field model and model prediction control to obtain display parameter control instructions;
[0039] Specifically, a potential field function is constructed according to the display content semantic area data and the display response trigger data to obtain the initial potential energy distribution data, including the brightness potential energy and the contrast potential energy. The purpose of constructing the potential field model is to simulate the energy distribution of different areas on the display screen. The brightness potential energy reflects the brightness distribution of different areas, while the contrast potential energy describes the contrast change of each area. The initial potential energy distribution data is gradient calculated to obtain the display parameter change trend data. Through the gradient calculation, the change direction of the brightness and contrast is obtained, indicating how the brightness and contrast should be adjusted to respond to the needs of the audience in the current display state. The brightness change direction indicates that the brightness needs to be increased or decreased at a certain moment, while the contrast change direction indicates how to adjust the contrast in different areas to optimize the visual effect. Through calculation, the adjustment direction of the display parameters is predicted. Based on the display parameter change trend data, a state space model is constructed to describe the state evolution process of the display screen under different control conditions. It can predict the change trend of the display parameters in the future time period and obtain the brightness prediction value and the contrast prediction value. Through the prediction value, the future display state is estimated according to the current state, and the display is adjusted according to the prediction value. Display constraint analysis is performed on the display state prediction data to ensure that the adjustment of display parameters does not exceed the preset range, thereby avoiding visual discomfort or technical failure caused by over-adjustment. The parameter constraint data obtained by the constraint analysis includes brightness range constraints and contrast range constraints. Dynamic range calibration is performed based on the parameter constraint data. According to the characteristics and actual needs of the display screen, the actual output values of brightness and contrast are adjusted to meet the constraint requirements, and display parameter calibration data, including brightness calibration parameters and contrast calibration parameters, are obtained. Model predictive control calculation is performed based on the display parameter calibration data. Model predictive control is an optimization control method that uses the mathematical model of the system to predict the future state and calculates a multi-step control sequence based on this. Through model predictive control, a series of optimal control actions are predicted, and the display parameters are gradually adjusted according to these actions. The obtained multi-step control sequence includes a brightness control sequence and a contrast control sequence, each sequence indicating how the brightness and contrast should be adjusted in multiple time steps in the future to achieve the best display effect. Parameter optimization processing is performed according to the multi-step control sequence to obtain the optimal control parameters. Through the optimization algorithm, the optimal brightness and contrast control parameters are selected under the premise of satisfying the display constraints to achieve the most ideal display effect. The optimal control parameters include an optimal brightness parameter and an optimal contrast parameter. Based on the optimal control parameters, corresponding display parameter control instructions are generated, including a brightness control instruction and a contrast control instruction.
[0040] Step 500: Analyze the interaction effect based on the display parameter control instruction to obtain a display screen parameter optimization strategy;
[0041] Specifically, the display response trigger data is subjected to a time series correlation analysis to obtain the interactive response time series data. By analyzing the time relationship between the trigger signal and the response, the trigger delay data and the response duration data are obtained. The trigger delay data describes the time delay of the interaction between the audience and the display screen, while the response duration data indicates the duration of the display response. The display parameter control instructions are subjected to a control effect analysis to evaluate the effects of brightness and contrast adjustment. Through this analysis, the parameter control effect data is obtained, including the brightness control effect data and the contrast control effect data. The brightness control effect data reflects whether the brightness adjustment achieves the expected visual effect and whether the appropriate brightness level is provided in different scenes; while the contrast control effect data evaluates whether the contrast adjustment effectively improves the displayed details and enhances the layering and clarity of the image. According to the audience area division map, the interactive response time series data is regionally mapped. According to the audience's viewing angle position, the interactive response time series data of different areas are divided to obtain the regional interactive analysis data. The regional interactive analysis data includes the interactive response distribution data of different viewing areas, which can reflect the interactive performance of the display area. The regional interactive analysis data is subjected to viewing angle weighting processing. The viewing angle weight data includes horizontal viewing angle weight and vertical viewing angle weight, which reflect the difference in interactive response under different viewing angles. The horizontal viewing angle weight and vertical viewing angle weight indicate the difference in the degree of interactive response and display effect of the audience when viewing the display screen from different angles. For example, some angles result in poor audience experience due to the viewing angle limitation of the display screen or the uneven distribution of screen brightness. By weighting the viewing angle weight data, these viewing angle differences can be compensated, and the parameter control effect data can be compensated and calculated to obtain regional compensation data. The regional compensation data includes brightness compensation parameters and contrast compensation parameters, which represent the optimized values of brightness and contrast after adjusting the weights of different viewing angles. Based on the regional compensation data, multi-region display quality analysis is performed to evaluate the overall performance of the display effect. Regional display quality data is obtained, including display clarity data and color reproduction data. Display clarity data reflects the detail performance of different regions, including edge sharpness, image clarity, etc.; color reproduction data reflects the ability of the display screen to reproduce colors in different regions, including color vividness, accuracy, etc. The regional display quality data and interactive response timing data are comprehensively scored to obtain interactive experience scoring data. The interactive experience scoring data includes the display quality score and the interactive response score. The display quality score indicates the display effect after compensation and optimization, while the interactive response score evaluates the immediacy and smoothness of the interaction. Based on the scoring data, a parameter optimization strategy for the display is generated, including brightness optimization rules and contrast optimization rules. These optimization rules provide specific parameter settings for the adjustment of the display, ensuring that the best visual effects and user experience can be provided in a variety of interactive scenarios, ultimately achieving highly accurate interactive experience optimization.
[0042] Step 600: Perform interactive mode analysis and multi-region display uniformity evaluation on the display screen parameter optimization strategy to obtain display screen interactive configuration data.
[0043] Specifically, the display parameter optimization strategy is analyzed in time series by the first time series sliding window to obtain the first optimization rule sequence data. Through the time series analysis, the change trend of the optimization strategy in different time periods is identified, and the rule sequence is generated to reflect the time series characteristics of the display parameter adjustment. The display parameter optimization strategy is analyzed in time series by the second time series sliding window to obtain the second optimization rule sequence data. By comparing and combining the results of the two analyses, the first optimization rule sequence data and the second optimization rule sequence data are combined to generate comprehensive optimization rule sequence data. Based on the optimization rule sequence data, the first-level interaction rule graph and the second-level interaction rule graph are constructed. The first-level interaction rule graph describes the association relationship between single interaction rules. These rules involve the adjustment of single display parameters such as brightness and contrast. By analyzing the mutual relationship between these rules, the best single rule adjustment path is found. The second-level interaction rule graph describes the association relationship between combined interaction rules, which refers to how different display parameters and interactive actions are combined to affect the final display effect. By fusing the first-level and second-level interaction rule graphs, more complex and accurate interaction mode association data are obtained. Based on the interaction pattern association data, the rule association is clustered by the first directed graph clustering algorithm to obtain the first interaction pattern category. Cluster analysis can reveal the similarity and association between different interaction rules, so as to classify them into corresponding interaction pattern categories. The hierarchy of the rules is clustered by the second directed graph clustering algorithm to obtain the second interaction pattern category. By combining the first interaction pattern category and the second interaction pattern category, the interaction pattern category data is obtained. Based on the interaction pattern category data, the display area is divided into blocks by the adaptive grid division algorithm to obtain the first area division data. The division method is automatically adjusted according to the characteristics of the display area to meet the display requirements of different areas. The display area is divided into blocks by the audience density perception algorithm to obtain the second area division data. The distribution density of the audience will affect the division method of the display area, so as to ensure the best display effect in the high-density area. After weighted fusion of the first area division data and the second area division data, the area division data is obtained. The uniformity measurement calculation is performed on the area division data to obtain the local display uniformity data. The second uniformity measurement calculation is performed to obtain the global display uniformity data. The local display uniformity data is multi-scale fused with the global display uniformity data to obtain the display uniformity data. Based on the display uniformity data, the local area compensation parameters are calculated by the first compensation model to obtain the first compensation data, which reflects how to adjust and optimize according to the display effect of the local area. At the same time, the global area is compensated by the second compensation model to obtain the second compensation data. After the two are nonlinearly weighted, the uniformity compensation data is obtained. The uniformity compensation data and the interactive mode category data are fused.The first fusion network is used to perform feature-level fusion, and the uniformity compensation data is fused with the interactive mode category data to obtain the first control strategy data. The second fusion network is used to perform decision-level fusion, and these data are further integrated to obtain the second control strategy data. According to actual needs, the first control strategy data and the second control strategy data are adaptively selected to generate the most suitable interactive control strategy data. Based on the interactive control strategy data, the first configuration data and the second configuration data are generated, wherein the first configuration data controls the dynamic adjustment of the display parameters to ensure that the display effect can respond to the needs of the audience in real time; the second configuration data controls the real-time switching of the interactive rules to optimize the interactive experience. The two parts of the configuration data are combined to finally obtain the interactive configuration data of the display screen.
[0044] In the embodiment of the present application, a complete set of visual interaction methods for LED display screens is established by combining deep learning technology with multimodal interactive control. The method uses a dual-branch Transformer module and a multi-layer visual feature encoding network to achieve accurate recognition of audience group behavior and semantic understanding of display content; accurately captures the audience's interactive intentions through manifold distance audience behavior clustering and perspective difference impact processing; based on the potential field model and model predictive control, the display parameter optimization method realizes the dynamic adjustment of the display effect; through interactive effect analysis and multi-region display uniformity evaluation, the consistency of visual experience in different viewing areas is guaranteed. The present invention significantly improves the interactive intelligence level of LED display screens, effectively solves the technical difficulties of traditional methods in group behavior recognition, interactive intention understanding, display parameter optimization and visual experience uniformity, and provides a new technical solution for the intelligent interaction of LED display screens in large venues. The present invention shows good technical effects in practical applications, effectively improving the audience's interactive experience and viewing experience.
[0045] In a specific embodiment, the process of executing step 100 may specifically include the following steps:
[0046] Perform multi-scale Gaussian filtering on the audience area image data to obtain multi-layer image pyramid data, and perform local contrast calculation based on the multi-layer image pyramid data to obtain regional light intensity distribution data;
[0047] Perform nonlinear mapping and histogram equalization on the regional light intensity distribution data to obtain illumination balanced image data, and perform feature extraction on the illumination balanced image data to obtain basic feature data;
[0048] The basic feature data is enhanced by a multi-layer coding layer of a dual-branch Transformer module to obtain enhanced feature data, wherein the multi-layer coding layer includes a Transformer layer and a hole convolution layer;
[0049] Based on the enhanced feature data, key point detection and spatial coordinate mapping are performed to obtain audience posture key point data, which includes head position data, torso direction data and arm movement data;
[0050] Perform radial distortion correction and perspective transformation on the key point data of the audience's posture to obtain corrected posture data;
[0051] A spatial density distribution matrix is constructed based on the corrected posture data, and the audience group posture heat map data is obtained through kernel density calculation. The audience group posture heat map data includes the spatial distribution density data and activity level data of the audience group.
[0052] Specifically, multi-scale Gaussian filtering is performed on the audience area image data to obtain multi-layer image pyramid data. The image is blurred multiple times using Gaussian filters to obtain image representations at different scales. Assume that the original image is , after multi-scale Gaussian filtering, the image pyramid is obtained It is expressed by the following formula:
[0053] ;
[0054] in, Represents the pyramid The image data of the layer, is a standard deviation The Gaussian filter kernel is used, and the symbol * represents the convolution operation. The image size of each layer is gradually reduced to form a pyramid structure from low resolution to high resolution. Local contrast calculation is performed based on the multi-layer image pyramid data to obtain regional light intensity distribution data. Local contrast calculation depends on the brightness change of the local image window. For example, the following formula is used to calculate the local contrast: ;
[0055] in, is the position in the image The pixel value of and Represent the maximum and minimum pixel values of the local window of the image, respectively. Represents the local contrast of the position. By calculating the local contrast, the light intensity distribution data of each area in the image is obtained. Nonlinear mapping and histogram equalization are performed on the regional light intensity distribution data to obtain the illumination balanced image data. Nonlinear mapping uses the sigmoid function or other nonlinear functions to transform the pixel values to improve the contrast and brightness distribution of the image. For example, using the sigmoid function: ;
[0056] in, is the image after nonlinear mapping, is the slope of the mapping, is the offset, set to the average brightness value of the image. After nonlinear mapping, the brightness distribution of the image is improved. Perform histogram equalization to adjust the grayscale distribution of the image to make the contrast of the image more uniform. Calculate the cumulative histogram of the image ,in Represents the pixel value in the image The frequency distribution of . By using the cumulative histogram, the transformation function is calculated , so that the pixel values of the transformed image are evenly distributed. Use the transformation function to map the pixel values of the original image to obtain a lighting balanced image. Feature extraction is performed on the lighting balanced image to obtain basic feature data. Feature extraction usually involves edge detection, corner detection or other texture-based features. For example, algorithms such as SIFT (scale-invariant feature transform) and SURF (speeded up robust features) are used. The basic feature data is feature enhanced through the multi-layer encoding layer of the dual-branch Transformer module to obtain enhanced feature data. The dual-branch Transformer module contains two independent network paths, one of which is used to process image features and the other is used to process other semantic information. The multi-layer encoding layer consists of multiple Transformer layers and hole convolution layers to capture global dependencies and local feature information in the image. Assume that the input basic feature data is , enhanced feature data after processing by the encoding layer It is expressed by the following formula: ;
[0057] Among them, Transformer It is a Transformer layer operation based on the self-attention mechanism, DilatedConv It is a dilated convolutional layer operation that can expand the receptive field and extract more extensive feature information. Based on the enhanced feature data, key point detection and spatial coordinate mapping are performed to obtain the key point data of the audience's posture. The goal of key point detection is to identify the key parts of the audience from the image, such as the position of the head, torso, and arms. Through deep learning methods, such as the posture estimation algorithm based on convolutional neural networks, the key points in the image are identified. Assume that the detected key point data is , which includes the head position , Torso Direction and arm movements . Perform radial distortion correction and perspective transformation to eliminate image deformation caused by camera distortion or perspective difference. The radial distortion correction uses the following formula: ;
[0058] in, is the corrected key point, is the coefficient of radial distortion, is the distance from the key point to the center of the image. The perspective transformation is processed through the perspective matrix to convert the image from one perspective to another, thereby ensuring the accuracy of the posture data. The spatial density distribution matrix is constructed based on the corrected posture data, and the audience group posture heat map data is obtained through kernel density estimation. Spatial density distribution matrix Describe the audience density in different areas of the image. Use the kernel density estimation formula: ;
[0059] in, is the number of viewers, It is The posture coordinates of each audience member, is the kernel function, using Gaussian kernel function. The obtained heat map Contains data on the spatial distribution density and activity level of the audience group.
[0060] In a specific embodiment, the process of executing step 200 may specifically include the following steps:
[0061] The LED display screen image is converted from RGB to HSV color space, and multi-scale decomposition is performed through wavelet transform to obtain a high dynamic range feature image, which includes low-frequency component data, medium-frequency component data and high-frequency component data;
[0062] The high dynamic range feature image is input into a three-layer cascade convolutional network for hierarchical feature extraction. The first layer of the three-layer cascade convolutional network uses 64 3×3 convolution kernels and ReLU activation function to extract basic texture features. The second layer uses 128 5×5 convolution kernels and parameterized ReLU activation function to extract regional structure features. The third layer uses 256 7×7 convolution kernels and LeakyReLU activation function with a negative slope of 0.2 to extract global semantic features, and obtain three-layer convolution features.
[0063] The spatial attention mechanism and channel attention mechanism are processed on the three-level convolutional features to obtain the feature attention weight data;
[0064] The feature attention weight data is input into the bidirectional visual-language attention module, where the visual-to-language attention submodule uses an 8-head attention mechanism and a 1024-dimensional query-key-value matrix to calculate the feature mapping relationship, and the language-to-visual attention submodule uses a cross-attention mechanism based on cosine similarity and a 512-dimensional context vector to fuse semantic information to obtain a bidirectional attention output feature map;
[0065] The bidirectional attention output feature map is reconstructed through a four-layer deconvolution network. The first layer of the four-layer deconvolution network uses a 4×4 deconvolution kernel with 256 channels and a step size of 2, the second layer uses a 4×4 deconvolution kernel with 128 channels and a step size of 2, the third layer uses a 4×4 deconvolution kernel with 64 channels and a step size of 2, and the fourth layer uses a 4×4 deconvolution kernel with 32 channels and a step size of 2 to obtain the reconstructed feature map;
[0066] The reconstructed feature map is subjected to multi-scale feature fusion to obtain a fused feature map, which is then upsampled 16 times. The bicubic interpolation kernel function is used to restore the feature resolution from 32×32 to the original 512×512 resolution to obtain pixel-level feature mapping data.
[0067] Semantic segmentation is performed based on pixel-level feature mapping data to obtain display content semantic region data, which includes region category labels and interactive attribute annotations.
[0068] Specifically, the LED display screen image is converted from RGB to HSV color space. HSV color space can more effectively separate the brightness (value V) and color information (hue H and saturation S) of the image. RGB to HSV conversion is achieved through the following formula:
[0069] ;
[0070] ;
[0071] ;
[0072] in, are the values of the red, green, and blue channels of the image in the RGB space, It is the color. is the saturation, is brightness. The image is decomposed into multiple scales by wavelet transform to obtain high dynamic range (HDR) feature image. Wavelet transform can decompose the image into sub-images of different frequencies and obtain image features at different scales. High dynamic range feature image includes low frequency component, medium frequency component and high frequency component. Assume that the original image is , then the wavelet transform is decomposed by the following formula:
[0073] ;
[0074] ;
[0075] ;
[0076] in, Represents low-frequency components, which contain the general structural information of the image; Represents the mid-frequency component, which contains the medium detail information of the image; Represents high-frequency components, including image details and edge information. The high dynamic range feature image is input into a three-layer cascade convolutional network for feature hierarchical extraction. The first layer of the three-layer cascade convolutional network uses 64 The convolution kernel is combined with the ReLU activation function to extract the basic texture features of the image. The mathematical representation of the convolution operation is: ;
[0077] in, is the convolution kernel, is the low-frequency component image, is the bias term, is the output feature map. The purpose of the first convolution layer is to extract the basic texture features of the image, such as edges and simple shapes. The second level uses 128 The convolution kernel is combined with the parameterized ReLU activation function (PReLU) to extract regional structural features. The formula of PReLU is as follows: ;
[0078] in, is a training parameter used to control the slope of the negative part. The second convolution operation is: ;
[0079] in, is the convolution kernel of the second layer, is the output feature map of the first layer. The second layer extracts the structural information of the region, such as the boundary of the object or the distribution characteristics of the region. The third layer uses 256 The convolution kernel is combined with the LeakyReLU activation function with a negative slope (0.2) to extract global semantic features. The formula of LeakyReLU is as follows: ;
[0080] in, , which is used to control the slope of the negative part to avoid the gradient vanishing problem. The third convolution operation is: ;
[0081] in, is the convolution kernel of the third layer, is the output feature map of the second layer, is the output global semantic feature map. The third layer extracts global image features, such as large-scale object shapes and semantic information. The three-level convolutional features are processed by spatial attention mechanism and channel attention mechanism to enhance the representation ability of the features. The spatial attention mechanism adjusts the weight of each position by paying attention to different spatial positions of the feature map, while the channel attention mechanism adjusts the channel weight according to the different importance of each channel. Assume that the spatial attention weight is , the channel attention weight is , then the feature attention weight data is obtained by the following formula:
[0082] ;
[0083] Where ⊙ represents the element-by-element multiplication operation. In this way, the network can pay more attention to important spatial regions and channels, improving the effectiveness of feature representation. The feature attention weight data is input into the bidirectional visual-language attention module. This module includes a visual-to-language attention submodule and a language-to-visual attention submodule. The visual-to-language attention mechanism calculates the feature mapping relationship through an 8-head attention mechanism and a 1024-dimensional query-key-value matrix. The calculation formula of the multi-head attention mechanism is as follows: ;
[0084] in, is the query matrix, is the key matrix, is the value matrix, is the dimension of the key matrix. The language-to-vision attention submodule adopts a cross-attention mechanism based on cosine similarity and fuses semantic information through a 512-dimensional context vector. Through the interaction of these two submodules, a bidirectional attention output feature map is obtained. The bidirectional attention output feature map is reconstructed through a four-layer deconvolution network. Each layer of the deconvolution network uses a convolution kernel with a different number of channels and stride to gradually reconstruct a high-resolution representation of the image. The operation of the four-layer deconvolution network is as follows: ;
[0085] in, It is the reconstructed feature map after deconvolution. DeConv represents the deconvolution operation. The first layer uses 256 channels and a step size of 2. Deconvolution kernel, the second layer uses 128 channels and a stride of 2 Deconvolution kernel, the third layer uses 64 channels and a stride of 2 Deconvolution kernel, the fourth layer uses 32 channels and a stride of 2 Deconvolution kernel. Perform multi-scale feature fusion on the reconstructed feature map to obtain a fused feature map. Merge features of different scales to obtain richer feature information. Upsample the fused feature map by 16 times, and use the bicubic interpolation kernel function to restore the feature map from 32×32 resolution to the original 512×512 resolution. Finally, pixel-level feature mapping data is obtained. Based on these pixel-level feature mapping data, semantic segmentation is performed to obtain the semantic region data of the display content. The image is segmented into multiple semantic regions, each of which represents a part of the image with specific meaning. The obtained display content semantic region data includes region category labels and interactive attribute annotations.
[0086] In a specific embodiment, the process of executing step 300 may specifically include the following steps:
[0087] Extract features from the audience group posture heat map data to obtain a group feature vector, which includes audience spatial distribution features and group activity features;
[0088] Performing regional attribute analysis on the displayed content semantic region data to obtain a regional feature vector, which includes a regional category label and an interactive attribute tag;
[0089] Based on the group feature vector and the region feature vector, similarity measurement is calculated to obtain the group-region association matrix, and the group-region association matrix is subjected to manifold distance calculation to obtain audience behavior clustering data, which includes group behavior type and activity level score;
[0090] Calculating the viewing angle deviation of the audience behavior clustering data based on the audience seat layout diagram to obtain viewing angle difference data, the viewing angle difference data including horizontal viewing angle deviation and vertical viewing angle deviation;
[0091] Performing compensation processing on the audience behavior clustering data according to the perspective difference data to obtain perspective correction data, and performing feature fusion on the perspective correction data and the regional feature vector to obtain interactive feature data;
[0092] A response decision analysis is performed based on the interaction feature data to obtain display response trigger data, which includes an interaction trigger signal and response type information.
[0093] Specifically, key spatial information and behavioral features are extracted from the heat map data. The posture heat map of the audience group reflects the density, attention distribution and behavioral status of the audience in different areas. For the extraction of the group's spatial distribution characteristics, the spatial characteristics of the group are obtained by calculating the density distribution in the heat map. Assume that the heat map data is ,in Indicates at location The audience density at the location. The spatial distribution characteristics of the group are calculated by the following formula: ;
[0094] This formula obtains the overall spatial distribution characteristics by summing all the positions in the heat map. This characteristic quantity is used to describe the concentration of the group in the display area, such as the concentration of the audience in front of or in the middle of the stage. By analyzing the local peaks in the heat map, the activity hotspots of the group are extracted to obtain the group activity characteristics. The group activity is measured by calculating the proportion of high-density areas in the heat map. For example, a threshold is defined To filter out active areas, calculate their ratio to the total heat map area: ;
[0095] in, is an indicator function, when Greater than threshold When , the function value is 1, otherwise it is 0. Reflects the activity level of the group. Analyze the semantic area data of the displayed content. The semantic area data of the displayed content includes the attribute information of each area on the display screen, such as the category label and interactive attribute mark of the area. For example, one area represents the advertising area, and another area is the interactive area. The audience's interaction in different areas may be different. Assume that the semantic data of the area is , which contains the region category label and interactive attribute tags . Regional feature vector By encoding these attributes. For example, use one-hot encoding to convert the category label into a vector and numerically process the interactive attribute mark. The regional feature vector is represented as: ;
[0096] This vector reflects the type of region and its interactive characteristics. Based on the group feature vector and the regional feature vector, the similarity measure is calculated to obtain the group-region association matrix. The similarity measure uses Euclidean distance or cosine similarity to measure the similarity between group features and regional features. For example, the cosine similarity between the group feature vector and the regional feature vector is calculated by the following formula: ;
[0097] in, is the dot product of the group feature vector and the regional feature vector, and is their modulus. By calculating the similarity between all groups and regions, we get the group-region association matrix , each element of the matrix represents the similarity between a group and a region. The manifold distance is calculated on the group-region association matrix to eliminate behavioral biases in different regions or perspectives, ensuring better clustering analysis of audience behavior. Manifold learning techniques, such as local linear embedding (LLE) or ISOMAP, help reveal the nonlinear relationship between group behavior and regional characteristics. Assuming that the local linear embedding algorithm is chosen, the manifold distance is calculated through the following steps: between each group and region pair, the similarity of the local neighborhood is calculated; based on the similarity, an adjacency graph is constructed; and the manifold embedding is calculated using the local linear embedding algorithm. In this way, group behavior clustering data is obtained, which includes group behavior types (e.g., concentrated, dispersed) and activity scores. The perspective deviation of the audience behavior clustering data is calculated based on the audience seating layout diagram. The audience's behavior is affected by their viewing angle, and the perspective deviation of the group behavior is adjusted according to the audience's seat position. Assume that the layout of the seats is , the viewing angle deviation of each seat position is calculated by the following formula:
[0098] ;
[0099] ;
[0100] in, are the coordinates of the seat, is the location of the display screen. is the distance from the audience to the display screen. Through calculation, the horizontal and vertical viewing angle deviations of each seat position are obtained. Based on the viewing angle difference data, the audience behavior clustering data is compensated. For example, for audiences with large viewing angle deviations, their behavior scores are adjusted so that their behavior is more consistent with the actual observed regional characteristics. The compensated audience behavior data is expressed as: ;
[0101] in, is the uncorrected audience behavior data. is a weight coefficient that indicates the degree of influence of perspective bias on behavior. is the viewing angle deviation. The audience behavior data after viewing angle correction is fused with the regional feature vector to obtain the interactive feature data. The interactive feature data contains the fusion information of group behavior and regional characteristics, reflecting the response of different regions to different group behaviors. Through the interactive feature data, response decision analysis is performed to determine the display content and interaction mode of the display screen. Assume that the output of the response decision analysis is , including interaction trigger signals and response type information:
[0102] ;
[0103] in, It is a decision function based on interactive feature data, which is trained using a machine learning model. At this stage, by analyzing the relationship between group behavior and regional features, appropriate responses are generated, such as changing display content, triggering interactive signals, etc., and finally display response trigger data is obtained.
[0104] In a specific embodiment, the process of executing step 400 may specifically include the following steps:
[0105] A potential field function is constructed according to the display content semantic area data and the display response trigger data to obtain initial potential energy distribution data, the initial potential energy distribution data including brightness potential energy and contrast potential energy;
[0106] Performing gradient calculation on the initial potential energy distribution data to obtain display parameter change trend data, the display parameter change trend data including brightness change direction and contrast change direction;
[0107] A state space model is constructed based on the display parameter change trend data to obtain display state prediction data, the display state prediction data includes a brightness prediction value and a contrast prediction value, and a display constraint analysis is performed on the display state prediction data to obtain parameter constraint data, the parameter constraint data includes a brightness range constraint and a contrast range constraint;
[0108] Performing dynamic range calibration on the parameter constraint data to obtain display parameter calibration data, the display parameter calibration data including brightness calibration parameters and contrast calibration parameters, and performing model predictive control calculation based on the display parameter calibration data to obtain a multi-step control sequence, the multi-step control sequence including a brightness control sequence and a contrast control sequence;
[0109] Parameter optimization processing is performed according to a multi-step control sequence to obtain optimal control parameters, which include optimal brightness parameters and optimal contrast parameters, and display parameter control instructions are generated based on the optimal control parameters, which include brightness control instructions and contrast control instructions.
[0110] Specifically, a potential field function is constructed based on the semantic area data of the display content and the display response trigger data. The potential field function is a mathematical model that describes energy distribution and interaction forces. It effectively combines the semantic area of the display content and the interactive behavior of the audience, thereby dynamically adjusting the display parameters in different areas. The potential field function is used to describe the energy distribution in space and can reflect how the brightness and contrast requirements change in different locations or areas. Given the semantic area data of the display content and display response trigger data , by constructing a potential field function to simulate the energy distribution. Assume that each point in the display area There is a potential energy value related to brightness and contrast. The following potential field function is used to describe the potential energy distribution of the display area: ;
[0111] in, is in position The total potential energy at It is the semantic area data of the display content at that location, reflecting the display content requirements of that location. It is the response trigger data, which indicates the feedback of the audience interaction. and is a weight coefficient used to balance the influence of semantic requirements and interactive response. Through this potential field function, the initial potential energy distribution of each area of the display screen is calculated. The initial potential energy distribution is divided into two parts: brightness potential energy and contrast potential energy, which represent the energy distribution of brightness and contrast in the display area respectively. and contrast potential They are expressed by the following formulas respectively:
[0112] ;
[0113] ;
[0114] in, and are weight coefficients used to adjust brightness and contrast respectively. Gradient calculation is performed on the initial potential energy distribution data to calculate the change trend of display parameters, especially the change direction of brightness and contrast. The gradient describes the direction and magnitude of potential energy change in the display area. For the potential energy function , whose gradient is calculated using the following formula: ;
[0115] in, and The potential energy is and The partial derivative in the direction indicates the direction and magnitude of the potential energy change. Based on this gradient, the direction of change of brightness and contrast is obtained:
[0116] ;
[0117] ;
[0118] These gradient values are used to guide the adjustment of display parameters to ensure that the change trend of brightness and contrast meets the regional requirements and the interactive feedback of the audience. The state space model is constructed based on the display parameter change trend data. The state space model is a mathematical model used to describe the process of system state changing over time. It can predict the brightness and contrast values of the display state at different time points. The state space model is expressed in the following form:
[0119] ;
[0120] ;
[0121] in, Indicates that the system is at time The state vector contains the brightness and contrast values, are control inputs (e.g. feedback from audience behavior), is the state transfer matrix, which indicates how the system state is transferred from the previous moment to the current moment. is the control matrix, which represents how the input affects the system state, is the output matrix, representing the output of the system (such as brightness and contrast), and It is a direct input matrix. By constructing a state space model, the display state is predicted to obtain the brightness prediction value and the contrast prediction value. Display constraint analysis is performed on the display state prediction data to ensure that the display parameters (brightness and contrast) do not exceed the preset range. For example, the brightness value must be within a certain range, usually [0,255] (assuming that the brightness is represented by 8 bits). Similarly, the contrast is within a reasonable range, such as [0,100]. Through constraint analysis, parameter constraint data is obtained. Dynamic range calibration is performed on the parameter constraint data. By adjusting the brightness and contrast of the display, the display effect in different areas is ensured to be consistent. Assuming that the predicted values and constraint data of brightness and contrast are obtained, the final brightness calibration parameters and contrast calibration parameters are calculated using the following dynamic range calibration formula:
[0122] ;
[0123] ;
[0124] The calibrated brightness and contrast parameters ensure that the display effect of the display meets the predetermined quality standards. Through the model predictive control (MPC) method, a multi-step control sequence is calculated to optimize the control strategy of brightness and contrast. MPC is an optimization control method that predicts the system behavior in the future and achieves the goal by optimizing the current control input. Assume that it is necessary to The goal of MPC is to minimize the following cost function: ;
[0125] in, and Respectively indicate time Brightness and contrast, and are the reference brightness and contrast values, and is the weight coefficient of the cost function. By optimizing this cost function, a brightness control sequence and a contrast control sequence are obtained. According to the multi-step control sequence, parameter optimization processing is performed to obtain optimal control parameters, which include optimal brightness parameters and optimal contrast parameters, and display parameter control instructions are generated based on the optimal control parameters, which include brightness control instructions and contrast control instructions.
[0126] In a specific embodiment, the process of executing step 500 may specifically include the following steps:
[0127] Performing a time series correlation analysis on the display response trigger data to obtain interactive response time series data, the interactive response time series data including trigger delay data and response duration data;
[0128] Analyzing the control effect of the display parameter control instruction to obtain parameter control effect data, wherein the parameter control effect data includes brightness control effect data and contrast control effect data;
[0129] Performing regional mapping on the interactive response time series data according to the audience area division map to obtain regional interactive analysis data, the regional interactive analysis data including interactive response distribution data of different viewing areas;
[0130] Performing viewing angle weighting processing on the regional interaction analysis data to obtain viewing angle weight data, the viewing angle weight data includes horizontal viewing angle weight and vertical viewing angle weight, and performing compensation calculation on the parameter control effect data based on the viewing angle weight data to obtain regional compensation data, the regional compensation data includes brightness compensation parameters and contrast compensation parameters;
[0131] Performing multi-region display quality analysis according to the regional compensation data to obtain regional display quality data, the regional display quality data including display clarity data and color reproduction data;
[0132] The regional display quality data and the interactive response timing data are comprehensively scored to obtain interactive experience scoring data, which includes a display quality score and an interactive response score. A display parameter optimization strategy is generated based on the interactive experience scoring data, which includes a brightness optimization rule and a contrast optimization rule.
[0133] Specifically, the display response trigger data is analyzed for timing correlation to obtain interactive response timing data, including trigger delay and response duration data. The trigger delay data reflects the time difference from the trigger event to the display response, while the response duration data indicates the duration of the response process. These two indicators can effectively reflect the real-time and stability of the interaction between the display and the audience. For example, if the trigger delay is long, it means that there is a performance bottleneck in the system and the response speed needs to be optimized. The control effect of the display parameter control instructions is analyzed to obtain data on the brightness and contrast control effects. Brightness control effect and contrast control effects The evaluation is performed by comparing the difference between the target value and the actual value. The specific calculation formula is:
[0134] ;
[0135] ;
[0136] in, and are the actual brightness and contrast values, respectively. and are the target brightness and contrast values. These control effect data reflect the gap between the actual performance of the display and the target, helping to further adjust the control strategy and improve the display effect. After obtaining the interactive response timing data and display control effect data, regional interaction analysis is performed. In order to evaluate the interactive effects of different regions, the interactive response timing data is mapped with the audience area division map to obtain regional interaction analysis data. The interactive data of different regions are matched with the display area, and the interactive response distribution data of different regions are obtained by analyzing the interaction situation of each region. These regional interaction data provide regional feedback for display effect optimization. These data are weighted by viewing angle, taking into account that different viewing angles have different effects on the display effect. The responses at different angles are weighted to adjust the display effect according to the actual viewing angle of the audience, and the viewing angle weight data is obtained. The viewing angle weight data includes horizontal viewing angle weight and vertical viewing angle weight. By calculating the angle difference, the influence of the viewing angle on the interactive data is obtained, and then the parameter control effect is compensated. The compensated brightness and contrast data are used to optimize the display effect, especially in regions with different angles. Based on the compensated brightness and contrast data, multi-region display quality analysis is performed to evaluate display quality indicators such as display clarity and color reproduction. Display clarity reflects the sharpness of the image, while color reproduction reflects the accuracy of the color of the display in different areas. Based on the comprehensive scoring of regional display quality data and interactive response timing data, the interactive experience scoring data is obtained. The interactive experience scoring data includes display quality score and interactive response score. Based on these two scores, the parameter optimization strategy of the display is generated. The optimization strategy includes brightness optimization rules and contrast optimization rules. The display parameters are dynamically adjusted according to the audience's behavior data and changes in display content to ensure that the visual effect of the display reaches the best state in different audience interaction environments.
[0137] In a specific embodiment, the process of executing step 600 may specifically include the following steps:
[0138] Performing a timing analysis on the display parameter optimization strategy through a first timing sliding window to obtain first optimization rule sequence data; performing a timing analysis through a second timing sliding window to obtain second optimization rule sequence data; combining the first optimization rule sequence data with the second optimization rule sequence data to obtain optimization rule sequence data;
[0139] Based on the optimized rule sequence data, a first-level interaction rule graph and a second-level interaction rule graph are constructed, wherein the first-level interaction rule graph describes the association relationship between single interaction rules, and the second-level interaction rule graph describes the association relationship between combined interaction rules, and the first-level interaction rule graph and the second-level interaction rule graph are merged to obtain interaction pattern association data;
[0140] The interaction pattern association data is clustered by rule association using a first directed graph clustering algorithm to obtain a first interaction pattern category; the rule hierarchical clustering is performed using a second directed graph clustering algorithm to obtain a second interaction pattern category; the first interaction pattern category and the second interaction pattern category are combined to obtain interaction pattern category data;
[0141] Based on the interactive mode category data, the display area is divided into blocks using an adaptive grid division algorithm to obtain first area division data; the display area is divided into blocks using an audience density perception algorithm to obtain second area division data; the first area division data and the second area division data are weightedly fused to obtain area division data;
[0142] Performing a first uniformity measurement calculation on the regional division data to obtain local display uniformity data; performing a second uniformity measurement calculation to obtain global display uniformity data; performing multi-scale fusion on the local display uniformity data and the global display uniformity data to obtain display uniformity data;
[0143] Based on the display uniformity data, a local area compensation parameter is calculated by a first compensation model to obtain first compensation data; a global area compensation parameter is calculated by a second compensation model to obtain second compensation data; and the first compensation data and the second compensation data are nonlinearly weighted combined to obtain uniformity compensation data;
[0144] The uniformity compensation data and the interactive mode category data are subjected to feature-level fusion through a first fusion network to obtain first control strategy data; the decision-level fusion is performed through a second fusion network to obtain second control strategy data; the first control strategy data and the second control strategy data are adaptively selected to obtain interactive control strategy data;
[0145] Based on the interactive control strategy data, first configuration data and second configuration data are generated, wherein the first configuration data controls the dynamic adjustment of display parameters, and the second configuration data controls the real-time switching of interactive rules, and the first configuration data and the second configuration data are combined to obtain the display screen interactive configuration data.
[0146] Specifically, the time series data of the display screen is analyzed through the first time series sliding window to obtain the first optimization rule sequence data. Based on historical data and real-time feedback, the analysis data is continuously updated on the time axis through the sliding window method. Assume that there is a set of time series data , the first time series sliding window It is expressed as: ;
[0147] in Indicates the current moment, is the size of the window. The sliding window process analyzes the changing trend of the time series data by updating the window at each time point. By performing sliding window analysis on the time series data, the first optimization rule sequence data is obtained, and the optimization strategy and system response behavior in different time periods are extracted. Similarly, the second time series sliding window By analyzing the time series data in a similar way, we can obtain the second optimization rule sequence data: ;
[0148] By comparing the optimization sequence data obtained from the two sliding windows, the two optimization rule sequence data are combined to obtain a unified optimization rule sequence data. , used to guide the adjustment of display parameters. The combination process uses weighted averaging or other fusion techniques, such as weighted summation: ;
[0149] in, and are the first and second optimization rule sequence data respectively, is the weight coefficient, which controls the relative importance of the two sequence data. Based on the obtained optimization rule sequence data, an interaction rule graph is constructed. First-level interaction rule graph , describing the association between single interaction rules. The nodes of the interaction rule graph represent different interaction rules, and the edges represent the relationship between these rules. Each node represents a specific display parameter or a control strategy, and the weight of the edge represents the degree of association or influence between them. For example, the node Indicates brightness adjustment, node Indicates contrast adjustment, side It indicates the mutual influence between these two parameters. Similarly, the second-level interaction rule diagram The association relationship between the combined interaction rules is described. By analyzing the effects of different rule combinations, higher-level interaction rules are obtained. The first-level interaction rule graph and the second-level interaction rule graph are fused to obtain the interaction pattern association data. , used to represent the overall behavior pattern between different interaction rule combinations. This data helps the system identify and optimize interaction patterns in complex interactive environments and improve user experience. Cluster analysis is performed on the interaction pattern association data. The first directed graph clustering algorithm is used to cluster the interaction patterns based on rule association to obtain the first interaction pattern category. . The problem is simplified by grouping similar interaction rules into the same category. In this process, the similarity between rules is calculated using a metric in graph theory such as cosine similarity: ;
[0150] in, and Represent two regular vectors, represents the dot product of vectors, and is the modulus of the rule vector. By calculating the similarity, similar rules are grouped together. The second directed graph clustering algorithm is used to cluster the hierarchical relationship of the rules to obtain the second interaction pattern category This process helps refine the hierarchy of interaction patterns, allowing the system to identify interaction strategies at different levels. and Combined to obtain the final interaction mode category data . Based on interaction pattern category data , the display area is divided into blocks using an adaptive grid division algorithm to obtain the first area division data The method automatically adjusts the division of the display area according to the physical layout of the display screen and the characteristics of the interactive mode, ensuring that each area can be effectively optimized. Similarly, the audience density perception algorithm is used to divide the display area into blocks to obtain the second area division data , which takes into account the density and behavior of viewers at different locations, thereby optimizing the division of display areas. and the second area division data Perform weighted fusion to obtain the final regional division data On this basis, the uniformity measurement calculation is performed on the regional division data to obtain the local display uniformity data. and global display uniformity data These metrics help analyze the uniformity and consistency of the display area, ensuring that the display effect in each area meets the expected quality standards. and global display uniformity data Perform multi-scale fusion to obtain display uniformity data Based on display uniformity data , the local area is compensated by the first compensation model to obtain the local area compensation parameter data At the same time, the global area is compensated by the second compensation model to obtain the global area compensation parameter data The two compensation data are combined through nonlinear weighting to obtain the final uniformity compensation data. , ensuring that the display effect in all areas can reach the best state. and interaction mode category data The first control strategy data is obtained by performing feature-level fusion through the first fusion network. Then the decision-level fusion is performed through the second fusion network to obtain the second control strategy data By adaptively selecting these two strategy data, we finally get the interactive control strategy data , used to guide the dynamic adjustment of the display screen and the real-time switching of the interaction rules. Based on the interactive control strategy data , generate the first configuration data and the second configuration data , respectively controlling the dynamic adjustment of display parameters and the real-time switching of interaction rules. Combining these two, the final display screen interaction configuration data is obtained. , this configuration data can optimize the display effect in real time according to the audience's behavior and changes in displayed content, providing the best interactive experience.
[0151] The visual interaction method of the LED display screen in the embodiment of the present application is described above. The visual interaction device 10 of the LED display screen in the embodiment of the present application is described below. Figure 2 In one embodiment of the present application, a visual interaction device 10 of an LED display screen includes:
[0152] The recognition module 11 is used to perform illumination balance and group posture recognition on the audience area image data to obtain audience group posture heat map data;
[0153] The extraction module 12 is used to perform high dynamic range feature extraction and bidirectional visual language attention decoding on the LED display screen to obtain semantic area data of the display content;
[0154] Processing module 13, used for performing manifold distance audience behavior clustering and perspective difference impact processing on audience group posture heat map data and display content semantic area data to obtain display response trigger data;
[0155] A control module 14, for performing potential field model and model prediction control on display parameters based on display response trigger data to obtain display parameter control instructions;
[0156] An analysis module 15 is used to analyze the interaction effect based on the display parameter control instruction and obtain a display screen parameter optimization strategy;
[0157] The evaluation module 16 is used to perform interactive mode analysis and multi-region display uniformity evaluation on the display screen parameter optimization strategy to obtain display screen interactive configuration data.
[0158] Through the collaborative cooperation of the above-mentioned components, a complete set of visual interaction methods for LED display screens is established by combining deep learning technology with multimodal interactive control. This method uses a dual-branch Transformer module and a multi-layer visual feature encoding network to achieve accurate recognition of audience group behavior and semantic understanding of display content; through manifold distance audience behavior clustering and perspective difference impact processing, the audience's interactive intention is accurately captured; the display parameter optimization method based on potential field model and model predictive control realizes dynamic adjustment of display effect; through interactive effect analysis and multi-region display uniformity evaluation, the consistency of visual experience in different viewing areas is guaranteed. The present invention significantly improves the interactive intelligence level of LED display screens, effectively solves the technical difficulties of traditional methods in group behavior recognition, interactive intention understanding, display parameter optimization and visual experience uniformity, and provides a new technical solution for the intelligent interaction of LED display screens in large venues. The present invention shows good technical effects in practical applications, effectively improving the audience's interactive experience and viewing experience.
[0159] See also Figure 3 , Figure 3 This is a schematic block diagram of the structure of an electronic device 300 provided in an embodiment of the present application. The electronic device 300 includes a processor 301 and a memory 302. The processor 301 and the memory 302 are connected via a device bus 303, wherein the memory 302 may include a non-volatile storage medium and an internal memory.
[0160] The non-volatile storage medium can store a computer program. The computer program includes program instructions, and when the program instructions are executed by the processor 301, the processor 301 can execute any of the above-mentioned visual interaction methods of the LED display screen.
[0161] The processor 301 is used to provide computing and control capabilities to support the operation of the entire electronic device 300 .
[0162] The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium. When the computer program is executed by the processor 301, the processor 301 can execute any of the above-mentioned visual interaction methods of the LED display screen.
[0163] Those skilled in the art will understand that Figure 3 The structure shown in the figure is only a block diagram of a partial structure related to the present application scheme, and does not constitute a limitation on the electronic device 300 involved in the present application scheme. The specific electronic device 300 may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0164] It should be understood that the processor 301 may be a central processing unit (CPU), and the processor 301 may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0165] It should be noted that technicians in the relevant field can clearly understand that for the convenience and simplicity of description, the specific working process of the electronic device 300 described above can refer to the corresponding process of the visual interaction method of the aforementioned LED display screen, and will not be repeated here.
[0166] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by one or more processors, the one or more processors implement the visual interaction method of the LED display screen provided in the embodiment of the present application.
[0167] The computer-readable storage medium may be an internal storage unit of the electronic device 300 in the aforementioned embodiment, such as a hard disk or memory of the electronic device 300. The computer-readable storage medium may also be an external storage device of the electronic device 300, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped with the electronic device 300.
[0168] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0169] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions to enable an electronic device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc., various media that can store program codes.
[0170] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A visual interaction method for an LED display screen, characterized in that: The method comprises: Perform illumination balance and group posture recognition on the audience area image data to obtain audience group posture heat map data; Perform high dynamic range feature extraction and bidirectional visual language attention decoding on the LED display screen to obtain the semantic area data of the display content; Performing manifold distance audience behavior clustering and viewing angle difference impact processing on the audience group posture heat map data and the display content semantic area data to obtain display response trigger data; Based on the display response trigger data, a potential field model and a model prediction control are performed on the display parameters to obtain a display parameter control instruction; Based on the display parameter control instruction, an interaction effect analysis is performed to obtain a display screen parameter optimization strategy; specifically, the strategy includes: performing a timing correlation analysis on the display response trigger data to obtain interaction response timing data, wherein the interaction response timing data includes trigger delay data and response duration data; performing a control effect analysis on the display parameter control instruction to obtain parameter control effect data, wherein the parameter control effect data includes brightness control effect data and contrast control effect data; performing regional mapping on the interaction response timing data according to an audience area division map to obtain regional interaction analysis data, wherein the regional interaction analysis data includes interaction response distribution data of different viewing areas; performing viewing angle weighting processing on the regional interaction analysis data to obtain viewing angle weight data, wherein the viewing angle weight data The method comprises the following steps: performing compensation calculation on the parameter control effect data based on the perspective weight data to obtain regional compensation data, wherein the regional compensation data includes brightness compensation parameters and contrast compensation parameters; performing multi-region display quality analysis based on the regional compensation data to obtain regional display quality data, wherein the regional display quality data includes display clarity data and color restoration data; performing comprehensive scoring processing on the regional display quality data and the interactive response timing data to obtain interactive experience scoring data, wherein the interactive experience scoring data includes a display quality score and an interactive response score; and generating a display screen parameter optimization strategy based on the interactive experience scoring data, wherein the display screen parameter optimization strategy includes a brightness optimization rule and a contrast optimization rule; The display screen parameter optimization strategy is subjected to interactive mode analysis and multi-region display uniformity evaluation to obtain display screen interactive configuration data.
2. The visual interaction method of the LED display screen according to claim 1, characterized in that: The step of performing illumination balancing and group posture recognition on the audience area image data to obtain audience group posture heat map data includes: Performing multi-scale Gaussian filtering on the audience area image data to obtain multi-layer image pyramid data, and performing local contrast calculation based on the multi-layer image pyramid data to obtain regional light intensity distribution data; Performing nonlinear mapping and histogram equalization on the regional light intensity distribution data to obtain illumination balanced image data, and performing feature extraction on the illumination balanced image data to obtain basic feature data; Performing feature enhancement on the basic feature data through a multi-layer coding layer of a dual-branch Transformer module to obtain enhanced feature data, wherein the multi-layer coding layer includes a Transformer layer and a hole convolution layer; Perform key point detection and spatial coordinate mapping based on the enhanced feature data to obtain audience posture key point data, wherein the audience posture key point data includes head position data, torso direction data and arm movement data; Performing radial distortion correction and perspective transformation on the audience posture key point data to obtain corrected posture data; A spatial density distribution matrix is constructed based on the corrected posture data, and audience group posture heat map data is obtained through kernel density calculation. The audience group posture heat map data includes spatial distribution density data and activity level data of the audience group.
3. The visual interaction method of the LED display screen according to claim 2, characterized in that: The high dynamic range feature extraction and bidirectional visual language attention decoding of the LED display screen are performed to obtain the semantic area data of the display content, including: The LED display screen image is converted from RGB to HSV color space, and multi-scale decomposition is performed through wavelet transform to obtain a high dynamic range feature image, wherein the high dynamic range feature image includes low frequency component data, medium frequency component data and high frequency component data; Inputting the high dynamic range feature image into a three-layer cascade convolutional network for hierarchical feature extraction, the first layer of the three-layer cascade convolutional network uses 64 3×3 convolution kernels and ReLU activation function to extract basic texture features, the second layer uses 128 5×5 convolution kernels and parameterized ReLU activation function to extract regional structure features, and the third layer uses 256 7×7 convolution kernels and a LeakyReLU activation function with a negative slope of 0.2 to extract global semantic features, thereby obtaining three-layer convolution features; Performing spatial attention mechanism and channel attention mechanism processing on the three-level convolutional features to obtain feature attention weight data; The feature attention weight data is input into a bidirectional visual-language attention module, wherein the visual-to-language attention submodule uses an 8-head attention mechanism and a 1024-dimensional query-key-value matrix to calculate feature mapping relationships, and the language-to-visual attention submodule uses a cross-attention mechanism based on cosine similarity and a 512-dimensional context vector to fuse semantic information to obtain a bidirectional attention output feature map; Reconstructing the bidirectional attention output feature map through a four-layer deconvolution network, wherein the first layer of the four-layer deconvolution network adopts a 4×4 deconvolution kernel with 256 channels and a step size of 2, the second layer adopts a 4×4 deconvolution kernel with 128 channels and a step size of 2, the third layer adopts a 4×4 deconvolution kernel with 64 channels and a step size of 2, and the fourth layer adopts a 4×4 deconvolution kernel with 32 channels and a step size of 2, to obtain a reconstructed feature map; Perform multi-scale feature fusion on the reconstructed feature map to obtain a fused feature map, upsample the fused feature map by 16 times, and use a bicubic interpolation kernel function to restore the feature resolution from 32×32 to the original 512×512 resolution to obtain pixel-level feature mapping data; Semantic segmentation is performed based on the pixel-level feature map data to obtain display content semantic region data, wherein the display content semantic region data includes region category labels and interactive attribute annotations.
4. The visual interaction method of the LED display screen according to claim 3, characterized in that: The performing of manifold distance audience behavior clustering and perspective difference impact processing on the audience group posture heat map data and the display content semantic area data to obtain display response trigger data includes: Extracting features from the audience group posture heat map data to obtain a group feature vector, wherein the group feature vector includes audience spatial distribution features and group activity features; Performing regional attribute analysis on the displayed content semantic region data to obtain a regional feature vector, wherein the regional feature vector includes a regional category label and an interactive attribute tag; Calculating a similarity measure based on the group feature vector and the region feature vector to obtain a group-region association matrix, and performing a manifold distance calculation on the group-region association matrix to obtain audience behavior clustering data, wherein the audience behavior clustering data includes a group behavior type and an activity level score; Calculating the viewing angle deviation of the audience behavior clustering data based on the audience seat layout diagram to obtain viewing angle difference data, wherein the viewing angle difference data includes a horizontal viewing angle deviation and a vertical viewing angle deviation; Performing compensation processing on the audience behavior clustering data according to the perspective difference data to obtain perspective correction data, and performing feature fusion on the perspective correction data and the regional feature vector to obtain interactive feature data; A response decision analysis is performed based on the interaction feature data to obtain display response trigger data, wherein the display response trigger data includes an interaction trigger signal and response type information.
5. The visual interaction method of the LED display screen according to claim 4, characterized in that: The method of performing potential field model and model prediction control on display parameters based on the display response trigger data to obtain display parameter control instructions includes: Constructing a potential field function according to the display content semantic area data and the display response trigger data to obtain initial potential energy distribution data, wherein the initial potential energy distribution data includes brightness potential energy and contrast potential energy; Performing gradient calculation on the initial potential energy distribution data to obtain display parameter change trend data, wherein the display parameter change trend data includes a brightness change direction and a contrast change direction; Building a state space model based on the display parameter change trend data to obtain display state prediction data, the display state prediction data including a brightness prediction value and a contrast prediction value, and performing display constraint analysis on the display state prediction data to obtain parameter constraint data, the parameter constraint data including a brightness range constraint and a contrast range constraint; Performing dynamic range calibration on the parameter constraint data to obtain display parameter calibration data, the display parameter calibration data including brightness calibration parameters and contrast calibration parameters, and performing model predictive control calculation based on the display parameter calibration data to obtain a multi-step control sequence, the multi-step control sequence including a brightness control sequence and a contrast control sequence; Parameter optimization processing is performed according to the multi-step control sequence to obtain optimal control parameters, the optimal control parameters including optimal brightness parameters and optimal contrast parameters, and display parameter control instructions are generated based on the optimal control parameters, the display parameter control instructions including brightness control instructions and contrast control instructions.
6. The visual interaction method of the LED display screen according to claim 1, characterized in that: The interactive mode analysis and multi-region display uniformity evaluation of the display parameter optimization strategy are performed to obtain the display interactive configuration data, including: Performing a time series analysis on the display screen parameter optimization strategy through a first time series sliding window to obtain first optimization rule sequence data; performing a time series analysis through a second time series sliding window to obtain second optimization rule sequence data; combining the first optimization rule sequence data and the second optimization rule sequence data to obtain optimization rule sequence data; constructing a first-level interaction rule graph and a second-level interaction rule graph based on the optimization rule sequence data, wherein the first-level interaction rule graph describes the association relationship between single interaction rules, and the second-level interaction rule graph describes the association relationship between combined interaction rules, and fusing the first-level interaction rule graph and the second-level interaction rule graph to obtain interaction mode association data; The interaction pattern association data is clustered by rules according to a first directed graph clustering algorithm to obtain a first interaction pattern category; the interaction pattern association data is clustered by rules according to a second directed graph clustering algorithm to obtain a second interaction pattern category; the first interaction pattern category and the second interaction pattern category are combined to obtain interaction pattern category data; Based on the interaction mode category data, the display area is divided into blocks using an adaptive grid division algorithm to obtain first area division data; the display area is divided into blocks using an audience density perception algorithm to obtain second area division data; the first area division data and the second area division data are weightedly fused to obtain area division data; Performing a first uniformity measurement calculation on the region division data to obtain local display uniformity data; performing a second uniformity measurement calculation to obtain global display uniformity data; performing multi-scale fusion on the local display uniformity data and the global display uniformity data to obtain display uniformity data; Based on the display uniformity data, a local area compensation parameter is calculated by a first compensation model to obtain first compensation data; a global area compensation parameter is calculated by a second compensation model to obtain second compensation data; and the first compensation data and the second compensation data are nonlinearly weighted combined to obtain uniformity compensation data; The uniformity compensation data and the interactive mode category data are subjected to feature-level fusion through a first fusion network to obtain first control strategy data; decision-level fusion is performed through a second fusion network to obtain second control strategy data; the first control strategy data and the second control strategy data are adaptively selected to obtain interactive control strategy data; Based on the interaction control strategy data, first configuration data and second configuration data are generated, wherein the first configuration data controls the dynamic adjustment of display parameters, and the second configuration data controls the real-time switching of interaction rules, and the first configuration data and the second configuration data are combined to obtain display screen interaction configuration data.
7. A visual interaction device for an LED display screen, characterized in that: The method for visual interaction of an LED display screen according to any one of claims 1 to 6 is used, wherein the visual interaction device of the LED display screen comprises: The recognition module is used to perform illumination balance and group posture recognition on the audience area image data to obtain audience group posture heat map data; The extraction module is used to perform high dynamic range feature extraction and bidirectional visual language attention decoding on the LED display screen to obtain the semantic area data of the display content; A processing module, used for performing manifold distance audience behavior clustering and viewing angle difference impact processing on the audience group posture heat map data and the display content semantic area data to obtain display response trigger data; A control module, used to perform potential field model and model prediction control on display parameters based on the display response trigger data to obtain display parameter control instructions; An analysis module, used to analyze the interaction effect based on the display parameter control instruction to obtain a display screen parameter optimization strategy; The evaluation module is used to perform interactive mode analysis and multi-region display uniformity evaluation on the display screen parameter optimization strategy to obtain display screen interactive configuration data.
8. An electronic device, characterized in that: The electronic device comprises: a memory and at least one processor, wherein instructions are stored in the memory; The at least one processor calls the instructions in the memory so that the electronic device executes the visual interaction method of the LED display screen as described in any one of claims 1-6.
9. A computer-readable storage medium having instructions stored thereon, characterized in that: When the instructions are executed by the processor, the visual interaction method of the LED display screen according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Intelligent interactive service method and system based on regional advertising media
CN111899677A
Object-level infrared-and-visible-light image fusion method based on fully convolutional neural network
WO2024174488A1