LED screen partition display method, device, medium and product

By accurately locating the user's focus area through multimodal interactive data, the LED screen partitioning display method logically divides the LED screen into high and low priority sub-screens, reduces the transmission frame rate, and generates intermediate transition frame images. This solves the problem of insufficient bandwidth in ultra-high-definition LED display systems and achieves efficient data transmission and smooth display effects.

CN122179614APending Publication Date: 2026-06-09SUZHOU ZHONGAO INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SUZHOU ZHONGAO INFORMATION TECH CO LTD
Filing Date
2026-03-20
Publication Date
2026-06-09

Smart Images

  • Figure CN122179614A_ABST
    Figure CN122179614A_ABST
Patent Text Reader

Abstract

This application provides a method, device, medium, and product for LED screen partitioned display, relating to the field of video signal processing technology. The method includes: firstly, acquiring multimodal interaction data between the user and the digital twin model displayed on the LED screen and determining the area of ​​interest, thereby dividing the LED sub-screen into a first LED sub-screen and a second LED sub-screen. Next, generating a first image data sequence and a second image data sequence. Finally, sending the first image data sequence to a first LED receiving card to drive the first LED sub-screen display; and sending the second image data sequence to a second LED receiving card, which performs interpolation calculations to generate intermediate transition frame images and generates a display drive signal equal to the standard display frame rate to control the second LED sub-screen display. This solution addresses the technical problem of reducing the bandwidth usage of the transmission link without changing the front-end physical display refresh rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video signal processing technology, specifically to a method, device, medium, and product for LED screen partitioned display. Background Technology

[0002] With the development of LED display technology, ultra-high-definition (4K / 8K) large screens are widely used in digital twin and virtual simulation scenarios, and users often need to control the screen in real time through interactive means.

[0003] In existing display systems, to maintain image integrity, transmitting devices typically employ a full-frame uniform transmission method. That is, regardless of changes in the screen content, the system continuously packages and sends all pixel data of the entire screen to the receiving end at a fixed high frame rate to ensure signal continuity.

[0004] However, at ultra-high resolutions, the data stream generated by this full-volume transmission mode is extremely large, easily reaching the bandwidth limit of physical cables (such as network cables and fiber optics). To maintain transmission, a large number of transmitting devices and cables are often required, leading to complex system installation, increased costs, and larger size. Furthermore, when the data throughput is close to its limit, network congestion can easily cause frame drops, stuttering, or even signal interruptions. Summary of the Invention

[0005] This application provides a method, device, medium, and product for LED screen partitioned display, to solve the technical problem in the prior art of how to reduce the bandwidth usage of the transmission link without changing the physical display refresh rate at the front end.

[0006] In a first aspect, this application provides a method for partitioned display of an LED screen, including: The system acquires multimodal interaction data between the user and the digital twin model displayed on the LED screen, and determines the area of ​​interest of the LED screen based on the multimodal interaction data. The LED screen includes several LED sub-screens. The LED sub-screen corresponding to the area of ​​interest is designated as the first LED sub-screen, and the LED sub-screens other than the first LED sub-screen are designated as the second LED sub-screen. The first image data sequence corresponding to each of the first LED sub-screens and the second image data sequence corresponding to each of the second LED sub-screens are generated according to the digital twin model; wherein, the first transmission frame rate corresponding to the first image data sequence is equal to the preset standard display frame rate, and the second transmission frame rate corresponding to the second image data sequence is lower than the first transmission frame rate. The first image data sequence is sent to the first LED receiving card corresponding to the first LED sub-screen, so that the first LED receiving card drives the first LED sub-screen to display according to the first image data sequence; The second image data sequence is sent to the second LED receiving card corresponding to the second LED sub-screen, so that each second LED receiving card performs interpolation calculation based on the second image data sequence to generate a number of intermediate transition frame images. The second image data sequence and the intermediate transition frame images are used to generate a display driving signal equal to the standard display frame rate, and the second LED sub-screen is controlled to display according to the display driving signal.

[0007] Optionally, the multimodal interaction data includes: the user's eye movement image and the user's voice interaction commands, and determining the attention area of ​​the LED screen based on the multimodal interaction data specifically includes: Calculate the physical line of sight based on the eye movement image; The semantic gaze point is obtained by parsing the voice interaction command through a pre-trained semantic parsing model. The intent confidence of the voice interaction command is calculated using a pre-trained NLP model; When the confidence level of the intent is greater than a preset threshold, the coordinates of the semantic gaze point are determined as the center coordinates of the region of interest. When the confidence level of the intent is less than or equal to the preset threshold, the coordinates of the physical line of sight landing point are determined as the center coordinates of the area of ​​interest. The region of interest is determined based on the preset radius and the center coordinates.

[0008] Optionally, a second image data sequence corresponding to each of the second LED sub-screens is generated based on the digital twin model, specifically including: Image frame data corresponding to the second LED sub-screen is generated based on the digital twin model; Extract the attributes of text objects and dynamic chart objects located within the display area of ​​the second LED sub-screen in the digital twin model; Based on the pixel positions of the text object attributes and the dynamic chart object attributes in the image frame data, a binarized mask layer aligned with the pixels of the image frame data is generated. The image frame data and the binarization mask layer are encapsulated into the second image data sequence, wherein the binarization mask layer is used to instruct the second LED receiving card, when generating the intermediate transition frame, to reuse the pixel value of the corresponding position in the previous frame image data of the intermediate transition frame as the pixel value of the intermediate transition frame for the target area covered by the binarization mask layer.

[0009] Optionally, the method further includes: Obtain event data generated by the background monitoring system; When an alarm event exists in the event data, the target area corresponding to the alarm event in the digital twin model is obtained; The target region is marked as the region of interest.

[0010] Optionally, before sending the first image data sequence to the first LED receiving card corresponding to the first LED sub-screen, the method further includes: Calculate the total transmission bandwidth utilization required to send the first image data sequence and the second image data sequence; If the total transmission bandwidth utilization rate is greater than the preset safety threshold, steps S1-S4 are executed repeatedly until the theoretical total transmission bandwidth utilization rate is less than or equal to the preset safety threshold, thus obtaining the final first LED sub-screen and the final second LED sub-screen. S1: Iteratively reduce the area of ​​the region of interest according to a preset reduction step size; S2: Determine the theoretical first LED sub-screen and the theoretical second LED sub-screen based on the reduced area of ​​interest; S3: Generate the theoretical first image data sequence corresponding to the theoretical first LED sub-screen and the theoretical second image data sequence corresponding to the theoretical second LED sub-screen; S4: Calculate the theoretical total transmission bandwidth utilization rate required to send the theoretical first image data sequence and the theoretical second image data sequence; Generate the final first image data sequence corresponding to the final first LED sub-screen and the final second image data sequence corresponding to the final second LED sub-screen; The final first image data sequence is determined as the first image data sequence, and the final second image data sequence is determined as the second image data sequence.

[0011] Optionally, generating a second image data sequence corresponding to each of the second LED sub-screens based on the digital twin model further includes: Extract the pixel motion vector data corresponding to the second LED sub-screen from the rendering pipeline of the digital twin model; The pixel motion vector data is encapsulated into the second image data sequence, wherein the pixel motion vector data includes the displacement direction and distance of corresponding pixels between adjacent frame images, and is used to calculate the intermediate transition frame image.

[0012] Optionally, the method further includes: When the target first LED sub-screen is determined to be the second LED sub-screen, the target first LED sub-screen is confirmed as the first LED sub-screen within a preset time period; When the preset time period is exceeded, the target first LED sub-screen will be identified as the second LED sub-screen.

[0013] In a second aspect, embodiments of this application provide an LED screen partitioned display device, which includes: one or more processors and a memory; the memory is coupled to the one or more processors, and the memory is used to store computer program code, which includes computer instructions, and the one or more processors call the computer instructions to cause the LED screen partitioned display device to perform the method described in the first aspect and any possible implementation thereof.

[0014] Thirdly, embodiments of this application provide a computer program product containing instructions that, when the computer program product is run on an LED screen partition display device, cause the LED screen partition display device to perform the method described in the first aspect and any possible implementation thereof.

[0015] Fourthly, embodiments of this application provide a computer-readable storage medium including instructions that, when executed on an LED screen partition display device, cause the LED screen partition display device to perform the method described in the first aspect and any possible implementation thereof.

[0016] In summary, one or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages: 1. By adopting the above technical solution, the system first uses multimodal interactive data to accurately locate the user's focus area, logically dividing the LED screen into a first LED sub-screen and a second LED sub-screen. The first LED sub-screen, which is the core of the user's visual experience, maintains a standard display frame rate transmission, ensuring the core visual experience; the second LED sub-screen, which is not of interest, reduces the transmission frame rate, reducing data throughput from the source. Crucially, the second LED receiving card generates intermediate transition frame images through interpolation calculations and reconstructs a display drive signal conforming to the standard display frame rate at the terminal. This processing method effectively reduces the bandwidth occupation of the transmission link without changing the front-end physical display refresh rate, alleviating the load pressure of massive pixel data transmission and achieving dynamic matching between transmission bandwidth resources and human visual characteristics.

[0017] 2. By adopting the above technical solution, a dual positioning mechanism is constructed by combining the calculation of physical gaze points from eye movement images with the parsing of semantic gaze points from voice interaction commands. By introducing an NLP model to calculate intent confidence, the system can intelligently determine the user's current dominant interaction modality: prioritizing semantic needs under high intent confidence, and reverting to physical gaze instinct under low confidence. This confidence-based decision logic solves the misjudgment problem caused by inconsistencies between user gaze and operational intent under a single modality, ensuring the accuracy of the attention area determination. This allows high frame rate transmission resources to be accurately allocated to the area where the user truly intends to interact, avoiding resource waste or degradation in the display quality of key information due to incorrect partitioning.

[0018] 3. By adopting the above technical solution, when generating the second image data sequence, the system deeply utilizes the data characteristics of the digital twin model to extract the attributes of text objects and dynamic chart objects and generate a binarized mask layer. This binarized mask layer serves as auxiliary control information sent along with the image data, guiding the second LED receiving card to perform differentiated processing during interpolation calculations. For the target area covered by the mask, the pixel values ​​of the previous frame are reused instead of linear interpolation, effectively avoiding the blurring or artifacts produced by traditional interpolation algorithms when processing high-frequency changing images. This mechanism ensures that even in low transmission frame rate mode, the edges of text and charts in non-interested areas remain clear and sharp, guaranteeing the readability and accuracy of the data content.

[0019] 4. By adopting the above technical solution, pixel motion vector data is directly extracted from the rendering pipeline of the digital twin model, obtaining accurate pixel-level displacement direction and distance information. This data is encapsulated into a second image data sequence, enabling the second LED receiving card to perform motion compensation interpolation based on the actual motion trajectory when generating intermediate transition frame images, rather than relying solely on pixel mixing between preceding and following frames. This interpolation method based on rendering source data eliminates the delay and error of traditional image post-processing for calculating motion vectors, significantly improving the continuity of dynamic images in the time domain. This allows the second LED sub-screen to still present a smooth and natural dynamic display effect under low bandwidth transmission conditions, reducing image stuttering or ghosting. Attached Figure Description

[0020] Figure 1 This is a flowchart illustrating one embodiment of the LED screen partition display method in this application. Figure 2 This is another flowchart illustrating the LED screen partition display method in the embodiments of this application; Figure 3 This is a schematic diagram of the physical device structure of an LED screen partition display device in the embodiments of this application. Detailed Implementation

[0021] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.

[0022] In the description of the embodiments of this application, the words "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design that is described as "for example" or "for instance" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design options. Rather, the use of the words "for example" or "for instance" is intended to present the relevant concepts in a specific manner.

[0023] In the description of the embodiments of this application, the term "multiple" means two or more. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, a feature defined with "first," "second," or "third" may explicitly or implicitly include one or more of that feature. The terms "comprising," "including," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized. It should be noted that all data collection in this scheme is conducted after obtaining user consent.

[0024] This application provides a method for partitioned display of an LED screen, referencing... Figure 1 , Figure 1 This is a flowchart of an LED screen partition display method provided in an embodiment of this application. The method includes: Step S101: Obtain multimodal interaction data between the user and the digital twin model displayed on the LED screen, and determine the area of ​​interest of the LED screen based on the multimodal interaction data. The LED screen includes several LED sub-screens. In this solution, the LED screen refers to a physical display terminal composed of multiple display units spliced ​​together to present visualized information, such as a P1.2 small-pitch LED splicing wall in a command center or an irregularly shaped LED screen in an exhibition. The digital twin model refers to a digitally mapped model corresponding to a physical entity or system, constructed in virtual space. This model is rendered and displayed in real time on the LED screen, such as a 3D city model reflecting real-time urban traffic flow in a smart city management system, or a virtual production line model reflecting equipment operating status in an industrial monitoring system. Multimodal interactive data refers to the collection of information gathered through various sensors or input devices regarding user interaction with the digital twin model, such as... The user gaze coordinates data collected by eye trackers, user voice command data collected by microphone arrays (such as "zoom in to view area A"), or spatial gesture operation data collected by gesture recognition radar; the attention area refers to the screen display range where the user's current visual attention is focused or the interaction intention is directed, such as the alarm pop-up area in the upper left corner of the screen where the user's gaze rests, or the 3D map area in the center of the screen selected by the user through voice command; LED sub-screen refers to the smallest independent physical splicing unit or logical driving unit that constitutes the LED screen, usually corresponding to a single LED cabinet or a single LED module, such as a standardized LED display cabinet with a size of 600mm×337.5mm.

[0025] This step is typically performed during the real-time response phase of human-computer interaction. Specifically, the system continuously monitors the user's interactive behavior. When the sensors capture an interaction signal, the system first analyzes and fuses the multimodal interaction data. For example, if the system simultaneously receives eye-tracking data indicating that the user is looking at the screen coordinates (X: 2000, Y: 1000) and the voice command "Query device parameters," the system will combine the eye-tracking coordinates with the projection position of the digital twin model on the screen to calculate the bounding box of the object the user intends to manipulate on the screen. Subsequently, based on the calculated bounding box or center coordinates, the system delineates the region of interest according to preset rules (such as a circular area centered on the gaze point or the outline area of ​​the selected digital object). For example, if the user is looking at "Turbine No. 2" in the digital twin model, the system marks the rectangular pixel area occupied by "Turbine No. 2" on the LED screen as the region of interest; if the user selects the right side of the screen with a gesture, the area enclosed by the gesture trajectory is the region of interest.

[0026] Step S102: The LED sub-screen corresponding to the area of ​​interest is determined as the first LED sub-screen, and the LED sub-screens other than the first LED sub-screen are determined as the second LED sub-screen; The first LED sub-screen refers to the set of LED sub-screens located within or spatially overlapping with the area of ​​interest. These sub-screens carry the user's current high-value visual information, such as the four LED boxes that cover the user's focal point. The second LED sub-screen refers to the set of remaining LED sub-screens that are not classified as the first LED sub-screen. These sub-screens usually display background information or content that is not currently of interest to the user, such as the remaining 50 LED boxes located at the edge of the screen that display static background images.

[0027] This step is performed after the area of ​​interest is determined and before data transmission. Specifically, the system maps the pixel coordinate range of the area of ​​interest determined in step S101 onto the physical hardware topology of the LED screen. The system traverses the physical location coordinates of each LED sub-screen in the LED screen and determines whether the display range of each LED sub-screen intersects with the area of ​​interest. If all or part of the display area of ​​an LED sub-screen falls within the area of ​​interest, that LED sub-screen is marked as the first LED sub-screen. For example, if the area of ​​interest is a circle with a radius of 500 pixels at the center of the screen, and the system detects that four LED boxes numbered ID_12, ID_13, ID_22, and ID_23 cover this circular area, then these four boxes are marked as the first LED sub-screen. Conversely, any LED sub-screen that does not overlap with the area of ​​interest is marked as the second LED sub-screen. Through this step, the physical LED screen is logically divided into high-priority display areas and low-priority display areas.

[0028] Step S103: Generate a first image data sequence corresponding to each of the first LED sub-screens and a second image data sequence corresponding to each of the second LED sub-screens according to the digital twin model; wherein, the first transmission frame rate corresponding to the first image data sequence is equal to the preset standard display frame rate, and the second transmission frame rate corresponding to the second image data sequence is lower than the first transmission frame rate. The first image data sequence refers to the video stream data containing continuous image frames generated for the first LED sub-screen, such as a real-time rendered video stream with a refresh rate of 60Hz generated for the boxes displaying "Turbine No. 2"; the second image data sequence refers to the video stream data generated for the second LED sub-screen, with a lower time sampling density, such as a video stream with a refresh rate of 10Hz generated for the boxes displaying "static sky background"; the first transmission frame rate represents the frame update frequency of the first image data sequence in the transmission channel, such as 60 frames per second (FPS); the standard display frame rate represents the smooth playback frame rate that the LED screen hardware terminal should ultimately present to the observer, usually 60Hz or 120Hz; the second transmission frame rate represents the frame update frequency of the second image data sequence in the transmission channel, which is downsampled, such as 15 frames per second or 30 frames per second.

[0029] This step is performed during the image rendering and encoding transmission phase. Specifically, the graphics workstation or rendering server renders the image based on the real-time state of the digital twin model. When generating transmission data, the system adopts a differentiated frame generation strategy based on the partitioning results of step S102. For the first LED sub-screen, the rendering engine renders each frame of the image completely at the standard display frame rate (e.g., 60Hz) and encapsulates it into a first image data sequence, ensuring high temporal resolution at the data source end. For example, a frame of image data containing dynamic dashboard reading changes is generated and sent every 16.6 milliseconds. For the second LED sub-screen, the rendering engine or encoder performs frame downsampling, extracting or rendering keyframes only at a lower second transmission frame rate (e.g., 10Hz) and encapsulating them into a second image data sequence. For example, a frame of image data containing distant buildings is generated and sent only every 100 milliseconds. This processing method achieves traffic splitting at the data source, allowing the transmission bandwidth to be concentrated on the first image data sequence, resulting in a high frame rate and higher resolution for the image viewed by the user, while significantly reducing the bandwidth occupied by the second image data sequence.

[0030] Step S104: Send the first image data sequence to the first LED receiving card corresponding to the first LED sub-screen, so that the first LED receiving card drives the first LED sub-screen to display according to the first image data sequence; In this embodiment, the first LED receiver card refers to the hardware control unit logically assigned to drive the first LED sub-screen at the current moment. This receiver card is not a general-purpose receiver card containing only basic SRAM cache, but rather an enhanced edge computing receiver card integrating a large-capacity frame cache (such as external DDR3 / DDR4 SDRAM chips) and a high-performance logic processing unit (such as a Kintex-7 series FPGA). It possesses certain image processing capabilities.

[0031] This step is executed during the transmission and presentation of high-priority data, and is performed in real time in conjunction with the dynamic changes in the area of ​​interest. Specifically, the system transmits the first image data sequence generated in step S103, which has a standard display frame rate (e.g., 60Hz), to the first LED receiving card in real time via a high-speed data transmission channel (such as fiber optic or Cat6 network cable). At this time, after receiving each frame of image data in the first image data sequence, the first LED receiving card does not perform multi-frame buffering through DDR cache, but directly uses on-chip cache for decoding, correction, and point-by-point brightness and color correction. It also generates corresponding row scan signals and column data signals strictly according to the original time interval of the first image data sequence (e.g., one frame every 16.6 milliseconds) to control the LED driver chip on the first LED sub-screen to display the image.

[0032] For example, in a smart city traffic command center scenario, the commander is retrieving real-time surveillance footage of an intersection where a traffic accident has occurred on a large screen. At this moment, the four LED receiver cards displaying the intersection's monitoring window are identified as the primary LED receiver cards. The system sends a 60Hz high-definition real-time monitoring video stream to these four cards. The primary LED receiver cards do not perform any buffering or interpolation of the image; they directly drive the screen at a 60Hz refresh rate. This ensures that the commander sees vehicle movement, traffic light changes, and traffic police hand gestures with zero latency and in real-time, without any algorithm-generated artifacts, thus guaranteeing the accuracy of command decisions.

[0033] Step S105: The second image data sequence is sent to the second LED receiving card corresponding to the second LED sub-screen, so that each second LED receiving card performs interpolation calculation based on the second image data sequence to generate a number of intermediate transition frame images. The second image data sequence and the intermediate transition frame images are used to generate a display driving signal equal to the standard display frame rate, and the second LED sub-screen is controlled to display according to the display driving signal.

[0034] The second LED receiver card refers to the hardware control unit logically assigned to drive the second LED sub-screen at the current moment. Its physical entity is exactly the same as the first LED receiver card (i.e., both are enhanced receiver cards with DDR cache). Interpolation calculation refers to the process of using the logic units (LUTs) inside the FPGA to perform weighted calculations based on the pixel grayscale values ​​of two known frames of images (such as linear blending), or to perform motion prediction based on a simple block matching algorithm, thereby deduce the pixel data at the missing moment on the time axis. The intermediate transition frame image refers to the virtual image frame generated by the above algorithm to fill the low frame rate gap in the second image data sequence. The display drive signal refers to the high refresh rate control signal with a frequency that conforms to the persistence of vision of the human eye, which is finally output to the LED driver chip.

[0035] This step is performed during the transmission compensation and presentation phase of low-priority data. Specifically, the system sends a low-frame-rate (e.g., 10Hz) second image data sequence to the second LED receiving card via the transmission channel. At this time, the second LED receiving card automatically activates the interpolation calculation logic. The second LED receiving card uses the onboard DDR SDRAM as a frame buffer to temporarily store the received two consecutive frames of second image data sequences (e.g., frame Frame_K at time T and frame Frame_K+1 at time T+100ms). Subsequently, the FPGA's internal arithmetic logic reads these two frames of data and calculates the pixel value of the intermediate transition frame based on the time weight between the two key frames at the current time (e.g., Frame_Mid = Frame_K × 0.5 + Frame_K+1 × 0.5). Finally, the second LED receiving card reassembles the original frame and the generated intermediate transition frame image in chronological order to generate a display drive signal that conforms to the standard display frame rate (e.g., 60Hz) to control the second LED sub-screen to emit light.

[0036] For example, in the aforementioned smart city traffic command center scenario, apart from the central accident intersection monitoring window, most of the remaining area of ​​the large screen displays a 3D electronic map of the entire city, featuring a slowly flowing traffic heatmap and flashing weather icons. The 50 enhanced receiver cards displaying these background map areas are designated as secondary LED receiver cards. The system sends only 10 frames (10Hz) of map data per second to these cards, significantly saving transmission bandwidth. The secondary LED receiver cards utilize onboard DDR cache for the preceding and following frames of map data and automatically generate the intermediate transition image using a linear interpolation algorithm, ultimately driving the screen to emit light at a frequency of 60Hz. Although this processing method causes the traffic heatmap on the map to appear tens of milliseconds later than the actual data (because it needs to wait for the next frame to calculate the interpolation), this slight delay is completely imperceptible to the commander for the static background map and the slowly changing heatmap, and the screen still appears smooth and flicker-free. This solves the problem of full-screen stuttering caused by insufficient bandwidth, allowing for higher frame rates and resolutions in areas of visual attention while ensuring a certain image quality in areas of non-visual attention.

[0037] The following is a more detailed description of the process of the method provided in this implementation.

[0038] Optional, see reference Figure 2 Steps S10101-S10106 are more specific steps of step S101; Step S10101: Calculate the physical line of sight based on the eye movement image; Among them, eye motion images refer to continuous frame image data containing user eye features (such as pupil position and corneal reflection point) captured by an infrared eye tracker or a high-resolution camera, such as grayscale eye video stream captured by the Tobii Pro eye tracker at a sampling rate of 120Hz; physical gaze point refers to the physical projection coordinates of the user's gaze on the LED screen plane, calculated based on the principles of geometric optics, and is usually expressed as (x, y) values ​​in the screen pixel coordinate system, such as screen coordinates (1920, 1080).

[0039] This step is performed during the preprocessing stage of multimodal interactive data and is typically synchronized with the user's gaze behavior in real time. Specifically, the system first processes the acquired eye movement images to extract the pupil center coordinates and corneal reflective spot (P-CR vector). Then, using a pre-calibrated spatial mapping matrix, the gaze vector in the eye coordinate system is mapped to the two-dimensional plane coordinate system of the LED screen. To eliminate the jitter interference caused by microsaccades, the system usually performs smoothing filtering (such as Kalman filtering) on ​​the raw calculation results. For example, when the user stares at the "system status bar" in the upper left corner of the screen, the system calculates a series of dense physical gaze point coordinates in real time, such as (100, 200), (102, 198), and (101, 201), and outputs a stable physical gaze point (101, 200) after filtering.

[0040] Step S10102: The semantic gaze point is obtained by parsing the voice interaction command through a pre-trained semantic parsing model; Among them, the pre-trained semantic parsing model refers to a deep learning model (such as BERT, GPT or a dedicated Slot Filling model) trained on a large amount of domain corpus, used to extract key entities and spatial orientation information from natural language text; voice interaction commands refer to the text sequence converted from natural language audio data containing the user's intention through the microphone, such as "enlarge the red alarm window in the middle"; semantic gaze point refers to the corresponding screen coordinate position obtained by looking up the target object or locative words described in the voice command in the digital twin model.

[0041] This step is performed during the data parsing phase following voice interaction. Specifically, the system first converts the user's voice audio into text. Then, the text is input into a pre-trained semantic parsing model, which identifies key entities (such as "alarm window") or directional words (such as "middle" or "bottom left corner") in the command. The system then retrieves the entity's attributes in the scene graph of the digital twin model to obtain its screen projection center coordinates in the currently rendered image. For example, if the user says "view generator number two," the model parses the entity "generator number two," and the system finds that "generator number two" is currently displayed at screen pixel coordinates (3000, 1500), which is the semantic gaze point. If the command is a vague direction such as "look to the right," the system uses the geometric center of a preset area to the right of the currently focused area as the semantic gaze point.

[0042] Step S10103: Calculate the intent confidence of the voice interaction command using a pre-trained NLP model; Among them, the pre-trained NLP model refers to the neural network model used for natural language understanding (NLU), which has the functions of intent recognition and classification, and is specifically used to determine whether the input speech text belongs to the predefined set of screen control instructions; the intent confidence score is the probability score of the model judging that the current speech instruction belongs to the category of "valid visual control instruction", which is usually a value between 0 and 1.

[0043] This step, performed concurrently with voice command parsing, filters out invalid speech and evaluates the execution value of commands. Specifically, the system inputs the speech text into the NLP model. The model first determines whether the text belongs to a non-interactive category such as "casual conversation," "interjections," or "ambient noise." If it belongs to a non-interactive category, the model directly outputs a very low intent confidence (approaching 0). If the text contains verbs or nouns (such as "view" or "map"), the model further analyzes its semantic structure and calculates the probability that the command corresponds to a specific screen operation (such as focus, zoom, or selection). For example, in a command center, if the commander says "access the intersection surveillance," the model recognizes a clear "access (operation) + surveillance (object)" structure, classifying it as a strong visual intent command and outputting an intent confidence of 0.92. If the commander is simply chatting with someone next to him saying "the weather is nice today," the model identifies the text as belonging to the "casual conversation" category, containing no visual guidance intent for the screen content, and therefore outputs an intent confidence of 0.05. Through this mechanism, the system can effectively filter out background speech interference unrelated to screen interaction.

[0044] Step S10104: When the confidence level of the intent is greater than a preset threshold, the coordinates of the semantic gaze point are determined as the center coordinates of the region of interest. Among them, the preset threshold refers to the probability limit value set by the system to determine whether the voice command is effective, such as 0.7 or 0.8; the center coordinates of the area of ​​interest refers to the final determined geometric reference point used to generate the high frame rate display area.

[0045] This step is performed during the decision fusion phase of multimodal data. Specifically, the system compares the intent confidence calculated in step S10103 with a preset threshold. If the intent confidence is higher than the threshold, it indicates that the user has expressed a strong and clear need for visual attention through voice. In this case, the system considers the information from the voice channel to be more representative of the user's true intent than the information from the eye-tracking channel (because the human eye may unconsciously scan, but voice commands are usually consciously issued). Therefore, the system ignores the current physical gaze point and forcibly uses the semantic gaze point parsed in step S10102 as the final center coordinates. For example, if the user's eyes are looking at the left side of the screen, but the voice command is "open the map on the right," the intent confidence is 0.9 (greater than 0.7). The system determines the center coordinates of the "map on the right" as the center coordinates of the area of ​​interest, ensuring that the right side of the screen is displayed at a high frame rate, responding to the user's voice request.

[0046] Step S10105: When the confidence level of the intent is less than or equal to the preset threshold, the coordinates of the physical line of sight landing point are determined as the center coordinates of the area of ​​interest. This step is performed during the decision fusion phase of multimodal data. Specifically, if the intent confidence level is lower than or equal to a preset threshold, it indicates that there is currently no effective voice command, or the voice command lacks clear spatial directionality. In this case, the system directly reads the physical gaze point calculated in step S10101 and assigns it as the center coordinate of the area of ​​interest. For example, if the user is browsing the screen in a silent state or engaging in a conversation unrelated to the screen content, the intent confidence level is extremely low. The system tracks the user's eye position in real time; wherever the user looks, the coordinates of that location become the center coordinates.

[0047] Step S10106: Determine the region of interest based on the preset radius and the center coordinates; The preset radius refers to the size parameter of a circular or rectangular area that needs to maintain high definition and high refresh rate, set according to the foveal vision characteristics of the human eye. For example, it is set as the screen pixel radius corresponding to a viewing angle of 5 degrees (such as 300 pixels). The area of ​​interest refers to the screen pixel range that will be allocated to the first LED sub-screen.

[0048] This step is performed after the coordinate decision in the region generation stage. Specifically, the system uses the center coordinates determined in step S10104 or S10105 as the center of a circle (or the center of a rectangle), combined with a preset radius, to generate a closed geometric region. All pixels within this region will be marked as "high priority". To prevent abrupt edge changes, the system sometimes sets a transition zone, but this step mainly focuses on determining the core region. For example, if the center coordinates are determined to be (1000, 1000) and the preset radius is 400 pixels, the system calculates a circular region with (1000, 1000) as the center and a radius of 400 pixels. This circular region is the final region of interest. Subsequently, this region information will be passed to subsequent steps for filtering the first LED sub-screen.

[0049] Optionally, steps S10301-S10304 are more specific steps for generating the second image data sequence; Step S10301: Render and generate image frame data corresponding to the second LED sub-screen based on the digital twin model; Among them, image frame data refers to the RGB color pixel matrix corresponding to the physical resolution of the second LED sub-screen, which is calculated and generated by the rendering engine (such as Unity 3D or Unreal Engine) according to the virtual camera's perspective.

[0050] This step is performed during the generation phase of the low frame rate data stream. Specifically, the graphics workstation or rendering server periodically triggers the rendering pipeline based on the currently set second transmission frame rate (e.g., 10Hz). The rendering engine performs rasterization rendering only on the urban background area covering the second LED sub-screen, generating a static image of the current moment. For example, on the large screen of a smart city traffic command center, the second LED sub-screen is responsible for displaying the static building complex, green belts, and slowly moving weather cloud map of the city's periphery. The rendering engine calculates the lighting and shadow effects of this area every 100 milliseconds, outputting a frame of image data with a resolution of 1920×1080. This data realistically reflects the visual state of the digital twin city at the sampling time, but due to the low sampling frequency, it does not yet possess continuous and smooth dynamic characteristics.

[0051] Step S10302: Extract the text object attributes and dynamic chart object attributes located within the display range of the second LED sub-screen in the digital twin model; Among them, text object attributes refer to the metadata of information elements presented in text form in the digital twin scene, including the bounding box coordinates of road name labels, area names, and status labels, as well as their projection positions in screen space; dynamic chart object attributes refer to the geometric outlines and position information of 2D or 3D chart elements in the scene that change in real time with the data (such as energy consumption dashboards floating above buildings and road congestion index bar charts).

[0052] This step is performed during the post-processing or parallel processing phase of the rendering process, leveraging the "data constructibility" characteristic of digital twin systems. Specifically, the system directly accesses the scene graph of the rendering engine or the UI layer data interface, traversing all visible rendering objects in the current frame. The system filters out objects tagged with "Text," "UI_Panel," or "Chart," and determines whether these objects fall within the display coordinate range of the second LED sub-screen. For objects that meet the criteria, the system extracts their precise screen pixel coordinate range. For example, in the urban background area (second LED sub-screen), a data label displaying "Today's Energy Consumption: 1200kWh" floats above a building. The system directly reads the rectangular area coordinates of this label on the screen (x: 500, y: 200, w: 150, h: 40) via API, without needing to perform complex OCR image recognition on the rendered image.

[0053] Step S10303: Generate a binarized mask layer aligned with the pixels of the image frame data according to the pixel positions of the text object attributes and the dynamic chart object attributes in the image frame data. The binarization mask layer refers to a black and white bitmap (or single-channel data matrix) with the same resolution as the image frame data, where each pixel contains only two state values: 0 or 1. Pixel alignment means that the coordinates (i, j) in the mask layer strictly correspond to the pixel at coordinates (i, j) in the image frame data.

[0054] Specifically, the system creates a blank matrix (initially all 0s) with the same size as the image frame data generated in step S10301. Based on the pixel positions of the road name text and energy consumption chart extracted in step S10302, the system marks the pixel values ​​of the corresponding areas in the matrix as 1 (representing "protected areas" or "statically preserved areas"), and keeps the pixel values ​​of the remaining background areas (such as building surfaces, roads, and the sky) as 0 (representing "interpolation areas"). For example, for the label area "Today's Energy Consumption: 1200kWh" mentioned above, the system fills the corresponding rectangular area of ​​the mask layer with the value 1. The finally generated binarized mask layer accurately outlines all the text and chart contours in the image that need to maintain high clarity and prevent interpolation blurring.

[0055] Step S10304: Encapsulate the image frame data and the binarization mask layer into the second image data sequence, wherein the binarization mask layer is used to instruct the second LED receiving card, when generating the intermediate transition frame, to reuse the pixel value of the corresponding position in the previous frame image data of the intermediate transition frame as the pixel value of the intermediate transition frame for the target area covered by the binarization mask layer. Encapsulation refers to the process of packaging and transmitting RGB image data with single-channel mask data, such as embedding the mask layer into the Alpha channel of the video signal, or sending it as an independent metadata packet with the video frame; the target region refers to the pixel region in the mask layer with a value of 1; multiplexing refers to performing "zero-order hold" on the time axis, that is, directly copying the value of the previous moment without performing mathematical calculations.

[0056] This step is performed during the data packaging and protocol definition phase before data transmission. Specifically, the sending end combines the image frame data with the corresponding binarized mask layer to form a complete second image data sequence, which is then sent to the second LED receiving card. The core of this step lies in defining the decoding logic of the receiving end: when the second LED receiving card calculates the intermediate transition frame at time T+Δt, it reads the mask layer pixel by pixel. If a pixel is 0 (background) in the mask layer, the receiving card performs linear interpolation on the preceding and following frames (e.g., (A+B) / 2), allowing the clouds and traffic to transition smoothly; if a pixel is 1 (text / graphic) in the mask layer, the receiving card forcibly copies the pixel value at that position from the previous frame (time T). For example, for the number "1200kWh", in the generated 5 intermediate transition frames, this part of the pixels always maintains a clear image of "1200kWh" until the next key frame arrives and it jumps to "1201kWh". This mechanism ensures that the city background is smooth in low frame rate interpolation mode, while key road names and energy consumption data remain sharp and clear, without becoming blurry or unreadable due to the interpolation algorithm.

[0057] Optionally, this solution also includes steps S106-S108; Step S106: Obtain event data generated by the background monitoring system; Among them, the background monitoring system refers to a server cluster or Internet of Things (IoT) platform that is independent of the LED display control system and is responsible for business logic processing and data acquisition. For example, a traffic flow monitoring platform for a smart city, a SCADA (Supervisory Control and Data Acquisition) system for a factory, or a security alarm host. Event data refers to a structured information package containing status changes or abnormal situations that is triggered by the background monitoring system according to preset rules. It usually contains fields such as event type, occurrence time, associated device ID, and severity level. For example, a JSON message: {type: "Traffic_Jam", level: "High", location_id: "Crossroad_05", time: "10:00:01"}.

[0058] This step is performed during the system's 24 / 7 background monitoring phase. Specifically, the central processing unit of the LED display control system listens to and subscribes to the data stream of the background monitoring system in real time through dedicated API interfaces, message queues (such as MQTT, RabbitMQ), or database triggers. The system continuously parses each received data message to determine whether any noteworthy state changes have occurred in the current business scenario. For example, in a smart city traffic command center, the system receives hundreds of traffic condition updates pushed by the traffic management bureau server every second, including intersection congestion indices, traffic accident reports, and traffic light malfunctions. This step ensures that the display system can perceive the dynamics of the physical world in real time, rather than simply passively rendering images.

[0059] Step S107: When an alarm event exists in the event data, obtain the target area corresponding to the alarm event in the digital twin model; Among them, alarm events refer to abnormal events in event data that exceed the preset safety threshold in severity level or belong to a specific high-priority category, such as "fire alarm", "equipment shutdown failure" or "severe traffic congestion"; target area refers to the geometric display range corresponding to the physical entity that triggered the alarm in the three-dimensional space or two-dimensional projection plane of the digital twin model.

[0060] This step is executed during the event-driven logic judgment and spatial mapping phase. Specifically, the system first filters the event data obtained in step S106. If a data item is detected as an "alarm event" (such as severe congestion), the system extracts the associated "device ID" or "geographic coordinates" (such as Crossroad_05). Subsequently, the system performs a reverse index in the scene database of the digital twin model to find the virtual model object corresponding to the ID (such as the intersection model group named "Crossroad_05"). After finding the model object, the system calculates the screen projection bounding box or center coordinates of the object from the current camera's perspective. For example, if the background system pushes an alarm for "severe congestion at the Crossroad_05 intersection," the system immediately locates the intersection model in the digital twin city and calculates that the intersection is currently displayed in the lower right corner of the LED screen, occupying a rectangular area with pixel coordinates (x: 2500, y: 1200) to (x: 3000, y: 1500). This rectangular area is the target area.

[0061] Step S108: Mark the target region as the region of interest; Here, "marking" refers to modifying the system's internal status register or rendering configuration table to upgrade the priority attribute of a specific area from "normal" to "high"; "area of ​​interest" refers to the core area defined in the aforementioned steps that will be allocated to the first LED sub-screen and transmitted and displayed at a standard display frame rate (such as 60Hz).

[0062] This step is performed during the dynamic adjustment phase of the display strategy. Specifically, once step S107 determines the target area corresponding to the alarm event, the system immediately triggers the overlay or merging of the current list of areas of interest. The system forcibly marks the pixel range of the target area as the new area of ​​interest. This means that regardless of whether the user's current gaze is on that area or whether the user has issued a voice command, the system will forcibly upgrade the display strategy for that area. For example, although the commander is looking at the chart on the left side of the screen (the original area of ​​interest is on the left), a severe congestion alarm suddenly breaks out at the "Crossroad_05" intersection in the lower right corner. After the system executes this step, it will immediately mark the intersection area in the lower right corner as an area of ​​interest as well. In the subsequent rendering and transmission cycle, the LED sub-screen corresponding to the intersection area will be switched to the first LED sub-screen, receiving a high frame rate data stream of 60Hz. This ensures that the sudden high-risk alarm image can be presented on the large screen in a clear, smooth, and low-latency state, actively attracting the commander's attention through visually dynamic effects (such as smooth flashing red light and smooth traffic animation), realizing an active interactive mode.

[0063] Optionally, this solution also includes steps S109-S112; Step S109: Calculate the total transmission bandwidth occupancy rate required to send the first image data sequence and the second image data sequence; The total transmission bandwidth utilization rate refers to the ratio of the data transmission rate consumed by all image data generated in the current frame (including the first image data sequence with a high frame rate and the second image data sequence with a low frame rate) in the physical transmission channel to the maximum physical bandwidth of the channel, usually expressed as a percentage; the preset security threshold is the upper limit of bandwidth utilization set to prevent network congestion, packet loss or system overheating, for example, set to 90% of the gigabit network bandwidth (i.e., 900Mbps).

[0064] This step is performed during the pre-verification phase before data transmission. Specifically, after generating the initial first and second image data sequences, the system does not send them immediately but first performs a virtual data volume count. The system calculates the instantaneous bitrate of the first part based on the total number of pixels, color depth (e.g., 24-bit or 32-bit), and first transmission frame rate (e.g., 60Hz) of the first LED sub-screen; similarly, it calculates the instantaneous bitrate of the second part based on the total number of pixels, color depth, and second transmission frame rate (e.g., 10Hz) of the second LED sub-screen. The two are added together to obtain the total instantaneous bitrate. Subsequently, the system divides this total instantaneous bitrate by the total bandwidth limit of the physical link. For example, in a smart city command center, the system calculates that the current high frame rate area (area of ​​interest) occupies 40% of the screen, and the low frame rate area occupies 60%. The calculated total data flow requires 950Mbps of bandwidth. Since the physical interface is a gigabit Ethernet port (1000Mbps), the calculated total transmission bandwidth utilization rate is 95%. This value will be sent to subsequent steps for judgment.

[0065] Step S110: If the total transmission bandwidth occupancy rate is greater than the preset safety threshold, repeat steps S1-S4 until the theoretical total transmission bandwidth occupancy rate is less than or equal to the preset safety threshold, thus obtaining the final first LED sub-screen and the final second LED sub-screen. This step, executed in the exception handling branch when bandwidth verification fails, is a typical iterative optimization process. Specifically, when the total transmission bandwidth utilization calculated in step S109 exceeds a preset safety threshold (e.g., exceeding 90%), it indicates that the current area of ​​interest is too large, and forcibly sending data would lead to network congestion. At this point, the system enters a WhileLoop logic, continuously fine-tuning parameters and recalculating until a balance is found that satisfies bandwidth limitations while preserving as large an area of ​​interest as possible. Once the condition is met, the loop terminates, and the determined partition status at this point is "Final First LED Sub-screen" and "Final Second LED Sub-screen".

[0066] S1: Iteratively reduce the area of ​​the region of interest according to a preset reduction step size; The preset shrinkage step size refers to the fixed distance or proportion parameter by which the boundary of the region of interest shrinks inward during each iteration, such as reducing the radius by 50 pixels or the area by 5% each time.

[0067] This step is the initial action of the loop. Specifically, the system reads the current region of interest parameters (such as radius R). In the first loop, the system subtracts a step size from the radius (R_new = R - 50); in the second loop, if the bandwidth is still insufficient, it continues to subtract the step size (R_new = R - 100). For example, the initial region of interest is a circle with a radius of 1000 pixels, covering half of the screen. Due to excessive bandwidth, the system shrinks it to a circle with a radius of 950 pixels in the first loop, attempting to reduce the total data volume by reducing the number of high frame rate pixels.

[0068] S2: Determine the theoretical first LED sub-screen and the theoretical second LED sub-screen based on the reduced area of ​​interest; Among them, the theoretical first / second LED sub-screen refers to the partition state temporarily assumed during the iterative calculation process, and has not yet taken effect.

[0069] This sub-step is performed after the region parameter adjustment. Referring to step S102, specifically, the system rescans the physical topology of the LED screen using the reduced new area of ​​interest (e.g., a circle with a radius of 950 pixels) from step S1. The system determines which LED sub-screens are still located within this reduced circle. Some sub-screens that were originally located at the edge of the circle may be excluded due to the circle's reduction in size. These excluded sub-screens are reclassified from "theoretically first LED sub-screen" (high frame rate) to "theoretically second LED sub-screen" (low frame rate). Through this step, the system effectively reduces the number of hardware units requiring high frame rate drive.

[0070] S3: Generate the theoretical first image data sequence corresponding to the theoretical first LED sub-screen and the theoretical second image data sequence corresponding to the theoretical second LED sub-screen; This sub-step is executed during the data volume estimation phase. Specifically, the system does not need to call the rendering engine to generate images, but instead performs mathematical calculations directly. The system counts the total number of pixels contained in the "theoretical first LED sub-screen," multiplies it by the color depth and the standard frame rate (60Hz), and obtains the theoretical data volume of the high frame rate portion; it counts the total number of pixels in the "theoretical second LED sub-screen," multiplies it by the color depth and the low frame rate (10Hz), and obtains the theoretical data volume of the low frame rate portion.

[0071] S4: Calculate the theoretical total transmission bandwidth utilization rate required to send the theoretical first image data sequence and the theoretical second image data sequence; This sub-step is the judgment node of the loop. Specifically, the system adds the two theoretical data amounts calculated in step S3 and divides them by the total physical link bandwidth to obtain the new bandwidth utilization rate. For example, after the first round of reduction (radius reduced from 1000 to 950), the calculated utilization rate drops from 95% to 91%. The system compares this 91% with the preset safety threshold (90%). If 91% is still greater than 90%, the next round of the loop is triggered, and the system jumps back to step S1 to continue reducing the radius; if after multiple rounds of reduction, the utilization rate finally drops to 89%, which is less than 90%, the loop is exited, and the current theoretical sub-screen state is locked as the final state, determining the final first LED sub-screen.

[0072] Step S111: Generate the final first image data sequence corresponding to the final first LED sub-screen and the final second image data sequence corresponding to the final second LED sub-screen; Among them, the final first LED sub-screen refers to the set of LED sub-screens that are finally confirmed as high frame rate display areas after bandwidth verification and possible area reduction adjustments; the final second LED sub-screen refers to the set of LED sub-screens that correspond to the final low frame rate display areas; the final first / second image data sequence refers to the image data stream that is re-rendered or extracted based on the adjusted partitioning scheme and is ready to be actually sent.

[0073] This step is performed during the data generation phase after the bandwidth adaptive adjustment is completed. Specifically, if the loop judgment logic in step S110 triggers the shrinking of the area of ​​interest (e.g., shrinking from a radius of 500 pixels to 300 pixels), or if the shrinking is not triggered (due to sufficient bandwidth), the system will obtain a confirmed and safe partitioning scheme. Based on this final partitioning scheme, the system calls the rendering engine or image extraction module again. For areas belonging to the "final first LED sub-screen," the system generates high frame rate data of 60Hz; for areas belonging to the "final second LED sub-screen," the system generates low frame rate data of 10Hz. For example, because the previously calculated 95% bandwidth exceeded the limit, after the system shrinks the area of ​​interest, the edge parts that originally belonged to the high frame rate area are now classified as low frame rate areas. The system regenerates the data accordingly, ensuring that the data stream generated this time will absolutely not overwhelm the network bandwidth.

[0074] Step S112: Determine the final first image data sequence as the first image data sequence, and determine the final second image data sequence as the second image data sequence; This step, executed in the final stage of data preparation, serves as the bridge between logical computation and physical transmission. Specifically, the system formally marks the "final first image data sequence" generated in step S111 as the "first image data sequence" to be sent, and the "final second image data sequence" as the "second image data sequence" to be sent. This assignment operation confirms that the data packets to be sent through the network interface have undergone bandwidth security checks. Subsequently, the system will jump back to steps S104 and S105 in the main process, driving the hardware transmitting card to actually send these data to the LED receiving card.

[0075] Through this series of steps, the system constructs a closed-loop control mechanism based on bandwidth constraints: First, an initial display partitioning scheme is formulated according to user needs. Then, the transmission load of this scheme is pre-calculated. If the load exceeds the physical link limit, the display area is automatically iteratively reduced until the load meets safety requirements, and finally, a verified image data stream is output. This mechanism ensures that the system eliminates the possibility of bandwidth overload before sending data, effectively preventing a surge in data volume caused by excessive user viewing area or simultaneous triggering of multiple alarms. This avoids packet loss, lag, or system crashes caused by network congestion, ensuring the stable operation of the LED display system.

[0076] Optionally, steps S10305-S10306 are another more specific steps for generating the second image data sequence; Step S10305: Extract pixel motion vector data corresponding to the second LED sub-screen from the rendering pipeline of the digital twin model; The rendering pipeline refers to the pipelined processing of the graphics processing unit (GPU) to convert a three-dimensional geometric model into a two-dimensional screen image, which typically includes stages such as vertex processing, rasterization, and fragment processing. Pixel motion vectors refer to two-dimensional vector field data generated during the rendering process that describes the positional changes of each pixel in screen space between the current frame and the previous frame. They are usually stored in the Velocity Pass channel of the G-Buffer (geometric buffer).

[0077] This step is executed during the parallel processing phase of low-frame-rate data stream generation, leveraging the low-level data access capabilities of modern graphics engines. Specifically, when the rendering engine generates image frame data for the second LED sub-screen, the system synchronously accesses the GPU's video memory buffer. The system directly reads the motion vector texture output by the rendering pipeline. This texture map has the same resolution as the RGB image, but instead of storing a color value, each pixel stores the displacement vector (Δx, Δy) of the object's surface on the screen. For example, in a digital twin factory scene, there is a rotating fan in the background area. For a specific pixel on the fan blades, the rendering engine precisely calculates that it has moved 3 pixels to the right and 2 pixels down in this frame relative to the previous frame. The system directly extracts this vector, which extremely accurately reflects the object's true motion trend without the need for estimation through image analysis algorithms.

[0078] Step S10306: Encapsulate the pixel motion vector data into the second image data sequence, wherein the pixel motion vector data includes the displacement direction and distance of corresponding pixels between adjacent frame images, and is used to calculate the intermediate transition frame image; This step is performed during the data packet transmission phase. Specifically, to transmit image content and control information simultaneously within limited bandwidth, the system uses channel multiplexing technology to encapsulate the data. The system typically quantizes and compresses the pixel motion vector data extracted in step S10305 (e.g., mapping it to a 7-bit integer) and concatenates it with the 1-bit binary mask data generated in step S10303. The concatenated 8-bit data is written into the Alpha channel (transparency channel) of the video signal. At this point, each pixel in the second image data sequence not only contains RGB color values ​​but also carries instructions via the Alpha channel regarding "whether interpolation (Mask)" and "how to perform interpolation (Vector)". Although this encapsulation method increases the data bit width per frame, because the overall transmission frame rate of the second image data sequence is extremely low (e.g., 10Hz), its total bandwidth usage is still far lower than the bandwidth required to directly transmit the original image at a high frame rate (e.g., 60Hz). After receiving the sequence, the second LED receiver card parses the Alpha channel data and performs pixel remapping of the RGB image according to the motion vector indication, thereby reconstructing a smooth intermediate transition frame locally. Through this processing, it ensures that non-interested areas maintain visual smoothness and clarity even under low bandwidth transmission, avoiding interference with the user's attention to the core area due to background image stuttering or blurring, while maintaining the consistency of the entire screen display.

[0079] Optionally, this solution also includes steps S113-S114; Step S113: When the target first LED sub-screen is determined to be the second LED sub-screen, the target first LED sub-screen is determined to be the first LED sub-screen within a preset time period; Here, the target first LED sub-screen refers to a display unit that was in a high frame rate display state (i.e., within the area of ​​interest) at the previous moment, but should be classified as a low frame rate display state (i.e., moved out of the area of ​​interest) at the current moment according to calculation; being determined as the second LED sub-screen means that, according to the real-time calculation logic of steps S101 to S102, the sub-screen is no longer within the user's current line of sight or voice command coverage range; the preset duration refers to the state switching buffer period or hysteresis time set by the system, such as 500 milliseconds or 1 second; being determined as the first LED sub-screen means forcibly maintaining its high frame rate transmission and display state, without immediately performing a degradation operation.

[0080] This step is performed during the transition phase of dynamic switching of display areas, aiming to introduce anti-jitter and visual persistence protection mechanisms. Specifically, when a user's eyes observe a screen, they often experience unconscious microsaccades or their gaze rapidly scans back and forth between two nearby targets. If the system responds to these minute changes in gaze completely in real time, it can cause some LED sub-screens located at the edge of the area of ​​interest to repeatedly switch between "high frame rate" and "low frame rate" modes at high frequency, resulting in screen brightness flickering or image tearing. Therefore, when the system detects that the user's gaze has moved away from a sub-screen that was originally in the core area (i.e., the sub-screen is determined to be downgraded to the second LED sub-screen), the system does not immediately execute the downgrade instruction, but instead starts a countdown timer. Within a preset duration (e.g., 500ms), the system continues to send 60Hz high frame rate data to that sub-screen, maintaining its display state as the first LED sub-screen.

[0081] Step S114: When the preset time period is exceeded, the target first LED sub-screen is determined as the second LED sub-screen; This step is performed during the final confirmation phase of the state transition. Specifically, if the user's gaze does not return to the sub-screen area after a preset time (e.g., 500ms), the system confirms that this is a genuine and stable gaze shift, rather than a brief eye-tracking interference. At this point, the system updates the state of the target first LED sub-screen to the second LED sub-screen. Subsequently, the system begins sending a low frame rate (e.g., 10Hz) second image data sequence to the sub-screen and instructs the receiving card to enable interpolation calculation. Through this "delayed confirmation" strategy, the system effectively smooths out the control signal jitter caused by rapid gaze movement, ensuring a smooth and stable display mode switching process and improving the user's visual comfort.

[0082] The LED screen partition display device in the embodiments of this invention is described below from the perspective of hardware processing. Please refer to [link / reference]. Figure 3 This is a schematic diagram of the physical device structure of an LED screen partition display device in the embodiments of this application.

[0083] It should be noted that, Figure 3 The structure of the LED screen partition display device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of the present invention.

[0084] like Figure 3 As shown, the LED screen partition display device includes a CPU 301, which can perform various appropriate actions and processes according to a program stored in the read-only memory ROM 302 or a program loaded from the storage section 308 into the random access memory RAM 303, such as executing the methods described in the above embodiments. The RAM 303 also stores various programs and data required for system operation. The CPU 301, ROM 302, and RAM 303 are interconnected via a bus 304. An I / O interface 305 is also connected to the bus 304.

[0085] The following components are connected to I / O interface 305: input section 306 including audio input devices, push-button switches, etc.; output section 307 including a liquid crystal display (LCD) and audio output devices, indicator lights, etc.; storage section 308 including a hard disk, etc.; and communication section 309 including a network interface card such as a LAN (Local Area Network) card, modem, etc. Communication section 309 performs communication processing via a network such as the Internet. Drive 310 is also connected to I / O interface 305 as needed. Removable media 311, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 310 as needed so that computer programs read from them can be installed into storage section 308 as needed.

[0086] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing computer programs for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 309, and / or installed from removable medium 311. When the computer program is executed by CPU 301, it performs the various functions defined in the present invention.

[0087] It should be noted that specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0088] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. Each block in a flowchart or block diagram may represent a module, program segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those shown in the drawings.

[0089] Specifically, the LED screen partition display device of this embodiment includes a processor and a memory. The memory stores a computer program. When the computer program is executed by the processor, it implements the LED screen partition display method provided in the above embodiment.

[0090] In another aspect, the present invention also provides a computer-readable storage medium, which may be included in the LED screen partition display device described in the above embodiments; or it may exist independently and not assembled into the LED screen partition display device. The storage medium carries one or more computer programs, which, when executed by a processor of the LED screen partition display device, cause the LED screen partition display device to implement the LED screen partition display method provided in the above embodiments.

[0091] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.

Claims

1. A method for partitioned display of an LED screen, characterized in that, The method includes: The system acquires multimodal interaction data between the user and the digital twin model displayed on the LED screen, and determines the area of ​​interest of the LED screen based on the multimodal interaction data. The LED screen includes several LED sub-screens. The LED sub-screen corresponding to the area of ​​interest is designated as the first LED sub-screen, and the LED sub-screens other than the first LED sub-screen are designated as the second LED sub-screen. The first image data sequence corresponding to each of the first LED sub-screens and the second image data sequence corresponding to each of the second LED sub-screens are generated according to the digital twin model; wherein, the first transmission frame rate corresponding to the first image data sequence is equal to the preset standard display frame rate, and the second transmission frame rate corresponding to the second image data sequence is lower than the first transmission frame rate. The first image data sequence is sent to the first LED receiving card corresponding to the first LED sub-screen, so that the first LED receiving card drives the first LED sub-screen to display according to the first image data sequence; The second image data sequence is sent to the second LED receiving card corresponding to the second LED sub-screen, so that the second LED receiving card performs interpolation calculations based on the second image data sequence to generate a number of intermediate transition frame images. The second image data sequence and the intermediate transition frame images are used to generate a display driving signal equal to the standard display frame rate, and the second LED sub-screen is controlled to display according to the display driving signal.

2. The method according to claim 1, characterized in that, The multimodal interaction data includes: the user's eye movement image and the user's voice interaction commands. Determining the area of ​​interest for the LED screen based on the multimodal interaction data specifically includes: Calculate the physical line of sight based on the eye movement image; The semantic gaze point is obtained by parsing the voice interaction command through a pre-trained semantic parsing model. The intent confidence of the voice interaction command is calculated using a pre-trained NLP model; When the confidence level of the intent is greater than a preset threshold, the coordinates of the semantic gaze point are determined as the center coordinates of the region of interest. When the confidence level of the intent is less than or equal to the preset threshold, the coordinates of the physical line of sight landing point are determined as the center coordinates of the area of ​​interest. The region of interest is determined based on the preset radius and the center coordinates.

3. The method according to claim 1, characterized in that, The second image data sequence corresponding to each of the second LED sub-screens is generated based on the digital twin model, specifically including: Image frame data corresponding to the second LED sub-screen is generated based on the digital twin model; Extract the attributes of text objects and dynamic chart objects located within the display area of ​​the second LED sub-screen in the digital twin model; Based on the pixel positions of the text object attributes and the dynamic chart object attributes in the image frame data, a binarized mask layer aligned with the pixels of the image frame data is generated. The image frame data and the binarization mask layer are encapsulated into the second image data sequence, wherein the binarization mask layer is used to instruct the second LED receiving card, when generating the intermediate transition frame, to reuse the pixel value of the corresponding position in the previous frame image data of the intermediate transition frame as the pixel value of the intermediate transition frame for the target area covered by the binarization mask layer.

4. The method according to claim 1, characterized in that, The method further includes: Obtain event data generated by the background monitoring system; When an alarm event exists in the event data, the target area corresponding to the alarm event in the digital twin model is obtained; The target region is marked as the region of interest.

5. The method according to claim 1, characterized in that, Before sending the first image data sequence to the first LED receiving card corresponding to the first LED sub-screen, the method further includes: Calculate the total transmission bandwidth utilization required to send the first image data sequence and the second image data sequence; If the total transmission bandwidth utilization rate is greater than the preset safety threshold, steps S1-S4 are executed repeatedly until the theoretical total transmission bandwidth utilization rate is less than or equal to the preset safety threshold, thus obtaining the final first LED sub-screen and the final second LED sub-screen. S1: Iteratively reduce the area of ​​the region of interest according to a preset reduction step size; S2: Determine the theoretical first LED sub-screen and the theoretical second LED sub-screen based on the reduced area of ​​interest; S3: Generate the theoretical first image data sequence corresponding to the theoretical first LED sub-screen and the theoretical second image data sequence corresponding to the theoretical second LED sub-screen; S4: Calculate the theoretical total transmission bandwidth utilization rate required to send the theoretical first image data sequence and the theoretical second image data sequence; Generate the final first image data sequence corresponding to the final first LED sub-screen and the final second image data sequence corresponding to the final second LED sub-screen; The final first image data sequence is determined as the first image data sequence, and the final second image data sequence is determined as the second image data sequence.

6. The method according to claim 1, characterized in that, The generation of a second image data sequence corresponding to each of the second LED sub-screens based on the digital twin model specifically includes: Extract the pixel motion vector data corresponding to the second LED sub-screen from the rendering pipeline of the digital twin model; The pixel motion vector data is encapsulated into the second image data sequence, wherein the pixel motion vector data includes the displacement direction and distance of corresponding pixels between adjacent frame images, and is used to calculate the intermediate transition frame image.

7. The method according to claim 1, characterized in that, The method further includes: When the target first LED sub-screen is determined to be the second LED sub-screen, the target first LED sub-screen is confirmed as the first LED sub-screen within a preset time period; When the preset time period is exceeded, the target first LED sub-screen will be identified as the second LED sub-screen.

8. An LED screen zoned display device, characterized in that, The LED screen partition display device includes: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to cause the LED screen partition display device to perform the method as described in any one of claims 1-7.

9. A computer-readable storage medium comprising instructions, characterized in that, When the instruction is executed on the LED screen partition display device, the LED screen partition display device performs the method as described in any one of claims 1-7.

10. A computer program product, characterized in that, When the computer program product is run on the LED screen partition display device, the LED screen partition display device performs the method as described in any one of claims 1-7.