Screen projection remote control method and system based on gap NEXT

By constructing a port mapping channel and a device adaptation mechanism, screen captures are performed on both the HarmonyOS NEXT real device and the simulator, solving the compatibility and dynamic perception issues of the screen projection remote control solution. This achieves a low-latency and highly stable remote control effect, improving user experience and operational accuracy.

CN121209818APending Publication Date: 2025-12-26BEIJING BANGCLE TECH CO LTD

Patent Information

Application Number
CN202511771383.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-28
Publication Date
2025-12-26

AI Technical Summary

Technical Problem

Existing screen mirroring and remote control solutions have compatibility issues with the HarmonyOS NEXT system. They cannot efficiently intercept system-level graphics function calls, resulting in high screen capture latency and high resource consumption. Furthermore, they lack dynamic perception capabilities in emulator remote control scenarios, leading to network transmission congestion and insufficient operation accuracy.

Method used

A port mapping channel is constructed to determine the device type, and different screen capture strategies are adopted for HarmonyOS NEXT real device and simulator environments. The original screen frames or simulator frames are obtained by injecting screen projection tools. Combined with video encoding and interactive command processing, low latency and high stability remote control are achieved.

Benefits of technology

It enables cross-platform, low-latency screen projection and interactive control, improving remote control accuracy and user experience in HarmonyOS NEXT devices and simulator environments, and ensuring smooth screen display and stable transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121209818A_ABST
    Figure CN121209818A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image data processing, in particular to a screen projection remote control method and system based on the gap NEXT. The method comprises the following steps: constructing a port mapping channel, and judging the type of current equipment; if the equipment type is a mobile phone, a preset screen projection tool is injected into a controlled end of the swan NEXT system, and screen picture content of the controlled end is extracted through the preset screen projection tool to form picture data; if the equipment type is a simulator, intercepting simulator operation frames according to a preset frame rate, and combining the simulator operation frames into a continuous image sequence to form picture data; performing video coding on the picture data, and transmitting the picture data to a control end in real time through a port mapping channel; and the control end receives, decodes, renders and displays to obtain a visual picture. By establishing an adaptive port mapping channel and a differentiated picture acquisition mechanism, cross-platform low-delay screen projection and high-precision remote control between the real machine and the simulator of the gap NEXT are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image data processing technology, and in particular to a remote control method and system for screen projection based on HarmonyOS NEXT. Background Technology

[0002] Existing screen mirroring and remote control solutions exhibit significant technical limitations and compatibility issues when dealing with the rapidly evolving HarmonyOS ecosystem, particularly the HarmonyOS NEXT system and its emulator. Firstly, most mainstream solutions on the market are based on the Android system's underlying graphics interface or general screen recording protocols. Their technical architecture is incompatible with HarmonyOS NEXT's self-developed system architecture and graphics services, making direct deployment and stable operation on HarmonyOS NEXT devices impossible. Even with compatibility layers, the inability to efficiently intercept system-level graphics function calls often results in high screen capture latency, high system resource consumption, and an inability to meet the requirements for smooth real-time interaction.

[0003] Secondly, in simulator remote control scenarios, existing technologies generally suffer from a lack of simplistic processing mechanisms. Most solutions rely solely on timed screenshots, lacking dynamic awareness of system load and network status, and failing to intelligently adjust output strategies when resources are scarce. This often results in excessively large amounts of encoded data, causing network congestion and screen lag on the control end, severely impacting user experience and operational efficiency. Furthermore, the insufficient precision in command mapping and parsing between the simulator and the control end makes it difficult to achieve the same operational accuracy as physical devices. Summary of the Invention

[0004] Therefore, it is necessary for the present invention to provide a remote control method and system for screen projection based on HarmonyOS NEXT, in order to solve at least one of the above-mentioned technical problems.

[0005] To achieve the above objectives, a remote control method for screen projection based on HarmonyOS NEXT includes the following steps: Step S1: Construct a port mapping channel and determine the current device type; Step S2: If the device type is a mobile phone, a preset screen mirroring tool is injected into the controlled terminal of the HarmonyOS NEXT system. The screen content of the controlled terminal is extracted through the preset screen mirroring tool to form screen data. Step S3: If the device type is an emulator, then capture the emulator running frames according to the preset frame rate and combine them into a continuous image sequence to form screen data; Step S4: Perform video encoding on the image data and transmit it to the control terminal in real time through the port mapping channel; the control terminal receives and decodes the data, renders and displays the image to obtain a visual image; Step S5: Generate user interaction instructions on the control end and return them to the controlled end through the port mapping channel; the controlled end parses the user interaction instructions and executes the corresponding operations.

[0006] Preferably, the present invention also provides a HarmonyOS NEXT-based screen projection remote control system for executing the above-described HarmonyOS NEXT-based screen projection remote control method, wherein the HarmonyOS NEXT-based screen projection remote control system includes: The port mapping and device identification module is used to build a port mapping channel and determine the current device type; The function is injected into the screen capture module. If the device type is a mobile phone, a preset screen projection tool is injected into the controlled end of the HarmonyOS NEXT system. The screen content of the controlled end is extracted through the preset screen projection tool to form screen data. The simulator frame capture module is used to capture simulator running frames according to a preset frame rate and combine them into a continuous image sequence to form screen data if the device type is a simulator. The video encoding / decoding and rendering transmission module is used to encode the image data and transmit it to the control terminal in real time through the port mapping channel; the control terminal receives and decodes the data, and renders and displays the image to obtain a visual image. The remote interactive command parsing and control module is used to generate user interaction commands at the control end and return them to the controlled end through a port mapping channel; the controlled end parses the user interaction commands and executes the corresponding operations.

[0007] This invention establishes differentiated screen capture and remote control strategies for HarmonyOS NEXT real devices and simulator environments by constructing port mapping channels and device type adaptive mechanisms, achieving cross-platform, low-latency screen projection and interactive control. On the mobile device side, an injection tool intercepts the underlying graphics functions of the HarmonyOS system, directly acquiring raw screen frames and avoiding the high latency and resource consumption of traditional screen recording protocols. On the simulator side, an adjustable frame rate screenshot compositing mechanism dynamically reduces frame rates during network congestion or high system load, significantly improving screen smoothness and transmission stability. The system as a whole employs end-to-end video encoding and encrypted channel transmission. The control end can decode and render in real time and achieve bidirectional synchronization of input events, forming a stable remote control closed loop, thereby effectively improving the remote control accuracy and user experience in HarmonyOS NEXT devices and simulator environments. Attached Figure Description

[0008] Other features, objects, and advantages of the invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 This is a flowchart illustrating the steps of a remote control method for screen projection based on HarmonyOS NEXT according to the present invention. Figure 2 Flowchart of differentiated screen capture for mobile phones and emulators in this embodiment of the invention; Figure 3 This is a schematic diagram of a system attribute simulation system based on HarmonyOS NEXT screen projection remote control according to the present invention. Detailed Implementation

[0009] The technical method of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0010] Furthermore, the accompanying drawings are merely illustrative of the invention and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor methods and / or microcontroller methods.

[0011] It should be understood that although the terms "first," "second," etc., may be used herein to describe various units, these units should not be limited by these terms. These terms are used merely to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, a first unit may be referred to as a second unit, and similarly, a second unit may be referred to as a first unit. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0012] To achieve the above objectives, please refer to Figures 1 to 3 This invention provides a remote control method for screen projection based on HarmonyOS NEXT, the method comprising the following steps: Step S1: Construct a port mapping channel and determine the current device type; Step S2: If the device type is a mobile phone, a preset screen mirroring tool is injected into the controlled terminal of the HarmonyOS NEXT system. The screen content of the controlled terminal is extracted through the preset screen mirroring tool to form screen data. Step S3: If the device type is an emulator, then capture the emulator running frames according to the preset frame rate and combine them into a continuous image sequence to form screen data; Step S4: Perform video encoding on the image data and transmit it to the control terminal in real time through the port mapping channel; the control terminal receives and decodes the data, renders and displays the image to obtain a visual image; Step S5: Generate user interaction instructions on the control end and return them to the controlled end through the port mapping channel; the controlled end parses the user interaction instructions and executes the corresponding operations.

[0013] Preferably, step S1 includes: Establish a port mapping channel and complete a two-way handshake; generate a device attribute dataset based on the system attribute data collected during the handshake process. The control unit parses the runtime environment type field in the device attribute dataset and determines the device category through the system kernel identifier and virtualization flag. If a real hardware serial number and system underlying driver identifier are detected, it is marked as mobile device type data; If a virtual display driver and emulation environment flag is detected, it is marked as emulator type data.

[0014] In this embodiment of the invention, the control terminal initiates the connection process by first sending a connection request to the controlled terminal via a secure network protocol (such as TCP protocol based on TLS encryption). After the connection is established, both parties perform a two-way handshake protocol. This process is not only a simple connectivity test, but also a security verification and information exchange.

[0015] During the handshake process, the controlling end sends a signed authentication data packet to the controlled end. After the controlled end verifies the signature, it collects a system attribute dataset from its own system layer and returns it to the controlling end. This dataset includes, but is not limited to, the following key fields: hardware platform identifier (ro.hardware), device product model (ro.product.model), virtualization environment presence identifier (sys.hypervisor.present), graphics driver vendor (graphics.vendor), device serial number (ro.serialno), and system kernel version (sys.os.kernel_version).

[0016] After receiving the device attribute dataset, the control unit passes it to the built-in device identification engine for parsing. The core judgment logic of the engine is based on the joint analysis of the system kernel identifier and virtualization flag bits.

[0017] Path 1: Identify the mobile phone The engine first checks the `ro.serialno` field. If the value of this field is a non-virtualized default string that conforms to the encoding rules of the International Mobile Equipment Identity (IMEI) or other real hardware serial numbers, a first positive weight is generated. Next, the engine parses the `graphics.vendor` field. If the value of this field is a known physical GPU driver vendor (such as "Kirin GPU" or "ARM Mali"), rather than a virtualized graphics driver (such as "QEMU VGA" or "VirtIO-GPU"), a second positive weight is generated. Simultaneously, the engine checks the `sys.hypervisor.present` field. If this field explicitly returns false or does not exist, indicating that the device is not running in a virtualized environment, a third positive weight is generated. When all the above positive weights are satisfied, and `sys.os.kernel_version` explicitly points to the HarmonyOS NEXT system kernel, the device identification engine ultimately determines that the device is "mobile device type data".

[0018] Path 2: Identify the simulator The engine checks the `sys.hypervisor.present` field. If this field returns true, a first strong pointing weight is generated, indicating that the device is in a virtualized environment. Next, the engine analyzes the `graphics.vendor` field. If this field value is a known emulator virtual display driver identifier (such as "QEMU VGA" or "VirtIO-GPU"), a second strong pointing weight is generated. Additionally, the engine checks the `ro.serialno` field. If this field value is empty, all zeros, or a default fixed serial number for the emulator (such as "EMULATOR123"), it further confirms that it is a virtual device. When the virtualization flag is detected as true, accompanied by a virtual display driver identifier, the device identification engine can determine that the device is "emulator type data," eliminating the need to wait for the hardware serial number timeout response, thus speeding up the identification process.

[0019] Most importantly, the specific steps for building a port mapping channel are as follows: The controlled terminal registers with the designated signaling server and reports its network address information; The control terminal initiates a request to the signaling server to connect to the controlled terminal; The signaling server exchanges address information between the two parties and assists them in performing NAT traversal and hole punching. Establish an end-to-end encrypted data transmission channel based on UDP between the controlled end and the control end.

[0020] In this embodiment of the invention, when the controlled terminal (such as a mobile phone or emulator running the HarmonyOS NEXT system) is started, the communication initialization module in the system will send a registration request to the preset signaling server.

[0021] During the registration process, the controlled terminal packages its own network attribute information into a network address information packet and reports it. The information includes: local IP address and port number; network type identifier (such as Wi-Fi, wired, cellular network, etc.); NAT type parameters (Full Cone, Restricted Cone, Symmetric, etc.); encrypted handshake identifier and device unique identifier (DeviceID).

[0022] After receiving the registration message, the signaling server records the address information of the controlled device in the internal device mapping table and returns a registration confirmation response, indicating that the controlled device can accept remote connections.

[0023] When the control terminal (such as a computer running a remote control management program) starts up, the connection management module sends a connection request message to the same signaling server. The request field contains the Device ID and authentication credentials of the target controlled terminal.

[0024] After verifying the identity of the controlling end, the signaling server retrieves the current network address of the target controlled end from the mapping table and assigns a session channel number (Session ID) to both parties.

[0025] In this process, the signaling server acts only as an intermediary control channel and does not directly convert audio, video or control data, thereby reducing system latency and bandwidth load.

[0026] The signaling server sends the public network address information of the control end and the controlled end to each other for subsequent point-to-point channel establishment.

[0027] After receiving each other's addresses, both parties complete the NAT traversal process through the UDP hole punching mechanism.

[0028] Specifically, both parties send empty UDP probe packets to each other's public IP addresses almost simultaneously. These packets carry: Session ID; timestamp; handshake nonce; and encryption negotiation field.

[0029] Once the probe packet successfully penetrates the NAT gateways of both parties, a bidirectional UDP data path can be established. If the initial hole punching fails, the system will attempt multi-port probing and timed retry strategies to adapt to different network environments.

[0030] To enhance the success rate of penetration, the system automatically enables the relay policy (TURN mode) when a symmetric NAT environment is detected, and the signaling server temporarily proxies the communication of the first packet to ensure that the channel is successfully established.

[0031] After successful NAT hole punching, the control end and the controlled end establish an end-to-end encrypted data transmission link through a UDP channel.

[0032] During the channel establishment phase, both parties first complete the key exchange and authentication process.

[0033] Based on the random number negotiated during the handshake phase and the unique device identifier, both parties execute the Diffie-Hellman (DH) key exchange algorithm to generate a symmetric encryption key; The AES-GCM algorithm is used to encrypt subsequent data packets during transmission, and an integrity check code (HMAC) is attached to the packet header.

[0034] This encryption mechanism prevents data from being intercepted or tampered with during transmission over the public network, ensuring the security of remote control operations and screen display transmission.

[0035] After the channel is established, the system maintains the transmission session context in the memory of both parties, including: the current port number and transmission status; bandwidth usage, packet loss rate and round-trip time (RTT); connection heartbeat cycle and reconnection threshold.

[0036] The control end and the controlled end will communicate through this channel to transmit all subsequent video data and interactive commands, including video encoded streams, input event command packets, and operation feedback frames.

[0037] During communication, both parties maintain the validity of the channel through a heartbeat detection mechanism.

[0038] Every preset period (e.g., 3 seconds), both parties exchange heartbeat packets, carrying the latest network status and frame sequence number information. If no response is received from the other party for several consecutive periods, the channel is determined to be interrupted, and the system automatically triggers a reconnection process, including: re-registering with the signaling server; re-requesting the peer's address; and performing a second NAT hole punching and key renegotiation.

[0039] This mechanism ensures that the port mapping channel can still self-recover and remain effective in the event of network fluctuations or routing changes.

[0040] Preferably, step S2 includes: If the device type is a mobile phone, the preset screen mirroring tool will be injected into the HarmonyOS NEXT system layer of the controlled device; The pre-set screen mirroring tool obtains the function call context, key parameters, and return values ​​in real time through Hook, forming function call data; Determine the original screen frame data based on function call data; Extract the screen content from the original screen frame data to form screen data.

[0041] In this embodiment of the invention, after determining that the device type is a mobile phone, the control terminal pushes a pre-prepared screen mirroring tool, compiled as an ARM64 architecture dynamic link library (.so file), to a designated system directory (e.g., / system / lib64 / ) on the controlled mobile phone through an established, high-privilege ADB debugging connection. Subsequently, using process injection technology (e.g., through the ptrace mechanism or by modifying the system environment variable LD_PRELOAD), the screen mirroring tool is dynamically loaded into the address space of a critical graphics compositing service process (such as graphic_compositor_service) in the HarmonyOS NEXT system. This injection process ensures that the tool runs at the system's underlying layer and has permission to access core graphics resources.

[0042] After the screen mirroring tool is successfully injected, its built-in Hook manager begins to work. This manager utilizes Inline Hook or PLT / GOT Hook technology to redirect the key function in the graphics compositing service responsible for submitting the final frame to the display buffer. In this embodiment, the key function is GraphicBufferMapper::commitFrame(Surface surface, const FrameData&frameData).

[0043] When the system is about to refresh the screen and calls the commitFrame function, the execution flow will first jump to the proxy processing function set by the screen mirroring tool.

[0044] The proxy function performs the following operations: It retrieves the function's call context in real time, including thread ID, timestamp, and other information; it captures key parameters, extracting the `frameData` parameter, which contains a handle to a `GraphicBuffer` object pointing to the final screen image data. This handle is a core identifier for the HarmonyOS system's underlying management of graphics memory; and it captures the return value, recording the status code returned after the original function's successful execution.

[0045] All of this information together constitutes the function call data and is temporarily stored in a shared memory area.

[0046] The screen mirroring tool's frame extraction module locates the crucial GraphicBuffer handle by analyzing the aforementioned function call data. Subsequently, the module locks the graphics buffer by calling a system-privileged graphics interface provided by the HarmonyOS NEXT system (such as GraphicBuffer::Lock or AHardwareBuffer_lock).

[0047] After the locking operation is successful, the frame extraction module can directly access the original address of the buffer in memory and read the raw screen frame data, which is usually in RGB_888 or RGBA_8888 format and is uncompressed. This data represents the most complete and real-time frame that the system will display on the screen.

[0048] The frame extraction module encapsulates the acquired raw screen frame data, along with its metadata such as resolution and color format, into a custom data packet structure. This screen data is then sent to the control terminal for encoding and transmission via the previously established port mapping channel.

[0049] Preferably, step S3 includes: If the device type is an emulator, the emulator's running frames are periodically captured according to the preset frame rate to generate a set of static image frames; The static image frame set is combined according to the time series to form a continuous image sequence as picture data; During the assembly process, the system load and network status are monitored in real time, and the output frame rate of the continuous image sequence is dynamically adjusted based on the detection results.

[0050] In this embodiment of the invention, after determining that the device type is an emulator, the system starts a dedicated frame capture service process on the host machine (the computer running the emulator). This service first reads the preset capture frame rate in the configuration file (set to 60fps in this embodiment) and calculates the capture period (approximately 16.67ms / frame) based on the frame rate.

[0051] The service establishes a connection through the virtual display interface (VirtIO-GPU display backend) provided by the HarmonyOS emulator to obtain access to the emulator's virtual display's framebuffer. Simultaneously, the service starts two monitoring threads: System load monitoring thread: continuously collects host machine CPU utilization and available memory; Network status monitoring thread: Continuously measures network round-trip time (RTT) and packet loss rate to the control terminal using the ping command.

[0052] The frame capture service operates as follows: a high-precision timer (Linux timerfd) is used to precisely trigger frame capture operations with a period of 16.67ms; each time it is triggered, a snapshot of the front buffer of the current virtual display is obtained through an ioctl system call; the acquired raw RGB data is converted into the standardized RGBA8888 format to generate static image frames; a high-precision timestamp (nanosecond level) and sequence number are added to each static image frame; and consecutive static image frames are stored sequentially into a circular buffer according to the timestamps to form a continuous image sequence.

[0053] During the image sequence assembly process, the system performs real-time resource awareness and frame rate adjustment.

[0054] CPU utilization thresholds: Level 1 warning threshold 70%, Level 2 severe threshold 85%; Network RTT thresholds: Level 1 warning threshold 50ms, Level 2 severe threshold 100ms; Packet loss rate thresholds: Level 1 warning threshold 3%, Level 2 severe threshold 8%.

[0055] When any metric is detected to exceed the first-level threshold, a gradual frame reduction process is initiated: Level 1 frame drop (mild resource shortage): The output frame rate is reduced from 60fps to 45fps; a "2 out of 3" frame dropping mode is adopted, retaining 2 frames out of every 3 frames; non-critical frames (B-frames that are non-reference frames in the encoding sequence) are dropped first.

[0056] Level 2 frame rate reduction (moderate resource shortage): The output frame rate is reduced from 45fps to 30fps; a "choose one of two" frame dropping mode is adopted; the GOP (Group of Pictures) length is increased and the keyframe density is reduced.

[0057] Level 3 frame drop (severe resource shortage): The output frame rate is reduced from 30fps to 15fps; a "4-out-of-1" frame dropping mode is adopted; only the continuity of key frames (I-frames) is guaranteed to ensure that the basic picture can be decoded.

[0058] When all monitoring indicators are below 80% of the first-level threshold for 5 consecutive cycles, the system initiates a gradual recovery process: gradually restoring the output frame rate by increasing it by 15fps every 2 seconds; continuously monitoring system indicators during the recovery process to prevent oscillations; and finally restoring to the initial target frame rate of 60fps.

[0059] To ensure image quality during frame rate reduction, the system implements the following optimizations: always prioritize the complete transmission of I-frames to maintain basic image recognizability; temporarily increase the frame rate when significant scene changes are detected through inter-frame difference analysis; and appropriately improve the encoding quality of single frames while reducing the frame rate to maintain the overall visual experience.

[0060] Preferably, dynamically adjusting the output frame rate of a continuous image sequence based on the detection results includes: If either the system CPU utilization or the network transmission latency is greater than or equal to a preset threshold, the system is deemed to be under resource pressure, and a frame reduction process is initiated. During the frame downsizing process, image frames that are not reference frames in the coding structure are discarded. When the system CPU utilization and network transmission latency are both less than the preset threshold and remain below the threshold for a period of time, the system resources are determined to be sufficient, and the output frame rate of the continuous image sequence is increased according to the preset frame rate increment until the preset capture frame rate is reached.

[0061] In this embodiment of the invention, the system establishes a multi-dimensional resource monitoring system and continuously collects the following key indicators: System CPU utilization: The global CPU utilization is obtained by reading / proc / stat, and user mode, system mode and idle time are distinguished; Network transmission delay: The round-trip time (RTT) between the control end and the network is measured using ICMP ping packets at a sampling frequency of 10 times / second; Additional monitoring metrics include system memory usage, GPU load, and network packet loss rate, which serve as auxiliary decision-making parameters.

[0062] The system presets the following hierarchical threshold system: Core threshold parameters: CPU utilization threshold (CT) 75%; network latency threshold (NT) 80ms; stable duration (ST) 3 seconds.

[0063] Auxiliary threshold parameters: Memory utilization threshold (MT) 85%, GPU load threshold (GT) 90%; Packet loss rate threshold (PT) 5%. The system determines a resource shortage state when any of the following conditions are met: Primary condition: CPU utilization ≥ CT or network latency ≥ NT; Auxiliary conditions: Memory usage ≥ MT and GPU load ≥ GT; Emergency condition: Network packet loss rate ≥ PT (immediately trigger frame rate reduction).

[0064] Once a resource shortage is triggered, the system initiates a tiered frame rate reduction process: Level 1 frame rate reduction (mild resource constraints): Target frame rate from 60fps to 45fps; the dropping strategy is to drop 1 B-frame (non-reference frame) every 4 frames; the encoding parameters are adjusted so that the GOP size is changed from 60 to 90.

[0065] Level 2 frame rate reduction (moderate resource shortage): target frame rate from 45fps to 30fps; the dropping strategy is to drop 1 B-frame and 1 P-frame every 3 frames; the encoding parameters are adjusted so that the GOP size is changed from 90 to 120.

[0066] Level 3 frame rate reduction (severe resource constraints): target frame rate from 30fps to 15fps; the discarding strategy is to retain only I-frames and key P-frames, and discard all B-frames; the encoding parameters are adjusted so that the GOP size is changed from 120 to 180.

[0067] The system identifies non-reference frames through the following mechanisms: Frame type analysis: Perform motion estimation and scene analysis on the image sequence before encoding; Dependency mapping: Establish an inter-frame dependency graph and mark the reference relationships of each frame; Priority sorting: High priority: I frame (key reference frame); Medium priority: P frame used as a reference; Low priority: B frame and non-reference P frame (preferred discard object).

[0068] When the system simultaneously meets the following conditions, it is determined that the resources are abundant: CPU occupancy rate < CT × 0.8 (i.e., < 60%); Network latency < NT × 0.7 (i.e., < 56ms); Sustained stable time ≥ ST (3 seconds).

[0069] After the resource abundance state is confirmed, the system starts the frame rate increase process: First-stage recovery: Target frame rate from 15fps to 30fps; The recovery timing is immediately executed after 3 seconds of stability; Frame quality optimization is to maintain a relatively high single-frame quality while increasing the frame rate.

[0070] Second-stage recovery: Target frame rate from 30fps to 45fps; The recovery timing is executed after continuing to be stable for 3 seconds; Encoding optimization is to gradually restore the normal GOP structure.

[0071] Final-stage recovery: Target frame rate from 45fps to 60fps (preset capture frame rate); The recovery timing is executed after being stable for another 3 seconds; Complete recovery means that all encoding parameters return to the initial optimized state.

[0072] To prevent frequent state switching, the system introduces the following protection measures: The frame rate decrease threshold is higher than the recovery threshold, forming a buffer interval; Any state must be maintained for at least 2 seconds before switching; The frame rate change adopts smooth transition to avoid sudden changes.

[0073] Preferably, performing video encoding on the picture data and transmitting it to the control end in real time through the port mapping channel includes: Input the picture data into the video encoding unit, and the video encoding unit adaptively selects encoding parameters according to the dynamic evaluation result of the network condition; Perform video encoding on the picture data based on the encoding parameters to generate an encoded video stream; The encoded video stream is transmitted to the control end in real time through the port mapping channel.

[0074] In the embodiment of the present invention, for the picture data generated by the controlled end, to achieve low-latency and high-stability remote display transmission, the system is configured with a video encoding unit and a port mapping transmission module at the controlled end, and the two cooperate to complete the compression and transmission of the picture data.

[0075] After the controlled device completes the acquisition of the screen content, the resulting image data is imported into the video encoding unit. The video encoding unit contains a frame buffer management module, a network status assessment module, and an adaptive encoding control module. Among them, the network status assessment module continuously collects network operation indicators, including bandwidth utilization, packet loss rate, average round-trip time (RTT), and jitter parameters, and calculates the available network bandwidth and transmission stability index periodically.

[0076] The adaptive coding control module dynamically selects the appropriate set of coding parameters based on the evaluation results. The coding parameters include inter-frame prediction mode, quantization step size (QP), rate control mode (CBR / VBR), and GOP structure configuration.

[0077] For example, when the network bandwidth is detected to be lower than the set threshold, the adaptive encoding control module will reduce the output bit rate and adjust the GOP length to reduce the transmission load; if the network condition is detected to return to stability, it will gradually restore to the target frame rate and bit rate to ensure a balance between image clarity and real-time performance.

[0078] After the encoding control module completes the parameter selection, the frame buffer management module sends the image data to the H.265 or H.264 compression engine according to the encoding timestamp order to perform image compression processing and form an encoded video stream that conforms to the Real-Time Transport Protocol (RTP) encapsulation format.

[0079] To ensure display synchronization between consecutive frames, the encoding unit has a built-in timing calibration mechanism that marks the key frames (I-frames) and reference frames (P / B-frames) with sequence identification and timestamps, thereby maintaining the continuity of the picture in the subsequent decoding stage.

[0080] The generated encoded video stream is sent to the control terminal via the port mapping transmission module. This module performs address mapping, stream fragmentation, and reassembly of the data based on the port mapping channel established in step S1.

[0081] The port mapping transmission module adopts a lightweight transmission protocol stack and reduces screen stuttering caused by network jitter through UDP multi-threaded asynchronous sending mechanism and sliding window buffering strategy.

[0082] Meanwhile, the module inserts a stream sequence number and a checksum field into the data header so that the control end can quickly detect and correct packet loss or sequence errors after receiving the data.

[0083] On the control end, the received encoded video stream will enter the corresponding video decoding unit for decoding and rendering, thereby realizing the real-time display of the controlled end's image on the control end.

[0084] The entire transmission path keeps the port mapping logical channel open at the application layer. If a network quality degradation or connection interruption is detected, the transmission module can automatically trigger the bandwidth reassessment and error recovery mechanism to restore the normal transmission rate without interrupting screen projection.

[0085] Preferably, the control terminal receives and decodes the data, and renders and displays the resulting visual image, including: After receiving the encoded video stream, the control terminal decodes it using the corresponding matching video decoding unit, outputs the original frame data, and renders and displays the original frame data as a visual image through a preset graphics rendering interface.

[0086] In this embodiment of the invention, after the video encoding unit of the controlled end sends the encoded video stream through the port mapping channel, the video stream receiving module of the control end continuously listens for data stream input on the corresponding logical port.

[0087] The receiving module verifies and reassembles each video data packet based on the established two-way handshake parameters (including device identifier, frame number, and check field).

[0088] When continuous packet loss or out-of-order frames are detected, the system restores the order by reordering the packet buffer queue and triggers a "frame reduction rendering mode" in the case of severe packet loss.

[0089] In frame rate reduction mode, the control unit automatically discards B-frames or P-frames that are not critical reference frames to ensure smoothness of the video and prevent high latency buildup.

[0090] To adapt to the network bandwidth and CPU performance of different devices, the control unit can dynamically adjust the size of the receive buffer and the queue depth. For example, a deeper buffer queue is used in high-bandwidth wired networks to ensure frame integrity, while the buffer depth is shortened in wireless networks to reduce decoding latency.

[0091] The reassembled encoded video stream enters the video decoding unit. This unit automatically matches the corresponding decoder instance based on the encoding standard used by the controlled device (such as H.264 / AVC or H.265 / HEVC).

[0092] In a typical implementation, the video decoding unit calls hardware acceleration interfaces (such as DXVA on the Windows platform, VA-API on the Linux platform, or MediaCodec interface on the HarmonyOS PC) to perform frame-level decoding operations in parallel with the GPU, thereby recovering the original frame data without significantly consuming CPU resources.

[0093] After each frame is decoded, the system writes the output raw frame data to the frame buffer management module. This module maintains a multi-level buffer structure: The pre-decoding buffer stores frames to be displayed; the display buffer holds the currently rendered frame; and the recycling buffer reclaims resources of already displayed frames.

[0094] When the frame buffer management module receives network status notifications or user operation delays, it can proactively discard expired frames to achieve frame synchronization adjustment and prevent image ghosting or frame errors.

[0095] After obtaining the raw frame data, the control terminal calls the preset graphics rendering interface module to complete the screen output.

[0096] This module supports rendering interfaces such as OpenGL, Vulkan, or Direct3D, and can automatically switch the rendering backend according to the operating system of the control terminal.

[0097] The rendering interface module first creates a frame buffer object (FBO) based on the resolution information of the video frame (such as aspect ratio, pixel format, and color space), and then maps the decoded frame to a GPU texture.

[0098] By performing YUV to RGB color space conversion and pixel interpolation resampling operations through fragment shaders, the output image remains clear and undistorted across different monitor resolutions.

[0099] Subsequently, the rendering interface module outputs the screen in the display thread at a fixed refresh rate (such as 60Hz) or an adaptive refresh method based on the VSync signal, thereby presenting the controlled screen in real time in the control window.

[0100] To achieve low-latency response for remote control interaction, the control terminal synchronously monitors the display timestamp (PTS) of the decoded frames and the network feedback time during the rendering process. If the cumulative latency exceeds a preset threshold (e.g., 150ms), the system automatically triggers a frame dropping strategy, retaining only keyframes and the most recent reference frame for rendering.

[0101] At the same time, the control terminal will report the sequence number and rendering status of the current display frame to the status monitoring module of the controlled terminal, so as to synchronize the subsequent operation response screen and form a rendering-feedback closed loop.

[0102] Preferably, generating user interaction commands at the control end and returning them to the controlled end via a port mapping channel includes: The control unit loads an event capture layer in the rendering window to listen for user-generated behavioral events and form a set of raw interactive events. The multi-source event structure in the original interaction event set is parsed and uniformly transformed into a standardized operation format, which is then encapsulated into a standard interaction instruction package. The control terminal performs proportional mapping on the coordinate fields of the standard interactive command packet based on the resolution information in the current screen data sequence and the screen ratio parameters of the controlled terminal, and the mapping calculation forms a coordinate mapping data table; The control terminal returns standard interactive command packets and coordinate mapping data tables to the controlled terminal through the port mapping channel, forming a transmission command queue.

[0103] In this embodiment of the invention, the control terminal loads an event capture layer at the top of the rendering window. This layer listens to the user's multi-source input behavior through the operating system's input interface (such as Raw Input on the Windows platform, the evdev interface on the Linux platform, or the input middleware API of HarmonyOS).

[0104] The types of events monitored include mouse clicks, drags, swipes, scroll wheels, keyboard input, and touchpad operations.

[0105] When a user performs any operation in the control window, the event capture layer records parameters such as event type, occurrence timestamp, relative coordinates, key code, pressure value (for touch input), and modifier key state, and writes them into the original interactive event set in the form of an event object.

[0106] The event capture layer uses a multi-threaded asynchronous listening mechanism to ensure that input events do not block screen refresh in a high frame rate rendering environment.

[0107] Meanwhile, by using event dejittering algorithms and duplicate event filtering strategies, redundant operations within a very short time are eliminated, reducing transmission bandwidth usage.

[0108] The original set of interactive events is parsed into a standardized operation format by the interactive instruction parsing module.

[0109] The parsing module first identifies the event source type and then maps events from different sources (such as mouse clicks and touch clicks) to standard operation categories, such as "click", "swipe", or "input", based on a unified operation abstraction model.

[0110] Subsequently, the parsing module extracts key fields (such as coordinates, press status, character input content, etc.) based on the operation category to form a standard interactive instruction package.

[0111] The standard interactive command package adopts a unified command data structure definition, which mainly includes: Command header: containing event sequence number, timestamp and check code; Command body: recording operation type, event parameter set and coordinate information; Command tail: recording status flag bits and reserved extended fields.

[0112] This structure ensures cross-platform compatibility between different control systems and can be consistently resolved on mobile devices, emulators, or tablets.

[0113] To ensure that the user's operation coordinates in the control window correspond precisely to the physical coordinates of the controlled screen, this embodiment sets up a coordinate mapping module in the control terminal.

[0114] After each frame is rendered, the module automatically reads the resolution information (e.g., 1920×1080) of the current frame data sequence and the actual screen parameters of the controlled terminal (e.g., 2340×1080).

[0115] The module calculates the scaling factor matrix based on the difference in the aspect ratio between the two components, and performs the mapping calculation according to the following formula: ; in, The coordinates of the control terminal operation point, The coordinates of the controlled end are... , To determine the size of the rendering window for the control panel, , The screen size of the controlled device.

[0116] The mapping results form a coordinate mapping data table, which contains the correspondence between multiple consecutive operations and is used for a fast index for subsequent instruction return.

[0117] If a change in screen ratio is detected (such as switching between portrait and landscape modes), the coordinate mapping module will automatically recalculate the scaling factor and update the data table to ensure the accuracy of the interaction.

[0118] The generated standard interactive instruction package and the corresponding coordinate mapping data table are transmitted back to the controlled end via the instruction transmission module.

[0119] The transmission module reuses the port mapping channel established above and uses a reliable transmission mechanism (such as a custom data stream protocol based on TCP tunnel encapsulation) to complete bidirectional data synchronization.

[0120] Each return data packet contains a command body, a coordinate data table, and a complete check field, and a transmission sequence number is generated before sending to support out-of-order reordering.

[0121] When the network is unstable, the transmission module will automatically adjust the sending rate and enable the retransmission queue mechanism to avoid input delays or command loss.

[0122] After receiving the transmission instruction queue, the controlled terminal parses out the interactive operation data block and triggers the corresponding system interface, thereby realizing remote clicks, swipes, inputs and other actual operations.

[0123] Most importantly, the multi-source event structure in the original interaction event set is parsed and uniformly transformed into a standardized operation format, including: Identify the source type of behavioral events, including mouse events, touchpad events, touchscreen events, physical keyboard events, and virtual keyboard events; Extract the corresponding event parameters based on different event source types. Event parameters include, but are not limited to, mouse click type, screen coordinates, key values, swipe trajectory, and pressure value. The extracted event parameters are mapped to a predefined standardized event enumeration set and encapsulated into an instruction package with a unified header structure.

[0124] In this embodiment of the invention, when the event capturing layer detects a user operation, the original set of interactive events is sent to the event recognition module.

[0125] The event recognition module identifies the source type of each behavioral event based on the event header field and device identification information of the input subsystem.

[0126] In a typical implementation, the source types include: Mouse events: These originate from physical mouse or Bluetooth wireless mouse input and include actions such as clicking, moving, and scrolling. Touchpad Event: Originates from the laptop's touchpad or external touchpad, and supports multi-finger swipe and zoom operations; Touchscreen event: Originates from input on an external touchscreen display or tablet screen, and has coordinate and pressure parameters; Physical Keyboard Event: Key input from a USB or Bluetooth keyboard; Virtual Keyboard Event: This event originates from input from the virtual keyboard controls displayed in the control panel's UI.

[0127] The event recognition module quickly categorizes each event by reading the Device Handle and Event Type Code fields of the operating system input interface (such as Windows Raw Input, Linux evdev interface, or HarmonyOS input service) and establishes an index relationship for subsequent parameter extraction.

[0128] After source identification, the raw event data is passed to the parameter extraction module. This module uses differentiated parsing templates for different source types to extract key parameter information from the event structure.

[0129] The extracted event parameters include, but are not limited to: Mouse event parameters: Click type (left click, right click, middle click), action state (pressed / released), screen coordinates (X, Y), scroll wheel offset; Touchpad event parameters: number of touch points, swipe trajectory coordinate sequence, swipe speed, and two-finger zoom ratio; Touch screen event parameters: touch point coordinates (X, Y), pressure value, contact area, action indicator (Down / Move / Up); Keyboard event parameters: key value (KeyCode), scan code (ScanCode), key combination flag (Shift, Ctrl, Alt), key state (Press / Release); Virtual keyboard event parameters: input character encoding (Unicode), input region index, and language layout identifier.

[0130] The parameter extraction module employs an event filtering and deduplication mechanism during the parsing process. By comparing timestamp differences with action flags, it eliminates identical events that are repeatedly generated within extremely short intervals, ensuring that the generated event sequence is clean and free of jitter.

[0131] Meanwhile, the module establishes a trajectory buffer queue for sliding and multi-touch events, and optimizes the sliding path through time series interpolation and smoothing algorithms (such as Bezier curve interpolation or Kalman filtering) to improve the accuracy and consistency of operation reproduction.

[0132] The extracted event parameters are sent to the instruction encapsulation module, which performs a unified mapping based on a predefined standardized event enumeration set.

[0133] The standardized event enumeration set is a set of operation behavior codes defined internally by the system to eliminate the differences in events between different hardware input devices and unify them into a common control format.

[0134] The instruction encapsulation module fills the event parameters into the corresponding instruction structure according to the mapping result and generates a standard interactive instruction package with a unified header format.

[0135] The structure of a standard interactive instruction package includes: Command Header: Contains event sequence number, source type, timestamp, and checksum; Instruction Body: Records standardized event codes and their parameter sets; Tail: Contains operation flags, frame synchronization identifiers, and reserved extended fields.

[0136] The generated command packets are uniformly encapsulated using a custom lightweight binary format to reduce network transmission load. Each command packet is written to the command buffer queue after encapsulation, awaiting execution by the controlled terminal via a port mapping channel in subsequent steps.

[0137] To ensure the correct sequence of multi-source events after fusion, the system maintains a global timestamp index table during the encapsulation phase. Each instruction packet is assigned a monotonically increasing sequence number upon generation and written to the time series synchronization buffer.

[0138] Before packaging, the sending module on the control end sorts the data according to the serial number to ensure that various input events (such as clicks and keyboard inputs) can be reproduced on the controlled end in strict accordance with the user's operation sequence.

[0139] When the input event rate is detected to be higher than the network transmission capacity, the system will trigger a dynamic frame dropping strategy, prioritizing the retention of events with direct interactive significance (such as clicks and confirmation keys) and discarding redundant movement events to reduce latency and maintain smooth remote control operation.

[0140] Preferably, the controlled terminal parsing user interaction commands and executing corresponding operations includes the following steps: The controlled end reads data packets from the transmission instruction queue, parses the results, and generates interactive operation data blocks. The controlled terminal triggers the system interface based on the instruction category of the interactive operation data block to generate an operation response state set. The instruction categories include touch, keyboard, and swipe. The control terminal receives the operation response status set, marks the operation result, and generates feedback screen data.

[0141] In this embodiment of the invention, after the control terminal returns an interactive instruction through the port mapping channel, the instruction parsing module of the controlled terminal continuously listens to the transmission instruction queue on the receiving port.

[0142] Whenever a new data packet is received, the module first verifies the data integrity based on the sequence number and check field in the packet header. If a sequence loss or check failure is detected, a retransmission request is automatically triggered.

[0143] Once the data packet passes verification, it enters the instruction buffer. The parsing module parses the packet body according to the predefined data structure and extracts fields such as instruction type, event parameters, coordinate information, and status flags.

[0144] The parsed results are structured into Interaction Operation Blocks, where each block corresponds to a user action, such as clicking, long pressing, swiping, or keyboard input.

[0145] Interactive operation data blocks are described using a unified format: operation identifier; instruction category field; coordinate parameter field; action parameter; execution priority and timestamp information.

[0146] This structure ensures that instructions can be uniformly parsed and scheduled under different input scenarios.

[0147] After receiving the interactive operation data block, the operation scheduling module of the controlled end will distribute it to the corresponding system interface call path according to the "instruction category field".

[0148] The system predefines three types of operation interfaces: Touch class: When the instruction type is touch-related, the operation scheduling module calls the system function hook that is pre-attached by the injection tool. This function is located in the input event distribution layer of the HarmonyOS NEXT system and is used to receive externally injected touch events.

[0149] The module encapsulates coordinate parameters and action flags into a system-level input event structure (such as InputEventStruct) and injects it into the system event queue through a virtual input driver interface to achieve an operation behavior equivalent to a real user click.

[0150] For example, when a click event occurs on the control end, the controlled end will trigger the TouchDown→TouchUp event sequence at the corresponding coordinate position, thereby enabling the selection of interface elements or the clicking of buttons.

[0151] Keyboard class: When the instruction type is keyboard type, the module calls the system input subsystem interface (Input ManagerService), encapsulates the key code (KeyCode) and status flag (Press / Release) returned by the control terminal into a keyboard input event, and writes it to the input buffer queue through the system message mechanism.

[0152] This mechanism can simulate text input, shortcut key triggering, and combination key response, and is suitable for scenarios such as remote text editing and debugging input.

[0153] Slide class: When the instruction type is sliding, the module generates a sliding trajectory dataset according to the path sequence in the coordinate mapping table. The trajectory consists of several discrete coordinate points.

[0154] The system interface trigger module injects sliding events frame by frame at preset time intervals, so that the controlled end presents a smooth and continuous sliding effect on the screen.

[0155] This mechanism supports complex gesture operations such as single-finger swipe, two-finger zoom, and long-distance scrolling.

[0156] After triggering various interfaces, the system will generate an OperationResponse Set based on the returned results. This set records the execution status code (such as success, rejection, waiting, exception), execution time, and target component feedback identifier for each operation.

[0157] After receiving the operation response status set, the feedback generation module on the controlled end binds it to the current screen status.

[0158] The module calls the screen capture interface or image rendering cache to extract the target area (i.e. the area affected by the operation) in the latest frame and generates feedback frame data.

[0159] For example, when a user clicks the application icon on the controlled device from the control device, the feedback screen will display the first frame of the application after it is launched; when a swipe operation is performed, the feedback screen will display the scrolling effect of the controlled device's screen.

[0160] The feedback screen data and the operation response status set are encapsulated together into a feedback data packet, which is returned to the control terminal in real time through the port mapping channel.

[0161] After receiving the data, the control unit performs rendering updates and marks the operation results according to the state set to achieve closed-loop synchronization of input and output.

[0162] Mobile Scenario: The controlled device is attached to the input distribution layer of the HarmonyOS NEXT system via an injection tool. Mouse click events from the control device are converted to screen coordinates via proportional mapping, triggering a system-level touch interface on the controlled device. After parsing and executing the instruction, the controlled device immediately generates a new interface frame and sends it back to the control device, achieving real-time screen projection and remote control synchronization.

[0163] Simulator-side scenario: The controlled end of the simulator receives swipe and keyboard input commands via a virtual display driver. After the command is executed, the simulator engine renders a new frame, the system captures the latest screenshot, and sends it back via a port mapping channel. The control end completes the rendering update at the receiving end and automatically discards expired feedback frames when network latency is high to ensure visual continuity and operational response speed.

[0164] Please see Figure 2The process begins at the "Start / End" node and proceeds to stage S1, where a port mapping channel is constructed and the device is determined to be either a mobile device or an emulator. If it is a mobile device, the process proceeds to stage S2, where a preset screen mirroring tool is injected into the controlled end of the HarmonyOS NEXT system to extract screen content. If it is an emulator, the process proceeds to stage S3, where running frames are captured according to a preset frame rate and combined into a continuous image sequence to form screen data. After stage S4, the screen data is encoded and transmitted to the control end via the port mapping channel. The control end receives, decodes, and renders the data. Subsequently, the control end generates user interaction commands and returns them to the controlled end via the port mapping channel. The controlled end parses the commands, and each stage in the process contains corresponding loop or jump logic.

[0165] Please see Figure 3 The present invention also provides a system attribute simulation system 100 based on HarmonyOS NEXT screen projection remote control, used to execute the above-mentioned system attribute simulation method based on HarmonyOS NEXT screen projection remote control, wherein the system attribute simulation system 100 based on HarmonyOS NEXT screen projection remote control includes: The port mapping and device identification module 101 is used to construct a port mapping channel and determine the current device type; The function injection screen acquisition module 102 is used to inject a preset screen projection tool into the controlled end of the HarmonyOS NEXT system if the device type is a mobile phone, and extract the screen content of the controlled end through the preset screen projection tool to form screen data. The simulator frame capture module 103 is used to capture simulator running frames according to a preset frame rate and combine them into a continuous image sequence to form screen data if the device type is a simulator. The video encoding / decoding and rendering transmission module 104 is used to perform video encoding on the image data and transmit it to the control terminal in real time through the port mapping channel; the control terminal receives and decodes the data, and renders and displays the image to obtain a visual picture. The remote interactive command parsing and control module 105 is used to generate user interactive commands at the control end and return them to the controlled end through the port mapping channel; the controlled end parses the user interactive commands and executes the corresponding operations.

[0166] Therefore, the embodiments should be considered as exemplary and non-limiting in all respects, and the scope of the invention is not limited by the foregoing description. Thus, all changes falling within the meaning and scope of the equivalents of the application are intended to be included within the scope of the invention.

[0167] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features of the invention herein.

Claims

1. A remote control method for screen projection based on HarmonyOS NEXT, characterized in that, Includes the following steps: Step S1: Construct a port mapping channel and determine the current device type; Step S2: If the device type is a mobile phone, a preset screen mirroring tool is injected into the controlled terminal of the HarmonyOS NEXT system. The screen content of the controlled terminal is extracted through the preset screen mirroring tool to form screen data. Step S3: If the device type is an emulator, then capture the emulator running frames according to the preset frame rate and combine them into a continuous image sequence to form screen data; Step S4: Perform video encoding on the image data and transmit it to the control terminal in real time through the port mapping channel; the control terminal receives and decodes the data, renders and displays the image to obtain a visual image; Step S5: Generate user interaction instructions on the control end and return them to the controlled end through the port mapping channel; the controlled end parses the user interaction instructions and executes the corresponding operations.

2. The remote control method for screen projection based on HarmonyOS NEXT according to claim 1, characterized in that, Step S1 includes: Establish a port mapping channel and complete a two-way handshake; generate a device attribute dataset based on the system attribute data collected during the handshake process. The control unit parses the runtime environment type field in the device attribute dataset and determines the device category through the system kernel identifier and virtualization flag. If a real hardware serial number and system underlying driver identifier are detected, it is marked as mobile device type data; If a virtual display driver and emulation environment flag is detected, it is marked as emulator type data.

3. The remote control method for screen projection based on HarmonyOS NEXT according to claim 1, characterized in that, Step S2 includes: If the device type is a mobile phone, the preset screen mirroring tool will be injected into the HarmonyOS NEXT system layer of the controlled device; The pre-set screen mirroring tool obtains the function call context, key parameters, and return values ​​in real time through Hook, forming function call data; Determine the original screen frame data based on function call data; Extract the screen content from the original screen frame data to form screen data.

4. The remote control method for screen projection based on HarmonyOS NEXT according to claim 1, characterized in that, Step S3 includes: If the device type is an emulator, the emulator's running frames are periodically captured according to the preset frame rate to generate a set of static image frames; The static image frame set is combined according to the time series to form a continuous image sequence as picture data; During the assembly process, the system load and network status are monitored in real time, and the output frame rate of the continuous image sequence is dynamically adjusted based on the detection results.

5. The remote control method for screen projection based on HarmonyOS NEXT according to claim 4, characterized in that, Dynamically adjusting the output frame rate of a continuous image sequence based on detection results includes: If either the system CPU utilization or the network transmission latency is greater than or equal to a preset threshold, the system is deemed to be under resource pressure, and a frame reduction process is initiated. During the frame downsizing process, image frames that are not reference frames in the coding structure are discarded. When the system CPU utilization and network transmission latency are both less than the preset threshold and remain below the threshold for a period of time, the system resources are determined to be sufficient, and the output frame rate of the continuous image sequence is increased according to the preset frame rate increment until the preset capture frame rate is reached.

6. The remote control method for screen projection based on HarmonyOS NEXT according to claim 1, characterized in that, The video data is encoded and transmitted in real time to the control terminal via a port mapping channel, including: The video data is input to the video encoding unit, which adaptively selects encoding parameters based on the dynamic evaluation results of the network conditions. Based on the encoding parameters, video encoding is performed on the image data to generate an encoded video stream; The encoded video stream is transmitted to the control terminal in real time through a port mapping channel.

7. The remote control method for screen projection based on HarmonyOS NEXT according to claim 1, characterized in that, The control terminal receives and decodes the data, renders and displays the resulting visual image, including: After receiving the encoded video stream, the control terminal decodes it using the corresponding matching video decoding unit, outputs the original frame data, and renders and displays the original frame data as a visual image through a preset graphics rendering interface.

8. The remote control method for screen projection based on HarmonyOS NEXT according to claim 1, characterized in that, Generating user interaction commands on the control end and returning them to the controlled end via a port mapping channel includes: The control unit loads an event capture layer in the rendering window to listen for user-generated behavioral events and form a set of raw interactive events. The multi-source event structure in the original interaction event set is parsed and uniformly transformed into a standardized operation format, which is then encapsulated into a standard interaction instruction package. The control terminal performs proportional mapping on the coordinate fields of the standard interactive command packet based on the resolution information in the current screen data sequence and the screen ratio parameters of the controlled terminal, and the mapping calculation forms a coordinate mapping data table; The control terminal returns standard interactive command packets and coordinate mapping data tables to the controlled terminal through the port mapping channel, forming a transmission command queue.

9. The remote control method for screen projection based on HarmonyOS NEXT according to claim 1, characterized in that, The controlled terminal parses user interaction commands and executes corresponding operations, including the following steps: The controlled end reads data packets from the transmission instruction queue, parses the results, and generates interactive operation data blocks. The controlled terminal triggers the system interface based on the instruction category of the interactive operation data block to generate an operation response state set. The instruction categories include touch, keyboard, and swipe. The control terminal receives the operation response status set, marks the operation result, and generates feedback screen data.

10. A screen projection remote control system based on HarmonyOS NEXT, characterized in that, The system for executing the HarmonyOS NEXT-based screen mirroring remote control method as described in claim 1, wherein the HarmonyOS NEXT-based screen mirroring remote control system comprises: The port mapping and device identification module is used to build a port mapping channel and determine the current device type; The function is injected into the screen capture module. If the device type is a mobile phone, a preset screen projection tool is injected into the controlled end of the HarmonyOS NEXT system. The screen content of the controlled end is extracted through the preset screen projection tool to form screen data. The simulator frame capture module is used to capture simulator running frames according to a preset frame rate and combine them into a continuous image sequence to form screen data if the device type is a simulator. The video encoding / decoding and rendering transmission module is used to encode the image data and transmit it to the control terminal in real time through the port mapping channel; the control terminal receives and decodes the data, and renders and displays the image to obtain a visual image. The remote interactive command parsing and control module is used to generate user interaction commands at the control end and return them to the controlled end through a port mapping channel; the controlled end parses the user interaction commands and executes the corresponding operations.

Citation Information

Patent Citations

  • Screen projection method, device, storage medium and computer equipment

    CN110515572A

  • Screen projection processing method, device and apparatus and computer readable storage medium

    CN111629239A

  • Screen projection method, system and device based on Android equipment

    CN118803341A

  • Ultra-low time delay high-definition screen projection method based on star flash technology and open source gap system

    CN119603495A

Cited By

  • Direct3D rendering model compatible method based on dynamic template pool

    CN121614179A