Mobile phone screen projection method and device based on intelligent hardware
By constructing a multi-source projection data acquisition and encoding system, and combining it with the smart hardware "Yunbao", we have achieved unified acquisition and standardized encoding of multiple types of data. This solves the problems of single projection data acquisition, low encoding format compatibility, and insufficient transmission optimization in existing technologies, and realizes the stability and synchronization of multi-device collaborative display.
Patent Information
- Application Number
- CN202511873467.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-12
- Publication Date
- 2026-01-23
AI Technical Summary
Existing technologies cannot achieve unified collection and standardized encoding of multiple types of screen projection data. The encoding format has low compatibility and lacks transmission optimization mechanisms for different screen projection scenarios, resulting in poor stability and scenario adaptability, and making it impossible to achieve compatible output of screen projection data across different systems and devices.
A multi-source projection data acquisition and encoding system is constructed. Smart hardware (such as 'Yunbao') is used to collect screen recordings from mobile phones, front and rear camera images, and dual-channel audio. RTSP and RTMP encoding is performed, and transmission stability adjustment factors are added. The protocol is parsed and device identifiers are embedded to generate data to be processed with device identifiers. Receive buffering and priority determination are performed to generate a projection data priority queue. Decoding and multimodal processing are performed to achieve multi-system compatible output.
It achieves stability and compatibility for mobile phone screen projection in multiple scenarios, ensures the real-time performance and reliability of projection data, solves the synchronization problem of multi-device collaborative display, and improves the accuracy and efficiency of screen projection.
Smart Images

Figure CN121397294A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, in particular to a mobile phone screen projection method and device based on intelligent hardware. BACKGROUND
[0002] The prior art cannot realize unified collection and standardized coding of multiple types of projection data, and has two key defects: single data collection dimension: most schemes only support mobile phone screen recording or single camera picture collection, and cannot synchronously obtain multi-source data such as "screen picture + front and rear cameras + double audio" (such as "anchor portrait + desktop courseware + explanation audio" cooperative projection in a live broadcast scene), and additional equipment is required for auxiliary collection, increasing hardware cost and operation complexity; low coding format compatibility: a fixed coding format (only supporting H.264) is adopted, which is not optimized for projection scenes, and real-time scenes (outdoor live broadcast) require low-delay coding, but the coding delay of the existing scheme is ≥100 ms, and multi-device scenes (simultaneous projection to a TV and a monitor) require multi-format output, but the existing scheme can only generate a single coding stream, resulting in the inability of some devices to decode (monitor devices only support MJPEG but receive H.264 streams).
[0003] The prior art lacks transmission optimization mechanisms for different projection scenes, resulting in low stability and scene adaptability: lack of connection adaptation and transmission optimization: no dynamic adaptation process is established between the mobile phone and the intelligent hardware, and the IP address and protocol parameters need to be manually configured, which has a high operation threshold; and no adjustment factor is added according to the protocol characteristics, the TSP protocol (real-time scene) does not control the delay, resulting in picture lag, and the RTMP protocol (multi-channel scene) does not handle packet loss, resulting in data loss (packet loss rate ≥8%); interactive and identification management are chaotic: the transmission data stream is not embedded with device identification, and data confusion may occur when multiple mobile phones are simultaneously projected (such as when multiple mobile phones are simultaneously projected, it is difficult to distinguish the data of each device); and the interactive mechanism is imperfect - the RTSP protocol does not bind the pull stream channel and the mobile phone identification, resulting in unauthorized devices being able to pull data, and the RTMP protocol does not synchronize the buffer capacity information, resulting in data overflow.
[0004] The prior art cannot realize compatible output of projection data in different systems and devices, and the core problems include: lack of format conversion and parameter adaptation: no format conversion is performed according to the signal compatibility requirements of different systems, and multiple system output scheduling conflicts: when multiple devices are projected simultaneously (such as projecting to a TV, a computer, and a monitor), no independent output link is allocated, resulting in link occupation conflicts (such as when an Ethernet link is occupied by a computer, a monitor device cannot receive data); and no timestamp calibration mechanism, the picture received by different devices deviates by ≥100 ms (such as the TV picture and the computer picture are not synchronized), affecting the cooperative display of multiple devices. SUMMARY
[0005] To solve the above technical problems, the present application provides the following technical solutions: A mobile phone screen projection method based on intelligent hardware, comprising: constructing a multi-source screen projection data acquisition and coding system; processing the coded data for connection adaptation and transmission optimization to generate transmission data stream adapted to multiple scenes, and adding a transmission stability adjustment factor; processing the transmission data stream for protocol analysis and device interaction to generate data to be processed with device identification; receiving and buffering the data to be processed and determining the priority to generate a screen projection data priority queue; decoding and multi-modal processing the priority queue data to generate screen projection result data; multi-system output and interactive coordination processing the screen projection result data to generate output data stream compatible across devices; grouping the screen projection requirements according to the screen projection scene restriction type, APP native screen projection constraint rules and device system differences, using the screen recording to screen projection technical path to process the APP restriction avoidance requirements, and combining the multi-system receiving adaptation standard, monitoring device signal compatibility requirements and intelligent hardware decoding capability information to construct a compatibility optimization model that breaks through the scene restrictions; based on the target adaptation parameter output of the compatibility optimization model, combining the dynamic characteristics of the mobile phone shooting scene and the real-time requirement association information of the live scene to generate a picture stability processing scheme and a scene expansion strategy, and generating target screen projection information.
[0006] A mobile phone screen projection device based on intelligent hardware, comprising: a construction module for constructing a multi-source screen projection data acquisition and coding system; a processing module for processing the coded data for connection adaptation and transmission optimization to generate transmission data stream adapted to multiple scenes, and adding a transmission stability adjustment factor; processing the transmission data stream for protocol analysis and device interaction to generate data to be processed with device identification; receiving and buffering the data to be processed and determining the priority to generate a screen projection data priority queue; decoding and multi-modal processing the priority queue data to generate screen projection result data; multi-system output and interactive coordination processing the screen projection result data to generate output data stream compatible across devices; grouping the screen projection requirements according to the screen projection scene restriction type, APP native screen projection constraint rules and device system differences, using the screen recording to screen projection technical path to process the APP restriction avoidance requirements, and combining the multi-system receiving adaptation standard, monitoring device signal compatibility requirements and intelligent hardware decoding capability information to construct a compatibility optimization model that breaks through the scene restrictions; based on the target adaptation parameter output of the compatibility optimization model, combining the dynamic characteristics of the mobile phone shooting scene and the real-time requirement association information of the live scene to generate a picture stability processing scheme and a scene expansion strategy, and generating target screen projection information.
[0007] Its beneficial effects are as follows: This invention uses smart hardware (such as "Yunbao") as the core to solve the problem of mobile phone screen projection in multiple scenarios. First, a multi-source data acquisition and encoding system is constructed to collect screen recordings, front and rear camera images, and dual-channel audio, which are then encoded using RTSP and RTMP. After encoding, the data undergoes connection adaptation (scanning for network configuration, parameter interaction) and transmission optimization, generating a data stream according to the protocol and adjustment factors; the protocol is parsed and interacted with, embedding device identifiers to obtain the data to be processed. The data is buffered and prioritized according to the scenario and cache depth, forming a queue for decoding, followed by audio optimization, video enhancement, and multi-source fusion to generate the projection result. The result undergoes format conversion and link scheduling to achieve multi-system compatible output. For group projection needs, screen recording to projection conversion overcomes APP limitations, integrates multiple standards to build a compatibility optimization model, combines scenario characteristics to generate stability solutions and expansion strategies, and finally outputs target projection information including compliance rate and accuracy. Attached Figure Description
[0008] Figure 1 A flowchart illustrating a mobile phone screen mirroring method based on smart hardware, provided as an embodiment of the present invention; Figure 2 This is a schematic diagram of a mobile phone screen projection device based on smart hardware, provided as an embodiment of the present invention. Detailed Implementation
[0009] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention. Figure 1 This application describes a mobile phone screen mirroring method based on smart hardware according to exemplary embodiments thereof.
[0010] In this application embodiment, a mobile phone screen projection method based on smart hardware is described, such as... Figure 1 As shown: S101, construct a multi-source projection data acquisition and encoding system.
[0011] In one implementation, multi-source raw data acquisition uses a mobile phone as the core acquisition terminal, supporting independent and parallel execution of two types of screen data (screen recording and camera capture) and synchronized audio data, meeting users' flexible needs in multiple scenarios. For screen recording (which can be executed independently), the entire screen of the mobile phone's display interface is captured in real time through the phone's underlying screen recording interface (not dependent on system screen recording controls, avoiding interception by the app). This includes the app's video playback interface, desktop operation interface, document presentation interface, etc. When users need to cast videos from the iQiyi app or PPT presentation documents from their phones, they can start the screen recording function independently. Without needing to associate it with camera operation, the corresponding screen and synchronized system audio can be captured in real time, ensuring that the cast screen and mobile phone operation are consistent in real time. Even if the app restricts native screen casting, the complete playback screen can still be obtained through screen recording. During the acquisition process, a background running mode is supported. After starting screen recording, the user can minimize the program to handle other tasks, while the screen recording function continues to run and capture screens. The program can be reopened to view the screen recording later.
[0012] For capturing footage from the front and rear cameras of a mobile phone (which can be executed independently), the system utilizes either the front or rear camera, selecting the shooting angle and resolution according to the scenario (e.g., the front camera for portrait photography in video conferences, and the rear camera for capturing the environment). It captures dynamic footage in real time and can run independently without needing to activate screen recording. In scenarios such as live streaming and on-site recording, users can activate the camera to capture footage independently and transmit it to the cloud recorder in real time. It also supports continuous background recording—even after starting recording and minimizing the program, the camera maintains a stable capture state, eliminating the need to constantly monitor the interface. During the capture process, it automatically adapts to the phone's optical image stabilization function to reduce handheld shooting shake and simultaneously captures ambient audio from the phone's microphone, meeting the need for independent recording of both visuals and sound.
[0013] The system supports dual-function parallel operation, allowing screen recording and camera capture to start and work simultaneously, enabling flexible scenarios of "background camera feed + foreground screen recording." Users can start the camera to capture footage and run it continuously in the background, while simultaneously opening a video app to watch videos or performing other tasks in the foreground. The screen recording function captures the foreground interface and system audio, with the two types of data collected independently and without interference. For example, if a user needs to fix their phone to film an outdoor scene (with the camera continuously capturing footage in the background), while simultaneously watching an online course and recording key segments, the system can simultaneously capture both video feeds and corresponding audio signals (ambient audio associated with the camera and audio associated with the course recording), meeting the needs of collaborative screen projection from multiple sources or separate storage.
[0014] Synchronous audio acquisition and encoding adaptation are supported, whether executing a single function independently or running two functions in parallel. Audio acquisition covers internal audio from the phone system (such as video sound effects during screen recording and audio recordings for document explanations) and ambient audio from the microphone (such as ambient sound during camera shooting). Preliminary processing includes noise reduction and volume normalization to ensure clear and uninterrupted audio. Encoding operations process the two types of image data separately, converting them to a high-compression, low-latency standard format using RDSD encoding, and then adapting to the RTSP or RTMP protocol according to the scenario—generating a single stream of data when running a single function, and generating two independent streams of data when running two functions in parallel, supporting individual or simultaneous reception by the cloud display, enabling flexible screen projection.
[0015] S102 performs connection adaptation and transmission optimization processing on the encoded data to generate a transmission data stream adapted to multiple scenarios and adds a transmission stability adjustment factor.
[0016] In one implementation, the mobile device sends a network configuration and binding request command by scanning the QR code of the smart hardware. The request command includes the mobile device identifier, the protocol type of the data to be transmitted, and the encoding format information. The mobile device triggers the network configuration and binding process by scanning the unique QR code generated by the smart hardware (such as "Yunbao"), sending the request command to the smart hardware. This request command contains three core pieces of information: first, the mobile device identifier, using the mobile phone's unique IMEI code (e.g., "861234567890123456"), ensuring that the smart hardware accurately identifies the mobile terminal initiating the request; second, the protocol type of the data to be transmitted, preset by the user according to the screen projection scenario (e.g., selecting RTSP protocol for single-person real-time live streaming, and selecting RTMP protocol for multi-camera collaborative screen projection); and third, the encoding format information, which by default carries an RDSD encoding identifier, specifying the encoding standard for subsequent data transmission. Users need to project the outdoor live stream captured by their phone's rear camera onto Yunbao. After scanning the Yunbao QR code, the phone will automatically send a request command containing the IMEI code "861234567890123456", the protocol type "RTSP", and the encoding format "RDSD", requesting to establish a network binding relationship with Yunbao.
[0017] After receiving the request command, the smart hardware returns the device IP address, protocol adaptation parameters, single / multi-channel projection support capability, and hardware decoding specifications to the mobile phone via a response command. The smart hardware (Yunbao) receives the request command from the mobile phone and returns a response command via the local area network. The response includes four key parameters: first, the smart hardware IP address (e.g., a fixed IP address within the local area network, "192.168.123.189"), providing the target address for subsequent data transmission; second, protocol adaptation parameters, such as the UDP transmission port "554" for the RTSP protocol and the TCP transmission port "1935" for the RTMP protocol; third, single / multi-channel projection support capability, specifying the number of projection channels currently supported by the hardware resources (e.g., "RTSP supports 1 channel, RTMP supports up to 4 channels"); and fourth, hardware decoding specifications, including supported resolutions (e.g., "up to 4K"), frame rates (e.g., "30 frames / second"), and encoding format compatibility (e.g., "only RDSD encoding is supported"). After receiving the request from the aforementioned outdoor live streaming scenario, Yunbao returns a response command containing the IP address "192.168.123.189", the RTSP protocol port "554", single-channel projection support capability, 4K resolution and RDSD decoding support information, informing the mobile device of the current hardware compatibility conditions.
[0018] The mobile device receives and parses the response information returned by the smart hardware, obtaining the smart hardware's IP address, the corresponding protocol's transmission port, the number of projection channels limit, and decoding capability parameters, and automatically matches the encoded data corresponding to the RTSP or RTMP protocol. After receiving the response command returned by the smart hardware, the mobile device extracts key parameters through its built-in parsing module and completes two types of core operations: First, it obtains basic transmission parameters, confirming the smart hardware's IP address, the corresponding protocol's transmission port (such as RTSP protocol port "554"), the number of projection channels limit (such as "currently only supports 1 RTSP projection channel"), and decoding capability parameters (such as "supports 4K / RDSD encoded data decoding"); second, it automatically matches encoded data, retrieving the corresponding data from the local encoding cache based on the parsed protocol type and decoding specifications (e.g., if the parsed protocol type is RTSP, it automatically matches real-time video data from the mobile phone's camera that has been encoded using RDSD), ensuring that the data format is compatible with the smart hardware's decoding capabilities. Example: After the mobile phone parses the Yunbao's response information, it confirms the IP "192.168.123.189", RTSP port "554", and 4K decoding support. It then automatically retrieves the 1080P / 30fps live video data captured by the mobile phone's rear camera and encoded with RDSD from the local cache, matching the Yunbao's protocol and decoding requirements.
[0019] The mobile app performs connection adaptation processing on the encoded data, generating link data compatible with the smart hardware's reception and displaying it in a list. The core of the link data is building a standardized transmission link: for the RTSP protocol, a complete pull address (e.g., "rtsp: / / 192.168.123.111 / av0") is generated, containing the mobile phone's IP address (e.g., "192.168.123.111") and the stream identifier (e.g., "av0"). For the RTMP protocol, a push address based on the smart hardware's IP address is generated (e.g., "rtsp: / / 192.168.123.189 / mobile / SN1", where "SN1" is a unique identifier after the mobile phone is bound to the smart device). The list displays the link status ("Pending Connection"), protocol type, and target address, allowing users to easily confirm the adaptation results. For the aforementioned outdoor live streaming scenario, the mobile phone generates an RTSP protocol pull address "rtsp: / / 192.168.123.111 / av0", which is displayed in the APP as a list "Link 1: RTSP protocol - rtsp: / / 192.168.123.111 / av0 - to be connected", so that the user can confirm and initiate the transmission.
[0020] The mobile app optimizes the transmission of the adapted link data based on the screen mirroring scenario requirements. A real-time adjustment factor is added for the RTSP protocol, and a reliability adjustment factor is added for the RTMP protocol, generating transmission data streams adapted to multiple scenarios. Dedicated adjustment factors are added for different protocols to generate transmission data streams adapted to various scenarios. For RTSP protocol real-time adjustment factors: For scenarios with high real-time requirements (such as outdoor live streaming and mobile camera shooting), a "UDP transmission latency control" factor is added to stabilize data transmission latency within 100ms; simultaneously, a "packet loss tolerance threshold" (e.g., "allow ≤3% packet loss rate") is enabled to avoid slight network fluctuations causing video stuttering, prioritizing real-time performance. In outdoor live streaming scenarios, the mobile phone adds a real-time adjustment factor to the RTSP link data, controlling UDP transmission latency to ≤100ms and allowing a 3% packet loss rate, ensuring that the live stream received by the cloud camera is synchronized with the mobile phone's shooting in real time, with no significant delay.
[0021] RTMP protocol reliability adjustment factors: For scenarios with high requirements for multi-channel projection and data integrity (such as multi-camera conference projection and multi-mobile desktop collaboration), a "TCP packet loss retransmission" factor is added, with a retransmission timeout of "500ms" to ensure that lost data segments can be retransmitted. Simultaneously, a "flow control threshold" (such as "maximum bitrate per channel 2Mbps") is added to prevent single-channel data from consuming too much bandwidth, causing stuttering in multi-channel projection and prioritizing transmission reliability. A company uses three mobile phones for multi-camera conference projection (phone 1 captures the presenter, phone 2 captures the PPT, and phone 3 captures the audience), all using the RTMP protocol. The mobile devices add a reliability adjustment factor to each data stream, configuring a 500ms packet loss retransmission timeout and a 2Mbps single-channel bitrate limit to ensure stable presentation of all three streams on the cloud platform without data loss or image gaps.
[0022] S103 performs protocol parsing and device interaction processing on the transmitted data stream to generate data to be processed with device identifier.
[0023] In one implementation, the transmitted data stream undergoes protocol header parsing. The protocol parsing module of the smart hardware extracts the protocol identifier field from the data stream to distinguish between RTSP and RTMP transmission protocols. For RTSP protocol data streams, the stream address field is additionally parsed to extract the mobile app server IP and port information. For RTMP protocol data streams, the push address field is parsed to extract the mobile device SN code and the smart hardware receiving channel identifier, generating protocol type identifier data. The smart hardware, through its built-in protocol parsing module, performs protocol header parsing on the received transmitted data stream. The core function is to distinguish the protocol type and extract key parameters to generate protocol type identifier data. Specific operations and examples are as follows: The parsing module scans the protocol identifier field in the data stream frame header (such as the "RTSP / 1.0" field for RTSP protocol and the "handshake packet" identifier for RTMP protocol) to determine the corresponding transmission protocol type of the data stream. Example: When a cloud-based device receives a data stream, parsing the frame header reveals the "RTSP / 1.0" identifier, determining that the data stream is an RTSP protocol; upon receiving another data stream, the "handshake packet" feature in the frame header indicates that it is an RTMP protocol.
[0024] For the RTSP protocol data stream, the stream address field is further parsed to extract the mobile app server's IP and port information. The stream address field is typically formatted as "rtsp: / / [Mobile IP]:[Port] / [Stream Identifier]". After parsing, the mobile IP (i.e., the app server's IP) and port number are separated. Parsing the RTSP protocol stream address "rtsp: / / 192.168.123.111:554 / av0" extracts the mobile app server's IP as "192.168.123.111", port as "554", and stream identifier as "av0", which are used for subsequent stream pull requests.
[0025] For RTMP protocol data streams, the push address field is parsed to extract the mobile device's serial number (SN) and the smart hardware receiving channel identifier. The push address field is typically formatted as "rtmp: / / [Yunbao IP] / [Channel Identifier] / [Mobile SN]". After parsing, the mobile SN (uniquely identifying the mobile device) and the smart hardware receiving channel identifier are separated. For example, parsing the RTMP protocol push address "rtmp: / / 192.168.123.189 / mobile / SN123456" extracts the mobile device's SN as "SN123456" and the smart hardware receiving channel identifier as "mobile", thus clarifying the mobile terminal and Yunbao receiving channel corresponding to this data stream. After parsing, the protocol type, extracted IP, port, SN, and channel identifier information are integrated to generate protocol type identifier data, providing parameter support for subsequent device interactions.
[0026] Device interaction is processed based on protocol type identifiers and parsing parameters. For the RTSP protocol, the smart hardware sends a pull request command to the parsed mobile phone IP via UDP, carrying the smart hardware ID and pull port information. After receiving the request, the mobile app returns an authorization response containing the unique identifier of the mobile device and its encoding format. The smart hardware binds the mobile phone identifier in the response to the local pull channel. For the RTMP protocol, the smart hardware sends a data reception ready command to the mobile phone via TCP, carrying the smart hardware's receive buffer capacity and multi-channel projection channel allocation information. After receiving the command, the mobile phone returns confirmation information containing the mobile device model and data stream frame rate. The smart hardware associates the mobile phone identifier with the corresponding push channel, generating device association data. The smart hardware initiates device interaction for RTSP and RTMP protocols respectively based on the protocol type identifier data, completing command sending and response reception, and generating device association data. Specific operations and examples are as follows. For RTSP protocol device interactions, the smart hardware initiates a pull request. Based on the parsed mobile app server IP and port, it sends a pull request command to the mobile phone via UDP protocol. The command carries the smart hardware ID (e.g., "Yunbao43000100") and the pull port (e.g., "554"), informing the mobile phone that it needs to transmit RTSP data streams to this port. Example: Yunbao sends a UDP pull request to IP "192.168.123.111" and port "554", carrying the smart hardware ID "Yunbao43000100" and the pull port "554". After receiving the pull request, the mobile app returns an authorization response containing the mobile device's unique identifier (e.g., IMEI code "861234567890123456") and encoding format ("RDSD"). After receiving the response, the smart hardware binds the mobile phone's IMEI code to its local RTSP pull channel (e.g., "Channel 1") to ensure that it only receives RTSP data streams from this mobile phone in the future. Example: Yunbao receives the IMEI code "861234567890123456" and "RDSD" encoding format returned by the mobile phone, binds the IMEI code to the local "Channel 1", and only receives the RTSP stream of the mobile phone through this channel.
[0027] For RTMP protocol device interaction, the smart hardware sends a data reception ready command. Based on the parsed mobile device SN code and receive channel identifier, it sends the data reception ready command to the mobile phone via TCP protocol. The command carries the smart hardware's receive buffer capacity (e.g., "4GB") and multi-channel projection channel allocation information (e.g., "Assign RTMP channel 2 to this mobile phone"), informing the mobile phone that it can push streams to the specified channel. Example: For the mobile phone SN code "SN123456", Yunbao sends a ready command via TCP, carrying the buffer capacity "4GB" and the channel allocation information "RTMP channel 2", allowing the mobile phone to push streams to channel 2. After receiving the ready command, the mobile phone returns confirmation information including the mobile device model (e.g., "Mate50") and data stream frame rate (e.g., "30 frames / second"). After receiving the confirmation, the smart hardware associates the mobile phone SN code with the corresponding RTMP streaming channel (e.g., "channel 2") to ensure accurate identification of the mobile phone to which the data stream received by this channel belongs. Yunbao receives the "Mate50" model and "30 frames / second" frame rate returned by the mobile phone, associates the SN code "SN123456" with "RTMP channel 2", and confirms that the data stream of channel 2 comes from the mobile phone. After completing the interaction, it integrates the association relationship between the smart hardware and the mobile phone (channel-device identifier binding) and generates device association data to provide a basis for subsequent identifier embedding.
[0028] The device-associated data and the parsed data stream undergo identification embedding processing. For a single RTSP data stream, a three-dimensional identifier of Smart Hardware ID - Mobile Phone IMEI Code - Pull Channel Number is embedded. For multiple RTMP data streams, in addition to embedding a three-dimensional identifier of Smart Hardware ID - Mobile Phone SN Code - Pull Channel Number, a data stream priority identifier is also embedded. Simultaneously, the identifier format is uniformly set to a 16-byte binary field and embedded in the data stream frame header, generating unprocessed data with device identification. For a single RTSP data stream, a three-dimensional identifier of "Smart Hardware ID - Mobile Phone IMEI Code - Pull Channel Number" is embedded. The Smart Hardware ID uniquely identifies the receiving device, the Mobile Phone IMEI Code uniquely identifies the sending mobile phone, and the Pull Channel Number clearly identifies the receiving channel of the data stream within the smart hardware. The combination of these three elements ensures accurate traceability of a single data stream. Furthermore, the identifier format is uniformly set to a 16-byte binary field and embedded in the data stream frame header. Yunbao embeds a 3D identifier into the RTSP protocol data stream. The smart hardware ID "Yunbao43000100" (4 bytes), the mobile phone IMEI code "861234567890123456" (8 bytes), and the streaming channel number "1" (4 bytes) are integrated into a 16-byte binary field and embedded in the frame header of each data stream. This identifier can be used to quickly locate the device and channel to which the data stream belongs.
[0029] For RTMP multi-stream data, in addition to embedding a three-dimensional identifier (functionally identical to the RTSP three-dimensional identifier, only replacing the IMEI code with the SN code) of "Smart Hardware ID - Mobile SN Code - Streaming Channel Number", a data stream priority identifier is also embedded. The priority identifier is set according to scenario requirements (e.g., "1" for the main live stream and "2" for auxiliary streams) for subsequent priority scheduling. Similarly, the unified identifier format is a 16-byte binary field embedded in the frame header. Yunbao embeds an identifier for a specific RTMP protocol data stream (the main live stream): Smart Hardware ID "Yunbao43000100" (4 bytes), Mobile SN Code "SN123456" (4 bytes), Streaming Channel Number "2" (4 bytes), and Priority Identifier "1" (4 bytes), integrated into a 16-byte binary field embedded in the frame header. This enables data stream traceability and provides a priority basis for subsequent queue scheduling. Through identifier embedding, each data stream carries a unique identifier containing device, channel, and priority (RTMP protocol) information, generating pending data with device identifiers, providing a basis for subsequent receive buffering and priority determination.
[0030] S104 performs receive buffering and priority determination on the data to be processed, and generates a priority queue for screen projection data.
[0031] In one implementation, the data to be processed is buffered via a FIFO (First-In-First-Out) receiver to generate a priority identifier. Specifically, the Yunbao device uses a multi-channel independent FIFO to buffer the data to be processed, including the data with device identifiers. The FIFO's data_count port is used to monitor the data write depth of each channel in real time, simultaneously extracting the device identifier and priority identifier from the data stream frame header to generate temporary buffered data. The Yunbao device uses a multi-channel independent FIFO (First-In-First-Out) buffering mechanism to buffer the data to be processed and simultaneously extract key identifiers to generate temporary buffered data. For multi-channel independent FIFO buffer allocation, an independent FIFO channel is allocated to each data stream based on its device identifier (e.g., "Smart Hardware ID - Mobile Phone IMEI Code - Pull Channel Number" for RTSP single-channel data, and "Smart Hardware ID - Mobile Phone SN Code - Push Channel Number" for RTMP multi-channel data), avoiding data buffer conflicts between different devices. Yunbao receives two streams of data to be processed. One stream uses the RTSP protocol (device identifier "Yunbao43000100-861234567890123456-1") and is allocated the FIFO1 channel. The other stream uses the RTMP protocol (device identifier "Yunbao43000100-SN123456-2") and is allocated the FIFO2 channel, thus enabling independent buffering of the two streams of data.
[0032] Using the data_count port of each FIFO channel, the data write depth (i.e., the amount of data currently stored in the FIFO) is collected in real time to provide a data volume basis for subsequent priority determination. The data_count port of FIFO1 channel detected that the current cached data volume is 320 frames, and the cached data volume of FIFO2 channel is 410 frames. The data depth information is fed back to the priority determination module in real time. While caching data, the 16-byte binary identifier field of the data stream frame header is parsed simultaneously to extract the device identifier (to confirm the mobile phone and channel to which the data belongs) and the priority identifier (priority information inherent in RTMP data, such as "1" representing the main live screen), and associated with the FIFO channel and data depth to generate temporary cached data. The data frame header identifier of FIFO2 channel is parsed to extract the device identifier "Yunbao43000100-SN123456-2" and the priority identifier "1". Combined with the data depth of 410 frames monitored by data_count, temporary cached data "FIFO2-Device Identifier XXX-Priority 1-Data Depth 410" is generated.
[0033] The depth of the FIFO cache data is comprehensively assessed based on the scenario requirements. Specifically, for live streaming scenarios with screen switching needs, if a cache depth of one of the RTMP multi-channel data streams reaches 400 frames and is identified as the main live stream screen, it is marked as priority one; if the cache depth of one RTSP single-channel data stream reaches 350 frames and is used for real-time shooting with a mobile camera, it is marked as priority two; if the RTMP data cache depth reaches 200 frames in non-real-time recording scenarios, it is marked as priority three; if any channel's video data cache reaches 8 lines and triggers a picture-in-picture function request, it is forcibly upgraded to priority one. Combining the FIFO cache data depth and screen projection scenario requirements, each data stream is prioritized to clarify the order of data processing. Priority one determination (highest priority): For live streaming scenarios with screen switching needs, if a cache depth of one of the RTMP multi-channel data streams reaches 400 frames and is identified as the "main live stream screen," it is marked as priority one; furthermore, if any channel's video data cache reaches 8 lines (corresponding to approximately 0.27 seconds / 30 frames) and triggers a picture-in-picture function request, it is forcibly upgraded to priority one. Example 1: The RTMP data (live main screen) cached in FIFO2 channel has a depth of 410 frames, which meets the condition of "depth ≥ 400 frames + live main screen identifier", and is marked as priority one; Example 2: The RTSP data (moving camera shooting) cached in FIFO1 channel has a depth of only 150 frames, but triggers the picture-in-picture function request and is forcibly upgraded to priority one.
[0034] Priority 2 (Medium Priority): For real-time shooting scenarios with mobile cameras (such as outdoor live streaming, on-site inspection), if the RTSP single-channel data buffer depth reaches 350 frames, it is marked as Priority 2. Example: The RTSP data buffered in FIFO1 channel (outdoor scene shooting with the phone's rear camera) reaches a depth of 360 frames, meeting the conditions of "RTSP protocol + depth ≥ 350 frames + mobile shooting scenario", and is marked as Priority 2 (when picture-in-picture request is not triggered). Priority 3 (Low Priority): For non-real-time recording scenarios (such as conference screen recording, course recording), if the RTMP data buffer depth reaches 200 frames, it is marked as Priority 3. Example: The RTMP data buffered in FIFO3 channel (conference content recorded by a mobile phone) reaches a depth of 220 frames, meeting the conditions of "RTMP protocol + depth ≥ 200 frames + non-real-time recording scenario", and is marked as Priority 3.
[0035] Data with priority identifiers is queued and scheduled to generate a priority queue for projection data. Based on the data's priority identifier, all priority-based data is queued and scheduled to generate a priority queue for projection data, ensuring that high-priority data is processed first. The basic sorting logic follows "Priority 1 > Priority 2 > Priority 3," with data of the same priority sorted in descending order of cache depth (the greater the cache depth, the higher the priority, avoiding data overflow). If a picture-in-picture request is triggered to upgrade priority, that data path is directly inserted at the beginning of the current queue. Example: Yunbao currently caches 3 data paths: "Priority 1 (FIFO2, 410 frames)," "Priority 2 (FIFO1, 360 frames)," and "Priority 3 (FIFO3, 220 frames)," with a sorted queue order of "FIFO2 → FIFO1 → FIFO3." If FIFO1 data triggers a picture-in-picture request to upgrade to priority 1, the queue is updated in real-time to "FIFO1 → FIFO2 → FIFO3."
[0036] Yunbao's scheduling module reads data sequentially from the corresponding FIFO channels according to priority queues, while simultaneously monitoring the FIFO data depth: if a high-priority data channel completes processing (FIFO data depth drops to 0), the next data channel is immediately scheduled; if the high-priority data buffer depth continues to increase (e.g., FIFO2 data depth increases from 410 frames to 450 frames), its processing time is extended to ensure complete data transmission. Example: Scheduling in the "FIFO2→FIFO1→FIFO3" queue, data is first read and processed from FIFO2. When the FIFO2 data depth drops to 50 frames, FIFO1 data reading is simultaneously initiated (double buffering) to avoid data processing interruption; FIFO3 data is scheduled after FIFO1 processing is complete, ensuring that high-priority data occupies hardware resources first. Through this process, a priority queue of projection data, sorted by priority and adapted to the scenario requirements, is ultimately generated, providing ordered data input for subsequent decoding and multimodal processing.
[0037] S105 decodes and performs multimodal processing on the priority queue data to generate the projection result data.
[0038] In one implementation, the decoding process takes priority queue data as input. Based on the device identifier and protocol type in the data frame header, the corresponding decoding module is invoked to restore the encoded data to the original audio and video data, ensuring that the data format can be used for subsequent multimodal processing. The Yunbao device automatically matches the corresponding decoding module based on the protocol type identifier (RTSP / RTMP) and encoding format identifier (RDSD) in the data frame header of the priority queue. For RDSD encoded data of the RTSP protocol, the dedicated RTSP-RDSD decoding module is invoked; for RDSD encoded data of the RTMP protocol, the dedicated RTMP-RDSD decoding module is invoked, while simultaneously initializing decoding parameters (such as resolution, frame rate, and audio sampling rate). If a data frame in the priority queue is identified as "RTSP-RDSD-1080P-30 frames", the Yunbao device invokes the RTSP-RDSD decoding module, initializing the resolution to 1920×1080, the frame rate to 30 frames / second, and the audio sampling rate to 44.1kHz, ensuring that the decoding parameters are consistent with those used during encoding.
[0039] The decoding module extracts encoded frames from the priority queue data, parses the compressed data in RDSD encoding format (such as intra-frame compression and inter-frame compression data), and restores it to the original video frames (YUV format) and audio frames (PCM format) using an inverse compression algorithm. Simultaneously, it verifies data integrity by checking the three-dimensional identifier in the frame header (such as "Smart Hardware ID-Mobile Phone IMEI Code-Streaming Channel Number" in RTSP data). If the identifier is missing or incorrect, a retransmission request is triggered. It parses the RDSD encoded data of the RTMP protocol, extracts the compressed video frame data, and restores it to the original 1080P / YUV420 format video frames using an RDSD inverse compression algorithm. The audio data is restored to 16bit / 44.1kHz PCM format. Simultaneously, it verifies the "Yunbao43000100-SN123456-2" identifier in the frame header to confirm that the data comes from the specified mobile phone and channel, and that there are no integrity issues.
[0040] To address potential synchronization discrepancies in audio and video data (such as audio-visual asynchrony caused by transmission delays), the decoding module employs a built-in timestamp correction mechanism to extract the timestamps of video and audio frames and adjust the data output timing. If the video frame timestamp lags behind the audio frame by more than 50ms, empty video frames are inserted to fill the gap; if the audio frame lags, audio segments are extended using an audio interpolation algorithm to ensure that the audio-visual synchronization discrepancy is ≤30ms. Example: When decoding the main screen data of a live stream, if the video frame timestamp lags behind the audio frame by 80ms, the decoding module automatically inserts two empty video frames (corresponding to approximately 67ms) and fine-tunes the output rhythm of subsequent video frames to correct the audio-visual synchronization discrepancy to within 20ms, meeting the viewing experience requirements of live streaming scenarios.
[0041] Multimodal processing optimizes the decoded raw audio and video data, enhances the video, and fuses multi-source data to generate projection results that can be directly used for projection output, taking into account the needs of the projection scenario. For the decoded PCM format audio data, noise reduction, volume adjustment, and channel adaptation are performed. Adaptive noise reduction algorithms (such as spectral subtraction) remove ambient noise (such as fan noise and background conversations); based on the projection scenario requirements (such as improving voice clarity in a meeting scenario), audio equalizer parameters are adjusted (enhancing voice clarity in the 1kHz-3kHz frequency band); and mono audio is automatically converted to stereo for multi-speaker output devices. Example: Processing on-site presentation audio captured by a mobile phone microphone, adaptive noise reduction algorithms remove ambient noise in the exhibition hall (improving the signal-to-noise ratio by 15dB after noise reduction), enhance the 1kHz-3kHz frequency band to improve voice clarity, and convert mono audio to stereo to ensure uniform sound distribution when projected onto a dual-speaker device in a meeting room.
[0042] For decoded YUV format video data, image quality enhancement, jitter correction, and format adaptation are performed. In non-real-time scenarios (such as recording and playback), an AI image quality enhancement model (based on Yunbao's 6T computing power, supporting super-resolution reconstruction) is used to upscale 720P video to 1080P and repair blurry details. In real-time scenarios (such as live streaming), a lightweight jitter correction algorithm is used to compensate for slight image jitter based on the phone's gyroscope data (embedded in the data frame header). At the same time, the video size is automatically scaled according to the output device resolution (such as 1080P for TVs and 720P for monitoring equipment) to adapt to the display needs of different devices. Example: For outdoor live streaming video shot by a mobile phone (decoded to 720P with slight jitter), Yunbao uses a lightweight jitter correction algorithm to compensate for image offset based on the gyroscope data in the frame header, while scaling the video to 1080P to adapt to the needs of casting to a 1080P TV. After processing, the image jitter is reduced by 60%, and the clarity meets live streaming standards.
[0043] If multiple data streams exist in the priority queue (such as the main screen + auxiliary screen in RTMP multi-channel projection), multi-source data fusion processing is performed. For picture-in-picture scenarios, the high-priority main screen (such as the live stream main screen) is used as the background layer, and the low-priority auxiliary screen (such as the PPT screen) is used as the foreground layer. The size (such as 1 / 4 of the screen size) and position (such as the lower right corner) of the foreground layer are adjusted to ensure that the two screens are unobstructed and displayed synchronously. For multi-screen split-screen scenarios, multiple data streams are allocated to screen areas according to priority (such as priority one occupying the left half of the screen, and priorities two and three each occupying the upper and lower areas of the right half of the screen). At the same time, multiple audio streams are synchronized, and the multiple audio streams are merged into a single output through an audio mixing algorithm to avoid sound overlap and confusion. Example: In the priority queue, there are two data streams (Priority 1: the host's portrait image, and Priority 2: the PPT presentation image). Yunbao uses the host's portrait image as the background layer to fill the entire screen, and the PPT presentation image as the foreground layer, scaled to 1 / 4 size and displayed in the lower right corner, creating a picture-in-picture effect. In terms of audio, the host's audio and the PPT background music are merged using a mixing algorithm, and the volume ratio of the two is adjusted to 3:1 to ensure that the host's voice is clear and the background music does not interfere with the presentation.
[0044] S106 performs multi-system output and interactive coordination processing on the screen projection result data to generate a cross-device compatible output data stream.
[0045] In one implementation, to address the signal compatibility requirements of different receiving systems (Linux, Windows, monitoring equipment, etc.), the projection result data (including video frames, audio frames, and control signals) undergoes format conversion and parameter adaptation to ensure the data can be recognized by the target system. Based on the video encoding format and resolution supported by the receiving system, the YUV format video frames in the projection result data are converted to the target format. For Windows systems, they are converted to H.264 encoded MP4 format with a resolution adapted to the commonly used 1920×1080 resolution of monitors; for Linux systems, the original YUV format is retained to adapt to professional video processing software; for monitoring equipment, they are converted to MJPEG encoded standard definition format (720×576) to adapt to the real-time decoding capabilities of the monitoring equipment. For example, a company needs to output projection result data to a Windows-based conference monitor. Yunbao converts 1080P / YUV420 format video frames to H.264 / MP4 format while maintaining a resolution of 1920×1080, ensuring direct playback on the monitor without format compatibility issues.
[0046] Based on the audio interface type and sampling rate requirements of the receiving system, PCM format audio data is converted to the corresponding format. For display devices with speakers (such as smart TVs), it is converted to dual-channel AAC format with a sampling rate of 44.1kHz; for monitoring devices, it is converted to mono G.711 format to adapt to the audio storage requirements of the monitoring system; for the audio processing module of the Linux system, the original PCM format is retained to support subsequent noise reduction secondary processing. Example: The 16bit / 44.1kHz PCM audio data in the screen projection result is converted to dual-channel AAC format when output to a smart TV, and converted to mono G.711 format when output to a monitoring device, ensuring that different devices can correctly parse the audio.
[0047] The system extracts interactive control signals (such as screen switching commands and picture-in-picture on / off commands) from the screen mirroring results and converts them into control protocols supported by the receiving system. For devices supporting the HDMI-CEC protocol (such as TVs), the control signals are encapsulated into HDMI-CEC commands; for web-based control interfaces, they are converted into HTTP request commands; and for monitoring devices, they are converted into ONVIF protocol commands, enabling cross-system control interaction. Users send a "Turn on Picture-in-Picture" command through the Yunbao web interface. Yunbao converts this control signal into an HTTP request command and transmits it to the screen mirroring receiving software on the Windows system. Upon receiving the command, the software immediately executes the picture-in-picture function, achieving cross-system control.
[0048] Based on the requirements of the projection scenario (such as single-system dedicated output or multi-system parallel output), the output links are scheduled and conflict avoidance is implemented to ensure that there are no link conflicts when multiple systems receive data simultaneously. Yunbao allocates independent output links to different receiving systems through its multi-channel output modules. For single-system output scenarios (such as projection only to a TV), one HDMI link is used; for multi-system parallel output scenarios (such as simultaneous projection to a TV, Windows computer, and monitoring equipment), HDMI, Ethernet, and RS485 links are allocated separately to avoid link resource contention. For example, at a certain exhibition, it was necessary to simultaneously project the mobile live stream to an on-site TV (HDMI link), a remote Windows computer (Ethernet link), and a local monitoring device (RS485 link). Yunbao allocated independent links to the three systems to ensure synchronized output without lag or disconnection.
[0049] For scenarios involving parallel output from multiple systems, a timestamp calibration mechanism synchronizes the data stream output timing of each link. A global timestamp is added to the header of the projection result data frame, and each output link adjusts its data transmission rhythm according to the timestamp to ensure that the deviation between the received image and audio is ≤50ms. The latency for a TV receiving data via the HDMI link is 20ms, and the latency for a Windows computer receiving data via the Ethernet link is 40ms. Yunbao uses timestamp calibration to advance the output rhythm of the Ethernet link by 20ms, ensuring that the images received by the two systems are synchronized with no significant time difference.
[0050] When multiple systems request the same output link (e.g., two Windows computers simultaneously requesting an Ethernet link), a link queuing mechanism is activated. Link usage is allocated based on system priority (e.g., office computers have higher priority than regular computers), with lower-priority systems entering a waiting queue and automatically transmitting once the link becomes available. If a link fails (e.g., an HDMI cable disconnects), it automatically switches to a backup link (e.g., an Ethernet link) and sends a link switching notification to the user. For example, if two Windows computers simultaneously request an Ethernet link, the system determines the office computer has higher priority and allocates the link first, while the regular computer enters a waiting queue. If the HDMI link suddenly disconnects, the system automatically switches the output to the backup Ethernet link to ensure uninterrupted screen mirroring on the TV, and simultaneously pushes a notification to the user via the web interface stating "HDMI link failed, switched to Ethernet link."
[0051] A multi-system interactive feedback mechanism is established to monitor output status in real time and handle anomalies, ensuring the stability and reliability of cross-device output. Yunbao receives real-time status information from the receiving system (such as "data reception normal," "decoding failure," "link disconnected") through feedback modules in each output link, and displays the output status of each system in real time on the web interface (green indicates normal, red indicates anomaly). If Yunbao detects a "decoding failure" report from a Windows computer, the web interface immediately marks that system status in red and displays the message "Decoding failed, please check the encoding format," helping users quickly locate the problem.
[0052] For different anomaly types, corresponding handling strategies are activated. If the receiving system reports "decoding failure," Yunbao automatically reduces the video encoding complexity (e.g., converting high-bitrate H.264 to low-bitrate). If it reports "link disconnection," it automatically retryes the connection (up to 3 times), switching to a backup link if the retry fails. If it reports "data packet loss," it activates a TCP retransmission mechanism for the RTMP protocol and a packet loss compensation algorithm (inserting empty frames) for the RTSP protocol. If the monitoring device reports "data packet loss rate of 10%," Yunbao activates a packet loss compensation algorithm for the RTSP protocol, inserting empty frames to fill the packet loss gaps, while simultaneously reducing the video bitrate to control the packet loss rate below 3%, ensuring continuous monitoring footage. It supports collaborative control of projection results data from multiple systems. When one system initiates a control command (e.g., pause the screen, adjust resolution), Yunbao synchronizes the command to all output systems, ensuring consistent display status across multiple systems. When a user initiates a "pause screen" command on a Windows computer, Yunbao not only pauses the local screen but also synchronizes the command to the TV and monitoring equipment, enabling simultaneous pause of all three system screens and avoiding interactive confusion caused by inconsistent statuses.
[0053] S107 groups screen casting requirements based on the types of screen casting scenario restrictions, native APP screen casting constraints, and differences in device systems. It adopts a screen recording to screen casting technology path to address APP restriction avoidance requirements in a targeted manner. It integrates multi-system receiving adaptation standards, monitoring device signal compatibility requirements, and smart hardware decoding capability information to build a compatibility optimization model that breaks through scenario limitations.
[0054] In one implementation, the limitations and requirements of screen projection scenarios are categorized and analyzed to generate requirement grouping information. Specifically, a scenario limitation identification module analyzes the types of screen projection scenario limitations, native app screen projection constraints, and device system differences. Screen projection requirements are divided into three categories: app limitation avoidance requirements, multi-system compatibility requirements, and functional expansion requirements. The scenario limitation identification module breaks down and analyzes the limiting factors and user needs in the screen projection process, clarifies the limitation types, classifies the requirement categories, and generates requirement grouping information, providing a basis for subsequent targeted processing. The scenario limitation identification module scans the core limiting factors in the screen projection scenario, including three key types of information: first, the type of screen projection scenario limitation, such as network fluctuation limitations for outdoor live streaming and multi-device synchronization limitations for conference screen projection; second, native app screen projection constraints, such as iQiyi prohibiting native screen projection and bank apps prohibiting screenshots and screen recordings; and third, device system differences, such as Windows systems supporting H.264 encoding, Linux systems supporting YUV raw format, and monitoring devices only supporting MJPEG encoding. Analysis of the "screen mirroring of iQiyi videos from mobile phone to TV" scenario reveals that the APP's native screen mirroring constraint rule is "prohibit native screen mirroring", the device system difference is "TV supports H.264 encoding, mobile phone outputs RDSD encoding", and the scenario restriction type is "requires real-time low-latency transmission".
[0055] Based on the types of restrictions identified, screen mirroring requirements are categorized into three types: first, app restriction avoidance requirements, which involve overcoming app restrictions on native screen mirroring and screenshots, such as bypassing iQiyi's native screen mirroring ban to achieve video screen mirroring; second, multi-system compatibility requirements, which involve achieving compatible output of screen mirroring data across different systems (Windows, Linux, monitoring devices), such as adapting the same screen mirroring data to both TVs and monitoring devices; and third, functional expansion requirements, which involve adding additional screen mirroring functions, such as image enhancement, picture-in-picture, and audio noise reduction. In a certain enterprise's screen mirroring requirements, "overcoming Tencent Video's native screen mirroring restrictions" falls under the app restriction avoidance requirement, "simultaneously mirroring to Windows computers and Linux servers" falls under the multi-system compatibility requirement, and "AI image quality enhancement for the mirrored screen" falls under the functional expansion requirement. These three types of requirements are generated into independent requirement grouping information.
[0056] Targeted technologies are employed to address the need for circumventing app restrictions. For apps that prohibit native screen mirroring, a screen recording-to-screen mirroring approach is used. This involves reading the screen data stream from the phone's underlying interface, converting it to an IP stream via RDSD encoding, and then transmitting it to the Cloud Computing device. For scenarios where apps prohibit screenshots, if the underlying data stream cannot be obtained, only the image captured by the phone's camera is used as the screen mirroring signal source, generating a solution to circumvent the restriction. Different technical approaches are used based on different constraint rules to overcome native screen mirroring and screenshot limitations, generating solutions to bypass these restrictions. The "screen recording-to-screen mirroring" approach does not rely on the phone's native screen mirroring controls. It directly reads the screen data stream through the phone's underlying interfaces (such as Android's MediaProjectionAPI and iOS's ReplayKit), avoiding app interception. The read data stream is converted to an IP stream via RDSD encoding and transmitted to the Cloud Computing device via RTSP / RTMP protocols, and then the Cloud Computing device adapts and outputs it to the target device. Users need to cast iQiyi videos that are prohibited from native casting. The mobile phone reads the iQiyi playback screen data stream through the underlying interface, compresses it into a 1080P / 30fps IP stream through RDSD encoding, and pushes it to Yunbao via the RTMP protocol. Yunbao converts it into the H.264 format supported by the TV before casting, thus bypassing the app's native casting restriction.
[0057] If an app intercepts the reading of underlying screen data streams through system permissions (such as some banking apps), the screen recording path is abandoned. Instead, only the phone's camera is used as the projection signal source. The phone's front or rear camera is used to capture the app's display in real time. The captured image is then encoded with RDSD and transmitted to the cloud-based screen mirroring device. Although the image clarity may be affected by the shooting angle, this bypasses the screenshot restriction. For example, a certain banking app prohibits screenshots and underlying screen recording. Users need to project their billing information from within the app. The phone's rear camera is used to capture the billing information on the screen. The captured image is then encoded with RDSD and transmitted to the cloud-based device via the RTSP protocol. The cloud-based device receives the image and projects it onto the computer, displaying the billing information and circumventing the app's screenshot restriction.
[0058] By integrating multi-dimensional adaptation standards and hardware capability information, a compatibility optimization model is constructed. This model integrates multi-system reception adaptation standards, monitoring device signal compatibility requirements, and smart hardware decoding capabilities to establish a mapping relationship between requirements, technology, and hardware. The model incorporates a scenario judgment module that automatically matches technical paths and hardware resources based on the input requirement type. It also integrates a dynamic adjustment mechanism that automatically updates adaptation parameters when device system changes or APP restriction upgrades are detected. The basic model architecture integrates three types of core information: first, multi-system reception adaptation standards, such as Windows supporting MP4 format, Linux supporting YUV format, and monitoring devices supporting MJPEG format; second, monitoring device signal compatibility requirements, such as frame rate ≤ 25 frames / second and resolution ≤ 720P; and third, smart hardware decoding capability information, such as Yunbao supporting RDSD / H.264 decoding, a maximum of 4 projection channels, and 6T computing power. A "demand-technology-hardware" mapping relationship is established based on three types of information. For example, the "APP restriction avoidance requirement" maps to the "screen recording to screen casting + RDSD encoding" technology and the "Yunbao RDSD decoding module" hardware. The "multi-system compatibility requirement" maps to the "format conversion + multi-link output" technology and the "Yunbao multi-channel output module" hardware. For the "multi-system compatibility requirement (screen casting to TV and monitoring equipment)," the mapping technology path is "RDSD encoding to H.264 (TV adaptation) + RDSD encoding to MJPEG (monitoring adaptation) + dual-link output," and the hardware resources are "Yunbao H.264 decoding module + MJPEG encoding module + HDMI / RS485 link."
[0059] The model incorporates two core modules: a scenario judgment module and a dynamic adjustment mechanism. The first module automatically matches the technical path and hardware resources based on the input requirement type, requiring no manual intervention. For example, inputting "APP restriction avoidance requirement (iQiyi screen casting)" will automatically match "screen recording to screen casting + RTMP streaming" technology with the "Yunbao RTMP receiving module." The second module monitors device system changes (such as TV system upgrades) and APP restriction upgrades (such as iQiyi adding screen recording detection) in real time, automatically updating adaptation parameters. If it detects that iQiyi APP upgrades and intercepts the underlying screen recording interface, the dynamic adjustment mechanism automatically switches the technical path from "underlying screen recording" to "camera shooting," while simultaneously updating Yunbao encoding parameters to ensure normal transmission of the cast screen.
[0060] S108, based on the target adaptation parameter output of the compatibility optimization model, combines the dynamic characteristics of the mobile phone shooting scene with the real-time requirements of the live streaming scene to generate a screen stabilization processing scheme and scene expansion strategy, and generates target projection information.
[0061] In one implementation, the target adaptation parameters output by the compatibility optimization model and the dynamic characteristics of the mobile phone shooting scene are quantitatively analyzed to generate a basic quantitative factor for image stability. By quantitatively decomposing the target adaptation parameters output by the compatibility optimization model and the dynamic characteristics of the mobile phone shooting scene, key influencing factors are extracted and assigned values to generate a basic quantitative factor reflecting image stability, providing data support for subsequent scheme evaluation. The target adaptation parameters output by the compatibility optimization model include encoding bitrate, transmission protocol, resolution, etc. A quantitative standard is set for each type of parameter: encoding bitrate ≥4Mbps is scored as 1 point, 2-4Mbps as 0.6 points, and <2Mbps as 0.3 points; RTSP protocol (low latency) is scored as 0.9 points, RTMP protocol (high reliability) as 0.6 points; resolution ≥1080P is scored as 0.8 points, 720P as 0.5 points, and <720P as 0.2 points. Example: The model outputs target adaptation parameters for the "outdoor mobile live streaming" scenario as "RTSP protocol, bitrate 5Mbps, 1080P resolution", which are quantized to 0.9, 1 and 0.8 points respectively. The total parameter quantization score is (0.9+1+0.8) / 3=0.9 points.
[0062] Analyze dynamic interference factors during mobile phone shooting, including image shake amplitude, network fluctuation frequency, and light change degree. Shake amplitude ≤ 5° is scored as 0.8 points, 5-15° as 0.5 points, and > 15° as 0.2 points; network fluctuation frequency ≤ 5 times / minute is scored as 0.9 points, 5-15 times / minute as 0.6 points, and > 15 times / minute as 0.3 points; light change amplitude ≤ 200 lux is scored as 0.7 points, 200-500 lux as 0.4 points, and > 500 lux as 0.1 points. Example: In an outdoor live streaming scene, the mobile phone shooting image shake amplitude is 3°, network fluctuation is 3 times / minute, and light change is 150 lux. After quantization, these are scored as 0.8, 0.9, and 0.7 points respectively. The total feature quantization score = (0.8 + 0.9 + 0.7) / 3 = 0.8 points. The total score of parameter quantization and the total score of feature quantization are combined with a weight of 4:6 to generate a basic quantization factor for image stability. The calculation formula is "basic quantization factor = total score of parameter quantization × 0.4 + total score of feature quantization × 0.6". In the above outdoor live streaming scenario, the basic quantization factor = 0.9 × 0.4 + 0.8 × 0.6 = 0.84 points. The closer the factor is to 1, the better the basic conditions for image stability.
[0063] Spatiotemporal matching and priority analysis of the real-time requirements of live streaming scenarios with the hardware resource parameters of Yunbao are performed to generate scenario adaptation quantification factors. The real-time requirements of live streaming scenarios include transmission latency requirements and frame rate requirements. The hardware resource parameters of Yunbao include computing power (6T) and decoding frame rate limit (30 frames / second). Matching standards are set: transmission latency requirement ≤100ms and Yunbao supports UDP transmission (latency ≤80ms) is scored as 1 point, requirement 100-200ms and Yunbao supports TCP transmission (latency ≤150ms) is scored as 0.7 points, and requirement >200ms is scored as 0.4 points; frame rate requirement ≤30 frames / second (Yunbao decoding limit) is scored as 0.9 points, and >30 frames / second is scored as 0.3 points. Example: The real-time requirements for the "multi-camera live streaming" scenario are "transmission latency ≤ 150ms, frame rate 25 frames / second". Yunbao supports TCP transmission (latency ≤ 150ms) and decoding frame rate 30 frames / second. After matching, it scores 0.7 points and 0.9 points respectively. The total matching score is (0.7 + 0.9) / 2 = 0.8 points.
[0064] Prioritize requirements based on their importance within the scenario. Core requirements (such as transmitting the main live stream screen) are awarded 1 point, important requirements (such as projecting auxiliary screens) are awarded 0.7 points, and general requirements (such as non-real-time recording) are awarded 0.4 points. In multi-camera live streaming, "real-time transmission of the main screen" is a core requirement, earning 1 point; "projecting the audience's screen" is an important requirement, earning 0.7 points. Here, we take the core requirement as an example, awarding it 1 point for priority.
[0065] The spatiotemporal matching score and the demand priority score are combined with a weighted average of 7:3 to generate a scene adaptation quantification factor. The calculation formula is "Scene adaptation quantification factor = total matching score × 0.7 + priority score × 0.3". In the above multi-camera live streaming scenario, the scene adaptation quantification factor = 0.8 × 0.7 + 1 × 0.3 = 0.86 points. The closer the factor is to 1, the higher the adaptation between scene requirements and hardware resources.
[0066] Based on the basic quantization factor of image stability and the quantization factor of scene adaptation, combined with the hierarchical architecture of the compatibility optimization model (requirement layer - technology layer - hardware layer), the image processing solutions (such as jitter correction and image quality enhancement) and scene expansion directions (such as multi-system output and new function additions) are evaluated and integrated to generate solution-strategy correlation quantization features. For image processing solutions (such as lightweight jitter correction and AI image quality enhancement), effect weights are set in combination with the basic quantization factor. When the basic quantization factor is ≥0.8, the lightweight jitter correction solution is weighted at 0.9 points (adapting to slight jitter), and the AI image quality enhancement solution is weighted at 0.7 points (supported by Yunbao 6T computing power); when the basic quantization factor is 0.5-0.8, the jitter correction solution is weighted at 0.6 points, and the image quality enhancement solution is weighted at 0.4 points. For outdoor live streaming scenarios, the basic quantization factor is 0.84 (≥0.8), the lightweight jitter correction solution is weighted at 0.9 points, the AI image quality enhancement solution is weighted at 0.7 points, and the solution quantization score is (0.9+0.7) / 2=0.8 points.
[0067] For scenario expansion directions (such as simultaneous screen mirroring to TVs and monitoring equipment, adding picture-in-picture functionality), feasibility weights are set based on the scenario adaptation quantification factor. When the scenario adaptation quantification factor is ≥0.8, the feasibility weight for multi-system output is 0.8 points (supported by Yunbao multi-channel), and the weight for picture-in-picture functionality is 0.9 points (matching hardware decoding capabilities). When the factor is between 0.5 and 0.8, the weights for all expansion directions are reduced by 0.3 points. For multi-camera live streaming scenarios, the adaptation quantification factor is 0.86 (≥0.8), the weight for multi-system output is 0.8 points, the weight for picture-in-picture functionality is 0.9 points, and the expansion quantification score is (0.8 + 0.9) / 2 = 0.85 points. The solution quantification score and the expansion quantification score are merged with a weight of 6:4 to generate a correlated quantification feature. The calculation formula is "correlated quantification feature = solution score × 0.6 + expansion score × 0.4". In outdoor live streaming scenarios, the correlation quantification feature score is 0.8 × 0.6 + 0.85 × 0.4 = 0.82. The higher the feature value, the stronger the synergy between the image processing solution and the scene expansion direction.
[0068] Based on a compatibility optimization model, the quantitative characteristics of the solution-strategy correlation are analyzed and processed to generate target projection information including screen stability compliance rate, scene adaptation accuracy rate, and function execution efficiency. A stability compliance threshold of 0.7 (basic quantitative factor ≥ 0.7) is set, and the percentage of projection scenarios with a basic quantitative factor ≥ 0.7 is statistically analyzed. Based on the model analysis of 100 outdoor live streaming projection scenarios, 85 scenarios had a basic quantitative factor ≥ 0.7, resulting in a screen stability compliance rate of 85 / 100 × 100% = 85%. An adaptation compliance threshold of 0.75 (scene adaptation quantitative factor ≥ 0.75) is set, and the percentage of scenarios with an adaptation quantitative factor ≥ 0.75 and no actual projection anomalies is statistically analyzed. In 80 multi-camera live streaming scenarios, 72 scenarios had an adaptation quantitative factor ≥ 0.75 and normal projection, resulting in a scene adaptation accuracy rate of 72 / 80 × 100% = 90%.
[0069] For image processing solutions and scene expansion functions, the response time of functions (such as jitter correction startup time and picture-in-picture loading time) is statistically analyzed and compared with hardware performance thresholds (such as ≤200ms) to calculate efficiency. The average response time for lightweight jitter correction performed by Yunbao is 150ms, with a hardware threshold of 200ms. Therefore, the function execution efficiency is calculated as (200-150) / 200×100%+100% (normal function operation) = 125% (the portion exceeding the threshold is proportionally increased, up to a maximum of 120%), and 120% is ultimately taken. The picture-in-picture loading time is 180ms, with an efficiency of (200-180) / 200×100%+100% = 110%. Therefore, the function execution efficiency is calculated as (120%+110%) / 2 = 115%.
[0070] like Figure 2As shown, a mobile phone screen projection device based on smart hardware includes: a construction module 201 for constructing a multi-source screen projection data acquisition and encoding system; a processing module 202 for performing connection adaptation and transmission optimization processing on the encoded data, generating a transmission data stream adapted to multiple scenarios, and adding a transmission stability adjustment factor; performing protocol parsing and device interaction processing on the transmission data stream to generate data to be processed with device identifiers; performing receiving buffering and priority determination on the data to be processed to generate a screen projection data priority queue; decoding and multimodal processing on the priority queue data to generate screen projection result data; and performing multi-system output and... Interactive coordination and processing generate cross-device compatible output data streams; combining the screen casting scenario limitation types, APP native screen casting constraint rules, and device system differences, screen casting requirements are grouped, and the technical path of screen recording to screen casting is used to specifically handle APP limitation avoidance requirements. By integrating multi-system receiving adaptation standards, monitoring device signal compatibility requirements, and smart hardware decoding capability information, a compatibility optimization model that breaks through scenario limitations is constructed; based on the target adaptation parameter output of the compatibility optimization model, combined with the dynamic characteristics of mobile phone shooting scenarios and the real-time requirements of live streaming scenarios, a screen stability processing scheme and scenario expansion strategy are generated, and target screen casting information is generated.
[0071] A computing device includes a memory for storing computer program instructions and a processor for executing the computer program instructions, wherein when the computer program instructions are executed by the processor, the device is triggered to execute any of the smart hardware-based mobile phone screen mirroring methods.
[0072] The methods and / or embodiments in this application can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowchart. When the computer program is executed by a processing unit, it performs the functions defined in the methods of this application.
[0073] It should be noted that the computer-readable medium described in this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0074] It will be apparent to those skilled in the art that this application is not limited to the details of the exemplary embodiments described above, and that this application can be implemented in other specific forms without departing from the spirit or essential characteristics of this application. Therefore, the embodiments should be regarded as exemplary and non-limiting in all respects, and the scope of this application is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be embraced within this application.
Claims
1. A method for screen mirroring from a mobile phone based on smart hardware, characterized in that, include: A multi-source projection data acquisition and encoding system is constructed. The raw data includes screen recordings from mobile phones and real-time images captured by the front and rear cameras of mobile phones. The encoding operations cover RTSP protocol adaptation encoding and RTMP protocol adaptation encoding. At the same time, it supports synchronous acquisition and preliminary processing of audio signals. The encoded data is processed for connection adaptation and transmission optimization to generate a transmission data stream adapted to multiple scenarios, and a transmission stability adjustment factor is added. The system performs protocol parsing and device interaction processing on the transmitted data stream to generate data to be processed with device identifiers. The system performs receive buffering and priority determination on the data to be processed, and generates a priority queue for screen projection data. The priority queue data is decoded and processed in a multimodal manner to generate the projection result data. Perform multi-system output and interactive coordination processing on the screen projection results data to generate a cross-device compatible output data stream; By grouping screen casting requirements according to the types of screen casting scenario restrictions, the native screen casting constraints of the APP, and the differences in device systems, the technical path of screen recording to screen casting is adopted to address the APP restriction avoidance requirements. By integrating multi-system receiving adaptation standards, monitoring device signal compatibility requirements, and smart hardware decoding capability information, a compatibility optimization model that breaks through scenario restrictions is constructed. Based on the target adaptation parameter output of the compatibility optimization model, combined with the dynamic characteristics of the mobile phone shooting scene and the real-time requirements of the live streaming scene, a screen stability processing scheme and scene expansion strategy are generated to generate target projection information.
2. The mobile phone screen projection method based on smart hardware according to claim 1, characterized in that, The encoded data undergoes connection adaptation and transmission optimization processing to generate a transmission data stream adaptable to multiple scenarios, and transmission stability adjustment factors are added, including: The mobile device sends a network configuration and binding request command by scanning the QR code of the smart hardware. The request command contains the mobile device identifier, the protocol type and encoding format information of the data to be transmitted. After receiving the request command, the smart hardware returns the device IP address, protocol adaptation parameters, single / multi-channel screen projection support capability, and hardware decoding specification information to the mobile phone through the response command; The mobile device receives and parses the response information returned by the smart hardware, obtains the smart hardware's IP address, the corresponding protocol's transmission port, the number of projection channels limit, and decoding capability parameters, and automatically matches the encoded data corresponding to the RTSP or RTMP protocol. The mobile device performs connection adaptation processing on the encoded data, generates link data that is adapted to be received by smart hardware, and displays it in a list; The mobile app optimizes the transmission of the adapted link data according to the needs of the screen mirroring scenario. It adds a real-time adjustment factor for the RTSP protocol and a reliability adjustment factor for the RTMP protocol to generate a transmission data stream adapted to multiple scenarios.
3. The mobile phone screen projection method based on smart hardware according to claim 1, characterized in that, The transmitted data stream undergoes protocol parsing and device interaction processing to generate data to be processed with device identifiers, including: The protocol header of the transmitted data stream is parsed. The protocol identification field in the data stream is extracted by the protocol parsing module of the smart hardware to distinguish between RTSP and RTMP transmission protocols. For RTSP protocol data streams, the stream address field is parsed to extract the IP and port information of the mobile APP server. For RTMP protocol data streams, the push address field is parsed to extract the mobile device SN code and the smart hardware receiving channel identifier to generate protocol type identification data. Device interaction is processed based on protocol type identifiers and parsing parameters. For the RTSP protocol, the smart hardware sends a pull request command to the parsed mobile phone IP via the UDP protocol, carrying the smart hardware ID and pull port information. After receiving the request, the mobile APP returns an authorization response containing the unique identifier of the mobile device and the encoding format. The smart hardware binds the mobile device identifier in the response to the local pull channel. For the RTMP protocol, the smart hardware sends a data reception ready command to the mobile phone via the TCP protocol, carrying the smart hardware receive buffer capacity and multi-channel projection channel allocation information. After receiving the command, the mobile phone returns confirmation information containing the mobile device model and data stream frame rate. The smart hardware associates the mobile device identifier with the corresponding push channel and generates device association data. The device-associated data and the parsed data stream are subjected to identification embedding processing. For RTSP single-channel data streams, a three-dimensional identifier of smart hardware ID-mobile phone IMEI code-pull channel number is embedded. For RTMP multi-channel data streams, in addition to embedding a three-dimensional identifier of smart hardware ID-mobile phone SN code-pull channel number, a data stream priority identifier is also embedded. At the same time, the unified identifier format is a 16-byte binary field, which is embedded in the data stream frame header position to generate unprocessed data with device identifier.
4. The mobile phone screen projection method based on smart hardware according to claim 1, characterized in that, The system performs receive buffering and priority determination on the data to be processed, and generates a priority queue for screen projection data, including: The data to be processed is buffered by FIFO reception and priority identifier is generated. Specifically, the data to be processed with device identifier is buffered by multi-channel independent FIFO of smart hardware. The data_count port of FIFO is used to monitor the data write depth of each channel in real time, and the device identifier and priority identifier of data stream frame header are extracted synchronously to generate temporary buffer data. The depth of the FIFO cache data is comprehensively assessed based on the scenario requirements. Specifically, for live streaming scenarios with screen switching requirements, if the cache depth of one of the RTMP multi-channel data reaches 400 and is identified as the main live stream screen, it is marked as priority one; if the cache depth of a single RTSP data channel reaches 350 and is used for real-time shooting scenarios with mobile cameras, it is marked as priority two; if the RTMP data cache depth reaches 200 in non-real-time recording scenarios, it is marked as priority three; if any channel of video data caches 8 lines and triggers a picture-in-picture function request, it is forcibly upgraded to priority one. Data with priority identifiers are sorted and scheduled to generate a priority queue for projection data.
5. The mobile phone screen projection method based on smart hardware according to claim 1, characterized in that, By grouping screen casting requirements according to the types of limitations in the casting scenario, the native casting constraints of the app, and differences in device systems, and using a screen recording to casting technology approach, the need to circumvent app limitations is addressed in a targeted manner. This involves integrating multi-system reception adaptation standards, monitoring device signal compatibility requirements, and smart hardware decoding capabilities to construct a compatibility optimization model that overcomes scenario limitations, including: The limitations and requirements of screen casting scenarios are classified and analyzed to generate requirement grouping information. Among them, the scenario limitation identification module analyzes the type of screen casting scenario limitation, the native screen casting constraint rules of the APP, and the differences between the device system. The screen casting requirements are divided into three categories: APP limitation avoidance requirements, multi-system compatibility requirements, and function expansion requirements. Targeted technologies are used to address the need to circumvent app restrictions. For apps that prohibit native screen mirroring, a screen recording to screen mirroring approach is employed, reading the screen data stream from the phone's underlying system and transmitting it to the smart hardware. For scenarios where apps prohibit screenshots, if the underlying data stream cannot be obtained, only the phone's camera footage is used as the screen mirroring signal source, generating a solution to circumvent the restriction. By integrating multi-dimensional adaptation standards and hardware capability information, a compatibility optimization model is constructed. This model integrates multi-system receiving adaptation standards, monitoring equipment signal compatibility requirements, and smart hardware decoding capability information to establish a mapping relationship between requirements, technology, and hardware. The model has a built-in scenario judgment module that automatically matches technical paths and hardware resources based on the type of input requirements. It also integrates a dynamic adjustment mechanism that automatically updates adaptation parameters when device system changes or APP restriction upgrades are detected.
6. The mobile phone screen projection method based on smart hardware according to claim 5, characterized in that, Based on the target adaptation parameter output of the compatibility optimization model, and combined with the dynamic characteristics of the mobile phone shooting scene and the real-time requirements of the live streaming scene, a screen stabilization processing scheme and scene expansion strategy are generated, and target projection information is generated, including: The target adaptation parameters output by the compatibility optimization model and the dynamic features of the mobile phone shooting scene are quantitatively analyzed and processed to generate the basic quantitative factor for image stability. Spatiotemporal matching and demand priority analysis are performed on the real-time demand information of live streaming scenarios and the parameters of smart hardware resources to generate scenario adaptation quantification factors. Based on the basic quantization factor of image stability and the quantization factor of scene adaptation, combined with the hierarchical architecture of the compatibility optimization model, the image processing scheme and the scene expansion direction are fused to generate scheme-strategy related quantization features. Based on the compatibility optimization model, the quantitative characteristics of the scheme-strategy association are analyzed and processed to generate target projection information including screen stability compliance rate, scene adaptation accuracy rate, and function execution efficiency.
7. A mobile phone screen projection device based on smart hardware, characterized in that, The device includes: The building module is used to construct a multi-source projection data acquisition and encoding system; The processing module performs connection adaptation and transmission optimization on the encoded data, generating a transmission data stream adapted to multiple scenarios and adding a transmission stability adjustment factor; it performs protocol parsing and device interaction processing on the transmission data stream, generating data to be processed with device identifiers; it performs receive buffering and priority determination on the data to be processed, generating a priority queue for projection data; it decodes and performs multimodal processing on the priority queue data, generating projection result data; it performs multi-system output and interaction coordination processing on the projection result data, generating a cross-device compatible output data stream; it groups projection requirements based on projection scenario limitation types, APP native projection constraint rules, and device system differences, and uses a screen recording to projection technology path to specifically address APP limitation avoidance requirements, integrating multi-system receive adaptation standards, monitoring device signal compatibility requirements, and smart hardware decoding capability information to construct a compatibility optimization model that breaks through scenario limitations; based on the target adaptation parameters output of the compatibility optimization model, combined with the dynamic characteristics of mobile phone shooting scenarios and the real-time requirements of live streaming scenarios, it generates a screen stability processing scheme and scenario expansion strategy, generating target projection information.
8. An electronic device, characterized in that, include: First processor; and memory for storing executable instructions of the first processor; The first processor is configured to execute the mobile phone screen projection method based on smart hardware according to any one of claims 1 to 6 by executing the executable instructions.
9. A computing device, the device comprising a memory for storing computer program instructions and a processor for executing the computer program instructions, wherein, When the computer program instructions are executed by the processor, the device is triggered to execute the mobile phone screen projection method based on smart hardware as described in any one of claims 1 to 6.