Screen transmission equipment and screen transmission method

By recognizing the user's selected area and operable objects in the screen sharing device, and dynamically adjusting the selected area by combining boundary and semantic features, the problem of frequent selection area adjustments in screen sharing is solved, improving efficiency and security, and adapting to changes in screen content.

CN120872207APending Publication Date: 2025-10-31HISENSE VISUAL TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510795923.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

In existing screen sharing technologies, users cannot flexibly select specific screen areas to share, and when the screen content changes, they need to frequently adjust the selection area manually, resulting in low efficiency.

Method used

A screen sharing device and method are provided. By identifying the user's selected area and operable objects in real-time image data, the selected area of ​​the screen sharing data is dynamically adjusted. By combining boundary and semantic features, the selected area is matched and motion is predicted to achieve adaptive adjustment of the selected area. Sensitive information is blurred when necessary.

Benefits of technology

It improves the efficiency and accuracy of screen sharing, reduces the frequency of users manually adjusting the selection area, enhances information security and real-time performance, and adapts to dynamic changes in screen content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120872207A_ABST
    Figure CN120872207A_ABST
Patent Text Reader

Abstract

The embodiment of the invention belongs to a screen sharing technology, and provides screen transmission equipment and a screen transmission method, and the method comprises the steps: determining screen transmission data in real-time image data in response to a screen transmission instruction, and transmitting the screen transmission data to the screen transmission equipment; the screen transmission instruction comprises any one of the following selected areas: a first selected area and a second selected area. According to the scheme, the selected area on which the screen transmission data is based is determined in the real-time image data, and the selected area can be the first selected area required by a user or the second selected area which is used for identifying the operable object of the first selected area and is determined based on the operable object after view conversion; according to the method, the selected area of the screen transmission data can be dynamically adjusted along with the change of the shared content, the continuity of the shared content is improved, the situation that a user frequently and manually adjusts the selected area is reduced, and therefore the screen sharing efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of screen sharing technology, and in particular to a screen sharing device and screen sharing method. Background Technology

[0002] With the rapid development of internet technology and communication tools, screen sharing has become an indispensable part of modern digital communication. Through screen sharing, users can remotely view and collaboratively process documents, presentation slides, software applications, or other digital content, thereby promoting instant communication and collaboration, and improving communication efficiency and productivity.

[0003] Users typically cannot customize the selection area during screen sharing; they can only share the entire screen or a specific application window, rather than flexibly choosing a specific screen region to share. Even in solutions where the selection area can be customized, users still need to manually readjust the selection area when the screen content changes to ensure the correct content is shared, thus reducing the efficiency of screen sharing. Summary of the Invention

[0004] This application provides a screen sharing device and a screen sharing method to improve the efficiency of screen sharing.

[0005] In a first aspect, embodiments of this application provide a screen sharing device, including: a display screen configured to display real-time image data; and a processing module connected to the display screen, configured to: in response to a screen sharing command, determine screen sharing data from the real-time image data and send the screen sharing data to the device to be projected; the screen sharing command includes any one of the following selection areas: a first selection area and a second selection area; wherein, the first selection area is input by the user, and the second selection area is determined by the processing module based on the operable object identified in the first selection area and the operable object after view transformation, and the operable object includes any one of the following: an image, a cell, or a table.

[0006] The technical solutions involved in the above embodiments have the following advantages or beneficial effects: In the solution of this example, the selection area on which the screen sharing data is based is determined in the real-time image data. It can be the first selection area required by the user, or the second selection area determined by recognizing operable objects in the first selection area and based on operable objects after view transformation. This realizes that the selection area of ​​the screen sharing data can be dynamically adjusted with the changes in the shared content. While improving the continuity of the shared content, it reduces the situation where users frequently manually adjust the selection area, thereby improving the efficiency of screen sharing.

[0007] In some embodiments of this application, the processing module is configured to: in response to a screen transmission command including a first selection area, determine first image data located in the first selection area in real-time image data, and extract a first boundary feature and a first semantic feature from the first image data; extract at least one boundary feature and at least one semantic feature from the real-time image data; determine a second boundary feature and a second semantic feature that have the highest similarity to the first boundary feature and the first semantic feature among the at least one boundary feature and the at least one semantic feature; and use the display area corresponding to the second boundary feature and the display area corresponding to the second semantic feature as the second selection area.

[0008] The technical solutions involved in the above embodiments have the following advantages or beneficial effects: In this example, the solution extracts the first boundary features and the first semantic features in the first selection area, searches for the features with the highest similarity in the entire real-time image data, and determines the area most relevant to the user's initial selection as the second selection area. Based on the features of the user's initial selection area, it can track and locate content areas with high relevance in the entire screen, so that the determination of the second selection area is not limited by the initial position and can adapt to the dynamic changes of the shared content, thereby improving the accuracy of the shared content.

[0009] In some embodiments of this application, the processing module is configured to: determine the coordinates of a first region corresponding to a second boundary feature, and determine the coordinates of a second region corresponding to a second semantic feature; and perform weighted fusion of the first region coordinates and the second region coordinates to obtain the region coordinates corresponding to the second selected area.

[0010] The technical solutions involved in the above embodiments have the following advantages or beneficial effects: In the solution of this example, the coordinates of the second selection area are obtained by weighted fusion of the coordinates of the first area corresponding to the second boundary feature and the coordinates of the second area corresponding to the second semantic feature, so that the determined second selection area can continuously include changing text or borders, and achieve content consistency when continuously sharing content.

[0011] In some embodiments of this application, the processing module is configured to: when an operable object in the first selected area responds to a user operation to perform a view transformation, perform motion prediction on the second boundary feature, and perform a correction operation on the coordinates of the first region corresponding to the second boundary feature based on the motion prediction result, the correction operation including expanding or shrinking the coordinates of the first region.

[0012] The technical solutions involved in the above embodiments have the following advantages or beneficial effects: In the solution of this example, when the content in the first selected area changes, the coordinates of the first area corresponding to the second boundary feature are corrected based on the motion prediction result of the second boundary feature. This enables the update of the second area coordinates to track the rhythm and speed of the content change in the first selected area in real time, reducing the position deviation caused by update lag, thereby improving the real-time performance of shared content.

[0013] In some embodiments of this application, the processing module is configured to: extract sensitive images from real-time image data based on a sensitive information database, or determine sensitive text corresponding to sensitive semantic features in a second semantic feature; and use the display area corresponding to the sensitive image or the display area corresponding to the sensitive text as a mask area to blur the screen transmission data.

[0014] The technical solutions involved in the above embodiments have the following advantages or beneficial effects: The solution in this example extracts sensitive images from real-time image data based on a sensitive information database, or determines the sensitive text corresponding to the sensitive semantic features in the second semantic features. It can identify and locate potential sensitive information in screen-shared content, and cover the sensitive information with a blurred mask area, thereby reducing the possibility of leakage of sensitive information and improving information security.

[0015] In some embodiments of this application, the processing module is configured to: classify sensitive images or sensitive text into sensitivity levels, wherein the sensitivity level corresponds one-to-one with the degree of blurring; use the display area corresponding to the sensitive image or the display area corresponding to the sensitive text as a mask area, and blur the screen transmission data according to the degree of blurring corresponding to the sensitivity level.

[0016] The technical solutions involved in the above embodiments have the following advantages or beneficial effects: The solution in this example, by classifying sensitive images or sensitive text into sensitivity levels, and with each sensitivity level corresponding to a corresponding degree of blurring, achieves hierarchical management of sensitive information protection. Strong blurring can be applied to extremely sensitive content, while weaker blurring can be applied to information with lower sensitivity. Under the premise of improving security, the identifiability of the original information can be preserved, thereby improving the readability of the content while protecting privacy.

[0017] In some embodiments of this application, the processing module is configured to: when a sensitive image or sensitive text is classified into a first sensitivity level, expand the mask area of ​​the sensitive image or sensitive text.

[0018] The technical solutions involved in the above embodiments have the following advantages or beneficial effects: The solution in this example automatically expands its mask area when a sensitive image or sensitive text is classified into the first sensitivity level, thereby further enhancing the protection of highly sensitive information and improving information security.

[0019] In some embodiments of this application, the processing module is configured to: in a first case, increase the sensitivity level of the sensitive image or sensitive text; wherein the first case is when the user performs at least one of the following: the frequency of inputting the first selection area is greater than a first frequency threshold, the speed of view transformation of the operable object is greater than a first speed, and the preset application software or system interface is opened.

[0020] The technical solutions involved in the above embodiments have the following advantages or beneficial effects: The solution in this example, by monitoring the user's operation behavior, increases the sensitivity level of sensitive images or sensitive text when the user performs a specific operation, realizes the perception of potential security risks or intentions, improves the protection strength of the corresponding sensitive information, and thus realizes the reliability of privacy and security protection.

[0021] In some embodiments of this application, the processing module is configured to: stop sending screen sharing data to the device to be projected when no operable object is detected in the first or second selection area.

[0022] The technical solutions involved in the above embodiments have the following advantages or beneficial effects: The solution in this example determines whether to continue transmitting screen sharing data by detecting the existence of an operable object in the first or second selection area. When no operable object is detected, the system automatically stops sending screen sharing data, which reduces the amount of screen sharing data sent, saves network bandwidth and computing resources, and avoids resource waste and potential information leakage risks caused by the transmission of screen sharing data with irrelevant content, thereby improving the adaptability of screen sharing.

[0023] Secondly, embodiments of this application provide a screen sharing method based on the screen sharing device described above. The method includes: in response to a screen sharing command, determining screen sharing data in real-time image data and sending the screen sharing data to the screen sharing device. The screen sharing command includes any one of the following selection areas: a first selection area and a second selection area. The first selection area is input by the user, and the second selection area is determined by identifying operable objects in the first selection area and calculating based on the operable objects after view transformation. The operable objects include any one of the following: images, cells, and tables.

[0024] The technical solutions involved in the above embodiments have the following advantages or beneficial effects: In the solution of this example, the selection area on which the screen sharing data is based is determined in the real-time image data. It can be the first selection area required by the user, or the second selection area determined by recognizing operable objects in the first selection area and based on operable objects after view transformation. This realizes that the selection area of ​​the screen sharing data can be dynamically adjusted with the changes in the shared content. While improving the continuity of the shared content, it reduces the situation where users frequently manually adjust the selection area, thereby improving the efficiency of screen sharing. Attached Figure Description

[0025] Figure 1 This is a schematic diagram illustrating the operational scenario between the screen sharing device and the control device provided in the embodiments of this application;

[0026] Figure 2 A flowchart illustrating the screen sharing method provided in this application embodiment;

[0027] Figure 3 A flowchart illustrating the screen sharing method provided in this application embodiment;

[0028] Figure 4 A flowchart illustrating the screen sharing method provided in this application embodiment;

[0029] Figure 5 A flowchart illustrating the screen sharing method provided in this application embodiment;

[0030] Figure 6 A flowchart illustrating the screen sharing method provided in this application embodiment;

[0031] Figure 7 A flowchart illustrating the screen sharing method provided in this application embodiment;

[0032] Figure 8 A flowchart illustrating the screen sharing method provided in this application embodiment;

[0033] Figure 9 A flowchart illustrating the screen sharing method provided in this application embodiment;

[0034] Figure 10 This is a schematic diagram of the screen sharing device provided in an embodiment of this application. Detailed Implementation

[0035] To make the objectives and implementation methods of this application clearer, the exemplary implementation methods of this application will be clearly and completely described below with reference to the accompanying drawings of the exemplary embodiments of this application. Obviously, the exemplary embodiments described are only some embodiments of this application, and not all embodiments.

[0036] It should be noted that the brief descriptions of terms in this application are only for the convenience of understanding the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise stated, these terms should be understood in their ordinary and common meaning.

[0037] The terms "first," "second," "third," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar or related objects or entities, and do not necessarily imply a specific order or sequence, unless otherwise specified. It should be understood that such terms are interchangeable where appropriate.

[0038] The terms “comprising” and “having”, and any variations thereof, are intended to cover but not exclude inclusion, for example, a product or device that includes a range of components is not necessarily limited to all of the components that are clearly listed, but may include other components that are not clearly listed or that are inherent to such product or device.

[0039] In the description of this application, it should be noted that, unless otherwise expressly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection between two components. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances.

[0040] The screen sharing device provided in this application can have various implementation forms, such as smart TVs, tablets, mobile phones, computers, etc. Figure 1 This is one specific implementation of the screen sharing device of this application.

[0041] Figure 1 This is a schematic diagram illustrating the operational scenario between the screen sharing device and the control device provided in an embodiment of this application. Figure 1 As shown, the user can operate the screen transmission device 200 through the smart device 300 or the control device 100.

[0042] In some embodiments, the control device 100 may be a remote control. Communication between the remote control and the screen sharing device includes infrared protocol communication, Bluetooth protocol communication, and other short-range communication methods, controlling the screen sharing device 200 wirelessly or via wired means. Users can control the screen sharing device 200 by inputting user commands through buttons on the remote control, voice input, control panel input, etc.

[0043] In some embodiments, a smart device 300 (such as a mobile terminal, tablet computer, computer, laptop computer, etc.) can also be used to control the screen sharing device 200. For example, an application running on the smart device can be used to control the screen sharing device 200.

[0044] In some embodiments, the screen sharing device may receive instructions not through the aforementioned smart devices or control devices, but through touch or gestures.

[0045] In some embodiments, the screen sharing device 200 can also be controlled in ways other than the control device 100 and the smart device 300. For example, it can be controlled by directly receiving the user's voice commands through a module configured inside the screen sharing device 200 for acquiring voice commands, or it can be controlled by receiving the user's voice commands through a voice control device set outside the screen sharing device 200.

[0046] In some embodiments, the screen sharing device 200 also communicates with the server 400. The screen sharing device 200 may communicate via a local area network (LAN), wireless local area network (WLAN), and other networks. The server 400 may provide various content and interactive features to the screen sharing device 200. The server 400 may be a cluster or multiple clusters, and may include one or more types of servers.

[0047] Screen sharing technology, as a key tool for modern communication and collaboration, has become an important bridge connecting global teams. It allows users to share computer screen content in real time, making remote collaboration and information exchange more intuitive and efficient. In remote work, screen sharing helps team members conduct real-time project discussions and presentations; in online education, it provides teachers with the convenience of displaying courseware and conducting interactive teaching; in the field of technical support, screen sharing enables technicians to directly view user problems and provide immediate solutions. By eliminating geographical barriers, screen sharing technology greatly improves communication efficiency and collaboration effectiveness, becoming an indispensable part of modern work and learning environments.

[0048] Mainstream screen sharing technologies are primarily based on full-screen capture or application window positioning. Their drawback is that users cannot precisely define the sharing boundaries using dynamic coordinates. Even in some screen sharing solutions that offer rectangular area selection, the underlying implementation still relies on fixed pixel coordinate positioning. When the application interface within the shared area scrolls, displays pop-ups, or updates content, the original selection area often fails to adjust adaptively, leading to frequent content misalignment or missing information within the shared area. This mechanical area locking mechanism forces users to repeatedly cycle through "selection area - content verification - secondary adjustment." Therefore, the current challenge is how to improve screen sharing efficiency while allowing users to automatically define the projection area.

[0049] The technical content provided in this application aims to solve the aforementioned technical problems in related technologies. In the screen sharing device and method provided in the embodiments of this application, in response to a screen sharing command, screen sharing data is determined from real-time image data and sent to the screen sharing device; the screen sharing command includes any one of the following selection areas: a first selection area and a second selection area. The solution of this application determines the selection area on which the screen sharing data is based in the real-time image data. This can be the first selection area required by the user, or a second selection area determined based on the operable object after view transformation by identifying operable objects in the first selection area. This enables the selection area of ​​the screen sharing data to be dynamically adjusted according to changes in the shared content, improving the continuity of the shared content while reducing the need for users to frequently manually adjust the selection area, thereby improving the efficiency of screen sharing.

[0050] The following example, using the processing module of a screen sharing device as the main execution unit, illustrates how screen sharing devices can improve the efficiency of screen sharing.

[0051] The technical solutions of this application will be described in detail below with reference to specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0052] Figure 2 This is a flowchart illustrating the screen sharing method provided in an embodiment of this application. Figure 2 As shown, the method includes the following steps:

[0053] S101. In response to the screen sharing command, determine the screen sharing data in the real-time image data; the screen sharing command includes any one of the following selection areas: a first selection area and a second selection area; wherein, the first selection area is input by the user, and the second selection area is determined by identifying operable objects in the first selection area and calculating based on the operable objects after the view transformation, and the operable objects include any one of the following: image, cell, table.

[0054] S102. Send the screen projection data to the device to be projected.

[0055] In some embodiments, users can directly send screen-sharing commands using the interactive interface corresponding to the screen-sharing device, or send screen-sharing commands by calling the screen-sharing device's application programming interface through other devices or software. The screen-sharing command, in addition to including a first selection area and a second selection area, may also include: the identifier or IP address of the device to be projected, and transmission parameters such as transmission protocol, resolution, scaling ratio and encoding format, timestamp, and encryption algorithm used to encrypt the projected data. It should be understood that when a user sends a screen-sharing command including the first selection area to the processing module, the processing module itself can generate a screen-sharing command including the second selection area based on the first selection area to determine the next screen-sharing data.

[0056] In some embodiments, the processing module can acquire real-time image data from the screen sharing device based on a driver interface or an application programming interface. Optionally, to improve performance, the captured real-time image data can be temporarily stored in a cache for faster processing and transmission.

[0057] In some embodiments, the first selection area can be implemented in various ways. For example, the user can drag and draw a rectangle on the screen using a mouse / stylus, and the range of the first selection area can be displayed in real time. For instance, if the screen sharing device is a touchscreen device, the user can also zoom / swipe with their finger to determine and adjust the boundaries of the first selection area. Optionally, the user can also directly input pixel coordinates, such as the upper left corner (x1, y1) and the lower right corner (x2, y2), to select the first selection area. In practical applications, the shape of the first selection area can be various, such as a rectangle, an irregular closed shape input by the user, or the shape of the corresponding operable outer contour within the first selection area.

[0058] As an optional embodiment, the processing module may only accept screen casting commands sent by the user that include the first selected area. That is, when the user sends a screen casting command that includes the first selected area A, the processing module continues to determine the screen casting data based on the first selected area A until the user sends a screen casting command that includes the first selected area B again.

[0059] For example, the processing module can flexibly transfer screen data based on the selected area in the screen casting command. For instance, when a user sends a screen casting command including a first selected area A, the transfer data is first determined based on the first selected area A. When the operable object in the first selected area A undergoes a view transformation, the processing module can determine a second selected area based on the operable object after the view transformation, and then determine the transfer data based on the second selected area. When the user sends a screen casting command again including a first selected area B, the transfer data is determined based on the first selected area B. When the operable object in the first selected area B undergoes a view transformation, the processing module can determine a second selected area based on the operable object after the view transformation, and then determine the transfer data based on the second selected area. The operable objects in the first selected areas A and B can be the same or different; this application does not limit this.

[0060] It should be noted that the operable objects in the first selection area refer to objects whose view can be changed by the user through manipulation. When the operable object is an image, this image can be from a webpage, a document, or an image viewer window. Depending on the manipulation method, the user can zoom, pan, or rotate the image. Similarly, operable objects can also be cells or tables. Optionally, operable objects can also be vector graphics, text blocks, videos in a media player, etc.

[0061] In some embodiments, the identification of operable objects in the first selection area can be achieved by locating operable objects using a pre-trained model or by recognizing structured features (such as edges and text layout) in the image. For example, when detecting tables, line distribution and text alignment rules are combined to ensure accurate identification of cell boundaries. In practical applications, deep learning models such as convolutional neural networks can be used to identify operable objects within the first selection area. Optionally, for structured data, OCR (Optical Character Recognition) can also be used to identify operable objects within the first selection area.

[0062] As an alternative embodiment, the object to be manipulated within the first selection area can also be determined through view structure analysis, such as locating operable controls like form input boxes in a webpage through the application's UI tree.

[0063] In some embodiments, when a user performs view transformation operations such as zooming or scrolling on a identified operable object, this can be determined by capturing gestures or mouse events (such as two-finger zooming or dragging displacement) in real time. Optionally, a pre-trained model can be used to track feature points of the operable object to determine whether the operable object within the first selection area has undergone a view transformation.

[0064] Furthermore, the area containing the operable object after the view transformation can be used as a second selection area to generate a projection command to update the projection data. For example, after a user zooms in on a table, the selection area can be adjusted according to the new size to ensure that the zoomed-in content is fully displayed during projection, without requiring the user to repeatedly select the area.

[0065] In the screen sharing method provided in this application embodiment, in response to a screen sharing command, screen sharing data is determined in real-time image data and sent to the screen sharing device; the screen sharing command includes any one of the following selection areas: a first selection area and a second selection area. The solution of this application determines the selection area on which the screen sharing data is based in real-time image data. This can be the first selection area required by the user, or a second selection area determined by identifying operable objects in the first selection area and tracking the changes in the view of the operable object. This enables the selection area of ​​the screen sharing data to be dynamically adjusted according to changes in the shared content, improving the continuity of the shared content while reducing the need for users to frequently manually adjust the selection area, thereby improving the efficiency of screen sharing.

[0066] Figure 3 This is a flowchart illustrating the screen sharing method provided in an embodiment of this application. Figure 3 As shown, the method also includes:

[0067] S201. In response to a screen transmission command including a first selection area, determine the first image data located in the first selection area in the real-time image data, and extract the first boundary feature and the first semantic feature from the first image data;

[0068] Determining the second selection area based on operable objects after view transformation includes:

[0069] S202. Extract at least one boundary feature and at least one semantic feature from real-time image data;

[0070] S203. Determine the second boundary feature and the second semantic feature that have the highest similarity to the first boundary feature and the first semantic feature among at least one boundary feature and at least one semantic feature;

[0071] S204. The display area corresponding to the second boundary feature and the display area corresponding to the second semantic feature shall be used as the second selection area.

[0072] In some embodiments, when extracting boundary features, image processing techniques are primarily used to identify the edges, lines, or shape contours of objects. For example, edge detection algorithms such as the Canny algorithm and Hough transform are used to find straight lines, curves, or geometric shapes within a selection area. Exemplary boundary features might be geometric information such as the border lines of a table within the selection area, the outline of an icon, or the right-angled edges of a window, such as the direction and length of a set of straight lines or the curvature of a curve.

[0073] In some embodiments, semantic features focus on the meaning of image content, such as extracting text within a selection area through text recognition, or using deep learning models to identify the type and function of interface elements such as icons and buttons. For example, semantic features might be specific text contained within the selection area, such as the function identifier of a "Confirm" button or icon, or the type of interface element, such as an input box or scroll bar.

[0074] Furthermore, after the user specifies the first selection area, the processing module searches the entire screen for the region most similar to the first selection area, i.e., the second selection area. Here, "second boundary features" refer to the geometric features most similar in shape to the first selection area (such as similar border lines), and "second semantic features" refer to features that match the meaning of the content in the first selection area (such as buttons that also contain the text "Delete"). The processing module determines the most relevant second selection area by matching these two types of features. In practical applications, if the user selects a table area (the first selection area), the system extracts the table's border lines as boundary features and the table header text such as "Serial Number" and "Name" as semantic features. When the user switches screens, the system searches for regions containing similar border lines and header text as the second selection area. The proposed solution extracts first boundary features and first semantic features from the first selection area, searches for the features with the highest similarity in the entire real-time image data, and determines the area most relevant to the user's initial selection as the second selection area. Based on the features of the user's initial selection area, it can track and locate highly relevant content areas in the entire screen, making the determination of the second selection area not limited by the initial position and able to adapt to the dynamic changes of the shared content, thereby improving the accuracy of the shared content.

[0075] Figure 4 This is a flowchart illustrating the screen sharing method provided in an embodiment of this application. Figure 4 As shown, S204 includes:

[0076] S301. Determine the coordinates of the first region corresponding to the second boundary feature, and determine the coordinates of the second region corresponding to the second semantic feature;

[0077] S302. Perform weighted fusion of the coordinates of the first region and the coordinates of the second region to obtain the coordinates of the region corresponding to the second selection area.

[0078] In some embodiments, for operable objects including text and images, the coordinates of the bounding box of a button icon detected using the YOLO (You Only LookOnce) model can be used as the coordinates of the first region corresponding to the second boundary feature, and the coordinates of the text box detected using OCR can be used as the coordinates of the second region corresponding to the second semantic feature. Optionally, for the boundary feature coordinates of table borders, the vertical and horizontal lines of the table can be detected using Hough transform, and the coordinates of the two outermost vertical lines and the two outermost horizontal lines can be determined as the coordinates of the first region corresponding to the second boundary feature.

[0079] For example, weights can be dynamically adjusted based on feature reliability. Higher boundary feature weights emphasize shape matching, while higher semantic weights emphasize content matching. In practical applications, for structured interfaces like Excel, a boundary weight of 0.7 can be set if table line detection accuracy is high, and a semantic weight of 0.3 can be set if text within cells may change dynamically. As an optional example, for dynamic content interfaces like video players, a boundary weight of 0.4 can be set if the control bar shape may be obscured by the progress bar, and a semantic weight of 0.6 can be set if "play" / "pause" text recognition is required.

[0080] In this example, the coordinates of the second selection area are obtained by weighted fusion of the coordinates of the first region corresponding to the second boundary feature and the coordinates of the second region corresponding to the second semantic feature. This ensures that the determined second selection area can continuously include changing text or borders, achieving content consistency during continuous content sharing.

[0081] Figure 5 This is a flowchart illustrating the screen sharing method provided in an embodiment of this application. Figure 5 As shown, the method also includes:

[0082] S401. When the operable object in the first selected area responds to the user's operation and performs a view transformation, motion prediction is performed on the second boundary features to obtain the motion prediction result.

[0083] S402. Based on the motion prediction results, the coordinates of the first region corresponding to the second boundary feature are corrected. The correction operation includes expanding or shrinking the coordinates of the first region.

[0084] In some embodiments, when a user drags a table horizontally and causes the boundary features to move laterally, the optical flow method can be used to calculate the movement vector of the boundary corner points of two adjacent frames, and the displacement of the next frame can be predicted based on the movement speed.

[0085] As an optional embodiment, to address the diffusion of road boundary features caused by zooming in on the map, feature combinations of boundary points can be extracted to determine an initial feature matrix, and the scaled intermediate feature matrix can be matched. Based on the matrix parameters of the two feature matrices and the Kalman filter algorithm, the position of the future updated feature matrix can be predicted.

[0086] In practical applications, when the list is detected to be scrolling downwards, the prediction boundary will continue to shift downwards. For example, if the bottom of the selection area was originally 800 pixels on the screen, and the prediction is that it will shift down another 30 pixels in the next frame, the bottom of the selection area will be automatically expanded to 845 pixels (original position + 1.5 times the predicted distance buffer). This way, when the user scrolls quickly, the corrected selection area can cover the content to be displayed in advance, preventing the target area from being suddenly lost. Optionally, if the predicted boundary exceeds the screen edge, for example, the right side is predicted to be 1200 pixels, but the actual screen width is only 1000 pixels, the system will actively shrink the right boundary back to 800 pixels. This preserves the core area (such as the main roads on the map) while avoiding selecting invalid areas outside the screen due to over-prediction, ensuring that the target area is always visible.

[0087] In this example, when the content in the first selected area changes, the coordinates of the first area corresponding to the second boundary feature are corrected based on the motion prediction results of the second boundary feature. This allows the update of the second area coordinates to track the rhythm and speed of the content changes in the first selected area in real time, reducing the positional deviation caused by update lag, thereby improving the real-time performance of shared content.

[0088] Figure 6 This is a flowchart illustrating the screen sharing method provided in an embodiment of this application. Figure 6 As shown, the method also includes:

[0089] S501. Based on the sensitive information database, extract sensitive images after performing image recognition on real-time image data, or determine the sensitive text corresponding to the sensitive semantic features in the second semantic features;

[0090] S502. Use the display area corresponding to the sensitive image or the display area corresponding to the sensitive text as a mask area to blur the screen transmission data.

[0091] It should be noted that this solution is executed after the second selection area is determined and before the screen sharing data is sent. In some embodiments, sensitive images and sensitive text may be obtained based on a sensitive information database, or only sensitive images may be obtained, or only sensitive text may be obtained, depending on the second boundary features and second semantic features in the second selection area, as well as the user's privacy protection requirements.

[0092] In some embodiments, a database containing sensitive image templates and sensitive keywords is established through a combination of manual annotation and automated crawling. In practical applications, users can also add specific sensitive information to the sensitive information database to achieve privacy identification and protection for specific scenarios. For example, in enterprise data security scenarios, templates of sensitive images such as employee ID cards and work badges are manually collected, and samples are taken from different angles and under different lighting conditions; or keywords such as "salary" and "contract number" are extracted from internal documents. In financial industry scenarios, images of the front and back of bank cards are collected, and data augmentation is used to generate blurred, tilted, and other deformed samples; or sensitive fields in bank agreements are crawled to form a regular expression rule base.

[0093] In some embodiments, after determining the mask area, the screen transmission data can be blurred using methods such as Gaussian blur or pixel mosaic to update the screen transmission data.

[0094] This example solution uses image recognition based on a sensitive information database to extract sensitive images or determine the sensitive text corresponding to sensitive semantic features in the second semantic features. It can identify and locate potential sensitive information in screen-shared content, and use a blurred mask area to cover the sensitive information, thereby reducing the possibility of leakage of sensitive information and improving information security.

[0095] Figure 7 This is a flowchart illustrating the screen sharing method provided in an embodiment of this application. Figure 7 As shown, the method also includes:

[0096] S601. Sensitivity levels are classified for sensitive images or sensitive text, with each sensitivity level corresponding to a one-to-one degree of blurring.

[0097] Specifically, S502 includes:

[0098] S602. Use the display area corresponding to the sensitive image or the display area corresponding to the sensitive text as a mask area, and blur the screen transmission data according to the blurring degree corresponding to the sensitivity level.

[0099] In some embodiments, the low-sensitivity level can be set to use a slight Gaussian blur, preserving the general outline but obscuring details; the medium-sensitivity level can be set to cover the middle few digits of the card number or ID number with a dense mosaic; the high-sensitivity level can be set to completely cover the information with a black block, displaying only the field name. Optionally, text replacement or dynamic pixelation technology can also be used to ensure that the information cannot be recovered even if the user takes a screenshot, thus improving privacy and security.

[0100] As an optional embodiment, five sensitivity levels can be set corresponding to 1 to 5 levels of non-Gaussian blurring, or the data can be directly removed when the blurring level is the highest. Optionally, differential privacy technology can also be used to ensure that the blurring process is irreversible, such as when the PSNR (Peak Signal-to-Noise Ratio) is <20dB and the reconstruction error is >95%.

[0101] The solution presented in this example categorizes sensitive images or texts into different sensitivity levels, with each level corresponding to a specific degree of blurring. This achieves hierarchical management of sensitive information protection. Strong blurring can be applied to extremely sensitive content, while weaker blurring can be applied to less sensitive information. This approach enhances security while preserving the original information's identifiability, thereby improving readability while protecting privacy.

[0102] As yet another example, based on any example, the screen sharing method also includes: when a sensitive image or sensitive text is classified as the first sensitivity level, expanding the mask area of ​​the sensitive image or sensitive text.

[0103] It should be noted that the first sensitivity level is the most sensitive level and the level with the highest degree of privacy ambiguity.

[0104] In some embodiments, when an ID card photo is detected as belonging to the highest level of sensitivity (Level 1), the masking area is automatically expanded. For example, if it originally only obscured the ID number area, the upgraded masking will cover the entire ID card. In some embodiments, if a bank card number is identified as belonging to the highest level of sensitivity, the text masking area will be expanded. For example, if the original processing only covered the middle 8 digits, the upgraded masking will cover the entire card number area. In practical applications, the masking area boundary can be expanded by 5 pixels to enhance privacy protection.

[0105] This example demonstrates a solution that enhances information security by automatically expanding the mask area of ​​sensitive images or text when they are classified as the first level of sensitivity.

[0106] As yet another example, based on any example, the screen sharing method further includes: in a first case, increasing the sensitivity level of the sensitive image or sensitive text; wherein the first case is when the user performs at least one of the following: the frequency of inputting the first selection area is greater than a first frequency threshold, the speed of view transformation of the operable object is greater than a first speed, and the preset application software or system interface is opened.

[0107] It is understandable that the first scenario is an abnormal user operation. In some embodiments, if the frequency of inputting the first selection area is too high, such as more than 3 times / second, when it is detected that the user selects multiple bank card number areas, such as selecting 4 different card numbers within 1 second, it is determined to be an abnormal information collection behavior, and the sensitivity level of the card number area is raised from low to high. Optionally, raising the sensitivity level can also expand the blurred mask area to the entire bank card image.

[0108] In some embodiments, if the view transformation speed is abnormal, such as exceeding 500 pixels / second, it is assumed that the user wants to display more details by zooming or to display more private data by scrolling quickly, and the sensitivity level is increased in this case.

[0109] In some embodiments, the preset application software or system interface includes, for example, a chat program, screen recording software, a task manager, and user-defined privacy folders and development files.

[0110] As an optional embodiment, abnormal user operations may also include: high-frequency copying of sensitive data from the clipboard, such as automatically replacing the ID card number in the clipboard with "***" when it is detected that the user has copied 5 ID card numbers in a row, and restricting subsequent copying operations.

[0111] The technical solutions involved in the above embodiments have the following advantages or beneficial effects: The solution in this example, by monitoring the user's operation behavior, increases the sensitivity level of sensitive images or sensitive text when the user performs a specific operation, realizes the perception of potential security risks or intentions, improves the protection strength of the corresponding sensitive information, and thus realizes the reliability of privacy and security protection.

[0112] As yet another example, based on any example, the screen sharing method further includes the step of stopping the transmission of screen sharing data to the device to be projected when no operable object is detected in the first or second selection area.

[0113] In some embodiments, for slideshow presentations, preset operable objects, such as presentation images, are continuously detected in the first selection area. If the image is not detected multiple times consecutively by an image recognition algorithm such as the YOLO model, and the background of the selection area is significantly different from the button features (e.g., a solid color background replaces the image), it is determined that "no operable object exists." At this point, sending data for that area to the screen sharing device is immediately stopped, and a prompt such as "Operation controls are missing, screen sharing has been interrupted" appears.

[0114] In some embodiments, when the operable object is covered by the interface of other application software and cannot be detected, the process of sending screen sharing data to the device to be projected can also be stopped.

[0115] This example solution determines whether to continue transmitting screen sharing data by detecting the presence of an operable object in the first or second selection area. When no operable object is detected, the system automatically stops transmitting screen sharing data, reducing the amount of screen sharing data transmitted, saving network bandwidth and computing resources, and avoiding resource waste and potential information leakage risks caused by transmitting screen sharing data with irrelevant content, thereby improving the adaptability of screen sharing.

[0116] Figure 8 This is a flowchart illustrating the screen sharing method provided in an embodiment of this application. Figure 8 As shown, after determining the first boundary features and first semantic features based on the first selected region, the real-time image data processing flow can include: on the one hand, boundary feature extraction (e.g., using the KCF (Kernelized Correlation Filters) algorithm or the HOG (Histogram of Oriented Gradients) algorithm) can be performed to obtain the second boundary feature most similar to the first boundary feature. Motion prediction is continuously performed on the second boundary feature (e.g., using a Kalman filter), and the coordinates of the first region corresponding to the second boundary feature are corrected. On the other hand, semantic feature extraction (e.g., using a Bidirectional Long Short-Term Memory (BiLSTM) network) can be performed, and semantic association is performed through an attention mechanism to obtain the second semantic feature most similar to the first semantic feature, as well as the coordinates of the second region corresponding to the second semantic feature. Coordinate fusion is performed based on pre-determined weights, i.e., the corrected first region coordinates × α + the second region coordinates × β, and the fusion result is used as the region coordinates corresponding to the second selected region.

[0117] Figure 9 This is a flowchart illustrating the screen sharing method provided in an embodiment of this application. Figure 9 As shown, after determining the screen transmission data in the second selection area, based on the sensitive information database, sensitive text corresponding to the sensitive semantic features is determined in the second semantic features (such as keyword matching through regular expressions or semantic relevance calculation). At the same time, image recognition is performed on the real-time image data in the second selection area to extract sensitive images, and regional coordinate positioning is performed to determine the display area corresponding to the sensitive image and the display area corresponding to the sensitive text. Sensitivity levels are classified for sensitive images or sensitive text, and when abnormal user operations are detected, the sensitivity level corresponding to the sensitive image or sensitive text is increased. Finally, the display area corresponding to the sensitive image or the display area corresponding to the sensitive text is used as a mask area, and the screen transmission data is blurred according to the blurring degree corresponding to the classified sensitivity level.

[0118] In the screen sharing method provided in this application embodiment, in response to a screen sharing command, screen sharing data is determined in real-time image data and sent to the screen sharing device; the screen sharing command includes any one of the following selection areas: a first selection area and a second selection area. The solution of this application determines the selection area on which the screen sharing data is based in real-time image data. This can be the first selection area required by the user, or a second selection area determined by identifying operable objects in the first selection area and tracking the changes in the view of the operable object. This enables the selection area of ​​the screen sharing data to be dynamically adjusted according to changes in the shared content, improving the continuity of the shared content while reducing the need for users to frequently manually adjust the selection area, thereby improving the efficiency of screen sharing.

[0119] Figure 10 This is a schematic diagram of the screen sharing device provided in an embodiment of this application. Figure 10 As shown, the device includes:

[0120] Display screen 91 is configured to display real-time image data;

[0121] The processing module 92, connected to the display screen, is configured to: in response to a screen sharing command, determine screen sharing data from real-time image data and send the screen sharing data to the device to be projected; the screen sharing command includes any one of the following selection areas: a first selection area and a second selection area; wherein, the first selection area is input by the user, and the second selection area is determined by the processing module 92 based on the operable object identified in the first selection area and determined based on the operable object after the view transformation, and the operable object includes any one of the following: an image, a cell, or a table.

[0122] In some embodiments of this application, the processing module 92 is further configured to: in response to a screen transmission command including a first selection area, determine first image data located in the first selection area in real-time image data, and extract a first boundary feature and a first semantic feature from the first image data; specifically, the processing module 92 is configured to: extract at least one boundary feature and at least one semantic feature from the real-time image data; determine a second boundary feature and a second semantic feature that have the highest similarity to the first boundary feature and the first semantic feature among the at least one boundary feature and the at least one semantic feature; and use the display area corresponding to the second boundary feature and the display area corresponding to the second semantic feature as the second selection area.

[0123] In some embodiments of this application, the processing module 92 is specifically configured to: determine the coordinates of the first region corresponding to the second boundary feature, and determine the coordinates of the second region corresponding to the second semantic feature; perform weighted fusion of the first region coordinates and the second region coordinates to obtain the region coordinates corresponding to the second selection area.

[0124] In some embodiments of this application, the processing module 92 is further configured to: when the operable object in the first selected area responds to the user's operation to perform a view transformation, perform motion prediction on the second boundary feature, and perform a correction operation on the coordinates of the first region corresponding to the second boundary feature based on the motion prediction result, the correction operation including expanding or shrinking the coordinates of the first region.

[0125] In some embodiments of this application, the processing module 92 is further configured to: extract sensitive images from real-time image data based on a sensitive information database, or determine sensitive text corresponding to sensitive semantic features in a second semantic feature; and use the display area corresponding to the sensitive image or the display area corresponding to the sensitive text as a mask area to blur the screen transmission data.

[0126] In some embodiments of this application, the system is further configured to: classify sensitive images or sensitive text into sensitivity levels, wherein the sensitivity level corresponds one-to-one with the degree of blurring; the processing module 92 is specifically configured to: use the display area corresponding to the sensitive image or the display area corresponding to the sensitive text as a mask area, and blur the screen transmission data according to the degree of blurring corresponding to the sensitivity level.

[0127] In some embodiments of this application, the processing module 92 is further configured to: when a sensitive image or sensitive text is classified into a first sensitivity level, expand the mask area of ​​the sensitive image or sensitive text.

[0128] In some embodiments of this application, the processing module 92 is further configured to: in a first case, increase the sensitivity level corresponding to the sensitive image or sensitive text; wherein the first case is when the user performs at least one of the following: the frequency of inputting the first selection area is greater than a first frequency threshold, the speed of view transformation of the operable object is greater than a first speed, and the preset application software or system interface is opened.

[0129] In some embodiments of this application, the processing module 92 is further configured to: stop sending screen sharing data to the device to be projected when no operable object is detected in the first selection area or the second selection area.

[0130] The screen sharing device provided in this application embodiment can execute the screen sharing method in the above method embodiment. Its implementation principle and technical effect are similar, and will not be described again here.

[0131] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

[0132] For ease of explanation, the above description has been provided in conjunction with specific embodiments. However, the above exemplary discussion is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. Various modifications and variations can be obtained based on the above teachings. The selection and description of the above embodiments are for the purpose of better explaining the principles and practical applications, thereby enabling those skilled in the art to better utilize the embodiments and various different variations of the embodiments suitable for specific application considerations.

Claims

1. A screen transmission device, characterized in that, include: The display screen is configured to display real-time image data; The processing module, connected to the display screen, is configured to: in response to a screen sharing command, determine screen sharing data from the real-time image data and send the screen sharing data to the device to be projected; the screen sharing command includes any one of the following selection areas: a first selection area and a second selection area; wherein, the first selection area is input by the user, and the second selection area is determined by the processing module based on the operable object identified in the first selection area and the operable object after the view transformation, and the operable object includes any one of the following: an image, a cell, or a table.

2. The screen transmission device according to claim 1, characterized in that, The processing module is configured as follows: In response to a screen sharing command that includes a first selection area, a first image data located in the first selection area in the real-time image data is determined, and a first boundary feature and a first semantic feature are extracted from the first image data; Extract at least one boundary feature and at least one semantic feature from the real-time image data; Determine the second boundary feature and the second semantic feature that have the highest similarity to the first boundary feature and the first semantic feature among the at least one boundary feature and the at least one semantic feature; The display area corresponding to the second boundary feature and the display area corresponding to the second semantic feature are used as the second selection area.

3. The screen transmission device according to claim 2, characterized in that, The processing module is configured as follows: Determine the coordinates of the first region corresponding to the second boundary feature, and determine the coordinates of the second region corresponding to the second semantic feature; The coordinates of the first region and the coordinates of the second region are weighted and fused to obtain the coordinates of the region corresponding to the second selected area.

4. The screen transmission device according to claim 3, characterized in that, The processing module is configured as follows: When an operable object in the first selected area responds to a user operation and performs a view transformation, motion prediction is performed on the second boundary feature, and the coordinates of the first region corresponding to the second boundary feature are corrected based on the motion prediction result. The correction operation includes expanding or shrinking the coordinates of the first region.

5. The screen transmission device according to claim 2, characterized in that, The processing module is configured as follows: Based on the sensitive information database, sensitive images are extracted after image recognition of the real-time image data, or sensitive text corresponding to the sensitive semantic features is determined in the second semantic features; The display area corresponding to the sensitive image or the display area corresponding to the sensitive text is used as a mask area to blur the screen transmission data.

6. The screen transmission device according to claim 5, characterized in that, The processing module is configured as follows: The sensitive images or sensitive texts are classified into sensitivity levels, wherein the sensitivity levels correspond one-to-one with the degree of blurring. The display area corresponding to the sensitive image or the display area corresponding to the sensitive text is used as a mask area, and the screen transmission data is blurred according to the blurring degree corresponding to the sensitivity level.

7. The screen transmission device according to claim 6, characterized in that, The processing module is configured as follows: When the sensitive image or sensitive text is classified as the first sensitivity level, the mask area of ​​the sensitive image or sensitive text is expanded.

8. The screen transmission device according to claim 6, characterized in that, The processing module is configured as follows: In a first case, the sensitivity level corresponding to the sensitive image or sensitive text is increased; wherein the first case is when the user performs at least one of the following: the frequency of inputting the first selection area is greater than a first frequency threshold, the speed of view transformation of the operable object is greater than a first speed, and a preset application software or system interface is opened.

9. The screen transmission device according to any one of claims 1-8, characterized in that, The processing module is configured as follows: When no operable object is detected in the first or second selection area, the transmission of screen data to the device to be projected is stopped.

10. A screen sharing method, characterized in that, The method is based on the screen sharing device as described in any one of claims 1-9; the method includes: In response to a screen sharing command, screen sharing data is determined from real-time image data and sent to the screen sharing device; the screen sharing command includes any one of the following selection areas: a first selection area and a second selection area; wherein, the first selection area is input by the user, and the second selection area is determined by identifying operable objects in the first selection area and calculating based on the operable objects after view transformation, and the operable objects include any one of the following: images, cells, and tables.

Citation Information

Cited By

  • Information access control in workspace ecosystems

    US20250252201A1