Screen sharing methods, devices, computer equipment, and storage media
By acquiring the performance parameters of the video sharing device during screen sharing, extracting and sending metadata for resolution enhancement, the problem of balancing video smoothness and quality under bandwidth limitations is solved, achieving high-quality screen sharing.
Patent Information
- Application Number
- CN202111166590.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-30
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2041-09-30
AI Technical Summary
Existing screen sharing methods struggle to balance video smoothness and quality under bandwidth constraints.
By obtaining the performance parameters of the video sharing device, extracting the metadata of the shared video frames, and sending the low-resolution video frames and metadata to the video receiving device, the resolution enhancement processing is performed using the metadata to improve video quality.
While maintaining a good frame rate, it improved the clarity of shared video, achieving high-quality screen sharing under bandwidth constraints.
Smart Images

Figure CN115883768B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data communication, and in particular to a screen sharing method, apparatus, computer device, and storage medium. Background Technology
[0002] Screen sharing is a common scenario in online meetings, with presentations and videos being the most frequently shared objects. For shared video streams, maintaining a smooth streaming experience by keeping the frame rate within a good range and preserving frame quality is crucial.
[0003] However, current screen sharing methods, due to bandwidth limitations, struggle to balance video smoothness and quality. Summary of the Invention
[0004] In a first aspect, embodiments of this application provide a screen sharing method, the method comprising:
[0005] Obtain the performance parameters of the video sharing device;
[0006] Based on the performance parameters, the corresponding metadata extraction operation is performed on the shared video frames of the video sharing device to obtain the metadata in the shared video frames; the metadata includes the target content information of the shared video frames.
[0007] The low-resolution video frame corresponding to the shared video frame, along with the metadata, is sent to a video receiving device associated with the video sharing device. The video receiving device then uses the metadata to perform resolution enhancement processing on the target content information corresponding to the low-resolution video frame and displays the processed video frame.
[0008] Secondly, embodiments of this application provide another screen sharing method, the method comprising:
[0009] The system receives a low-resolution video frame corresponding to a shared video frame sent by a video sharing device, as well as metadata of the shared video frame. The metadata is obtained by the video sharing device through a metadata extraction operation on the shared video frame based on the performance parameters of the video sharing device. The metadata includes target content information of the shared video frame.
[0010] Based on the metadata, the target content information in the low-resolution video frame is subjected to resolution enhancement processing to obtain the processed video frame.
[0011] The processed video frames are displayed.
[0012] Thirdly, embodiments of this application provide a screen sharing device, the device comprising:
[0013] The acquisition module is used to acquire the performance parameters of the video sharing device;
[0014] The extraction module is used to perform corresponding metadata extraction operations on the shared video frames of the video sharing device based on the performance parameters, so as to obtain the metadata in the shared video frames; the metadata includes the target content information of the shared video frames;
[0015] The sending module is used to send the low-resolution video frame corresponding to the shared video frame and the metadata to the video receiving device associated with the video sharing device, so that the video receiving device can perform resolution enhancement processing on the target content information corresponding to the low-resolution video frame through the metadata and display the processed video frame.
[0016] Fourthly, embodiments of this application also provide another screen sharing device, the device comprising:
[0017] The receiving module is used to receive a low-resolution video frame corresponding to a shared video frame sent by a video sharing device, as well as the metadata of the shared video frame; the metadata is obtained by the video sharing device based on the performance parameters of the video sharing device by performing a corresponding metadata extraction operation on the shared video frame; the metadata includes the target content information of the shared video frame.
[0018] The processing module is used to perform resolution enhancement processing on the target content information in the low-resolution video frame according to the metadata, so as to obtain the processed video frame.
[0019] The display module is used to display the processed video frames.
[0020] Fifthly, embodiments of this application provide a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:
[0021] Obtain the performance parameters of the video sharing device;
[0022] Based on the performance parameters, the corresponding metadata extraction operation is performed on the shared video frames of the video sharing device to obtain the metadata in the shared video frames; the metadata includes the target content information of the shared video frames.
[0023] The low-resolution video frame corresponding to the shared video frame and the metadata of the shared video frame are sent to the video receiving device associated with the video sharing device, so that the video receiving device can perform resolution enhancement processing on the target content information corresponding to the low-resolution video frame through the metadata and display the processed video frame.
[0024] Sixthly, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, performs the following steps:
[0025] Obtain the performance parameters of the video sharing device;
[0026] Based on the performance parameters, the corresponding metadata extraction operation is performed on the shared video frames of the video sharing device to obtain the metadata in the shared video frames; the metadata includes the target content information of the shared video frames.
[0027] The low-resolution video frame corresponding to the shared video frame and the metadata of the shared video frame are sent to the video receiving device associated with the video sharing device, so that the video receiving device can perform resolution enhancement processing on the target content information corresponding to the low-resolution video frame through the metadata and display the processed video frame.
[0028] The aforementioned screen sharing method, apparatus, computer device, and storage medium first extract metadata from the video frames to be shared based on the performance parameters of the video sharing device. This yields metadata containing the target content information of the shared video frames. Then, the low-resolution video frame corresponding to the shared video frame, along with its metadata, is sent to the associated video receiving device. The video receiving device then uses the metadata to enhance the resolution of the target content information in the low-resolution video frame and displays the processed video frame. This solution, on the one hand, sends the video frames to be shared to the video receiving device in a low-resolution format to ensure smoothness of the shared video; on the other hand, by extracting metadata from the shared video frames and sending it to the video receiving device, the receiving device can enhance the resolution of the low-resolution video frames using the metadata, thus ensuring the clarity of the shared video. Therefore, during screen sharing, both a good frame rate and high-quality shared video are guaranteed. Attached Figure Description
[0029] Figure 1 This is a diagram illustrating the application environment of a screen sharing method in one embodiment;
[0030] Figure 2 This is a flowchart illustrating a screen sharing method in one embodiment;
[0031] Figure 3a This is a schematic diagram of a shared video frame displayed by a video receiving device in one embodiment;
[0032] Figure 3b This is another schematic diagram showing a shared video frame displayed by a video receiving device in one embodiment;
[0033] Figure 3c This is yet another schematic diagram of a shared video frame displayed by a video receiving device in one embodiment;
[0034] Figure 4 This example illustrates metadata extraction operations under different performance parameters in one embodiment.
[0035] Figure 5 This is a schematic diagram illustrating object detection on a shared screen using an object detector in one embodiment;
[0036] Figure 6 This is a flowchart illustrating a screen sharing method in another embodiment;
[0037] Figure 7 This is a structural block diagram of a screen sharing device in one embodiment;
[0038] Figure 8 This is a structural block diagram of a screen sharing device in another embodiment;
[0039] Figure 9 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0040] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0041] It should be noted that the terms "first," "second," "third," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in sequences other than those illustrated or described herein.
[0042] The screen sharing method provided in this application can be applied to, for example... Figure 1The application environment shown includes a video sharing device 102 and a video receiving device 104. The video sharing device 102 communicates with the video receiving device 104 via a network. The video sharing device 102 can represent the provider of shared content on the shared screen, and the video receiving device 104 can represent the recipient of the shared content. The video sharing device 102 and the video receiving device 104 can be, but are not limited to, various personal computers, laptops, smartphones, and tablets. In the application scenario of this application, the video sharing device 102 performs metadata extraction operations on the shared video frames based on its own performance parameters to obtain the metadata in the shared video frames. It then sends the corresponding low-resolution video frame along with the metadata to the video receiving device 104. After receiving the low-resolution video frame and the metadata, the video receiving device 104 performs resolution enhancement processing on the target content information corresponding to the low-resolution video frame using the metadata and displays the processed video frame.
[0043] In one embodiment, such as Figure 2 As shown, a screen sharing method is provided, which can be applied to Figure 1 Taking video sharing device 102 as an example, the following steps are included:
[0044] Step S202: Obtain the performance parameters of the video sharing device.
[0045] Among them, performance parameters can be used to characterize the processing capabilities of video sharing devices. Specifically, performance parameters can be the performance parameters of components such as processors and CPUs (Central Processing Units), such as the CPU's clock speed and external clock speed.
[0046] Step S204: Based on performance parameters, perform corresponding metadata extraction operations on the shared video frames of the video sharing device to obtain the metadata in the shared video frames; the metadata includes the target content information of the shared video frames.
[0047] The target content information can be text information in the shared video frame, important content information in the shared video frame, etc. The important content information can be determined based on the context of the previous video frame of the shared video frame.
[0048] In practical implementation, the video sharing device 102 will have different processing capabilities under different performance parameters. Therefore, multiple parameter ranges can be pre-defined, and each parameter range can specify a corresponding metadata type to be extracted. The metadata type can include text information, object information of predefined objects, and object information of voice objects, etc. After obtaining the performance parameters of the video sharing device 102, the target parameter range corresponding to these performance parameters is determined. The target metadata type corresponding to the target parameter range is taken as the metadata type to be extracted by the video sharing device 102. Furthermore, the metadata extraction method corresponding to this target metadata type is used to perform metadata extraction operations on the shared video frames to obtain metadata containing the target content information of the shared video frames.
[0049] Step S206: Send the low-resolution video frame corresponding to the shared video frame and the metadata of the shared video frame to the video receiving device associated with the video sharing device, so that the video receiving device can perform resolution enhancement processing on the target content information corresponding to the low-resolution video frame through the metadata and display the processed video frame.
[0050] The low-resolution video frames have a resolution lower than the original resolution of the shared video frames.
[0051] In practical implementation, due to bandwidth limitations, to ensure the smoothness of the shared video, after extracting the metadata from the shared video frame, the resolution of the shared video frame needs to be adjusted before sending it to the video receiving device 104. Specifically, the resolution of the shared video frame needs to be lowered to obtain a low-resolution video frame, which is then sent to the video receiving device 104 associated with the video sharing device 102 to ensure the smoothness of the shared video. Simultaneously, while ensuring the smoothness of the shared video, to guarantee video quality, the extracted metadata from the shared video frame can be sent together with the low-resolution video frame to the video receiving device 104. This allows the video receiving device 104 to perform resolution enhancement processing on the corresponding target content information in the low-resolution video frame using the metadata, thereby improving the resolution of the video frame displayed by the video receiving device.
[0052] In the aforementioned screen sharing method, firstly, based on the performance parameters of the video sharing device, metadata extraction is performed on the video frames to be shared to obtain metadata containing the target content information of the shared video frames. Then, the low-resolution video frame corresponding to the shared video frame, along with its metadata, is sent to the video receiving device associated with the video sharing device. The video receiving device then uses the metadata to perform resolution enhancement processing on the target content information in the low-resolution video frame and displays the processed video frame. This solution, on the one hand, sends the video frames to be shared to the video receiving device in a low-resolution format to ensure the smoothness of the shared video; on the other hand, by extracting the metadata from the shared video frames and sending it to the video receiving device, the video receiving device can use the metadata to enhance the resolution of the low-resolution video frames to ensure the clarity of the shared video. Thus, during screen sharing, both a good frame rate and the quality of the shared video are guaranteed.
[0053] In one embodiment, step S204 specifically includes: when it is determined that the performance parameters of the video sharing device are within a first parameter range, performing character recognition operation on the shared video frame to obtain the text information of the shared video frame as the metadata of the shared video frame.
[0054] The first parameter range can be understood as the range of performance parameters with the weakest processing power. For example, taking the CPU external frequency as the performance parameter, the first parameter range can be 60-90MHz.
[0055] The text information can include information such as the text content of the characters, the text font, the text color, and the text position.
[0056] In practical applications, character recognition operations can be achieved using optical character recognition (OCR) algorithms.
[0057] In the specific implementation, after obtaining the performance parameters of the video sharing device, the performance parameters are matched one by one with each parameter range. When it is determined that the performance parameters of the video sharing device are in the first performance parameter range with the weakest processing capability, since the processing capability is the weakest, the video sharing device can perform some simple metadata extraction operations. That is, the video sharing device can perform character recognition operations on the shared video frame through the optical character recognition algorithm to obtain the text information in the shared video frame as the metadata of the shared video frame.
[0058] In one embodiment, step S204 further includes: when it is determined that the performance parameters of the video sharing device are within the second parameter range, performing character recognition operation on the shared video frame to obtain the text information of the shared video frame; and, according to the predefined object, performing object detection on the shared video frame to obtain the object information of the predefined object, and using the text information and the object information of the predefined object as the metadata of the shared video frame.
[0059] The second parameter range can be understood as the range of performance parameters with weaker processing power. For example, taking the CPU external frequency as the performance parameter, the second parameter range can be 90-120MHz.
[0060] Among them, the predefined objects can be objects that are considered very important, such as faces, logos, etc., and the predefined objects can also have set characteristics, such as a child's face or a logo containing a circle.
[0061] Among them, object information can be information that can characterize the features of an object. Specifically, object information can be the object's position information, pixel information, color information, etc. in the shared video frame.
[0062] In specific implementation, once it is determined that the performance parameters of the video sharing device are in the second performance parameter range with weaker processing capabilities, in addition to performing character recognition operations on the shared video frames to obtain the text information of the shared video frames, object detection can also be performed on the shared video frames according to predefined objects to obtain the object information of the predefined objects. Furthermore, the predefined objects can be sorted according to the chronology followed in the predefined set, and the text information and the object information of the sorted predefined objects can be used as the metadata of the shared video frames.
[0063] For example, if the predefined object is a face, then all faces are detected from the shared video frames, and the positional information of each face is determined as the object information of the predefined object "face". Similarly, if the predefined object is a logo containing a circle, then all logos containing a circle are detected from the shared video frames, and the positional, color, and shape information of the logos containing circles are obtained as the object information of the predefined object "logo containing a circle".
[0064] In one embodiment, step S204 further includes: when it is determined that the performance parameters of the video sharing device are within the third parameter range, performing character recognition operation on the shared video frame to obtain the text information of the shared video frame; and determining the voice object based on the voice of the user account during screen sharing, obtaining the object information of the voice object, and using the text information and the object information of the voice object as metadata of the shared video frame.
[0065] The third parameter range can be understood as the range of performance parameters with the strongest processing capability. For example, taking the CPU external frequency as the performance parameter, the third parameter range can be 120-200MHz.
[0066] In practice, once the performance parameters of the video sharing device are determined to be within the third performance parameter range with the strongest processing capability, in addition to performing character recognition operations on the shared video frames to obtain the text information of the shared video frames, voice objects can also be considered. The voice objects are determined based on the voice of the user account during screen sharing, and the object information of the voice objects is obtained. The text information and the object information of the voice objects are used as the metadata of the shared video frames.
[0067] More specifically, the process of obtaining object information of a voice object includes: during screen sharing, detecting the user account's voice in real time using a speech-to-text model, obtaining the voice objects mentioned by the user account, and determining the position of the voice objects in the video frames in order to extract the object information of the voice objects.
[0068] For example, during screen sharing, the video receiving device displays low-resolution video frames such as... Figure 3a As shown in the image, the two objects "dog" and "cat" are both at low resolution. When the user says something about "dog," such as "dogs are our best friends," the speech-to-text model will detect the presence of the "dog" object in the speech, determine its position in the video frame, extract its pixel and position information, and send this information as the "dog's" object information to the video receiving device. The receiving device then uses this information, along with the received pixel data, to perform resolution enhancement processing on the "dog" object in the currently displayed low-resolution video frame, increasing its resolution to display it as shown in the image. Figure 3b The image of the "dog" is clearly displayed until the speech-to-text model detects that the user account mentions other objects. For example, if the user account then says, "Cats are our friends too," the focus will shift, and the video sharing device will continue to extract the object information of the "cat" and send it to the video receiving device. The video receiving device will then perform resolution enhancement processing on the "cat" object in the currently displayed low-resolution video frame based on the "cat" object information, increasing the resolution of the "cat" object to display it as clearly as... Figure 3c The image shown is a clear picture of a "cat".
[0069] In the above embodiments, when the performance parameters of the video sharing device are in different parameter ranges, different metadata is extracted, and the corresponding metadata extraction operation is performed according to the processing capability of the video sharing device, so as to avoid problems such as metadata extraction failure due to performance parameter limitations.
[0070] In one embodiment, to facilitate understanding of the embodiments of this application by those skilled in the art, the above step S204 will be described below with reference to specific examples in the accompanying drawings. Figure 4 The methods for metadata extraction under different performance parameters are shown below:
[0071] (1) After acquiring the shared video frame, the video sharing device first applies the OCR (Optical Character Recognition) algorithm to perform character recognition operation on the shared video frame to obtain the text information of the shared video frame.
[0072] (2a) If it is determined that the performance parameters of the video sharing device are in the first parameter range with the weakest corresponding processing capability (i.e.) Figure 4 As shown in the example (under very low resource conditions), the text information is sent as metadata, along with the low-resolution video frame corresponding to the shared video frame and the text information, to the video receiving device.
[0073] (2b) If it is determined that the performance parameters of the video sharing device are not within the first parameter range, then an object detection algorithm is applied to the shared video frames to perform object detection, specifically:
[0074] i. If it is determined that the performance parameters of the video sharing device fall within the second performance parameter range, which corresponds to weaker processing power (i.e. Figure 4 As shown in the example under low resource conditions, object detection is performed based on predefined objects, and low-resolution video frames, text information, and object information of predefined objects are sent to the video receiving device.
[0075] ii. If it is determined that the performance parameters of the video sharing device are within the third performance parameter range, which corresponds to the strongest processing capability (i.e. Figure 4 As shown in the standard resource example, speech objects can be detected in real time through the speech-to-text model, and low-resolution video frames, text information, and speech object information can be sent to the video receiving device.
[0076] In one embodiment, before the step of obtaining the object information of the voice object, the method further includes: periodically detecting objects in the shared video frame to obtain object information of multiple detected objects; and storing the object information of the multiple detected objects into the memory of the video sharing device.
[0077] The steps for obtaining the object information of the speech object described above specifically include: matching the target object corresponding to the speech object from multiple detection objects; and obtaining the object information of the target object from memory as the object information of the speech object.
[0078] The object information of the detected object may include information such as object type and bounding box position.
[0079] In a practical implementation, an object detector can be run on the shared screen framework. This detector periodically checks for objects in the shared video frames, for example, running every k seconds to obtain object information such as object type and bounding box position. Figure 5 The diagram illustrates object detection on a shared screen using an object detector. The bounding box positions 502 and 504 of the two objects "dog" and "cat" in the image are detected and stored as object information in the memory of the video sharing device. Therefore, the object information of the multiple detected objects stored in memory is updated every k seconds. Later, when a voice object is detected from the user's voice, it is matched against the multiple detected objects stored in memory to determine the target object that matches the voice object. The object information of the target object is then retrieved from memory as the object information of the voice object mentioned by the user's account.
[0080] In this embodiment, shared video frames are periodically detected, and the object information of multiple detected objects is stored in the memory of the video sharing device. This allows the object information of the voice object to be directly retrieved from memory when a voice object is detected from the user's voice, thereby improving the acquisition rate of object information.
[0081] In one embodiment, the step of obtaining object information of a speech object further includes: if no target object corresponding to the speech object is matched from multiple detection objects, converting the multiple detection objects and the words corresponding to the speech object into vectors of fixed dimensions respectively; obtaining the cosine distance between the vector corresponding to the speech object and the vectors corresponding to each detection object; and detecting speech objects in video frames based on the object information of detection objects whose cosine distance to the speech object is less than a threshold, thereby obtaining the object information of the speech object.
[0082] In specific implementation, when no target object corresponding to the speech object is matched among multiple detection objects stored in memory, speech object detection can be performed using a word vector model to improve detection efficiency. Specifically, the speech object and the words corresponding to each detection object stored in memory can be converted into fixed-dimensional vectors, and the cosine distance between the vector corresponding to the speech object and the vectors corresponding to each detection object can be calculated. The cosine distances are sorted in ascending order according to their numerical values, and a target cosine distance less than a set threshold is determined from all the cosine distances. Based on the object information of the detection object corresponding to the target cosine distance, the speech object is detected in the shared video frame to obtain the object information of the speech object.
[0083] Understandably, when the cosine distance between a set of vectors is less than a set threshold, it means that these vectors are not too far apart in the vector space, and similar objects can be detected quickly. For example, "cow" and "bull" are not too far apart in the vector space. Therefore, detecting speech objects in video frames based on the object information of detected objects whose cosine distance to speech objects is less than a threshold can improve the detection efficiency of speech objects.
[0084] For example, an object detector performs a detection every k seconds. It has already detected objects A, B, C, and D, and stored their information in memory. After detecting object E from the user's voice, it matches object E with the stored objects A, B, C, and D. If no match is found, it converts the detected objects A, B, C, and D, as well as the words corresponding to the voice object E, into fixed-dimensional vectors, and calculates the cosine distance d between the vector corresponding to the voice object E and the vectors corresponding to the detected objects A, B, C, and D. AE d BE d CE d DE If, from these four cosine distances, the target cosine distance less than a set threshold is determined to be d... AE If the target cosine distance is A, then the speech object E is detected based on the object information of the detection object A, and the object information of the speech object E is obtained.
[0085] In this embodiment, when no target object corresponding to the speech object is matched among the multiple detection objects stored in memory, the multiple detection objects and the words corresponding to the speech object are converted into vectors of fixed dimensions respectively. Based on the object information of the detection objects whose cosine distance with the speech object is less than a threshold, the speech object is detected, which can improve the detection efficiency of the speech object and thus increase the rate of obtaining the object information of the speech object.
[0086] In another embodiment, such as Figure 6 As shown, a screen sharing method is also provided, which can be applied to Figure 1 Taking the video receiving device 104 as an example, the following steps are included:
[0087] Step S602: Receive the low-resolution video frame corresponding to the shared video frame sent by the video sharing device, as well as the metadata of the shared video frame; the metadata is obtained by the video sharing device based on the performance parameters of the video sharing device by performing corresponding metadata extraction operations on the shared video frame; the metadata includes the target content information of the shared video frame.
[0088] Step S604: Perform resolution enhancement processing on the target content information in the low-resolution video frame based on the metadata to obtain the processed video frame.
[0089] Step S606: Display the processed video frames.
[0090] In a specific implementation, after receiving the low-resolution video frame corresponding to the shared video frame sent by the video sharing device 102, and the metadata of the shared video frame, the video receiving device 104 can perform resolution enhancement processing on the corresponding text information and important content information in the low-resolution video frame by using the metadata containing the text information and important content information of the shared video frame, so as to improve the resolution of the text information and important content information in the processed video frame, thereby improving the clarity of the shared video frame displayed by the video receiving device.
[0091] The screen sharing method provided in this embodiment, in addition to receiving low-resolution video frames corresponding to the shared video frames sent by the video sharing device to ensure the smoothness of the shared video, also receives metadata extracted from the shared video frames by the video sharing device based on its performance parameters. This metadata allows for resolution enhancement processing of the target content information in the low-resolution video frames, improving the clarity of the shared video frames displayed by the video receiving device. Thus, during screen sharing, both a good frame rate and the quality of the shared video are guaranteed.
[0092] In one embodiment, before performing step S604 above, the method further includes: performing super-resolution processing on the low-resolution video frame to obtain the processed video frame.
[0093] Step S604 above includes: performing resolution enhancement processing on the processed video frames based on metadata.
[0094] Super-resolution (SR) technology refers to the reconstruction of a corresponding high-resolution image from an observed low-resolution image.
[0095] In practice, to reduce unnecessary resource consumption, video sharing devices may omit metadata extraction for unimportant information or information without specific clarity requirements. Therefore, the metadata of the shared video frames received by the video receiving device may not contain all the information within those frames. To ensure that the content information in the shared video frames whose metadata was not extracted can also be displayed at a relatively high resolution, the video receiving device, upon receiving a low-resolution video frame, can first perform super-resolution processing on the low-resolution video frame using a super-resolution model to reconstruct the corresponding high-resolution image as the processed video frame. Further, the metadata is used to perform resolution enhancement processing on the text information and important content information in the processed video frame, thereby ensuring accurate reproduction of the text information and important content information in the shared video frame.
[0096] In this embodiment, before performing resolution enhancement processing on low-resolution video frames using metadata, super-resolution processing is first performed on the low-resolution video frames using a super-resolution model to reconstruct the high-resolution images corresponding to the low-resolution video frames. This allows the content information in the low-resolution video frames for which metadata has not been extracted to be displayed at a relatively high resolution, further improving the quality of the shared video frames displayed by the video receiving device.
[0097] In one embodiment, the step of performing super-resolution processing on low-resolution video frames to obtain processed video frames further includes: obtaining performance parameters of the video receiving device; if the performance parameters of the video receiving device are in the fourth parameter range, then performing super-resolution processing on the video receiving device; if the performance parameters of the video receiving device are in the fifth parameter range, then sending the low-resolution video frames to the server so that the server performs super-resolution processing on the low-resolution video frames; and receiving the processed video frames returned by the server.
[0098] The fourth parameter range can be understood as the range of performance parameters with relatively strong processing capabilities.
[0099] The fifth parameter range can be understood as the range of performance parameters corresponding to weaker processing capabilities.
[0100] In practice, running the super-resolution model requires the video receiving device to have a certain processing capability. Therefore, before performing super-resolution processing on low-resolution video frames, it is necessary to determine whether the video receiving device has the capability to perform super-resolution processing based on its performance parameters. Specifically, if the performance parameters of the video receiving device are in the fourth parameter range, which corresponds to stronger processing capability, it indicates that the video receiving device can perform super-resolution processing. Therefore, the low-resolution video frames can be super-resolution processed on the video receiving device using the super-resolution model. If the performance parameters of the video receiving device are in the fifth parameter range, which corresponds to weaker processing capability, it indicates that the video receiving device has difficulty performing super-resolution processing. Therefore, the video receiving device can send the low-resolution video frames to the server, allowing the server to perform super-resolution processing on the low-resolution video frames. The video receiving device then receives the processed video frames returned by the server.
[0101] In this embodiment, based on the parameter range corresponding to the performance parameters of the video receiving device, it is determined whether the video receiving device has the ability to perform super-resolution processing. Therefore, different strategies are adopted to perform super-resolution processing on low-resolution video frames. This can avoid the problem that the performance parameters of the video receiving device are limited, resulting in the inability to perform super-resolution processing on low-resolution video frames, which would affect the clarity of the shared video frames displayed by the video receiving device.
[0102] It should be understood that, although Figure 2 , Figure 4 and Figure 6 The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order in which these steps are executed, and they can be performed in other orders. Figure 2 , Figure 4 and Figure 6 At least some of the steps in the process may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but may be executed at different times. The execution order of these steps or stages is not necessarily sequential, but may be executed in turn or alternately with other steps or at least some of the steps or stages in other steps.
[0103] In one embodiment, such as Figure 7 As shown, a screen sharing device is provided, including: an acquisition module 702, an extraction module 704, and a sending module 706, wherein:
[0104] Module 702 is used to acquire performance parameters of the video sharing device;
[0105] The extraction module 704 is used to perform corresponding metadata extraction operations on the shared video frames of the video sharing device based on performance parameters, so as to obtain the metadata in the shared video frames; the metadata contains the target content information of the shared video frames.
[0106] The sending module 706 is used to send the low-resolution video frame corresponding to the shared video frame and the metadata of the shared video frame to the video receiving device associated with the video sharing device, so that the video receiving device can perform resolution enhancement processing on the target content information corresponding to the low-resolution video frame through the metadata and display the processed video frame.
[0107] In one embodiment, the extraction module 704 is further configured to perform character recognition operation on the shared video frame when it is determined that the performance parameters of the video sharing device are within the first parameter range, to obtain the text information of the shared video frame as the metadata of the shared video frame.
[0108] In one embodiment, the extraction module 704 is further configured to perform character recognition on the shared video frame to obtain the text information of the shared video frame when the performance parameters of the video sharing device are determined to be within the second parameter range; and to perform object detection on the shared video frame according to a predefined object to obtain the object information of the predefined object, and to use the text information and the object information of the predefined object as the metadata of the shared video frame.
[0109] In one embodiment, the extraction module 704 is further configured to perform character recognition on the shared video frame to obtain the text information of the shared video frame when the performance parameters of the video sharing device are determined to be in the third parameter range; and to determine the voice object based on the voice of the user account during screen sharing, obtain the object information of the voice object, and use the text information and the object information of the voice object as the metadata of the shared video frame.
[0110] In one embodiment, the above-mentioned device further includes a detection module, which is used to periodically detect objects in the shared video frames to obtain object information of multiple detected objects; store the object information of multiple detected objects in the memory of the video sharing device; the extraction module 704 is also used to match the target object corresponding to the voice object from the multiple detected objects; and obtain the object information of the target object from the memory as the object information of the voice object.
[0111] In one embodiment, the extraction module 704 is further configured to, if no target object corresponding to the speech object is matched from multiple detection objects, convert the multiple detection objects and the words corresponding to the speech object into vectors of fixed dimensions respectively; obtain the cosine distance between the vector corresponding to the speech object and the vector corresponding to each detection object; and detect the speech object in the video frame based on the object information of the detection objects whose cosine distance to the speech object is less than a threshold, thereby obtaining the object information of the speech object.
[0112] In one embodiment, such as Figure 8 As shown, a screen sharing device is provided, including: a receiving module 802, a processing module 804, and a display module 806, wherein,
[0113] The receiving module 802 is used to receive the low-resolution video frame corresponding to the shared video frame sent by the video sharing device, as well as the metadata of the shared video frame; the metadata is obtained by the video sharing device based on the performance parameters of the video sharing device and performing corresponding metadata extraction operations on the shared video frame; the metadata includes the target content information of the shared video frame.
[0114] The processing module 804 is used to perform resolution enhancement processing on the target content information in the low-resolution video frame based on metadata, so as to obtain the processed video frame.
[0115] Display module 806 is used to display the processed video frames.
[0116] In one embodiment, the processing module 804 is further configured to perform super-resolution processing on the low-resolution video frame to obtain the processed video frame; and to perform resolution enhancement processing on the processed video frame based on metadata.
[0117] It should be noted that the screen sharing device of this application corresponds one-to-one with the screen sharing method of this application. The technical features and beneficial effects described in the embodiments of the screen sharing method described above are also applicable to the embodiments of the screen sharing device. For details, please refer to the description in the embodiments of the method of this application. It will not be repeated here.
[0118] Furthermore, each module in the aforementioned screen sharing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in the computer device in hardware form, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0119] In one embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 9As shown, the computer device includes a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a screen sharing method. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad mounted on the computer device casing, or an external keyboard, touchpad, or mouse.
[0120] Those skilled in the art will understand that Figure 9 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0121] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.
[0122] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.
[0123] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical storage, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0124] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0125] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. A screen sharing method, characterized in that, Applied to video sharing devices, the method includes: Obtain the performance parameters of the video sharing device; Based on the performance parameters, a corresponding metadata extraction operation is performed on the shared video frames of the video sharing device to obtain the metadata in the shared video frames; the metadata includes the target content information of the shared video frames; wherein, the step of performing a corresponding metadata extraction operation on the shared video frames of the video sharing device based on the performance parameters to obtain the metadata in the shared video frames includes: determining the target parameter range corresponding to the performance parameters, taking the target metadata type corresponding to the target parameter range as the metadata type to be extracted by the video sharing device, and using a metadata extraction method corresponding to the target metadata type to perform the metadata extraction operation on the shared video frames to obtain the metadata containing the target content information of the shared video frames; The low-resolution video frame corresponding to the shared video frame, along with the metadata, is sent to a video receiving device associated with the video sharing device. The video receiving device then uses the metadata to perform resolution enhancement processing on the target content information corresponding to the low-resolution video frame and displays the processed video frame.
2. The method according to claim 1, characterized in that, Based on the performance parameters, the metadata extraction operation is performed on the shared video frames of the video sharing device to obtain the metadata in the shared video frames, including: When the performance parameters of the video sharing device are determined to be within the first parameter range, character recognition is performed on the shared video frame to obtain the text information of the shared video frame, which serves as the metadata of the shared video frame.
3. The method according to claim 1, characterized in that, The step of performing corresponding metadata extraction operations on the shared video frames based on the performance parameters to obtain the metadata in the shared video frames also includes: When it is determined that the performance parameters of the video sharing device are within the second parameter range, a character recognition operation is performed on the shared video frame to obtain the text information of the shared video frame; Furthermore, based on a predefined object, object detection is performed on the shared video frame to obtain the object information of the predefined object, and the text information and the object information of the predefined object are used as the metadata of the shared video frame.
4. The method according to claim 1, characterized in that, The step of performing metadata extraction operations on the shared video frames of the video sharing device based on the performance parameters to obtain the metadata in the shared video frames further includes: When it is determined that the performance parameters of the video sharing device are within the third parameter range, character recognition is performed on the shared video frame to obtain the text information of the shared video frame; Additionally, the voice object is determined based on the user account's voice during screen sharing, the object information of the voice object is obtained, and the text information and the object information of the voice object are used as metadata of the shared video frame.
5. The method according to claim 4, characterized in that, Before obtaining the object information of the voice object, the process also includes: Periodically detect objects in the shared video frames to obtain object information for multiple detected objects; The object information of the multiple detected objects is stored in the memory of the video sharing device; The step of obtaining the object information of the voice object includes: From the plurality of detected objects, a target object corresponding to the speech object is matched; The object information of the target object is obtained from the memory and used as the object information of the voice object.
6. The method according to claim 5, characterized in that, The step of obtaining the object information of the voice object further includes: If no target object corresponding to the speech object is matched from the plurality of detection objects, then the plurality of detection objects and the words corresponding to the speech object are converted into vectors of fixed dimensions respectively. Obtain the cosine distance between the vector corresponding to the speech object and the vector corresponding to each of the detection objects; Based on the object information of the detected object whose cosine distance to the speech object is less than a threshold, the speech object is detected in the video frame to obtain the object information of the speech object.
7. A screen sharing method, characterized in that, Applied to a video receiving device, the method includes: The system receives a low-resolution video frame corresponding to a shared video frame sent by a video sharing device, along with metadata of the shared video frame. The metadata is obtained by the video sharing device performing a metadata extraction operation on the shared video frame based on its performance parameters. The metadata includes target content information of the shared video frame. Specifically, the metadata containing the target content information of the shared video frame is obtained by the video sharing device determining a target parameter range corresponding to the performance parameters, using the target metadata type corresponding to the target parameter range as the metadata type to be extracted by the video sharing device, and performing the metadata extraction operation on the shared video frame using a metadata extraction method corresponding to the target metadata type. Based on the metadata, the target content information in the low-resolution video frame is subjected to resolution enhancement processing to obtain the processed video frame. The processed video frames are displayed.
8. The method according to claim 7, characterized in that, Before performing resolution enhancement processing on the low-resolution video frame based on the metadata, the method further includes: The low-resolution video frames are subjected to super-resolution processing to obtain the processed video frames; The resolution enhancement processing of the low-resolution video frame based on the metadata includes: The processed video frames are then subjected to resolution enhancement processing based on the metadata.
9. A screen sharing device, characterized in that, Applied to video sharing devices, the device includes: The acquisition module is used to acquire the performance parameters of the video sharing device; An extraction module is configured to perform corresponding metadata extraction operations on the shared video frames of the video sharing device based on the performance parameters, to obtain metadata in the shared video frames; the metadata includes target content information of the shared video frames; wherein, the extraction module is further configured to determine the target parameter range corresponding to the performance parameters, take the target metadata type corresponding to the target parameter range as the metadata type to be extracted by the video sharing device, and use a metadata extraction method corresponding to the target metadata type to perform the metadata extraction operation on the shared video frames, to obtain the metadata containing the target content information of the shared video frames; The sending module is used to send the low-resolution video frame corresponding to the shared video frame and the metadata to the video receiving device associated with the video sharing device, so that the video receiving device can perform resolution enhancement processing on the target content information corresponding to the low-resolution video frame through the metadata and display the processed video frame.
10. A screen sharing device, characterized in that, Applied to a video receiving device, the apparatus includes: A receiving module is configured to receive a low-resolution video frame corresponding to a shared video frame sent by a video sharing device, and metadata of the shared video frame; the metadata is obtained by the video sharing device performing a corresponding metadata extraction operation on the shared video frame based on the performance parameters of the video sharing device; the metadata includes target content information of the shared video frame; wherein, the metadata containing the target content information of the shared video frame is obtained by the video sharing device determining a target parameter range corresponding to the performance parameters, taking the target metadata type corresponding to the target parameter range as the metadata type to be extracted by the video sharing device, and performing the metadata extraction operation on the shared video frame using a metadata extraction method corresponding to the target metadata type; The processing module is used to perform resolution enhancement processing on the target content information in the low-resolution video frame according to the metadata, so as to obtain the processed video frame. The display module is used to display the processed video frames.
11. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 8.
12. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Regions of interest in video frames
US20070086669A1
Image processing method and apparatus, electronic device, and computer readable storage medium
WO2021017811A1