Conference video flow control method, device and electronic equipment
By identifying and segmenting the conference videos and reducing the image quality area processing, the problem of poor smoothness of high-quality videos when the network is not good is solved, and efficient transmission and fluency of videos are achieved.
Patent Information
- Application Number
- CN202311682263.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-08
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2043-12-08
AI Technical Summary
In the case of poor network conditions, high-quality conference videos lead to poor video fluency, affecting the normal progress of the conference.
The conference video is identified through the first user equipment, divided into areas that need to reduce the image quality and areas that maintain the original image quality, and processing the areas that reduce the image quality to generate and transmit videos to ensure the integrity and fluency of the video.
While ensuring video quality, it improves transmission efficiency and fluency, reduces unnecessary traffic consumption and storage space, and improves user experience.
Smart Images

Figure CN117692680B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of data processing, and in particular to a conference video flow control method, device and electronic equipment. Background Art
[0002] A video conferencing system is a system that allows two or more individuals or groups in different locations to exchange audio, video, and file information through transmission lines and user devices, enabling instant and interactive communication to achieve the purpose of the meeting.
[0003] The two main factors that determine the effectiveness of conference video are image quality and video fluency. Currently, conference video often places excessive emphasis on high-quality image quality, often resulting in large video files, requiring higher bandwidth and a more stable network connection. However, when network conditions are poor, high-quality conference video can result in poor video fluency.
[0004] Therefore, a conference video flow control method, device and electronic equipment are urgently needed. Summary of the Invention
[0005] The present application provides a conference video flow control method, device and electronic device, which are convenient for solving the problem of poor video fluency caused by high-definition conference video when the network conditions are poor.
[0006] In a first aspect of the present application, a conference video flow control method is provided, which is applied to a first user device, and the method includes: when the network status of the second user device reaches a preset standard, obtaining a conference video, wherein the conference video is collected by a camera connected to the first user device; identifying the conference video to obtain an identification result; according to the identification result, dividing the conference video into a first area and a second area, the first area being an area where the image quality needs to be reduced according to a preset rule, and the second area being an area where the original image quality is maintained according to the preset rule; lowering the image quality of the first area to obtain a third area; splicing the third area with the second area to obtain a transmission video; and sending the transmission video to the second user device so that the second user device displays the transmission video.
[0007] Using the above technical solution, the first user device first checks the network status of the second user device. Only when the network status meets preset standards will the conference video be retrieved. Next, the first user device identifies the conference video and quickly segments it into a first region and a second region. The first region is where the image quality needs to be reduced according to preset rules, thereby reducing unnecessary data consumption and storage space usage. Based on the identification results, the first user device automatically lowers the image quality of the first region to meet network transmission requirements while maintaining the image quality of the second region. This improves transmission efficiency and smoothness while ensuring video quality. The first user device then splices the lowered-quality third region with the original-quality second region to generate the complete transmitted video. This ensures video integrity and continuity, preventing local quality changes from affecting the overall viewing experience. Finally, the first user device sends the processed transmitted video to the second user device for normal display and viewing. This effectively solves the problem of poor smoothness caused by transmitting high-quality video under poor network conditions.
[0008] Optionally, when the network status of the second user device reaches a preset standard, obtaining the conference video specifically includes: responding to a conference request sent by the second user device, the conference request including the network status; monitoring the network status of the second user device, the network status of the second user device including the network speed; if it is determined that the network speed of the second user device is lower than a preset threshold, obtaining the conference video.
[0009] By adopting the above technical solution, when the second user device sends a meeting request, the first user device responds in real time and monitors the device's network status. This ensures that the conference video can be acquired and transmitted in a timely manner, and the meeting start is not delayed due to waiting for a stable network connection. The first user device not only receives the meeting request but also monitors the network status of the second user device. This provides a comprehensive understanding of the user's network environment, allowing appropriate policy adjustments based on network conditions. The first user device will only begin acquiring the conference video when it determines that the second user device's network speed is below a preset threshold, ensuring smooth and stable video playback even in low-speed conditions. By monitoring the network status in real time and adjusting the method for acquiring the conference video based on the preset threshold, the user experience can be greatly improved. Even in poor network conditions, users can smoothly participate in the meeting without being affected by lags or delays.
[0010] Optionally, the identifying of the conference video to obtain an identification result specifically includes: performing portrait recognition on the conference video to obtain a first identification result, the first identification result including the portrait information contained in the conference video; based on the first identification result, cutting out the portrait area in the conference video to obtain a first video; performing text recognition on the first video to obtain a second identification result, the second identification result including the text information contained in the conference video; based on the second identification result, cutting out the text area in the conference video to obtain a second video; performing image recognition on the second video to obtain a third identification result, the third identification result including the image information contained in the conference video.
[0011] By employing the above technical solution, various types of information contained in conference videos can be extracted through portrait recognition, text recognition, and image recognition. This information provides important reference for subsequent video processing and analysis. Portrait recognition accurately identifies people in the video, text recognition identifies textual information in the video, and image recognition identifies the image content in the video. These recognition results can improve the accuracy of understanding the video content. Through these recognition processes, unnecessary areas in the conference video can be removed, resulting in a more concise video image. This facilitates subsequent video processing and analysis. By performing recognition processing on the conference video in a step-by-step manner, multiple recognition tasks can be performed in parallel, improving processing efficiency.
[0012] Optionally, the conference video is divided into a first area and a second area based on the recognition result, specifically including: setting the portrait area to the first area based on the first recognition result; setting the text area and the area corresponding to the image information to the second area based on the second recognition result and the third recognition result.
[0013] By employing the above technical solution, the first user device can segment the conference video into different regions, enabling targeted processing based on the characteristics of each region. Furthermore, by processing different regions of the video separately, the first user device can handle multiple tasks in parallel, improving processing efficiency. This segmentation method can be flexibly adjusted based on actual needs. The first user device can automatically segment the conference video into different regions based on different recognition results, without manual intervention, demonstrating high intelligence.
[0014] Optionally, dividing the conference video into a first area and a second area based on the recognition result specifically also includes: setting the portrait area and the area corresponding to the image information as the first area based on the first recognition result and the third recognition result; setting the text area as the second area based on the second recognition result.
[0015] By adopting the above technical solution, by combining the first recognition result and the third recognition result to determine the area corresponding to the portrait area and the image information, the people and image content in the conference video can be more accurately identified. At the same time, by determining the text area based on the second recognition result, the text information in the conference video can be more accurately identified. By dividing the conference video into more detailed areas, more detailed processing can be performed based on the characteristics of each area. Through a more detailed segmentation method, the needs of users can be better met. By setting the area corresponding to the portrait area and the image information as the first area, users can be provided with more intuitive and richer video content for more detailed processing and analysis.
[0016] Optionally, lowering the image quality of the first region to obtain the third region specifically includes: acquiring color distribution information of the first region; and performing color quantization processing on the first region using a color quantization algorithm according to the color distribution information to obtain the third region.
[0017] By adopting the above technical solution, by reducing the image quality of the first region, video bandwidth usage can be reduced, making network transmission smoother. Reducing image quality can also reduce video storage space, making video storage and backup more efficient. By obtaining color distribution information in the first region and applying a color quantization algorithm to perform color quantization processing, the video's primary color information can be preserved, ensuring that the image quality of the third region still meets certain visual requirements. The color quantization algorithm can quickly downsample the image, thereby improving processing efficiency.
[0018] Optionally, the method also includes: receiving a conference video setting request sent by the second user device, the conference setting request including a fourth area; according to the conference video setting request, increasing the image quality of the fourth area to obtain a target video; and sending the target video to the second user device so that the second user device displays the target video.
[0019] By adopting the technical solution, the first user device can meet the needs of different users in a targeted manner by receiving a conference video setting request from the second user device. The user can independently select the area where the image quality needs to be improved so as to better watch the conference video. By improving the image quality of the fourth area, the image quality of the area can be made clearer and more delicate, thereby improving the video viewing experience. The first user device only improves the image quality of the fourth area requested by the user, and does not change the image quality of other areas, thereby maintaining the integrity of the video content. The first user device can receive and process the conference video setting request of the second user device in real time, and send the processed target video to the user device in a timely manner, ensuring the real-time nature of the video transmission. The user can actively send a conference video setting request to the first user device, which enhances the interactivity between the user and the system and enables the user to participate in the meeting more actively.
[0020] In the second aspect of the present application, a conference video flow control device is provided, which includes an acquisition module and a processing module, wherein the acquisition module is used to acquire the conference video when the network status of the second user device reaches a preset standard, and the conference video is collected by the camera connected to the first user device; the processing module is used to identify the conference video and obtain an identification result; the processing module is also used to divide the conference video into a first area and a second area according to the identification result, the first area is an area where the image quality needs to be reduced according to preset rules, and the second area is an area where the original image quality is maintained according to the preset rules; the processing module is also used to lower the image quality of the first area to obtain a third area; the processing module is also used to splice the third area with the second area to obtain a transmission video; the processing module is also used to send the transmission video to the second user device so that the second user device displays the transmission video.
[0021] In a third aspect of the present application, an electronic device is provided, which includes a processor, a memory, a user interface, and a network interface. The memory is used to store instructions, the user interface and the network interface are both used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device performs the method described above.
[0022] In a fourth aspect of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores instructions. When the instructions are executed, the method described above is executed.
[0023] In summary, one or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:
[0024] 1. First, the first user device will check the network status of the second user device, and only when the network status meets the preset standard will the conference video be obtained. Next, the first user device identifies the conference video and can quickly divide it into the first area and the second area. The first area is the area where the image quality needs to be reduced according to the preset rules, which can reduce unnecessary traffic consumption and storage space usage. Based on the recognition results, the first user device will automatically lower the image quality of the first area to adapt to the network transmission requirements, while keeping the image quality of the second area unchanged. This can improve transmission efficiency and smoothness while ensuring video quality. The first user device will splice the third area with lowered image quality with the second area that maintains the original image quality to generate a complete transmission video. This can ensure the integrity and continuity of the video and avoid affecting the overall viewing experience due to changes in local image quality. Finally, the first user device will send the processed transmission video to the second user device so that the user can display and watch it normally. This makes it easy to solve the problem of poor fluency caused by transmitting high-definition video when the network conditions are poor;
[0025] 2. Through portrait recognition, text recognition, and image recognition, various types of information contained in conference videos can be extracted. This information can provide important reference for subsequent video processing and analysis. Portrait recognition can accurately identify the people in the video, text recognition can identify the text information in the video, and image recognition can identify the image content in the video. These recognition results can improve the accuracy of understanding the video content. Through these recognition processes, unnecessary areas in the conference video can be cut out to obtain a simpler video picture. This can facilitate subsequent video processing and analysis operations. By performing step-by-step recognition processing on the conference video, multiple recognition tasks can be performed in parallel, improving processing efficiency;
[0026] 3. By reducing the image quality of the first region, video bandwidth usage can be reduced, enabling smoother network transmission. Lowering image quality also reduces video storage space, making video storage and backup more efficient. By acquiring color distribution information from the first region and applying a color quantization algorithm to perform color quantization, the primary color information of the video can be preserved, ensuring that the image quality of the third region still meets certain visual requirements. The color quantization algorithm can quickly downsample the image, thereby improving processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 A flowchart of a conference video flow control method provided in an embodiment of the present application.
[0028] Figure 2 A module diagram of a conference video flow control device provided in an embodiment of the present application.
[0029] Figure 3 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application.
[0030] Explanation of the reference numerals: 21, acquisition module; 22, processing module; 31, processor; 32, communication bus; 33, user interface; 34, network interface; 35, memory. DETAILED DESCRIPTION
[0031] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments.
[0032] In the description of the embodiments of this application, words such as "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "for example" or "for instance" in the embodiments of this application should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "for example" or "for instance" is intended to present the relevant concepts in a concrete manner.
[0033] In the description of the embodiments of the present application, the term "multiple" means two or more. For example, multiple systems refer to two or more systems, and multiple screen terminals refer to two or more screen terminals. In addition, the terms "first" and "second" are used for descriptive purposes only and are not to be understood as indicating or implying relative importance or implicitly indicating the indicated technical features. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized.
[0034] A video conferencing system is a modern communication tool that, through the use of advanced electronic equipment and communication technologies, breaks geographical limitations, enabling two or more individuals or groups in different locations to interact and communicate in real time. This system not only includes high-quality video and audio transmission, but also enables data file sharing and interaction, significantly improving the efficiency and effectiveness of meetings.
[0035] Among the many factors that determine the effectiveness of conference video, video image quality and video fluency are two of the most critical. Image quality directly affects the viewing experience of conference participants, while video fluency is directly related to whether the meeting proceeds smoothly.
[0036] Current video conferencing systems often place excessive emphasis on high-quality images, resulting in large video files. This requires higher bandwidth and a more stable network connection to ensure video transmission. If the network is poor, even the highest quality video cannot guarantee smooth video playback. Any video freezes or delays can severely disrupt the smooth progress of the meeting.
[0037] In order to solve the above technical problems, this application provides a conference video flow control method, referring to Figure 1 , Figure 1 A flow chart of a conference video flow control method provided in an embodiment of the present application. The conference video flow control method is applied to a first user device and includes steps S110 to S160, which are as follows:
[0038] S110. When the network status of the second user device reaches a preset standard, obtain a conference video, where the conference video is captured by a camera connected to the first user device.
[0039] Specifically, the preset standard is a network speed standard or a transmission time standard. For example, when the network speed is poor or the transmission time is long due to network fluctuations, the first user device will actively obtain the conference video captured by the camera. The user devices include a first user device and a second user device. The first user device and the second user device are two devices for video conferencing. The first user device and the second user device are connected via a wired or wireless connection. The types of user devices include, but are not limited to: Android system devices, mobile operating system (iOS) devices developed by Apple, personal computers (PCs), World Wide Web (web) devices, virtual reality (VR) devices, augmented reality (AR) devices, and other devices. In the embodiment of the present application, the first user device and the second user device are preferably both computers.
[0040] In one possible implementation, when the network status of the second user device reaches a preset standard, obtaining the conference video specifically includes: responding to a conference request sent by the second user device, the conference request including the network status; monitoring the network status of the second user device, the network status of the second user device including the network speed; if it is determined that the network speed of the second user device is lower than a preset threshold, obtaining the conference video.
[0041] Specifically, this paragraph means that when the second user device sends a meeting request containing network status information, the first user device will respond to this request. The first user device will continuously monitor the network status of the second user device, including the network upload and download speeds, etc. If the first user device detects that the network speed of the second user device is lower than a preset threshold, it will start to obtain the conference video. The preset threshold is the user-defined setting corresponding to the first user device. For example, suppose there are two users A and B who work on a conference video system. User B's network speed is greater than 2Mbps. User A, as the first user device, detects that the network status of user B's second user device meets the preset conditions, and then starts to control the camera to obtain the conference video, which is then transmitted to user B.
[0042] S120: Identify the conference video and obtain a recognition result.
[0043] Specifically, the first user device recognizes the conference video by analyzing and processing it using computer vision and artificial intelligence technologies to extract key information or identify specific content. This recognition process can include different tasks such as portrait recognition, text recognition, and voice recognition.
[0044] In one possible implementation, a conference video is identified to obtain an identification result, specifically including: performing portrait recognition on the conference video to obtain a first recognition result, the first recognition result including the portrait information contained in the conference video; based on the first recognition result, cutting out the portrait area in the conference video to obtain a first video; performing text recognition on the first video to obtain a second recognition result, the second recognition result including the text information contained in the conference video; based on the second recognition result, cutting out the text area in the conference video to obtain a second video; performing image recognition on the second video to obtain a third recognition result, the third recognition result including the image information contained in the conference video.
[0045] Specifically, the first user device analyzes the conference video using portrait recognition technology to detect and identify all portraits appearing in the video. This portrait information constitutes the first recognition result. Based on the first recognition result, the portrait area in the conference video is removed, leaving a video clip that does not contain portraits, thereby obtaining the first video. Next, the first user device performs text recognition processing on the first video to extract all text information appearing in the video. This text information constitutes the second recognition result. Based on the second recognition result, the text area in the first video is removed, leaving a video clip that does not contain text, which is the second video. Finally, image recognition processing is performed on the second video to extract all image information appearing in the video. This image information constitutes the third recognition result.
[0046] The first user device determines the portrait area, text area, and image information corresponding areas based on the outline, thereby ensuring the integrity of the portrait and text. Secondly, portrait information includes, for example, the current speaker and meeting participants; text information includes, for example, the text display of the meeting agenda and meeting design; and image information refers to other information in the conference video besides portrait information and text information, which may include blank areas or graphic display areas. This multi-level recognition processing makes it easy to quickly extract key figures, text, and image information from the conference video, facilitating subsequent image quality analysis.
[0047] S130 , based on the recognition result, dividing the conference video into a first area and a second area, where the first area is an area where the image quality needs to be reduced according to a preset rule, and the second area is an area where the original image quality needs to be maintained according to the preset rule.
[0048] Specifically, the first user device divides the conference video into two areas: a first area and a second area, based on the results of previously performed recognition tasks such as portrait recognition, text recognition, and image recognition. The first area is where the image quality needs to be reduced according to preset rules. This area can be understood as an area of lesser importance, where the image quality has less impact on the overall meeting. In other words, the image quality in this area will be reduced to ensure smooth video playback. The second area is where the image quality needs to be maintained according to preset rules. This is because the image quality in this area affects the overall reception of the meeting. Considering the effectiveness of the meeting, the image quality in this area will remain unchanged to preserve the original quality of the video. For example, if the first area corresponds to the portrait area of the current speaker, and the focus of the meeting is on viewing the digital information in the text portion of the second area, the image quality of the first area will be reduced, while the image quality of the second area will remain unchanged. This adapts to network transmission requirements, reduces network congestion and latency, and improves the viewing experience for remote participants.
[0049] In a possible implementation, the conference video is divided into a first area and a second area based on the recognition results, specifically including: setting the portrait area to the first area based on the first recognition result; setting the text area and the area corresponding to the image information to the second area based on the second recognition result and the third recognition result.
[0050] Specifically, the above process is a method for determining the first area and the second area provided by an embodiment of the present application. For example, suppose that in a company conference video, the participants are signing an important contract. At this time, the importance of the participant's portrait quality is lower than the importance of the text quality and the image quality, wherein the image at this time is a contract photo provided by the participant, etc. Therefore, the first area is the participant's portrait area, and the second area is the area corresponding to the text information and the contract photo. This segmentation process helps to improve the accuracy and efficiency of subsequent data analysis. Thus, the first user device can perform targeted processing based on the characteristics of each area by dividing the conference video into different areas. Moreover, the first user device processes different areas of the video separately, and can process multiple tasks in parallel to improve processing efficiency. This segmentation method can be flexibly adjusted according to actual needs. The first user device can automatically divide the conference video into different areas based on different recognition results without manual intervention, and has strong intelligence.
[0051] In a possible implementation, the conference video is divided into a first area and a second area based on the recognition results, specifically including: setting the portrait area and the area corresponding to the image information as the first area based on the first recognition result and the third recognition result; and setting the text area as the second area based on the second recognition result.
[0052] Specifically, the above process is another method for determining the first area and the second area provided by an embodiment of the present application. For example, assuming that in a company conference video, the participants are signing an important contract. At this time, the conference video only includes the portrait area and the text area of the participants, and the image area is a blank page. The importance of the participant's portrait quality and the importance of the image area are both lower than the importance of the text quality. Therefore, the first area is the image area corresponding to the participant's portrait area and the blank page, and the second area is the text area. Thus, by dividing the conference video into more detailed areas, more detailed processing can be performed according to the characteristics of each area. A more detailed segmentation method can better meet the needs of users. By setting the area corresponding to the portrait area and the image information as the first area, users can be provided with more intuitive and richer video content for more detailed processing and analysis.
[0053] S140: Lower the image quality of the first area to obtain a third area.
[0054] Specifically, after the first user device determines the first region and the second region, the image quality of the first region is lowered to obtain the third region, while the image quality of the second region remains unchanged. Methods for lowering image quality include, but are not limited to: reducing resolution, such as reducing the resolution of a video to a lower level, for example, reducing a 1080p video in the first region to 720p or lower; reducing frame rate, such as reducing the frame rate of a video to a lower level, for example, reducing a 30 fps video in the first region to 15 fps or lower; compression encoding, such as reducing the size of a video file using video compression encoding technology, thereby reducing image quality, such as using a compression encoder such as H.264 or H.265 to compress the segmented video corresponding to the first region to reduce image quality; adding noise, such as adding noise to the segmented video corresponding to the first region to degrade the image quality of the video, which can blur details of the video, thereby reducing image quality; and blurring, such as blurring details of the segmented video corresponding to the first region using blurring technology, thereby reducing image quality, such as using Gaussian blurring or mean blurring to process the video.
[0055] In a possible implementation, lowering the image quality of the first region to obtain the third region specifically includes: obtaining color distribution information of the first region; and performing color quantization processing on the first region using a color quantization algorithm according to the color distribution information to obtain the third region.
[0056] Specifically, the color quantization algorithm determines which pixels to color quantize based on the position, size, and color distribution of the portrait area. This can be achieved by analyzing the image's color histogram. The color quantization algorithm can also select pixels for color quantization based on characteristic information within the portrait area. For example, the algorithm can identify facial features and perform color quantization based on these features to enhance facial detail. By analyzing this data, the color quantization algorithm can effectively downsample the image quality of the portrait area in conference videos, reducing video quality while still maintaining a good viewing experience. By reducing the image quality of the first area, video bandwidth usage can be reduced, ensuring smoother network transmission. Reducing image quality also reduces video storage space, making video storage and backup more efficient. By obtaining color distribution information for the first area and applying the color quantization algorithm to the image, the primary color information of the video can be preserved, ensuring that the image quality of the third area still meets certain visual requirements. The color quantization algorithm can quickly downsample the image, thereby improving processing efficiency.
[0057] S150: Splice the third area with the second area to obtain a transmission video.
[0058] S160: Send the transmission video to the second user equipment, so that the second user equipment displays the transmission video.
[0059] Specifically, the above process involves the first user device splicing the third area with the second area to form a transmission video, and then sending the transmission video to the second user device so that the second user device can display the transmission video. The second user device can display the video normally and clearly see the portrait information after color quantization processing and the text information that maintains the original image quality. Therefore, through this processing method, the first user device can segment the conference content and transmit the processed video to the second user device for display. This splicing process not only improves the transmission efficiency of the video, but also facilitates the viewing experience of the second user device.
[0060] In this way, the first user device checks the network status of the second user device and only retrieves the conference video when the network status meets preset standards. Next, the first user device identifies the conference video and quickly segments it into a first region and a second region. The first region is where the image quality needs to be reduced according to preset rules, thereby reducing unnecessary data consumption and storage space usage. Based on the identification results, the first user device automatically lowers the image quality of the first region to meet network transmission requirements while maintaining the image quality of the second region. This improves transmission efficiency and smoothness while ensuring video quality. The first user device then splices the lowered-quality third region with the original-quality second region to generate the complete transmitted video. This ensures video integrity and continuity, preventing local quality changes from affecting the overall viewing experience. Finally, the first user device sends the processed transmitted video to the second user device for normal display and viewing. This effectively solves the problem of poor smoothness caused by transmitting high-quality video under poor network conditions.
[0061] In one possible implementation, a conference video setting request sent by a second user device is received, where the conference setting request includes a fourth area; based on the conference video setting request, the image quality of the fourth area is increased to obtain a target video; and the target video is sent to the second user device so that the second user device displays the target video.
[0062] Specifically, the above process involves the first user device receiving a conference video setup request sent by the second user device, increasing the image quality of the fourth area based on the request, and then sending the target video to the second user device for display. This solves the problem of the user corresponding to the first user device being unable to clearly view some relatively important video areas when watching the conference video. This not only ensures the real-time nature of video transmission, but also allows the user to actively send a conference video setup request to the first user device, which enhances the interactivity between the user and the conference video system and enables the user to participate more actively in the meeting.
[0063] This application also provides a conference video flow control device, referring to Figure 2 , Figure 2 A module diagram of a conference video flow control device provided in an embodiment of the present application. The conference video flow control device is a first user device, and the first user device includes an acquisition module 21 and a processing module 22, wherein the acquisition module 21 is used to acquire a conference video when the network status of the second user device reaches a preset standard, and the conference video is collected by a camera connected to the first user device; the processing module 22 is used to identify the conference video and obtain an identification result; the processing module 22 is also used to divide the conference video into a first area and a second area according to the identification result, the first area being an area where the image quality needs to be reduced according to a preset rule, and the second area being an area where the original image quality is maintained according to the preset rule; the processing module 22 is also used to lower the image quality of the first area to obtain a third area; the processing module 22 is also used to splice the third area with the second area to obtain a transmission video; the processing module 22 is also used to send the transmission video to the second user device so that the second user device displays the transmission video.
[0064] In one possible implementation, the acquisition module 21 acquires the conference video when the network status of the second user device reaches a preset standard, specifically including: the acquisition module 21 responds to the conference request sent by the second user device, and the conference request includes the network status; the processing module 22 monitors the network status of the second user device, and the network status of the second user device includes the network speed; if the processing module 22 determines that the network speed of the second user device is lower than the preset threshold, the conference video is acquired.
[0065] In one possible implementation, the processing module 22 identifies the conference video and obtains an identification result, which specifically includes: the processing module 22 performs portrait recognition on the conference video to obtain a first recognition result, and the first recognition result includes the portrait information contained in the conference video; the processing module 22 cuts out the portrait area in the conference video based on the first recognition result to obtain a first video; the processing module 22 performs text recognition on the first video to obtain a second recognition result, and the second recognition result includes the text information contained in the conference video; the processing module 22 cuts out the text area in the conference video based on the second recognition result to obtain a second video; the processing module 22 performs image recognition on the second video to obtain a third recognition result, and the third recognition result includes the image information contained in the conference video.
[0066] In a possible implementation, the processing module 22 divides the conference video into a first area and a second area based on the recognition results, specifically including: the processing module 22 sets the portrait area to the first area based on the first recognition result; the processing module 22 sets the text area and the area corresponding to the image information to the second area based on the second recognition result and the third recognition result.
[0067] In a possible implementation, the processing module 22 divides the conference video into a first area and a second area based on the recognition results, specifically including: the processing module 22 sets the portrait area and the area corresponding to the image information as the first area based on the first recognition result and the third recognition result; the processing module 22 sets the text area as the second area based on the second recognition result.
[0068] In a possible implementation, the processing module 22 lowers the image quality of the first area to obtain the third area, specifically including: the acquisition module 21 obtains the color distribution information of the first area; the processing module 22 uses a color quantization algorithm to perform color quantization on the first area according to the color distribution information to obtain the third area.
[0069] In one possible implementation, the acquisition module 21 receives a conference video setting request sent by the second user device, where the conference setting request includes a fourth area; the processing module 22 increases the image quality of the fourth area according to the conference video setting request to obtain a target video; the processing module 22 sends the target video to the second user device so that the second user device displays the target video.
[0070] It should be noted that the above embodiments provide devices that implement their functions using only the division of the above functional modules as examples. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the device and method embodiments provided in the above embodiments are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.
[0071] This application also provides an electronic device, referring to Figure 3 , Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device may include: at least one processor 31, at least one network interface 34, a user interface 33, a memory 35, and at least one communication bus 32.
[0072] The communication bus 32 is used to realize the connection and communication between these components.
[0073] The user interface 33 may include a display screen (Display) and a camera (Camera). Optionally, the user interface 33 may also include a standard wired interface and a wireless interface.
[0074] The network interface 34 may optionally include a standard wired interface or a wireless interface (such as a WI-FI interface).
[0075] The processor 31 may include one or more processing cores. Using various interfaces and circuits, the processor 31 connects to various components within the server. It executes instructions, programs, code sets, or instruction sets stored in the memory 35, as well as accesses data stored in the memory 35, to perform various server functions and process data. Optionally, the processor 31 may be implemented using at least one of the following hardware forms: a digital signal processing (DSP), a field-programmable gate array (FPGA), or a programmable logic array (PLA). The processor 31 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU primarily processes the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing content displayed on the display; and the modem handles wireless communications. It is understood that the modem may also be implemented as a separate chip, rather than integrated into the processor 31.
[0076] Among them, the memory 35 may include a random access memory (RAM) or a read-only memory (Read-Only Memory). Optionally, the memory 35 includes a non-transitory computer-readable storage medium. The memory 35 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 35 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store data involved in the above-mentioned various method embodiments, etc. The memory 35 may also be optionally at least one storage device located away from the aforementioned processor 31. As Figure 3 As shown, the memory 35 as a computer storage medium may include an operating system, a network communication module, a user interface module and an application program of a conference video stream control method.
[0077] exist Figure 3 In the electronic device shown, the user interface 33 is mainly used to provide an input interface for the user and obtain data input by the user; and the processor 31 can be used to call an application program storing a conference video flow control method in the memory 35. When executed by one or more processors, the electronic device executes one or more methods in the above embodiments.
[0078] It should be noted that for the aforementioned method embodiments, for simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required for this application.
[0079] The present application also provides a computer-readable storage medium storing instructions, which, when executed by one or more processors, enable an electronic device to execute one or more of the methods described in the above embodiments.
[0080] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0081] In the several embodiments provided in this application, it should be understood that the disclosed devices can be implemented in other ways. For example, the device embodiments described above are merely schematic, such as the division of units, which is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some service interface, and the indirect coupling or communication connection of devices or units can be electrical or other forms.
[0082] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0083] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0084] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of this application, or the portion that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the method of this application. The aforementioned memory includes various media that can store program code, such as USB flash drives, mobile hard drives, magnetic disks, or optical disks.
[0085] The above is only an exemplary embodiment of the present disclosure and cannot be used to limit the scope of the present disclosure. That is, any equivalent changes and modifications made according to the teachings of the present disclosure are still within the scope of the present disclosure. After considering the disclosure of the specification and the truth of practice, those skilled in the art will easily think of other embodiments of the present disclosure. This application is intended to cover any variation, use or adaptive change of the present disclosure, which follows the general principles of the present disclosure and includes common knowledge or customary technical means in the art that are not recorded in the present disclosure. The description and examples are to be regarded as exemplary only, and the scope and spirit of the present disclosure are defined by the claims.
Claims
1. A conference video flow control method, characterized in that: Applied to a first user equipment, the method includes: When the network status of the second user device meets a preset standard, obtaining a conference video, where the conference video is captured by a camera connected to the first user device; Identify the conference video and obtain an identification result; According to the recognition result, the conference video is divided into a first area and a second area, the first area is an area where the image quality needs to be reduced according to a preset rule, and the second area is an area where the original image quality is maintained according to the preset rule; Lowering the image quality of the first area to obtain a third area; splicing the third area with the second area to obtain a transmission video; Sending the transmission video to a second user device, so that the second user device displays the transmission video; When the network condition of the second user's device meets the preset standard, the conference video is obtained, specifically including: responding to a conference request sent by the second user equipment, wherein the conference request includes a network status; monitoring a network status of the second user equipment, where the network status of the second user equipment includes a network speed; If it is determined that the network speed of the second user equipment is lower than a preset threshold, obtaining the conference video; The step of lowering the image quality of the first region to obtain the third region specifically includes: Acquiring color distribution information of the first area; performing color quantization processing on the first region using a color quantization algorithm according to the color distribution information to obtain the third region; receiving a conference video setting request sent by the second user equipment, where the conference setting request includes a fourth area; According to the conference video setting request, the image quality of the fourth area is increased to obtain a target video; sending the target video to the second user device so that the second user device displays the target video; The identifying of the conference video to obtain an identification result specifically includes: Performing portrait recognition on the conference video to obtain a first recognition result, where the first recognition result includes portrait information contained in the conference video; According to the first recognition result, cutting out the portrait area in the conference video to obtain a first video; Performing text recognition on the first video to obtain a second recognition result, where the second recognition result includes text information contained in the conference video; According to the second recognition result, cutting out the text area in the conference video to obtain a second video; Performing image recognition on the second video to obtain a third recognition result, where the third recognition result includes image information contained in the conference video; The step of dividing the conference video into a first area and a second area according to the recognition result specifically includes: According to the first recognition result, the portrait area is set as the first area; According to the second recognition result and the third recognition result, setting the text area and the area corresponding to the image information as the second area; The step of dividing the conference video into a first area and a second area according to the recognition result further includes: According to the first recognition result and the third recognition result, the portrait area and the area corresponding to the image information are set as a first area; According to the second recognition result, the text area is set as the second area.
2. A conference video flow control device, characterized in that: The device executes the method according to claim 1, and the conference video flow control device includes an acquisition module (21) and a processing module (22), wherein: The acquisition module (21) is used to acquire a conference video when the network status of the second user device reaches a preset standard, wherein the conference video is collected by a camera connected to the first user device; The processing module (22) is used to identify the conference video and obtain an identification result; The processing module (22) is further configured to divide the conference video into a first area and a second area according to the recognition result, wherein the first area is an area where the image quality needs to be reduced according to a preset rule, and the second area is an area where the original image quality is maintained according to the preset rule; The processing module (22) is further configured to lower the image quality of the first region to obtain a third region; The processing module (22) is further configured to splice the third area with the second area to obtain a transmission video; The processing module (22) is further configured to send the transmission video to a second user device so that the second user device displays the transmission video.
3. An electronic device, characterized in that: The electronic device comprises a processor (31), a memory (35), a user interface (33) and a network interface (34), wherein the memory (35) is used to store instructions, the user interface (33) and the network interface (34) are both used to communicate with other devices, and the processor (31) is used to execute the instructions stored in the memory (35) so that the electronic device executes the method according to claim 1.
4. A computer-readable storage medium, characterized in that The computer-readable storage medium stores instructions, which, when executed, perform the method according to claim 1 .
Citation Information
Patent Citations
Regional variable-resolution online teaching method, system and equipment and storage medium
CN113099254A