Video conference picture recognition coding scheme and system and computer equipment

By segmenting, text recognition and coding video images in the video conferencing system, the problem of sensitive information leakage in video conferencing is solved, and the privacy data protection and efficiency of video conferencing are improved.

CN120070228APending Publication Date: 2025-05-30GUANGZHOU MAILING INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311636561.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-30
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing video conferencing systems are prone to leaking company sensitive or confidential information during shooting conferences.

Method used

By obtaining the video images captured by the video conferencing terminal, dividing them into several block images, and text recognition is performed on each block image to determine whether sensitive information exists. If it exists, code it; if it does not exist, it remains as it is. The coded block image is spliced ​​with the default block image, the coded video image is generated and sent to the second video conference terminal.

Benefits of technology

It effectively protects private data, avoids the leakage of sensitive data, and improves the security and efficiency of video conferencing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070228A_ABST
    Figure CN120070228A_ABST
Patent Text Reader

Abstract

The invention relates to a video conference picture recognition coding method and system and computer equipment. The video conference picture identifying and coding method comprises the following steps: acquiring a video image shot by a first video conference terminal, and segmenting the video image into a plurality of block images; performing character recognition on each block image, judging whether sensitive information exists in the block image or not, if the sensitive information exists in the block image, coding the block image and acquiring a coded block image, and if the sensitive information does not exist in the block image, keeping the block image unchanged and acquiring a default block image; and splicing the coding block image and the default block image to obtain a coding video image, and sending the coding video image to a second video conference terminal. The video conference picture identifying and coding method and system and the computer equipment have the advantages that privacy data security is guaranteed, and sensitive data leakage is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of video coding, and in particular to a method, a system and a computer device for identifying and coding video conference images. Background Art

[0002] With the development of science and technology, video communication is widely used in people's lives. To improve the convenience of communication, video conferencing is often used for communication between companies and customers. Using video conferencing, participants can hear the voices of other meeting rooms, see the images, movements and expressions of on-site participants in other meeting rooms, and can also send electronic presentation content. However, in existing video conferences, the images captured by most cameras are clear and distinct. During the meeting, the document content in the company meeting room can also be clearly captured, which is likely to cause the leakage of the company's sensitive information or confidential information. Summary of the Invention

[0003] Based on this, the purpose of the present application is to provide a method, a system and a computer device for identifying and coding video conference images, which have the advantages of ensuring the security of privacy data and avoiding the leakage of sensitive data.

[0004] A method for identifying and coding video conference images includes the following steps:

[0005] Obtain a video image captured by a first video conference terminal, and divide the video image into a plurality of block images;

[0006] Perform text recognition on each of the block images to determine whether there is sensitive information in the block image. If there is sensitive information in the block image, then code the block image to obtain a coded block image. If there is no sensitive information in the block image, then the block image remains unchanged to obtain a default block image;

[0007] Stitch the coded block image and the default block image to obtain a coded video image, and send the coded video image to a second video conference terminal.

[0008] A system for identifying and coding video conference images includes:

[0009] Obtain a video image captured by a first video conference terminal, and divide the video image into a plurality of block images;

[0010] Perform text recognition on each of the block images to determine whether there is sensitive information in the block image. If there is sensitive information in the block image, then code the block image to obtain a coded block image. If there is no sensitive information in the block image, then the block image remains unchanged to obtain a default block image;

[0011] Stitch the coded block image and the default block image to obtain a coded video image, and send the coded video image to a second video conferencing terminal.

[0012] A computer device includes: a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the video conferencing screen recognition and coding method as described above are implemented.

[0013] A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the steps of the video conferencing screen recognition and coding method as described above are implemented.

[0014] The video conferencing screen recognition and coding method described in this application obtains a video image captured by a first video conferencing terminal, divides the video image into several block images, performs character recognition on each block image, determines whether sensitive information exists in the block image, codes the block image with sensitive information to obtain a coded block image, keeps the block image without sensitive information unchanged to obtain a default block image, and finally stitches the coded block image and the default block image to obtain a coded video image, and sends the coded video image to a second video conferencing terminal. In the video conferencing screen recognition and coding method in this application, by identifying and coding sensitive information in the video conference, the privacy data security is ensured and the leakage of sensitive data is avoided. When performing character recognition, sensitive information judgment, and coding on the video image, after the video image is divided into several block images, multiple processing threads can be called to process each image block in parallel at the same time, without waiting in line, and the amount of data processed by each image block is much less than the amount of data processed by an entire video image, so the processing efficiency is improved, which is particularly suitable for application scenarios with extremely high requirements for immediacy such as video conferences, and improves the efficiency of video conferences.

[0015] For better understanding and implementation, the present application will be described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 It is a flowchart of the steps of a video conferencing screen recognition and coding method in an embodiment of the present application;

[0017] Figure 2 It is a schematic diagram of a video conferencing screen recognition and coding method in an embodiment of the present application;

[0018] Figure 3 It is a flowchart of the steps of determining whether sensitive information exists in the block image in an embodiment of the present application;

[0019] Figure 4The flowchart of the steps for determining whether there is sensitive information in the block image in another embodiment of the present application;

[0020] Figure 5 The structural schematic diagram of the video conference screen recognition and coding system in the embodiment of the present application;

[0021] Figure 6 The schematic diagram of the computer device for the video conference screen recognition and coding method in the embodiment of the present application. Detailed implementation manners

[0022] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all the implementation manners consistent with the present application. On the contrary, they are only examples of the devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0023] The terms used in the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. The singular forms of "a", "the", and "said" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0024] It should be understood that although the terms first, second, third, etc. may be used in the present application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" / "when" as used herein may be interpreted as "when" or "while" or "in response to determining".

[0025] Please refer to Figure 1 and Figure 2 , Figure 1 The flowchart of the steps of a video conference screen recognition and coding method in the embodiment of the present application, Figure 2 The schematic diagram of a video conference screen recognition and coding method in the embodiment of the present application.

[0026] A video conference screen recognition and coding method includes the following steps:

[0027] S101, obtaining the video image captured by the first video conference terminal, and dividing the video image into a plurality of block images;

[0028] S102. Perform character recognition on each of the block images, and determine whether there is sensitive information in the block image. If there is sensitive information in the block image, then perform blurring on the block image to obtain a blurred block image. If there is no sensitive information in the block image, then the block image remains unchanged, and a default block image is obtained.

[0029] S103. Stitch the blurred block image and the default block image together to obtain a blurred video image, and send the blurred video image to a second video conference terminal.

[0030] In the video conference screen recognition and blurring method described in this application, by obtaining the video image captured by a first video conference terminal, dividing the video image into a plurality of block images, performing character recognition on each block image, determining whether there is sensitive information in the block image, performing blurring on the block image with sensitive information to obtain a blurred block image, keeping the block image without sensitive information unchanged to obtain a default block image, finally stitching the blurred block image and the default block image together to obtain a blurred video image, and sending the blurred video image to a second video conference terminal. In the video conference screen recognition and blurring method in this application, by recognizing and blurring sensitive information in the video conference, it ensures the security of privacy data, avoids the leakage of sensitive data, and improves the efficiency of the video conference.

[0031] For step S101, obtain the video image captured by the first video conference terminal, and divide the video image into a plurality of block images;

[0032] Among them, the first video conference terminal is a device for initiating or participating in a video conference. The first video conference terminal includes a camera device, a display device, and an audio device. Among them, the audio device includes a recording device and a speaker. In one embodiment, the first video conference terminal includes a smart phone, a mobile tablet, and / or a computer.

[0033] The video image is obtained by being captured by the first video conference terminal. In this embodiment, after participating in or initiating a video conference through the first video conference terminal, the video image is captured through the camera device of the first video conference terminal. The video image has a preset resolution. In one embodiment, the preset resolution of the video image can be set by the user.

[0034] The block image is a small area image generated by dividing the video image. In one embodiment, the video image is evenly divided into a plurality of block images of the same size, or the video image can also be divided into a plurality of block images of different sizes.

[0035] In one embodiment, the step of dividing the video image into a plurality of block images includes:

[0036] Obtain the image resolution of the video image and obtain a preset resolution;

[0037] Calculate the division ratio of the length and width of the video image according to the image resolution and the preset resolution;

[0038] Divide the video image into a plurality of block images according to the division ratio.

[0039] Wherein, the image resolution is the resolution of the video image. In this embodiment, the image resolution is obtained by recognizing the video image. In other embodiments, the image resolution can also be obtained according to the resolution parameters of the imaging device of the first video conferencing terminal. The preset resolution is a pre-set resolution. In this embodiment, the video image is divided into a plurality of block images according to the preset resolution.

[0040] The division ratio is a division scheme for the length and width of the video image calculated according to the image resolution and the preset resolution. The division ratio includes a length division ratio and a width division ratio. Calculating the division ratio of the length and width of the video image means calculating the ratios of the length and width in the image resolution to the length and / or width in the preset resolution respectively. In one embodiment, if the video image resolution is 1200*800 and the preset resolution is 100*200, then the length division ratio is 1200 / 100 = 12, and the width division ratio is 800 / 200 = 4; in another embodiment, if the video image resolution is 1200*800 and the preset resolution is 100*200, then the length division ratio is 1200 / 200 = 6, and the width division ratio is 800 / 100 = 8.

[0041] According to the segmentation ratio, the video image is segmented into a number of block images, that is, the length and width of the video image are segmented according to the length segmentation ratio and the width segmentation ratio in the segmentation ratio respectively to obtain the block images. In one embodiment, if the resolution of the video image is 1200*800, the preset resolution is 100*200, the length segmentation ratio is 1200 / 100 = 12, and the width segmentation ratio is 800 / 200 = 4. The length of the video image is segmented into 12 parts according to the length segmentation ratio, and the width of the video image is segmented into 4 parts according to the width segmentation ratio, and finally 48 block images with a resolution of 100*200 are obtained. In another embodiment, if the resolution of the video image is 1200*800 and the preset resolution is 100*200, then the length segmentation ratio is 1200 / 200 = 6, and the width segmentation ratio is 800 / 100 = 8. The length of the video image is segmented into 6 parts according to the length segmentation ratio, and the width of the video image is segmented into 8 parts according to the width segmentation ratio, and finally 48 block images with a resolution of 100*200 are obtained.

[0042] In other embodiments, the video image can also be divided according to a fixed ratio to obtain a number of the block images. For example, a ratio such as one-fourth, one-fifth or one-tenth of the length and width of the video image is used as the length and width of the block image to divide the video image.

[0043] In this embodiment, by obtaining the video image captured by the first video conferencing terminal and segmenting the video image into a number of block images, it is convenient to quickly and accurately identify each part of the video image, and improve the accuracy of video image coding.

[0044] For step S102, perform text recognition on each block image to determine whether there is sensitive information in the block image. If there is sensitive information in the block image, then code the block image to obtain a coded block image. If there is no sensitive information in the block image, then the block image remains unchanged to obtain a default block image;

[0045] Among them, the text is recognized by performing text recognition on the block image. In one embodiment, the text recognition algorithm includes Optical Character Recognition (OCR). In other embodiments, the text recognition algorithm further includes end-to-end text recognition based on deep learning.

[0046] The sensitive information is information that needs to be kept confidential and affects the company's data security. The sensitive data includes some company data and discussion content that appears in the office area and is not suitable for public disclosure. In one embodiment, the sensitive data includes at least one of an account number, a password, and a verification code. The sensitive information can be preset by the user, and the content of the sensitive information can be changed by logging in to the management account. In this embodiment, by performing optical character recognition on each of the block images, it is determined whether there is sensitive information in the block image, and the sensitive information is blurred to protect privacy and avoid leakage of sensitive data.

[0047] Please refer to Figure 3 , Figure 3 which is a flowchart of the steps for determining whether there is sensitive information in the block image in an embodiment of the present application. In one embodiment, performing optical character recognition on each of the block images to determine whether there is sensitive information in the block image includes the following steps:

[0048] S201, obtaining the number of texts in the block image by performing optical character recognition on the block image;

[0049] S202, obtaining the sensitive data level of the block image according to the number of texts;

[0050] S203, if the sensitive data level is lower than the preset sensitive data level, it is determined that the block image does not include sensitive information, and if the sensitive data level is higher than the preset sensitive data level, it is determined that the block image includes sensitive information.

[0051] For steps S201 to S203, wherein, the number of texts is the number of words obtained by performing optical character recognition on the block image to obtain the block image.

[0052] The sensitive data level is the sensitive data level of the block image obtained according to the number of texts. In one embodiment, the number of texts is used as the sensitive data level of the block image. For example, by performing optical character recognition on the block image, the number of texts obtained in the block image is 200, then the sensitive data level is 200. The preset sensitive data level is the sensitive data level set in advance and is used to determine whether the block image includes sensitive information. For example, it is set that the preset sensitive data level is 150. When the sensitive data level of the block image is 200, it is determined that the block image includes sensitive information.

[0053] In other embodiments, the sensitive data level may also be different levels divided according to the number of texts appearing in the block image. For example, when the number of texts is set to be greater than 1 and less than 50, the sensitive data level is 1; when the number of texts is greater than 50 and less than 100, the sensitive data level is 2; when the number of texts is greater than 100, the sensitive data level is 3. When the sensitive data level is 3, the block image includes sensitive information.

[0054] In this embodiment, by performing text recognition on the block image, obtaining the sensitive data level of the block image and comparing it with the preset sensitive data level, it is determined whether the block image includes sensitive information. The solution in this embodiment does not require precise recognition of each text in the block image, improving the efficiency of text recognition of the block image.

[0055] The solution in this embodiment is applied to a video conference, and can quickly recognize and censor the block images containing a large amount of text content in the video image, without requiring precise recognition of the text content, improving the efficiency of image transmission in the video conference and ensuring the real-time nature of the video conference image transmission.

[0056] Please refer to Figure 4 , Figure 4 which is the flowchart of the steps for determining whether there is sensitive information in the block image in another embodiment of the present application. In another embodiment, the steps of performing text recognition on each block image and determining whether there is sensitive information in the block image include:

[0057] S301, by performing text recognition on the block image, obtaining the text information in the block image, and comparing the text information with a preset sensitive information text library;

[0058] S302, if the text information matches the sensitive information in the sensitive information text library, it is determined that there is sensitive information in the block image, otherwise, it is determined that there is no sensitive information in the block image.

[0059] For steps S301 to S302, the preset sensitive information text library is a database that is set in advance and includes a number of sensitive information. In one embodiment, the preset text library includes account numbers, passwords, and / or verification codes.

[0060] The text information matches the sensitive information in the sensitive information text library, that is, there is information in the text information that is the same as the content in the sensitive information text library. For example, if there is a set of passwords "123456" in the preset sensitive information text library and the same password "123456" also exists in the text information, it is determined that the text information matches the preset text library and there is sensitive information in the block image.

[0061] In this embodiment, by performing optical character recognition on the block image, text information in the block image is obtained, and the text information is compared with the preset text library. When the text information matches the preset text library, the block image includes sensitive information. The solution in this embodiment can accurately recognize the text information and the preset text library, determine the sensitive information in the block image, and improve the accuracy of blurring.

[0062] In one embodiment, for the block image determined not to include sensitive information in step S203, the solution in steps S301 to S302 can be further used to further recognize the text information in the block image, determine whether the block image includes sensitive keywords, and further determine whether the block image includes sensitive information.

[0063] In one embodiment, if there is sensitive information in the block image, the steps of blurring the block image to obtain a blurred block image include:

[0064] If there is sensitive information in the block image, obtain the sensitive keywords in the block image, perform blurring processing on the sensitive keywords, and obtain a blurred block image.

[0065] Among them, the sensitive keywords include the content in the block image that matches the sensitive information text library. In this embodiment, when there is sensitive information in the block image, the content that matches the preset text library in the text information is used as the keyword, and only the keyword is blurred to obtain a blurred block image, which improves the accuracy of video conference screen recognition and blurring.

[0066] In one embodiment, the blurring processing of the keyword includes the following steps:

[0067] Perform blurring processing or text replacement processing on the keyword to obtain a blurred block image.

[0068] Among them, the blurring processing blurs the keyword, and the text replacement replaces the keyword with preset text content. In this embodiment, the keyword is blurred by means of blurring processing or text replacement to obtain a blurred block image.

[0069] In another embodiment, if there is sensitive information in the block image, the steps of blurring the block image to obtain a blurred block image include:

[0070] If there is sensitive information in the block image, replace the block image with a preset image to obtain the blurred block image.

[0071] Among them, the preset image is an image that is set in advance to have the same size as the block image, and the coded block image is the image obtained after coding the block image. In this embodiment, when there is sensitive information in the block image, the preset image is directly used to replace the block image, and the preset image is used as the coded block image, which improves the efficiency of image coding and reduces the complexity of image processing.

[0072] It should be noted that the image coding scheme described in this application includes the scheme described in the above embodiment, but is not limited to the scheme described in the above embodiment. Other image coding schemes that can be thought of by those skilled in the art are also applicable to this application. For example, the block image is coded by means of mosaic processing, image covering processing, etc.

[0073] For step S103, the coded block image and the default block image are spliced to obtain a coded video image, and the coded video image is sent to the second video conference terminal.

[0074] Among them, the coded video image is an image generated by coding the sensitive information in the video image, and the second video conference terminal is a terminal that participates in or initiates a video conference. In this embodiment, the first video conference terminal is communicatively connected to the second video conference terminal, and the first video conference terminal sends the coded video image to the second video conference terminal. In other embodiments, there may be multiple second video conference terminals, and the first video conference terminal simultaneously sends the coded video image to the multiple video conference terminals.

[0075] In this embodiment, when all the text in the block image is recognized, the coded block image and the default block image are spliced again according to the format of the block image to obtain a coded video image, and finally the coded video image is sent to the second video conference terminal.

[0076] In one embodiment, the video conference screen recognition and coding method further includes the following steps:

[0077] Send the video image to the first video conference terminal.

[0078] Among them, the video image is the image captured by the first video conference terminal. In this embodiment, the video image is directly sent to the first video conference terminal and displayed on the display device of the first video conference terminal. Only the image sent to the second video conference terminal is a coded video image, so as to avoid displaying the coded video image on the first video conference terminal.

[0079] In other embodiments, the coded video image may also be sent to the first video conferencing terminal for display.

[0080] For the video conferencing screen recognition and coding method described in this application, by obtaining the video image captured by the first video conferencing terminal, dividing the video image into several block images, performing character recognition on each block image, determining whether there is sensitive information in the block image, and coding the block image with sensitive information to obtain a coded block image, keeping the block image without sensitive information unchanged to obtain a default block image, and finally splicing the coded block image and the default block image to obtain a coded video image, and sending the coded video image to the second video conferencing terminal.

[0081] For the video conferencing screen recognition and coding method in this application, when performing character recognition, sensitive information judgment, and coding on the video image, after the video image is divided into several block images, multiple processing threads can be called to simultaneously perform parallel processing on each image block without waiting in line, and the amount of data processed for each image block is much less than the amount of data processed for an entire video image, improving the processing efficiency. It is particularly suitable for application scenarios such as video conferencing with extremely high requirements for immediacy, ensuring the security of private data, avoiding the leakage of sensitive data, and eliminating the need to tidy up the office or work station before the meeting to prevent the leakage of sensitive data, thus improving the efficiency of video conferencing.

[0082] Please refer to Figure 5 , Figure 5 , which is a schematic structural diagram of the video conferencing screen recognition and coding system described in the embodiments of this application. This application also discloses a video conferencing screen recognition and coding system, including:

[0083] An image acquisition module 11, configured to acquire the video image captured by the first video conferencing terminal and divide the video image into several block images;

[0084] An image recognition and coding module 12, configured to perform character recognition on each of the block images, determine whether there is sensitive information in the block image, if there is sensitive information in the block image, code the block image to obtain a coded block image, and if there is no sensitive information in the block image, keep the block image unchanged to obtain a default block image;

[0085] An image splicing module 13, configured to splice the coded block image and the default block image to obtain a coded video image, and send the coded video image to the second video conferencing terminal.

[0086] It should be noted that when the video conference screen recognition and coding system provided in the above embodiments executes the video conference screen recognition and coding method, only the division of the above functional modules is used for illustration. In actual applications, the above functions can be assigned to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. The video conference screen recognition and coding system provided in the above embodiments is used to execute the video conference screen recognition and coding method described in the above embodiments. Its operation method and principle are the same as those of the video conference screen recognition and coding method described above. That is, the video conference screen recognition and coding system provided in the above embodiments and the video conference screen recognition and coding method belong to the same concept. The implementation process is detailed in the above method embodiments and will not be repeated here.

[0087] Please refer to Figure 6 , Figure 6 , which is a schematic diagram of a computer device for the video conference screen recognition and coding method in an embodiment of the present application. As Figure 6 shown, the computer device 21 includes: a control device 211, a memory 212, and a computer program 213 stored in the memory 212 and executable on the control device 211, for example: a video conference screen recognition and coding program; when the control device 211 executes the computer program 213, the video conference screen recognition and coding method described in the above embodiments can be implemented.

[0088] Among them, the control device 211 includes a processor, and the processor may include one or more processing cores. The processor uses various interfaces and lines to connect various parts within the computer device 21, and by running or executing instructions, programs, code sets, or instruction sets stored in the memory 212, and by calling data in the memory 212, it executes various functions of the computer device 21 and processes data. Optionally, the processor may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor may integrate a combination of one or more of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. Among them, the CPU mainly processes the operating system, user interface, and application programs, etc.; the GPU is responsible for rendering and drawing the content required to be displayed on the touch display screen; the modem is used to process wireless communication. It can be understood that the above modem may not be integrated into the processor and may be implemented separately by a single chip.

[0089] Among them, the memory 212 may include a Random Access Memory (RAM), or may also include a Read-Only Memory. Optionally, the memory 212 includes a non-transitory computer-readable storage medium. The memory 212 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 212 may include a program storage area and a data storage area. Among them, the program storage area can store instructions for implementing an operating system, instructions for at least one function (such as touch instructions, etc.), instructions for implementing the above various method embodiments, etc.; the data storage area can store the data involved in the above various method embodiments. Optionally, the memory 212 may also be at least one storage device located far from the aforementioned processor.

[0090] The embodiment of the present application also provides a readable storage medium. The computer-readable storage medium can store multiple instructions. These instructions are suitable for being loaded and executed by a control device to perform the method steps of the above embodiment. The specific execution process can refer to the specific description of the above embodiment and will not be elaborated here.

[0091] The above embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the present application.

Claims

1. A method for identifying and censoring video conference images, characterized in that, it includes the following steps: Obtain the video image captured by the first video conference terminal, and divide the video image into several block images; Perform text recognition on each of the block images to determine whether there is sensitive information in the block image. If there is sensitive information in the block image, censor the block image to obtain a censored block image. If there is no sensitive information in the block image, the block image remains unchanged to obtain a default block image; Stitch the censored block image and the default block image to obtain a censored video image, and send the censored video image to the second video conference terminal.

2. The method for identifying and censoring video conference images according to claim 1, characterized in that, Performing text recognition on each of the block images to determine whether there is sensitive information in the block image includes the following steps: Obtain the number of texts in the block image by performing text recognition on the block image; Obtain the sensitive data level of the block image according to the number of texts; If the sensitive data level is lower than the preset sensitive data level, it is determined that the block image does not include sensitive information. If the sensitive data level is higher than the preset sensitive data level, it is determined that the block image includes sensitive information.

3. The method for identifying and censoring video conference images according to claim 1 or 2, characterized in that, Performing text recognition on each of the block images to determine whether there is sensitive information in the block image further includes the following steps: Obtain the text information in the block image by performing text recognition on the block image, and compare the text information with a preset sensitive information text library; If the text information matches the sensitive information in the sensitive information text library, it is determined that there is sensitive information in the block image. Otherwise, it is determined that there is no sensitive information in the block image.

4. The method for identifying and censoring video conference images according to claim 3, characterized in that, If there is sensitive information in the block image, censoring the block image to obtain a censored block image includes the following steps: If there is sensitive information in the block image, obtain the sensitive keywords in the block image, and perform censoring processing on the sensitive keywords to obtain a censored block image, where the sensitive keywords include the content in the block image that matches the sensitive information text library.

5. The method for identifying and censoring video conference images according to claim 4, characterized in that, Performing censoring processing on the keywords includes the following steps: Perform blurring processing or text replacement processing on the keywords to obtain a censored block image.

6. The method for identifying and censoring video conference images according to claim 1, characterized in that, If there is sensitive information in the block image, censoring the block image to obtain a censored block image includes the following steps: If there is sensitive information in the block image, replace the block image with a preset image to obtain the censored block image.

7. The video conference screen recognition and coding method according to claim 1, characterized in that, dividing the video image into a plurality of block images, comprising the following steps: obtaining the image resolution of the video image and obtaining a preset resolution; calculating the segmentation ratio of the length and width of the video image according to the image resolution and the preset resolution; dividing the video image into a plurality of block images according to the segmentation ratio.

8. A video conference screen recognition and coding system, characterized in that, comprising: an image acquisition module, configured to acquire a video image captured by a first video conference terminal and divide the video image into a plurality of block images; an image recognition and coding module, configured to perform character recognition on each of the block images, determine whether there is sensitive information in the block image, if there is sensitive information in the block image, then code the block image to obtain a coded block image, if there is no sensitive information in the block image, then the block image remains unchanged, and obtain a default block image; an image splicing module, configured to splice the coded block image and the default block image to obtain a coded video image, and send the coded video image to a second video conference terminal.

9. A computer device, comprising: a processor, a memory, and a computer program stored in the memory and executable on the processor, characterized in that when the processor executes the computer program, the steps of the video conference screen recognition and coding method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: when the computer program is executed by a processor, the steps of the video conference screen recognition and coding method according to any one of claims 1 to 7 are implemented.