Image processing method, electronic equipment, conference system and storage medium

By grading the importance of image content and coding with different code rates, the problem of image resolution and clarity reduction in multi-camera meeting scenes is solved, improving the user experience.

CN120358356APending Publication Date: 2025-07-22BOE TECHNOLOGY GROUP CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510713433.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

In multi-camera meeting scenarios, as the number of cameras increases, the main processor's encoding and decoding capabilities are insufficient, resulting in a decrease in image resolution and clarity, affecting the user experience.

Method used

By grading the importance of the image content, using different encoding and decoding schemes, high-code rate encoding is used for important areas, and low-code rate encoding is used for non-important areas, ensuring the resolution and clarity of important content.

Benefits of technology

With limited encoding and decoding resources, the resolution and clarity of important display content are improved and the user experience is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120358356A_ABST
    Figure CN120358356A_ABST
Patent Text Reader

Abstract

The invention provides an image processing method, electronic equipment, a conference system and a storage medium. The image processing method comprises the following steps: acquiring current audio data and a current frame image acquired by an audio and video acquisition unit in a target space; a first position of a current spokesman in a target space is acquired according to the current audio data, target identification is performed on the current frame image, a target identification result comprises a target attribute and a target position, and the target attribute comprises a person; determining a second position of the current spokesman in the current frame image according to the first position and the target recognition result, and setting different importance levels for different areas of the current frame image according to the target recognition result and the second position, wherein the importance level of an area including the current spokesman in the current frame image is greater than that of an area not including the current spokesman; and according to the importance levels of the different areas of the current frame image, coding the different areas according to different code rates, wherein the code rates are in positive correlation with the importance levels.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing technologies. More specifically, it relates to an image processing method, an electronic device, a conference system, and a storage medium. Background Art

[0002] With the continuous innovation of display technologies and the continuous expansion of application scenarios, display products are gradually developing towards diversification, complexity, and functional integration. In a conference scenario, conference display devices are no longer limited to simple display functions, but integrate multiple functions such as camera image recognition and shooting.

[0003] Generally, in a conference scenario, a conference display device needs to access multiple cameras. The video data collected by these cameras needs to be encoded and decoded in a main processor and finally presented on a conference display screen. However, limited by the fixed encoding and decoding capabilities of the main processor, when the number of accessed cameras increases and the content to be displayed on the screen simultaneously increases, the resolution and clarity of the screen display content will be significantly reduced, seriously affecting the user's visual experience and conference effect. Summary of the Invention

[0004] The purpose of the present disclosure is to provide an image processing method, an electronic device, a conference system, and a storage medium to solve at least one of the above technical problems.

[0005] To achieve the above purpose, the present disclosure adopts the following technical solutions:

[0006] The first aspect of the present disclosure provides an image processing method, including the following steps:

[0007] Obtain current audio data and a current frame image collected by an audio-video acquisition unit in a target space;

[0008] Obtain a first position of a current speaker in the target space according to the current audio data, and perform object recognition on the current frame image. The object recognition result includes object attributes and object positions, and the object attributes include people;

[0009] Determine a second position of the current speaker in the current frame image according to the first position and the object recognition result, and set different importance levels for different regions of the current frame image according to the object recognition result and the second position, where the importance level of the region including the current speaker in the current frame image is greater than the importance level of the region not including the current speaker;

[0010] Encode different regions of the current frame image at different bitrates according to the importance levels of different regions of the current frame image, and the bitrate is positively correlated with the importance level.

[0011] Optionally, the step of encoding different regions of the current frame image at different bitrates according to the importance levels of different regions of the current frame image includes:

[0012] Determine whether the area ratio of the non-important regions in the current frame image is greater than a first preset ratio threshold, where the non-important regions are regions with an importance level less than a first preset level threshold;

[0013] If the determination result is greater than the first preset ratio threshold, use the key frame encoding method to encode different regions of the current frame image at different bitrates according to the importance level;

[0014] If the determination result is less than or equal to the first preset ratio threshold, determine whether the ratio of the difference region between the current frame image and the previous frame image to the total area of the current frame image is greater than a second preset ratio threshold;

[0015] If the determination result is greater than the second preset ratio threshold, use the key frame encoding method to encode different regions of the current frame image at different bitrates according to the importance level;

[0016] If the determination result is less than or equal to the second preset ratio threshold, use the predicted frame encoding method to encode different regions of the current frame image at different bitrates according to the importance level.

[0017] Optionally, before the step of determining whether the area ratio of the non-important regions in the current frame image is greater than the first preset ratio threshold, the following steps are further included:

[0018] Determine whether the number of image frames between the current frame image and the target frame image is greater than a preset frame number threshold, where the target frame image is the previous image encoded using the key frame encoding method;

[0019] If the determination result is yes, use the key frame encoding method to encode different regions of the current frame image at different bitrates according to the importance level;

[0020] If the determination result is no, perform the step of determining whether the area ratio of the non-important regions in the current frame image is greater than the first preset ratio threshold.

[0021] Optionally, the audio and video acquisition unit includes a camera, the current frame image is the image currently captured by the camera, and the step of setting different importance levels for different regions of the current frame image according to the target recognition result and the second position includes:

[0022] Set the importance level of the image content at the second position of the current frame image to the highest level;

[0023] Set an importance level for other objects other than the current speaker in the target recognition result according to the relationship between the other objects and the current speaker, and the importance level of the other objects is less than the highest level;

[0024] Divide the current frame image into multiple sub-regions according to a preset grid, and the importance level of any sub-region is determined by the importance level of the objects in the sub-region;

[0025] The step of encoding different regions of the current frame image at different bitrates according to the importance levels of different regions of the current frame image includes: encoding different sub-regions of the current frame image at different bitrates according to the importance levels of different sub-regions of the current frame image.

[0026] Optionally, the object attribute further includes an object, the highest level is level one, and the step of setting an importance level for other objects other than the current speaker in the target recognition result according to the relationship between the other objects and the current speaker includes:

[0027] Set the importance level of the people other than the current speaker in the current frame image to level two;

[0028] Set the importance level of the objects having an interaction relationship with the current speaker in the current frame image to level two;

[0029] Set the importance level of the objects having no interaction relationship with the current speaker and having an interaction relationship with people other than the current speaker in the current frame image to level three;

[0030] Set the importance level of the objects having no interaction relationship with any person in the current frame image to level four.

[0031] Optionally, the interaction relationship is determined by the contact relationship between the object and the person and the object type. The object type includes a first type of object and a second type of object. The object being a first type of object and having a contact relationship with the person indicates that the object has an interaction relationship with the person. The object being a first type of object and having no contact relationship with the person indicates that the object has no interaction relationship with the person. The object being a second type of object indicates that the object has no interaction relationship with any person.

[0032] Optionally, the audio-video acquisition unit includes multiple cameras, and the current frame image is formed by splicing multiple sub-images currently captured by the multiple cameras. The step of setting different importance levels for different regions of the current frame image according to the target recognition result and the second position includes:

[0033] Set the importance level of the first sub-image in the current frame image to level one, and the first sub-image is the sub-image containing the current speaker;

[0034] Obtain the voice - sounding time of each second - level sub - image in the current frame image, and set an importance level for each second - level sub - image according to the voice - sounding time. The voice - sounding time is positively correlated with the importance level. For any second - level sub - image, the voice - sounding time of this second - level sub - image represents the cumulative voice - sounding time of the person in this second - level sub - image within a preset time window. The second - level sub - image is a sub - image other than the first - level sub - image in the current frame image;

[0035] The step of encoding different regions of the current frame image at different bitrates according to the importance levels of different regions of the current frame image includes: encoding different sub - images of the current frame image at different bitrates according to the importance levels of different sub - images.

[0036] Optionally, the step of encoding different sub - images of the current frame image at different bitrates according to the importance levels of different sub - images includes:

[0037] For any sub - image, determine whether the sub - image is a third - level sub - image. The third - level sub - image is a sub - image whose proportion of the importance level is greater than a preset proportion threshold;

[0038] If the judgment result is that the sub - image is a third - level sub - image, determine whether the proportion of the area of the non - important region in this sub - image is greater than a first preset proportion threshold. The non - important region is a region whose importance level is less than a first preset level threshold;

[0039] If the judgment result is greater than the first preset proportion threshold, use the key - frame encoding method to encode different regions of this sub - image at different bitrates according to the importance level;

[0040] If the judgment result is less than or equal to the first preset proportion threshold, determine whether the ratio of the difference region between this sub - image and the previous - frame sub - image to the total area of this sub - image is greater than a second preset proportion threshold;

[0041] If the determination result is greater than the second preset proportion threshold, use the key - frame encoding method to encode different regions of this sub - image at different bitrates according to the importance levels of different regions in this sub - image;

[0042] If the judgment result is less than or equal to the second preset proportion threshold, use the predictive - frame encoding method to encode different regions of this sub - image at different bitrates according to the importance levels of different regions in this sub - image;

[0043] If the judgment result is that the sub - image is not a third - level sub - image, use the key - frame encoding method or the predictive - frame encoding method to determine the bitrate of this sub - image according to the importance level of this sub - image and use this bitrate for encoding.

[0044] Optionally, when using the key-frame coding method or the predictive-frame coding method, the steps of determining the bit rate of the sub-image according to the importance level of the sub-image and performing coding using the bit rate include:

[0045] Determine whether the number of sub-images between the sub-image and the target-frame sub-image is greater than a preset number-of-frames threshold, where the target-frame sub-image is the previous sub-image captured by the same camera that was encoded using the key-frame coding method;

[0046] If the determination result is yes, use the key-frame coding method to determine the bit rate of the sub-image according to the importance level of the sub-image and perform coding using the bit rate;

[0047] If the determination result is no, use the predictive-frame coding method to determine the bit rate of the sub-image according to the importance level of the sub-image and perform coding using the bit rate.

[0048] Optionally, after encoding different sub-images at different bit rates according to the importance levels of different sub-images of the current-frame image, the following steps are further included:

[0049] Decode the encoded sub-images;

[0050] Display the decoded first sub-image in the main display area;

[0051] Display the decoded second sub-image in the secondary display area. For any second sub-image, the higher the importance level of the second sub-image, the larger the area occupied by the second sub-image in the secondary display area.

[0052] Optionally, after the step of displaying the decoded first sub-image in the main display area, the following steps are further included:

[0053] Detect the duration after the current speaker stops speaking;

[0054] Determine whether the duration is less than a first time threshold;

[0055] When the determination result is yes, control the main display area to continuously display the first sub-image;

[0056] When the determination result is no, control the display content of the main display area to switch to the sub-image including the latest current speaker.

[0057] Optionally, the main display area includes multiple sub-display areas, and the image processing method further includes:

[0058] Display the sub-images corresponding to the latest multiple current speakers in the multiple sub-display areas respectively.

[0059] Optionally, the audio-video acquisition unit includes a plurality of microphones, and the positions of the plurality of microphones are different in the two-dimensional plane of the target space. The step of obtaining the first position of the current speaker in the target space according to the current audio data includes:

[0060] Obtain the distances between the microphones in the two-dimensional plane and the time differences of the sounds detected by the microphones;

[0061] Determine the first position of the current speaker in the two-dimensional plane according to the distances and the time differences.

[0062] A second aspect of the present disclosure provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the above-mentioned image processing method are implemented.

[0063] A third aspect of the present disclosure provides a conference system, including an audio-video acquisition unit, an image processing unit, and a display device. The audio-video acquisition unit is used to acquire audio data and images in a target space. The image processing unit is used to implement the steps of the above-mentioned image processing method. The display device is used to display the images acquired by the audio-video acquisition unit.

[0064] A fourth aspect of the present disclosure provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of the above-mentioned image processing method are implemented.

[0065] A fifth aspect of the present disclosure provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the above-mentioned image processing method are implemented.

[0066] The beneficial effects of the present disclosure are as follows:

[0067] The image processing method according to the embodiments of the present disclosure, when obtaining the current audio data and the current frame image, first locates the current speaker by using the collected current audio data to determine the first position of the current speaker in the target space, then performs object recognition on the current frame image to obtain each object included in the current frame image and the target position of each object, and then determines the second position of the current speaker in the current frame image according to each target position and the first position, and sets different importance levels for different regions of the current frame image, where the importance level of the region including the current speaker in the current frame image is greater than that of the region not including the current speaker, and finally encodes different regions at different bitrates according to the importance levels of different regions of the current frame image, and the bitrate is positively correlated with the importance level. In this way, the importance levels of different regions of the current frame image are set according to the image content, and different bitrates are allocated to different regions according to the importance levels, that is, different sizes of encoding and decoding resources are adaptively allocated to different display contents. In this way, the resolution and clarity of important display contents in the image can be ensured under the condition of limited encoding and decoding resources, the visibility of the video during the meeting is improved, and the user experience is good. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] The following further describes in detail the specific embodiments of the present disclosure with reference to the drawings.

[0069] Figure 1 is a flowchart of the image processing method provided by the embodiments of the present disclosure;

[0070] Figure 2 is a schematic diagram of the principle of obtaining the first position of the current speaker in the target space according to the current audio data;

[0071] Figure 3 is a schematic diagram of dividing the current frame image into multiple sub-regions when the audio-visual acquisition unit includes one camera;

[0072] Figure 4 is a schematic diagram of the current frame image when the audio-visual acquisition unit includes multiple cameras;

[0073] Figure 5 is a flowchart of encoding different regions at different bitrates according to the importance levels of different regions of the current frame image;

[0074] Figure 6 is a flowchart of encoding different regions at different bitrates according to the importance levels of different regions of the current frame image when the video acquisition unit includes multiple cameras;

[0075] Figure 7 is a schematic diagram of the display area being divided into a main display area and a secondary display area;

[0076] Figure 8 Schematic diagram of multiple sub - images displayed in the main display area in a multi - person discussion scenario;

[0077] Figure 9 Block diagram of an embodiment of a conference system provided by an embodiment of the present disclosure;

[0078] Figure 10 Block diagram of another embodiment of a conference system provided by an embodiment of the present disclosure. Detailed implementation manners

[0079] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present disclosure. Obviously, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present disclosure without creative efforts shall fall within the scope of protection of the present disclosure.

[0080] Unless otherwise defined, the technical terms or scientific terms used in the present disclosure shall have the ordinary meanings understood by those of ordinary skill in the art to which the present disclosure pertains. The terms "first", "second", and similar terms used in the present disclosure do not denote any order, quantity, or importance, but are only used to distinguish different components. Similarly, terms such as "a", "an", or "the" do not denote a quantity limitation, but mean that there is at least one. The terms "including" or "comprising" and similar terms mean that the elements or items appearing before this term cover the elements or items listed after this term and their equivalents, without excluding other elements or items. The terms "connected" or "coupled" and similar terms are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms "upper", "lower", "left", "right", etc. are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.

[0081] To better understand the technical solutions of the image processing method, electronic device, conference system, and storage medium of the present disclosure, the design concept of the present disclosure will be briefly introduced first.

[0082] In a conference scenario, the core processor not only undertakes video encoding and decoding functions, but also integrates various functions such as data processing and transmission communication. Whether single-camera acquisition or multi-camera acquisition is adopted, the image content captured by the camera includes unimportant information such as background and non-speakers. In the single-camera acquisition scenario, if video encoding and decoding are performed on all image content without discrimination, it will greatly increase the data processing volume of the core processor. This not only leads to an increase in device power consumption and temperature, but also may interfere with the normal operation of other data processing functions and affect the overall performance of the device. In the multi-camera acquisition scenario, as the number of cameras increases, the encoding data of multiple cameras is superimposed, which easily exceeds the upper limit of the encoding and decoding capabilities of the core processor. In related technologies, to ensure that images can be normally displayed, the image resolution is usually reduced. This processing method will reduce the resolution of all displayed content, including important displayed content, seriously affecting the clear presentation and viewing effect of information, and the user experience is poor.

[0083] To solve the above technical problems, the embodiments of the present disclosure provide an image processing method, and the core points mainly include two points: (1) perform importance grading on image content, and adopt different encoding and decoding schemes for image content with different importance levels; (2) for the multi-camera acquisition scenario, design different display methods according to the importance level of image content.

[0084] Please refer to Figure 1 , Figure 1 for the flowchart of the image processing method provided by the embodiments of the present disclosure. As Figure 1 shown, it includes the following steps:

[0085] Step S101, obtain the current audio data and the current frame image collected by the audio-video acquisition unit in the target space.

[0086] In the embodiments of the present disclosure, the target space may be a conference room provided with an audio-video acquisition unit. Specifically, in implementation, the target space may include one conference room, or may include multiple conference rooms at different locations. For example, one of the multiple conference rooms is used as the main venue, and the others are used as branch venues, and an audio-video acquisition unit is provided in each conference room.

[0087] Among them, the audio-video acquisition unit includes an audio acquisition unit and a video acquisition unit.

[0088] The audio acquisition unit includes multiple microphones, and the multiple microphones are arranged at different positions in the conference room for collecting audio data. The above-mentioned current audio data represents the audio data collected by the multiple microphones at the current moment.

[0089] The video acquisition unit includes one or more cameras for acquiring images of the meeting room. The above-mentioned current frame image represents a frame of image acquired by the camera at the current moment, and multiple consecutive frames of images acquired by the camera constitute a video stream. Specifically, when the video acquisition unit includes multiple cameras, each camera is usually distributed at different positions or angles in the meeting room, and each camera is used to acquire images of a certain local area or a certain perspective of the meeting room.

[0090] Step S102, obtain the first position of the current speaker in the target space according to the current audio data, and perform object recognition on the current frame image. The object recognition result includes object attributes and object positions, and the object attributes include people.

[0091] Optionally, the audio-video acquisition unit includes multiple microphones, and the positions of the multiple microphones are different in the two-dimensional plane of the target space. The step of obtaining the first position of the current speaker in the target space according to the current audio data includes: (11) obtaining the distances between the microphones in the two-dimensional plane and the time differences of the sounds detected by the microphones; (12) determining the first position of the current speaker in the two-dimensional plane according to the distances and time differences.

[0092] Exemplarily, please refer to Figure 2 , Figure 2 FIG. is a schematic diagram of the principle of obtaining the first position of the current speaker in the target space according to the current audio data. In the embodiments of the present disclosure, the two-dimensional plane of the target space can be understood as the x-y horizontal plane of the meeting room, where the audio acquisition unit includes 3 microphones, denoted as M1, M2, and M3 respectively. The 3 microphones are equidistantly placed along the x-axis direction in the meeting room, and the distance between adjacent two microphones is d. Assume that the distance between the sound source S and the microphone M1 is r1, the distance between the sound source S and the microphone M2 is r2, the distance between the sound source S and the microphone M3 is r3, and the angle between the line connecting the sound source S and the microphone M1 and the x-axis is ɑ. Assume that the propagation speed of sound is c, then the time difference t 12 for the sound source S to reach the microphone M1 and the microphone M2 is:

[0093] t 12 =(r1 - r2) / c (Formula 1)

[0094] The time difference t 13 for the sound source S to reach the microphone M1 and the microphone M3 is:

[0095] t 13 =(r1 - r3) / c (Formula 2)

[0096] According to the relationship between the sides and angles of a triangle, it can be obtained that

[0097]

[0098] In Formulas 1, 2, 3, and 4, since d, t 12 , t 13 , and c are all known variables, therefore, according to the above four formulas, the values of r1, r2, r3, and ɑ can be obtained, and thus the coordinate values of the sound source S in the x-y two-dimensional plane coordinate system can be determined. Since the sound source S also represents the current speaker, therefore, according to the time differences t 12 , t 13 in the current audio data, combined with the fixed spacing d and the sound propagation speed c, the position of the current speaker in the target space can be determined, and this position is also the first position.

[0099] Among them, the targets in the image include people. The target recognition result includes target attributes and target positions. The target attributes are used to represent the target type. For example, when the target is a person, the target attribute is also a person. The target position is used to represent the position coordinates of the target in the current frame of the image, that is, the position coordinates of the target in the image coordinate system.

[0100] Step S103, determine the second position of the current speaker in the current frame of the image according to the first position and the target recognition result, and set different importance levels for different regions of the current frame of the image according to the target recognition result and the second position, where the importance level of the region including the current speaker in the current frame of the image is greater than the importance level of the region not including the current speaker.

[0101] Among them, the first position of the current speaker in the target space is represented by the position coordinates of the current speaker in the conference room coordinate system. Here, the conference room coordinate system can also be understood as the coordinate system where the two-dimensional plane of the conference room is located. The target position in the target recognition result represents the position of the target in the image, that is, the position in the camera image coordinate system. For any determined conference room, the image coordinate system of the camera in it and the conference room coordinate system are both known and unchanged. Therefore, a mapping relationship between the image coordinate system and the conference room coordinate system can be established in advance. After the first position is determined, this mapping relationship can be used to map the first position to the image coordinate system, that is, map the position of the current speaker in the conference room coordinate system to the position of the current speaker in the image coordinate system, that is, the second position. By comparing the mapped second position with each target position in the target recognition result, it can be determined which target in the image is the current speaker.

[0102] During the meeting, the image of the speaker is very important display content. Therefore, for any frame of image, it is necessary to first ensure the clarity and resolution of the image in the area where the speaker is located in this frame, while the importance of the image content of other participants in non-speaking states or objects in the meeting room is secondary. In the embodiments of the present disclosure, different importance levels are set for different regions of the current frame image according to the target recognition result and the second position. The simplest implementation method is to set the importance level of the second position where the current speaker is located to the highest level, and set the importance levels of other regions to lower levels. For example, the highest level is level one, and the lower levels are level two, level three, etc.

[0103] Step S104, encode different regions at different bitrates according to the importance levels of different regions of the current frame image, and the bitrate is positively correlated with the importance level.

[0104] Among them, the bitrate refers to the amount of binary data contained in the video data stream per unit time, usually measured in bits per second (bps), such as kbps, Mbps, etc. It directly reflects the size of the data volume after video compression and is a key parameter affecting video quality and file size. The higher the bitrate, the smaller the compression ratio of the video, the larger the amount of data that needs to be stored or transmitted per unit time, the larger the video file size, the higher the picture quality, and the clearer the picture quality; conversely, the lower the bitrate, the smaller the file size, and a low bitrate may cause the picture to be distorted or blurred.

[0105] In the embodiments of the present disclosure, for any region in the current frame image, if the importance level of this region is higher, the display content of this region needs to have higher resolution and clarity. Therefore, a higher bitrate is required for encoding. And the lower the importance level, it means that the display content of this region is less important. Therefore, lower resolution and clarity can be adopted, and at this time, a lower bitrate can be used for encoding, that is, the importance level is positively correlated with the bitrate.

[0106] Compared with related technologies, in the image processing method according to the embodiments of the present disclosure, when the current audio data and the current frame image are obtained, first, the current speaker is located by using the collected current audio data to determine the first position of the current speaker in the target space, then the current frame image is subjected to target recognition to obtain each target included in the current frame image and the target position of each target, and then, according to each target position and the first position, the second position of the current speaker in the current frame image is determined, and different importance levels are set for different regions of the current frame image, wherein the importance level of the region including the current speaker in the current frame image is greater than the importance level of the region not including the current speaker, and finally, different regions are encoded at different bit rates according to the importance levels of different regions of the current frame image, and the bit rate is positively correlated with the importance level. In this way, the importance levels of different regions of the current frame image are set according to the image content, and different bit rates are allocated to different regions according to the importance levels, that is, different sizes of encoding and decoding resources are adaptively allocated to different display contents. In this way, the resolution and clarity of important display contents in the image can be ensured when the encoding and decoding resources are limited, the visibility of the video during the meeting is improved, and the user experience is good.

[0107] In a possible implementation manner, the audio-video acquisition unit includes a camera, and the current frame image is the image currently captured by the camera. The steps of setting different importance levels for different regions of the current frame image according to the target recognition result and the second position include (21) to (24), which are specifically as follows:

[0108] (21) Set the importance level of the image content at the second position of the current frame image to the highest level. Exemplarily, the highest level is level one.

[0109] (22) Set the importance level for other targets other than the current speaker according to the relationship between other targets other than the current speaker and the current speaker in the target recognition result. The importance level of other targets is less than the highest level. Among them, setting the importance level for other targets is also setting the importance level for the image data at the position where other targets are located. Exemplarily, the importance level of other targets can be set to level two, level three, level four, etc.

[0110] (23) Divide the current frame image into multiple sub-regions according to a preset grid, and the importance level of any sub-region is determined by the importance level of the targets in the sub-region.

[0111] Among them, the preset grid can be understood as the unit size. When a frame of image is divided into multiple sub-regions by using the preset grid, the number of sub-regions is usually of the order of two digits, such as more than a dozen or dozens. Exemplarily, please refer to Figure 3, assume that the video acquisition unit only includes a camera C1, and the resolution of the image captured by the camera C1 is 1920 (columns) * 1080 (rows). If the size of the preset grid is 240 * 180, then the current frame image can be divided into 8 * 6 = 48 sub-regions Q at this time.

[0112] Among them, the importance level of any sub-region is determined by the importance level of the target within that region. Exemplarily, if the importance level of the target within a certain sub-region is level one, then the importance level of this sub-region is level one; if the importance level of the target within a certain sub-region is level two, then the importance level of this sub-region is level two; if a certain sub-region includes two targets, and the importance levels of the two targets are level two and level three respectively, then the importance level of this sub-region can be set to level two.

[0113] (24) The step of encoding different regions at different bitrates according to the importance levels of different regions of the current frame image includes: encoding different sub-regions at different bitrates according to the importance levels of different sub-regions of the current frame image.

[0114] Exemplarily, when the importance level is level one, the bitrate K = k1 is set; when the importance level is level two, the bitrate K = k2 is set; when the importance level is level three, the bitrate K = k3 is set; when the importance level is level four, the bitrate K = k4 is set, where k1 > k2 > k3 > k4.

[0115] In some embodiments, the targets in the image include not only people but also objects, that is, the targets in the image are mainly divided into two categories: people and objects. Correspondingly, the target attributes in the target recognition result include people and objects. Among them, person recognition mainly uses image recognition algorithms based on biological features such as faces and limbs; object recognition mainly uses supervised learning in machine learning. Specifically, during implementation, first, a training image set containing common objects in the meeting room (such as computers, mice, water cups, mobile phones, desks, chairs, walls, green plants, etc.) is established, then the object recognition model is trained using the training image set, and then the trained object recognition model is used to recognize the objects in the current image frame.

[0116] Optionally, the steps of setting the importance levels for other targets according to the relationship between other targets other than the current speaker and the current speaker in the target recognition result include (31) to (34), specifically as follows:

[0117] (31) Set the importance level of the people other than the current speaker in the current frame image to level two;

[0118] (32) Set the importance level of the objects having an interaction relationship with the current speaker in the current frame image to level two;

[0119] (33) Set the importance level of the object in the current frame image that has no interaction relationship with the current speaker but has an interaction relationship with a person other than the current speaker to level three;

[0120] (34) Set the importance level of the object in the current frame image that has no interaction relationship with any person to level four.

[0121] In the embodiments of the present disclosure, when the video acquisition unit includes only one camera, the current frame image is the image captured by this camera at this time. When classifying the image content according to importance, set the importance level of the current speaker to level one, the importance level of other persons to level two. For the object of the objects in the meeting room, further set different importance levels according to the interaction relationship between the object and the person. Specifically, if the object has an interaction relationship with the current speaker, set the importance level of this object to level two. If the object has no interaction relationship with the current speaker but has an interaction relationship with other persons, then set the importance level of this object to level three. If the object has no interaction relationship with any person, then set the importance level of this object to level four.

[0122] Optionally, the interaction relationship is determined by the contact relationship between the object and the person and the object type. The object type includes the first type of object and the second type of object. The object being the first type of object and having a contact relationship with the person indicates that this object has an interaction relationship with this person. The object being the first type of object and having no contact relationship with the person indicates that this object has no interaction relationship with this person. The object being the second type of object indicates that this object has no interaction relationship with any person.

[0123] In the embodiments of the present disclosure, the objects in the meeting room are divided into two categories. The first type of object refers to objects such as computers, mice, water cups, mobile phones, desks, and chairs that have a close relationship with people. The second type of object mainly refers to the constituent elements of the background environment of the meeting room, such as background walls, background plants, etc.

[0124] For the first type of object, if it has a contact relationship with the person, then it is determined that this object has an interaction relationship with the person. Among them, whether the object and the person have a contact relationship can be determined according to the target positions of each target in the target recognition result. For example, for person A1 and object B1, if the position of person A1 in the current frame image intersects with the position of object B1 or the distance between the two positions is less than the preset distance threshold, then it can be considered that person A1 and object B1 have a contact relationship. Otherwise, it is considered that person A1 and object B1 have no contact relationship.

[0125] For the second type of object, directly determine that it has no interaction relationship with any person.

[0126] It can be understood that, in the embodiments of the present disclosure, setting importance levels for different targets all refers to setting the importance level of the image data at the position of the target in the current frame image to a determined importance level.

[0127] In a possible implementation manner, the audio-visual acquisition unit includes a plurality of cameras, and the current frame image is formed by splicing a plurality of sub-images currently acquired by the plurality of cameras. Exemplarily, please refer to Figure 4 , Figure 4 FIG. is a schematic diagram of a current frame image acquired when the audio-visual acquisition unit includes a plurality of cameras. Figure 4 Taking the audio-visual acquisition unit including four cameras C1, C2, C3, and C4 as an example, assuming that the original images acquired by each camera are represented as sub-images, and the sub-images acquired by cameras C1, C2, C3, and C4 are respectively denoted as P1, P2, P3, and P4, then the current frame image is an image formed by splicing sub-images P1, P2, P3, and P4.

[0128] Among them, the plurality of cameras may be located in the same conference room or in multiple different conference rooms. For example, when there is only one conference room in the conference scenario, at this time, the plurality of cameras may be distributed at different positions in the conference room to collect specific images of different perspectives of the conference room; when the conference scenario includes a main venue and branch venues, the plurality of cameras are distributed in the conference rooms in the main venue and branch venues, and generally, the conference room in the main venue is larger and usually has a plurality of cameras installed therein, while the conference room in the branch venue is smaller and usually has one camera installed therein. In the embodiments of the present disclosure, regardless of whether the conference room includes one or multiple, the sub-images acquired by each camera are all used to splice and form the current frame image, and the processing principle is the same. The difference is that when obtaining the first position of the current speaker in the target space according to the current audio data, if the plurality of cameras are distributed in one conference room, the principle of determining the first position is the same as that Figure 2 shown in the figure, if the plurality of cameras are distributed in multiple conference rooms, then it is first necessary to determine the conference room where the audio acquisition unit is located through the current audio data, and then based on Figure 2 the principle shown in the figure to determine the position of the current speaker in the conference room.

[0129] At this time, the steps of setting different importance levels for different regions of the current frame image according to the target recognition result and the second position include (41) to (42), specifically as follows:

[0130] (41) Set the importance level of the first sub-image in the current frame image to level one, where the first sub-image is the sub-image containing the current speaker.

[0131] (42)Obtain the voice - sounding time of each second sub - image in the current frame image, and set an importance level for each second sub - image according to the voice - sounding time. The voice - sounding time is positively correlated with the importance level. For any second sub - image, the voice - sounding time of this second sub - image represents the cumulative voice - sounding time of the person in this second sub - image within a preset time window. The second sub - image is a sub - image other than the first sub - image in the current frame image.

[0132] In the embodiment of the present disclosure, a time - counting module is provided. The time - counting module is used to record the information of the sub - image where the current speaker is located detected in real - time, and statistically obtain the cumulative voice - sounding time of each sub - image within the preset time window according to the recorded information within the preset time window.

[0133] Exemplarily, the preset time window is T. Assume that the current speaker at the current moment is U1, the sub - image where the current speaker U1 is located is sub - image P1. And within the preset time window T before the current moment, a total of 3 participants U2, U3, U4 have spoken. The sub - image where participant U2 is located is P2, the sub - image where participant U3 is located is P3, and the sub - image where participant U4 is located is P4. The speaking time of participant U2 within the preset time window T is t1, the speaking time of participant U3 within the preset time window T is t2, and the speaking time of participant U4 within the preset time window T is t3. Then the cumulative voice - sounding time corresponding to sub - image P2 statistically obtained by the time - counting module is t1, the cumulative voice - sounding time corresponding to sub - image P3 is t2, and the cumulative voice - sounding time corresponding to sub - image P4 is t3. If t1>t2>t3, then the importance level of sub - image P1 can be set as the first level, the importance level of sub - image P2 as the second level, the importance level of sub - image P3 as the third level, and the importance level of sub - image P4 as the fourth level.

[0134] In the embodiment of the present disclosure, setting different importance levels for different regions of the current frame image is carried out with sub - images as the basic unit. Specifically, when implementing, the importance level of the sub - image where the current speaker is located, that is, the first sub - image, is set as the first level. For other sub - images other than the first sub - image, that is, the second sub - images, the importance level is set according to the voice - sounding time of each sub - image.

[0135] Correspondingly, the step of encoding different regions at different bitrates according to the importance levels of different regions of the current frame image includes: encoding different sub - images at different bitrates according to the importance levels of different sub - images of the current frame image.

[0136] Exemplarily, continue to take the importance level of sub-image P1 as level one, the importance level of sub-image P2 as level two, the importance level of sub-image P3 as level three, and the importance level of sub-image P4 as level four as an example. Then, the bitrates corresponding to sub-images P1, P2, P3, and P4 can be set as k1, k2, k3, and k4 respectively, where k1 > k2 > k3 > k4.

[0137] In the related art, when a video is encoded, it includes an Intra-coded frame (abbreviated as I-frame) and a Predictive-coded frame (abbreviated as P-frame). Among them, an I-frame is a completely independent frame. When encoding, it does not depend on the information of other frames. It encodes a single-frame image through intra-frame compression technology and uses the intra-frame spatial redundancy information to reduce the data volume. The I-frame contains complete image information, and only the data of the current I-frame is required for decoding to reconstruct the complete image. While a P-frame is a predictive-coded frame that depends on the previous frame. It reduces the data volume by recording the difference information between the current frame and the previous frame. When encoding, the P-frame only stores the difference information from the previous frame and does not need to store the complete image data separately. The compression efficiency of P-frames is relatively high because they mainly store the changing information rather than the complete image data, thus significantly reducing the data volume. And different compression degrees can be set during I-frame encoding, and the compression degree can be represented by the bitrate; while the compression efficiency during P-frame encoding is mainly determined by the fact that the P-frame only records the difference information from the previous frame.

[0138] In the embodiments of the present disclosure, for the current frame image, multiple conditions are judged to determine whether the current frame image adopts the key-frame encoding method or the predictive-frame encoding method, that is, I-frame encoding or P-frame encoding. Among them, the judgment conditions include but are not limited to: the number of frames between the current frame image and the previous I-frame, the proportion of the area of non-important regions in the current frame image, and the degree of difference between the current frame image and the previous frame image.

[0139] Optionally, as Figure 5 shown, the step of encoding different regions at different bitrates according to the importance levels of different regions of the current frame image includes:

[0140] Step S201, judge whether the number of image frames between the current frame image and the target frame image is greater than a preset frame number threshold. The target frame image is the previous image encoded by the key-frame encoding method. If the judgment result is yes, execute step S205. If the judgment result is no, execute step S202.

[0141] Among them, the previous image encoded using the key frame encoding method can also be understood as the previous I-frame image in the video. Exemplarily, if the preset frame number threshold is 25, when the number of intervals between the current frame image and the previous I-frame is greater than 25, the current frame image is directly encoded using the key frame encoding method, that is, step S205 is executed. Conversely, if the number of intervals between the current frame image and the previous I-frame is less than or equal to 25, the area ratio of the non-important regions in the current frame image is further determined, that is, step S202 is executed.

[0142] Step S202: Determine whether the area ratio of the non-important regions in the current frame image is greater than the first preset ratio threshold. The non-important regions are regions with an importance level lower than the first preset level threshold. If the determination result is greater than the first preset ratio threshold, step S205 is executed. Conversely, if the determination result is less than or equal to the first preset ratio threshold, step S203 is executed.

[0143] Among them, the higher the importance level, the more important the content of this part of the image, and higher resolution and clarity are required. On the contrary, the lower the importance level, the less important the content of this part of the image, and a greater degree of compression can be performed.

[0144] The non-important regions refer to regions with an importance level lower than the first preset level threshold. For example, if the first preset level threshold is the second level, then the third level, fourth level, or lower-level regions are non-important regions. The area ratio of the non-important regions refers to the ratio of the area occupied by the non-important regions to the total area of the current frame image.

[0145] Exemplarily, the first preset ratio threshold is 50%, and the first preset level threshold is the second level. In the current frame image, if the ratio of the area occupied by the regions with the third and fourth importance levels to the total area of the current frame image is greater than 50%, it means that in the current frame image, the area of the non-important display content regions is large, while the area of the important display content regions is small. That is, in the entire image, the area ratio of the important display content regions is small. At this time, the key frame encoding method can be used for encoding, that is, step S205 is executed. When using key frame encoding, since the area ratio of the important display content regions is small, the image quality can be relatively low at this time. Therefore, a relatively small bit rate can be set overall, that is, a greater degree of compression is performed, so that the compressed file is as small as possible, especially applicable to the scenario where the compression degree is only determined by the size of the difference information during predictive frame encoding. Conversely, if the area ratio of the important display content regions in the current frame image is large, it is necessary to further determine the difference between the current frame image and the previous frame image, that is, step S203 is executed.

[0146] It can be understood that in other embodiments, the first preset ratio threshold can also be set to other values according to the actual scenario.

[0147] Step S203: Determine whether the ratio of the difference area between the current frame image and the previous frame image to the total area of the current frame image is greater than a second preset ratio threshold. If the determination result is greater than the second preset ratio threshold, execute Step S205; if the determination result is less than or equal to the second preset ratio threshold, execute Step S204.

[0148] Exemplarily, the second preset ratio threshold is 30%. It can be understood that in other embodiments, the second preset ratio threshold can also be set to other values according to the actual scenario.

[0149] If the ratio of the difference area between the current frame image and the previous frame image to the total area of the current frame image is greater than 30%, it indicates that the difference between the current frame image and the previous frame image is relatively large. At this time, the key frame coding method is adopted, that is, execute Step S205. On the contrary, it indicates that the difference between the current frame image and the previous frame image is relatively small. At this time, the predictive frame coding method is adopted, that is, execute Step S204.

[0150] Step S204: Adopt the predictive frame coding method to encode different regions of the current frame image at different bitrates according to the importance level. When using the predictive frame coding method for encoding, encoding different regions of the current frame image at different bitrates specifically means that for the difference area between the current frame image and the previous frame image, different regions within the difference area are encoded at different bitrates.

[0151] In the related art, when using the predictive frame coding method for encoding, the compression degree is determined by the size of the difference information. And since only the difference information between the current frame and the previous frame is stored, the amount of information is small. Therefore, usually no further compression is performed. In the embodiments of the present disclosure, when using the predictive frame coding method for encoding, on the one hand, only the data of the difference area is stored. On the other hand, different bitrates are set for different regions within the difference area according to the importance level to achieve further data compression.

[0152] It can be understood that when using the predictive frame coding method for encoding in the embodiments of the present disclosure, the encoding method in the related art can also be adopted, that is, only considering storing the difference information between the current frame and the previous frame, and not performing further compression according to the importance level of the difference area.

[0153] Step S205: Adopt the key frame coding method to encode different regions of the current frame image at different bitrates according to the importance level.

[0154] For the key frame coding method and the predicted frame coding method, their coding processes both include three main processes: transformation (Transform), quantization (Quantization), and entropy coding (Entropy Coding). Among them, transformation is to convert the residual data (spatial domain) in the pixel domain into the frequency domain to concentrate the energy for compression; quantization is to divide the frequency domain coefficients after transformation by the quantization step size, round after integerization to reduce the data precision, discard the high-frequency information insensitive to the human eye, and greatly reduce the data volume; entropy coding is used to perform lossless compression on the quantized coefficients to further reduce redundancy.

[0155] In the image processing method of the embodiments of the present disclosure, when the number of image frames between the current frame image and the previous I frame is greater than a certain number, the proportion of the area of the important display content in the current frame image is relatively small, or although the proportion of the area of the important display content is relatively large but the difference from the previous frame image is relatively large, I frame coding is adopted; when the proportion of the area of the important display content is relatively large but the difference from the previous frame image is relatively small, P frame coding is adopted. Such a setting can ensure the image quality of the important display content on the basis of reducing the data volume and saving encoding and decoding resources, retain more details for the important display content, and make it clearly presented.

[0156] Correspondingly, when the video acquisition unit includes multiple cameras, the steps of encoding different sub-images at different bit rates according to the importance levels of different sub-images of the current frame image are as Figure 6 shown, including:

[0157] Step S301, for any sub-image, determine whether the sub-image is a third sub-image, where the third sub-image is a sub-image whose proportion of the importance level is greater than the preset level proportion threshold. If the determination result is that the sub-image is a third sub-image, execute step S302; otherwise, execute step S308.

[0158] In the embodiments of the present disclosure, for the sub-images collected by each camera, the importance level of the first sub-image where the current speaker is located is set to level one. For the remaining other sub-images, the importance levels are set according to the voice time corresponding to the sub-images. If the voice times corresponding to each sub-image are different, then there are at most n importance levels in the end, where n represents the number of sub-images, that is, the number of cameras in the video acquisition unit.

[0159] For any sub-image, the proportion a of the importance level of the sub-image refers to the ratio of the importance level of the sub-image to all importance levels. Exemplarily, if the importance level of the sub-image is i and the current frame image includes a total of n importance levels, then the proportion a of the importance level of the sub-image is a = i / n. The preset proportion threshold is a fixed value set in advance, which is a value greater than 0 and less than 1. For example, if the preset proportion threshold is 30%, then the third sub-image represents a sub-image with an importance level proportion greater than or equal to 30%. This third sub-image can also be understood as the top 30% of the most important sub-images. For the top 30% of the most important sub-images, the same scheme as that in Figure 5 the embodiment shown is adopted when encoding them, that is, steps S302 to S306 are executed. For the remaining 70% of the sub-images, the encoding scheme is as shown in steps S307 to S309.

[0160] Step S302: Determine whether the number of sub-images between the sub-image and the target frame sub-image is greater than the preset frame number threshold. The target frame sub-image is the sub-image that was encoded using the key frame encoding method in the sub-images captured by the same camera and is the previous one. If the judgment result is yes, execute step S306; if the judgment result is no, execute step S303.

[0161] In the embodiments of the present disclosure, the current frame image is formed by splicing multiple sub-images. For a certain number of sub-images with a higher importance level, that is, the third sub-images, it is determined whether to use I-frame encoding or P-frame encoding based on the number of frames between the sub-image and the previous I-frame, the proportion of the area of the non-important region in the sub-image, and the degree of difference between the sub-image and the previous frame sub-image. The target frame sub-image refers to the sub-image that was encoded using I-frame encoding in the sub-images captured by the same camera and is the previous one.

[0162] Step S303: Determine whether the proportion of the area of the non-important region in the sub-image is greater than the first preset proportion threshold. The non-important region is a region with an importance level less than the first preset level threshold. If the judgment result is yes, execute step S306; otherwise, execute step S304.

[0163] Step S304: Determine whether the ratio of the difference region between the sub-image and the previous frame sub-image to the total area of the sub-image is greater than the second preset proportion threshold. If the judgment result is yes, execute step S306; otherwise, execute step S305. Here, the previous frame sub-image is the previous frame sub-image of the current sub-image in the sub-images captured by the same camera.

[0164] Among them, the specific principles of step S303 and step S304 can refer to the descriptions of step S202 and step S203.

[0165] Step S305: Using the predictive frame coding method, encode different regions of the sub-image at different bitrates according to the importance levels of different regions in the sub-image.

[0166] Step S306: Using the key frame coding method, encode different regions of the sub-image at different bitrates according to the importance levels of different regions in the sub-image.

[0167] Step S307: Determine whether the number of sub-images between the sub-image and the target frame sub-image is greater than a preset frame number threshold. If the determination result is yes, execute Step S308; if the determination result is no, execute Step S309.

[0168] Step S308: Using the key frame coding method, encode different regions of the sub-image at different bitrates according to the importance levels of different regions in the sub-image.

[0169] Step S309: Using the predictive frame coding method, encode different regions of the sub-image at different bitrates according to the importance levels of different regions in the sub-image.

[0170] It can be understood that in the embodiments of the present disclosure, in order to implement Steps S302 to S306, in the process of setting different importance levels for different regions of the current frame image according to the target recognition result and the second position, in addition to setting the importance level for each sub-image at the sub-image level in Steps (41) to (42), it is also necessary to further set the importance level for different regions of each sub-image according to the target recognition result. The principle is as shown in Steps (31) to (34), which will not be elaborated here.

[0171] Compared with the related art, in the image processing method of the embodiments of the present disclosure, for sub-images with an importance level ratio in the top 30%, when the number of sub-images between the sub-image and the previous I-frame is greater than a certain number, the area ratio of the important display content in the sub-image is relatively small, or the area ratio of the important display content is relatively large but the difference from the previous frame sub-image is relatively large, I-frame coding is used; when the area ratio of the important display content is relatively large but the difference from the previous frame sub-image is relatively small, P-frame coding is used; for sub-images with an importance level ratio not in the top 30%, when the number of sub-images between the sub-image and the previous I-frame is greater than a certain number, I-frame coding is used, and when it is less than or equal to a certain number, P-frame coding is used. Such settings can ensure the image quality of important display content, retain more details of important display content, and make it clearly presented on the basis of reducing the data volume and saving encoding and decoding resources.

[0172] In a possible implementation, after encoding different sub-images at different bitrates according to the importance levels of different sub-images in the current frame image, steps (51) to (53) are further included, specifically as follows:

[0173] (51) Decode the encoded sub-images.

[0174] (52) Display the decoded first sub-image in the main display area.

[0175] (53) Display the decoded second sub-image in the secondary display area. For any second sub-image, the higher the importance level of the second sub-image, the larger the area occupied by the second sub-image in the secondary display area.

[0176] Exemplarily, as Figure 6 shown in steps S310 to S312, for the top 30% of the most important sub-images, when decoding, it is determined whether the importance level of the sub-image is level one. If the sub-image is level one, that is, the sub-image is the first sub-image, then decode it to the main display area for display. If the sub-image is not level one, it means the sub-image is not the first sub-image, then decode it to the secondary display area for display. For the remaining 70% of the sub-images, directly decode them to the secondary display area for display when decoding.

[0177] Among them, the area of the main display area is larger than the area of the secondary display area, that is, a larger display area is allocated to the first sub-image with the highest importance level to highlight the image content where the current speaker is located. Exemplarily, the ratio of the area of the main display area to the overall display area of the display device is greater than or equal to 80%, and the second sub-image is displayed in the secondary display area. Specifically in implementation, the secondary display area is divided into multiple sub-secondary display areas, and each second sub-image is displayed in a sub-secondary display area, and for different second sub-images, different areas of sub-secondary display areas are allocated according to the importance level of the second sub-image. Exemplarily, as Figure 7 shown, when the video acquisition unit includes 7 cameras, the overall display area is divided into a main display area and 6 sub-secondary display areas.

[0178] In the image processing method of the embodiments of the present disclosure, when the video acquisition unit includes multiple cameras, in addition to using different encoding methods and different bitrates for different sub-images in the encoding stage according to the importance levels of different regions of the current frame image, the display of the current frame image on the display device after decoding is also improved. Specifically, different areas of display regions are allocated to different sub-images according to the importance level. Such a setting can highlight the key display content, that is, a larger display area is allocated to the image content with a high importance level to preferentially ensure its clarity.

[0179] Optionally, after the step of displaying the decoded first sub-image in the main display area, steps (61) to (64) are further included, specifically as follows:

[0180] (61) Detect the duration after the current speaker stops speaking.

[0181] (62) Determine whether the duration is less than a first time threshold. If the determination result is yes, execute step (63); otherwise, if the determination result is no, execute step (64).

[0182] (63) Control the main display area to continuously display the first sub-image.

[0183] (64) Control the display content of the main display area to switch to the sub-image including the latest current speaker.

[0184] In a conference scenario, there may be a scenario where multiple people are discussing. In this scenario, the speaker changes frequently. At this time, if the main display area displays the first sub-image where the current speaker is located in real time, it may cause the display content in the main display area to switch frequently. Frequent switching may cause delays or lags in the display content, which is likely to cause discomfort to the viewer.

[0185] To solve this problem, the embodiment of the present disclosure will time the time after the current speaker stops speaking, that is, detect the duration after the current speaker stops speaking. For example, if the current speaker is U1, the image processing method of the embodiment of the present disclosure will obtain the duration after the current speaker U1 stops speaking in real time. If the duration is less than the first time threshold te, for example, the first time threshold te is 5 seconds, then at this time, the main display area still displays the sub-image where the current speaker U1 is located, that is, the image captured by the corresponding camera in real time. Only when the duration that the current speaker U1 does not speak exceeds the first time threshold te, the display content of the main display area may switch to the sub-image where the latest speaker is located. In this way, it is possible to avoid frequently switching the display content of the main display area in a scenario where multiple people are discussing, avoid phenomena such as display delays or lags, and improve the viewing experience of users.

[0186] Optionally, the main display area includes a plurality of sub-main display areas, and the image processing method further includes: respectively displaying the sub-images corresponding to the latest plurality of current speakers in the plurality of sub-main display areas.

[0187] In the embodiment of the present disclosure, in a scenario where multiple people are discussing, the sub-images of the latest plurality of speakers are all displayed in the main display area. Such a setting can facilitate the participants to intuitively see information such as which are the current main speakers and the status of each speaker, and enhance the viewing experience of the participants. Exemplarily, please refer to Figure 8 , Figure 8Schematic diagram of a main display area including multiple sub-main display areas Figure 8 In the illustrated example, the video acquisition unit includes 10 cameras. Among them, the main display area is divided into 4 sub-main display areas. That is, the main display area displays the sub-images corresponding to the latest four current speakers in real time. For example, at a certain moment, 4 people are participating in a discussion. At this time, the images of the 4 speakers are displayed in the main display area. When the 5th person starts speaking at the next moment, the sub-image corresponding to the person who has not spoken for the longest time among the 4 people is removed from the main display area, and the sub-image of the 5th person is moved from the secondary display area to the main display area.

[0188] It should be noted that the number of sub-main display areas included in the main display area can be set by the user or automatically adjusted according to the situation during a multi-person discussion in the meeting.

[0189] In a possible implementation manner, the image processing method further includes: detecting whether the audio acquisition unit in the branch venue conference room is turned on. When the audio acquisition unit in the branch venue conference room is turned on, the sub-images captured by the cameras in the branch venue conference room are displayed in the main display area.

[0190] The embodiments of the present disclosure are mainly applicable to the scenario where the branch venue conference room includes one camera. During the meeting, the audio acquisition unit in the branch venue conference room is usually in a mute state. When a participant in the branch venue wants to speak, they will turn on the audio acquisition unit, that is, turn on the voice function. At this time, in order to avoid the picture display lagging behind the voice, in this embodiment, when the audio acquisition unit in the branch venue is turned on, the image captured by the camera in the branch venue is displayed in the main display area.

[0191] Based on the same inventive concept, a second aspect of the present disclosure provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the above-mentioned image processing method are implemented.

[0192] Based on the same inventive concept, a third aspect of the present disclosure provides a conference system, including an audio-video acquisition unit 10, an image processing unit 20, and a display device 30. The audio-video acquisition unit 10 is used to acquire audio data and images in the target space. The image processing unit 20 is used to implement the steps of the above-mentioned image processing method. The display device 30 is used to display the images acquired by the audio-video acquisition unit 10.

[0193] In the embodiments of the present disclosure, as Figure 9As shown, the audio and video acquisition unit 10 includes the audio acquisition unit 110 and the video acquisition unit 120 as shown above. Among them, the video acquisition unit 120 includes n cameras, where n is greater than or equal to 1, and the n cameras are respectively denoted as camera 1, camera 2... camera n. The image processing unit 20 is used to implement the steps of the image processing method as described above, and the display device 30 is used to display the images acquired by the video acquisition unit 120 during the meeting, that is, the video.

[0194] In a specific implementation, the image processing unit 20 may include an audio detection module 210, a time counting module 220, a data processing module 230, an encoding module 240, a decoding module 250, and a display control module 260.

[0195] Among them, the audio detection module 210 is connected to the audio acquisition unit 110 and is used to detect the first position of the current speaker in the target space according to the current audio data;

[0196] The time counting module 220 is connected to the audio detection module 210 and is used to record the information of the sub-image where the current speaker is located detected in real time, and statistically obtain the cumulative speaking time of each sub-image within the preset time window according to the recorded information within the preset time window.

[0197] The data processing module 230 is connected to the audio detection module 210, the time counting module 220, and the video acquisition unit 120, and is used to perform target recognition on the current frame image acquired by the video acquisition unit 120, determine the second position of the current speaker in the current frame image according to the first position detected by the audio detection module 210 and the target recognition result, and set different importance levels for different regions of the current frame image according to the target recognition result and the second position.

[0198] The encoding module 240 is used to encode different regions at different bit rates according to the importance levels of different regions of the current frame image. For example, the acquired current frame image is compressed and encoded, and during the compression and encoding process, the bit rates of different regions are adaptively allocated and processed according to the processing results of the data processing module 230.

[0199] The decoding module 250 is used to decode the encoded current frame image. For example, the decoding module 250 decompresses the video compressed and encoded by the encoding module 240 and converts the video from the compressed format into the format for image display.

[0200] The display control module 260 is used to control the decoded current frame image to be displayed on the display device 30. Exemplarily, the display control module 260 is used to control the decoded first sub-image to be displayed in the main display area and the decoded second sub-image to be displayed in the secondary display area; Exemplarily, the display control module 260 respectively displays the sub-images corresponding to the latest multiple current speakers in multiple sub-main display areas.

[0201] Optionally, Figure 9 In the illustrated embodiment, the encoding module 240 adopts an independent encoding module. However, in other embodiments, as Figure 10 shown, the encoding module 240 may also be integrated into the video acquisition unit 120. At this time, the data processing module 230 feeds back the importance levels corresponding to different regions of the current frame image to the encoding modules at the positions of the respective cameras of the video acquisition unit 120. The encoding modules at the positions of the respective cameras encode different regions at different bitrates according to the importance levels.

[0202] Based on the same inventive concept, a fourth aspect of the present disclosure provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of the image processing method described above are implemented.

[0203] In a specific implementation process, the computer storage medium may include: Universal Serial Bus Flash Drive (USB for short), mobile hard disk, Read Only Memory (ROM for short), Random Access Memory (RAM for short), magnetic disk or optical disc and other storage media that can store program codes.

[0204] Based on the same inventive concept, a fifth aspect of the present disclosure provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the image processing method described above are implemented. Since the principle of solving problems by the above computer program is similar to the principle of the image processing method, the implementation of the above computer program can refer to the implementation of the image processing method, and the repeated parts will not be elaborated.

[0205] A computer program product may employ any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. A readable storage medium may, for example, but not be limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the foregoing. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CDROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0206] Obviously, the above embodiments of the present disclosure are merely examples for clearly illustrating the present disclosure, and are not intended to limit the implementation manners of the present disclosure. For those of ordinary skill in the art, other different forms of changes or variations can be made based on the above description. It is impossible to enumerate all implementation manners here. Any obvious changes or variations derived from the technical solutions of the present disclosure still fall within the protection scope of the present disclosure.

Claims

1. An image processing method, characterized in that, Including the following steps: Obtain the current audio data and the current frame image collected by the audio-video acquisition unit in the target space; Obtain the first position of the current speaker in the target space according to the current audio data, and perform target recognition on the current frame image. The target recognition result includes target attributes and target positions, and the target attributes include people; Determine the second position of the current speaker in the current frame image according to the first position and the target recognition result, and set different importance levels for different regions of the current frame image according to the target recognition result and the second position, where the importance level of the region including the current speaker in the current frame image is greater than that of the region not including the current speaker; Encode different regions of the current frame image at different bitrates according to the importance levels of different regions of the current frame image, and the bitrate is positively correlated with the importance level.

2. The image processing method according to claim 1, wherein The step of encoding different regions of the current frame image at different bitrates according to the importance levels of different regions of the current frame image includes: Judge whether the area ratio of the non-important region in the current frame image is greater than a first preset ratio threshold, and the non-important region is a region with an importance level less than a first preset level threshold; If the judgment result is greater than the first preset ratio threshold, use the key frame encoding method to encode different regions of the current frame image at different bitrates according to the importance level; If the judgment result is less than or equal to the first preset ratio threshold, judge whether the ratio of the difference region between the current frame image and the previous frame image to the total area of the current frame image is greater than a second preset ratio threshold; If the determination result is greater than the second preset ratio threshold, use the key frame encoding method to encode different regions of the current frame image at different bitrates according to the importance level; If the judgment result is less than or equal to the second preset ratio threshold, use the predicted frame encoding method to encode different regions of the current frame image at different bitrates according to the importance level.

3. The image processing method according to claim 2, wherein Before the step of judging whether the area ratio of the non-important region in the current frame image is greater than the first preset ratio threshold, it also includes: Judge whether the number of image frames between the current frame image and the target frame image is greater than a preset number of frames threshold, and the target frame image is the previous image encoded by the key frame encoding method; If the judgment result is yes, use the key frame encoding method to encode different regions of the current frame image at different bitrates according to the importance level; If the judgment result is no, execute the step of judging whether the area ratio of the non-important region in the current frame image is greater than the first preset ratio threshold.

4. The image processing method according to any one of claims 1 to 3, characterized in that, The audio-video acquisition unit includes a camera, and the current frame image is the image currently collected by the camera. The step of setting different importance levels for different regions of the current frame image according to the target recognition result and the second position includes: Set the importance level of the image content at the second position of the current frame image to the highest level; Set an importance level for other targets according to the relationship between other targets other than the current speaker in the target recognition result and the current speaker, and the importance level of other targets is less than the highest level; Divide the current frame image into multiple sub-regions according to a preset grid, and the importance level of any sub-region is determined by the importance level of the targets in that sub-region; The step of encoding different regions of the current frame image at different bitrates according to the importance levels of different regions of the current frame image includes: encoding different sub-regions of the current frame image at different bitrates according to the importance levels of different sub-regions of the current frame image.

5. The image processing method according to claim 4, characterized in that, The target attribute further includes an object, the highest level is level one, and the step of setting an importance level for other targets according to the relationship between other targets other than the current speaker in the target recognition result and the current speaker includes: Set the importance level of the person other than the current speaker in the current frame image to level two; Set the importance level of the object having an interaction relationship with the current speaker in the current frame image to level two; Set the importance level of the object having no interaction relationship with the current speaker and having an interaction relationship with a person other than the current speaker in the current frame image to level three; Set the importance level of the object having no interaction relationship with any person in the current frame image to level four.

6. The image processing method according to claim 5, wherein The interaction relationship is determined by the contact relationship between the object and the person and the object type. The object type includes a first type of object and a second type of object. The object being a first type of object and having a contact relationship with a person indicates that the object has an interaction relationship with the person. The object being a first type of object and having no contact relationship with a person indicates that the object has no interaction relationship with the person. The object being a second type of object indicates that the object has no interaction relationship with any person.

7. The image processing method according to any one of claims 1 to 3, characterized in that, The audio-video acquisition unit includes multiple cameras, and the current frame image is formed by splicing multiple sub-images currently captured by the multiple cameras. The step of setting different importance levels for different regions of the current frame image according to the target recognition result and the second position includes: Set the importance level of the first sub-image in the current frame image to level one, and the first sub-image is the sub-image containing the current speaker; Obtain the voice time of each second sub-image in the current frame image, and set an importance level for each second sub-image according to the voice time. The voice time is positively correlated with the importance level. For any second sub-image, the voice time of the second sub-image represents the cumulative voice time of the person in the second sub-image within a preset time window. The second sub-image is the sub-image other than the first sub-image in the current frame image; The step of encoding different regions of the current frame image at different bitrates according to the importance levels of different regions of the current frame image includes: encoding different sub-images of the current frame image at different bitrates according to the importance levels of different sub-images of the current frame image.

8. The image processing method according to claim 7, wherein The step of encoding different sub-images of the current frame image at different bitrates according to the importance levels of different sub-images of the current frame image includes: For any sub-image, determine whether the sub-image is a third sub-image, where the third sub-image is a sub-image with an importance level greater than a second preset level threshold; If the determination result is that the sub-image is a third sub-image, determine whether the area ratio of the unimportant region in the sub-image is greater than a first preset ratio threshold, where the unimportant region is a region with an importance level less than a first preset level threshold; If the determination result is greater than the first preset ratio threshold, use the key-frame coding method to encode different regions of the sub-image at different bitrates according to the importance level; If the determination result is less than or equal to the first preset ratio threshold, determine whether the ratio of the difference region between the sub-image and the previous-frame sub-image to the total area of the sub-image is greater than a second preset ratio threshold; If the determination result is greater than the second preset ratio threshold, use the key-frame coding method to encode different regions of the sub-image at different bitrates according to the importance level of different regions in the sub-image; If the determination result is less than or equal to the second preset ratio threshold, use the predictive-frame coding method to encode different regions of the sub-image at different bitrates according to the importance level of different regions in the sub-image; If the determination result is that the sub-image is not a third sub-image, use the key-frame coding method or the predictive-frame coding method to determine the bitrate of the sub-image according to the importance level of the sub-image and encode it using the bitrate; 9. The image processing method according to claim 8, wherein The steps of using the key-frame coding method or the predictive-frame coding method to determine the bitrate of the sub-image according to the importance level of the sub-image and encoding it using the bitrate include: Determine whether the number of sub-images between the sub-image and the target-frame sub-image is greater than a preset number-of-frames threshold, where the target-frame sub-image is the previous sub-image encoded using the key-frame coding method among the sub-images captured by the same camera; If the determination result is yes, use the key-frame coding method to determine the bitrate of the sub-image according to the importance level of the sub-image and encode it using the bitrate; If the determination result is no, use the predictive-frame coding method to determine the bitrate of the sub-image according to the importance level of the sub-image and encode it using the bitrate; 10. The image processing method according to claim 7, wherein After encoding different sub-images at different bitrates according to the importance levels of different sub-images of the current-frame image, it further includes: Decode the encoded sub-images; Display the decoded first sub-image in the main display area; Display the decoded second sub-image in the secondary display area. For any second sub-image, the higher the importance level of the second sub-image, the larger the area occupied by the second sub-image in the secondary display area.

11. The image processing method according to claim 10, wherein After the step of displaying the decoded first sub-image in the main display area, it further includes: Detect the duration after the current speaker stops speaking; Determine whether the duration is less than a first time threshold; When the determination result is yes, control the main display area to continuously display the first sub-image; When the determination result is no, control the display content of the main display area to switch to the sub-image including the latest current speaker.

12. The image processing method according to claim 10, wherein The main display area includes multiple sub-display areas, and the image processing method further includes: The sub-images corresponding to the latest multiple current speakers are respectively displayed in multiple sub-display areas.

13. The image processing method according to claim 1, wherein The audio-video acquisition unit includes a plurality of microphones, and the positions of the plurality of microphones are different in the two-dimensional plane of the target space. The step of obtaining the first position of the current speaker in the target space according to the current audio data includes: Obtaining the distances between the microphones in the two-dimensional plane and the time differences of the sounds detected by the microphones; Determining the first position of the current speaker in the two-dimensional plane according to the distances and the time differences.

14. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, the steps of the image processing method according to any one of claims 1-13 are implemented.

15. A conference system, characterized in that, It includes an audio-video acquisition unit, an image processing unit and a display device. The audio-video acquisition unit is used to acquire audio data and images in the target space. The image processing unit is used to implement the steps of the image processing method according to any one of claims 1-13. The display device is used to display the images acquired by the audio-video acquisition unit.

16. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, the steps of the image processing method according to any one of claims 1-13 are implemented.

17. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, the steps of the image processing method according to any one of claims 1-13 are implemented.

Citation Information

Cited By

  • Online conference live broadcast method and system, terminal and medium

    CN121262394A