Video conference speaker positioning method, system and device and storage medium
By analyzing the audio signals captured by the conference microphones and combining the intensity and layout information to calculate the speaker's position, the camera's shooting is dynamically adjusted, solving the problem that cameras cannot track speakers in traditional video conferencing and improving the smoothness and interactive experience of video conferencing.
Patent Information
- Application Number
- CN202511132487.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-13
- Publication Date
- 2025-11-18
AI Technical Summary
In traditional video conferencing, cameras cannot automatically track speakers, resulting in a poor interactive experience and high latency, especially when multiple people are speaking, with frequent screen switching.
The system collects audio signals through conference microphones, combines intensity analysis and microphone layout information to calculate the speaker's position, and controls camera shooting based on priority scores, dynamically adjusting weights to distinguish between single and multi-person conversation scenarios.
It enables the camera to track speakers in real time during video conferences, maintaining high fluency and improving the intuitiveness and efficiency of information transmission.
Smart Images

Figure CN120980340A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of video conference, and particularly relates to a video conference speaker positioning method, device and system. BACKGROUND
[0002] In a traditional video conference system, video picture transmission communication is usually performed through a fixed-position camera. In this way, the camera picture is fixed and cannot automatically track the picture of the person who is speaking according to the conference progress. When multiple people participate in speaking at the conference site, the remote conference participants can only see the picture of a fixed visual angle and cannot clearly focus on the person who is speaking, resulting in poor interactive experience and less intuitive and efficient information transmission of the video conference.
[0003] Nowadays, the camera visual angle is manually adjusted to focus the lens of the video conference on the speaker, but manual adjustment has a large delay, which affects the conference fluency, and when multiple people speak at the same time, the camera cannot intelligently select and track the target, resulting in frequent picture switching. SUMMARY
[0004] The application provides a video conference speaker positioning method, device and system, which can solve the problem that the camera cannot be aimed at the person who is speaking in the prior art when a video conference is performed, resulting in large delay of the video conference.
[0005] The first aspect of the application provides a video conference speaker positioning method, which comprises the following steps:
[0006] transmitting the collected voice signal to the conference host through the conference microphone;
[0007] analyzing the voice signal by the conference host, and obtaining the position coordinates of the current speaker by combining the intensity analysis result and the microphone layout information;
[0008] comparing the priority scores corresponding to the position coordinates, and controlling the camera to take a picture of the current speaker according to the comparison result; wherein the priority score is determined by a preset microphone priority, a speaking activity weight and a voice feature weight.
[0009] The above scheme analyzes the intensity of the received voice signal, and determines the position of the speaker in the conference room in combination with the microphone layout information of the conference room. In order to distinguish between a single speaker scene and a multiple speaker scene, the camera is controlled to aim the lens at the person whose speaking content is most important according to the speaking activity, the fixed microphone priority and the relevance of the speaking content to the conference theme, so as to timely switch the picture to the speaker in the video conference, so that the video conference can always maintain high fluency.
[0010] In a possible implementation of the first aspect, the voice signal is subjected to intensity analysis by the conference host, and the position coordinates of the current speaker are obtained in combination with the intensity analysis result and the microphone layout information, specifically as follows:
[0011] The voice signal is subjected to noise filtering to obtain a noise-reduced signal in which environmental noise is removed.
[0012] The signal intensity is extracted from the noise-reduced signal, and the time difference of arrival of the noise-reduced signal at the conference microphones is calculated.
[0013] In combination with the signal intensity, the time difference, and the microphone layout information, the position coordinates of the current speaker are calculated.
[0014] The above scheme first filters the voice signal, and then analyzes which conference microphone is closest to the speaker according to the signal intensity of the filtered voice signal and the time difference of the conference microphone, to determine the position coordinates of the current speaker.
[0015] In a possible implementation of the first aspect, the voice signal is subjected to noise filtering to obtain a noise-reduced signal in which environmental noise is removed, specifically as follows:
[0016] The noise-reduction threshold of the conference microphone is adjusted according to the signal correlation between the conference microphones, wherein the adjustment strength of the noise-reduction threshold is related to the signal intensity of the voice signal.
[0017] The non-voice noise in the voice signal is identified based on a preset conference room noise sample library.
[0018] The non-voice noise is filtered by the adjusted noise-reduction threshold to obtain the noise-reduced signal.
[0019] The above scheme adjusts the noise-reduction threshold of the conference microphone closest to the current speaker, which retains more voice information in the noise reduction process. In combination with the characteristics of the conference room noise, the conference room-specific noise is identified from the voice signal and removed, thereby effectively achieving noise reduction of the voice signal.
[0020] In a possible implementation of the first aspect, the priority score is determined by a preset microphone priority, a speaking activity weight, and a voice feature weight, specifically as follows:
[0021] The microphone priority of each conference microphone is determined according to the seat order provided by the microphone layout information.
[0022] The cumulative speaking duration and the speaking frequency corresponding to each conference microphone are counted in real time to generate the speaking activity weight of the conference microphone.
[0023] extracting a speech feature from the speech signal, analyzing a volume peak and a speech spectrum change rate according to the speech feature, and generating a speech feature weight;
[0024] extracting a keyword related to a conference theme in the speech signal, and generating a relevance weight according to a frequency of the keyword;
[0025] calculating the priority score according to the microphone priority, the speaking activity weight, the speech feature weight, and the relevance weight.
[0026] The above scheme introduces dynamic weights to determine important speeches. The importance of different speakers is reflected by the microphone priority. According to the speech signal, the tone, content, and speaking activity of the speakers are comprehensively analyzed to adjust the corresponding dynamic weights and accurately determine the most important speaker.
[0027] In a possible implementation method of the first aspect, the method further includes:
[0028] The microphone priority is pre-set;
[0029] The higher the cumulative speaking duration or the speaking frequency, the higher the speaking activity weight;
[0030] When the speech spectrum change rate exceeds a first threshold value within a preset time, it is considered that the speech signal has an emphatic expression, and the speech feature weight is increased.
[0031] In a possible implementation method of the first aspect, the priority scores corresponding to the position coordinates are compared, and the camera is controlled to capture a picture of the current speaker according to the comparison result, specifically:
[0032] Only one priority score is greater than a second threshold value, indicating that the current scene is a single-person conversation scene.
[0033] There are multiple priority scores greater than a second threshold value, indicating that the current scene is a multi-person conversation scene.
[0034] In a single-person conversation scene, the camera is controlled to move to the position coordinate corresponding to the priority score and capture a picture of the current speaker.
[0035] In a possible implementation method of the first aspect, the multi-person conversation scene specifically includes:
[0036] The difference between the priority scores is calculated. If the difference is less than a third threshold value, the camera is switched to a wide angle and simultaneously captures pictures of the position coordinates corresponding to the priority scores.
[0037] If the difference is not less than a third threshold value, the camera is controlled to take a picture of the position coordinate with the highest priority score.
[0038] The above scheme considers the complex situation of a multi-person conversation scenario. If there is speech that is most important to the meeting in the multi-person conversation, the screen of the speaker is preferentially played on the video. If there are multiple important speeches at the same time, the screens of multiple speakers are simultaneously displayed through split-screen, thereby improving the intelligent level of the video conference.
[0039] The second aspect of the present application provides a video conference speaker positioning system, the system comprising: a signal transmission module, a signal analysis module, and a camera positioning module;
[0040] The signal transmission module is configured to transmit the collected speech signal to the conference host through the conference microphone.
[0041] The signal analysis module is configured to analyze the intensity of the speech signal through the conference host, and obtain the position coordinates of the current speaker in combination with the intensity analysis result and the microphone layout information.
[0042] The camera positioning module is configured to compare the priority scores corresponding to the position coordinates, and control the camera to take a picture of the current speaker according to the comparison result. The priority score is determined by a preset microphone priority, a speaking activity weight, and a speech feature weight.
[0043] The third aspect of the present application provides a terminal device, the device comprising: a terminal device comprising a processor and a memory, the memory storing a computer program, and the processor executing the computer program to implement the steps of the video conference speaker positioning method according to any one of the embodiments of the present application.
[0044] The fourth aspect of the present application provides a storage medium storing computer readable program code, which, when executed, implements the steps of the video conference speaker positioning method according to any one of the embodiments of the present application. BRIEF DESCRIPTION OF DRAWINGS
[0045] In order to more clearly illustrate the technical solutions of the present application, the following will briefly introduce the drawings needed in the embodiments. Obviously, the drawings described below are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.
[0046] Figure 1 is a specific flowchart of a video conference speaker positioning method provided by an embodiment of the present application;
[0047] Figure 2 is a structural diagram of a video conference speaker positioning system provided by an embodiment of the present application;
[0048] Figure 3 is a structural diagram of a terminal device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0049] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.
[0050] It should be understood that the step numbers used herein are only for the convenience of description, and are not intended to limit the execution sequence of the steps.
[0051] First Embodiment
[0052] Video conference speaking is generally achieved by collecting conference room video through a camera, and then transmitting the conference room video to the network. However, in a multi-person conversation scenario, because there are many speakers, the camera is difficult to quickly capture the person who is speaking, resulting in poor interactive experience of the video conference and insufficiently intuitive and efficient information transmission. Moreover, when multiple people speak at the same time, the camera cannot intelligently select and track the target, resulting in frequent switching of the picture and affecting the smoothness of the conference.
[0053] As shown in FIG. 1, to solve the problem that the camera cannot be aimed at the person who is speaking during a video conference, resulting in large delay of the video conference, a specific flowchart of a video conference speaker positioning method is provided by the first embodiment of the present application. The video conference speaker positioning method of the present embodiment includes steps S1 to S3, which are described in detail as follows. Figure 1
[0054] In step S1, the collected voice signal is transmitted to the conference host through the conference microphone.
[0055] In the embodiment of the present application, a plurality of conference microphones and cameras are installed in the conference room. The conference microphone is used to collect the voice signal of the conference site and transmit the voice signal to the conference host. The camera is connected to the conference host and is used to receive the control instruction sent by the conference host to complete the picture switching and zooming operation.
[0056] The conference host is used for built-in noise filtering module, personnel positioning module, priority processing module and camera control module. The noise filtering module is used for noise filtering processing on the received voice signal, the personnel positioning module is used for determining the position of the person speaking according to the processed voice signal, the priority processing module is used for processing according to the pre-set conference microphone priority when multiple people are talking at the same time, and the camera control module is used for controlling the camera to track the picture of the person speaking according to the position determined by the personnel positioning module or the processing result of the priority processing module, and adjusting the picture zoom.
[0057] In step S2, the conference host analyzes the intensity of the voice signal, and obtains the position coordinates of the current speaker by combining the intensity analysis result with the microphone layout information.
[0058] The noise filtering module uses a conference scene adaptive noise reduction algorithm to realize signal noise reduction, including adjusting the noise reduction threshold of the conference microphone using a microphone array cooperative noise reduction mechanism, and identifying non-speech noise according to a conference room noise sample library.
[0059] Specifically, according to the microphone array cooperative noise reduction mechanism, the signal correlation between the conference microphones is analyzed, the microphone corresponding to the current speaker is located, and is recorded as the active microphone. The noise reduction threshold of the active microphone is reduced, and the signal gain effect of the non-active microphone is also reduced, achieving strong suppression of the non-active microphone, while reducing background noise interference and preserving more speech details in the voice signal.
[0060] For example, when it is detected that the conference microphone A is the active microphone, the noise filtering module reduces the noise reduction threshold of the conference microphone A to preserve more speech details, and reduces the signal gain of other microphones by 30% to 50% to reduce background noise interference, solving the problem of "misfiltering speaking sound" or "insufficient background noise suppression" in traditional single-channel noise reduction.
[0061] The conference room noise sample library is constructed by noise specific to the conference room, such as book flipping sound, keyboard clicking sound, seat moving sound, and other common non-speech noise characteristics in the conference. Based on the conference room noise sample library, the noise characteristics of the voice signal are identified to determine the non-speech noise in the voice signal that matches the conference room noise sample library.
[0062] For example, when a short pulse signal in the 500-2000Hz frequency band is detected in the voice signal, it is matched with the book flipping sound feature in the conference room noise sample library, and the frequency band notch filter is started to eliminate this type of noise without affecting the voice signal, avoiding non-speech noise triggering camera mis-tracking.
[0063] Then, the non-speech noise is filtered according to the adjusted noise reduction threshold to obtain a noise reduction signal. The signal intensity is extracted from the noise reduction signal, and the time difference of the noise reduction signal reaching different conference microphones is calculated. Then, the position coordinates of the current speaker are calculated in combination with the known microphone layout information.
[0064] The above noise filtering method can not only remove environmental noise, but also adapt to the scene characteristics of "multi-microphone cooperation" and "dynamic change of speaking state" in video conferencing. Compared with existing general audio noise reduction technology, the method has significant improvement in speech clarity maintenance, conference scene noise targeted filtering, and synergy with picture tracking function, which provides technical support for the accuracy of subsequent speaker positioning and the reliability of camera tracking.
[0065] In step S3, the priority scores corresponding to the position coordinates are compared, and the camera is controlled to capture the picture of the current speaker according to the comparison result.
[0066] The priority processing module calculates the priority scores of each speech signal according to the pre-set microphone priority and the dynamically adjusted speaking activity weight and speech feature weight at a predetermined frequency. According to the priority scores, the speaker to be positioned by the camera and the video display mode are determined. The microphone priority is determined according to the seat order provided by the microphone layout information. Generally, the more important the participant, the higher the corresponding microphone priority.
[0067] The speaking activity weight is determined based on the cumulative speaking time and the number of speeches of the conference microphone. According to these data, the speaking activity weight of each conference microphone is determined.
[0068] For example, the priority processing module collects the cumulative speaking time and the number of speeches of each speaker in a preset length of time window. If there is a person whose cumulative speaking time accounts for more than 40% or the number of continuous speeches is greater than or equal to 3, the speaking activity weight of the conference microphone of the person is increased by 20% to 30%. Through the above adjustment method, even if the preset microphone priority is low, the camera may still preferentially position the speaker because the speaker can continuously dominate the discussion. This is because the higher the cumulative speaking time or the number of speeches, the higher the speaking activity weight.
[0069] Meanwhile, the speech features are extracted from the speech signal, and the volume peak and the speech spectrum change rate are analyzed according to the speech features to generate the speech feature weight.
[0070] Specifically, the volume peak and the speech spectrum change rate are extracted from the noise-reduced speech signal to determine the semantic features of the speaker. For example, if the speed of speech is detected to be accelerated, the volume is suddenly increased (e.g., the volume is more than 50% of the average level and lasts for more than 2 seconds), it indicates that the speech signal of the speaker has an expressive expression, and then the corresponding speech feature weight is increased.
[0071] Further, the application also proposes to use the correlation weight to evaluate the relevance of the content in the speech signal to the conference topic. First, the keywords related to the conference topic are extracted from the speech signal, and the keywords related to the conference topic in the speech signal are matched with the conference topic to calculate the corresponding correlation weight.
[0072] For example, the current conference topic is set as "project progress report" and "budget approval", the keywords related to the conference topic are extracted from the speech signal through speech recognition, including "budget amount" and "delay risk", and these words are matched with the conference topic to calculate the corresponding correlation weight. This ensures that even if the speaker's microphone priority is the same as that of others, the speaker can also get more attention because of the content of the speech.
[0073] Finally, the corresponding priority score is calculated according to the microphone priority, the speaking activity weight, the speech feature weight, and the correlation weight. The priority score can be used to represent the priority of the speaker at the camera, and the higher the priority, the more likely the picture is taken.
[0074] The application combines the static microphone priority and the dynamic weight together to obtain the following formula for calculating the priority score:
[0075] Priority score = microphone priority x 50% + activity weight x 20% + attention weight x 20% + topic matching weight x 10%;
[0076] The priority processing module updates the priority score every 1 to 2 seconds.
[0077] The application extends the static control of the pickup permission to a multi-dimensional dynamic evaluation index of the importance of the speech, and combines the real-time state of the conference (speech behavior, content, agenda) to intelligently adjust the camera, so that the lens tracking is more suitable for the speaking picture of the conference.
[0078] As an improvement of the above scheme, the application also compares the priority scores to determine the picture shooting mode.
[0079] Currently, the conference scene is divided into single-person conversation scene and multi-person conversation scene. When there is one priority score greater than a set threshold, it indicates that the current scene is a single-person conversation scene; if there are multiple priority scores greater than a set threshold, it indicates that the current scene is a multi-person conversation scene.
[0080] In the single-person conversation scenario, the camera is controlled to move to the position coordinate corresponding to the priority score and take a picture of the current speaker.
[0081] However, in the complex case of the multi-person conversation scenario, the difference between the plurality of priority scores needs to be calculated first. If the difference is less than 5%, the camera is switched to wide-angle and pictures are taken of the position coordinates corresponding to the plurality of priority scores simultaneously, and the plurality of speaking pictures are displayed through split-screen to realize simultaneous display of multi-person speaking. If the difference is not less than 5%, the camera is controlled to take a picture of the position coordinate corresponding to the highest priority score.
[0082] In addition, in the single-person conversation scenario, the speech fidelity weight in the noise filtering module is automatically increased, and the prosody and details of the speech are preferentially retained in the filtering process of the speech signal. In the multi-person conversation scenario, the noise filtering module is switched to a noise suppression priority mode to enhance the filtering of cross-crosstalk noise and ensure that the speech signal of the main speaker is more prominent.
[0083] The implementation of the present application has the following beneficial effects:
[0084] The embodiment of the present application determines the position of the speaker in the conference room by analyzing the intensity of the received speech signal and combining the microphone layout information of the conference room. In order to distinguish between single-person speaking scenarios and multi-person speaking scenarios, the speaker in the video conference is switched to the person who is currently speaking by controlling the camera to aim at the person whose speaking content is most important according to the speaking activity, the fixed microphone priority, and the relevance of the speaking content to the conference theme, so that the video conference can always maintain high smoothness.
[0085] Second embodiment
[0086] Further, in order to execute the video conference speaker positioning system corresponding to the method embodiment to realize the corresponding functions and technical effects, Figure 2 A structural diagram of a video conference speaker positioning system is provided. For ease of illustration, only the part related to the present embodiment is shown. The video conference speaker positioning system provided by the embodiment of the present application comprises:
[0087] The signal transmission module 201 is configured to transmit the collected speech signal to the conference host through the conference microphone.
[0088] In the embodiment of the present application, a plurality of conference microphones and cameras are installed in the conference room. The conference microphone is used to collect the speech signal of the conference site and transmit the speech signal to the conference host; the camera is connected to the conference host and is used to receive the control instruction sent by the conference host to complete the picture switching and zooming operation.
[0089] The conference host is used for built-in noise filtering module, personnel positioning module, priority processing module and camera control module. The noise filtering module is used for noise filtering processing on the received voice signal, the personnel positioning module is used for determining the position of the person speaking according to the processed voice signal, the priority processing module is used for processing according to the pre-set conference microphone priority when multiple people are talking at the same time, and the camera control module is used for controlling the camera to track the picture of the person speaking according to the position determined by the personnel positioning module or the processing result of the priority processing module, and adjusting the picture zoom.
[0090] The conference microphone collects the voice signal in the conference room and transmits it to the conference host, and the conference host filters the noise of the voice signal through the noise filtering module to obtain a noise-reduced signal.
[0091] The signal analysis module 202 is used for intensity analysis of the voice signal by the conference host, and the position coordinates of the current speaker are obtained by combining the intensity analysis result and the microphone layout information.
[0092] In the embodiment of the application, the voice signal is filtered to obtain a noise-reduced signal that removes environmental noise;
[0093] The signal strength is extracted from the noise-reduced signal and the time difference of the noise-reduced signal reaching the conference microphone is calculated;
[0094] The position coordinates of the current speaker are calculated by combining the signal strength, the time difference and the microphone layout information.
[0095] The camera positioning module 203 is used for comparing the priority score corresponding to the position coordinates, and controlling the camera to take pictures of the current speaker according to the comparison result; wherein the priority score is determined by the pre-set microphone priority, the speaking activity weight and the voice feature weight.
[0096] In the embodiment of the application, only when there is one priority score greater than the second threshold, it means that the current is a single-person conversation scene;
[0097] When there are multiple priority scores greater than the second threshold, it means that the current is a multi-person conversation scene;
[0098] In the single-person conversation scene, the camera is controlled to move to the position coordinates corresponding to the priority score, and take pictures of the current speaker.
[0099] In some embodiments, the signal analysis module 202 specifically comprises:
[0100] The noise filtering module adopts a conference scene adaptive noise reduction algorithm to realize signal noise reduction, including adjusting the noise reduction threshold of the conference microphone by using a microphone array cooperative noise reduction mechanism, and identifying non-speech noise according to a conference room noise sample library.
[0101] Specifically, according to the microphone array cooperative noise reduction mechanism, the signal correlation between the conference microphones is analyzed, the microphone corresponding to the current speaker is located, and is recorded as an active microphone. The noise reduction threshold of the active microphone is reduced, and the signal gain effect of the non-active microphone is also reduced, achieving strong suppression of the non-active microphone, while reducing background noise interference and preserving more speech details in the speech signal.
[0102] For example, when it is detected that the conference microphone A is the active microphone, the noise filtering module reduces the noise reduction threshold of the conference microphone A to preserve more speech details, and reduces the signal gain of other microphones by 30% to 50% to reduce background noise interference, solving the problem of "misfiltering speech" or "insufficient background noise suppression" in traditional single-channel noise reduction.
[0103] The conference room noise sample library is constructed by noise specific to the conference room, such as book flipping sound, keyboard tapping sound, seat moving sound, and other common non-speech noise characteristics in the conference. Based on the conference room noise sample library, noise feature recognition is performed on the speech signal to determine non-speech noise in the speech signal that matches the conference room noise sample library.
[0104] For example, when a short pulse signal in the 500-2000 Hz frequency band is detected in the speech signal, it is indicated that the book flipping sound feature in the conference room noise sample library is matched, and a frequency band notch filter is started to eliminate this type of noise without affecting the speech signal, avoiding non-speech noise triggering camera false tracking.
[0105] Then, the non-speech noise is filtered according to the adjusted noise reduction threshold to obtain the noise reduction signal. The signal strength is extracted from the noise reduction signal, and the time difference of the noise reduction signal reaching different conference microphones is calculated, and the position coordinates of the current speaker are calculated in combination with the known microphone layout information.
[0106] The above noise filtering method not only removes environmental noise, but also adapts to the scene characteristics of "multi-microphone cooperation" and "dynamic change of speaking state" in video conferencing. Compared with existing general audio noise reduction technology, it has significant improvement in speech clarity preservation, conference scene noise targeted filtering, and synergy with picture tracking function, providing technical support for the accuracy of subsequent speaker positioning and the reliability of camera tracking.
[0107] In some embodiments, the camera positioning module 203, in particular:
[0108] The priority processing module calculates the priority scores of the voice signals according to the preset microphone priorities and the dynamically adjusted speaking activity weights and voice feature weights at a predetermined frequency. According to the priority scores, the speaker to be positioned by the camera and the video display mode are determined. The microphone priorities are determined according to the seat order provided by the microphone layout information, and generally, the more important the participant, the higher the corresponding microphone priority.
[0109] For the speaking activity weight, the cumulative speaking time and the speaking frequency of the conference microphone are determined, and the speaking activity weight of each conference microphone is determined according to the data.
[0110] For example, the priority processing module collects the cumulative speaking time and the speaking frequency of each speaker in a preset length of time window. If there is a person whose cumulative speaking time accounts for more than 40% or whose continuous speaking frequency is greater than or equal to 3 times, the speaking activity weight of the conference microphone of the person is increased by 20% to 30%. Through the above adjustment method, even if the preset microphone priority is low, the camera can still be positioned preferentially because the speaker can continuously dominate the discussion. This is because the higher the cumulative speaking time or the speaking frequency, the higher the speaking activity weight.
[0111] Meanwhile, the voice features are extracted from the voice signals, and the volume peak and the voice spectrum change rate are analyzed according to the voice features to generate the voice feature weight.
[0112] Specifically, the volume peak and the voice spectrum change rate are extracted from the noise-reduced voice signals to determine the semantic features of the speaker. For example, if the speed is detected to be accelerated, the volume is suddenly increased (such as the volume exceeding the average level by 50% and lasting for more than 2 seconds), it indicates that the voice signal of the speaker has an emphasis expression, and the corresponding voice feature weight is increased.
[0113] Further, the correlation weight is used to evaluate the relevance of the content in the voice signal to the conference topic. First, the keywords related to the conference topic are extracted from the voice signal, and the keywords related to the conference topic in the voice signal are matched with the conference topic to calculate the corresponding correlation weight. This ensures that even if the microphone priority of the speaker is the same as that of others, the speaker can still get more attention because of the content of the speech.
[0114] For example, the current conference topics are set as “project progress report” and “budget approval”, the keywords related to the conference topic are extracted from the voice signal through voice recognition, including “budget amount” and “delay risk”, and these words are matched with the conference topic to calculate the corresponding correlation weight. This ensures that even if the microphone priority of the speaker is the same as that of others, the speaker can still get more attention because of the content of the speech.
[0115] Finally, a corresponding priority score is calculated according to the microphone priority, the speaking activity weight, the voice feature weight, and the relevance weight. The priority score can be used to represent the priority of the speaker at the camera, and the higher the priority, the more likely the picture is taken.
[0116] The embodiment of the present application combines the static microphone priority and the dynamic weight together to obtain the following formula for calculating the priority score:
[0117] Priority score = microphone priority x 50% + activity weight x 20% + attention weight x 20% + topic matching weight x 10%;
[0118] The priority processing module updates the priority score every 1 to 2 seconds.
[0119] The embodiment of the present application extends the static control of the pickup permission to a multi-dimensional dynamic evaluation index of the speaking importance, and combines the real-time state of the conference (speaking behavior, content, agenda) to intelligently adjust the camera, so that the lens tracking is more suitable for the speaking picture of the conference.
[0120] As an improvement of the above-mentioned scheme, the embodiment of the present application also determines the picture shooting mode by comparing the priority scores.
[0121] Currently, the conference scene is divided into a single-person conversation scene and a multi-person conversation scene. When there is one priority score greater than a set threshold, it indicates that the current scene is a single-person conversation scene; if there are multiple priority scores greater than the set threshold, it indicates that the current scene is a multi-person conversation scene.
[0122] In the single-person conversation scene, the camera is controlled to move to the position coordinates corresponding to the priority score, and the current speaker is taken a picture.
[0123] However, in the complex case of the multi-person conversation scene, the difference between multiple priority scores also needs to be calculated. If the difference is less than 5%, the camera is switched to wide-angle and the position coordinates corresponding to the priority scores are taken a picture at the same time, and multiple speaking pictures are displayed through split screen to realize the simultaneous display of multiple speakers. If the difference is not less than 5%, the camera is controlled to take a picture of the position coordinates corresponding to the highest priority score.
[0124] In addition, in the single-person conversation scene, the voice fidelity weight in the noise filtering module is automatically increased, and the prosody and details of the voice are preferentially retained in the filtering process of the voice signal. In the multi-person conversation scene, the noise filtering module is switched to a noise suppression priority mode to enhance the filtering of cross-crosstalk noise, ensuring that the voice signal of the main speaker is more prominent.
[0125] Implementing the embodiments of this application has the following beneficial effects:
[0126] This application embodiment determines the location of the speaker in the conference room by analyzing the strength of the received voice signal and combining it with the microphone layout information of the conference room. In order to distinguish between single-person and multi-person speaking scenarios, the camera is also controlled to focus on the person most important in the current speech based on the speaker's activity level, fixed microphone priority, and the relevance of the speech content to the conference topic. This enables timely switching of the camera to the speaker during the video conference, ensuring that the video conference maintains a high degree of smoothness.
[0127] Furthermore, Figure 3 This is a structural diagram of a terminal device provided in one embodiment of this application. Figure 3 As shown, the terminal device 3 of this embodiment includes: at least one processor 30 (in... Figure 3 (Only one is shown in the image) and a memory 31 and a computer program 32 stored in the memory 31 and executable on the at least one processor, wherein when the processor 30 executes the computer program 32, it can implement the steps of a video conferencing speaker positioning method according to any one of the embodiments of this application.
[0128] The terminal device 3 may be a computing device such as a desktop computer, a cloud server, or a laptop computer, and the computing device may include, but is not limited to, a processor 30 and a memory 31. Figure 3 This is merely an example of terminal device 3 and does not constitute a limitation on terminal device 3. It may include more or fewer components than those shown in the figure.
[0129] This application provides a storage medium that stores computer-readable program code, which, when executed, implements the steps of the above-described video conferencing speaker location method.
[0130] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of this application. It should be understood that the above descriptions are merely specific embodiments of this application and are not intended to limit the scope of protection of this application. In particular, it should be noted that any modifications, equivalent substitutions, or improvements made by those skilled in the art within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A method for locating speakers in a video conference, characterized in that, include: The collected audio signals are transmitted to the conference host via the conference microphone; The conference host performs intensity analysis on the voice signal, and combines the intensity analysis results with microphone layout information to obtain the current speaker's position coordinates; The priority scores corresponding to the location coordinates are compared, and the camera is controlled to capture the image of the speaker based on the comparison results; wherein, the priority score is determined by preset microphone priority, speaking activity weight and voice feature weight.
2. The video conferencing speaker positioning method according to claim 1, characterized in that, The process involves analyzing the intensity of the voice signal using the conference host and combining the intensity analysis results with microphone layout information to obtain the current speaker's position coordinates. The speech signal is subjected to noise filtering to obtain a noise-reduced signal with environmental noise removed; Extract the signal strength and the time difference between the noise-reduced signal and the arrival time of the noise-reduced signal at the conference microphone from the noise-reduced signal; By combining the signal strength, the time difference, and the microphone layout information, the current speaker's position coordinates are calculated.
3. The video conferencing speaker positioning method according to claim 2, characterized in that, The noise filtering of the speech signal to obtain a noise-reduced signal with environmental noise removed specifically involves: Based on the signal correlation between the conference microphones, the noise reduction threshold of the conference microphones is adjusted; wherein, the adjustment degree of the noise reduction threshold is related to the signal strength of the voice signal; Based on a pre-set conference room noise sample library, noise feature recognition is performed on the speech signal to identify non-speech noise in the speech signal; The non-speech noise is filtered by the adjusted noise reduction threshold to obtain the noise-reduced signal.
4. The video conferencing speaker positioning method according to claim 1, characterized in that, The priority score is determined by preset microphone priority, speaking activity weight, and voice feature weight, specifically: Based on the seating order provided by the microphone layout information, determine the microphone priority of each conference microphone; The cumulative speaking time and number of times each conference microphone is used to generate the speaking activity weight of the conference microphone in real time. Speech features are extracted from the speech signal, and the volume peak and speech spectrum change rate are analyzed based on the speech features to generate the speech feature weights. Extract keywords related to the meeting topic from the speech signal, and generate relevance weights based on the frequency of the keywords; The priority score is calculated based on the microphone priority, the speaking activity weight, the voice feature weight, and the relevance weight.
5. The video conferencing speaker positioning method according to claim 4, characterized in that, Also includes: The microphone priority is preset; The higher the cumulative speaking time or the number of times speaking, the higher the weight of speaking activity. When the rate of change of the speech spectrum exceeds a first threshold within a preset time, the speech signal is considered to have an emphatic expression, and the weight of the speech feature is increased.
6. The video conferencing speaker positioning method according to claim 1, characterized in that, The comparison of priority scores corresponding to the location coordinates, and the control of the camera to capture the image of the speaker based on the comparison result, specifically involves: If there is only one priority score greater than the second threshold, it indicates that the current scenario is a single-person conversation. If multiple priority scores are greater than the second threshold, it indicates that the current scenario is a multi-person conversation. In a one-person conversation scenario, the camera is controlled to move to the position coordinates corresponding to the priority score and to capture the image of the current speaker.
7. The video conferencing speaker positioning method according to claim 6, characterized in that, The multi-person conversation scenario is specifically as follows: Calculate the difference between the priority scores. If the difference is less than a third threshold, switch the camera to wide-angle and simultaneously capture the image at the position coordinates corresponding to the priority scores. If the difference is not less than the third threshold, then the camera is controlled to capture an image of the location coordinates with the highest priority score.
8. A video conferencing speaker positioning system, characterized in that, include: Signal transmission module, signal analysis module, and camera positioning module; The signal transmission module is used to transmit the collected voice signal to the conference host via the conference microphone; The signal analysis module is used to perform intensity analysis on the voice signal through the conference host, and combine the intensity analysis results with the microphone layout information to obtain the position coordinates of the current speaker; The camera positioning module is used to compare the priority scores corresponding to the location coordinates, and control the camera to capture the image of the speaker based on the comparison results; wherein, the priority score is determined by preset microphone priority, speaking activity weight and voice feature weight.
9. A terminal device, characterized in that, The device includes a processor and a memory, the memory storing a computer program, and the processor executing the computer program to implement the steps of a video conferencing speaker positioning method according to any one of claims 1 to 7.
10. A storage medium, characterized in that, The storage medium stores computer-readable program code that, when executed, implements the steps of a video conferencing speaker location method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Method and device for adjusting conventioneer image display in multi-screen video conference
CN104038725A
Video conference shooting device shooting method and picture display method
CN111586341A
Conference speaker positioning method and device, equipment and storage medium
CN116156100A
Intelligent conference management method and system
CN119893030A