Processing method and processing device

By identifying and suppressing discussion behaviors in video conferences, using multimodal big model and voiceprint technology to process video streams, the problem of private conversation interference is solved and meeting efficiency and user experience is improved.

CN120499335APending Publication Date: 2025-08-15LENOVO (BEIJING) LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510699620.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

In video conferences, private conversations and interference behaviors of participants affect the efficiency and communication effects of meetings, and it is difficult for the existing technology to effectively identify and suppress these behaviors.

Method used

By acquiring video streams, extracting image and audio features, using multimodal big models to identify discussion behaviors, and suppressing the audio and image data of the target object through voiceprint features and image processing techniques, adjusting the camera direction to avoid interference areas, and outputting text information of discussion behaviors.

Benefits of technology

Effectively reduce sudden interference and noise in video conferences, improve conference concentration and efficiency, enhance user experience, and ensure clear audio and video quality and communication effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120499335A_ABST
    Figure CN120499335A_ABST
Patent Text Reader

Abstract

The invention provides a processing method and a processing device, and relates to the technical field of terminal equipment. The processing method comprises the following steps: acquiring a video stream of a video conference, wherein the video stream is obtained by shooting a plurality of participating objects of the video conference in the same physical space; the video stream is detected, a target behavior is determined, and the target behavior represents a discussion behavior between participants of the video conference; and determining a target object with the target behavior from the plurality of participating objects, and performing suppression processing on target data of the target object in the video stream.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of terminal equipment, and in particular to a processing method and a processing device. Background Art

[0002] With the popularity of remote work and online meetings, video conferencing has become an important communication tool for businesses and individuals. However, in practice, participants may exhibit some behaviors that affect the progress of the meeting. Summary of the Invention

[0003] In view of the above problems, the present disclosure provides a processing method and a processing device.

[0004] According to a first aspect of the present disclosure, a processing method is provided, comprising: obtaining a video stream of a video conference, where the video stream is obtained by shooting multiple participants in the video conference in the same physical space; detecting the video stream to determine whether a target behavior has occurred, where the target behavior represents a discussion behavior between the participants in the video conference; determining a target object in which the target behavior has occurred from the multiple participants, and suppressing target data of the target object in the video stream.

[0005] According to an embodiment of the present disclosure, a video stream is detected to determine whether a target behavior has occurred, including: extracting image features and audio features from the video stream; inputting the image features, audio features, and target prompt words into a target model to obtain an output result of the target model, wherein the target prompt words are used to guide the target model to perform reasoning; and determining the target behavior in the video stream based on the output result.

[0006] According to an embodiment of the present disclosure, the target data includes target audio data; and the target data of the target object in the video stream is suppressed, including: processing the target audio data in the video stream, so that in the processed video stream, the volume corresponding to the target audio data is smaller than the volume corresponding to the target audio data before processing.

[0007] According to an embodiment of the present disclosure, processing target audio data in a video stream includes: detecting the video stream to determine a main speaker among multiple participants, wherein the main speaker represents a participant who is speaking in a video conference; obtaining a voiceprint feature of the main speaker; and processing the target audio data using the voiceprint feature.

[0008] According to an embodiment of the present disclosure, target audio data is processed using voiceprint features, including: extracting mixed audio data from a video stream, and converting the mixed audio data into time-frequency spectrum features, wherein the mixed audio data includes target audio data and audio data of a main speaker; fusing the voiceprint features of the main speaker with the time-frequency spectrum features to obtain fused features; processing the fused features using a deep neural network model to obtain a time-frequency mask having multiple time-frequency units, wherein each time-frequency unit represents the probability that the time-frequency unit is dominated by the main speaker; multiplying the time-frequency mask with the time-frequency spectrum features to obtain the time-frequency spectrum features of the main speaker; converting the time-frequency spectrum features of the main speaker into time-domain audio data, and replacing the mixed audio data in the video stream with the time-domain audio data.

[0009] According to an embodiment of the present disclosure, the target data includes target image data; suppressing the target data of the target object in the video stream includes: processing the target image data in the video stream, and the processed video stream does not include the target image data.

[0010] According to an embodiment of the present disclosure, processing target image data in a video stream includes: performing segmentation processing on video frames in the video stream.

[0011] According to an embodiment of the present disclosure, the video stream is obtained by collecting the first direction by the target device, and the method also includes: in response to the occurrence of the target behavior, controlling the collection direction of the target device to be adjusted from the first direction to the second direction, and the collection area corresponding to the target device in the second direction is an area outside the area where the target object is located.

[0012] According to an embodiment of the present disclosure, the method further includes: obtaining audio information corresponding to the discussion behavior; converting the audio information into text information; and outputting the text information.

[0013] The second aspect of the present disclosure provides a processing device, including: a video stream acquisition module, used to acquire a video stream of a video conference, where the video stream is obtained by shooting multiple participants in the video conference in the same physical space; a target behavior detection module, used to detect the video stream and determine the occurrence of a target behavior, wherein the target behavior represents the discussion behavior between the participants in the video conference; and a suppression processing module, used to determine the target object that has the target behavior from the multiple participants and suppress the target data of the target object in the video stream.

[0014] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The above and other objects, features and advantages of the present disclosure will become more apparent through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, in which:

[0016] Figure 1 The following schematically illustrates an application scenario of the processing method according to an embodiment of the present disclosure;

[0017] Figure 2 A flowchart of a processing method according to an embodiment of the present disclosure is schematically shown;

[0018] Figure 3 The following schematically shows a principle diagram of determining a target behavior in a video stream according to an embodiment of the present disclosure;

[0019] Figure 4 The following schematically shows a flow chart of processing target audio data according to an embodiment of the present disclosure;

[0020] Figure 5 The block diagram schematically shows a processing device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0021] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the detailed description below, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessary confusion of the concepts of the present disclosure.

[0022] The terms used herein are only for describing specific embodiments and are not intended to limit the present disclosure. The terms "comprise," "include," etc. used herein indicate the presence of features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0023] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0024] When expressions such as "at least one of A, B, and C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).

[0025] Figure 1 The application scenario diagram of the processing method according to the embodiment of the present disclosure is schematically shown.

[0026] like Figure 1 As shown, the application scenario 100 may be a video conferencing scenario, including a processing device 101 for conducting a video conference and a plurality of participants 102A, 102B, 102C, 102D, and 102E (only for example) participating in the video conference in the same physical space.

[0027] The processing device 101 may be any electronic device with a display screen and video conferencing capabilities, including but not limited to a smartphone, a tablet computer, a laptop computer, a desktop computer, and the like.

[0028] For example, the processing device 101 is configured with a target device for capturing multiple participants 102A, 102B, 102C, 102D, and 102E participating in the video conference in the same physical space. For example, the target device may be a camera, a camera, or a drone.

[0029] The multiple participants 102A, 102B, 102C, 102D, and 102E are participants in the video conference in the same physical space. The multiple participants being in the same physical space means that the participants are located in the same place, such as a conference room, studio, office, or open office workstation.

[0030] Each participant is a person who participates in the video conference. According to the role classification, the multiple participants 102A, 102B, 102C, 102D, 102E may include a host, a speaker, a participant, a recorder, a technical support staff, an observer, etc.

[0031] In view of this, an embodiment of the present disclosure provides a processing method, which can be executed by the processing device 101 or by software installed in the processing device 101.

[0032] Please continue reading Figure 1Taking a video conference scenario as an example, first, the processing device 101 can obtain a video stream of the video conference, which is obtained by shooting multiple participating objects 102A, 102B, 102C, 102D, and 102E in the same physical space of the video conference.

[0033] Next, the processing device 101 may detect the video stream and determine whether a target behavior occurs, wherein the target behavior represents a discussion behavior between the participants of the video conference. For example, by detecting the video stream, it is determined that a discussion behavior occurs between the participants 102D and 102E.

[0034] Then, the processing device 101 can identify target objects 102D and 102E that are engaging in target behaviors from among the multiple participants and suppress the target data of target objects 102D and 102E in the video stream. In this way, when multiple participants are videoconferencing in the same conference room, target behaviors such as private conversations can be identified during the video conference. Furthermore, the target objects engaging in the target behaviors in the video stream can be identified and their target data suppressed, thereby reducing undesirable factors such as sudden interruptions and noise during the video conference, maintaining the focus of the video conference, and improving meeting efficiency and the participant experience.

[0035] It should be understood that Figure 1 The number of processing devices, participating objects, and target objects in the figure is merely illustrative. Any number of processing devices, participating objects, and target objects may be provided according to implementation requirements.

[0036] The following will be based on Figure 1 The scene described by Figures 2 to 4 The processing method of the embodiment of the present disclosure is described in detail.

[0037] Figure 2 The flowchart of the processing method according to the embodiment of the present disclosure is schematically shown.

[0038] like Figure 2 As shown, the processing method of this embodiment includes operations S210 to S230, and the processing method can be executed by the above-mentioned processing device.

[0039] In operation S210 , a video stream of a video conference is obtained. The video stream is obtained by shooting a plurality of participants in the video conference in the same physical space.

[0040] The video stream of a video conference refers to a continuous sequence of video data transmitted in real time over the network, which is used to transmit the images in the physical space to other online participating devices.

[0041] For example, the video stream may be a current video stream captured in real time during a video conference, or a historical video stream captured and recorded during a historical period.

[0042] For example, the video stream of the video conference can be captured by a target device pre-configured by the processing device. The target device can capture multiple participants in one direction or in multiple directions.

[0043] In operation S220 , the video stream is detected to determine whether a target behavior occurs, wherein the target behavior represents a discussion behavior between participants in the video conference.

[0044] Target behaviors refer to discussion activities between at least two participants, such as whispering, private conversations, disruptive movements, and gestures. These behaviors can disrupt the video conference, potentially affecting meeting efficiency, disrupting communication, and even infringing on the rights of others.

[0045] In operation S230, a target object that performs a target behavior is determined from the plurality of participating objects, and target data of the target object in the video stream is suppressed.

[0046] After determining that the target behavior occurs during the video conference, the target object where the target behavior occurs can be located, and the target data of the target object can be suppressed, thereby effectively reducing the impact of the target behavior on the video conference and ensuring efficient communication.

[0047] Through the processing method provided by the embodiment of the present disclosure, when multiple participants conduct a video conference in the same conference room, target behaviors such as private conversations during the video conference can be identified, and then the target object performing the target behavior in the video stream can be identified, and the target data of the target object can be suppressed, thereby reducing sudden interference, noise and other adverse factors during the video conference, maintaining the focus of the video conference, and improving the meeting efficiency and participation experience.

[0048] Figure 3 The diagram schematically shows a principle diagram for determining a target behavior in a video stream according to an embodiment of the present disclosure.

[0049] like Figure 3 As shown, in some embodiments, the above operation S220 detects the video stream and determines whether the target behavior occurs, which may further include: extracting image features and audio features from the video stream; inputting the image features, audio features and target prompt words into the target model to obtain the output result of the target model, wherein the target prompt words are used to guide the target model to perform reasoning; and determining the target behavior in the video stream based on the output result.

[0050] First, image and audio features are extracted from the video stream. These features can comprehensively reflect the characteristics of multiple participants in the video conference. For example, image features include facial expressions, posture characteristics, and clothing of multiple participants; audio features include speech content and the voice characteristics of the speaker.

[0051] Next, the image features, audio features, and target prompt are fed into the target model to obtain the target model's output. The target prompt guides the target model's reasoning. By editing the target prompt, the target model can be invoked to perform video understanding and obtain the output.

[0052] The target prompt can be core task information entered through the target model's interactive interface, such as "Please invoke the target model to detect the video stream based on image and audio features to determine whether a discussion is occurring between participants in the video stream." The content of the target prompt can be adjusted based on the needs of the actual video conferencing scenario and can be flexibly increased or decreased. It is understood that more comprehensive the target prompt content, the more it stimulates the target model's reasoning capabilities, resulting in more accurate output results.

[0053] The disclosed embodiments do not specifically limit the method for obtaining the target prompt words. For example, the target prompt words can be obtained through course learning, searched from professional books or materials, obtained from community or forum exchanges, or obtained from ready-made and available templates or libraries.

[0054] The target model can handle time-sensitive multimodal data and is a large multimodal model. This model jointly processes multiple modal data, such as text, images, audio, and video, to mine semantic connections between different data forms and achieve intelligent information processing and generation.

[0055] For example, the target model can detect the lip movements of each participant in the video stream to perform target behavior recognition.

[0056] The output result of the target model can reflect whether a discussion behavior occurs between the participants in the video stream, and thus the target behavior in the video stream can be determined based on the output result.

[0057] The disclosed embodiments do not specifically limit the type of output results of the target model. For example, according to the output modality, the output results of the target model can be single-modal output or cross-modal combined output. Among them, single-modal output includes text generation, image generation, or audio and video generation, and cross-modal combined output includes mixed text and image results, synchronized audio and video output, multi-modal interactive response, etc.

[0058] In an example, taking the target model as a multimodal large model and the output result of the target model as a mixed result of text and images, when image features, audio features and target prompt words are input into the multimodal large model, the output result of the multimodal large model can be synchronously extracted text labels and image labels. These two types of labels reflect whether discussion behavior occurs between the participants in the video stream from different levels.

[0059] After obtaining the text and image tags, the content of these two types of tags can be matched. Based on the matching results, it can be determined whether a discussion activity is occurring between participants in the video stream. For example, a definition can be defined: only when both the text and image tags indicate a discussion activity between participants in the video stream can the target activity be determined to be present in the video stream; otherwise, the target activity cannot be determined. In this way, the target activity in the video stream can be determined based on the output results.

[0060] The processing method provided by the embodiments of the present disclosure combines the image features and audio features in the video stream, and uses the prompt engineering and video understanding technology of a large multimodal model to perform video action recognition, thereby identifying target behaviors such as private conversations during video conferences. This can compensate for comprehension biases through multimodal data, enhance the depth of comprehension, and improve the accuracy of target behavior recognition, thereby having the advantage of processing massive amounts of video data.

[0061] In some embodiments, the target data in the above operation S230 includes target audio data; suppressing the target data of the target object in the video stream includes: processing the target audio data in the video stream, so that in the processed video stream, the volume corresponding to the target audio data is less than the volume corresponding to the target audio data before processing.

[0062] For example, processing the target audio data in the video stream may include reducing the volume of the target audio data or shielding the target audio data, wherein shielding the target audio data may be understood as reducing the volume of the target audio data to zero.

[0063] Through the processing method provided by the embodiment of the present disclosure, after locating the target object that performs the target behavior, the volume corresponding to the audio data of the target object can be suppressed, thereby reducing the semantic confusion and interference of the noise generated by the target behavior on the video conference, improving voice clarity, reducing misunderstandings and communication barriers, improving meeting efficiency, and enhancing user experience.

[0064] Figure 4 The flowchart of processing target audio data according to an embodiment of the present disclosure is schematically shown.

[0065] like Figure 4As shown, in some embodiments, the above-mentioned processing of the target audio data in the video stream may further include operations S401 to S403.

[0066] In operation S401, a video stream is detected to determine a main speaker among a plurality of participants, wherein the main speaker represents a participant who is speaking in a video conference.

[0067] The main speakers are divided according to their roles. The main speaker can be the main speaker of the video conference, who is responsible for the output of core content and can convey information through screen sharing, document / video presentation.

[0068] For example, by detecting the video stream, image features and audio features can be extracted from it, and then the main speaker of the video conference can be determined based on the image features and audio features.

[0069] For example, the above-mentioned target model may be used to detect the video stream and determine the main speaker among multiple participants.

[0070] In operation S402, the voiceprint feature of the speaker is obtained.

[0071] Voiceprint features can represent the uniqueness of the speaker. Voiceprint features can include, for example, Mel-frequency cepstral coefficients (MFCC), fundamental frequency, formant distribution, pronunciation habits, etc.

[0072] In operation S403, the target audio data is processed using the voiceprint feature.

[0073] For example, a filter can be constructed or a mask can be generated based on the voiceprint features of the main speaker to distinguish the audio data of the main speaker in the video stream from the target audio data of the target object performing the target behavior, thereby reducing the volume of the target audio data in the video stream.

[0074] The processing method provided by the disclosed embodiments detects and locates the main speaker among multiple participants in the video stream. The main speaker's voiceprint characteristics are then obtained and used to process the target audio data in the video stream, reducing the volume of the target audio data in the video stream. Based on the uniqueness of voiceprint characteristics, the disclosed embodiments utilize the main speaker's voiceprint characteristics to suppress the voices of discussants. This allows for precise distinction between the main speaker and the target participant, suppressing only the target audio data of the target participant, preserving the complete sound quality and more voice details of the video stream and adapting to complex sound environments.

[0075] Furthermore, the above operation S402 of obtaining the voiceprint feature of the speaker may include: extracting audio data of the speaker from the video stream; and obtaining the voiceprint feature of the speaker that matches the audio data of the speaker from a pre-built voiceprint database.

[0076] Furthermore, the above-mentioned voiceprint database is pre-constructed in the following manner: for each of the multiple participants, audio sample data of the participant for a preset time period before the video conference is collected; the voiceprint sample features of the participant are extracted from the audio sample data using a voiceprint model; and the audio sample data of the multiple participants are associated with the corresponding voiceprint sample features and stored in the voiceprint database.

[0077] Before a video conference, you can collect audio samples of the same preset duration for each participant to facilitate comparisons between different participants. This preset duration can be set based on the actual acoustic environment. By collecting audio for this preset duration, you can ensure that the audio samples of each participant are referenceable and reflect the unique voice characteristics of that participant.

[0078] The voiceprint model can extract the voiceprint sample features of each participant from the audio sample data. The embodiments of this disclosure do not specifically limit the type of voiceprint model. For example, the voiceprint model can be a d-vector model, an x-vector model, or an ECAPA-TDNN model.

[0079] The voiceprint database stores audio sample data of multiple participants, and the audio sample data of each participant is associated with and stores the voiceprint sample features of the participant.

[0080] Furthermore, the above-mentioned obtaining of the voiceprint features of the speaker that match the audio data of the speaker from a pre-constructed voiceprint database includes: matching the audio data with multiple audio sample data stored in the voiceprint database to obtain matching target audio sample data; and determining the voiceprint sample features corresponding to the target audio sample data in the voiceprint database as the voiceprint features of the speaker.

[0081] For example, the aforementioned matching of audio data with multiple audio sample data stored in a voiceprint database includes calculating the similarity between each audio sample data in the voiceprint database and the audio data of the speaker; and selecting the audio sample data with the greatest similarity as the target audio sample data for matching. In this case, the participant corresponding to the audio sample data can be designated as the speaker, and the voiceprint sample features of the participant can be extracted as the voiceprint features of the speaker. It should also be noted that the disclosed embodiments do not specifically limit the method for calculating the similarity described above.

[0082] In some embodiments, the above operation S403 uses voiceprint features to process the target audio data, which may further include: extracting mixed audio data from the video stream, and converting the mixed audio data into time-frequency spectrum features, wherein the mixed audio data includes the target audio data and the audio data of the main speaker; fusing the voiceprint features of the main speaker with the time-frequency spectrum features to obtain a fused feature; using a deep neural network model to process the fused feature to obtain a time-frequency mask having multiple time-frequency units, wherein each time-frequency unit represents the probability that the time-frequency unit is dominated by the main speaker; multiplying the time-frequency mask with the time-frequency spectrum features to obtain the time-frequency spectrum features of the main speaker; converting the time-frequency spectrum features of the main speaker into time-domain audio data, and replacing the mixed audio data in the video stream with the time-domain audio data.

[0083] For example, when converting mixed audio data into time-frequency spectrum features, a deep neural network (DNN) model can be constructed to predict the time-frequency mask. The speaker's voiceprint features are then fused with the time-frequency spectrum features. This fusion can be done by concatenating the speaker's voiceprint features with the time-frequency spectrum features as input to the DNN model; using the speaker's voiceprint features as query vectors in the attention mechanism to weight the features in the DNN model; or using the speaker's voiceprint features to adjust the parameters of the batch normalization (BN) layer in the DNN model.

[0084] For example, the DNN model takes fused features as input and outputs a time-frequency mask, where each element represents the probability that the corresponding time-frequency unit belongs to the main speaker. The time-frequency mask is then multiplied by the time-frequency spectral features to separate the mixed audio data and obtain the time-frequency spectral features of the separated main speaker. The inverse short-time Fourier transform (ISTFT) can then be used to convert the time-frequency spectral features of the main speaker into time-domain audio data. During the ISTFT process, the phase spectrum of the mixed speech can be used, or the Griffin-Lim algorithm can be used to reconstruct the phase spectrum.

[0085] Through the processing method provided by the embodiments of the present disclosure, sound denoising is performed using voiceprint features, which can accurately separate the audio data of the main speaker, perform targeted repair based on the vocal characteristics of the main speaker, retain more details, reduce voice distortion, achieve more accurate and natural noise suppression, and adapt to complex sound environments.

[0086] In some embodiments, the target data in the above operation S230 includes target image data; suppressing the target data of the target object in the video stream includes: processing the target image data in the video stream, and the processed video stream does not include the target image data.

[0087] Through the processing method provided by the embodiment of the present disclosure, after locating the target object that performs the target behavior, the target image data of the target object in the video stream can be removed, thereby removing the interfering image in the video conference, enhancing the image clarity, purifying the shared screen, improving the quality of the video conference, and improving the user experience.

[0088] Furthermore, the above-mentioned processing of the target image data in the video stream includes: cutting the video frames in the video stream.

[0089] For example, the video stream of a video conference can be read frame by frame in the time dimension, and each read video frame can be segmented. The image data of the main speaker and the target image data of the target object can be distinguished according to the segmentation results, and then only the target image data in each frame image can be removed.

[0090] Through the processing method provided by the embodiments of the present disclosure, the visual content in the video stream can be finely reconstructed from the time dimension, the information expression of the speaker can be enhanced, visual interference can be reduced, the video quality can be improved, and the user experience can be enhanced.

[0091] In some embodiments, the video stream is obtained by the target device collecting data in a first direction. The processing method also includes: in response to the occurrence of the target behavior, controlling the collection direction of the target device to be adjusted from the first direction to the second direction, and the collection area corresponding to the target device in the second direction is an area outside the area where the target object is located.

[0092] For example, a video stream for a video conference can be obtained by a target device pre-configured by the processing device capturing multiple participants in real time along a first direction. Based on this, if it is determined that the target behavior has occurred during the video conference, the target device's capture direction can be adjusted from the first direction to a second direction, so that the target device's capture area after the adjustment is outside the area where the target object is located, that is, the target device no longer captures the target object.

[0093] For example, the target device can be a camera, a camera, a drone, etc. Taking the target device as an example, when it is determined that the target behavior occurs during the video conference, the processing device can control the camera to automatically move, so that the camera's collection area after movement deviates from the area where the target object is located, and the target object performing the target behavior is no longer captured.

[0094] Through the processing method provided by the embodiment of the present disclosure, when the target object performing the target behavior in the video stream is identified, some interfering images in the video conference can be removed by switching the dynamic picture of the target object, thereby improving the video quality, enhancing the image information, and improving the professionalism, privacy and communication efficiency of the video conference.

[0095] In some embodiments, the processing method further includes: obtaining audio information corresponding to the discussion behavior; converting the audio information into text information; and outputting the text information.

[0096] For example, if a discussion activity is detected in a video conference, audio information corresponding to the discussion activity is extracted from the video stream. The audio information may be a voice, music, or ambient sound.

[0097] The embodiments of the present disclosure do not specifically limit the method for converting audio information into text information. For example, the method may be to use a voice writing tool, manual input, or use OCR (Optical Character Recognition), speech recognition software or applications, etc.

[0098] After the text information is converted, it can be output. The presently disclosed embodiments do not limit the output terminal of the text information. For example, the text information can be output to the device of the aforementioned speaker to remind the speaker that some participants are currently discussing the video conference, thereby prompting the speaker to adjust their speaking strategy or wait for the discussion to end.

[0099] For another example, the text information may be output to a display screen of the processing device to provide a visual warning to the target object that has engaged in the discussion behavior, or to alert all participating objects to the bad behavior.

[0100] For another example, the text information can be output to other devices in communication with the processing apparatus, so that the other devices can effectively monitor or manage the entire video conference and evaluate the conference quality and user experience.

[0101] Through the processing method provided by the embodiment of the present disclosure, when the target object having discussion behavior in the video stream is identified, the audio information corresponding to the discussion behavior is obtained, and text information is output through voice recognition. The text information can be used for message notification reminders later, so that each participating object maintains good meeting behavior, keeps the focus of the video conference, and improves meeting efficiency and participation experience.

[0102] The above description is merely an example, and the embodiments of the present disclosure are not limited thereto. In other embodiments, the processing method further includes: using a multimodal large model to separately identify the first position information of the speaker and the second position information of the target object from the video stream; enhancing the audio data corresponding to the first position information in the video stream, and weakening the audio data corresponding to the second position information in the video stream.

[0103] For example, the aforementioned large multimodal model can be used to identify the directions of the speaker and the discussants. Then, beamforming technology can be used to enhance the sound signal from the speaker's direction while suppressing noise from the discussants' direction. Beamforming coordinates the signals received or transmitted by multiple sensors (such as microphone arrays or speaker arrays) to enhance sound signals from a specific direction and suppress interference from other directions.

[0104] Based on the above processing method, the present disclosure also provides a processing device, which will be combined with Figure 5 The device is described in detail.

[0105] Figure 5 The block diagram schematically shows a processing device according to an embodiment of the present disclosure.

[0106] like Figure 5 As shown, the processing device 500 according to this embodiment includes a video stream acquisition module 510 , a target behavior detection module 520 and a suppression processing module 530 .

[0107] The video stream acquisition module 510 is used to acquire the video stream of the video conference. The video stream is obtained by shooting multiple participants in the video conference in the same physical space.

[0108] The target behavior detection module 520 is used to detect the video stream and determine whether a target behavior occurs, wherein the target behavior represents the discussion behavior between participants in the video conference.

[0109] The suppression processing module 530 is used to determine the target object that performs the target behavior from multiple participating objects, and perform suppression processing on the target data of the target object in the video stream.

[0110] It should be noted that the embodiment of the device part is similar to the embodiment of the method part, and the technical effects achieved are also similar. For specific details, please refer to the above-mentioned method embodiment part, which will not be repeated here.

[0111] According to embodiments of the present disclosure, any multiple of the video stream acquisition module 510, target behavior detection module 520, and suppression processing module 530 can be implemented in a single module, or any one of these modules can be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules can be combined with at least part of the functionality of other modules and implemented in a single module. According to embodiments of the present disclosure, at least one of the video stream acquisition module 510, target behavior detection module 520, and suppression processing module 530 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or can be implemented in hardware or firmware through any other reasonable means of circuit integration or packaging, or can be implemented in any one of the three implementation methods of software, hardware, and firmware, or any suitable combination of any of these. Alternatively, at least one of the video stream acquisition module 510, target behavior detection module 520, and suppression processing module 530 can be at least partially implemented as a computer program module that, when executed, performs the corresponding functionality.

[0112] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions configured to implement the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of the boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0113] Those skilled in the art will appreciate that the features described in the various embodiments of the present disclosure may be combined and / or coupled in various ways, even if such combinations or couplings are not explicitly described in the present disclosure. In particular, the features described in the various embodiments of the present disclosure may be combined and / or coupled in various ways without departing from the spirit and teachings of the present disclosure. All such combinations and / or couplings fall within the scope of the present disclosure.

[0114] The above describes the embodiments of the present disclosure. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although each embodiment has been described separately above, this does not mean that the measures in each embodiment cannot be advantageously used in combination. Without departing from the scope of the present disclosure, those skilled in the art may make various substitutions and modifications, which should all fall within the scope of the present disclosure.

Claims

1. A processing method comprising: Acquire a video stream of a video conference, where the video stream is obtained by filming multiple participants of the video conference in the same physical space; Detecting the video stream to determine whether a target behavior occurs, wherein the target behavior represents a discussion behavior between participants in the video conference; A target object that has the target behavior is determined from the multiple participating objects, and target data of the target object in the video stream is suppressed.

2. The method according to claim 1, wherein detecting the video stream to determine whether the target behavior occurs comprises: extracting image features and audio features from the video stream; Inputting the image features, the audio features, and the target prompt word into a target model to obtain an output result of the target model, wherein the target prompt word is used to guide the target model to perform reasoning; The target behavior in the video stream is determined according to the output result.

3. The method according to claim 1, wherein the target data includes target audio data; and the step of suppressing the target data of the target object in the video stream comprises: The target audio data in the video stream is processed, and in the video stream after the processing, the volume corresponding to the target audio data is smaller than the volume corresponding to the target audio data before the processing.

4. The method according to claim 3, wherein processing the target audio data in the video stream comprises: Detecting the video stream to determine a main speaker among the multiple participants, wherein the main speaker represents a participant who is speaking in the video conference; Obtaining the voiceprint characteristics of the speaker; The target audio data is processed using the voiceprint feature.

5. The method according to claim 4, wherein processing the target audio data using the voiceprint feature comprises: Extracting mixed audio data from the video stream, and converting the mixed audio data into time-frequency spectrum features, wherein the mixed audio data includes the target audio data and audio data of a main speaker; fusing the speaker's voiceprint features with the time-frequency spectrum features to obtain fused features; Processing the fused features using a deep neural network model to obtain a time-frequency mask having a plurality of time-frequency units, wherein each of the time-frequency units represents a probability that the time-frequency unit is dominated by the main speaker; Multiplying the time-frequency mask by the time-frequency spectrum feature to obtain the time-frequency spectrum feature of the speaker; The time-frequency spectrum characteristics of the main speaker are converted into time-domain audio data, and the mixed audio data in the video stream is replaced with the time-domain audio data.

6. The method according to claim 1, wherein the target data comprises target image data; and the step of suppressing the target data of the target object in the video stream comprises: The target image data in the video stream is processed, and the processed video stream does not include the target image data.

7. The method according to claim 6, wherein processing the target image data in the video stream comprises: The video frames in the video stream are cut.

8. The method according to claim 1, wherein the video stream is obtained by capturing the first direction by the target device, and the method further comprises: In response to the target behavior occurring, the collection direction of the target device is controlled to be adjusted from the first direction to a second direction, and the collection area corresponding to the target device in the second direction is an area outside the area where the target object is located.

9. The method according to claim 1, further comprising: Obtaining audio information corresponding to the discussion behavior; Converting the audio information into text information; The text information is output.

10. A processing device comprising: A video stream acquisition module is used to acquire a video stream of a video conference, wherein the video stream is obtained by shooting multiple participants of the video conference in the same physical space; a target behavior detection module, configured to detect the video stream and determine whether a target behavior occurs, wherein the target behavior represents a discussion behavior between participants in the video conference; The suppression processing module is used to determine the target object that has performed the target behavior from the multiple participating objects, and perform suppression processing on the target data of the target object in the video stream.