Automatic conference grouping method, electronic equipment and computer readable storage medium

By acquiring the audio frames and status of participants in online meetings and using a large model to automatically group them, the problem of low efficiency in manual grouping in online meetings is solved, and efficient and accurate grouping is achieved.

CN121125373APending Publication Date: 2025-12-12ZHEJIANG DAHUA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510933701.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-07
Publication Date
2025-12-12

AI Technical Summary

Technical Problem

In existing technologies, the grouping of participants in online meetings is inefficient, and manual grouping is usually used, which is also inefficient.

Method used

By acquiring the target audio frames of the participants, their participation status and group status are determined. Using a large model combined with text information, historical group information, and conflicting groups of objects, the participants are automatically grouped and the target meeting group is output.

Benefits of technology

It achieves high efficiency and accuracy in automatic grouping, thus improving grouping efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121125373A_ABST
    Figure CN121125373A_ABST
Patent Text Reader

Abstract

The invention discloses an automatic conference grouping method, electronic equipment and a computer readable storage medium. According to the application, the conference participation state of each participating object is determined through the obtained target audio frame of each participating object; determining at least two to-be-grouped objects from the participating objects according to the participating state of each participating object and the grouping state of each participating object; and then text information represented by collected target audio data of to-be-grouped objects, historical grouping information of the to-be-grouped objects and the obtained object conflict group are input into the large model to obtain a target conference group output by the large model. Therefore, the text information, the historical grouping information and the object conflict group are input into the large model, and the target conference group is output by using the large model, so that the to-be-grouped objects are automatically grouped, and the grouping efficiency and accuracy are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of speech processing, and particularly relates to a conference automatic grouping method, an electronic device and a computer readable storage medium. BACKGROUND

[0002] Online conference is a core tool for modern work, education and socialization. To realize the video and audio communication of each participant in the online conference, a multipoint control unit (MCU) is needed for mixing and forwarding the stream. Each participant in the conference establishes a connection with the center unit, uploads multimedia stream through the connection, and downloads the mixed stream of other conference participants to realize the video and audio communication in the online conference.

[0003] When the number of conference participants is large or there are more than one conference topics, it is necessary to group the participants in the current conference to enable separate discussion within the group. Currently, the participants in the online conference are usually grouped manually, which is low in efficiency. SUMMARY

[0004] The technical problem solved by the present application is to provide a conference automatic grouping method, an electronic device and a computer readable storage medium, which can improve the grouping efficiency.

[0005] To solve the above technical problem, the present application provides a conference automatic grouping method, comprising: acquiring target audio frames of each participant in the current conference; determining the participation state of each participant according to the target audio frames of each participant; determining at least two grouping objects from each participant according to the participation state of each participant and the grouping state of each participant; inputting the text information represented by the collected target audio data of the grouping objects, the historical grouping information of each grouping object and the acquired object conflict group into a large model to obtain the target conference group output by the large model.

[0006] In one embodiment, the step of inputting the text information representing the target audio data of the objects to be grouped, the historical grouping information of each object to be grouped, and the acquired object conflict group into a large model to obtain the target conference group output by the large model includes: initially grouping each object to be grouped using the text information corresponding to each object to be grouped through the large model to obtain multiple initial conference groups, wherein the initial conference groups include already grouped objects; performing group verification processing on the already grouped objects according to the conference theme of the initial conference group to which each already grouped object belongs and the historical conference theme represented by the historical grouping information of the already grouped objects to obtain a first verification result; in response to the first verification result indicating that the already grouped objects are correctly grouped, verifying whether there are any objects conflicting with the already grouped objects in the initial conference group to which the already grouped objects belong, to obtain a second verification result; in response to the second verification result indicating that there are no objects conflicting with the already grouped objects in the initial conference group to which the already grouped objects belong, verifying the number of objects in the initial conference group to which the already grouped objects belong, to obtain a third verification result; in response to the third verification result indicating that the verification is passed, determining the initial conference group as the target conference group.

[0007] In one embodiment, the step of performing group verification processing on the grouped objects based on the meeting topic of the initial meeting group to which each grouped object belongs and the historical meeting topic represented by the historical grouping information of the grouped objects to obtain a first verification result includes: obtaining the similarity between the meeting topic of the initial meeting group to which each grouped object belongs and the historical meeting topic represented by the historical grouping information of the grouped objects; in response to the similarity being greater than or equal to a preset similarity threshold, obtaining a first verification result indicating that the grouped objects are correctly grouped; in response to the similarity being less than the preset similarity threshold, obtaining the relevance between the text information corresponding to each grouped object and the meeting topic of the initial meeting group to which the grouped object belongs; in response to the relevance being greater than or equal to a preset relevance threshold, obtaining a first verification result indicating that the grouped objects are correctly grouped; and in response to the relevance being less than the preset relevance threshold, obtaining a first verification result indicating that the objects to be grouped are incorrectly grouped.

[0008] In one embodiment, the step of verifying whether there is an object conflicting with the grouped object in the initial meeting group of the grouped object based on the object conflict group, and obtaining a second verification result, includes: determining whether the grouped object is the same as a conflicting object in the object conflict group; if so, identifying the conflicting object other than the grouped object in the corresponding object conflict group as the target conflicting object; determining whether there is a grouped object in the initial meeting group of the grouped object that is the same as the target conflicting object; if not, obtaining a second verification result indicating that there is no object conflicting with the grouped object in the initial meeting group of the grouped object; if so, obtaining a second verification result indicating that there is an object conflicting with the grouped object in the initial meeting group of the grouped object.

[0009] In one embodiment, the step of obtaining the target audio frame of each participant in the current meeting includes: obtaining the initial audio of each participant in the current meeting; performing periodic sampling processing on the initial audio to obtain multiple initial audio frames; and filtering each initial audio frame according to the audio decibel of each obtained initial audio frame to obtain the target audio frame.

[0010] In one embodiment, before the step of inputting the text information representing the target audio data of the objects to be grouped, the historical grouping information of each object to be grouped, and the obtained object conflict group into the large model to obtain the target conference group output by the large model, the method further includes: obtaining a first ratio between the number of target audio frames of each object to be grouped within the preset time period and the number of initial audio frames within the preset time period; in response to the existence of at least two objects to be grouped whose first ratio is greater than a first preset value, the corresponding objects to be grouped are determined as conflicting objects in the object conflict group.

[0011] In one embodiment, the step of determining the participation status of each participant based on the target audio frames of each participant includes: obtaining a second ratio between the number of target audio frames of each participant and the acquisition duration of the initial audio in which the target audio frames are located; determining that the participation status of the participant is active if the second ratio is greater than or equal to a second preset value; and determining that the participation status of the participant is silent if the second ratio is less than the second preset value.

[0012] In one embodiment, after the step of inputting the text information representing the target audio data of the objects to be grouped, the historical grouping information of each object to be grouped, and the obtained object conflict group into the large model to obtain the target conference group output by the large model, the method further includes: traversing each target conference group, selecting a target sharing object from the grouped objects of the currently traversed target conference group; mixing the audio of the other grouped objects in the currently traversed target conference group except for the target sharing object to obtain mixed audio; and sending the mixed audio and the obtained video of the other grouped objects to the target sharing object.

[0013] To address the aforementioned technical problems, this application provides an electronic device, including a memory and a processor, wherein the memory stores program instructions, and the processor retrieves the program instructions from the memory to execute the aforementioned automatic conference grouping method.

[0014] To address the aforementioned technical problems, this application provides a computer-readable storage medium, comprising: storing program data, which, when executed by a processor, is used to implement the aforementioned automatic conference grouping method.

[0015] The above scheme acquires the target audio frames of each participant in the current meeting; determines the participation status of each participant based on the target audio frames; identifies at least two participants to be grouped based on their participation status and grouping status; and then inputs the text information represented by the target audio data of the participants to be grouped, the historical grouping information of each participant, and the acquired object conflict groups into a large model to obtain the target meeting group output by the large model. Thus, by inputting text information, historical grouping information, and object conflict groups into the large model, and using the large model to output the target meeting group, automatic grouping of participants to be grouped is achieved, improving grouping efficiency. At the same time, the accuracy of the large model's grouping is also improved. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly described below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort, wherein:

[0017] Figure 1 This is a flowchart illustrating an exemplary embodiment of the automatic meeting grouping method shown in this application;

[0018] Figure 2 yes Figure 1A flowchart illustrating an exemplary embodiment of step S140 in the automatic meeting grouping method is shown.

[0019] Figure 3 This is a schematic diagram illustrating an exemplary embodiment of the process of acquiring a target conference group using a large model, as shown in this application;

[0020] Figure 4 This is an application diagram illustrating an exemplary embodiment of the automatic meeting grouping method shown in this application;

[0021] Figure 5 This is a block diagram illustrating an automatic conference grouping device according to an exemplary embodiment of this application;

[0022] Figure 6 This is a schematic diagram of the structure of an embodiment of the electronic device provided in this application;

[0023] Figure 7 This is a schematic diagram of an embodiment of the computer-readable storage medium provided in this application. Detailed Implementation

[0024] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It is understood that the specific embodiments described herein are only for explaining this application and not for limiting it. Furthermore, it should be noted that, for ease of description, only the parts related to this application are shown in the accompanying drawings, not all structures. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0025] First, it's important to note that online meetings are a core tool for modern work, education, and social interaction. To enable video and audio communication among participants in an online meeting, a multi-device control unit (MCU) is needed for mixing and forwarding the streams. Each participant establishes a connection with the central unit, uploading multimedia streams and downloading the mixed streams from other participants to facilitate video and audio communication. When there are many participants or multiple topics, it may be necessary to group participants for individual discussions within smaller groups. Currently, manual grouping is commonly used, which is inefficient.

[0026] Based on this, this application provides a method for automatic conference grouping, an electronic device, and a computer-readable storage medium. For details, please refer to [link / reference needed]. Figure 1 , Figure 1 This is a flowchart illustrating an exemplary embodiment of a meeting automatic grouping method shown in this application.

[0027] The executing entity of an automatic conference grouping method can be a terminal device, a server, or other processing device. The terminal device can be a computer, mobile device, terminal, computing device, vehicle-mounted device, etc. The executing entity of the automatic conference grouping method can also be a conference automatic grouping device. In some possible implementations, the automatic conference grouping method can be implemented by a processor calling computer-readable instructions stored in memory. The executing entity of the automatic conference grouping method can also be a big data cluster. A big data cluster is a computer system architecture formed by multiple computers connected through a network. The big data cluster can be deployed on a private cloud built with K8S (Kubernetes, a container orchestration engine).

[0028] Specifically, one method for automatically grouping meetings according to this embodiment includes the following steps:

[0029] Step S110: Obtain the target audio frames of each participant in the current meeting.

[0030] The participants refer to those who attend the meeting.

[0031] The current meeting refers to the meeting that is currently in progress.

[0032] The target audio frame refers to the audio frame selected from the initial audio frames acquired.

[0033] The automatic conference grouping device acquires the target audio frames for each participant in the current conference. Specifically, the automatic conference grouping device samples the audio of each participant in the current conference at intervals to obtain initial audio frames; it then preprocesses the initial audio frames to obtain the target audio frames. For example, the automatic conference grouping device can preprocess the initial audio frames by performing speech signal compensation on the initial audio frames using a preset filter to obtain the target audio frames.

[0034] Step S120: Determine the participation status of each participant based on the target audio frame of each participant.

[0035] Participation status refers to the state of a participant in the current meeting. Participation status can be divided into active and inactive status. An active status indicates that the participant frequently speaks in the current meeting, while an inactive status indicates that the participant does not speak or speaks infrequently.

[0036] The automatic conference grouping device determines the participation status of each participant based on the target audio frames of each participant. Specifically, if the data of the target audio frames of each participant within a preset time period is greater than or equal to a third preset value, the automatic conference grouping device determines the active state as the participation status of the corresponding participant; if the data of the target audio frames of each participant within a preset time period is less than the third preset value, the automatic conference grouping device determines the inactive state as the participation status of the corresponding participant.

[0037] Step S130: Determine at least two objects to be grouped from each participant based on their participation status and grouping status.

[0038] Group status refers to the state of a participant's presence in a particular meeting group during the current meeting. Group status can include both grouped and ungrouped states.

[0039] The automatic meeting grouping device determines at least two participants to be grouped based on their participation status and grouping status. As an example, the device determines a participant as a to-be-grouped participant if the participant's participation status is active and their grouping status is ungrouped. As another example, the device determines a participant as a to-be-grouped participant if the number of participants with active participation status and ungrouped status is greater than or equal to a fourth preset value. The fourth preset value can be 2 or 6.

[0040] In one embodiment, the automatic meeting grouping device sequentially determines whether the participation status of each participant is active. If so, it further determines whether the grouping status of the corresponding participant is ungrouped. If so, the corresponding participant is identified as a participant to be grouped. Alternatively, the automatic meeting grouping device sequentially determines whether the grouping status of each participant is ungrouped. If so, it further determines whether the participation status of the corresponding participant is active. If the participation status of a participant is active, the corresponding participant is identified as a participant to be grouped.

[0041] Step S140: Input the text information representing the target audio data of the objects to be grouped, the historical grouping information of each object to be grouped, and the obtained object conflict groups into the large model to obtain the target conference group output by the large model.

[0042] Target audio data refers to the audio data of the objects to be grouped in the current meeting. As one example, the automatic meeting grouping device collects the audio data of each object to be grouped within a preset time period to obtain the target audio data. As another example, the automatic meeting grouping device collects the audio data of each object to be grouped in the current meeting at preset time intervals to obtain the target audio data. As yet another example, the automatic meeting grouping device collects the audio data of each object to be grouped in the current meeting at preset time intervals to obtain initial audio data; noise is then removed from the initial audio data to obtain the target audio data. For example, the automatic meeting grouping device acquires the initial audio data according to a preset semantic observation time sliding window, and filters out audio frequencies below a first preset decibel and above a second preset decibel from the initial audio data to obtain the target audio data, where the first preset decibel is less than the second preset decibel. The preset semantic observation time sliding window can be 3 seconds. Under normal speaking speed, it is usually 120-160 words / minute. Therefore, a medium-length sentence, such as 15 words, takes about 3 seconds under normal speaking speed. So, setting it to collect initial audio data every 3 seconds to obtain text information can more comprehensively obtain the text information of each object to be grouped.

[0043] Text information refers to information encoded in written language forms such as characters, symbols, and numbers. Specifically, an automatic conference grouping device performs text conversion processing on target audio data to obtain text information. For example, the automatic conference grouping device can use speech recognition technology to convert target audio data into text information. In one embodiment, the text information can be a text vector, such as [text1, text2, ..., textn]. The automatic conference grouping device collects audio from each target group at preset time intervals to obtain at least one target audio data; it then performs speech recognition on each target audio data to obtain multiple texts in the text vector; in response to no target audio data being collected, the text in the text vector is empty, meaning the corresponding target group has not spoken.

[0044] Historical grouping information refers to the grouping information of objects to be grouped in historical meetings. Historical grouping information may include the objects to be grouped, the historical meeting groups to which the objects belong, and the meeting topics of the historical meeting groups. Specifically, a preset storage module stores the historical grouping information of each object to be grouped, and the automatic meeting grouping device identifies the grouping information received from the preset storage module as historical grouping information.

[0045] A conflict group is a group that includes conflicting objects. A conflicting object is an object to be grouped that has audio conflicts with other objects to be grouped. Conflicting objects in a conflict group have target audio frames within a preset time period. The number of conflicting objects in a conflict group can be 0, 2, or more than 3. Specifically, the automatic conference grouping device samples the initial audio of each object to be grouped at intervals within the preset time period to obtain initial audio frames; it filters the initial audio frames to obtain the target audio frames for each object to be grouped within the preset time period; it determines the participation status of objects to be grouped whose number of target audio frames is greater than the preset number of audio frames as speaking status; in response to the number of objects to be grouped in speaking status being greater than or equal to a first preset number of objects, the corresponding objects to be grouped are determined as conflicting objects, and the conflicting objects are grouped into a conflicting object group. The preset time period can be 2 seconds.

[0046] In one embodiment, the preset number of audio frames can be 1, and the first preset number of objects can be 2. The automatic conference grouping device samples the initial audio of each object to be grouped within a preset time period at intervals to obtain initial audio frames and acquires the audio decibel of each initial audio frame. It determines whether the audio decibel of each initial audio frame is greater than the first preset decibel and less than the second preset decibel. If so, the corresponding initial audio frame is determined as the target audio frame. It determines whether the number of target audio frames of each object to be grouped within the preset time period is greater than 1. If so, the participation status of the corresponding object to be grouped is determined to be speaking. Otherwise, the participation status of the corresponding object to be grouped is determined to be silent. It determines whether the number of objects to be grouped in speaking state is greater than or equal to 2. If so, the objects to be grouped in speaking state are determined as conflict objects, and the conflict objects are combined into a conflict object group.

[0047] The large model refers to a general-purpose artificial intelligence model trained based on deep learning. The large model includes multiple pre-set validation conditions and prompts. Validation conditions can include: the meeting topic of the group containing the grouped object being consistent with historical topics; or the relevance between the meeting topic of the group containing the group and the meeting topic represented by the text information of the grouped object being greater than a pre-set relevance threshold. Another validation condition is that there are no conflicts among the grouped objects within the same meeting group. A further validation condition is that the number of grouped objects in the meeting group is greater than a second pre-set number of objects, where the second pre-set number of objects can be one.

[0048] The prompts in the large model are text information, historical grouping information, and object conflict groups. For example, historical grouping information, LastThemeGroup_V = {{{LastThemeGroup_V}}}, represents the historical meeting groups to which each object to be grouped belongs and the meeting topics of those groups. If the historical meeting groups are empty, it indicates that the objects to be grouped are in their first grouping. For example, text information, AudioText_V = {{{AudioText_V}}}, stores the text of each object to be grouped [text1, text2, ..., textn] from beginning to end. If an object to be grouped has not spoken recently, the text information is empty. The large model's prompts also include: Please group the objects to be grouped.

[0049] A target meeting group is a meeting group obtained by grouping the objects to be grouped. A target meeting group includes multiple grouped objects.

[0050] The automatic conference grouping device inputs the text information representing the target audio data of the objects to be grouped, the historical grouping information of each object, and the acquired object conflict groups into a large model to obtain the target conference groups output by the large model. Thus, by leveraging the natural language generalization capability of the large model, high-accuracy automatic conference grouping can be achieved without re-pre-training the algorithm model or performing other semantic analysis and clustering algorithms. Only text information, historical grouping information, object conflict groups, and preset verification conditions are required for feedback.

[0051] As can be seen, by acquiring the target audio frames of each participant in the current meeting; determining the participation status of each participant based on the target audio frames; identifying at least two participants to be grouped based on their participation status and grouping status; and then inputting the text information represented by the target audio data of the participants to be grouped, the historical grouping information of each participant, and the acquired object conflict groups into the large model, the target meeting group output by the large model is obtained. Therefore, by inputting text information, historical grouping information, and object conflict groups into the large model, and using the large model to output the target meeting group, automatic grouping of participants to be grouped is achieved, improving grouping efficiency.

[0052] Based on the above embodiments, please refer to Figure 2 , Figure 2 yes Figure 1 The illustrated flowchart shows an exemplary embodiment of step S140 in the automatic conference grouping method. Specifically, step S140 includes the following steps in which the text information representing the target audio data of the objects to be grouped, the historical grouping information of each object to be grouped, and the obtained object conflict groups are input into the large model to obtain the target conference group output by the large model:

[0053] Step S210: Initially group each object to be grouped using the text information corresponding to each object in the large model to obtain multiple initial meeting groups, which include the already grouped objects.

[0054] The initial meeting group refers to the group obtained after the first grouping of the objects to be grouped.

[0055] Grouped objects refer to the objects in the meeting group obtained after grouping the objects to be grouped according to the grouping results of the large model.

[0056] The automatic meeting grouping device uses a large model to initially group objects based on their corresponding text information, resulting in multiple initial meeting groups. As one example, the device performs text analysis on the text information of each object to be grouped, obtaining the text topic; objects with the same text topic are identified as already grouped objects in the same initial meeting group, and their corresponding text topics are determined as the meeting topic of the initial meeting group. As another example, the device extracts keywords from the text information of each object to be grouped, obtaining the keywords for each object; it then determines the relevance between the keywords and preset meeting topics, identifying the preset meeting topic with the highest relevance as the meeting topic for the corresponding object; and objects with the same meeting topic are identified as already grouped objects in the same initial meeting group.

[0057] Step S220: Perform group verification processing on the grouped objects according to the meeting topic of the initial meeting group to which each grouped object belongs and the historical meeting topic represented by the historical grouping information of the grouped objects, and obtain the first verification result.

[0058] As an example, the automatic meeting grouping device obtains the similarity between the meeting topic of the initial meeting group to which each grouped object belongs and the historical meeting topic represented by the historical grouping information of the grouped object; in response to the similarity being greater than or equal to a preset similarity threshold, a first verification result indicating that the grouped object is correctly grouped is obtained; in response to the similarity being less than the preset similarity threshold, a first verification result indicating that the grouped object is incorrectly grouped is obtained.

[0059] As another example, the automatic meeting grouping device obtains the similarity between the meeting topic of the initial meeting group to which each grouped object belongs and the historical meeting topic represented by the historical grouping information of the grouped object; in response to the similarity being greater than or equal to a preset similarity threshold, a first verification result indicating that the grouped objects are correctly grouped is obtained; in response to the similarity being less than the preset similarity threshold, the device obtains the relevance between the text information corresponding to each grouped object and the meeting topic of the initial meeting group to which the grouped object belongs; in response to the relevance being greater than or equal to a preset relevance threshold, a first verification result indicating that the grouped objects are correctly grouped is obtained; in response to the relevance being less than the preset relevance threshold, a first verification result indicating that the object to be grouped is incorrectly grouped is obtained.

[0060] The automatic meeting grouping device obtains the relevance between the text information corresponding to each grouped object and the meeting topic of the initial meeting group to which the grouped object belongs. Specifically, the automatic meeting grouping device extracts keywords from the text information corresponding to each grouped object, obtains the keywords corresponding to each grouped object, calculates the relevance between the keywords corresponding to each grouped object and the meeting topic of the initial meeting group to which the grouped object belongs, and obtains the relevance between the text information corresponding to each grouped object and the meeting topic of the initial meeting group to which the grouped object belongs.

[0061] Furthermore, in response to the first verification result indicating that the grouped objects were grouped incorrectly, the automatic conference grouping device returns to step S210 to perform initial grouping of each object to be grouped using the text information corresponding to each object to be grouped through the large model, thereby obtaining multiple initial conference groups.

[0062] Step S230: In response to the first verification result indicating that the grouped objects are correctly grouped, verify whether there are any objects in the initial meeting group where the grouped objects are located that conflict with the grouped objects, and obtain the second verification result.

[0063] As an example, the automatic meeting grouping device determines whether the conflicting object in the object conflict group is the same as the grouped object in the same initial meeting group. If so, it obtains a second verification result indicating that there is an object in the initial meeting group where the grouped object is located that conflicts with the grouped object; otherwise, it obtains a second verification result indicating that there is no object in the initial meeting group where the grouped object is located that conflicts with the grouped object.

[0064] As another example, the automatic meeting grouping device determines whether the grouped object is the same as a conflicting object in the object conflict group; if so, it identifies the conflicting object in the corresponding object conflict group other than the grouped object as the target conflicting object; it determines whether there is a grouped object in the initial meeting group where the grouped object is located that is the same as the target conflicting object; if not, it obtains a second verification result indicating that there is no object in the initial meeting group where the grouped object is located that conflicts with the grouped object; if so, it obtains a second verification result indicating that there is an object in the initial meeting group where the grouped object is located that conflicts with the grouped object.

[0065] Furthermore, in response to the second verification result indicating that there are objects in the initial meeting group where the grouped objects are located that conflict with the grouped objects, the automatic meeting grouping device returns to step S210 to perform initial grouping of each object to be grouped using the large model with the text information corresponding to each object to be grouped, thereby obtaining multiple initial meeting groups.

[0066] Step S240: In response to the second verification result indicating that there are no objects conflicting with the grouped objects in the initial meeting group where the grouped objects are located, the number of objects in the initial meeting group where the grouped objects are located is verified to obtain the third verification result.

[0067] Specifically, the automatic meeting grouping device counts the number of objects in the initial meeting group to which each grouped object belongs. If the number of objects in the initial meeting group to which each grouped object belongs is greater than or equal to the second preset number of objects, a third verification result of successful verification is obtained; if the number of objects in the initial meeting group to which each grouped object belongs is less than the second preset number of objects, a third verification result of failed verification is obtained.

[0068] Furthermore, in response to the third verification result indicating verification failure, the automatic meeting grouping device returns to step S210 to perform initial grouping of each object to be grouped using the text information corresponding to each object to be grouped through the large model, thereby obtaining multiple initial meeting groups.

[0069] Step S250: In response to the third verification result indicating that the verification is passed, the initial meeting group is determined as the target meeting group.

[0070] If the automatic meeting grouping device responds to the third verification result indicating that the verification is passed, it will determine the initial meeting group as the target meeting group and output the target meeting group through the large model.

[0071] In one embodiment, combined with Figure 3As shown, historical grouping information, text information, and object conflict groups are input into the large model. The large model performs initial grouping to obtain the initial meeting group. The first verification condition checks whether the meeting topic of the initial meeting group is consistent with the historical meeting topic. If not, the initial meeting group needs adjustment, and the system returns to the large model, reporting that the verification failed and triggering the large model to regroup. If yes, it indicates that the meeting topics of the grouped objects in the initial meeting group maintain consistency. The second verification condition checks whether the initial meeting group avoids object conflict groups. If not, the initial meeting group needs adjustment, and the system returns to the large model, reporting that the initial meeting group did not avoid object conflict groups and triggering the large model to regroup. If yes, it indicates that there are no conflicts among the grouped objects in the initial meeting group. The third verification condition checks whether the number of grouped objects in the initial meeting group is greater than or equal to the second preset value. If not, the initial meeting group needs adjustment, and the system returns to the large model, reporting that the number of objects in the initial meeting group is insufficient and triggering the large model to regroup. If yes, the initial meeting group meets the conditions and is output as the target meeting group.

[0072] The detailed steps for the large model to verify whether the initial meeting group maintains the theme inertia according to the first verification condition can be found in step S220, and will not be repeated here; the detailed steps for the large model to verify whether the initial meeting group avoids object conflict groups according to the second verification condition can be found in step S230, and will not be repeated here; the detailed steps for the large model to verify whether the number of grouped objects in the initial meeting group is greater than or equal to the second preset number of objects according to the third verification condition can be found in step S230, and will not be repeated here.

[0073] The steps of the automatic conference grouping device to obtain the target audio frame of each participant in the current conference include: obtaining the initial audio of each participant in the current conference; performing periodic sampling processing on the initial audio to obtain multiple initial audio frames; and filtering each initial audio frame according to the audio decibel of each obtained initial audio frame to obtain the target audio frame.

[0074] The initial audio frame refers to the audio frame sampled from the initial audio.

[0075] The automatic conference grouping device performs periodic sampling processing on the initial audio to obtain multiple initial audio frames. Specifically, the automatic conference grouping device packages the initial audio according to a preset packaging period to obtain multiple initial audio frames. In one embodiment, the preset packaging period is 40ms, and the automatic conference grouping device packages the 1-second initial audio into 40ms audio frames, resulting in 1s / 40ms = 25 audio frames.

[0076] Audio decibels refer to the volume decibel value of audio. Specifically, the automatic conference grouping device collects the volume decibel value of each initial audio frame according to a preset audio activity detection cycle, and stores the volume decibel value in the extension header of the corresponding initial audio frame to obtain the audio decibel value of the initial audio frame.

[0077] The automatic conference grouping device filters each initial audio frame based on its audio decibel level to obtain the target audio frame. Specifically, the device removes initial audio frames with an audio decibel level lower than a first preset decibel level and those with an audio decibel level higher than a second preset decibel level, thus obtaining the target audio frame. Filtering the initial audio frames removes noise, which helps improve the quality of the target audio frame and reduce interference.

[0078] In one embodiment, the first preset decibel level can be 10 dB, and the second preset decibel level can be 85 dB. The automatic conference grouping device inputs each initial audio frame into a filter, filters out initial audio frames with a decibel level below 10 dB and those with a decibel level above 85 dB, and outputs the target audio frame. In this way, by eliminating weak sounds below 10 dB and strong sounds above 85 dB, audio frames within the normal speaking range can be obtained, reducing interference and improving the quality of the target audio frame.

[0079] Before the step of inputting the text information representing the target audio data of the objects to be grouped, the historical grouping information of each object to be grouped, and the obtained object conflict groups into a large model to obtain the target conference group output by the large model, the automatic conference grouping device further includes: obtaining the object conflict groups. As an example, the automatic conference grouping device obtains a first ratio between the number of target audio frames of each object to be grouped within a preset time period and the number of initial audio frames within the preset time period; in response to the existence of at least two objects to be grouped whose first ratio is greater than a first preset value, the corresponding objects to be grouped are identified as conflicting objects in the object conflict group. The first preset value can be 60%.

[0080] As another example, the automatic conference grouping device collects the decibel levels of each object to be grouped according to a preset simultaneous speaking detection time sliding window interval, obtaining multiple initial decibel values ​​in the decibel vector of each object to be grouped; it then obtains a third ratio between the number of initial decibel values ​​within a preset decibel range in the decibel vector of each object to be grouped and the total number of initial decibel values; at least two objects to be grouped whose third ratio is greater than a first preset value are identified as conflicting objects in the same object conflict group. The preset decibel range can be greater than the first preset decibel and less than the second preset decibel.

[0081] For example, the total number of initial decibel values ​​satisfies the following formula:

[0082] m = TSameTalk * 1s / PacketPeriod

[0083] In the above formula, m represents the total number of initial decibel values, TSameTalk represents the preset simultaneous speech detection time sliding window, and PacketPerio represents the audio frame packing period. When the audio frame packing period is 40ms, then m = Tsemantic * 1s / 40ms = 25 * TSameTalk. The preset simultaneous speech detection time sliding window can be 2s.

[0084] In one embodiment, the decibel vector of the first object to be grouped, A1, is [A1_dB_t1, A1_dB_t2, ..., A1_dB_tm]. A third ratio is obtained between the number of initial decibel values ​​in A1's decibel vector that are greater than a first preset decibel and less than a second preset decibel, and the total number of initial decibel values ​​in A1's decibel vector. The decibel vector of the second object to be grouped, A2, is [A2_dB_t1, A2_dB_t2, ..., A2_dB_tn]. A third ratio is obtained between the number of initial decibel values ​​in A2's decibel vector that are greater than a first preset decibel and less than a second preset decibel. The third ratio between the number of initial decibel values ​​of the second preset decibel and the total number of initial decibel values ​​in the decibel vector of A2; when the third ratio corresponding to A1 and the third ratio corresponding to A2 are both greater than 60%, it is considered that the first grouping object and the second grouping object are in the speaking state during the same time period. However, objects in the same conference group will not speak at the same time in a short period of time. That is, the first grouping object and the second grouping object are not in the same conference group. The first grouping object and the second grouping object are conflicting objects in the object conflict group. Then the object conflict group is (A1, A2).

[0085] The automatic grouping device for the meeting determines the participation status of each participant based on the target audio frames of each participant, including: obtaining a second ratio between the number of target audio frames of each participant and the acquisition duration of the initial audio of the target audio frame; determining that the participation status of the participant is active if the second ratio is greater than or equal to a second preset value; and determining that the participation status of the participant is silent if the second ratio is less than the second preset value.

[0086] In one embodiment, the automatic conference grouping device collects initial audio at preset audio activity statistics sliding window intervals, and obtains the number of target audio frames in the currently collected initial audio; it then obtains a second ratio between the number of target audio frames in the currently collected initial audio and the collection duration of the currently collected initial audio, compares the second ratio with a second preset value, and if the second ratio is greater than or equal to the second preset value, the participant's participation status is determined to be active; otherwise, the participant's participation status is determined to be inactive. The preset audio activity statistics sliding window period can be 20 seconds.

[0087] After the automatic conference grouping device inputs the text information representing the target audio data of the objects to be grouped, the historical grouping information of each object to be grouped, and the obtained object conflict groups into a large model to obtain the target conference group output by the large model, the method further includes: traversing each target conference group, selecting the target sharing object from the grouped objects of the currently traversed target conference group; mixing the audio of the other grouped objects in the currently traversed target conference group except for the target sharing object to obtain the mixed audio; and sending the mixed audio and the obtained video of the other grouped objects to the target sharing object.

[0088] The target sharing object refers to the object that receives audio and video shared by other objects.

[0089] In one embodiment, the automatic conference grouping device groups each object to be grouped, and after obtaining multiple target conference groups, the automatic conference grouping device acquires the audio and video of each grouped object in real time, and traverses each target conference group. From the grouped objects of the currently traversed target conference group, it sequentially selects one grouped object as the target sharing object; mixes the audio of the other grouped objects in the currently traversed target conference group except for the target sharing object, to obtain the mixed audio; and sends the mixed audio and the acquired video of the other grouped objects to the target sharing object.

[0090] In one embodiment, the large model considers and groups the input prompts and preset verification conditions, outputting target meeting groups as [Topic 1: A1, A3, A5, A7] and [Topic 2: A2, A4, A6, A8]. [Topic 1: A1, A3, A5, A7] indicates that the target meeting group with meeting topic 1 includes grouped members A1, A3, A5, and A7; [Topic 2: A2, A4, A6, A8] indicates that the target meeting group with meeting topic 2 includes grouped members A2, A4, A6, and A8. Target sharing objects are selected sequentially from each target meeting group. The audio of the other grouped objects in the target meeting group (excluding the target sharing object) is mixed to obtain the mixed audio. The mixed audio and the acquired videos of the other grouped objects are then sent to the target sharing object. For example, if A1 is selected as the target sharing object from the currently traversed [Topics 1: A1, A3, A5, A7], then the audio from A3, A5, and A7 is mixed, and the resulting mixed audio is sent to A1. The video from A3, A5, and A7 is also sent to A1. Therefore, it is unnecessary to mix all objects from A1 to A8; audio and video only need to be forwarded within the same target conference group object, reducing the bandwidth pressure on mixing and traffic forwarding.

[0091] Combination Figure 4As shown, the automatic conference grouping device includes an MCU server, an MCU server Mixer, an MCU conference grouping conflict detection module, an MCU server decibel filter, a speech recognition module, and a large model.

[0092] Step 1: The automatic grouping device presets various thresholds through the MCU server, including the semantic observation time sliding window, the fourth preset value, the audio activity statistics sliding window, and the simultaneous speaking detection time sliding window.

[0093] Step 2: The automatic conference grouping device collects the initial audio of each participant and packages the audio frames in the initial audio according to the preset packaging cycle to obtain multiple initial audio frames; it collects the volume decibel value of each initial audio frame according to the audio activity statistics sliding window; it stores each volume decibel value in the extension header of the corresponding initial audio frame to obtain the audio decibel of each initial audio frame, and sends it to the MCU server.

[0094] Step 3: The MCU server sends the initial audio frame and audio decibel level to the MCU server decibel filter. The MCU server decibel filter filters out noise from each initial audio frame to obtain the target audio frame. The target audio frame is then sent to the MCU server.

[0095] Step 4: The MCU server obtains the participation status of each participant based on the target audio frames within the preset time period, and counts the number of participants whose participation status is active and whose group status is ungrouped. If the number of participants in both active and ungrouped states is less than the fourth preset value, the initial audio of all participants is sent to the MCU server Mixer.

[0096] Step 5: The MCU server Mixer performs mixing on the initial audio and sends the mixed audio to each participating object.

[0097] Step 6: When the number of objects that are simultaneously in an active state and an ungrouped state is greater than or equal to the fourth preset value, the MCU server determines the participating objects that are simultaneously in an active state and an ungrouped state as objects to be grouped, obtains the target audio data of each object to be grouped according to the semantic observation time sliding window, and sends the target audio data of each object to the speech recognition module.

[0098] Step 7: The speech recognition module performs speech recognition processing on the target audio data to obtain text information, and sends the text information of each object to be grouped to the MCU server.

[0099] Step 8: The MCU server collects the volume decibel value of the target audio frame of each object to be grouped according to the simultaneous speaking detection time sliding window, and sends the volume decibel value of the target audio frame of each object to be grouped to the MCU conference group conflict detection module.

[0100] Step 9: The MCU conference group conflict detection module performs conflict detection on the target audio frame, obtains the object conflict group, and sends the object conflict group to the MCU server.

[0101] Step 10: The MCU server sends the text information of each object to be grouped, the historical grouping information of each object to be grouped, and the object conflict groups to the large model.

[0102] Step 11: The large model considers and groups the input text information of each object to be grouped, the historical grouping information of each object to be grouped, the object conflict groups, and the preset verification conditions to obtain multiple target meeting groups, and sends the multiple target meeting groups to the MCU server.

[0103] Step 12: The MCU server sends the audio of grouped objects within the same target conference group to the MCU server Mixer.

[0104] Step 13: After the MCU server Mixer performs the mixing process, it sends the mixed audio to the grouped objects in each target conference group.

[0105] Step 14: The MCU server stores multiple target conference groups for use as historical group information in the next grouping process.

[0106] In one embodiment, the large model receives text information of each object to be grouped, historical grouping information of each object to be grouped, and object conflict groups through a text input module; a text preprocessing module preprocesses the text information, historical grouping information of each object to be grouped, and object conflict groups; a semantic parsing module parses the preprocessed text, specifically including lexical analysis and syntactic analysis, to obtain word vector representations and dependency trees, respectively; a context encoder encodes the word vector representations and dependency trees, a position encoding module performs position encoding on the encoded results, and a multi-head self-attention calculation module processes the results output by the position encoding module. The calculation is performed; the calculation results of the multi-head self-attention calculation module are processed through the residual connection + layer normalization module; the output results of the residual connection + layer normalization module are processed through the feedforward neural network; the output results of the feedforward neural network are processed through the residual connection + layer normalization module; the output results of the residual connection + layer normalization module are semantically processed through the deep semantic representation module; the output results of the deep semantic representation module are processed through the semantic understanding output module. The semantic understanding output module specifically includes intent recognition, sentiment analysis, and entity extraction on the output results of the deep semantic representation module to obtain the initial meeting group.

[0107] Figure 5 This is a block diagram illustrating an automatic conference grouping device, as shown in an exemplary embodiment of this application.Figure 5 As shown, the exemplary automatic meeting grouping device 500 includes: a participant status determination module 510, a target group determination module 520, and a target meeting group acquisition module 530. Specifically:

[0108] The participation status determination module 510 determines the participation status of each participant based on the target audio frame of each participant.

[0109] The grouping object determination module 520 is used to determine at least two grouping objects from each participant based on the participation status and grouping status of each participant.

[0110] The target conference group acquisition module 530 is used to input the text information representing the target audio data of the objects to be grouped, the historical grouping information of each object to be grouped, and the acquired object conflict groups into the large model to obtain the target conference group output by the large model.

[0111] In this exemplary automatic meeting grouping device, the target audio frames of each participant in the current meeting are acquired; the participation status of each participant is determined based on the target audio frames; at least two participants to be grouped are identified from the participants based on their participation status and grouping status; then, the text information represented by the target audio data of the participants to be grouped, the historical grouping information of each participant to be grouped, and the acquired object conflict groups are input into a large model to obtain the target meeting group output by the large model. Thus, by inputting text information, historical grouping information, and object conflict groups into the large model and using the large model to output the target meeting group, automatic grouping of participants to be grouped is achieved, improving grouping efficiency.

[0112] The functions of each module can be found in the implementation example of the automatic grouping method for meetings, and will not be repeated here.

[0113] To implement the automatic meeting grouping method of the above embodiments, this application proposes another electronic device, please refer to [link / reference needed]. Figure 6 , Figure 6 This is a schematic diagram of the structure of an embodiment of the electronic device provided in this application.

[0114] Electronic device 600 includes memory 601 and processor 602, wherein memory 601 and processor 602 are coupled together.

[0115] The memory 601 is used to store program data, and the processor 602 is used to execute the program data to implement the automatic meeting grouping method of the above embodiment.

[0116] In this embodiment, processor 602 can also be referred to as CPU (Central Processing Unit). Processor 602 may be an integrated circuit chip with signal processing capabilities. Processor 602 can also be a general-purpose processor, digital signal processor, application-specific integrated circuit, field-programmable gate array or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component. The general-purpose processor can be a microprocessor, or processor 602 can be any conventional processor.

[0117] This application also provides a computer-readable storage medium, such as Figure 7 As shown, the computer-readable storage medium 700 is used to store program data 701, which, when executed by a processor, is used to implement the automatic meeting grouping method as described in the method embodiments of this application.

[0118] The methods involved in the automatic grouping method embodiments of this application, when implemented as software functional units and sold or used as independent products, can be stored in a device, such as a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0119] The above description is merely an embodiment of this application and does not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. A method for automatically grouping meetings, characterized by, The method comprises: acquiring target audio frames of each participant in a current conference; determining a conference participation state of each participant according to the target audio frames of each participant; determining at least two grouping objects from each participant according to the conference participation state of each participant and a grouping state of each participant; inputting collected text information represented by target audio data of the grouping objects, historical grouping information of each grouping object, and acquired object conflict groups into a large model to obtain a target conference group output by the large model.

2. The method for automatically grouping a conference according to claim 1, wherein, The step of inputting the collected text information represented by the target audio data of the grouping objects, the historical grouping information of each grouping object, and the acquired object conflict groups into the large model to obtain the target conference group output by the large model comprises: initially grouping each grouping object according to corresponding text information of each grouping object through the large model to obtain a plurality of initial conference groups, wherein the initial conference groups include grouped objects; performing grouping verification processing on the grouped objects according to a conference theme of an initial conference group in which each grouped object is located and a historical conference theme represented by historical grouping information of the grouped object to obtain a first verification result; in response to the first verification result representing that the grouping of the grouped object is correct, verifying whether there is an object conflicting with the grouped object in the initial conference group in which the grouped object is located according to the object conflict groups to obtain a second verification result; in response to the second verification result representing that there is no object conflicting with the grouped object in the initial conference group in which the grouped object is located, verifying the number of objects in the initial conference group in which the grouped object is located to obtain a third verification result; in response to the third verification result representing that the verification is passed, determining the initial conference group as the target conference group.

3. The method of claim 2, wherein, The step of performing grouping verification processing on the grouped objects according to a conference theme of an initial conference group in which each grouped object is located and a historical conference theme represented by historical grouping information of the grouped object to obtain a first verification result comprises: acquiring a similarity between the conference theme of the initial conference group in which each grouped object is located and the historical conference theme represented by the historical grouping information of the grouped object; in response to the similarity being greater than or equal to a preset similarity threshold, obtaining a first verification result representing that the grouping of the grouped object is correct; in response to the similarity being less than the preset similarity threshold, acquiring a correlation between corresponding text information of each grouped object and the conference theme of the initial conference group in which the grouped object is located; in response to the correlation being greater than or equal to a preset correlation threshold, obtaining a first verification result representing that the grouping of the grouped object is correct; in response to the correlation being less than the preset correlation threshold, obtaining a first verification result representing that the grouping of the grouping object is incorrect.

4. The method for automatically grouping a conference according to claim 2, wherein, The step of verifying whether there is an object conflicting with the grouped object in the initial conference group in which the grouped object is located according to the object conflict groups to obtain a second verification result comprises: determining whether the grouped object is identical to a conflict object in the object conflict group, and if so, determining a target conflict object as a conflict object in the corresponding object conflict group other than the grouped object; determining whether there is a grouped object identical to the target conflict object in the initial conference group in which the grouped object is located; if not, obtaining a second verification result indicating that there is no object conflicting with the grouped object in the initial conference group in which the grouped object is located; if so, obtaining a second verification result indicating that there is an object conflicting with the grouped object in the initial conference group in which the grouped object is located.

5. The method for automatically grouping a conference of claim 1, wherein, The step of obtaining the target audio frame of each participant object in the current conference includes: obtaining initial audio of each participant object in the current conference; performing periodic sampling processing on the initial audio to obtain a plurality of initial audio frames; performing filtering processing on each initial audio frame according to the audio decibel of each initial audio frame to obtain the target audio frame.

6. The method for automatically grouping a conference according to claim 5, wherein, Before the step of obtaining the target conference group output by the large model in the input large model of the text information represented by the target audio data of the to-be-grouped object collected, the historical grouping information of each to-be-grouped object, and the obtained object conflict group, the method further includes: obtaining a first ratio between the number of target audio frames of each to-be-grouped object in the preset time period and the number of initial audio frames in the preset time period; in response to the first ratio corresponding to at least two to-be-grouped objects being greater than a first preset value, determining the corresponding to-be-grouped object as a conflict object in the object conflict group.

7. The method for automatically grouping a conference of claim 1, wherein, The step of determining the participation state of each participant object according to the target audio frame of each participant object includes: obtaining a second ratio between the number of target audio frames of each participant object and the collection duration of the initial audio in which the target audio frame is located; in response to the second ratio being greater than or equal to a second preset value, determining the participation state of the participant object as an active state; in response to the second ratio being less than the second preset value, determining the participation state of the participant object as a silent state.

8. The method for automatically grouping a conference of claim 1, wherein, After the step of obtaining the target conference group output by the large model in the input large model of the text information represented by the target audio data of the to-be-grouped object collected, the historical grouping information of each to-be-grouped object, and the obtained object conflict group, the method further includes: traversing each target conference group to select a target sharing object from the grouped objects in the currently traversed target conference group; performing audio mixing processing on the audio of other grouped objects in the currently traversed target conference group other than the target sharing object to obtain mixed audio; sending the mixed audio and the video of the other grouped objects obtained to the target sharing object.

9. An electronic device, comprising: comprise: a memory and a processor, wherein the memory stores program instructions, and the processor retrieves the program instructions from the memory to execute the method of any one of claims 1-8.

10. A computer-readable storage medium, characterized in that, comprise: program data stored, which is executed by a processor to implement the method of any one of claims 1-8.