Voice separation and paragraph affiliation method and system for teleconference scene
By constructing a logical map of claims-arguments and main lines-controversial points on the remote conference platform, the problem of difficulty in distinguishing the content of speakers' speeches in multi-party meetings was solved, and clear attribution and efficient editing of meeting records were achieved.
Patent Information
- Application Number
- CN202511003591.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-07-21
AI Technical Summary
In remote conferences with multiple participants, existing technologies make it difficult to effectively distinguish speakers and the logical structure of their speeches, resulting in unclear attribution of meeting minutes and inefficient editing.
By matching the voice processing engine and temporarily accessing the target conference window on the conference platform, two-way voice separation processing is performed, and a logical map of the proposition-argument relationship and the main line-dispute point is constructed. Automated meeting records are generated based on the voice logical map, and storage management is carried out in combination with permission settings.
It realizes the two-way semantic logic separation and structured attribution of remote conference voice content, improving the accuracy of meeting records and the efficiency of automatic typesetting.
Smart Images

Figure CN120808757A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of speech recognition, in particular to a speech separation and paragraph attribution method and system for remote meeting scenarios. BACKGROUND
[0002] Under the background of the increasing popularity of modern remote office and collaboration, remote meetings have become an important form of team communication, project reporting and business negotiation. However, in a multi-party simultaneous meeting scenario, due to the existence of cross, overlap and frequent switching of speeches, traditional speech recognition and transcription systems often have difficulty in accurately distinguishing the identities of the speakers and the corresponding semantic structures, especially in complex contexts such as controversial discussions and viewpoint collisions, the attribution and logical division of speech content are even more chaotic. This not only affects the completeness and readability of the meeting records, but also greatly increases the burden of post-processing and archiving. SUMMARY
[0003] The present application provides a speech separation and paragraph attribution method and system for remote meeting scenarios, which solves the technical problem that in the prior art, the speech interlacing in multi-party meetings is difficult to effectively distinguish the speakers and the logical structure of their speeches, resulting in unclear attribution of meeting records and low editing efficiency.
[0004] In a first aspect, the present application provides a speech separation and paragraph attribution method for remote meeting scenarios, the method comprising:
[0005] Upon entry of multiple user terminals into the meeting, the conference platform triggers voice recording of the target conference window to generate conference audio; performs matching of the voice processing engine and temporary access of the target conference window, and according to the first engine of the temporary access, performs double-pass speech separation processing on the conference audio to construct a speech logical graph, wherein the double-pass speech separation processing includes first logical graph construction based on argument-persuasion relationship and second logical graph construction based on main line-dispute point; according to the second engine of the temporary access, matches the target conference template, and according to the speech logical graph, performs speech segment attribution based on automated conference record layout to generate a target conference record; stores the target conference record in a platform database for retrieval and management with permission settings.
[0006] In a second aspect, the present application provides a speech separation and paragraph attribution system for remote meeting scenarios, the system comprising:
[0007] The audio generation module generates conference audio by triggering voice recording of the target conference window when the multi-party user terminal enters the conference. The voice processing module performs matching of the voice processing engine and temporary access of the target conference window, performs two-way voice separation processing on the conference audio according to a first engine of the temporary access, and constructs a voice logical graph, wherein the two-way voice separation processing includes first logical graph construction based on a claim-argument relationship and second logical graph construction based on a main line-dispute point. The record generation module matches a target conference template according to a second engine of the temporary access, performs voice segment attribution based on automatic conference record layout according to the voice logical graph, and generates a target conference record. The management module stores the target conference record in a platform database and manages the target conference record by permission setting.
[0008] The one or more technical solutions provided in the present application have at least the following technical effects or advantages:
[0009] The audio generation module generates conference audio by triggering voice recording of the target conference window when the multi-party user terminal enters the conference. The voice processing module performs matching of the voice processing engine and temporary access of the target conference window, performs two-way voice separation processing on the conference audio according to a first engine of the temporary access, and constructs a voice logical graph, wherein the two-way voice separation processing includes first logical graph construction based on a claim-argument relationship and second logical graph construction based on a main line-dispute point. The record generation module matches a target conference template according to a second engine of the temporary access, performs voice segment attribution based on automatic conference record layout according to the voice logical graph, and generates a target conference record. The management module stores the target conference record in a platform database and manages the target conference record by permission setting. The technical problem of difficulty in effectively distinguishing speakers and their speech content logical structure in voice interlacing in the multi-party conference in the prior art, resulting in unclear conference record attribution and low editing efficiency, is solved, two-way semantic logical separation and structured attribution of remote conference voice content are achieved, and the technical effects of improving conference record accuracy and automatic layout efficiency are achieved. BRIEF DESCRIPTION OF DRAWINGS
[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0011] Figure 1 The voice separation and paragraph attribution method flowchart for the remote conference scenario provided by the embodiments of the present application is shown in the following figure:
[0012] Figure 2A system structure diagram of speech separation and paragraph attribution for a remote conference scenario is provided in the embodiments of the present application.
[0013] Reference signs: audio generation module 11, speech processing module 12, record generation module 13, management module 14. DETAILED DESCRIPTION
[0014] The present application provides a speech separation and paragraph attribution method and system for a remote conference scenario, which solves the technical problem that in the prior art, speech interlacing in a multi-party conference is difficult to effectively distinguish the speaker and the logical structure of the speech content, resulting in unclear conference record attribution and low editing efficiency.
[0015] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative labor fall within the scope of protection of the present application.
[0016] It should be noted that the terms “comprise” and “have” are intended to cover non-exclusive inclusion, for example, a process, method, system, product or server comprising a series of steps or units does not have to be limited to only those steps or units clearly listed, but can include other steps or modules not clearly listed or inherent to the process, method, product or device.
[0017] Embodiment one, as shown in the present application, a speech separation and paragraph attribution method for a remote conference scenario is provided, wherein the method comprises: Figure 1 When multiple user terminals join the conference, the conference platform triggers voice recording of the target conference window to generate conference audio.
[0018] Before the remote conference starts, the conference platform establishes a conference session instance and generates a unique target conference window identifier. As multiple user terminals complete the join operation based on the target conference window, the conference platform detects the access state of the multiple user terminals in real time and triggers the initialization process of the conference voice collection module. Specifically, the conference platform receives, synchronizes and preprocesses the audio stream data uploaded by the user terminals based on the built-in audio collection engine, and timestamps and channel encodes all access voice signals. After preprocessing the audio data stream, the platform merges the voice streams into a unified conference audio data stream and continuously records the data stream during the conference period to form a complete original conference audio.
[0019]
[0020] Performing matching of the voice processing engine and temporary access of the target conference window, performing double-pass voice separation processing on the conference audio according to the first engine of the temporary access, and constructing a voice logical graph, wherein the double-pass voice separation processing includes first logical graph construction based on claim-evidence relationship and second logical graph construction based on main line-dispute point.
[0021] After the conference audio is generated, the conference platform performs dynamic matching on the idle voice processing engine group by traversing the deployed voice processing engine array based on the access state and the conference window identifier, determines the target engine group suitable for the target conference window, and establishes temporary logical connection between the target voice processing engine group and the target conference window through the system protocol, so that the engine group can access and process the conference audio stream in real time. After the connection is established, the first engine in the matched engine group is called to perform double-pass voice separation processing. Specifically, the first engine performs semantic reconstruction and deep separation based on the pre-separated voice segments (such as voice streams preliminarily separated according to voiceprints or speaker channels) in the conference audio.
[0022] The double-pass voice separation processing includes:
[0023] Claim-evidence relationship extraction: structural elements such as arguments, reasons, supplementary explanations in the speech content are identified through natural language understanding algorithm to construct a first logical graph with “claim node” as the center and “evidence edge” connection, which is used to express the argument structure and logical chain of each speaker.
[0024] Main line-dispute point structure analysis: identify the topic main line and conflict or opinion difference point between speakers, construct a second logical graph with “topic main line node” as the core and “dispute point edge” connection, which is used to reflect the discussion trend and difference structure in the whole conference process.
[0025] Further, before performing matching of the voice processing engine, the construction of the voice processing engine array includes:
[0026] Deploying a class of engines for voice separation reconstruction, deploying a class of engines for voice layout attribution, combining the class of engines and the class of engines as a voice processing engine group; logically supervising and training the voice processing engine group to determine a first engine subarray; performing multi-bit mirror twin of the first engine subarray to form a voice processing engine array, and embedding the voice processing engine array in the conference platform.
[0027] Before performing the voice processing engine matching, the construction and deployment of the voice processing engine array are completed in advance. Specifically, two types of special voice processing engines are constructed according to functions: one type of engine is dedicated to performing conference voice separation and reconstruction processing, and has the capabilities of voiceprint recognition, semantic segmentation, voice structure extraction, etc.; the other type of engine is used for attribution reasoning and layout mapping of the processed voice segments, and has the capabilities of semantic attribution judgment, template matching and layout generation. The first type of engine and the second type of engine are combined in a one-to-one manner to form a plurality of voice processing engine groups, each engine group containing a complete set of pre-processing (separation) and post-processing (attribution) capabilities. The plurality of voice processing engine groups are uniformly supervised and trained, and the training process is based on historical conference data sets, and the labeled claim / dispute point data, the attribution relationship between the voice segments and the conference records are used to iteratively optimize the model parameters. Through the training process, a number of engine groups with the best performance are selected as the first engine sub-array, which is used for subsequent dynamic task allocation. After the first engine sub-array is constructed, the platform performs multiple image twin duplication, that is, multiple mirror engine instances with the same structure and function are constructed on different physical or virtual computing nodes, and the scheduling and load of the mirror engine instances are uniformly managed. Finally, all the mirror engines form a voice processing engine array, and the array is deployed in an embedded manner in the conference platform server environment, ensuring high availability, low latency, and high parallelism of voice processing capability in a multi-concurrent, multi-user conference scenario.
[0028] Further, for a multi-concurrent conference, the platform performs dynamic balancing of the computing power of the voice processing engine array.
[0029] For the running scenario of the platform simultaneously carrying multiple concurrent conference instances, in order to ensure stable execution of voice processing tasks in a high-concurrent environment, the conference platform performs a dynamic balancing of the computing power of the voice processing engine array. Specifically, the platform obtains the running state of each voice processing engine node in real time through a resource monitoring module, including current CPU occupancy, memory usage, engine load indicators, processing queue length, etc.; at the same time, it listens to the number of concurrent conferences, the amount of voice data and the processing priority of each conference under the current platform.
[0030] In the detection of multi-concurrent conference scenarios, the platform calls the resource scheduling unit, based on a scheduling strategy model (such as weighted minimum load, resource affinity scheduling, or priority-based task shunting strategy), to schedule and allocate each sub-engine instance in the engine array. When a certain engine group reaches a preset load threshold, the platform automatically allocates a new conference instance to an engine group with lower load or in standby state, and supports dynamic expansion or task migration of mirror twin nodes to achieve balanced distribution of overall computing resources of the platform. In addition, the platform can also flexibly adjust the allocation strategy of the voice processing engine according to conference priority, user rights, or real-time performance requirements, such as reserving high-performance node resources for key users or high-level conferences to ensure the stability and timeliness of voice processing.
[0031] Further, the matching of the voice processing engine is performed, including:
[0032] As multi-party users join the conference, the idle engine sub-array is matched by traversing the voice processing engine array, and the target voice processing engine is determined; according to the platform interaction protocol, a temporary connection is established between the target voice processing engine and the target conference window, wherein the connection is disconnected when the conference ends.
[0033] After the multi-party user terminal accesses the target conference window, the platform automatically triggers the voice processing resource scheduling process and performs the matching operation of the voice processing engine. Specifically, the current deployed voice processing engine array is traversed, the engine sub-array in the idle state or accessible state is identified based on the running state indicators (such as CPU load, memory occupation, processing thread proportion, and idle channel number) of each engine sub-array, and the candidate engine group is scored according to the voice data size, processing priority, and other parameters of the current conference to determine the most matched target voice processing engine group. According to the preset platform interaction protocol, the selected target voice processing engine group is temporarily logically bound to the target conference window to establish a task processing session channel, wherein the task processing session channel can realize bidirectional communication and state synchronization based on RESTful API, WebSocket, or gRPC communication protocol, and supports automatic authentication and task state feedback mechanism. After the temporary connection is established, the real-time audio stream collected by the conference window is automatically forwarded to the bound engine for processing.
[0034] During the conference, the target voice processing engine continuously receives and processes the audio stream from the target conference window, and synchronizes the structured output result to the platform database after each round of voice analysis and processing. When a conference end event (such as all users exiting the conference, the host closing the window, etc.) is detected, the platform will automatically trigger the engine unbinding operation to disconnect the temporary connection between the target voice processing engine and the conference window, and reset the engine resource state to "idle" for subsequent task scheduling.
[0035] Further, before performing the double-pass speech separation processing on the conference audio, comprising:
[0036] According to the voiceprint recognition plug-in, performing voiceprint recognition and separation on the conference audio to determine voiceprint separated speech, wherein the multi-party speaker is taken as the separation standard; for the multi-party user end, generating a multi-party identification code, labeling the voiceprint separated speech to determine pre-separated speech.
[0037] Before performing the double-pass speech separation processing on the conference audio, the platform needs to perform preliminary voiceprint recognition and speech labeling processing to realize basic differentiation of the multi-party user speech content. Specifically, the conference platform calls the built-in voiceprint recognition plug-in, and the voiceprint recognition plug-in performs voiceprint extraction and recognition processing on the real-time collected conference audio based on a deep neural network (such as a CNN-BLSTM structure or an ECAPA-TDNN voiceprint encoder). The conference platform performs time sequence segmentation on the entire conference audio, extracts the feature vector of each speech segment, and compares it with the pre-stored user voiceprint feature library in the platform database to identify the user identity corresponding to the speaker.
[0038] After completing the speaker identification, the platform takes the multi-party user speaker as the separation standard, divides the original conference audio into several voiceprint separated speeches according to the interval clustering of voiceprint features and the time period segmentation rule, and each piece of audio clearly corresponds to a specific speaker. At the same time, the conference platform generates a unique multi-party identification code (for example, a combined code based on the conference ID + user ID) for all participants, and embeds the identification code into the metadata of each voiceprint separated speech. Finally, after completing the labeling operation on all voiceprint separated speeches, the pre-separated speech is obtained.
[0039] Further, performing the double-pass speech separation processing on the conference audio, comprising:
[0040] Triggering the first engine to take the argument item as the identification target, take the claim-argument as the construction element, perform the re-separation reconstruction on the pre-separated speech to determine the first logical graph; taking the main line-dispute point as the construction element, performing the re-separation reconstruction on the pre-separated speech to determine the second logical graph; adding the first logical graph and the second logical graph into the speech logical graph.
[0041] By triggering the first engine, the semantic structure of the content of each speech in the pre-separated voice is analyzed as the identification target of the argument entry. The first engine identifies the elements such as argument sentence, supporting argument, refutation information, data reference, etc. in each voice segment based on a natural language understanding model (such as BERT+BiLSTM-CRF or a structured semantic recognition network based on Transformer), and constructs a first logical graph according to the semantic dependency relationship between the arguments and the supporting information. In the first logical graph, each node represents an independent argument or semantic unit, and the edge represents the association path between the argument and the supporting information.
[0042] With the main line-dispute point as the construction element, the first engine performs semantic re-separation and reconstruction processing on the same batch of pre-separated voice again. The first engine analyzes the development track of the topic main line embodied in the conference speech, as well as the differences, objections or dispute focuses related thereto, identifies the aggregation and opposition relationship of different user views under a certain topic, and further constructs a second logical graph. Among them, the main line is the main structure, and the dispute point is the branch node, forming a multi-party view cross structure around a specific issue.
[0043] The first logical graph and the second logical graph are added to the voice logical graph, which is a deep semantic structure representation of the conference voice and will be used to drive the semantic attribution, automatic typesetting and storage mapping of the conference record content in the subsequent core processes.
[0044] According to the temporarily accessed second engine, a target conference template is matched, and voice segment attribution based on automatic conference record typesetting is performed according to the voice logical graph to generate a target conference record.
[0045] The conference platform calls the temporarily accessed second engine to perform the conference template matching and voice segment attribution typesetting generation process. First, the conference template library embedded in the conference platform presets the structure templates of various typical conference scenarios, including but not limited to: round table discussion type, issue promotion type, report and debate type, brainstorming type, etc., each template contains corresponding paragraph structure, role format, content level label and logical attribution constraint condition.
[0046] The second engine performs template similarity calculation and matching decision mechanism based on the structure features and semantic labels (such as speech main line, dispute point node distribution, argument-argument density, etc.) extracted from the voice logical graph to determine the most suitable target conference template. In the matching process, a joint comparison strategy based on semantic embedding (such as Sentence-BERT) and structure vectorization (such as graph embedding) can be used to improve the matching accuracy.
[0047] According to the matched target conference template, a voice segment attribution mapping and automatic layout process is triggered. The second engine maps each speech node in the voice logic graph to a specified position unit in the template according to the attribution constraint condition of the conference template, and combines the speech subject identifier code, timestamp, argument number and other information to perform attribution alignment. Specifically, it includes: based on the topic main line grouping determined by the second logic graph, the paragraph division is performed, based on the argument-persuasion relationship determined by the first logic graph, the internal structure of the speech is determined, and the semantic attribution of the voice segment is completed. Finally, the conference platform fills the attributed voice content into the template by using the structured layout engine, and automatically generates the target conference record with uniform format, clear logic and clear content.
[0048] Further, matching the target conference template includes:
[0049] For a multi-element conference scenario, a conference template library is constructed, wherein the conference template library is embedded in a platform database; for a target conference scenario, a conference template library matching based on scene elements is performed to determine the target conference template.
[0050] Based on a multi-element conference scenario, a conference template library is constructed, and the conference template library is embedded and deployed as a structured knowledge component in a platform database. The conference template library contains multiple standardized template styles, each template is predefined in combination with conference attributes (such as participant roles, agenda structure, speech turn, topic promotion logic, etc.), and covers but is not limited to: project report type, expert review type, brainstorming type, task scheduling type, technical discussion type, etc. conference scenarios.
[0051] During the actual development of a remote conference, the platform analyzes the conference initial data by using the second engine, including user role information, conference title / topic keywords, historical template reference conditions, etc., to form a set of conference scene element features. Based on these scene elements, the platform performs a rapid matching operation of the template library, that is, adopts a scene vector embedding and template index comparison method, calls a strategy model based on semantic similarity calculation and template structure matching degree evaluation, and selects a target conference template that best fits the current conference context.
[0052] Further, the voice segment attribution based on the automatic conference record layout is performed according to the voice logic graph, including:
[0053] The second engine generates attribution constraint conditions by scanning the target conference template; the attribution constraint conditions are used to perform voice attribution determination of the first attribution data based on the second logic graph on the pre-separated voice; the first attribution data is fine-tuned according to the first logic graph to determine a voice attribution result; and the target conference template is automatically laid out according to the voice attribution result to determine the target conference record.
[0054] The second engine performs a structural scan and field analysis on the selected target meeting template, extracting formatting elements such as paragraph types, role nodes, topic order, and attribution tags. Based on this, it dynamically generates attribution constraints. These constraints define the categorization of speech content within the structure template. For example, speeches under a specific topic are prioritized within topic sections, controversial points are categorized within question discussion units, and different speakers are assigned different role zones.
[0055] Based on the attribution constraints and the pre-separated speech (labeled with voiceprint information and time series), and guided by the second logical graph (mainline-dispute point relationship), the speech segments are initially attributed to generate first attribution data. This first attribution data represents the initial corresponding position of each speech segment in the conference structure template.
[0056] The first attribution data is semantically fine-tuned using the first logical graph (claim-argument logical relationship) to improve the logical integrity and expression coherence of speech attribution, such as aligning argument fragments to their claim speech paragraphs, aggregating speech blocks with the same position, etc., and finally outputting the speech attribution results.
[0057] The conference platform automatically fills in and structures the target conference template based on the audio attribution results, and automatically generates target conference records with standard format and clear semantic classification.
[0058] The target meeting records are stored in the platform database and retrieved and managed based on permission settings.
[0059] After the target meeting record is generated, the conference platform implements a unified data archiving and rights management strategy. Specifically, the target meeting record is associated with its corresponding voice logic map (including the first logic map and the second logic map) and stored in the platform database.
[0060] During database storage, the system sets access rights based on user identity, meeting organizational structure, and speaking role classification, creating a multi-tiered access control strategy. For example, the meeting host has edit and approval permissions for all voice clips and structured records, while ordinary participants can only view their own speeches or authorized topic sections. Administrators have archiving, auditing, and global access permissions.
[0061] In addition, an index mapping table is constructed using node elements in the voice logic graph (such as speakers, topic labels, and controversial point locations), supporting targeted retrieval based on multiple criteria such as role, content, topic, and time. When a retrieval request is triggered, the index field is matched against the current user's permission tags, returning only meeting clips and structured content that comply with the access policy, effectively ensuring data security and controllability.
[0062] Further, the target conference record is stored in a platform database for access management according to permission settings, including:
[0063] The first logical graph and the second logical graph are used as intermediate layers to establish a mutual mapping with the target conference record, and a permission level is set and stored in the platform database; and the target conference record is directed to be accessed through a graph node element in the first logical graph and the second logical graph.
[0064] After the target conference record is generated, the first logical graph and the second logical graph are used as semantic intermediate layers to establish a one-to-one mapping or a many-to-many mapping relationship with the corresponding conference record content, to form a mutual mapping structure. The mutual mapping structure is implemented in the form of a mapping table or an index table, including but not limited to corresponding relationships such as speaker identification and speech segment number, logical node and speech content paragraph, graph path and conference topic time sequence relationship.
[0065] On this basis, according to user roles (such as a host, a recorder, a participant, etc.), semantic sensitivity (such as controversial content, decision items), conference levels and other parameters, an access permission level is set for each voice segment in the target conference record.
[0066] By writing the target conference record together with its logical graph index information into the platform database in the form of a composite of structured data and graph structure data, mixed graph-text retrieval and graph-level relationship analysis are supported.
[0067] When the calling or accessing operation of the conference content is performed, the platform can perform directed access on the target conference record according to the graph node elements (including node types, path structures, key semantic identifiers, etc.) in the first logical graph and the second logical graph. For example, when the content related to a specific topic is accessed, the "topic main line" node in the graph can be used for quick aggregation; when the view of a specific speaker is accessed, the "position node" and its speech mapping relationship can be used for backtracking extraction, so as to realize accurate content access and permission-controlled retrieval driven by logical structure.
[0068] In summary, the embodiments of the present application have at least the following technical effects:
[0069] With the multi-party user terminal joining the conference, the conference platform triggers voice recording of the target conference window to generate conference audio. Then, the matching of the voice processing engine and the temporary access of the target conference window are performed. According to the first engine of the temporary access, the conference audio is subjected to double-pass voice separation processing to construct a voice logical graph, wherein the double-pass voice separation processing includes first logical graph construction based on the argument-persuasion relationship and second logical graph construction based on the main line-disputed point. Then, according to the second engine of the temporary access, the target conference template is matched, and the voice segment attribution based on the automatic conference record layout is performed according to the voice logical graph to generate the target conference record. Finally, the target conference record is stored in the platform database for retrieval management with permission settings. The technical problems of the prior art that the voice interweaving in the multi-party conference is difficult to effectively distinguish the speaker and the logical structure of the speech content, resulting in unclear conference record attribution and low editing efficiency are solved. The technical effects of double-pass semantic logical separation and structured attribution of remote conference voice content are achieved, and the accuracy and automatic layout efficiency of the conference record are improved.
[0070] In the second embodiment, based on the same inventive concept as the voice separation and paragraph attribution method for the remote conference scenario in the foregoing embodiments, as shown in the following table, the present application provides a voice separation and paragraph attribution system for the remote conference scenario, wherein the system includes: Figure 2
[0071] The audio generation module 11 generates conference audio by triggering voice recording of the target conference window with the multi-party user terminal joining the conference. The voice processing module 12 performs matching of the voice processing engine and temporary access of the target conference window. According to the first engine of the temporary access, the conference audio is subjected to double-pass voice separation processing to construct a voice logical graph, wherein the double-pass voice separation processing includes first logical graph construction based on the argument-persuasion relationship and second logical graph construction based on the main line-disputed point. The record generation module 13 matches the target conference template according to the second engine of the temporary access, and performs voice segment attribution based on the automatic conference record layout according to the voice logical graph to generate the target conference record. The management module 14 stores the target conference record in the platform database for retrieval management with permission settings.
[0072] Further, the voice processing module 12 is configured to perform the following method:
[0073] A first type of engine is deployed for voice separation reconstruction, a second type of engine is deployed for voice layout attribution, the first type of engine and the second type of engine are combined as a voice processing engine group, the voice processing engine group is subjected to logical supervision training to determine a first engine subarray, the first engine subarray is subjected to multi-bit mirror twin to form a voice processing engine array, and the voice processing engine array is embedded and deployed in the conference platform.
[0074] Further, the voice processing module 12 is configured to perform the following method:
[0075] With the entry of multiple users into the conference, the idle engine sub-array matching is performed to determine the target voice processing engine by traversing the voice processing engine array; and the temporary connection between the target voice processing engine and the target conference window is established according to the platform interaction protocol, and the connection is disconnected when the conference is terminated.
[0076] Further, the voice processing module 12 is configured to perform the following method:
[0077] The voiceprint recognition and separation are performed on the conference audio according to the voiceprint recognition plug-in to determine voiceprint-separated speech, wherein the multi-party speaking subjects are taken as the separation standard; the multi-party identification codes are generated for the multi-party user terminals, and the voiceprint-separated speech is labeled to determine pre-separated speech.
[0078] Further, the voice processing module 12 is configured to perform the following method:
[0079] The pre-separated speech is reconstructed by complex separation according to the argument item as the recognition target and the claim-argument as the construction element to determine the first logical graph; the pre-separated speech is reconstructed by complex separation according to the main line-disputed point as the construction element to determine the second logical graph; and the first logical graph and the second logical graph are added to the voice logical graph.
[0080] Further, the record generation module 13 is configured to perform the following method:
[0081] The conference template library is constructed for the multi-element conference scene, wherein the conference template library is embedded in the platform database; and the conference template library matching based on the scene elements is performed for the target conference scene to determine the target conference template.
[0082] Further, the record generation module 13 is configured to perform the following method:
[0083] The second engine generates the attribution constraint condition by scanning the target conference template; the first attribution data is determined by performing the voice attribution determination based on the second logical graph on the pre-separated speech according to the attribution constraint condition; the voice attribution result is determined by fine-tuning the first attribution data according to the first logical graph; and the target conference record is determined by automatically typesetting the target conference template according to the voice attribution result.
[0084] Further, the voice processing module 12 is configured to perform the following method:
[0085] The first logical atlas and the second logical atlas are used as intermediate layers to establish mutual mapping with the target conference record, and a permission level is set and stored to a platform database; wherein the target conference record is directed to call through the atlas node elements in the first logical atlas and the second logical atlas.
[0086] Further, the management module 14 is used to execute the following method:
[0087] If it is a multi-concurrent conference, the platform computing power is dynamically balanced and distributed to the voice processing engine array.
[0088] It should be noted that the above-mentioned sequence of the embodiments of the present application is only for description, and does not represent the advantages and disadvantages of the embodiments. The above-mentioned specific embodiments of the present application are described. The processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are possible or can be advantageous.
[0089] The above-mentioned is only the preferred embodiment of the present application, and does not limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
[0090] The present application and the drawings are only exemplary descriptions of the present application, and are considered to cover any and all modifications, changes, combinations or equivalents within the scope of the present application. Obviously, those skilled in the art can make various modifications and changes to the present application without departing from the scope of the present application. Thus, if these modifications and changes of the present application belong to the scope of the present application and its equivalent technology, the present application intends to include these modifications and changes.
Claims
1. A speech separation and paragraph attribution method for remote conference scenarios, characterized by: The method comprises: As multiple users join the conference, the conference platform triggers the voice recording of the target conference window and generates conference audio; Performing matching of the voice processing engine and temporary access to the target conference window, performing two-way voice separation processing on the conference audio according to the temporarily accessed first engine, and constructing a voice logic graph, wherein the two-way voice separation processing includes constructing a first logic graph based on the proposition-argument relationship and constructing a second logic graph based on the main line-disputed point; According to the temporarily connected second engine, the target conference template is matched, and the voice segment attribution based on the automated conference record typesetting is performed according to the voice logic graph to generate the target conference record; The target meeting records are stored in the platform database and retrieved and managed based on permission settings.
2. The method for speech separation and paragraph attribution for remote conference scenarios according to claim 1, characterized in that: Before performing voice processing engine matching, the voice processing engine array is constructed, including: Deploy the first engine based on speech separation and reconstruction, and deploy the second engine based on speech typesetting and attribution. Combine the first and second engines to form a speech processing engine group. Performing logical supervision training on the speech processing engine group to determine a first engine sub-array; The first engine sub-array is multi-bit mirrored twinned to form a voice processing engine array, and the voice processing engine array is embedded and deployed on the conference platform.
3. The method for speech separation and paragraph attribution for remote conference scenarios according to claim 2, characterized in that: Perform matching on the speech processing engine, including: As multiple users join the conference, the voice processing engine array is traversed, and idle engine sub-arrays are matched to determine the target voice processing engine; According to the platform interaction protocol, a temporary connection is established between the target voice processing engine and the target conference window, wherein the connection is disconnected when the conference is terminated.
4. The method for speech separation and paragraph attribution for remote conference scenarios according to claim 2, wherein: Before performing two-way voice separation processing on the conference audio, the method includes: According to the voiceprint recognition plug-in, voiceprint recognition and separation are performed on the conference audio to determine voiceprint separation speech, wherein the separation standard is based on multiple speaking subjects; For multiple user terminals, multiple identification codes are generated, the voiceprint-separated speech is marked, and the pre-separated speech is determined.
5. The method for speech separation and paragraph attribution for remote conference scenarios according to claim 4, characterized in that: Performing two-way voice separation processing on the conference audio includes: The first engine is triggered to perform complex separation and reconstruction on the pre-separated speech with argument entries as recognition targets and propositions and arguments as construction elements to determine a first logical graph; Taking the main line-disputed point as construction elements, performing complex separation and reconstruction on the pre-separated speech to determine a second logical graph; The first logic graph and the second logic graph are added to the voice logic graph.
6. The method for speech separation and paragraph attribution for remote conference scenarios according to claim 4, characterized in that: Match the target conference template, including: Build a conference template library for multiple conference scenarios, where the conference template library is embedded in the platform database; For the target conference scene, a conference template library matching based on scene elements is performed to determine the target conference template.
7. The method for speech separation and paragraph attribution for remote conference scenarios according to claim 6, characterized in that: Performing voice segment attribution based on automated conference record typesetting according to the voice logic graph includes: The second engine generates affiliation constraints by scanning the target conference template; Performing voice attribution based on the second logical graph on the pre-separated voice with the attribution constraint condition to determine first attribution data; Fine-tune the first attribution data according to the first logical graph to determine a voice attribution result; According to the voice attribution result, the target conference template is automatically typeset to determine the target conference record.
8. The method for speech separation and paragraph attribution for remote conference scenarios according to claim 1, wherein: The target meeting minutes are stored in the platform database and accessed and managed using permission settings, including: Using the first and second logical graphs as the intermediate layers, a mutual mapping is established with the target meeting record, and permission levels are set, and the records are stored in the platform database; Among them, the target meeting records are retrieved in a targeted manner through the graph node elements in the first logical graph and the second logical graph.
9. The method for speech separation and paragraph attribution for remote conference scenarios according to claim 2, wherein: If there are multiple concurrent conferences, the platform computing power is dynamically and evenly distributed to the voice processing engine array.
10. A speech separation and paragraph attribution system for remote conference scenarios, characterized by: A system for implementing the speech separation and paragraph attribution method for remote conference scenarios according to any one of claims 1 to 9, comprising: Audio generation module: When multiple users join the conference, the conference platform triggers the voice recording of the target conference window and generates conference audio; The voice processing module matches the voice processing engine and temporarily connects to the target conference window. Based on the temporarily connected first engine, it performs two-way voice separation processing on the conference audio and constructs a voice logic graph. The two-way voice separation processing includes constructing a first logic graph based on the claim-argument relationship and a second logic graph based on the main line-disputed point. Record generation module: matches the target conference template according to the temporarily connected second engine, and performs voice segment attribution based on the automated conference record typesetting according to the voice logic graph to generate the target conference record; Management module: stores the target meeting records in the platform database and manages their access based on permission settings.
Citation Information
Patent Citations
Method and device for constructing knowledge graph based on conference record and processor
CN112487197A
Intelligent conference management method and device, computer device and storage medium
CN113127633A
Conference analysis method and device and storage medium
CN117933237A
Conference recording method based on Internet of Things
CN119993161A
RAG-based intelligent conference full-process management method and device, and medium
CN120146178A