RAG-based intelligent conference full-process management method and device, and medium
By adopting a full-process management method based on RAG in intelligent meetings, we collect and process speech and voice information in real time and generate personalized conference summary, the omissions and deviations of information transmission in complex conference scenarios are solved, and real-time interaction and personalized services are realized.
Patent Information
- Application Number
- CN202510057366.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-06-13
AI Technical Summary
The existing intelligent conference management method cannot meet the needs of real-time interaction and personalized services for participating users in the complex conference scenario where multiple people use one terminal device to attend, especially in the process of information transmission, omissions and deviations are prone to occur.
Using the intelligent full-process management method of RAG-based conferences, we collect speech speech information in real time in the current conference space, perform speech recognition and generate identification text information containing detailed text speech attributes, and conflict processing is carried out based on text speech attributes, update to the RAG conference knowledge base in real time, and generate personalized conference summary information through the real-time interactive question-and-answer interface.
It realizes accurate recording and information management in complex conference scenarios, ensures real-time and personalized information, improves the participation experience and conference efficiency, and meets the real-time interaction and personalized service needs of multiple people participating in the conference using terminal equipment.
Smart Images

Figure CN120146178A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of smart conference technology, and in particular to a RAG-based smart conference full-process management method, device and medium. Background Art
[0002] With the development of information technology, remote meetings and online meetings have been integrated into daily communication scenarios in enterprises and academia, becoming a normalized way of communication. The traditional meeting model mainly relies on manual recording of meeting content, organizing meeting minutes after the meeting, and then using tools to transmit and communicate information to achieve information sharing. The above traditional meeting model is inefficient, and it is easy to omissions in the process of information transmission. Participants' understanding of the content of the meeting is also prone to deviations, which makes it difficult for the results of the meeting to be effectively applied. In recent years, with the advancement of artificial intelligence technologies such as natural language processing and machine learning, intelligent conference systems have improved the process of meeting organization, recording and information management through automation.
[0003] However, most of the existing smart conference solutions focus on the realization of a single function, such as automatic speech recognition or simple meeting minutes generation tools that are only used for text transcription of conference recordings, lacking in-depth understanding and effective management of conference content. In actual conferences, there are complex conference scenarios where multiple people use one terminal device to attend the meeting. In such complex scenarios, multiple participants corresponding to the same terminal device participant terminal will have internal discussions during the conference speech, and there will also be situations where the terminal sets the participant terminal to participate in the main line of the conference content. If a single automatic speech recognition transcription is used, the content of the internal discussion is easily mistaken for the content of the meeting, resulting in inaccurate meeting records. In addition, in the above-mentioned complex conference scenarios, participants may come from multiple departments and organizations, and different participants have different concerns about the content of the meeting. The meeting minutes generated in a fixed form are in a unified format, without highlighting the key points, which affects the participants' acquisition of the content of the meeting. In addition, the existing smart conference needs to generate a summary after the meeting. If a participant joins the meeting midway, he cannot promptly learn the content of the meeting before joining the meeting, and the real-time interactivity is poor.
[0004] In summary, the existing intelligent conference management method uses the method of transcribing conference recordings to generate meeting minutes after the meeting, which cannot meet the needs of real-time interaction and personalized services for participants in complex conference scenarios where multiple people use one terminal device to attend the meeting. Summary of the invention
[0005] One or more embodiments of this specification provide a method, device, and medium for intelligent conference full-process management based on RAG to solve the following technical problems: The existing intelligent conference management method uses the method of transcribing the meeting recording text to generate meeting minutes after the meeting, which cannot meet the needs of real-time interaction and personalized services for participating users in complex meeting scenarios where multiple people use a single terminal device to participate in the meeting.
[0006] One or more embodiments of this specification adopt the following technical solutions:
[0007] One or more embodiments of this specification provide a method for intelligent conference full-process management based on RAG. The method includes: real-time collecting the speech information corresponding to each participating terminal in the current meeting space, performing speech recognition on each piece of the speech information to generate identification text information, where the identification text information includes speech text information and corresponding text speech attributes, and the text speech attributes include a speech terminal identifier and a unique identifier of the speaking user; according to the text speech attributes, performing conflict processing on the speech text information to determine the meeting knowledge increment information corresponding to the speech text information, so as to update the meeting knowledge increment information to the RAG meeting knowledge base corresponding to the current meeting space constructed in advance in real time; obtaining the real-time question-and-answer request information of the participating users through a pre-constructed real-time interaction question-and-answer interface, and based on the question-and-answer request information, generating real-time question-and-answer interaction information through a pre-constructed large language model and the RAG meeting knowledge base; generating meeting summary information in the current meeting space corresponding to each participating user based on the RAG meeting knowledge base and the real-time question-and-answer interaction information corresponding to each participating user.
[0008] One or more embodiments of this specification provide an intelligent conference full-process management device based on RAG, including:
[0009] At least one processor; and,
[0010] A memory communicatively connected to the at least one processor; wherein,
[0011] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above method.
[0012] A non-volatile computer storage medium provided by one or more embodiments of this specification stores computer-executable instructions, and the computer-executable instructions are set to: execute the above method.
[0013] The above at least one technical solution adopted in the embodiments of this specification can achieve the following beneficial effects: By collecting speech information in real time in the current meeting space and performing speech recognition to generate identification text information containing detailed text speech attributes, various speech situations in the meeting can be comprehensively and meticulously captured, ensuring that information in complex scenarios such as single-person speech or multiple people sharing a terminal for speech can be accurately recorded; By performing conflict processing on the speech text information based on the text speech attributes, the truly valuable incremental meeting knowledge information can be identified, avoiding knowledge chaos caused by conflicts between different participants' viewpoints and repeated expressions, etc., and ensuring that only high-quality knowledge that conforms to the meeting logic and has been screened will be updated to the RAG meeting knowledge base, so that the knowledge in the knowledge base always maintains accuracy, consistency, and keeps pace with the times, being dynamically updated in line with the actual progress of the meeting; With the pre-constructed real-time interactive Q&A interface, participating users can submit real-time Q&A request information at any time and quickly obtain real-time Q&A interactive information generated based on the language large model and the RAG meeting knowledge base. This convenient interaction method enables participating personnel to get targeted answers to the meeting content as soon as they have questions during the meeting without waiting for the meeting to end, greatly improving the participation experience; Moreover, for late-arriving users, even if they did not participate in the previous meeting, they can obtain the current meeting progress and the meeting knowledge of the unparticipated meeting process through the form of real-time Q&A interaction; Generating personalized meeting summary information based on the RAG meeting knowledge base and the real-time Q&A interactive information corresponding to each participating user fully considers the differences in the concerns and knowledge needs of different participants. Each participant's meeting summary is a customized version tailored to their own situation, enabling them to quickly review the key points of the meeting and absorb the knowledge useful to themselves, saving the time and effort of sorting out knowledge after the meeting; Through the real-time updated RAG meeting knowledge base, real-time interaction Q&A, and personalized meeting summary information after knowledge conflict processing, the needs for real-time interaction and personalized services of participating users in complex meeting scenarios where multiple people use one terminal device to participate in the meeting are met. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In order to more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments recorded in this specification. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings. In the drawings:
[0015] Figure 1 It is a schematic flowchart of a method for intelligent meeting full-process management based on RAG provided by an embodiment of this specification;
[0016] Figure 2Schematic diagram of an application scenario of an intelligent conference full-process management method based on RAG provided by an embodiment of this specification;
[0017] Figure 3 Schematic diagram of the structure of an intelligent conference full-process management device based on RAG provided by an embodiment of this specification. Detailed implementation manners
[0018] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments of this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this specification.
[0019] An embodiment of this specification provides an intelligent conference full-process management method based on RAG. It should be noted that the execution subject in the embodiment of this specification can be a server or any device with data processing capabilities. Figure 1 Schematic diagram of the process of an intelligent conference full-process management method based on RAG provided by an embodiment of this specification, as Figure 1 shown, mainly includes the following steps:
[0020] Step S101, in the current conference space, real-time collect the speech information corresponding to each participating terminal, perform speech recognition on each speech information, and generate identification text information.
[0021] Among them, the identification text information includes speech text information and corresponding text speech attributes, and the text speech attributes include a speech terminal identifier and a unique identifier of the speech user;
[0022] Retrieval-Augmented Generation (RAG) is a technical architecture that combines information retrieval and text generation. The main purpose is to use the retrieved relevant knowledge to improve the generation quality when generating text content, so that the generated answers are more accurate, targeted, and information-rich. Introducing the RAG retrieval enhancement system into the conference scenario supports real-time one-on-one intelligent Q&A interactions among conference participants during the conference and provides instant knowledge support.
[0023] Before collecting the speech information corresponding to each participating terminal in real time in the current meeting space, the method further includes: under the trigger of the meeting management terminal, constructing the current meeting space at the meeting creation node; obtaining the meeting participation information and meeting theme information of the target meeting, where the meeting participation information includes multiple participating terminal information, at least one participating user corresponding to each participating terminal, and the participation permissions of each participating user; based on the meeting participation information of the target meeting, setting the participating roles in the current meeting space, and setting the knowledge base framework of the RAG meeting knowledge base corresponding to the current meeting space through the meeting theme information and the participating roles.
[0024] In one embodiment of this specification, when the meeting organizing user triggers the meeting creation process on the meeting management terminal by clicking the creation button or other forms, the current meeting space can be constructed at the meeting creation node through the meeting initialization module pre-set in the management system. And various key information of the target meeting uploaded by the meeting organizing user is obtained. The key information here includes meeting theme information and meeting participant information. The meeting theme information here is used to represent the summary of the meeting content during the target meeting. The meeting participant information includes the participant terminal information of multiple participant terminals. For example, the participant terminal identifier, and each participant terminal corresponds to a participant account. The participant terminals here can be various types of mobile terminals. In addition to the participant terminal information, it also includes at least one participant user corresponding to each participant terminal. It should be noted that in complex meeting scenarios, online meetings occur between different organizations. That is to say, multiple users within an organization use the same participant terminal to join the meeting. Therefore, the participant user information corresponding to each participant terminal needs to be included, which can be in the form of participant user identifiers. In addition, in actual online meetings, the meeting permissions corresponding to different organizations are different. Based on the meeting participant information of the target meeting, the participant roles within the current meeting space are set. For example, the person responsible for technical explanation is set as the technical expert role, the person responsible for coordinating the meeting process and controlling the meeting rhythm is determined as the meeting host role, and ordinary information receivers and feedbackers are classified as general participant roles, etc. For different meeting themes and meeting participants, the knowledge base framework of the RAG meeting knowledge base within the meeting space is different. Through the meeting theme information and the participant role, the knowledge base framework of the RAG meeting knowledge base corresponding to the current meeting space is set. First, according to the meeting theme, multiple framework themes in the knowledge base framework are determined. For example, the meeting theme information is subjected to theme keyword extraction. Based on the extracted keywords, with the help of a knowledge graph or other resources for theme expansion, relevant synonyms, hypernyms, and hyponyms are mined, and multiple framework themes in the knowledge base architecture are respectively determined. The framework themes here are major themes and can be determined by hypernyms. There may be multiple sub-themes within each framework theme, and the corresponding ones can be relevant hyponyms. After constructing the knowledge base architecture, the main storage area for the speech data in the corresponding knowledge base architecture is set according to the participant role. The knowledge base framework of the RAG meeting knowledge base corresponding to the current meeting space is set in the above manner.
[0025] Through the above technical solution, the meeting organizing user triggers the meeting creation process on the meeting management terminal, quickly constructs the current meeting space, saving the time cost of meeting preparation; constructs a dedicated RAG meeting knowledge base framework for different meeting themes and participants, expands the framework themes that meet the knowledge needs of the participants according to the meeting theme, and then sets the main storage area for the speech data according to the participant role, making the knowledge storage organized, and the knowledge generated during the subsequent meeting process can be classified orderly according to the framework.
[0026] Perform speech recognition on each piece of the speech information of the speech, and generate identification text information, specifically including: determining the speaking participating terminal corresponding to the speech information of the speech to determine the speaking terminal identifier corresponding to each piece of the speech information of the speech; according to the speaking terminal identifier corresponding to each piece of the speech information of the speech, perform timbre recognition on at least one specified speech information belonging to the same speaking participating terminal to determine the timbre feature corresponding to each piece of the specified speech information; perform similarity matching on multiple such timbre features to determine the timbre similarity matching result between each piece of the specified speech information; based on the timbre similarity matching result between each piece of the specified speech information, determine the unique identifier of the speaking user of the speech information corresponding to the participating terminal; perform text conversion on each piece of the speech information of the speech to determine the corresponding speaking text information, and set the text speaking attribute corresponding to the speaking text information with the speaking terminal identifier and the unique identifier of the speaking user.
[0027] In one embodiment of the present specification, the speech information of the speech corresponding to each participating terminal is collected in real time within the current meeting space, and the speech information during the meeting can be collected in real time through the sound collection device of each participating terminal. It should be noted that the speech information of the speech corresponding to the participating terminal collected here is continuous speech information collected under continuous time stamps, that is, there is no speech situation of other participating terminals within the time period corresponding to the continuous time stamps. Among the speech information of the speech of the participating terminal collected, there may be a situation where only one speaker participates in the main discussion of the meeting for each participating terminal, and there will also be a situation where multiple speakers in the participating terminal conduct internal discussions. In the case of multiple speakers conducting internal discussions in the same participating terminal, the content of the internal discussion is non-main discussion content of the meeting, and only the conclusive expression of the meeting discussion content needs to be stored in the knowledge base. Therefore, it is necessary to identify the speech information of the speech of each participating terminal, and determine the actual situation corresponding to the speech information of the speech of each participating terminal through the identification text information. The identification text information includes the speaking text information and the corresponding text speaking attribute, and the text speaking attribute includes the speaking terminal identifier and the unique identifier of the speaking user. The unique identifier of the speaking user includes two identifiers: the speaking user is unique and the speaking user is not unique. The speaking user is unique means that multiple pieces of speech information corresponding to the same participating terminal come from the same speaking user, that is, the participating user; correspondingly, the speaking user is not unique means that multiple pieces of speech information corresponding to the same participating terminal come from different users, which may be that multiple users participate in the main discussion of the meeting or may be an internal discussion situation.
[0028] In one embodiment of this specification, when performing speech recognition on each piece of the speech information of the speech, the speech participating terminal corresponding to the speech information of the speech is determined to determine the speech terminal identifier corresponding to each piece of the speech information of the speech. In order to judge the speech uniqueness identifier in the text speech attribute corresponding to the speech information of the speech and determine whether the speech information of the speech is internal discussion content, it is found in the study of the meeting scenario that if multiple participating users use the same speech terminal for meeting speeches, the timbres corresponding to the speech voices collected by the same speech participating terminal are significantly different. Therefore, according to the speech terminal identifier corresponding to each piece of the speech information of the speech, timbre recognition is performed on at least one specified speech information belonging to the same speech participating terminal to determine the timbre feature corresponding to each specified speech information of the speech. For example, the speech signal corresponding to the specified speech information of the speech is converted to the Mel frequency domain, and then the Mel frequency cepstral coefficients are obtained through discrete cosine transform (DCT) to obtain the timbre feature corresponding to each specified speech information of the speech.
[0029] Through the feature similarity matching algorithm, the similarities of multiple such timbre features are matched, and according to the timbre similarity threshold corresponding to the pre-set corresponding feature similarity matching algorithm, the timbre similarity matching result between each specified speech information of the speech is determined. For example, if the cosine similarity matching algorithm is adopted, the timbre similarity threshold can be set to 0.8. If the timbre similarity between any two speeches of the speech is less than the timbre similarity threshold, it means that the timbre similarity matching result corresponding to the two speeches of the speech is non-similar timbre. As long as there is a timbre similarity matching result of non-similar timbre for any two speeches of the speech among the multiple specified speech information belonging to the same speech participating terminal, the speech user uniqueness identifier of the multiple specified speech information belonging to the same speech participating terminal is the non-unique user identifier. Correspondingly, if there is no timbre similarity matching result of non-similar timbre for any two speeches of the speech among the multiple specified speech information belonging to the same speech participating terminal, that is, they are all similar timbres, it is determined that the speech user uniqueness identifier of the multiple specified speech information belonging to the same speech participating terminal is the unique user identifier. Using the speech recognition technology, text conversion is performed on each piece of the speech information of the speech to determine the corresponding speech text information, and the text speech attribute corresponding to the speech text information is set with the speech terminal identifier and the speech user uniqueness identifier.
[0030] Through the above technical scheme, the unique identification of the speaking user is judged according to the timbre characteristics. When there are multiple speaking voice information on the same participating terminal, it can be accurately distinguished whether the voice information comes from the same user or different users; the actual situation corresponding to the speaking voice information is clearly distinguished, and its source (speaking terminal identification) and speaking subject situation (speaking user unique identification) are accurately recorded using identification text information, which provides a clear basis for the traceability and verification of knowledge in the knowledge base, ensures that the information stored in the knowledge base is accurate and reliable, avoids knowledge errors caused by unknown sources or confusion, and maintains the quality of the knowledge base knowledge system; taking into account the different speaking scenarios that may occur in actual meetings, such as a single speaker participating in the main line discussion and multiple speakers using the same terminal to speak and other complex situations, the technical scheme can be processed in a targeted manner, and the speaking voice information in various scenarios can be identified through reasonable speech recognition, timbre analysis and other means, thereby enhancing the adaptability to diverse conference scenarios.
[0031] Step S102, according to the text speech attribute, conflict resolution is performed on the speech text information to determine the incremental conference knowledge information corresponding to the speech text information, so as to update the incremental conference knowledge information in real time to the pre-built RAG conference knowledge base corresponding to the current conference space.
[0032] In complex conference scenarios, if multiple participants use the same terminal to speak, not all speech texts corresponding to the collected speech information need to be stored in the knowledge base. For example, only the summary of the internal branch discussion needs to be stored in the knowledge base. If the content of the internal branch discussion is added to the knowledge base, the knowledge base will be filled with a large number of non-critical and trivial internal discussion details.
[0033] In one embodiment of the present specification, based on the text speech attributes, the speech text information is conflict-processed to determine the incremental conference knowledge information corresponding to the speech text information that ultimately needs to be updated to the knowledge base, and the incremental conference knowledge information is updated in real time to the RAG conference knowledge base corresponding to the current conference space. It should be noted that the RAG conference knowledge base here already stores the conference knowledge corresponding to the content of the meeting.
[0034] According to the speech attribute of the text, conflict processing is performed on the speech text information to determine the incremental information of meeting knowledge corresponding to the speech text information, specifically including: judging whether there is a branch discussion conflict in at least one speech text information belonging to the same speech terminal identifier according to the speech text attribute corresponding to the speech text information; through the judgment result, determining the target judgment semantic information according to at least one speech text information to determine whether the target judgment semantic information is the main line conflict semantic; if the target judgment semantic information is the main line conflict semantic, obtaining the first speech participating terminal corresponding to the conflict semantic feature, and determining the retained semantic according to the current speech participating terminal corresponding to the target judgment semantic information and the first speech participating terminal; when the retained semantic is the target judgment semantic information, determining the target judgment semantic information as the incremental information of meeting knowledge, and deleting and archiving the conflict semantic feature.
[0035] In an embodiment of the present specification, first, according to the unique identifier of the speech user in the speech text attribute corresponding to the speech text information, it is determined whether the speech text information is unique to the speech user. If so, it means that there is no branch discussion conflict in at least one speech text information belonging to the same speech terminal identifier. If not, the volume feature of the speech voice information corresponding to the speech text is further extracted, and the average volume feature corresponding to the speech participating terminal is calculated according to the volume features of multiple speech voice information belonging to the same speech terminal identifier. Generally, in the actual scenarios of the main line discussion or internal branch discussion in a meeting, if the participating users are having an internal discussion, they generally lower their voices. On the contrary, if they participate in the main line discussion of the meeting, they generally get closer to devices such as microphones. That is to say, the volume of the main line discussion is greater than that of the internal branch discussion. When the average volume feature is lower than the preset main line discussion volume feature threshold, it indicates that there is an internal discussion among multiple participating users, and it is determined that there is a branch discussion conflict in at least one speech text information belonging to the same speech terminal identifier.
[0036] Through the above technical solutions, by performing conflict processing on the speech text information, the incremental information of meeting knowledge can be accurately determined, avoiding the mixing of duplicate, contradictory or irrelevant information into the knowledge base, so that the knowledge base stores only the truly valuable new knowledge after screening and discrimination, which helps to improve the quality of the content in the knowledge base and keep the knowledge base always concise and efficient; considering the different situations where the speech user may be unique or not unique in the meeting, as well as the differences in volume features and other aspects between the main line discussion and the internal branch discussion, it is possible to flexibly make corresponding judgments and processes according to the characteristics of these actual scenarios, adapting to various complex meeting speech and discussion scenarios, and being able to operate effectively whether it is a regular single-person speech or a complex situation where multiple people share a terminal, ensuring the effectiveness and accuracy of meeting knowledge management.
[0037] Based on the judgment result, determine the target judgment semantic information according to at least one speech text information, specifically including: if there is no branch discussion conflict in at least one speech text information belonging to the same speech terminal identifier, determine the target judgment semantic information with the current speech semantic feature corresponding to the speech text information; if there is a branch discussion conflict in at least one speech text information belonging to the same speech terminal identifier, determine multiple current speech semantic features belonging to the same speech terminal identifier according to the speech terminal identifier; generate a branch discussion text set corresponding to the speech terminal identifier according to the speech timestamp corresponding to each current speech semantic feature obtained in advance; summarize the branch discussion text set to determine the terminal speech semantic feature corresponding to the branch discussion text set, and determine the target judgment semantic information.
[0038] In an embodiment of the present specification, if there is no branch discussion conflict in at least one speech text information belonging to the same speech terminal identifier, it indicates that the obtained speech text information is all the main discussion content of the meeting. Then, determine the target judgment semantic information with the current speech semantic feature corresponding to the speech text information for subsequent main line conflict judgment. If there is a branch discussion conflict in at least one speech text information belonging to the same speech terminal identifier, it indicates that there is an internal discussion situation in the speech text information corresponding to this speech terminal identifier. Therefore, it is necessary to preprocess the speech text information. First, according to the speech terminal identifier, determine multiple current speech text information belonging to the same speech terminal identifier, and extract the semantic features of the current speech text information to obtain multiple current speech semantic features. When collecting speech information, record the speech timestamp corresponding to each speech information, that is, the speech timestamp corresponding to the corresponding speech semantic feature. Generate a branch discussion text set corresponding to the speech terminal identifier according to the speech timestamp corresponding to each current speech semantic feature. It should be noted that the branch discussion text set here is multiple current speech semantic features arranged according to the speech timestamp. Summarize the branch discussion text set to obtain the definite summary semantics of the internal discussion, that is, the terminal speech semantic feature, and determine the target judgment semantic information for subsequent main line conflict judgment with the summarized terminal speech semantic feature.
[0039] Through the above technical solution, by judging whether there is a conflict in the branch discussion, the speech text information can be clearly divided into the main discussion content of the meeting and the internal discussion content (branch discussion). For the speech text information with a conflict in the branch discussion, that is, the internal discussion situation exists, by extracting semantic features and generating a set of branch discussion texts based on the speech timestamp, and further summarizing to obtain the semantic features of the terminal speech, the scattered and disordered internal discussion content can be integrated and refined in an orderly manner, presenting the complex internal discussion in a simple and definite semantic form, making the knowledge system of the entire meeting more organized, and avoiding the messy internal discussion information from mixing into the meeting knowledge system and interfering with the understanding of the key content of the meeting; whether directly determining the target judgment semantic information based on the current speech semantic features corresponding to the main discussion content of the meeting or obtaining the semantic features of the terminal speech after processing the branch discussion content to determine the target judgment semantic information, it provides unified and appropriate judgment materials for the subsequent main line conflict judgment.
[0040] Determine whether the target judgment semantic information is the main line conflict semantic, specifically including: determining the corresponding reverse semantic features through the target judgment semantic information and a pre-constructed semantic inversion model; according to the reverse semantic features, performing semantic retrieval on the knowledge stored in the RAG meeting knowledge base, and when there is specified knowledge stored in the RAG meeting knowledge base that meets the preset requirements, determine whether the target judgment semantic information is the main line conflict semantic.
[0041] In one embodiment of this specification, after determining the target judgment semantic information, it is necessary to perform a conflict judgment on the target judgment semantic information corresponding to the real-time collected voice information and the stored knowledge corresponding to the completed part of the meeting progress, that is, the judgment of the main-line conflict semantics. First, use the pre-constructed semantic inversion model to perform semantic inversion on the target judgment semantic information to determine the corresponding reverse semantic features. The semantic inversion model here can be a deep learning-based model, which is trained by collecting a large number of pairs of positive and negative opposing sentences as training data, such as obtaining from debate records, opinion comparison articles, etc. A model using the Transformer architecture (such as BERT, GPT, etc.) is trained, with the original sentence as the input and the corresponding reverse sentence as the output. During the training process, the model can learn the semantic conversion rules in the sentence. Input the target judgment semantic information that needs to be semantically inverted into the already trained semantic inversion model. After receiving the input target judgment semantic information, the model performs semantic analysis and reasoning on the input statement based on the semantic conversion rules it learned during the training process, and generates a corresponding reverse semantic statement. The semantic features corresponding to this reverse semantic statement are the reverse semantic features. According to the reverse semantic features, perform semantic retrieval on the existing stored knowledge in the RAG meeting knowledge base corresponding to the meeting progress. When performing semantic retrieval, the target retrieval object is the semantics similar to the reverse semantic features. When there is designated stored knowledge that meets the preset requirements in the RAG meeting knowledge base, determine whether the target judgment semantic information is the main-line conflict semantics. It should be noted that the preset requirement here is that the similarity between the designated stored knowledge and the reverse semantic features exceeds the preset similarity judgment threshold, and the similarity judgment threshold can be set according to the threshold corresponding to the selected similarity algorithm. If there is designated stored knowledge that meets the preset requirements in the RAG meeting knowledge base, it means that there is knowledge similar to the reverse semantic features in the RAG meeting knowledge base, that is, there is semantics in the current RAG meeting knowledge base that conflicts with the target judgment semantic information, that is, the target judgment semantic information is the main-line conflict semantics.
[0042] Through the above technical solution, by performing a main-line conflict semantic judgment on the target judgment semantic information corresponding to the real-time collected voice information and the knowledge already stored in the database, it is possible to timely discover whether the newly generated viewpoints and content conflict with the existing knowledge in the knowledge base. This can effectively prevent conflicting knowledge from continuously accumulating in the knowledge base, ensure that the knowledge system in the knowledge base is logically coherent and consistent, avoid subsequent participants making wrong decisions or understandings based on conflicting knowledge, and maintain the accuracy of the knowledge base as a knowledge storage and reference basis; by using the method of combining a semantic inversion model with semantic retrieval, the main-line conflict semantics can be accurately identified. With the help of the pre-constructed semantic inversion model and semantic retrieval in the RAG conference knowledge base, the main-line conflict semantic judgment process does not require manual comparison and analysis, reducing the workload and cost of manual intervention and improving the efficiency of knowledge management work.
[0043] In one embodiment of this specification, if the target judgment semantic information is the main-line conflict semantic, the first speaking participating terminal corresponding to the conflict semantic feature is obtained. During the processes of collecting the speaking voice information and performing corresponding speech recognition, semantic analysis, etc., for each speaking semantic feature, the identification information of the corresponding speaking participating terminal is associated and recorded. After determining that the target judgment semantic information is the main-line conflict semantic, based on the previously established association relationship, trace back to find the speaking participating terminal corresponding to the speaking that generated the conflict semantic feature, and mark it as the first speaking participating terminal. Obtain the participating role information corresponding to the current speaking participating terminal and the first speaking participating terminal, and analyze the authorities and responsibilities of different roles in the meeting. For example, if the participating role corresponding to the first speaking participating terminal is an ordinary participant, while the current speaking participating terminal corresponds to the expert role corresponding to the target judgment semantic information, then when judging the retained semantics, it may be more inclined to refer to the speaking semantics corresponding to the role with higher authority, such as the expert role. When it is determined that the retained semantics is the target judgment semantic information, it indicates that the content expressed by the target judgment semantic information is considered to be more in line with the meeting after comparison and screening with the conflict semantics, so it is determined as the meeting knowledge increment information. This part of the information can be updated to the RAG meeting knowledge base according to the established knowledge management process in the future to supplement new knowledge content to the knowledge base. For the content determined to be the conflict semantic feature, in order to ensure the standardization of knowledge management and facilitate possible subsequent traceability and viewing, it cannot be simply directly deleted, but rather archived for deletion. It can be stored in a dedicated conflict semantic archive area (such as creating a specific table or document collection in the database to store such information), and record key information such as the detailed content of the conflict semantic feature, the corresponding speaking participating terminal, the time when the conflict occurred, and which target judgment semantic information it conflicts with. Moreover, in the meeting summary stage, the information in the conflict semantic archive area can be displayed to the user, and the user can screen the meeting knowledge to be added in the summary stage. Correspondingly, if the target judgment semantic information is a non-main-line conflict semantic, the target judgment semantic is directly used as the meeting knowledge increment information and updated to the RAG meeting knowledge base in real time.
[0044] Through the above technical solutions, by judging and retaining semantics based on the participating roles as the incremental information of the meeting knowledge, it can ensure that the new knowledge incorporated into the knowledge base is reasonably screened and more authoritative and accurate, which helps to build a high-quality knowledge system and avoid the mixing of low-quality or inaccurate information; the content with conflict semantic features is deleted and filed, and the relevant key information is recorded in detail, making the entire evolution process of the meeting knowledge traceable; updating the screened incremental information of the meeting knowledge to the RAG meeting knowledge base can make the knowledge base dynamically updated as the meeting progresses, always keeping in line with the actual progress of the meeting, ensuring that the knowledge in the knowledge base is the latest and most in line with the actual needs. When the knowledge base is used for knowledge query, reference and other operations subsequently, the participating personnel can obtain more valuable information, improving the utilization efficiency of the knowledge in the knowledge base.
[0045] Step S103: Through a pre-constructed real-time interactive Q&A interface, obtain the real-time Q&A request information of the participating users, and based on the real-time Q&A request information, generate real-time Q&A interaction information through a pre-constructed large language model and the RAG meeting knowledge base.
[0046] In one embodiment of this specification, a one-to-one real-time interactive Q&A interface between participating users and the RAG system is provided through the intelligent interaction module. In the backend service of the intelligent interaction module, a route is configured to listen for requests to the corresponding interface. When a request is received, the real-time Q&A request information of the participating user is obtained. The keyword extraction algorithm in natural language processing technology is used to extract key words from the real-time Q&A request information of the participating user. In order to more accurately obtain knowledge with similar semantics to the question, semantic retrieval technology can be adopted. Using pre-trained word vector models (such as Word2Vec, BERT, etc.), the questions of users and the knowledge entries in the knowledge base are both transformed into vector representations. Then, by calculating methods such as the cosine similarity between vectors, the knowledge content with the highest semantic similarity is found. The preliminary results retrieved are screened to remove content that clearly does not meet the requirements (such as too low relevance, outdated knowledge, etc.). Sorting can be performed according to various factors, such as descending order according to keyword matching degree, semantic similarity value, knowledge update time, etc., and the most relevant and valuable knowledge is presented first, providing high-quality materials for subsequent use of the large language model to generate high-quality responses. Select a suitable large language model, integrate the relevant knowledge content retrieved from the RAG conference knowledge base (which can be document fragments in text form, knowledge point summaries, etc.) with the original real-time Q&A request information of the participating user, and construct a text format suitable for input into the large language model. Use prompt engineering techniques to optimize the prompt text input into the large model, guiding the large model to generate responses more in line with expectations. For example, clearly require the large model to answer based on the knowledge in the provided conference knowledge base, remind it to answer in a well-organized and concise language that is easy to understand, and the general structure of the answer can be specified (such as listing the answers point by point), improving the quality and pertinence of the large model output. Send the constructed prompt text containing the knowledge base retrieval results and the question to the large language model. The knowledge base of the large language model is the RAG conference knowledge base. The large model will generate the corresponding real-time Q&A interaction information response text based on its powerful language generation ability and the given input information. Integrate the processed response content and the corresponding knowledge base source information, and return it to the participating user through the real-time interactive Q&A interface to complete the entire real-time Q&A interaction process.
[0047] In one embodiment of this specification, after receiving the real-time Q&A request information, the question can also be extracted and input into the large language model. The knowledge base in the large language model is the RAG conference knowledge base that is updated in real time. The large language model uses its language understanding and generation capabilities to output the corresponding answer text to generate the real-time Q&A interaction information.
[0048] Through the above technical solution, with the help of the real-time interactive Q&A interface, participating users can immediately submit questions when they have doubts. The system can quickly process and return answers based on the language large model and the RAG meeting knowledge base, without the need for participants to spend a lot of time searching and sorting out meeting-related knowledge by themselves, greatly improving the speed of obtaining information, making the meeting communication more smooth and efficient, and avoiding the slow pace of the meeting caused by difficult information search; by combining the targeted meeting knowledge stored in the RAG meeting knowledge base and the powerful semantic understanding and generation capabilities of the language large model, the generated real-time Q&A interaction information can more accurately meet the question needs of participating users; the RAG meeting knowledge base itself converges the meeting knowledge corresponding to the meeting progress, but if there is a lack of effective retrieval and application methods, this knowledge is easily idle. Through the form of real-time Q&A interaction, participating users can dig out valuable information from the knowledge base at any time; for users who join the meeting late, even if they have not participated in the previous meeting, they can obtain the current meeting progress and the meeting knowledge of the unparticipated meeting process through the form of real-time Q&A interaction.
[0049] Step S104: Generate meeting summary information within the current meeting space corresponding to each participating user based on the RAG meeting knowledge base and the real-time Q&A interaction information corresponding to each participating user.
[0050] The concerns of different participating users are different. Conventional meeting summary information is extensive and does not highlight key points. When users search for target information in the meeting summary information, they need to search item by item, which is relatively cumbersome, resulting in low utilization rate of the meeting summary information. In an embodiment of this specification, the RAG meeting knowledge base generated in real time according to the meeting process is used, and combined with the real-time Q&A interaction information corresponding to each participating user, to generate meeting summary information within the current meeting space corresponding to each participating user.
[0051] Through the above technical solution, different participating users participate in the meeting with different purposes and concerns. By combining their real-time Q&A interaction information during the meeting, the generated meeting summary can accurately focus on the content that individuals care about; abandoning the previous long and unfocused meeting summary mode, participating users do not need to search for target content line by line in a large amount of conventional information.
[0052] Based on the RAG meeting knowledge base and the real-time Q&A interaction information corresponding to each participating user, generate the meeting summary information within the current meeting space for each such participating user, specifically including: determining the meeting attention characteristics of the participating user through the real-time Q&A interaction information corresponding to each such participating user; adjusting the preset RAG meeting knowledge base architecture based on the meeting attention characteristics to generate a customized RAG meeting knowledge base architecture for each such participating user; integrating the meeting knowledge in the RAG meeting knowledge base according to the customized RAG meeting knowledge base architecture to generate the meeting summary information within the current meeting space for each such participating user.
[0053] In an embodiment of the present specification, the Q&A situation of the participating user usually reflects the user's meeting focus. Therefore, the real-time Q&A interaction information corresponding to each participating user is used to determine the meeting attention characteristics of the participating user. The meeting attention characteristics here include the meeting attention topics and the attention factors corresponding to each meeting attention topic. Determine the arrangement order of multiple meeting attention topics in descending order of the attention factors corresponding to each meeting attention topic, and set the un-involved architecture topics after the meeting attention topics. Adjust the preset RAG meeting knowledge base architecture in the above manner to generate a customized RAG meeting knowledge base architecture for each such participating user. Integrate the meeting knowledge in the real-time updated RAG meeting knowledge base according to the preset RAG meeting knowledge base architecture to obtain the topic meeting summary information under each architecture topic. Then, sort the above topic meeting summary information according to the topic order of the customized RAG meeting knowledge base architecture to generate the meeting summary information within the current meeting space for each such participating user.
[0054] Through the above technical solution, by analyzing the real-time Q&A interaction information of the participating user to determine the meeting attention characteristics, it is possible to accurately capture the unique attention points of each user in the meeting, covering specific attention topics and the attention factors corresponding to each topic. This makes the generated meeting summary information closely revolve around the content that the user truly cares about, avoiding the problem of excessive irrelevant information and lack of focus in traditional meeting summaries; sorting the topics according to the size of the attention factors corresponding to the meeting attention topics and reasonably placing the un-involved architecture topics at the back makes the presentation order of the meeting summary information more logical. What the user sees first is the topic content that they are most concerned about and has a higher importance level, facilitating them to sequentially understand various aspects of knowledge along the logical clue, avoiding the piling up of chaotic information, enabling the user to have a clearer and more organized understanding of the meeting knowledge, and helping to deepen the understanding of the overall meeting content.
[0055] Based on the real-time Q&A interaction information corresponding to each participating user, determine the meeting attention characteristics of the participating user, specifically including: obtaining multiple question descriptions in the real-time Q&A interaction information, extracting keywords from the multiple question descriptions, and determining multiple Q&A keywords corresponding to each participating user; matching the multiple Q&A keywords with the architecture themes in the RAG meeting knowledge base architecture to determine the matching architecture theme corresponding to each Q&A keyword, and determining the meeting attention theme corresponding to the participating user based on the matching architecture theme; counting the Q&A frequencies of the multiple Q&A keywords to determine the number of Q&A times corresponding to each Q&A keyword; determining the attention factor corresponding to each matching architecture theme according to the matching architecture theme corresponding to each Q&A keyword and the number of Q&A times corresponding to each Q&A keyword; and determining the meeting attention characteristics of the participating user based on the meeting attention theme and the corresponding attention factor.
[0056] In one embodiment of the present specification, obtain multiple question descriptions in the real-time Q&A interaction information. Question descriptions are usually the focus points of users. Extract keywords from the multiple question descriptions, determine multiple Q&A keywords corresponding to each participating user, match the Q&A keywords with the architecture themes in the RAG meeting knowledge base architecture, and determine which architecture theme the Q&A keyword belongs to. Then, the corresponding matching architecture theme is the meeting attention theme of the user. In addition, count the Q&A frequencies of the multiple Q&A keywords to determine the number of Q&A times corresponding to each Q&A keyword. Determine the attention factor of the matching architecture theme corresponding to this Q&A keyword according to the ratio of the number of Q&A times corresponding to each Q&A keyword to the total cumulative Q&A times. Determine the obtained meeting attention theme and the corresponding attention factor as the meeting attention characteristics of the participating user.
[0057] Through the above technical solution, by extracting keywords from the question descriptions in the real-time Q&A interaction information and matching them with the architecture themes in the RAG meeting knowledge base architecture, it is possible to accurately locate the specific attention themes of each participating user in the meeting. This analysis based on the actual question content of the user avoids subjective speculation and truly and accurately reflects the meeting content that the user wants to understand and attach importance to; on the basis of determining the attention theme, further count the occurrence frequency of the Q&A keywords, and calculate the attention factor of the corresponding matching architecture theme, realizing the quantitative measurement of the user's attention degree.
[0058] Through the technical solution provided in the embodiments of this specification, by collecting speech voice information in real time in the current conference space and performing speech recognition to generate identification text information containing detailed text speech attributes, it is possible to comprehensively and meticulously capture various speech situations in the meeting, ensuring that information in complex scenarios such as single-person speech or multi-person speech on a shared terminal can be accurately recorded; conflict resolution of speech text information based on text speech attributes can identify truly valuable incremental conference knowledge information, avoid knowledge confusion caused by conflicts of views and repeated expressions of different participants, and ensure that only high-quality knowledge that conforms to the logic of the meeting and has been screened will be updated to the RAG conference knowledge base, so that the knowledge in the knowledge base always maintains accuracy and consistency, and keeps pace with the times, and is dynamically updated in line with the actual progress of the meeting; with the help of a pre-built real-time interactive question-and-answer interface, participating users can submit real-time question-and-answer request information at any time, and quickly obtain real-time question-and-answer interactive information generated based on the language large model and the RAG conference knowledge base. The convenient interactive method allows participants to get targeted answers to the meeting content once they encounter questions during the meeting without waiting for the meeting to end, greatly improving the meeting experience. In addition, for late-joining users, even if they have not participated in previous meetings, they can obtain the current meeting progress and meeting knowledge that they have not participated in the meeting process through real-time question-and-answer interaction. Personalized meeting summary information is generated based on the RAG meeting knowledge base and the real-time question-and-answer interaction information corresponding to each participating user, fully considering the differences in the focus and knowledge needs of different participants. The meeting summary received by each participant is a customized version that fits their own situation, which can quickly review the key points of the meeting and absorb useful knowledge for themselves, saving time and energy in sorting out knowledge after the meeting. The real-time updated RAG meeting knowledge base, real-time interactive question-and-answer, and personalized meeting summary information after knowledge conflict resolution meet the needs of participants for real-time interaction and personalized services in complex meeting scenarios where multiple people use one terminal device to attend meetings.
[0059] Figure 2This is a schematic diagram of the application scenario of an intelligent conference full - process management method provided by the embodiments of this specification. The above - mentioned technical solution is applied to an intelligent conference management system, which includes a meeting initialization module, a real - time speech transcription module, a knowledge base construction module, an intelligent interaction module, a meeting summary module, and a knowledge retrieval module. The meeting initialization module is used to establish a meeting space and record the identity information of the participants, and at the same time initialize the RAG knowledge base system exclusive to the meeting. This module creates an independent meeting space by generating a unique meeting identifier, records information such as the names, identities, and permissions of the participants, and builds a customized RAG knowledge base framework based on the meeting theme and participant information, laying a foundation for subsequent knowledge storage and retrieval. The real - time speech transcription module is used to collect the speech information during the meeting, separate the speech through speaker recognition technology, and convert the speech into text information with speaker identification (each sentence in the text will be preceded by the speaker's name and other identifications) in real - time. This module captures the meeting audio in real - time, and at the same time combines noise reduction and audio enhancement technologies to ensure the quality of speech transcription, and converts the processed speech information into structured text with timestamps and speaker identifications.
[0060] The knowledge base construction module is used to perform structured processing and vectorization on the transcribed text, build a knowledge base index that supports semantic retrieval, and realize the real - time update of the knowledge base. This module first segments the text, and then uses text vectorization technology to convert the processed text into a high - dimensional vector representation, builds a knowledge base based on vector index, supports similarity retrieval at the semantic level, and at the same time realizes incremental update and conflict handling of knowledge to ensure the consistency and timeliness of the knowledge base. The intelligent interaction module is used to provide a real - time one - on - one question - answering interface between the meeting participants and the RAG system, and perform question analysis and answer generation based on the real - time meeting knowledge base. This module provides an independent question - answering interaction interface for each participant, through semantic understanding and intention analysis of the user - input question, retrieves relevant information from the real - time updated knowledge base in combination with the meeting context, uses a large - model to combine and reason about the retrieved content, generates accurate answers that meet the user's needs, and supports multi - round conversations and answer optimization.
[0061] The meeting summary module is used to analyze the meeting content using a large model, extract key information and important decisions, and automatically generate a structured meeting minutes. This module conducts multi-dimensional analysis on the entire meeting content, including theme context sorting, key point extraction, decision induction, etc. It identifies important content through semantic understanding and information extraction technologies, and automatically generates meeting minutes in a standardized format. It also supports content summarization at different granularities and personalized display methods. The knowledge retrieval module is used to support question-and-answer interactions based on natural language after the meeting, and achieve multi-dimensional retrieval and personalized display of meeting knowledge. This module provides flexible retrieval interfaces, supports various retrieval methods such as keyword-based, semantic similarity-based, and time series-based, can understand the user's natural language questions and convert them into precise retrieval requirements, provides complete answers through knowledge association analysis and reasoning supplementation, and organizes and displays information in a personalized manner according to the user's knowledge background and concerns.
[0062] As Figure 2 shown, in the meeting initialization stage, the system establishes a unique meeting identifier, records the information of the participants, and initializes the RAG knowledge base dedicated to this meeting. During the meeting, the meeting audio is collected through the real-time speech transcription module and the speaker is identified, and the speech is converted into text with speaker identification in real time. The transcribed text information is structured and vectorized, and updated to the meeting RAG knowledge base in real time. It supports the meeting participants to have real-time one-on-one Q&A with the RAG system through an independent interaction interface, and the system provides intelligent answers based on the current meeting knowledge base. After the meeting, the large model is used to conduct multi-dimensional analysis on the meeting content, automatically generate structured meeting minutes and store them in the knowledge base. A meeting knowledge retrieval module is established to support users to conduct question-and-answer interactions based on natural language after the meeting, and achieve efficient retrieval and reuse of meeting knowledge. By realizing the real-time capture, organization, and intelligent interaction of meeting knowledge, it solves the technical problems such as the lack of real-time knowledge support, difficulty in in-depth mining and reuse of content in traditional meeting systems, and improves the meeting efficiency and knowledge acquisition efficiency.
[0063] This embodiment of the specification also provides an intelligent meeting full-process management device based on RAG, as Figure 3 shown, the device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above method.
[0064] This embodiment of the specification also provides a non-volatile computer storage medium storing computer-executable instructions, and the computer-executable instructions are set to: execute the above method.
[0065] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the embodiments of the apparatus, device, and non-volatile computer storage medium, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the corresponding parts of the method embodiments for relevant content.
[0066] The specific embodiments of this specification have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the specific order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0067] The devices and media provided in the embodiments of this specification correspond one-to-one with the methods. Therefore, the devices and media also have beneficial technical effects similar to those of their corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be elaborated here.
[0068] Those skilled in the art should understand that the embodiments of this specification can be provided as methods, systems, or computer program products. Therefore, this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0069] This specification is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of this specification. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0070] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 specified in one block or multiple blocks.
[0071] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are performed on the computer or other programmable device to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 specified in one block or multiple blocks.
[0072] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0073] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). Memory is an example of computer-readable media.
[0074] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage, or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.
[0075] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising said element.
[0076] The above description is only one or more embodiments of this specification and is not intended to limit this specification. For those skilled in the art, various changes and modifications can be made to one or more embodiments of this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of one or more embodiments of this specification shall be included within the scope of the claims of this specification.
Claims
1. A RAG-based intelligent conference full-process management method, characterized in that: The method comprises: Collect speech voice information corresponding to each participating terminal in real time in the current conference space, perform speech recognition on each speech voice information, and generate identification text information, wherein the identification text information includes speech text information and corresponding text speech attributes, and the text speech attributes include a speech terminal identifier and a unique identifier of the speaking user; According to the text speech attribute, conflict processing is performed on the speech text information to determine the incremental conference knowledge information corresponding to the speech text information, so as to update the incremental conference knowledge information in real time to the pre-built RAG conference knowledge base corresponding to the current conference space; Acquire real-time question-and-answer request information of conference participants through a pre-built real-time interactive question-and-answer interface, and generate real-time question-and-answer interactive information based on the question-and-answer request information through a pre-built language model and the RAG conference knowledge base; Based on the RAG conference knowledge base and the real-time question-and-answer interaction information corresponding to each participating user, the conference summary information in the current conference space corresponding to each participating user is generated.
2. According to the RAG-based intelligent conference full-process management method of claim 1, it is characterized in that: Before collecting speech voice information corresponding to each participating terminal in real time in the current conference space, the method further includes: Under the triggering of the conference management terminal, the current conference space is constructed at the conference creation node; Acquire conference participant information and conference theme information of the target conference, wherein the conference participant information includes information of multiple participating terminals, at least one participating user corresponding to each participating terminal, and participation rights of each participating user; Based on the conference participant information of the target conference, the participant roles in the current conference space are set, and through the conference theme information and the participant roles, the knowledge base framework of the RAG conference knowledge base corresponding to the current conference space is set.
3. According to the RAG-based intelligent conference full-process management method of claim 1, it is characterized in that: Performing speech recognition on each of the speech voice information to generate identification text information specifically includes: Determine the speaking participant terminal corresponding to the speaking voice information, so as to determine the speaking terminal identifier corresponding to each speaking voice information; According to the speaking terminal identifier corresponding to each of the speaking voice information, timbre recognition is performed on at least one designated speaking voice information belonging to the same speaking conference participating terminal to determine the timbre feature corresponding to each of the designated speaking voice information; Perform similarity matching on the plurality of timbre features to determine a timbre similarity matching result between each of the designated speech voice information; determine a unique identifier of a speaker of the speech voice information corresponding to the participating terminal based on the timbre similarity matching result between each of the designated speech voice information; Each of the speech voice information is converted into text to determine the corresponding speech text information, and the text speech attribute corresponding to the speech text information is set with the speech terminal identifier and the unique identifier of the speaking user.
4. According to the RAG-based intelligent conference full-process management method of claim 1, it is characterized in that: According to the text speech attribute, conflict processing is performed on the speech text information to determine the incremental conference knowledge information corresponding to the speech text information, specifically including: According to the speech text attributes corresponding to the speech text information, determining whether at least one speech text information belonging to the same speech terminal identifier has a branch discussion conflict; Determine target judgment semantic information based on at least one speech text information through the judgment result to determine whether the target judgment semantic information is a main line conflict semantics; If the target judgment semantic information is the main line conflict semantics, then obtaining the first speaking conference participant terminal corresponding to the conflict semantic feature, and determining the reserved semantics according to the current speaking conference participant terminal and the first speaking conference participant terminal corresponding to the target judgment semantic information; When the retained semantics is the target judgment semantic information, the target judgment semantic information is determined to be conference knowledge increment information, and the conflicting semantic features are deleted and archived.
5. According to the RAG-based intelligent conference full-process management method of claim 4, it is characterized in that: Determining whether the target judgment semantic information is a main line conflict semantics specifically includes: Determine the corresponding reverse semantic features through the target judgment semantic information and the pre-built semantic reversal model; According to the reverse semantic features, a semantic search is performed on the stored knowledge in the RAG conference knowledge base. When there is designated stored knowledge that meets the preset requirements in the RAG conference knowledge base, it is determined whether the target judgment semantic information is the main line conflict semantics.
6. According to the RAG-based intelligent conference full-process management method of claim 4, it is characterized in that: Determining target judgment semantic information based on the judgment result and at least one speech text information, specifically including: If at least one speech text information belonging to the same speech terminal identifier does not have a branch discussion conflict, determining the target judgment semantic information based on the current speech semantic feature corresponding to the speech text information; If at least one speech text information belonging to the same speech terminal identifier has a branch discussion conflict, determining multiple current speech semantic features belonging to the same speech terminal identifier according to the speech terminal identifier; Generate a branch discussion text set corresponding to the speech terminal identifier according to the speech timestamp corresponding to each of the current speech semantic features obtained in advance; The branch discussion text set is summarized and concluded, the terminal speech semantic features corresponding to the branch discussion text set are determined, and the target judgment semantic information is determined.
7. According to the RAG-based intelligent conference full-process management method of claim 1, it is characterized in that: Based on the RAG conference knowledge base and the real-time question-and-answer interaction information corresponding to each participating user, generating conference summary information in the current conference space corresponding to each participating user, specifically including: Determining conference attention features of each conference participant through real-time question-and-answer interaction information corresponding to each conference participant; Based on the conference attention features, the preset RAG conference knowledge base architecture is adjusted to generate a customized RAG conference knowledge base architecture corresponding to each of the participating users; According to the customized RAG conference knowledge base architecture, the conference knowledge in the RAG conference knowledge base is integrated to generate conference summary information in the current conference space corresponding to each participating user.
8. According to the RAG-based intelligent conference full-process management method of claim 7, it is characterized in that: Determining the conference attention features of each conference participant through the real-time question-and-answer interaction information corresponding to each conference participant, specifically including: Acquire multiple question descriptions in the real-time question-and-answer interaction information, extract keywords from the multiple question descriptions, and determine multiple question-and-answer keywords corresponding to each of the conference participants; Match the multiple question-and-answer keywords with the architecture topics in the RAG conference knowledge base architecture, determine the matching architecture topic corresponding to each of the question-and-answer keywords, and determine the conference focus topic corresponding to the participating user based on the matching architecture topic; Performing question and answer frequency statistics on the multiple question and answer keywords to determine the number of questions and answers corresponding to each of the question and answer keywords; Determine the attention factor corresponding to each matching architecture theme according to the matching architecture theme corresponding to each question and answer keyword and the number of questions and answers corresponding to each question and answer keyword; The conference attention features of the conference participants are determined based on the conference attention topics and corresponding attention factors.
9. A RAG-based intelligent conference full-process management device, characterized in that: The device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform the method according to any one of claims 1 to 8.
10. A non-volatile computer storage medium storing computer executable instructions, characterized in that: The computer executable instructions are configured to execute the method according to any one of claims 1 to 8.
Citation Information
Cited By
Voice separation and paragraph affiliation method and system for teleconference scene
CN120808757A
Conference management method and device
CN120856488A
Collaborative dialogue driven specification and software automatic co-construction method and device
CN121935357A