AI digital human conference proxy method and device under off-line local area network and medium
By combining a localized topic knowledge base and user history data under offline LAN conditions, voice data is collected in real time and personalized responses are generated. This solves the problem of real-time personalized responses and topic prediction for AI digital human conference agents, achieving full-process agent security and interactive authenticity, and meeting the meeting needs when key decision-makers are absent.
Patent Information
- Application Number
- CN202511625945.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-07
- Publication Date
- 2026-02-27
AI Technical Summary
In an offline LAN environment, AI digital human meeting agents struggle to achieve real-time personalized response generation and meeting agenda prediction, failing to meet the full-process proxy needs when key decision-makers are absent. Furthermore, there are risks of data transmission leakage and insufficient authenticity in interactions.
By collecting voice data in real time and converting it into text, combined with a localized industry topic knowledge base and the target user's historical meeting trajectory, core topics are predicted and personalized responses are generated. Real-time responses are achieved in an offline environment using pre-loaded user feature packages. By combining multi-level topology structure and semantic logic matching, the relevance and timeliness of the responses are ensured.
It enables real-time personalized response generation and meeting topic prediction by AI digital human meeting agents in offline LAN environments, improving the authenticity and efficiency of meeting interactions and meeting the security and compliance requirements of the entire process of agency.
Smart Images

Figure CN121585789A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of artificial intelligence, in particular to an AI digital human conference proxy method and device under an offline local area network and a medium. BACKGROUND
[0002] In the current high-security demand and data-sensitive industries such as finance and medical treatment, important conferences need to rely on offline local area networks to carry out to completely avoid the risk of leakage caused by data transmission to the outside world, but core decision-makers often cannot attend due to time conflicts, and rely on conference recording and manual transmission, which can easily cause conference information transmission gaps and delay decision-making progress. An intelligent conference proxy solution that adapts to offline environments is urgently needed. Traditional AI conference proxy technology relies on the cloud to complete reasoning and data storage, and can realize related operations such as triggering user participation, basic communication, or fixed process execution. The digital human solution can output standardized voice and actions, and usually uses online large models to support related AI functions, and the overall operation relies on cloud computing power and data interaction.
[0003] However, traditional AI conference proxy technology relies on cloud reasoning and data storage and cannot run in offline local area networks, which does not meet the compliance requirements of industries such as finance and government that "conference data is not out of the domain" and has the risk of data transmission leakage; the proxy capability is fragmented, and can only trigger user participation, basic communication, or fixed processes, lacks the full-process proxy capability of pre-conference configuration, real-time interaction during the conference, and post-conference structured minutes, and cannot truly replace core personnel to attend the meeting; the digital human related solution lacks user style simulation function, can only output standardized voice and actions, and is difficult to restore the speaking habits of specific personnel, resulting in insufficient authenticity of conference interaction; and it is not optimized for the computing power limitation and data isolation requirements of offline local area networks, such as online large models that cannot run on edge nodes, resulting in complete failure of AI functions in offline scenarios. SUMMARY
[0004] The present application provides an AI digital human conference proxy method, device and medium under an offline local area network to solve the problem that real-time personalized reply generation and conference topic prediction of AI digital human conference proxy cannot be cooperatively implemented in offline local area network environment.
[0005] To achieve the above-mentioned purpose, the present application provides an AI digital human conference proxy method under an offline local area network, comprising:
[0006] real-time collection of voice data of a target conference, conversion of the voice data into real-time text;
[0007] in combination with a localized industry topic knowledge base and a historical conference track of a target user, prediction of a plurality of target core topics in the future according to the lexical density and semantic logic in the real-time text, to obtain a set of predicted target core topics;
[0008] If the real-time text detects mentions of the target user or matters to be decided, a first response voice is generated by combining the target user's historical speech data; if the target meeting enters the predicted target core topic set, a second response voice is generated by calling the pre-loaded target user feature package.
[0009] Based on the first or second response voice, the digital human is controlled to generate corresponding audio and video stream outputs.
[0010] This invention provides an instantly interpretable data source for subsequent analysis through real-time speech-to-text conversion, ensuring the entire process keeps pace with the meeting. By combining a localized industry-specific knowledge base and user meeting history, it avoids the limitations of offline environments where external data cannot be accessed online, and allows topic prediction based on lexical density and semantic logic to align with industry characteristics and user meeting habits, identifying core topics in advance. This design transforms topic prediction from an isolated process, providing clear direction for subsequent personalized responses and creating a synergy between topic prediction and response generation, thus solving the problem of their lack of coordination. When a mention of a target user or a decision-making matter is detected, a first response is generated using locally stored user history data. This utilizes offline personalized data to ensure the response aligns with user habits while meeting real-time requirements through an instant triggering mechanism. When the meeting reaches the predicted core topic, a second response is generated by calling a pre-loaded user feature package, directly transforming the previously predicted topic results into the basis for the response. This ensures that topic prediction is not an isolated preliminary step but forms a data loop with response generation. Both rely on locally stored user data in an offline environment, avoiding delays caused by network dependence. Finally, by converting the first response voice related to the real-time mentioned scenario and the second response voice related to the predicted core topic into a digital human audio and video stream output, the real-time personalized response generation and the meeting topic prediction results can be synchronously presented in an offline environment, thereby solving the problem of the two being difficult to coordinate.
[0011] Compared to existing technologies, this invention achieves topic prediction by combining a localized industry topic knowledge base and user historical meeting trajectories, providing advance preparation for real-time responses. Simultaneously, for scenarios involving mentioning users or pending decisions, and entering predicted topics, it generates responses by calling upon user historical data and pre-loaded user feature packages. This leverages localized data to ensure real-time performance in offline environments. Through the connection between prediction and scenario-specific responses, topic prediction provides precise support for personalized responses, achieving synergy between the two. Therefore, it solves the problem of the difficulty in achieving coordinated real-time personalized response generation and meeting topic prediction in offline LAN environments using AI digital human meeting agents.
[0012] As a preferred embodiment, the method for acquiring the localized industry topic knowledge base is as follows:
[0013] obtaining enterprise historical conference data in a target local area network;
[0014] Based on the enterprise historical conference data, a vertical industry topic graph library is obtained by building a multi-level topology structure of core topics, sub-topics and associated attributes;
[0015] Based on the enterprise historical conference data, for each topic in the vertical industry topic graph library, the corresponding high-frequency discussion dimension and user historical corpus are associated to obtain the localized industry topic knowledge base; wherein the high-frequency discussion dimension refers to a topic dimension whose preset keyword appearance frequency exceeds a preset number of times.
[0016] The preferred scheme is based on the enterprise historical conference data in the local area network, builds a multi-level topology structure of the vertical industry topic graph library, and associates the high-frequency discussion dimension and the user historical corpus, so that the knowledge base closely fits the enterprise business scenario. The focus of the high-frequency discussion dimension can reduce redundant information interference and make the subsequent topic prediction and reply generation more accurate; the localization feature ensures the efficiency of knowledge calling in offline environment, avoids the delay or data leakage risk caused by relying on external network, and provides knowledge support for digital human conference agent that fits the actual enterprise.
[0017] As a preferred scheme, based on the enterprise historical conference data, a vertical industry topic graph library is obtained by building a multi-level topology structure of core topics, sub-topics and associated attributes, specifically:
[0018] The enterprise historical conference data is subjected to data cleaning and industry attribute labeling to obtain basic materials;
[0019] Based on the basic materials, a number of industry core topics are obtained by double-dimensional weighted screening of decision correlation and discussion proportion, a number of conference dimensions are extracted from the text corresponding to the number of industry core topics as a number of sub-topics, and a hierarchical skeleton about the number of industry core topics and the number of sub-topics is constructed; wherein the decision correlation refers to the proportion of the number of co-occurrences of statistical terms and preset decision class words, and the discussion proportion refers to the proportion of the duration of the discussion paragraph corresponding to the terms in the total conference duration;
[0020] Based on the hierarchical skeleton, the historical discussion dimension data and user historical decision tendency data in the number of sub-topics are integrated into a multi-level topology structure to obtain the vertical industry topic graph library.
[0021] The preferred scheme can ensure the standardization of the basic material through data cleaning and industry attribute labeling; the weighted screening of the decision correlation degree and the discussion proportion avoids the misjudgment of the core issues caused by a single dimension, and ensures that the selected issues have both decision value and discussion heat. The hierarchical skeleton combines the historical discussion dimension and the user decision tendency, so that the graph library is structured and systematic, facilitating the subsequent quick matching of real-time conference content, and laying a structured knowledge foundation for accurately predicting core issues.
[0022] As a preferred scheme, in combination with the localized industry issue knowledge base and the historical conference track of the target user, a plurality of target core issues are predicted according to the word density and semantic logic in the real-time text, to obtain a set of predicted target core issues, specifically:
[0023] Taking the plurality of industry core issues in the localized industry issue knowledge base as a benchmark, the word appearance frequency proportion of the real-time text and the localized industry issue knowledge base is counted to obtain a word density matching degree;
[0024] The structural similarity between the real-time text semantic logic of the real-time text and the issue logic in the localized industry issue knowledge base is calculated by a semantic matching algorithm to obtain a semantic logic chain matching degree, and the word density matching degree and the semantic logic chain matching degree are weighted and summed to obtain a real-time text feature matching degree;
[0025] The historical conference track of the target user is called, and the similarity between the real-time text features of the real-time text and the features of the same type of issue discussion stage in the historical conference track is calculated according to an offline cosine similarity algorithm to obtain a user historical track similarity;
[0026] The real-time text feature matching degree and the user historical track similarity are weighted in two dimensions to obtain a comprehensive score of each issue in the plurality of industry core issues, and the top several issues according to the score from high to low are taken as the set of predicted target core issues.
[0027] In the preferred scheme, the word density matching degree quantifies the association of keywords, and the semantic logic matching degree captures the internal logic of the issue, and the real-time text feature matching degree obtained by weighting the two can ensure the fit with the knowledge base; the user historical track similarity is associated with the historical features of the same type of issue by the cosine similarity algorithm, reflecting the user's meeting habits. The comprehensive score mechanism of two-dimensional weighting makes the prediction result consistent with the real-time trend of the meeting and in line with the user's historical behavior, improving the accuracy of the prediction of the target core issue, enabling the digital person to prepare in advance and enhancing the pertinence of the proxy response.
[0028] As a preferred solution, whether the target meeting enters the set of pre-judgment target core topics is determined by real-time statistics of the matching word frequency proportion of the real-time text and a preset keyword library in the set of pre-judgment target core topics, calculation of the text semantic logic of the real-time text and the structural similarity of the preset typical logic template in the set of pre-judgment target core topics, and verification of the cosine similarity of the text features of the real-time text and the historical similar topic discussion stage features of the target user.
[0029] The preferred solution can quickly capture topic-related keyword signals by real-time statistics of the matching word frequency proportion, avoid single word misjudgment, calculate the structural similarity of the text semantic logic and the typical template to deeply understand the internal logical association of the topic discussion and overcome the surface limitation of keyword matching, and verify the cosine similarity of the text features and the historical similar topic stage features of the user to combine the user behavior habits and strengthen the personalized adaptability of the judgment. The three are combined to form a complementary verification mechanism, which not only ensures efficient real-time response in an offline local area network environment, but also accurately identifies the topic switching node, provides reliable basis for subsequent calling of the user feature package to generate reply voice that meets the needs, and thus improves the response accuracy and meeting participation fluency of the AI digital human meeting agent.
[0030] As a preferred solution, before the digital human generates the corresponding audio and video stream output according to the first reply voice or the second reply voice, the method further includes:
[0031] Extracting temporary decision-making tendency data and target new term preference data of a plurality of users from the voice data to generate a real-time feature increment package;
[0032] If the real-time meeting features in the real-time feature increment package conflict with the reply features of the first reply voice or the second reply voice, the reply logic of the first reply voice or the second reply voice is modified in combination with a time weight determination result. The time weight determination result includes that the recent feature weight is greater than the long-term feature weight, and the real-time meeting feature weight is greater than the historical static feature weight.
[0033] The temporary decision-making tendency and term preference extracted from the voice data can capture the real-time dynamics in the meeting. When the real-time features conflict with the pre-generated reply features, the reply logic can be adjusted according to the weight rules that the recent features are better than the long-term features and the real-time features are better than the historical static features, to ensure that the digital human reply meets the current meeting situation. This dynamic adjustment mechanism enhances the timeliness and flexibility of the reply, avoids the inappropriateness caused by mechanical dependence on historical data, and improves the on-site adaptability of the digital human agent.
[0034] As a preferred solution, before the real-time collection of the voice data of the target meeting, the method further includes:
[0035] obtaining a personalized configuration and uploading to a local area network server; wherein the personalized configuration includes the historical corpus of the target user and a conference ID;
[0036] controlling the edge node to load a large language model, a speech recognition model and a speech synthesis model into memory within a preset time before the start of the target conference;
[0037] controlling the conference access module of the edge node to initiate an access test, after verifying that the audio and video channel to the system where the target conference is located is unobstructed, log in to the target conference relying on the conference ID, and push the initial video stream of the digital person to the conference interface of the target conference.
[0038] The preferred scheme focuses on the preparation process before the conference, and solves the problems of model loading delay, poor conference access and other problems in offline environment. The uploading of personalized configuration ensures that the digital person fits the characteristics of the target user; the preloading of large models and audio and video processing models into memory within a preset time reduces the response delay in the conference; the access test verifies the audio and video channel to avoid technical failures in the conference; relying on the conference ID to log in and push the initial video stream realizes the seamless connection of the digital person and the conference system. These preparations guarantee the stability and efficiency of the conference agent in the offline local area network environment, and improve the user's experience of using the digital person agent.
[0039] As a preferred scheme, after the digital person generates the corresponding audio and video stream output according to the first reply voice or the second reply voice, it further includes:
[0040] integrating all summary segments in the target conference to generate a structured summary;
[0041] encrypting the structured summary based on a security management module, and sending a summary push notification to the terminal of the target user; wherein the security management module processes the structured summary through a preset encryption algorithm, and realizes data encryption through password authentication, device fingerprint authentication and dynamic password three-factor authentication.
[0042] The preferred scheme solves the problems of low conference information sorting efficiency and high data security risk through post-conference structured summary generation and encryption push. The structured summary integrates conference segments, which facilitates users to quickly review the core content and improves work efficiency; the security management module uses a preset encryption algorithm combined with password authentication, device fingerprint authentication and dynamic password three-factor authentication to strengthen data confidentiality in offline local area network environment, preventing sensitive conference information from being leaked. Pushing the summary notification to the user terminal in a timely manner ensures that the user can quickly obtain the conference results, taking into account the information utilization efficiency and security protection, and perfecting the closed-loop service of the digital person conference agent.
[0043] The application also provides an AI digital person conference agent device under an offline local area network, comprising a data module, a pre-judgment module, a voice module and a reply module.
[0044] The data module is configured to collect voice data of a target conference in real time, and convert the voice data into real-time text.
[0045] The pre-judgment module is configured to combine a localized industry topic knowledge base and a historical conference track of a target user, and pre-judge a plurality of target core topics according to the lexical density and semantic logic in the real-time text to obtain a pre-judged target core topic set.
[0046] The voice module is configured to generate a first reply voice in combination with historical data of the target user if the real-time text detects a mention of the target user or a pending matter, and generate a second reply voice by calling a preloaded target user feature package if the target conference enters the pre-judged target core topic set.
[0047] The reply module is configured to control a digital person to generate a corresponding audio and video stream output according to the first reply voice or the second reply voice.
[0048] The application also provides a storage medium having a computer program stored thereon, wherein the computer program is invoked and executed by a computer to implement the AI digital person conference agent method under an offline local area network. BRIEF DESCRIPTION OF DRAWINGS
[0049] Figure 1 FIG. 1 is a flow diagram of an AI digital person conference agent method under an offline local area network provided by an embodiment of the application;
[0050] Figure 2 FIG. 2 is a structural diagram of an AI digital person conference agent device under an offline local area network provided by an embodiment of the application. DETAILED DESCRIPTION
[0051] The technical solutions in the embodiments of the application will be described clearly and completely below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only part of the embodiments of the application, rather than all the embodiments of the application. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the application.
[0052] In the description of the present application, it should be understood that the terms "first" and "second" are only for the purpose of description, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features. Therefore, the features defined with "first" and "second" can explicitly or implicitly include one or more of the features. In the description of the present application, unless otherwise specified, the meaning of "several" is two or more.
[0053] The AI digital human conference agent method under the offline local area network provided by the embodiment of the present application aims to build a cloud-free dependent full-process conference agent architecture to meet the safety compliance requirements of "data not out of domain", realize real-time inference of large models and digital human interaction under the limited computing power of edge nodes, and achieve "pre-meeting configuration, meeting interaction, and post-meeting summary" full-process agent. Through technical optimization, AI digital human can accurately simulate the speaking style and decision-making tendency of a specific user to improve the authenticity of interaction, and realize structured processing of conference content to ensure the completeness of the summary and the accuracy of information transmission. Ultimately, it supports full-process conference agent when core personnel are absent, and takes into account safety compliance, interaction authenticity and efficiency improvement.
[0054] Embodiment I:
[0055] Please refer to Figure 1 The embodiment of the present application provides an AI digital human conference agent method under an offline local area network, including S1-S4, and the specific implementation steps are as follows:
[0056] S1, real-time acquisition of voice data of a target conference, and conversion of the voice data into real-time text.
[0057] The embodiment of the present application step S1 is specifically:
[0058] Obtain the personalized configuration information of the target user, which includes the historical corpus of the target user, the meeting ID and the interaction permission, and then upload these information to the local area network server; after the local area network server receives it, the personalized configuration will be stored in the encrypted database to ensure the security of the user's core data in the offline environment; wherein the target user is a user who is absent from the target conference, and the embodiment of the present application aims to provide full-process conference agent for this absent scenario;
[0059] Within a preset period before the formal start of the target meeting, the local area network server issues a model warming-up instruction to the edge computing node; after responding to the instruction, the edge node loads the large language model, the speech recognition model, and the speech synthesis model into the local memory, completes the initialization of the AI inference environment required for the meeting, and prepares for subsequent real-time processing; the preset period can be 30 minutes. The large language model can specifically use the DeepSeek model or the Llama 2 model; the speech recognition model can use the Vosk model or the CMU Sphinx offline model; and the speech synthesis model can use the Text-to-Speech model (TTS).
[0060] The edge node initiates an audio and video channel access test through the built-in conference access module to verify whether the audio and video data transmission between itself and the system where the target meeting is located is smooth; after the test passes, the edge node logs into the target meeting relying on the conference ID obtained in the early stage;
[0061] When the meeting officially starts, the local area network server issues a conference access instruction to the edge node; after receiving the instruction, the conference access module of the edge node automatically completes the login verification with the system where the target meeting is located, and then synchronously pushes the initial video stream of the digital human to the conference interface, ensuring that the digital human accesses the meeting scene in real time and completes the presentation preparation in the meeting start-up phase.
[0062] The multi-modal module collects the speech data of the target meeting in real time, converts the speech data into real-time text, and transmits the real-time text to the large model inference engine.
[0063] To adapt to the computing power limitation of the edge node in the offline local area network, the large language model needs to be optimized: first, the model pruning technique is used to accurately identify and remove redundant parameters in the model that are irrelevant to the meeting scenario, reducing invalid calculation; then, the INT8 quantization technique is used to compress the model data precision and storage volume, reducing the computing power load of the edge node; after double optimization, the model inference speed can be improved by 3 times, fully adapting to the computing power bearing capacity of RK3588;
[0064] At the same time, the large language model is fine-tuned based on the LoRA (Low Rank Adaptation) technique to achieve user individualization: only 3-5 hours of historical meeting corpus of the target user needs to be input, and the model can automatically extract two types of core user features: one is the language style feature, and the other is the decision feature; then, based on these features, the model weight parameters specific to the user are generated, and the weight parameters are encrypted and stored in the local area network server, providing technical support for generating reply content that fits the user's expression habits and decision-making style in subsequent meetings. The language style feature includes the use density of industry-specific terms, expression sentence preference, speed and tone correlation rules, etc., and the decision-making feature includes risk preference tendency in topic judgment, task priority selection logic, and stance tendency on controversial issues, etc.
[0065] The embodiment S1 mainly focuses on the preparation process before the meeting, solves the problems of model loading delay and poor meeting access in offline environment. The personalized configuration upload ensures that the digital person fits the target user characteristics; the preloading of large models and audio and video processing models to the memory within the preset time reduces the response delay in the meeting; the access test verifies the audio and video channel to avoid technical failures in the meeting; relying on the meeting ID login and pushing the initial video stream realizes the seamless connection of the digital person and the meeting system. These preparations guarantee the stability and efficiency of the meeting agent in the offline local area network environment, and improve the user experience of the digital person agent.
[0066] S2, in combination with the localized industry topic knowledge base and the historical meeting track of the target user, the future target core topics are predicted according to the word density and semantic logic in the real-time text, and the predicted target core topic set is obtained.
[0067] The embodiment step S2 includes S2.1-S2.2; wherein S2.1 is the process of establishing a localized industry topic knowledge base, and S2.2 is the process of obtaining a predicted target core topic set based on the localized industry topic knowledge base, specifically:
[0068] S2.1, obtain the historical meeting data of the target local area network, including meeting audio transcription text, meeting decision records, etc., obtain the enterprise historical meeting data, and do not access any external data throughout the process, from the source to guarantee data security and compliance; wherein the "target local area network" refers to the offline local area network built inside the enterprise or organization to which the target user belongs for important meetings. This local area network does not access external public networks, and only serves the internal data transmission and meeting collaboration of the enterprise or organization, and its core function is to store the historical meeting data of the enterprise or organization, providing a closed and secure data environment for the construction of the localized industry topic knowledge base.
[0069] The enterprise historical meeting data is preprocessed, including: first, invalid information is removed through data cleaning and text errors are corrected, and then combined with industry attributes to complete labeling, forming higher-precision basic materials. Among them, invalid information includes repeated speaking segments, irrelevant chatting content, etc.; "combined with industry attributes to complete labeling", for example, in the government field, label "policy implementation", "budget approval" related content, in the financial field, label "risk assessment", "product online" related segments.
[0070] Based on the basic materials, a number of industry core issues are obtained by double-dimension weighting screening of decision correlation and discussion proportion. The decision correlation refers to the proportion of the co-occurrence times of the statistical term and the preset decision class vocabulary. The discussion proportion refers to the proportion of the duration of the discussion paragraph corresponding to the term in the total conference duration. In this way, the core issues that meet the industry focus are screened out by combining the weight values of the two dimensions. At the same time, from the text corresponding to a number of industry core issues, the high-frequency discussion dimensions of the conference are extracted as sub-issues to obtain a number of sub-issues, and a hierarchical skeleton of “core issues and sub-issues” is built. The preset decision class vocabulary includes words such as “approval”, “pass”, “adjustment”, etc. The core issues that meet the industry focus include “policy landing” and “budget approval” in the government field, “risk assessment” and “product online” in the financial field, etc. The high-frequency discussion dimensions include “amount range”, “fund use”, “return period”, etc. associated with “budget approval”;
[0071] Based on the hierarchical skeleton, the historical discussion dimension data and the user historical decision tendency data in a number of sub-issues are integrated, and these data are sorted into a multi-level topology structure to finally form a vertical industry issue graph library that fits the industry characteristics and is closed-loop controllable. The historical discussion dimension data, such as the past discussion points of different “amount ranges”; the user historical decision tendency data, such as the decision preference and attitude logic of the user in the face of “budget over 1 million”.
[0072] Based on the historical conference data of enterprises in the local area network, the target key information corresponding to the target core issue is extracted, and the corresponding high-frequency discussion dimensions and user historical corpus are associated for each issue in the vertical industry issue graph library to obtain a localized industry issue knowledge base. The high-frequency discussion dimension refers to the issue dimension whose preset keyword appears more than a preset number of times.
[0073] In the process of establishing the vertical industry issue graph library, the embodiment S2.1 can ensure the standardization of the basic materials through data cleaning and industry attribute annotation; the weighting screening of the decision correlation and the discussion proportion avoids the misjudgment of the core issues caused by a single dimension, and ensures that the selected issues have both decision value and discussion heat. The hierarchical skeleton combined with the historical discussion dimension and the user decision tendency makes the graph library structured and systematic, which is convenient for subsequent rapid matching of real-time conference content, and lays a structured knowledge foundation for accurately predicting core issues;
[0074] Furthermore, based on historical meeting data within the local area network, a multi-level vertical industry topic graph library is built, linking high-frequency discussion dimensions with user historical data, ensuring the knowledge base closely aligns with enterprise business scenarios. Focusing on high-frequency discussion dimensions reduces redundant information interference, leading to more accurate subsequent topic prediction and response generation; localization ensures efficient knowledge retrieval in offline environments, avoiding delays or data leakage risks caused by reliance on external networks, providing digital human meeting agents with knowledge support tailored to enterprise needs.
[0075] S2.2. Using several core industry topics in the localized industry topic knowledge base as a reference, the proportion of "words that match the real-time text and the localized industry topic knowledge base" to the total frequency of "all words in the real-time text" is statistically analyzed. Based on this, the word density matching degree is calculated, and the topic direction related to the real-time discussion is initially identified.
[0076] The semantic matching algorithm calculates the structural similarity between the real-time text's semantic logic and the topic logic in the localized industry topic knowledge base, yielding the semantic logic chain matching degree. This is achieved by considering the semantic logic of the real-time text, such as the order of expression: "propose project requirements → assess implementation feasibility → discuss supporting budget." The lexical density matching degree and the semantic logic chain matching degree are weighted and summed to obtain the real-time text feature matching degree, which reflects the degree of relevance between the real-time text and the topics in the knowledge base.
[0077] The system retrieves the target user's historical meeting trajectory and calculates the similarity between the real-time text features of the real-time text and the discussion stage features of similar topics in the historical meeting trajectory using an offline cosine similarity algorithm. This yields the user's historical trajectory similarity. The historical meeting trajectory includes the discussion stage features of the user's past participation in similar topics. For example, when participating in "budget-related meetings," the focus is often on "risk control" topics after "feasibility analysis."
[0078] By performing a two-dimensional weighted calculation of real-time text feature matching degree and user historical trajectory similarity, a comprehensive score is assigned to each of several core industry topics in the localized industry topic knowledge base. Top-scoring topics (1-2) are selected from high to low scores as the predicted target core topic set, providing a basis for the advance response of the subsequent digital human conference agent. Furthermore, a preset keyword library and typical logic templates are established for the predicted target core topic set. The preset keyword library is a dedicated vocabulary set constructed based on high-frequency discussion dimensions of the corresponding predicted core topics in the localized industry topic knowledge base and key terms from user historical corpora. The typical logic templates are extracted from the common discussion logic chains of the predicted core topic in the target user's historical meeting trajectory. Both are used together to provide a benchmark for subsequent real-time verification of whether the target meeting has entered the predicted core topic set.
[0079] In this embodiment S2.2, keyword association is quantified by lexical density matching, and semantic logic matching captures the inherent logic of the topics. The real-time text feature matching degree obtained by weighting the two ensures a good fit with the knowledge base. User historical trajectory similarity is associated with historical features of similar topics through the cosine similarity algorithm, reflecting users' meeting habits. The comprehensive scoring mechanism with dual-dimensional weighting ensures that the prediction results are consistent with both the real-time trend of the meeting and the user's historical behavior, improving the accuracy of the prediction of the target core topics, allowing the digital human to prepare in advance, and enhancing the targeting of the proxy response.
[0080] S3. If the real-time text detects mentions of the target user or matters to be decided, the first response voice is generated by combining the target user's historical speech data; if the target meeting enters the predicted core topic set, the pre-loaded target user feature package is called to generate the second response voice.
[0081] Step S3 in this embodiment of the application is specifically as follows:
[0082] Deep analysis of real-time text is performed based on the large-model inference engine:
[0083] If the target user or decision-making matter is detected in the real-time text, the first response text that is adapted to the meeting context and fits the user's expression habits is generated by combining the previously stored historical data of the target user with the offline multimodal semantic association enhancement mechanism.
[0084] If the target meeting is determined to have entered the previously established set of predicted core topics, the control LAN server immediately pushes the target user feature package corresponding to that core topic to the edge computing nodes in real time. This feature package precisely contains key information about the user's past discussions on such topics, such as high-frequency technical terms and fixed decision-making logic for the "risk control" topic, ensuring that the feature package accurately matches the topic discussion needs. High-frequency technical terms include "bad debt rate threshold" and "stress test," while fixed decision-making logic includes "prioritizing the security of core business funds." After receiving the target user feature package, the edge node loads it into a temporary adapter dedicated to the user model. This adapter only occupies 5%-10% of the video memory resources, fully adapting to the computing power limitations of the RK3588 chip, avoiding excessive hardware resource consumption that could affect the operation of other meeting functions. Once the meeting formally enters the discussion phase of the predicted target core agenda, the large model inference engine at the edge nodes does not need to re-access the user's historical corpus for calculation. Instead, it can directly call upon the user features pre-loaded in the temporary adapter and quickly generate a second response text that conforms to the user's decision-making tendencies and professional expression style based on the offline multimodal semantic association enhancement mechanism. This further reduces the response latency to ≤200ms, significantly improving the real-time performance and smoothness of the meeting interaction. Furthermore, both the first and second response texts are transmitted to the multimodal module in real time, providing audio data for the subsequent generation of the corresponding audio and video streams for the digital human. Additionally, it should be noted that if the real-time text detects mentions of the target user or matters to be decided, and the target meeting enters the predicted target core agenda, the corresponding operation method for "the target meeting has entered the predicted target core agenda" will be prioritized.
[0085] If the analysis results of the large model inference engine show that the text is ordinary discussion content, such as no user mentions, no pending decisions, or non-core issues, then a structured minutes fragment will be automatically generated and temporarily stored in an encrypted storage area within the local area network to ensure the security and standardization of meeting information retention.
[0086] The multimodal module converts the first response text and the second response text into the first response speech and the second response speech, respectively.
[0087] The following provides a detailed explanation of how to determine whether a target meeting falls within the pre-defined core topic set, and the offline multimodal semantic association enhancement mechanism:
[0088] (1) Determining whether the target meeting falls into the predicted core agenda is achieved through three verification methods:
[0089] First, real-time text keyword matching verification. Based on the preset keyword library of the predicted core issues, the frequency ratio of matching words in the real-time text is counted. If the ratio exceeds the preset threshold, it can be used as a basis for the discussion content to be highly related to the core words of the predicted issues.
[0090] Second, real-time text semantic logic matching verification involves calling a semantic matching algorithm to calculate the structural similarity between the real-time text semantic logic and the typical logical templates in the predicted target core issues set. If the similarity meets the standard, it can further prove that the discussion logic framework is consistent with the core logic of the predicted issues.
[0091] Third, the user's historical meeting trajectory feature matching verification combines the characteristics of the target user's historical discussion of similar topics. The offline cosine similarity algorithm is used to calculate the similarity between the text features of the real-time text and the historical features. If the similarity meets the standard, it can help confirm that the current discussion stage matches the characteristics of the stage when the topic was entered in the past.
[0092] If all three verification results are met simultaneously, it can be determined that the target meeting has entered the predicted core topic set.
[0093] (2) Offline multimodal semantic association enhancement mechanism, specifically:
[0094] ① Localized Multimodal Data Processing Closed Loop: Led by the multimodal interaction module, a complete processing link for three types of data—text, vision, and behavior—is built in an offline environment, adapted to the computing power of the RK3588 edge node. At the text semantic level, the meeting speech is converted into real-time text using the Vosk open-source offline speech recognition model, and core keywords of the topic are extracted simultaneously to obtain text keywords. At the visual semantic level, the lightweight Tesseract offline OCR is used to recognize content such as Excel cost sheets and progress PPTs shared on the screen during the meeting, and accurately extract key visual information such as "cost overrun of 15%" and "R&D cycle delayed by 3 days". At the behavioral semantic level, the speaker's gestures and expressions are captured by the meeting terminal camera and converted into standardized behavioral labels such as "gestures emphasize high priority" and "frowning corresponds to questioning" using lightweight gesture recognition models such as MediaPipePose offline version.
[0095] ② Cross-dimensional semantic association algorithm: A "multimodal semantic association matrix" is introduced into the large model inference engine. First, text keywords, key visual information, and standardized behavioral labels are mapped to the same semantic space. Then, the target user's historical data, which is encrypted and stored on the local area network server, is called up. Combined with user-specific model weights finely tuned by LoRA small samples, response text that fits the user's expression style and decision-making tendency is generated. For example, if a meeting shows "shared document shows 'R&D cycle delayed' + speaker frowns", the algorithm will automatically associate it with the decision-making logic of "prioritizing the allocation of the assigned team when there is a delay" in the user's historical data and output a targeted response.
[0096] ③ Localized Destruction of Multimodal Data: After all raw multimodal data in the target meeting is processed into semantic tags at the edge node, the edge node security component immediately performs the destruction operation, and the destruction log is synchronously uploaded to the local area network server security management module for archiving. Only the semantic tag data is retained for cross-dimensional correlation calculations. After the calculations are completed, it is cleaned up along with the non-critical data of the meeting at a preset cycle after the meeting ends. This avoids the privacy leakage risk of raw data storage and fully complies with the compliance requirement of "meeting data never leaving the domain" in offline local area networks. The raw multimodal data includes video frames, screenshots of shared documents, etc.
[0097] In this embodiment, S3 can quickly capture topic-related keyword signals by statistically analyzing the frequency ratio of matched words in real time, avoiding misjudgment based on a single word. Calculating the structural similarity between the semantic logic of the text and typical templates allows for a deeper understanding of the inherent logical connections in the discussion of topics, overcoming the superficial limitations of keyword matching. Verifying the cosine similarity between text features and the characteristics of similar topics in the user's history enhances the personalized adaptability of the judgment by incorporating user behavior habits. The combination of these three elements forms a complementary verification mechanism, ensuring efficient real-time response in offline LAN environments and accurately identifying topic switching nodes. This provides a reliable basis for subsequently calling user feature packages to generate tailored response voices, thereby improving the response accuracy and meeting participation fluency of the AI digital human conference agent.
[0098] S4. Based on the first or second response voice, control the digital human to generate corresponding audio and video stream outputs.
[0099] Step S4 in this embodiment includes S4.1 to S4.3; wherein, S4.1 is the process of adjusting and optimizing the first and second response voices, S4.2 is the process of controlling the digital human to generate corresponding audio and video streams based on the final first or second response voice, and S4.3 is the process of data integration after the target meeting ends, specifically as follows:
[0100] S4.1 Extract temporary decision-making tendency data and target new terminology preference data from the voice data. Integrate the two types of data into a real-time feature increment package through a keyword weight dynamic adjustment algorithm. Simultaneously, to balance processing efficiency and real-time interaction requirements, an "incremental inference" optimization strategy is adopted: by marking the timestamps and content boundaries of the meeting text in real time, an incremental recognition mechanism for text processing is constructed. This controls the core large model in the large model inference engine to only perform inference calculations on newly added meeting text fragments, automatically filtering and skipping historical text content that has already been processed, avoiding redundant calculations that consume computing resources. Ultimately, the response generation delay is strictly controlled within 500ms, meeting the timeliness requirements of real-time interaction in meeting scenarios. Specifically, temporary decision-making tendency data, such as users' past preference for the "cost-first" principle, but the repeated emphasis in this meeting on "urgent project schedule, requiring priority to ensure delivery timeliness," temporarily adjusts the decision-making tendency to "schedule-first." Target new terminology preference data, such as users using "agile development iteration cycle" for the first time instead of the previously commonly used "project phase division" expression, reflects the real-time change in terminology habits, and the expression "agile development iteration cycle" will be used thereafter. The "keyword weight dynamic adjustment algorithm" adopts an offline calculation mode, which only updates the weight parameters of feature words in a targeted manner without triggering full model training, thus greatly reducing the consumption of computing power.
[0101] If the real-time meeting features in the real-time feature increment package conflict with the response features of the previously generated first or second response voice, a time-weighted judgment mechanism is triggered. This mechanism modifies the conflict response logic of the first or second response voice according to the core rule that "recent feature weight is greater than long-term feature weight, and real-time meeting feature weight is greater than historical static feature weight," resulting in the final first or second response voice. The time-weighted judgment result includes recent feature weight being greater than long-term feature weight, and real-time meeting feature weight being greater than historical static feature weight. Simultaneously, a detailed log of this feature conflict is automatically recorded and pushed to the user's terminal after the meeting for manual confirmation and secondary calibration, further preventing model misjudgments due to feature conflicts and ensuring the accuracy of personalized feature iteration. The detailed log includes information such as the content of the conflicting feature, the judgment basis, and the modification result. For example, a real-time meeting feature might be used where the user temporarily supports "cross-departmental collaboration to advance the project."
[0102] In this embodiment, S4.1 extracts temporary decision-making tendencies and terminology preferences from voice data, capturing real-time dynamics during meetings. When real-time features conflict with pre-generated response features, the response logic can be adjusted according to a weighting rule that recent features are superior to long-term features, and real-time features are superior to historical static features, ensuring that the digital human's responses are appropriate for the current meeting context. This dynamic adjustment mechanism enhances the timeliness and flexibility of responses, avoids inappropriateness caused by mechanically relying on historical data, and improves the on-site adaptability of the digital human agent.
[0103] S4.2 After acquiring the final first or second response voice, the multimodal module performs adaptive processing on the voice signal to ensure that the sound quality and speech rate meet the requirements of the meeting scenario. Simultaneously, the digital human rendering module is triggered to generate matching digital human body movements based on the context and rhythm of the response voice. The response voice and corresponding digital human movements are integrated into a coherent audio-visual stream, which is then output to the target meeting in real time through the meeting access module, ensuring the naturalness of the digital human interaction and the real-time nature of meeting participation. The "context of the response voice" includes, for example, a calm gesture corresponding to a decision-making response, and an interactive action corresponding to a discussion-based response.
[0104] The control module continuously monitors the dynamic status of the meeting while outputting audio and video streams, including key information such as speaker identity switching, participant screen sharing operations, and meeting topic transitions, and synchronizes this status data to the local area network server in real time. After the status data is lightweightly processed by the local area network server, it is immediately pushed to the target user's terminal, allowing the target user to remotely monitor the meeting progress in real time even if they are not directly participating in the meeting, thus achieving a collaborative connection between digital human proxy participation and remote user awareness.
[0105] S4.3 After the target meeting concludes, integrate all the minutes fragments from the target meeting to generate a structured minutes;
[0106] The security management module encrypts the structured minutes and sends a push notification of the minutes to the target user's terminal. The security management module processes the structured minutes using a preset encryption algorithm and uses triple authentication, including password authentication, device fingerprint authentication, and dynamic password authentication, to achieve data encryption.
[0107] After the target meeting concludes, the Blenderbot model automatically retrieves all temporarily stored minutes from the meeting. Through content correlation analysis and logical organization, the scattered minutes are integrated into a structured summary. This summary clearly includes basic meeting information, core discussion points, action items to be implemented, and key decision results. The structured summary is then transmitted to the security management module. The basic meeting information includes the meeting topic, time, and attendees.
[0108] After receiving the structured minutes, the security management module first performs low-level encryption on the structured minutes using a preset encryption algorithm. Then, it constructs a security protection system with a triple authentication mechanism of password authentication, device fingerprint authentication, and dynamic password. Under this dual protection, the encrypted structured minutes are stored on the local area network server. After storage is completed, a push notification is automatically sent to the target user's terminal stating "Meeting minutes have been generated and can be viewed with permission," ensuring that the target user can obtain the core meeting information in a timely manner.
[0109] Simultaneously, the local area network (LAN) server triggers a data archiving process: according to data security standards, the entire meeting recording, digital human interaction audio and video streams, and system operation logs are categorized and archived for storage; a periodic cleanup mechanism is also initiated to automatically delete non-critical data exceeding the preset retention period, thus optimizing LAN storage resource usage while meeting industry data retention compliance requirements. The "preset retention period" can be one year, and "non-critical data" includes duplicate temporary minutes and interaction records without critical information.
[0110] It should be noted that this embodiment adopts a three-layer hardware architecture of "edge computing node - local area network server - terminal device", coupled with five major software modules (conference access and control module, multimodal interaction module, large model inference engine, digital human rendering engine, and security management module) to build an offline closed-loop system. High-availability inference in an offline environment is achieved through model sharding and fault self-healing technology, avoiding conference interruptions. Its hardware architecture and core device parameters are as follows:
[0111] ① The edge computing nodes utilize the Zhiwei Industrial RK3588 chip edge box, which boasts 6 TOPS NPU computing power and supports INT8 quantization inference. A single node deploys a localized large model and digital human rendering engine for real-time conference interaction and inference computation, supporting 3-5 concurrent conference proxies with a power consumption of <10W. Alternatively, the NVIDIA Jetson Nano (4 TOPS computing power) can be used to replace the RK3588 chip.
[0112] ② The local area network server is a dual-machine hot standby physical server, responsible for user configuration management, conference resource scheduling, and data encryption and archiving. It communicates with edge nodes and terminal devices through the TCP / IP protocol.
[0113] ③ Terminal devices are divided into two categories: one is the meeting initiation terminal, such as a PC / meeting terminal, which supports Windows / Linux systems; the other is the user interaction terminal, such as a mobile phone / tablet, which supports Web / client access. These terminal devices are used for user configuration input, meeting status viewing, and minutes reception.
[0114] The following provides a detailed explanation of the collaborative work and workflow connection of the five major software modules appearing in the embodiments of this application:
[0115] ① Meeting Access and Control Module: This module serves as the central access hub for the entire system. On one hand, it encapsulates a general meeting control API and reserves dedicated adaptation interfaces for mainstream meeting systems such as Zoom and WeLink, allowing for flexible integration with commonly used meeting platforms across different enterprises. On the other hand, it accurately retrieves the meeting ID and automatically synchronized access password from the LAN server, then issues a "meeting access command" to the edge computing nodes to initiate the access process. Simultaneously, this module receives real-time feedback on the meeting's dynamic status from the edge nodes and pushes this status information synchronously to the target user's terminal, ensuring users can remotely monitor the meeting's progress in real time. The meeting's dynamic status includes whether the meeting has officially started, the current speaker's identity has changed, and participants have initiated screen sharing. Furthermore, in terms of meeting system integration options, in addition to the aforementioned mainstream platforms, Tencent Meeting API or Zoom Rooms SDK can also be used to further enrich the system's integration options.
[0116] ② Multimodal Interaction Module: This module acts as the hub for audio-text conversion and interaction. Its deployed Vosk open-source offline speech recognition model boasts advantages such as lightweight design, multi-language support, and high accuracy, adapting to the resource limitations of offline LANs and various scenario requirements. This module first receives real-time conference audio streams, accurately converts the speech into text using the Vosk model, and then transmits it to the large model inference engine. After receiving the large model's output response text, it uses locally deployed TTS engines such as eSpeak to convert the text into natural speech, simultaneously triggering the lip-syncing logic of the digital human rendering module to ensure consistency between the speech and the digital human's lip movements, enhancing the realism of the interaction.
[0117] ③ Large-Scale Inference Engine: This engine is the core decision-making and content generation engine of the system, integrating two types of specialized optimized models: one is a large language model processed by INT8 quantization, such as the DeepSeek model, which is mainly responsible for semantic understanding of meeting texts, simulation of user decision-making tendencies, and generation of response texts that fit the user's style; the other is the blenderbot-400M-distill lightweight model (referred to as the blenderbot model), which focuses on the structured organization of meeting content and the generation of minutes fragments. During operation, this engine obtains meeting texts from the multimodal interaction module, combines them with the "user personalized corpus" distributed by the local area network server, and calls the corresponding model as needed: when real-time response is required, the large language model is called to generate response text and transmit it to the multimodal interaction module; when meeting content needs to be recorded, the blenderbot model is called to generate structured minutes fragments for data archiving.
[0118] ④ Digital Human Rendering Engine: This engine, serving as the visual presentation carrier, employs a lightweight 3D rendering architecture, strictly controlling rendering latency to within 200ms to ensure smooth real-time interaction. The engine simultaneously receives two types of key data: first, voice rhythm data transmitted from the multimodal interaction module; second, emotion tendency labels output by the large-model inference engine. Based on this data, the engine precisely drives the digital human to complete facial expressions and body gestures, generating a coherent digital human video stream, which is then pushed to the target meeting system via the conference access and control module. Specifically, voice rhythm data is used to match the speed of lip movements; emotion tendency labels, including neutral statements and emphasis, are used to determine the style of facial expressions and movements.
[0119] ⑤ Security Management Module: This module is used to build a full-process security layer, constructing a solid security defense from three dimensions: data transmission, storage, and access. In the data transmission stage, the AES-256 encryption algorithm is used to protect all interactive data between edge nodes and the LAN server, preventing leakage during transmission. In the data storage stage, the national cryptographic algorithm SM4 is used to encrypt and store core data such as personalized user data and meeting minutes, ensuring static data security. In the access control stage, triple authentication of "password authentication + device fingerprint authentication + dynamic password" is implemented. Simultaneously, operation permissions are divided based on the RBAC (Role-Based Access Control) permission model, and operation logs for all modules are recorded in real time, facilitating subsequent security audits and issue tracing. For example, "operation permissions" may allow only system administrators to configure meeting access permissions and only target users to view exclusive meeting minutes.
[0120] The following provides a detailed explanation of model fragmentation and fault self-healing techniques:
[0121] ① Localized fragmented deployment of large models: The large language model after INT8 quantization is split into three independent fragments: "semantic understanding submodule, decision simulation submodule, and response generation submodule". The Blenderbot model is split into "memory fragment extraction submodule and structured integration submodule". All fragments are deployed in a distributed manner through 3-5 Zhiwei Industrial RK3588 chip edge nodes in the local area network. Each node runs only 1-2 fragments. Data is exchanged between fragments through the RTPS low-latency protocol.
[0122] ② Real-time fault detection and fragment migration: Deploy a fault detection submodule on the local area network server to monitor node status through "heartbeat packets and inference result verification"; if a node fails, such as semantic understanding fragment failure, the server immediately pushes the fragment preloading parameters to the backup node, and the backup node starts fragmentation within 1 second, synchronizing the current meeting user feature data, real-time feature increment packets and minutes fragments to avoid meeting interruption.
[0123] ③ Dynamic resource scheduling optimization: The server adjusts the sharding deployment according to the meeting load; if the meeting enters a decision-intensive period, the server automatically migrates the shard to an edge node with more computing power to ensure inference efficiency. Among them, the "decision-intensive period" refers to the specific stage in the meeting process where the matters to be decided appear in a concentrated manner and decision tendencies or conclusions need to be output frequently.
[0124] This embodiment, S4.3, addresses the issues of low efficiency in organizing meeting information and high data security risks by generating and encrypted post-meeting structured minutes. The structured minutes integrate meeting segments, allowing users to quickly access core content and improving work efficiency. The security management module employs a preset encryption algorithm combined with password authentication, device fingerprint authentication, and dynamic password triple authentication, strengthening data confidentiality in offline LAN environments and preventing the leakage of sensitive meeting information. Timely push of minutes notifications to user terminals ensures users can quickly access meeting results, balancing information utilization efficiency and security, and perfecting the closed-loop service of the digital human meeting agent.
[0125] Overall, the embodiments of this application have the following beneficial effects:
[0126] This application utilizes real-time speech-to-text conversion to provide an instantly interpretable data source for subsequent analysis, ensuring the entire process keeps pace with the meeting. By combining a localized industry-specific knowledge base and user meeting history, it avoids the limitations of offline environments where external data cannot be accessed online, and allows topic prediction based on lexical density and semantic logic to align with industry characteristics and user meeting habits, identifying core topics in advance. This design transforms topic prediction from an isolated process, providing clear direction for subsequent personalized responses and creating a synergy between topic prediction and response generation, thus resolving the issue of their lack of coordination. When a mention of a target user or a decision-making matter is detected, the first response speech is generated using locally stored user history data. This leverages offline personalized data to ensure the response aligns with user habits while meeting real-time requirements through an instant triggering mechanism. When the meeting reaches the predicted core topic, a second response speech is generated by calling a pre-loaded user feature package, directly transforming the previously predicted topic results into the basis for the response. This ensures that topic prediction is not an isolated preliminary step but forms a data loop with response generation. Both rely on locally stored user data in an offline environment, avoiding delays caused by network dependence. Finally, by converting the first response voice related to the real-time mention scenario and the second response voice related to the predicted core topic into a digital human audio and video stream output, the real-time personalized response generation and the meeting topic prediction results can be synchronously presented in an offline environment, thereby solving the problem of the two being difficult to coordinate.
[0127] In summary, this application enables end-to-end offline processing of meeting data, storing it entirely within the local area network. Combined with AES-256 transmission encryption and national cryptographic algorithm-based storage encryption, it fully meets the "data not leaving the domain" requirements of the Information Security Protection Standard 2.0 and industries such as finance and government, mitigating the data leakage risks of cloud-based solutions. Simultaneously, it achieves a closed loop of "automatic pre-meeting configuration, real-time in-meeting interaction, and intelligent post-meeting minutes," reducing ineffective meeting time for middle and senior management. The completeness of structured meeting minutes is 1.8 times that of manual recording, avoiding traditional information transmission gaps. Utilizing LoRA small-sample fine-tuning technology, the AI digital human can accurately simulate user speaking styles and decision-making tendencies, with interactive realism superior to existing standardized digital human solutions. The typical power consumption of edge computing nodes is less than 10W, reducing the average annual deployment cost of the system by more than 60% compared to professional meeting secretaries. It also supports 3-5 concurrent meeting proxies, resulting in higher hardware utilization. Furthermore, after INT8 quantization compression, large models can achieve real-time inference at edge nodes, addressing the pain point that existing online large models cannot run offline.
[0128] Example 2:
[0129] Please see Figure 2 The embodiments of this application provide an AI digital human conference agent device for offline local area networks, including a data module 10, a prediction module 20, a voice module 30 and a response module 40;
[0130] Among them, the data module 10 is used to collect the voice data of the target meeting in real time and convert the voice data into real-time text;
[0131] Prediction module 20 is used to combine a localized industry topic knowledge base and the target user's historical meeting trajectory to predict several future target core topics based on the word density and semantic logic in the real-time text, and obtain a set of predicted target core topics.
[0132] The voice module 30 is used to generate a first response voice if the target user or decision-making matter is detected in the real-time text, combined with the target user's historical speech data; if the target meeting enters the predicted target core topic set, it calls the pre-loaded target user feature package to generate a second response voice.
[0133] The response module 40 is used to control the digital human to generate corresponding audio and video stream outputs based on the first or second response voice.
[0134] It should be noted that the technical concept of this second embodiment is completely consistent with that of the first embodiment. The two maintain a high degree of synergy at the technical logic level. The specific technical details can be referred to the relevant description of the first embodiment, which will not be repeated here.
[0135] Example 3:
[0136] This application provides a computer-readable storage medium, which includes a stored computer program, wherein the computer program, when running, controls the device where the computer-readable storage medium is located to execute the AI digital human conference agent method under an offline local area network.
[0137] The AI digital human conference agent method under an offline local area network, when implemented as a software functional unit and used as an independent product, can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.
[0138] The above are preferred embodiments of the present invention. It should be noted that, for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications are also considered to be within the scope of protection of the present invention.
Claims
1. A method for AI digital human conference proxy under offline local area network, characterized in that, include: Real-time acquisition of audio data from the target conference, and conversion of the audio data into real-time text; By combining a localized industry topic knowledge base and the target user's historical meeting trajectory, several target core topics are predicted in the future based on the lexical density and semantic logic in the real-time text, resulting in a predicted set of target core topics; If the real-time text detects mentions of the target user or matters to be decided, a first response voice is generated by combining the target user's historical speech data; if the target meeting enters the predicted target core topic set, a second response voice is generated by calling the pre-loaded target user feature package. Based on the first or second response voice, the digital human is controlled to generate corresponding audio and video stream outputs.
2. The AI digital human conference proxy method under an offline local area network as described in claim 1, characterized in that, The method for obtaining the localized industry topic knowledge base is as follows: Acquire historical meeting data of enterprises within the target local area network; Based on the aforementioned historical meeting data of the enterprise, a vertical industry topic map library is obtained by constructing a multi-level topological structure of core topics, sub-topics, and related attributes. Based on the enterprise's historical meeting data, each topic in the vertical industry topic graph library is associated with a corresponding high-frequency discussion dimension and user historical corpus to obtain the localized industry topic knowledge base; wherein, the high-frequency discussion dimension refers to the topic dimension where the frequency of preset keywords exceeds a preset number of times.
3. The AI digital human conference proxy method under an offline local area network as described in claim 2, characterized in that, Based on the aforementioned historical meeting data of the enterprise, a vertical industry topic map library is obtained by constructing a multi-level topological structure of core topics, sub-topics, and related attributes, specifically as follows: The historical meeting data of the enterprise was cleaned and industry attribute labeled to obtain basic materials; Based on the aforementioned basic materials, several core industry topics are obtained by weighting and filtering based on decision relevance and discussion ratio. Several meeting dimensions are extracted from the text corresponding to these core industry topics as sub-topics, and a hierarchical framework for these core industry topics and sub-topics is constructed. Here, decision relevance refers to the proportion of co-occurrence of statistical terms and preset decision-related words, and discussion ratio refers to the proportion of the discussion paragraph time corresponding to the term to the total meeting time. Based on the hierarchical framework, the historical discussion dimension data and user historical decision tendency data in the several sub-topics are integrated into a multi-level topological structure to obtain the vertical industry topic map library.
4. The AI digital human conference proxy method under an offline local area network as described in claim 3, characterized in that, By combining a localized industry topic knowledge base and the target users' historical meeting records, and based on the lexical density and semantic logic in the real-time text, several future target core topics are predicted, resulting in a predicted set of target core topics, specifically: Based on the core industry topics in the localized industry topic knowledge base, the frequency ratio of words matching the real-time text and the localized industry topic knowledge base is statistically analyzed to obtain the word density matching degree. The semantic matching algorithm is used to calculate the structural similarity between the real-time text semantic logic and the topic logic in the localized industry topic knowledge base to obtain the semantic logic chain matching degree. The real-time text feature matching degree is obtained by weighted summation of the word density matching degree and the semantic logic chain matching degree. The historical meeting trajectory of the target user is invoked, and the similarity between the real-time text features of the real-time text and the discussion stage features of the same topic in the historical meeting trajectory is calculated according to the offline cosine similarity algorithm to obtain the user's historical trajectory similarity. By weighting the real-time text feature matching degree and the user's historical trajectory similarity in two dimensions, the comprehensive score of each topic in the several industry core topics is obtained, and the top several topics in descending order of score are selected as the predicted target core topic set.
5. The AI digital human conference proxy method under an offline local area network as described in claim 1, characterized in that, Determining whether the target meeting has entered the predicted target core topic set is achieved by real-time statistical analysis of the frequency ratio of matching words between the real-time text and the preset keyword library in the predicted target core topic set, calculating the structural similarity between the text semantic logic of the real-time text and the preset typical logic template in the predicted target core topic set, and verifying the cosine similarity between the text features of the real-time text and the characteristics of the target user's historical discussion of similar topics.
6. The AI digital human conference proxy method under an offline local area network as described in claim 1, characterized in that, Before controlling the digital human to generate corresponding audio and video stream outputs based on the first or second response voice, the method further includes: Extract temporary decision-making tendency data and target new terminology preference data of several users from the voice data to generate real-time feature increment packages; If the real-time conference features in the real-time feature increment package conflict with the response features of the first or second response voice, the response logic of the first or second response voice is modified based on the time weight determination result; wherein, the time weight determination result includes that the weight of recent features is greater than that of distant features, and that the weight of real-time conference features is greater than that of historical static features.
7. The AI digital human conference proxy method under an offline local area network as described in claim 1, characterized in that, Before the real-time acquisition of the target conference's voice data, the following is also included: The personalized configuration is obtained and uploaded to the local area network server; wherein, the personalized configuration includes the target user's historical data and meeting ID; Within a preset time period before the start of the target meeting, the control edge node loads the large language model, speech recognition model, and speech synthesis model into memory; The conference access module of the control edge node initiates an access test. After verifying that the audio and video path with the system where the target conference is located is smooth, it logs into the target conference based on the conference ID and pushes the initial video stream of the digital human to the conference interface of the target conference.
8. The AI digital human conference proxy method under an offline local area network as described in claim 1, characterized in that, After controlling the digital human to generate corresponding audio and video stream outputs based on the first or second response voice, the method further includes: Integrate all the minutes fragments from the target meeting to generate a structured minutes; The structured minutes are encrypted using a security management module, and a minutes push notification is sent to the target user's terminal. The security management module processes the structured minutes using a preset encryption algorithm, and uses triple authentication—password authentication, device fingerprint authentication, and dynamic password authentication—to achieve data encryption.
9. An AI digital human conference agent device for offline local area networks, characterized in that, It includes a data module, a prediction module, a voice module, and a response module; The data module is used to collect voice data of the target meeting in real time and convert the voice data into real-time text. The prediction module is used to combine a localized industry topic knowledge base and the target user’s historical meeting trajectory to predict several future target core topics based on the lexical density and semantic logic in the real-time text, thereby obtaining a set of predicted target core topics. The voice module is used to generate a first response voice based on the target user's historical speech data if the real-time text detects mentions of the target user or matters to be decided; and to generate a second response voice by calling a pre-loaded target user feature package if the target meeting enters a predicted core topic set. The response module is used to control the digital human to generate corresponding audio and video stream outputs based on the first response voice or the second response voice.
10. A storage medium, characterized in that, The storage medium stores a computer program, which is called and executed by a computer to implement an AI digital human conference agent method for offline local area networks as described in any one of claims 1 to 8.
Citation Information
Cited By
AI model lightweight optimization method and device suitable for resource constraint equipment
CN121835930A
An AI model lightweight optimization method and device suitable for a resource-constrained device
CN121835930B