A method and system for scene-aware real-time accompaniment and structured memory generation

By using a scene-aware real-time accompanying and structured memory generation system, the problems of multi-dimensional structured information extraction and high power consumption in existing technologies have been solved, achieving low power consumption, privacy protection, multi-scenario applicability, and efficient information recording.

CN122334482APending Publication Date: 2026-07-03CHENGDU RUIMIXIN TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHENGDU RUIMIXIN TECHNOLOGY CO LTD
Filing Date
2026-04-03
Publication Date
2026-07-03

AI Technical Summary

Technical Problem

Existing conversation recording tools cannot achieve multi-dimensional structured information extraction based on scene semantics, cannot meet the needs of everyday and fragmented scenarios, and suffer from information omissions and high power consumption.

Method used

The system employs a scene-aware real-time accompanying and structured memory generation system, including an audio and video acquisition module, a scene intent recognition and scheduling module, a sensor parameter dynamic configuration module, an external knowledge source retrieval module, a real-time dialogue perception engine, a voiceprint reference sample acquisition and encrypted storage module, a structured information extraction module, and a memory persistence and presentation module. Through a large language model, it performs intent classification, sensor parameter strategy calculation, and semantic analysis to achieve multi-dimensional structured data extraction and local privacy protection.

Benefits of technology

It enables low-power, multi-dimensional structured information extraction in different scenarios, improves the accuracy of role differentiation, meets the privacy compliance requirements of sensitive scenarios such as medical and business, reduces device power consumption and computing power requirements, and reduces invalid data transmission and computing resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122334482A_ABST
    Figure CN122334482A_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for real-time accompaniment and structured memory generation based on scene perception, relating to the fields of artificial intelligence, augmented reality, and mobile computing. The system includes core modules such as an audio and video acquisition module, a scene intent recognition and scheduling module, a sensor parameter dynamic configuration module, a real-time dialogue perception engine, and a structured information extraction module. This invention achieves dynamic optimization of hardware parameters driven by scene semantics, balancing perception accuracy and device power consumption. It realizes speaker recognition without training through a large language model, and adopts a local priority architecture to protect data privacy. It can achieve seamless continuous accompaniment after a single trigger, completing real-time understanding of dialogue content, multi-dimensional structured information extraction, and local memory generation. At the same time, through a voice activity adaptive acquisition mechanism and a runtime voice command reconfiguration mechanism, it achieves low-power acquisition and touchless parameter adjustment, making it suitable for various scenarios such as medical consultation, business negotiation, and daily social interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method and system for real-time accompaniment and structured memory generation based on scene perception. Background Technology

[0002] With the implementation of AI large-scale model technology and the popularization of AR smart glasses and smart wearable devices, users' demand for real-time assistance, complete recording and structured organization of conversation content in everyday scenarios such as medical consultation, business negotiation, daily social interaction and learning and education continues to increase. Existing conversation recording tools can only achieve basic speech-to-text conversion. They cannot extract multi-dimensional structured information based on contextual semantics, nor can they achieve multi-channel parallel information distribution from a single input. Users need to manually organize the conversation content afterward, which is inefficient and prone to information omission. In addition, they are only suitable for formal meeting scenarios and cannot meet the needs of everyday life and fragmented scenarios. Summary of the Invention

[0003] The purpose of this invention is to provide a method and system for real-time accompaniment and structured memory generation based on scene awareness, so as to solve the problems mentioned in the background art.

[0004] To achieve the above objectives, the present invention adopts the following technical solution: The system is based on scene-aware real-time companion and structured memory generation, including an audio and video acquisition module, a scene intent recognition and scheduling module, a sensor parameter dynamic configuration module, an external knowledge source retrieval module, a real-time dialogue perception engine, a voiceprint reference sample acquisition and encrypted storage module, a structured information extraction module, an intelligent routing module, and a memory persistence and presentation module.

[0005] As a further improvement to this technical solution: the scene intent recognition and scheduling module receives the natural language description input by the user, performs intent classification through a large language model, and outputs a standardized scene classification identifier.

[0006] As a further improvement to this technical solution: the sensor parameter dynamic configuration module receives the scene classification identifier output by the scene intent recognition and scheduling module, calculates and distributes the sampling parameter strategy of the wearable device's underlying hardware sensor through a large language model. The sampling parameter strategy includes the voice activity detection sensitivity threshold, the maximum duration of a single acquisition, the silent cut-off timeout duration, the visual sensor on / off status, the data transmission quality level, and the video stream acquisition mode.

[0007] As a further improvement to this technical solution: the voiceprint reference sample acquisition and encrypted storage module acquires short-term voice samples of the user, encrypts and stores the voice samples on the user's local device, and submits the voice samples as a component of multimodal input along with real-time dialogue audio to the large language model. The large language model completes speaker role labeling through acoustic feature comparison, distinguishing the user from other dialogue participants, without the need for voiceprint vector matching based on traditional signal processing.

[0008] As a further improvement to this technical solution: the external knowledge source retrieval module receives scene classification identifiers, obtains knowledge and rule information in the corresponding field, and injects the information into the AI ​​reasoning context. External knowledge sources include search engines, cloud knowledge graphs, and local pre-built vector libraries.

[0009] As a further improvement to this technical solution: the real-time dialogue perception engine collects environmental audio and video signals according to the issued sampling parameter strategy, completes speech recognition and multi-role labeling, and performs semantic analysis by combining the injected knowledge and rule information.

[0010] As a further improvement to this technical solution: the structured information extraction module submits the complete dialogue record to the large language model to extract multi-dimensional structured data, which includes summaries, key notes, task items, and financial records. The intelligent routing module synchronously distributes the structured data to the corresponding functional subsystems, which include a memory management subsystem, a diary generation subsystem, a financial management subsystem, a task management subsystem, and a location tagging subsystem.

[0011] As a further improvement to this technical solution: the memory persistence and presentation module encrypts and stores multi-dimensional structured data on the user's local device, generating visual memory information containing scene classification identifiers.

[0012] As a further improvement to this technical solution: the audio and video acquisition module includes a voice activity detection submodule. This submodule monitors the energy value of the audio signal in real time. When the energy value exceeds the preset voice activity threshold, recording is automatically started. When the continuous silence duration exceeds the preset silence threshold or the recording duration reaches the preset maximum acquisition duration, the current recording segment is automatically cut off and analysis and uploading are triggered. A dual-trigger cutoff mechanism is adopted to achieve adaptive acquisition based on voice activity. During periods without effective voice signals, a low-power silent listening state is maintained, reducing the transmission of invalid data and the consumption of computing resources.

[0013] As a further improvement to this technical solution: the sensor parameter dynamic configuration module supports runtime voice reconfiguration. During the real-time dialogue perception phase, the large language model analyzes the environmental audio and video content while detecting natural language configuration commands issued by the user that begin with a specific wake word. The detected parameter changes are included in the analysis results in the form of structured data fields and returned. The system automatically parses the structured data fields and hot-updates the sensor sampling parameter strategy of the wearable device, realizing touchless parameter adjustment during the perception process. It also supports voice action commands, allowing users to trigger system-level operations through natural language. System-level operations include ending the perception session and automatically generating a structured summary, and immediately triggering the analysis and uploading of the current content.

[0014] As a further improvement to this technical solution: the real-time dialogue perception engine includes output suppression and content authenticity assurance submodules. When no valid human speech is detected in the audio signal or the speech content is insufficient for accurate recognition, a silence flag is forcibly returned and any speculative dialogue content is prohibited from being generated. When the dialogue content is casual conversation and does not involve decision-making information that the user needs assistance with, the output of analysis suggestions is actively suppressed, and only the basic transcription function is retained. At the same time, during the speaker role labeling process, when the confidence level of acoustic feature comparison is lower than a preset threshold, the speech segment is forcibly labeled as a non-user role to prevent incorrect role assignment and realize adaptive output control based on dialogue value and content credibility.

[0015] The method based on scene-aware real-time accompaniment and structured memory generation includes the following steps: S1. Pre-preparation stage: Collect short-term voice samples of the user in advance, encrypt and store the voice samples on the user's local device, and complete the initial configuration of the system's basic operating environment. S2, Scene Triggering and Recognition Stage: The user inputs a scene description into the system using natural language. The system uses a large language model to recognize the user's scene intent and outputs a standardized scene classification label. S3, Parameter Configuration and Rule Injection Stage: Based on scene classification identifiers, the system calculates and distributes the sampling parameter strategy of the wearable device's underlying hardware sensors through a large language model. At the same time, based on scene classification identifiers, it obtains knowledge and rule information of the corresponding domain, injects the information into the AI ​​inference context, and completes the scene adaptation configuration. S4. In the real-time accompaniment and dialogue perception stage, the system monitors the energy value of the ambient audio signal in real time through the voice activity detection submodule during the continuous scene. When a valid voice signal is detected, recording is automatically started. When the continuous silence timeout or the maximum acquisition time is reached, the current recording segment is automatically cut off and analysis and uploading are triggered. While analyzing the audio and video content, the large language model detects the voice configuration commands and action commands issued by the user, automatically parses and hot-updates the sensor sampling parameter strategy or triggers system-level operations. When the credibility of the analysis results is insufficient or the dialogue does not involve decision-making information that the user needs to assist, the output suppression submodule actively suppresses the output of the analysis suggestions and maintains a silent state. S5. In the structured information extraction and distribution stage, after the scene ends, the system submits the complete dialogue record to the large language model, extracts multi-dimensional structured data, and distributes the structured data to the corresponding functional subsystems through the intelligent routing module. S6. In the memory storage and presentation stage, the system encrypts and stores multi-dimensional structured data on the user's local device, generates visual memory information containing scene classification labels, and completes the entire process.

[0016] Compared with the prior art, the beneficial effects of the present invention are: 1. This invention achieves intelligent dynamic mapping from scene semantics to hardware physical configuration through a sensor parameter dynamic configuration module. Based on the perception needs of different scenarios, it dynamically calculates and distributes sampling parameter strategies for wearable device sensors through a large language model. While ensuring scene perception accuracy, it significantly reduces device power consumption and extends the battery life of wearable devices, solving the problem of insufficient battery life caused by continuous high-load data acquisition in existing technologies. Through the multimodal capabilities of the large language model, it achieves training-free speaker recognition. Only short-term user speech samples are needed as reference; speaker role labeling can be completed through acoustic feature comparison of the large language model, eliminating the need for traditional MFCC feature extraction, voiceprint vector matching, and pre-training processes. This significantly reduces the computing power requirements and deployment threshold of edge devices, while improving the accuracy of role differentiation in multi-person dialogue scenarios.

[0017] 2. This invention adopts a local-first privacy protection architecture. All data extraction, processing, and storage are completed on the user's local device. The user's voiceprint biometric data is only temporarily loaded into local memory and is not persistently stored on any third-party server. This avoids the risk of centralized leakage of sensitive dialogue data and biometric data, and can meet the privacy compliance requirements of sensitive scenarios such as medical consultation and business negotiations. Through the external knowledge source retrieval module, knowledge and rule information in the corresponding field can be obtained in real time based on the scenario classification label. There is no need to perform model pre-training and function adaptation for specific scenarios. It can be quickly adapted to any scenario in open domains such as medical, business, chess and card games, and education, which greatly expands the applicability and scenario adaptation capability of the system.

[0018] 3. This invention achieves adaptive acquisition based on audio signal energy through a voice activity detection submodule. Recording is initiated only when a valid voice signal is detected. A dual-trigger cutoff mechanism of silent timeout and maximum acquisition duration is employed for automatic segmented uploading. Compared to the continuous acquisition mode of existing technologies, this reduces invalid data transmission and computational resource consumption by approximately 60%. Through a runtime voice reconfiguration mechanism, users can adjust sensor sampling parameters in real time during the perception process using natural language voice commands, without interrupting the perception process or performing touch operations, achieving zero-touch parameter adjustment for wearable devices. Through output suppression and content authenticity assurance submodules, the system can adaptively control the output and suppression suggestions based on the value of the dialogue content and the credibility of audio recognition, avoiding redundant interference information in non-critical dialogues and preventing the generation of speculative content due to insufficient audio quality.

[0019] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it according to the contents of the specification, the preferred embodiments of the present invention are described in detail below with reference to the accompanying drawings. Specific embodiments of the present invention are given in detail below with reference to the accompanying drawings. Attached Figure Description

[0020] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings: Figure 1 This is a schematic diagram of the structure of a scene-aware real-time accompanying and structured memory generation method and system. Detailed Implementation

[0021] The principles and features of the present invention are described below with reference to the accompanying drawings. The examples given are for illustrative purposes only and are not intended to limit the scope of the invention. The invention is described more specifically in the following paragraphs by way of example with reference to the accompanying drawings. It should be noted that the drawings are in a very simplified form and use non-precise proportions, and are only used to facilitate and clarify the illustration of the embodiments of the present invention.

[0022] Please see Figure 1 In this embodiment of the invention, the scene-aware real-time accompanying and structured memory generation system includes an audio and video acquisition module, a scene intent recognition and scheduling module, a sensor parameter dynamic configuration module, an external knowledge source retrieval module, a real-time dialogue perception engine, a voiceprint reference sample acquisition and encrypted storage module, a structured information extraction module, an intelligent routing module, and a memory persistence and presentation module. The scene intent recognition and scheduling module receives natural language descriptions input by users, performs intent classification through a large language model, and outputs standardized scene classification labels. The sensor parameter dynamic configuration module receives the scene classification identifier output by the scene intent recognition and scheduling module, calculates and distributes the sampling parameter strategy of the wearable device's underlying hardware sensor through the large language model. The sampling parameter strategy includes the voice activity detection sensitivity threshold, the maximum duration of a single acquisition, the silent cut-off timeout duration, the visual sensor on / off status, the data transmission quality level, and the video stream acquisition mode. The voiceprint reference sample acquisition and encrypted storage module acquires short-term voice samples of the user, encrypts and stores the voice samples on the user's local device, and submits the voice samples as part of the multimodal input along with the real-time dialogue audio to the large language model. The large language model completes the speaker role labeling through acoustic feature comparison, distinguishing the user from other dialogue participants, without the need for voiceprint vector matching based on traditional signal processing. The external knowledge source retrieval module receives scene classification identifiers, obtains knowledge and rule information in the corresponding fields, and injects the information into the AI ​​reasoning context. External knowledge sources include search engines, cloud knowledge graphs, and local pre-built vector libraries. The real-time dialogue perception engine collects environmental audio and video signals according to the issued sampling parameter strategy, completes speech recognition and multi-role labeling, and performs semantic analysis by combining the injected knowledge and rule information; The structured information extraction module submits the complete dialogue record to the large language model to extract multi-dimensional structured data, including summaries, key notes, task items, and financial records. The intelligent routing module synchronously distributes the structured data to the corresponding functional subsystems, including the memory management subsystem, diary generation subsystem, financial management subsystem, task management subsystem, and location tagging subsystem. The memory persistence and presentation module encrypts and stores multi-dimensional structured data on the user's local device, generating visual memory information containing scene classification labels; The audio and video acquisition module includes a voice activity detection submodule. This submodule monitors the energy value of the audio signal in real time. When the energy value exceeds the preset voice activity threshold, it automatically starts recording. When the continuous silence duration exceeds the preset silence threshold or the recording duration reaches the preset maximum acquisition duration, it automatically cuts off the current recording segment and triggers analysis and uploading. It adopts a dual-trigger cutoff mechanism to achieve adaptive acquisition based on voice activity. It maintains a low-power silent listening state during periods without effective voice signals, reducing the transmission of invalid data and the consumption of computing resources. The sensor parameter dynamic configuration module supports runtime voice reconfiguration. During the real-time dialogue perception phase, the large language model analyzes the environmental audio and video content while detecting natural language configuration commands issued by the user that begin with a specific wake word. The detected parameter changes are included in the analysis results as structured data fields and returned. The system automatically parses the structured data fields and hot-updates the sensor sampling parameter strategy of the wearable device, realizing touchless parameter adjustment during the perception process. It also supports voice action commands, allowing users to trigger system-level operations through natural language. System-level operations include ending the perception session and automatically generating a structured summary, and immediately triggering the analysis and uploading of the current content. The real-time dialogue perception engine includes output suppression and content authenticity assurance submodules. When no valid human speech is detected in the audio signal or the speech content is insufficient for accurate recognition, a silence flag is forcibly returned and any speculative dialogue content is prohibited from being generated. When the dialogue content is casual conversation and does not involve decision-making information that the user needs assistance with, the output of analysis suggestions is actively suppressed, and only the basic transcription function is retained. At the same time, during the speaker role labeling process, when the confidence of acoustic feature comparison is lower than a preset threshold, the speech segment is forcibly labeled as a non-user role to prevent incorrect role assignment and achieve adaptive output control based on dialogue value and content credibility. The method based on scene-aware real-time accompaniment and structured memory generation includes the following steps: S1. Pre-preparation stage: Collect short-term voice samples of the user in advance, encrypt and store the voice samples on the user's local device, and complete the initial configuration of the system's basic operating environment. S2, Scene Triggering and Recognition Stage: The user inputs a scene description into the system using natural language. The system uses a large language model to recognize the user's scene intent and outputs a standardized scene classification label. S3, Parameter Configuration and Rule Injection Stage: Based on scene classification identifiers, the system calculates and distributes the sampling parameter strategy of the wearable device's underlying hardware sensors through a large language model. At the same time, based on scene classification identifiers, it obtains knowledge and rule information of the corresponding domain, injects the information into the AI ​​inference context, and completes the scene adaptation configuration. S4. In the real-time accompaniment and dialogue perception stage, the system monitors the energy value of the ambient audio signal in real time through the voice activity detection submodule during the continuous scene. When a valid voice signal is detected, recording is automatically started. When the continuous silence timeout or the maximum acquisition time is reached, the current recording segment is automatically cut off and analysis and uploading are triggered. While analyzing the audio and video content, the large language model detects the voice configuration commands and action commands issued by the user, automatically parses and hot-updates the sensor sampling parameter strategy or triggers system-level operations. When the credibility of the analysis results is insufficient or the dialogue does not involve decision-making information that the user needs to assist, the output suppression submodule actively suppresses the output of the analysis suggestions and maintains a silent state. S5. In the structured information extraction and distribution stage, after the scene ends, the system submits the complete dialogue record to the large language model, extracts multi-dimensional structured data, and distributes the structured data to the corresponding functional subsystems through the intelligent routing module. S6. In the memory storage and presentation stage, the system encrypts and stores multi-dimensional structured data on the user's local device, generates visual memory information containing scene classification labels, and completes the entire process. Example 1: Medical Consultation Scenario Example: The application scenario of this embodiment is medical consultation in an English-speaking environment overseas. The main entities are AR smart glasses equipped with this system and a matching smartphone terminal. The specific implementation process is as follows: The first step is the pre-preparation stage; 3-5 seconds of normal speech samples of the user are collected in advance, the speech samples are encrypted using the AES-256 encryption algorithm, and the encrypted speech samples are stored in the local secure partition of the AR smart glasses, without being persistently stored on any third-party server, thus completing the initial configuration of the system's basic operating environment. The second step is the scene triggering and recognition stage. The user wears AR smart glasses and inputs a scene description into the system through natural language: "Help me see a doctor. I don't understand English very well." The system inputs the user's natural language description into the big language model, and completes the user's intent recognition and classification through the big language model, and outputs a standardized scene classification label for medical consultation. The third step is the parameter configuration and rule injection stage. The system inputs the scene classification label of medical consultation into the large language model. The large language model calculates and distributes the sampling parameter strategy of the underlying hardware sensor of the AR smart glasses. Specifically, the audio acquisition interval is set to 5 seconds, the single acquisition duration is set to 2 seconds, the visual sensor is enabled in keyframe capture mode, and the video stream adopts AI-triggered short video acquisition mode. When the AI ​​detects the doctor's display of examination reports, medicines, etc., it triggers a 5-10 second short video acquisition. The data transmission quality level is set to medium level. At the same time, the system uses the search engine and cloud medical knowledge graph in external knowledge sources to obtain knowledge in the fields of professional terminology rules, English-Chinese comparison information of commonly used medicines, and core precautions for consultation in real time. The acquired knowledge and rule information is injected into the AI ​​inference context of the large language model to complete the scene adaptation configuration. The fourth step is the real-time accompaniment and dialogue perception stage. During the consultation scenario, the system strictly follows the issued sampling parameter strategy, continuously collecting audio and video signals from the environment through the audio and video acquisition module of the AR smart glasses. The system performs speech recognition on the collected audio signals, converting the speech content into text data. Simultaneously, it decrypts the locally stored user voiceprint reference samples and temporarily loads them into the running memory, submitting them along with the real-time dialogue audio as multimodal inputs to the large language model. The large language model uses acoustic feature comparison to complete speaker role labeling, distinguishing the dialogue content between the user and the doctor / nurse. The system combines injected medical domain knowledge and rule information to perform real-time semantic analysis of the dialogue content, providing real-time Chinese translation of the doctor's English instructions, and then using AR smart glasses... The glasses' display module pushes information to the user. During the consultation, the voice activity detection submodule continuously monitors the ambient audio energy value. When it detects that the doctor or user has started speaking, it automatically starts recording. When the conversation is paused for more than 2.5 seconds or the maximum duration of a single recording segment is reached, it automatically cuts off and triggers AI analysis. During the consultation, the user can adjust the collection strategy through voice commands, such as saying "ReMi" to adjust the recording to 10 seconds. The large language model recognizes this configuration command while analyzing the current audio and automatically updates the maximum duration of a single collection to 10 seconds, without any touch operation. When non-critical conversations between doctors and nurses are collected, the output suppression submodule determines that the conversation does not involve decision-making information that the user needs assistance with, actively suppresses the analysis and suggests outputting suggestions, and only performs basic transcription recording. The fifth step is the structured information extraction and distribution stage. After the consultation scenario ends, the system submits the complete dialogue record of this consultation to the large language model. Multi-dimensional structured data is extracted from the dialogue record, including a summary of the consultation content, key notes of the doctor's diagnosis, drug names and dosage tasks, financial records of drug and examination fees, and location tags for follow-up visits. The system uses an intelligent routing module to synchronously distribute the extracted structured data to the corresponding functional subsystems. The summary and key notes are distributed to the memory management subsystem, the drug administration and follow-up visit tasks are distributed to the task management subsystem, the cost data is distributed to the financial management subsystem, and the follow-up visit location information is distributed to the location tagging subsystem. The sixth step is the memory storage and presentation stage. The system encrypts all structured data using the AES-256 algorithm and stores it on the local devices of the AR smart glasses and smartphones, without storing it in plaintext in the cloud. At the same time, it generates a visual memory card containing medical consultation scenario classification labels. The memory card contains a foldable, role-based display of the complete dialogue record. It also generates a system-level local reminder notification for follow-up consultation time, completing the entire process of this consultation scenario.

[0023] Example 2: Card and Board Game Entertainment Scenario Example: The application scenario of this embodiment is offline entertainment of Sichuan Mahjong. The main implementers are AR smart glasses equipped with this system and a matching smartphone terminal. The specific implementation process is as follows: The first step, the pre-preparation stage, is the same as in Example 1, which involves collecting, encrypting, and storing user voiceprint samples locally, and completing system initialization. The second step is the scene triggering and recognition stage; the user inputs a scene description into the system through natural language: "Help me play Sichuan Mahjong". The system completes the intent recognition through a large language model and outputs a standardized scene classification label: "Chess and Card Entertainment". The third step is the parameter configuration and rule injection stage. Based on scene classification and identification, the system calculates and distributes sensor sampling parameter strategies through a large language model. Specifically, the audio acquisition interval is set to 15 seconds, the single acquisition duration is set to 3 seconds, the visual sensor is enabled in intermittent photo mode, the continuous video recording function is disabled, and the card photo capture is only triggered at the end of each game of mahjong. The data transmission quality level is set to low level. At the same time, the system obtains the scoring rules and scoring rules of Sichuan mahjong from external knowledge sources and injects them into the AI ​​inference context. The fourth step is the real-time companion and dialogue perception stage. During the mahjong game, the system collects environmental audio and video signals according to the issued low-power sampling strategy, completes voice recognition and speaker role labeling, and records the wins and losses and chip exchange data of each game in real time in combination with the mahjong scoring rules. The fifth step is the structured information extraction and distribution stage. After the game ends, the system submits the complete dialogue and game record to the big language model to extract the game win / loss summary, the financial record of each game's win / loss, and the game time and location tags. The financial record is distributed to the financial management subsystem, the game summary is distributed to the memory management subsystem, and the location information is distributed to the location tagging subsystem through the intelligent routing module. Step 6, memory storage and presentation stage; the system encrypts and stores all data on the user's local device, generates a win / loss record memory card containing classification identifiers for chess and card game entertainment scenarios, and completes the entire process. Comparison of existing technologies in proportional card and board game entertainment scenarios: This comparative example uses the same hardware and application scenario as Example 2. The difference is that the sensor dynamic parameter configuration strategy of this invention is not adopted. Instead, the default normal audio and video acquisition mode of the prior art is used, that is, continuous audio acquisition and continuous visual camera capture are continuously enabled. The rest of the environment and operation are completely the same. In the same 3-hour chess and card game scenario test, the voice activity adaptive acquisition strategy of Example 2 of this invention reduces the number of AI analysis interface calls by about 60%, the amount of audio data transmission by about 63%, and the overall power consumption of the device by about 40% to 50% compared with the continuous acquisition mode of this comparative example.

[0024] Example 3: Business Negotiation Scenario Example: The application scenario of this embodiment is a cross-language business negotiation scenario. The executing entities are AR smart glasses equipped with this system and a matching smartphone terminal. The specific implementation process is as follows: The first step, the pre-preparation stage, is the same as in Example 1, which involves collecting, encrypting, and storing user voiceprint samples locally, and completing system initialization. The second step is the scenario triggering and recognition stage. The user inputs a scenario description into the system using natural language: "Help me conduct business negotiations and I need to record the key points of cooperation." The system uses a large language model to complete the intent recognition and outputs a standardized scenario classification label for business negotiations. The third step is the parameter configuration and rule injection stage. Based on scene classification and identification, the system calculates and distributes sensor sampling parameter strategies through a large language model. Specifically, the audio acquisition interval is set to 5 seconds, the single acquisition duration is set to 4 seconds, the visual sensor is enabled in continuous video recording mode, the video stream is automatically segmented according to topic switching, each segment is 30-60 seconds long, and the data transmission quality level is set to high level. At the same time, the system obtains the core contract terms and rules of business negotiations and the precautions for cooperation negotiations from external knowledge sources and injects them into the AI ​​inference context. The fourth step is the real-time accompaniment and dialogue perception stage; during the negotiation, the system collects environmental audio and video signals according to the issued sampling strategy, completes speech recognition, multi-role labeling and real-time semantic analysis, and extracts and pushes the core cooperation terms and demands of both parties in real time. The fifth step is the structured information extraction and distribution stage. After the negotiation, the system submits the complete dialogue record to the large language model to extract the negotiation content summary, key cooperation notes, contract tasks to be executed, business financial records, negotiation location and partner information tags, and distributes the corresponding data to each functional subsystem through the intelligent routing module. The sixth step is the memory storage and presentation stage; the system encrypts and stores all data on the user's local device, generates visual memory cards with classification labels for business negotiation scenarios, generates system-level local reminders for tasks to be performed, and completes the entire process.

[0025] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any way. Those skilled in the art can readily implement the present invention based on the description and drawings above. However, any modifications, alterations, and variations made by those skilled in the art without departing from the scope of the present invention using the disclosed technical content are equivalent embodiments of the present invention. Furthermore, any modifications, alterations, and variations made to the above embodiments based on the essential technology of the present invention are still within the protection scope of the present invention.

Claims

1. A scene-aware, real-time accompanying, and structured memory generation system, characterized in that: It includes an audio and video acquisition module, a scene intent recognition and scheduling module, a sensor parameter dynamic configuration module, an external knowledge source retrieval module, a real-time dialogue perception engine, a voiceprint reference sample acquisition and encrypted storage module, a structured information extraction module, an intelligent routing module, and a memory persistence and presentation module.

2. The scene-aware real-time accompanying and structured memory generation system according to claim 1, characterized in that, The scene intent recognition and scheduling module receives natural language descriptions input by the user, performs intent classification through a large language model, and outputs standardized scene classification identifiers.

3. The scene-aware real-time accompanying and structured memory generation system according to claim 1, characterized in that, The sensor parameter dynamic configuration module receives the scene classification identifier output by the scene intent recognition and scheduling module, calculates and distributes the sampling parameter strategy of the wearable device's underlying hardware sensors through a large language model. The sampling parameter strategy includes the voice activity detection sensitivity threshold, the maximum duration of a single acquisition, the silent cut-off timeout duration, the visual sensor on / off status, the data transmission quality level, and the video stream acquisition mode.

4. The scene-aware real-time accompanying and structured memory generation system according to claim 1, characterized in that, The voiceprint reference sample acquisition and encrypted storage module acquires short-term voice samples of the user, encrypts and stores the voice samples on the user's local device, and submits the voice samples as part of the multimodal input along with the real-time dialogue audio to the large language model. The large language model completes the speaker role labeling through acoustic feature comparison, distinguishing the user from other dialogue participants, without the need for voiceprint vector matching based on traditional signal processing.

5. The scene-aware real-time accompanying and structured memory generation system according to claim 1, characterized in that, The external knowledge source retrieval module receives scene classification identifiers, obtains knowledge and rule information in the corresponding fields, and injects the information into the AI ​​reasoning context. External knowledge sources include search engines, cloud knowledge graphs, and local pre-built vector libraries.

6. The scene-aware real-time accompanying and structured memory generation system according to claim 1, characterized in that, The real-time dialogue perception engine collects environmental audio and video signals according to the issued sampling parameter strategy, completes speech recognition and multi-role labeling, and performs semantic analysis by combining the injected knowledge and rule information.

7. The scene-aware real-time accompanying and structured memory generation system according to claim 1, characterized in that, The structured information extraction module submits the complete dialogue record to the large language model to extract multi-dimensional structured data, including summaries, key notes, task items, and financial records. The intelligent routing module synchronously distributes the structured data to the corresponding functional subsystems, including the memory management subsystem, diary generation subsystem, financial management subsystem, task management subsystem, and location tagging subsystem.

8. The scene-aware real-time accompanying and structured memory generation system according to claim 1, characterized in that, The memory persistence and presentation module encrypts and stores multi-dimensional structured data on the user's local device, generating visual memory information containing scene classification identifiers.

9. The scene-aware real-time accompanying and structured memory generation system according to claim 1, characterized in that, The audio and video acquisition module includes a voice activity detection submodule. This submodule monitors the energy value of the audio signal in real time. When the energy value exceeds the preset voice activity threshold, recording is automatically started. When the continuous silence duration exceeds the preset silence threshold or the recording duration reaches the preset maximum acquisition duration, the current recording segment is automatically cut off and analysis and uploading are triggered. A dual-trigger cutoff mechanism is adopted to achieve adaptive acquisition based on voice activity. During periods without effective voice signals, a low-power silent listening state is maintained, reducing the transmission of invalid data and the consumption of computing resources.

10. The scene-aware real-time accompanying and structured memory generation system according to claim 1, characterized in that, The sensor parameter dynamic configuration module supports runtime voice reconfiguration. During the real-time dialogue perception phase, the large language model analyzes the environmental audio and video content while detecting natural language configuration commands issued by the user that begin with a specific wake word. The detected parameter changes are included in the analysis results as structured data fields and returned. The system automatically parses the structured data fields and hot-updates the sensor sampling parameter strategy of the wearable device, realizing touchless parameter adjustment during the perception process. It also supports voice action commands, allowing users to trigger system-level operations through natural language. System-level operations include ending the perception session and automatically generating a structured summary, and immediately triggering the analysis and uploading of the current content.

11. The scene-aware real-time accompanying and structured memory generation system according to claim 1, characterized in that, The real-time dialogue perception engine includes output suppression and content authenticity assurance submodules. When no valid human speech is detected in the audio signal or the speech content is insufficient for accurate recognition, it forces a return to a silence flag and prohibits the generation of any speculative dialogue content. When the dialogue content is casual conversation and does not involve decision-making information that the user needs assistance with, the output of analysis suggestions is actively suppressed, and only the basic transcription function is retained. At the same time, during the speaker role labeling process, when the confidence of the acoustic feature comparison is lower than the preset threshold, the speech segment is forcibly labeled as a non-user role to prevent incorrect role assignment and achieve adaptive output control based on dialogue value and content credibility.

12. A method for scene-aware real-time accompaniment and structured memory generation, applied to the scene-aware real-time accompaniment and structured memory generation system according to any one of claims 1-11, characterized in that, Includes the following steps: S1. Pre-preparation stage: Collect short-term voice samples of the user in advance, encrypt and store the voice samples on the user's local device, and complete the initial configuration of the system's basic operating environment. S2, Scene Triggering and Recognition Stage: The user inputs a scene description into the system using natural language. The system uses a large language model to recognize the user's scene intent and outputs a standardized scene classification label. S3, Parameter Configuration and Rule Injection Stage: Based on scene classification identifiers, the system calculates and distributes the sampling parameter strategy of the wearable device's underlying hardware sensors through a large language model. At the same time, based on scene classification identifiers, it obtains knowledge and rule information of the corresponding domain, injects the information into the AI ​​inference context, and completes the scene adaptation configuration. S4. Real-time Accompaniment and Dialogue Awareness Stage: During the continuous scene, the system monitors the energy value of the ambient audio signal in real time through the voice activity detection submodule. When a valid voice signal is detected, recording is automatically started. When continuous silence exceeds the time limit or the maximum collection time is reached, the current recording segment is automatically cut off and analysis and uploading are triggered. While analyzing audio and video content, the large language model detects the user's voice configuration commands and action commands, automatically parses and hot-updates sensor sampling parameter strategies or triggers system-level operations; when the credibility of the analysis results is insufficient or the dialogue does not involve decision-making information that the user needs to assist, the output suppression submodule actively suppresses the output of analysis suggestions and remains silent. S5. In the structured information extraction and distribution stage, after the scene ends, the system submits the complete dialogue record to the large language model, extracts multi-dimensional structured data, and distributes the structured data to the corresponding functional subsystems through the intelligent routing module. S6. In the memory storage and presentation stage, the system encrypts and stores multi-dimensional structured data on the user's local device, generates visual memory information containing scene classification labels, and completes the entire process.