This invention discloses a method and
system for real-time accompaniment and structured memory generation based on scene
perception, relating to the fields of
artificial intelligence,
augmented reality, and
mobile computing. The
system includes core modules such as an audio and video acquisition module, a scene
intent recognition and scheduling module, a sensor parameter dynamic configuration module, a real-time dialogue
perception engine, and a structured
information extraction module. This invention achieves dynamic optimization of hardware parameters driven by scene
semantics, balancing
perception accuracy and device
power consumption. It realizes speaker recognition without training through a large
language model, and adopts a local priority architecture to protect data privacy. It can achieve seamless continuous accompaniment after a single trigger, completing real-time understanding of dialogue content, multi-dimensional structured
information extraction, and local memory generation. At the same time, through a
voice activity adaptive acquisition mechanism and a runtime voice command reconfiguration mechanism, it achieves low-power acquisition and touchless parameter adjustment, making it suitable for various scenarios such as
medical consultation, business negotiation, and daily social interaction.