Intelligent question-answering system and method based on multi-modal feature fusion
By employing quantum behavior acquisition and multimodal feature fusion technologies, the problem of insufficient input dimensions in intelligent question-answering systems under privacy protection regulations has been solved, enabling highly accurate and reliable real-time passenger flow analysis in scenarios such as commercial complexes.
Patent Information
- Application Number
- CN202511394527.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-28
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2045-09-28
AI Technical Summary
Existing multimodal feature fusion intelligent question answering systems suffer from insufficient input dimensions due to privacy protection regulations restricting the collection of biometric data, which affects the accuracy of abnormal behavior identification and the reliability of decision-making in real-time passenger flow analysis scenarios.
The quantum behavior acquisition module captures microscopic human behavioral characteristics through a Josephson junction array sensor, a single-photon avalanche diode matrix, and a superconducting quantum interference device, generating a non-Euclidean spatiotemporal behavioral feature tensor. Combined with a neural radiation field simulation module, a meta-learning feature encoding module, a game reasoning engine module, and a chaotic rule evolution module, it achieves the fusion of multimodal features and accurate identification of abnormal behaviors.
Expanding the dimensions of input data while adhering to privacy compliance improves the accuracy of abnormal behavior identification and the reliability of decision-making, significantly enhancing the real-time passenger flow analysis capabilities in scenarios such as commercial complexes.
Smart Images

Figure CN120892538B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, and in particular to an intelligent question-answering system and method based on multi-modal feature fusion. BACKGROUND
[0002] An intelligent question-answering system is an artificial intelligence application based on natural language processing and machine learning techniques that can automatically understand the semantic of user-posed questions, extract relevant data from knowledge bases or internet resources through information retrieval methods, and generate accurate and coherent natural language answers using reasoning mechanisms. This system relies on trained models to identify problem patterns and answer associations, such as providing solutions after analyzing user query intentions in customer service scenarios, thereby improving information acquisition efficiency and supporting diverse applications such as educational assistance and intelligent assistants.
[0003] Existing multi-modal feature fusion intelligent question-answering systems have the following technical pain points. After privacy protection regulations forcibly restrict the collection of biometric data, the system is forced to rely only on human behavior analysis algorithms for real-time passenger flow analysis, but the lack of identity features in behavior data leads to insufficient input dimensions. For example, in a commercial complex scenario, a passenger flow statistics system can only infer crowd density based on human posture and movement trajectory, but cannot verify individual identities or accurately track specific personnel trajectories through facial features. When abnormal behaviors such as lingering or gathering occur, the algorithm is prone to misjudging normal visitors as risk targets due to the lack of features, resulting in increased false positive rates and weakened system decision reliability. SUMMARY
[0004] To address the deficiencies in the prior art, the present application provides an intelligent question-answering system and method based on multi-modal feature fusion. The present application solves the technical problem of insufficient input dimensions in multi-modal feature fusion intelligent question-answering systems due to privacy protection regulations forcibly restricting the collection of biometric data, affecting the accuracy of abnormal behavior recognition and the reliability of decision-making in real-time passenger flow analysis scenarios.
[0005] To solve the above technical problems, the specific content of the present application is as follows:
[0006] In a first aspect, the present application provides an intelligent question-answering system based on multi-modal feature fusion, comprising:
[0007] A quantum behavior acquisition module is deployed to capture human microscopic behavior features using a quantum sensor array, generating a non-Euclidean space-time behavior feature tensor. The quantum sensor array includes a Josephson junction array sensor, a single-photon avalanche diode matrix, and a superconducting quantum interference device.
[0008] The neural radiance field simulation module receives the behavior feature tensor output by the quantumized behavior acquisition module, converts the behavior feature tensor into a differentiable rendering parameter, generates an initial rendering parameter, introduces a generative adversarial network architecture to create an adversarial sample data of the occluded crowded scene for the initial rendering parameter, designs a hyperbolic geometry embedding model to align the entity virtual coordinates, processes the adversarial sample data and outputs an enhanced spatiotemporal sequence;
[0009] The meta-learning feature coding module receives the enhanced spatiotemporal sequence output by the neural radiance field simulation module, extracts hierarchical features of the behavior features through a variational autoencoder, integrates a neural Turing machine architecture to store scene migration knowledge, fuses the hierarchical features to generate memory-enhanced features, optimizes the feature extractor parameters by implementing a model-agnostic meta-learning algorithm, processes the memory-enhanced features and outputs a 128-dimensional hyper vector;
[0010] The game reasoning engine module receives the 128-dimensional hyper vector output by the meta-learning feature coding module, simulates a group interaction strategy, and generates a group behavior probability matrix;
[0011] The chaotic rule evolution module receives the group behavior probability matrix output by the game reasoning engine module, analyzes the stability of the decision trajectory, and outputs a dynamic decision rule;
[0012] The cross-modal consensus module receives the dynamic decision rule output by the chaotic rule evolution module, fuses the behavior feature tensor of the quantumized behavior acquisition module, the enhanced spatiotemporal sequence output by the neural radiance field simulation module, and the group behavior probability matrix output by the game reasoning engine module, and outputs a hierarchical early warning signal. The hierarchical early warning signal is fed back to the quantumized behavior acquisition module to adjust the sampling frequency.
[0013] Further, the intelligent question and answer system based on multi-modal feature fusion of the present application, the quantumized behavior acquisition module comprises:
[0014] The quantum tunneling effect measurement unit processes the human muscle microtremor through a Josephson junction array sensor and outputs microtremor frequency data.
[0015] The gait optical flow field phase difference detection unit receives the microtremor frequency data output by the quantum tunneling effect measurement unit, processes the microtremor frequency data through a single-photon avalanche diode matrix, and outputs motion vector data. The superconducting quantum interference device unit receives the motion vector data output by the gait optical flow field phase difference detection unit, processes the motion vector data through a superconducting quantum interference device configured with a gradient meter structure, and outputs electromagnetic disturbance data.
[0016] The manifold learning dimension reduction unit receives the micro-tremor frequency data output by the quantum tunneling effect measurement unit, the motion vector data output by the gait optical flow field phase difference detection unit, and the electromagnetic disturbance data output by the superconducting quantum interference device unit, processes the micro-tremor frequency data, the motion vector data, and the electromagnetic disturbance data through a manifold learning algorithm, generates a behavior feature tensor, and outputs the behavior feature tensor to the neural radiance field simulation module.
[0017] Further, the neural radiance field simulation module of the intelligent question and answer system based on multi-modal feature fusion of the present application comprises:
[0018] The dynamic scene reconstruction unit receives the behavior feature tensor output by the quantized behavior collection module, processes the behavior feature tensor through a differentiable rendering, and outputs initial rendering parameters.
[0019] The synthetic data injection unit receives the initial rendering parameters output by the dynamic scene reconstruction unit, processes the initial rendering parameters through a generative adversarial network architecture, and outputs adversarial sample data.
[0020] The spacetime mapping function unit receives the adversarial sample data output by the synthetic data injection unit, processes the adversarial sample data through a hyperbolic geometry embedding model, and outputs an enhanced spacetime sequence to the meta-learning feature encoding module.
[0021] Further, the meta-learning feature encoding module of the intelligent question and answer system based on multi-modal feature fusion of the present application comprises:
[0022] The variational autoencoder unit receives the enhanced spacetime sequence output by the neural radiance field simulation module, processes the enhanced spacetime sequence through a capsule network structure, and outputs hierarchical features.
[0023] The memory enhancement network unit receives the hierarchical features output by the variational autoencoder unit, processes the hierarchical features through a neural Turing machine architecture, and outputs memory enhancement features.
[0024] The meta-learning strategy unit receives the memory enhancement features output by the memory enhancement network unit, processes the memory enhancement features through a model-agnostic meta-learning algorithm, and outputs a 128-dimensional hyper-vector to the game reasoning engine module. Further, the game reasoning engine module of the intelligent question and answer system based on multi-modal feature fusion of the present application comprises:
[0025] The agent policy network unit receives the 128-dimensional hyper-vector output by the meta-learning feature encoding module, processes the 128-dimensional hyper-vector through a graph attention mechanism, and outputs interaction policy data.
[0026] The Nash equilibrium solver unit receives the interaction policy data output by the agent policy network unit, processes the interaction policy data through a deep virtual game framework, and outputs equilibrium policy data.
[0027] The Monte Carlo tree search unit receives the equilibrium strategy data output by the Nash equilibrium solver unit, processes the equilibrium strategy data through time difference learning, and outputs a group behavior probability matrix to the chaotic rule evolution module.
[0028] Further, the intelligent question and answer system based on multi-modal feature fusion of the present application, the chaotic rule evolution module comprises:
[0029] The Lyapunov index analysis unit receives the group behavior probability matrix output by the game reasoning engine module, processes the group behavior probability matrix through a recurrent neural network, and outputs a stability index.
[0030] The risk pattern reconstruction unit receives the stability index output by the Lyapunov index analysis unit, processes the stability index through multiple fractal detrended fluctuation analysis, and outputs a reconstructed risk pattern.
[0031] The early warning threshold function optimization unit receives the reconstructed risk pattern output by the risk pattern reconstruction unit, processes the reconstructed risk pattern through adaptive decision boundary setting, and outputs a dynamic decision rule to the cross-modal consensus module.
[0032] Further, the intelligent question and answer system based on multi-modal feature fusion of the present application, the cross-modal consensus module comprises:
[0033] The cross-validation mechanism unit receives the dynamic decision rule output by the chaotic rule evolution module, processes the dynamic decision rule, the behavior feature tensor output by the quantumized behavior acquisition module, the enhanced spatiotemporal sequence output by the neural radiation field simulation module, and the group behavior probability matrix output by the game reasoning engine module through evidence theory fusion, and outputs a fusion result.
[0034] The chaotic rule correction unit receives the fusion result output by the cross-validation mechanism unit, processes the fusion result through nonlinear filtering, and outputs correction data.
[0035] The credibility weight allocation unit receives the correction data output by the chaotic rule correction unit, processes the correction data through a fuzzy logic controller, and outputs a four-dimensional early warning vector.
[0036] The four-dimensional early warning vector comprises an event type, a probability, an urgency, and a disposal suggestion.
[0037] Further, the intelligent question and answer system based on multi-modal feature fusion of the present application further comprises:
[0038] When the synthetic data injection unit of the neural radiation field simulation module outputs adversarial sample data, the memory enhancement network unit of the meta-learning feature encoding module processes historical scene parameters through a neural Turing machine architecture, and outputs enhanced features to the meta-learning strategy unit of the meta-learning feature encoding module.
[0039] Further, the intelligent question and answer system based on multi-modal feature fusion of the present application further comprises:
[0040] When the Monte Carlo tree search unit of the game reasoning engine module outputs the group behavior probability matrix, the risk pattern reconstruction unit of the chaotic rule evolution module processes the fractal dimension through multiple fractal detrended fluctuation analysis, and outputs the optimized decision rule to the early warning threshold function optimization unit of the chaotic rule evolution module;
[0041] The four-dimensional early warning vector output by the cross-modal consensus module includes a credibility weight and is output to the quantumized behavior collection module.
[0042] In a second aspect, the present application provides an intelligent question and answer method based on multi-modal feature fusion, applied to the intelligent question and answer system based on multi-modal feature fusion, comprising:
[0043] Step 1, deploying a quantum sensor array to capture human micro-behavior characteristics, generating a non-Euclidean space-time behavior feature tensor, the quantum sensor array including a Josephson junction array sensor, a single-photon avalanche diode matrix, and a superconducting quantum interference device;
[0044] Step 2, receiving the behavior feature tensor output by step 1, converting the behavior feature tensor into differentiable rendering parameters, generating initial rendering parameters, introducing a generative adversarial network architecture to create adversarial sample data for the initial rendering parameters to create a blocked crowded scene, designing a hyperbolic geometry embedding model to align entity virtual coordinates, processing the adversarial sample data and outputting an enhanced space-time sequence;
[0045] Step 3, receiving the enhanced space-time sequence output by step 2, extracting hierarchical features of the behavior features through a variational autoencoder, integrating a neural Turing machine architecture to store scene migration knowledge, fusing the hierarchical features to generate memory-enhanced features, implementing a model-agnostic meta-learning algorithm to optimize feature extractor parameters, processing the memory-enhanced features and outputting a 128-dimensional hyper-vector;
[0046] Step 4, receiving the hyper-vector output by step 3, simulating group interaction strategies, and generating a group behavior probability matrix;
[0047] Step 5, receiving the group behavior probability matrix output by step 4, analyzing the stability of the decision trajectory, and outputting a dynamic decision rule;
[0048] Step 6, receiving the dynamic decision rule output by step 5, fusing the behavior feature tensor of step 1, the enhanced space-time sequence output by step 2, and the group behavior probability matrix output by step 4, outputting a hierarchical early warning signal, and feeding back the hierarchical early warning signal to step 1 to adjust the sampling frequency.
[0049] The present application has the following advantages:
[0050] The beneficial effects of the present application are embodied in multi-level technical innovation: the quantumized behavior acquisition module captures muscle microtremor frequency through Josephson junction array sensors, analyzes motion vectors with single-photon avalanche diode matrix, and inverts electromagnetic disturbance patterns with superconducting quantum interference devices to generate non-Euclidean space-time behavior characteristic tensors, expand input data dimensions with non-biological features under privacy compliance; the neural radiation field simulation module decomposes behavior characteristic tensors using differentiable rendering technology, generates adversarial network architecture injection disturbances in occluded crowded scenes, and outputs enhanced space-time sequences with hyperbolic geometry embedding model alignment to compensate for data singularity defects caused by biological feature loss; the meta-learning feature encoding module extracts hierarchical features through capsule networks, the neural Turing machine integrates historical scene knowledge, and the model-agnostic meta-learning algorithm optimizes network parameters to output 128-dimensional hyper vectors, improving model cross-scene generalization capability; the game reasoning engine module models group interaction based on graph attention mechanism, converges Nash equilibrium with deep virtual game framework, and predicts behavior probability matrix with time-difference learning to achieve precise quantification of abnormal behavior evolution; the chaotic rule evolution module calculates Lyapunov stability index through recurrent neural network, reconstructs risk patterns with multi-fractal detrended fluctuation analysis, and adaptively determines boundary settings to dynamically optimize decision rules; the cross-modal consensus module fuses multi-modal data confidence using evidence theory, corrects system bias with nonlinear filtering, and generates four-dimensional early warning vectors with fuzzy logic controller to feedback quantum sensor sampling frequency, forming a closed-loop optimization mechanism for data acquisition and risk response, significantly improving abnormal behavior recognition accuracy and decision reliability in real-time passenger flow scenarios such as commercial complexes. BRIEF DESCRIPTION OF DRAWINGS
[0051] In order to more clearly illustrate the technical solutions of the present application, the following will briefly introduce the drawings needed in the embodiments. Obviously, for those skilled in the art, other drawings can also be obtained without creative labor under the premise of the drawings.
[0052] Figure 1 The flowchart of the intelligent question and answer method based on multi-modal feature fusion provided by the embodiments of the present application. DETAILED DESCRIPTION
[0053] In order to make the technical solutions of the present application clearer, the following will combine the specific embodiments of the present application and the corresponding drawings to clearly and completely describe the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application. The following will combine the drawings to specifically describe the present application provided by the embodiments of the present application. In order to better understand the purpose of the present application, the following will further describe the present application in detail.
[0054] In a first aspect, the application provides an intelligent question-answering system based on multi-modal feature fusion, comprising:
[0055] A quantumized behavior acquisition module is configured to capture human microscopic behavior features by deploying a quantum sensor array, generate a non-Euclidean space-time behavior feature tensor, and the quantum sensor array includes a Josephson junction array sensor, a single-photon avalanche diode matrix, and a superconducting quantum interference device.
[0056] A neural radiation field simulation module is configured to receive the behavior feature tensor output by the quantumized behavior acquisition module, convert the behavior feature tensor into differentiable rendering parameters, generate initial rendering parameters, introduce a generative adversarial network architecture to create adversarial sample data for the initial rendering parameters to create a crowded scene, design a hyperbolic geometry embedding model to align entity virtual coordinates, process the adversarial sample data and output an enhanced space-time sequence.
[0057] A meta-learning feature encoding module is configured to receive the enhanced space-time sequence output by the neural radiation field simulation module, extract hierarchical features of behavior features through a variational autoencoder, integrate a neural Turing machine architecture to store scene migration knowledge, fuse the hierarchical features to generate memory-enhanced features, implement a model-agnostic meta-learning algorithm to optimize feature extractor parameters, process the memory-enhanced features and output a 128-dimensional hyper-vector.
[0058] A game reasoning engine module is configured to receive the 128-dimensional hyper-vector output by the meta-learning feature encoding module, simulate group interaction strategies, and generate a group behavior probability matrix.
[0059] A chaotic rule evolution module is configured to receive the group behavior probability matrix output by the game reasoning engine module, analyze decision trajectory stability, and output dynamic decision rules.
[0060] A cross-modal consensus module is configured to receive the dynamic decision rules output by the chaotic rule evolution module, fuse the behavior feature tensor of the quantumized behavior acquisition module, the enhanced space-time sequence output by the neural radiation field simulation module, and the group behavior probability matrix output by the game reasoning engine module, and output a hierarchical early warning signal. The hierarchical early warning signal is fed back to the quantumized behavior acquisition module to adjust the sampling frequency.
[0061] The quantumization behavior acquisition module deploys a quantum sensor array to capture human microscopic behavior characteristics, generating a non-Euclidean space-time behavior characteristic tensor. The quantum sensor array is composed of a Josephson junction array sensor, a single-photon avalanche diode matrix, and a superconducting quantum interference device. The Josephson junction array sensor processes human muscle microtremor, outputs microtremor frequency data, which reflects the quantum-level characteristics of muscle activity. The gait optical flow field phase difference detection unit receives the microtremor frequency data, analyzes the motion vector through the single-photon avalanche diode matrix, and outputs the vector information of limb movement. The superconducting quantum interference device unit receives the motion vector data, and uses a device configured with a gradient meter structure to invert the electromagnetic disturbance mode, outputting electromagnetic disturbance data. The manifold learning dimension reduction unit integrates the microtremor frequency data, motion vector data, and electromagnetic disturbance data, applies a manifold learning algorithm for non-linear dimension reduction, and generates a behavior characteristic tensor, which is transmitted to the neural radiation field simulation module. This process realizes feature fusion in non-Euclidean space, avoids biological feature acquisition restrictions, and expands the input data dimension.
[0062] The neural radiation field simulation module receives the behavior characteristic tensor output by the quantumization behavior acquisition module, converts the behavior characteristic tensor into differentiable rendering parameters, and generates initial rendering parameters. The dynamic scene reconstruction unit uses differentiable rendering technology to decompose the radiation field parameters, processes the behavior characteristic tensor to output the initial rendering parameters. The synthetic data injection unit receives the initial rendering parameters, injects adversarial samples of occluded and crowded scenes through a generative adversarial network architecture, and outputs adversarial sample data. The space-time mapping function unit receives the adversarial sample data, aligns the real entity coordinates and virtual coordinates using a hyperbolic geometry embedding model, and outputs an enhanced space-time sequence to the meta-learning feature encoding module. This step compensates for the lack of real data through virtual scene enhancement, improving the model's generalization ability to complex environments.
[0063] The meta-learning feature encoding module receives the enhanced space-time sequence output by the neural radiation field simulation module, and extracts hierarchical features of the behavior characteristics through a variational autoencoder. The variational autoencoder unit uses a capsule network structure to analyze the spatial hierarchical relationship of the space-time sequence, and outputs hierarchical features. The memory enhancement network unit receives the hierarchical features, integrates a neural Turing machine architecture to store historical scene migration knowledge, and fuses the knowledge to output memory enhancement features. The meta-learning strategy unit receives the memory enhancement features, implements a model-agnostic meta-learning algorithm to optimize the feature extractor network weights, processes the memory enhancement features to output a 128-dimensional hyper-vector to the game reasoning engine module. This process strengthens the consistency of feature representation, supporting subsequent group interaction simulation.
[0064] The game reasoning engine module receives the 128-dimensional hyper vector output by the meta-learning feature encoding module, simulates the group interaction strategy, and generates a group behavior probability matrix. The agent strategy network unit calculates individual interaction weights through a graph attention mechanism, processes the 128-dimensional hyper vector, and outputs interaction strategy data. The Nash equilibrium solver unit receives the interaction strategy data, converges the strategy to a Nash equilibrium point using a deep virtual game framework, and outputs equilibrium strategy data. The Monte Carlo tree search unit receives the equilibrium strategy data, introduces a time-difference learning to predict the event evolution path, and outputs the group behavior probability matrix to the chaotic rule evolution module. This step realizes the modeling and prediction of group dynamic behavior, providing a probability basis for decision analysis.
[0065] The chaotic rule evolution module receives the group behavior probability matrix output by the game reasoning engine module, analyzes the stability of the decision trajectory, and outputs the dynamic decision rule. The Lyapunov index analysis unit applies a recurrent neural network to quantify the divergence of the decision trajectory, processes the group behavior probability matrix, and outputs a stability index. The risk pattern reconstruction unit receives the stability index, reconstructs the risk feature dimension through multiple fractal detrended fluctuation analysis, and outputs the reconstructed risk pattern. The early warning threshold function optimization unit receives the reconstructed risk pattern, optimizes the threshold function based on an adaptive decision boundary, and outputs the dynamic decision rule to the cross-modal consensus module. This module dynamically adjusts the decision boundary to enhance the response accuracy of the system to abnormal behavior.
[0066] The cross-modal consensus module receives the dynamic decision rule output by the chaotic rule evolution module, fuses the behavior feature tensor of the quantumized behavior acquisition module, the enhanced spatiotemporal sequence output by the neural radiation field simulation module, and the group behavior probability matrix output by the game reasoning engine module, and outputs a hierarchical early warning signal. The cross-validation mechanism unit fuses the confidence of multi-source data using evidence theory, processes the dynamic decision rule and the fusion result of multi-modal input and output. The chaotic rule correction unit receives the fusion result, applies nonlinear filtering technology to correct the reasoning bias, and outputs the correction data. The credibility weight allocation unit receives the correction data, allocates event weights through a fuzzy logic controller, and outputs a four-dimensional early warning vector including event type, probability, urgency, and disposal suggestion. The hierarchical early warning signal is fed back to the quantumized behavior acquisition module to adjust the sampling frequency of the quantum sensor in real time, forming a closed-loop optimization system.
[0067] In the overall process, each module is closely connected through data flow: the quantumized behavior acquisition provides original feature input, the neural radiation field simulation enhances data diversity, the meta-learning feature encoding optimizes feature representation, the game reasoning engine simulates group behavior, the chaotic rule evolution dynamically adjusts the decision rule, and the cross-modal consensus fuses multi-modal information to output an early warning signal. The feedback mechanism realizes adaptive optimization of the system, improving the accuracy of abnormal behavior recognition.
[0068] Specifically, the intelligent question and answer system based on multi-modal feature fusion of the application comprises:
[0069] The quantum tunneling effect measurement unit processes the human muscle microtremor through the Josephson junction array sensor and outputs microtremor frequency data.
[0070] The gait optical flow field phase difference detection unit receives the microtremor frequency data output by the quantum tunneling effect measurement unit, processes the microtremor frequency data through the single-photon avalanche diode matrix, and outputs motion vector data.
[0071] The superconducting quantum interference device unit receives the motion vector data output by the gait optical flow field phase difference detection unit, processes the motion vector data through the superconducting quantum interference device configured with a gradient meter structure, and outputs electromagnetic disturbance data.
[0072] The manifold learning dimension reduction unit receives the microtremor frequency data output by the quantum tunneling effect measurement unit, the motion vector data output by the gait optical flow field phase difference detection unit, and the electromagnetic disturbance data output by the superconducting quantum interference device unit, processes the microtremor frequency data, the motion vector data, and the electromagnetic disturbance data through a manifold learning algorithm, generates a behavior feature tensor, and outputs the behavior feature tensor to the neural radiation field simulation module.
[0073] The quantum tunneling effect measurement unit in the quantumized behavior acquisition module captures quantum-level features of human muscle microtremor through the Josephson junction array sensor. Based on the superconducting quantum interference principle, when the human muscle produces micron-level tremor, the magnetic flux change in the superconducting ring will cause the Josephson junction to generate quantum tunneling current. The current oscillation frequency is mapped to the muscle tremor frequency, and the microtremor frequency data is output. This data reflects the microscopic quantum characteristics of neuromuscular activity, avoiding the privacy restrictions of existing biometric feature acquisition.
[0074] The gait optical flow field phase difference detection unit receives the microtremor frequency data output by the quantum tunneling effect measurement unit, and analyzes the motion vector through the single-photon avalanche diode matrix. The single-photon avalanche diode matrix emits near-infrared laser pulses, and modulates the pulse phase according to the microtremor frequency data. After the laser is reflected by the human joint, a photon flow field is formed, and a three-dimensional optical flow field of gait motion is constructed by analyzing the arrival time difference of the reflected photons, and limb motion vector data is output. The motion vector data includes spatial displacement direction and speed information, realizing the conversion from quantum characteristics to macroscopic motion characteristics.
[0075] The quantum tunneling effect measurement unit and the gait optical flow field phase difference detection unit form a hierarchical feature extraction chain: the muscle microtremor frequency as the quantum representation of limb movement drives the phase modulation of the photon pulse; the reflected optical flow field phase difference inverts the accurate movement trajectory. The two units are tightly coupled through data flow, and the microtremor frequency collected by the Josephson junction array is the input basis for the analysis of the optical flow field, and the motion vector data is the extended expression of the microtremor in the macro movement dimension. This "quantum and classical" dual-level feature fusion mechanism provides multi-scale input for subsequent electromagnetic disturbance inversion.
[0076] The superconducting quantum interference device unit receives the motion vector data output by the gait optical flow field phase difference detection unit, and processes the motion data through the superconducting quantum interference device of the gradient meter structure. The large-scale movement of the human body causes biological electromagnetic field disturbance, and the gradient meter structure detects the magnetic field gradient change corresponding to the motion vector; the change of the magnetic flux in the superconducting ring causes the critical current of the Josephson junction to deviate, and the disturbance mode is analyzed by inverting the electromagnetic field partial differential equation, and the electromagnetic disturbance data is output. The electromagnetic disturbance data reflects the electromagnetic field interference characteristics when the group moves, and forms a space-time complement with the motion vector.
[0077] The manifold learning dimension reduction unit receives the microtremor frequency data, the motion vector data and the electromagnetic disturbance data, and fuses and processes them through the Riemann manifold learning algorithm. The three types of heterogeneous data are mapped to the Riemann manifold space, and the geodesic distance is used to calculate the similarity of the data points; the original data topological structure is maintained in the low-dimensional manifold through the isometric embedding algorithm, and a non-Euclidean space-time behavior feature tensor integrating quantum features, motion features and electromagnetic features is generated, and output to the neural radiation field simulation module. This process realizes the space-time alignment and dimension compression of multi-source heterogeneous features, and provides structured input for virtual scene reconstruction.
[0078] Specifically, the intelligent question and answer system based on multi-modal feature fusion of the present application, the neural radiation field simulation module comprises:
[0079] A dynamic scene reconstruction unit receives the behavior feature tensor output by the quantized behavior acquisition module, processes the behavior feature tensor through differentiable rendering, and outputs initial rendering parameters;
[0080] A synthetic data injection unit receives the initial rendering parameters output by the dynamic scene reconstruction unit, processes the initial rendering parameters through a generative adversarial network architecture, and outputs adversarial sample data;
[0081] A space-time mapping function unit receives the adversarial sample data output by the synthetic data injection unit, processes the adversarial sample data through a hyperbolic geometry embedding model, and outputs an enhanced space-time sequence to the meta-learning feature encoding module.
[0082] The input of the neural radiance field simulation module is the behavior feature tensor output by the quantized behavior acquisition module. The dynamic scene reconstruction unit receives the behavior feature tensor, and decomposes the tensor into radiance field parameters through a differentiable rendering technology. The differentiable rendering is based on a neural radiance field framework, and converts the behavior feature tensor into density field parameters and color field parameters; initial rendering parameters representing the geometric structure and optical properties of the virtual scene are generated by integrating the volume rendering equation. The non-Euclidean space-time features in the behavior feature tensor are mapped into differentiable rendering bases to support subsequent scene enhancement processing.
[0083] The synthetic data injection unit receives the initial rendering parameters output by the dynamic scene reconstruction unit, and processes the parameters through a generative adversarial network architecture. The generative adversarial network includes a generator and a discriminator component, the generator synthesizes virtual scene data with occlusion and crowd interference based on the initial rendering parameters, and the discriminator evaluates the authenticity of the synthetic data; the output of the generator is optimized through adversarial training to generate adversarial sample data. The adversarial sample data introduces uncertainty in the real scene, such as local occlusion or high-density crowd interference, to make up for the shortcomings of the original behavior features and improve the robustness of the model.
[0084] The spatio-temporal mapping function unit receives the adversarial sample data output by the synthetic data injection unit, and processes the data through a hyperbolic geometry embedding model. The hyperbolic geometry model uses a Poincare disk to represent the entity coordinates in the adversarial sample data; the similarity between the entity and the virtual coordinates is calculated using hyperbolic distance, and coordinate alignment is achieved through geodesic projection, and an enhanced spatio-temporal sequence is output to the meta-learning feature encoding module. This process unifies the spatio-temporal scale, realizes the diversity of the enhanced spatio-temporal sequence, and eliminates coordinate bias, providing a structured input for feature encoding.
[0085] The behavior feature tensor is converted into initial rendering parameters through differentiable rendering, providing a basis for scene reconstruction; the adversarial sample data injection enhances the complexity of the data; the hyperbolic geometry embedding optimizes the spatio-temporal consistency, and finally outputs an enhanced sequence suitable for meta-learning. The data flow from parameter generation to sample enhancement to coordinate alignment gradually improves the realism and generalization ability of the virtual scene, supporting subsequent feature extraction tasks.
[0086] Specifically, the intelligent question and answer system based on multi-modal feature fusion of the present application comprises:
[0087] The variational autoencoder unit receives the enhanced spatio-temporal sequence output by the neural radiance field simulation module, processes the enhanced spatio-temporal sequence through a capsule network structure, and outputs hierarchical features;
[0088] The memory enhancement network unit receives the hierarchical features output by the variational autoencoder unit, processes the hierarchical features through a neural Turing machine architecture, and outputs memory enhancement features;
[0089] The meta-learning strategy unit receives the memory-enhanced features output by the memory-enhanced network unit, processes the memory-enhanced features through a model-agnostic meta-learning algorithm, and outputs a 128-dimensional hyper-vector to the game reasoning engine module.
[0090] The meta-learning feature encoding module receives the enhanced spatio-temporal sequence output by the neural radiation field simulation module, and processes the sequence through a variational autoencoder unit. The variational autoencoder unit adopts a capsule network structure to analyze the spatio-temporal features. The capsule network establishes a hierarchical relationship of features through a dynamic routing mechanism: the primary capsules identify local spatio-temporal patterns, and the high-level capsules aggregate local features to form overall behavior representations. The variational encoder learns the distribution of the latent space, and outputs hierarchical features that preserve the spatio-temporal structure of behaviors by constraining the distribution of latent variables through KL divergence. The hierarchical features include the spatial combination relationship of behavior elements, providing a structured input for memory enhancement.
[0091] The memory-enhanced network unit receives the hierarchical features output by the variational autoencoder unit, and processes the features through a neural Turing machine architecture. The neural Turing machine includes a read-write controller and an external memory matrix. The read-write controller analyzes the hierarchical features to generate addressing signals. The content addressing mechanism matches historical scenario parameters, and the position addressing mechanism locates relevant memory blocks. The scenario transfer knowledge in the memory matrix is read and fused with the current features to output memory-enhanced features that integrate historical experience. This process realizes cross-scenario knowledge transfer and improves the consistency of feature representation.
[0092] The meta-learning strategy unit receives the memory-enhanced features output by the memory-enhanced network unit, and optimizes the feature extractor through a model-agnostic meta-learning algorithm. The algorithm constructs a multi-task support set, calculates the gradient of the feature extractor parameters on the support set, and updates the network weights based on the gradient information to enable the feature extractor to quickly adapt to new scenario distributions. The memory-enhanced features are processed to generate a 128-dimensional hyper-vector with strong generalization, which is output to the game reasoning engine module. The 128-dimensional hyper-vector compresses the essential features of behaviors and provides a robust input for group interaction modeling.
[0093] The hierarchical features extracted by the capsule network serve as the input basis for memory enhancement. The historical knowledge fused by the neural Turing machine strengthens the semantics of the features. The hyper-vector optimized by the meta-learning algorithm has the ability of scenario transfer. The data flow is transformed from the spatio-temporal sequence to the hierarchical features through memory enhancement, forming a progressive feature optimization path, and realizing the discriminability and stability of behavior features in complex scenarios.
[0094] Specifically, the intelligent question-answering system based on multi-modal feature fusion of the present application includes:
[0095] The agent strategy network unit receives the 128-dimensional hyper-vector output by the meta-learning feature encoding module, processes the 128-dimensional hyper-vector through a graph attention mechanism, and outputs interaction strategy data.
[0096] The Nash equilibrium solver unit receives the interaction strategy data output by the agent policy network unit, processes the interaction strategy data through a deep virtual game framework, and outputs equilibrium strategy data.
[0097] The Monte Carlo tree search unit receives the equilibrium strategy data output by the Nash equilibrium solver unit, processes the equilibrium strategy data through a time-difference learning, and outputs a group behavior probability matrix to the chaotic rule evolution module.
[0098] The game reasoning engine module receives the 128-dimensional hyper vector output by the meta-learning feature encoding module, and processes the features through the agent policy network unit. The agent policy network unit models the group interaction relationship using a graph attention mechanism: maps the 128-dimensional hyper vector to agent node features, constructs a group relationship graph, calculates the interaction weight between nodes through a multi-head attention mechanism, generates an adjacency matrix representing the individual influence distribution, and outputs interaction strategy data including cooperation and competition strategies. The interaction strategy data reflects the micro decision-making motivation within the group and provides input basis for equilibrium solving.
[0099] The Nash equilibrium solver unit receives the interaction strategy data output by the agent policy network unit, and optimizes the strategy through a deep virtual game framework. The framework builds a multi-agent reinforcement learning environment and updates the agent policy network parameters using a policy gradient algorithm; simulates the strategy confrontation process in the virtual game round, dynamically converges to the Nash equilibrium point through the best response, and outputs the equilibrium strategy data in the stable state. The equilibrium strategy data represents the steady-state solution of group interaction and eliminates the prediction bias caused by strategy conflict.
[0100] The Monte Carlo tree search unit receives the equilibrium strategy data output by the Nash equilibrium solver unit, and predicts behavior evolution through time-difference learning. Based on the equilibrium strategy data, the search tree root node is initialized, and the decision tree branches are expanded using Monte Carlo simulation; the node state value is evaluated using time-difference learning, and the action value function is updated combining the Bellman equation; the path selection strategy is optimized through back propagation, and the group behavior probability matrix representing the spatio-temporal behavior distribution probability is output to the chaotic rule evolution module. The matrix quantifies the trigger probability of group abnormal behavior in different scenarios, supporting dynamic decision rule generation.
[0101] The interaction strategy extracted by the graph attention mechanism is the input of the virtual game, and the equilibrium strategy output by the virtual game is the initialization basis of the Monte Carlo tree search. The data goes through three levels of progression: individual interaction modeling, group equilibrium solving, and behavior evolution prediction, forming a complete chain from micro decision-making to macro behavior deduction. The output group behavior probability matrix integrates strategy stability and evolution dynamic characteristics, providing a probabilistic input basis for subsequent chaotic analysis.
[0102] Specifically, the intelligent question and answer system based on multi-modal feature fusion of the application comprises:
[0103] The Lyapunov index analysis unit receives the group behavior probability matrix output by the game reasoning engine module, processes the group behavior probability matrix through a recurrent neural network, and outputs a stability index;
[0104] The risk pattern reconstruction unit receives the stability index output by the Lyapunov index analysis unit, processes the stability index through multi-fractal detrended fluctuation analysis, and outputs a reconstructed risk pattern;
[0105] The early warning threshold function optimization unit receives the reconstructed risk pattern output by the risk pattern reconstruction unit, processes the reconstructed risk pattern through adaptive decision boundary setting, and outputs a dynamic decision rule to the cross-modal consensus module.
[0106] The chaotic rule evolution module receives the group behavior probability matrix output by the game reasoning engine module, and processes the matrix through the Lyapunov index analysis unit. The Lyapunov index analysis unit reconstructs the phase space of the decision system using a recurrent neural network, and inputs the group behavior probability matrix into the recurrent neural network time module. The evolution process of the decision trajectory is simulated through hidden state transmission, the exponential average of the divergence rate of adjacent trajectories is calculated, and a stability index representing the stability of the system is output. The stability index quantifies the sensitivity of the decision process to the initial conditions, and provides dynamic feature input for risk pattern reconstruction.
[0107] The risk pattern reconstruction unit receives the stability index output by the Lyapunov index analysis unit, and processes the index data through multi-fractal detrended fluctuation analysis. The multi-fractal detrended fluctuation analysis performs multi-scale segmentation on the stability index sequence, calculates the scale index of the fluctuation function under different scales, extracts the singular intensity distribution through Hurst index spectrum analysis, identifies the critical features of the system at the edge of chaos, and outputs a reconstructed risk pattern representing the risk pattern in multiple dimensions. The reconstructed risk pattern reveals the complex dynamic characteristics of the system, and supports decision boundary optimization.
[0108] The early warning threshold function optimization unit receives the reconstructed risk pattern output by the risk pattern reconstruction unit, and processes the pattern features through adaptive decision boundary setting. Based on the fractal characteristics of the reconstructed risk pattern, a strange attractor phase space is constructed, the fractal dimension of the attractor boundary is calculated, the topological structure of the decision boundary is dynamically adjusted according to the fractal dimension, a dynamic decision rule that adapts to the changes of the risk pattern is generated through a parameterized threshold function, and the dynamic decision rule is output to the cross-modal consensus module. The dynamic decision rule realizes the real-time evolution of the early warning threshold, and improves the abnormal detection sensitivity.
[0109] The group behavior probability matrix is converted into a stability index through Lyapunov analysis; the stability index is reconstructed into a risk pattern through multifractal analysis; and the risk pattern drives the adaptive adjustment of the threshold function. Data flow is reconstructed from stability quantification to risk dimension, and then to boundary optimization, forming a progressive chain of "dynamic feature extraction, risk dimension mapping, and decision rule generation", and realizing the dynamic evolution of the early warning mechanism with system complexity.
[0110] Specifically, the intelligent question and answer system based on multi-modal feature fusion comprises a cross-modal consensus module:
[0111] The cross-validation mechanism unit receives the dynamic decision rules output by the chaotic rule evolution module, processes the behavior feature tensor output by the quantumized behavior acquisition module, the enhanced spatiotemporal sequence output by the neural radiation field simulation module, and the group behavior probability matrix output by the game reasoning engine module through evidence theory fusion, and outputs the fusion result;
[0112] The chaotic rule correction unit receives the fusion result output by the cross-validation mechanism unit, processes the fusion result through nonlinear filtering, and outputs the corrected data;
[0113] The credibility weight allocation unit receives the corrected data output by the chaotic rule correction unit, processes the corrected data through a fuzzy logic controller, and outputs a four-dimensional early warning vector;
[0114] The four-dimensional early warning vector includes event type, probability, urgency, and disposal suggestion.
[0115] The cross-modal consensus module receives the dynamic decision rules output by the chaotic rule evolution module, and fuses multiple sources of input through the cross-validation mechanism unit. The cross-validation mechanism unit processes the dynamic decision rules, the behavior feature tensor output by the quantumized behavior acquisition module, the enhanced spatiotemporal sequence output by the neural radiation field simulation module, and the group behavior probability matrix output by the game reasoning engine module using Dempster-Shafer evidence theory. The basic probability assignment function is constructed to quantify the confidence of each modality, and the evidence conflict is eliminated through orthogonal rules to calculate the joint confidence interval and output the fusion result. The fusion result integrates multi-modal decision basis and eliminates single data source bias.
[0116] The chaotic rule correction unit receives the fusion result output by the cross-validation mechanism unit, and processes the data through nonlinear filtering. An extended Kalman filter algorithm is used to model the state equation of the dynamic system, and the fusion result is used as the observation input; the residual error between the predicted state and the observed state is fed back and corrected through the Kalman gain matrix, eliminating the trajectory deviation caused by the nonlinearity of the system, and outputting the corrected data after deviation correction. The corrected data maintains the stability of the decision trajectory and provides accurate input for weight allocation.
[0117] The credibility weight distribution unit receives the correction data output by the chaotic rule correction unit, and processes the data through a fuzzy logic controller. The input variable is the event confidence and the conflict factor of the correction data, and the output variable is the weight coefficient; the Gaussian membership function is defined to quantify the fuzzy rule, and the four-dimensional weight distribution is calculated by the barycenter method, and the four-dimensional early warning vector including the event type, occurrence probability, emergency level and disposal measures is output. The four-dimensional early warning vector realizes risk grading quantization and supports closed-loop system feedback regulation.
[0118] The data flow in the module is closely connected: the evidence theory fusion output is used as the input of the nonlinear filtering, and the filtering correction result drives the fuzzy logic distribution. The multi-modal input is processed through three levels of "fusion, correction and weighting", forming a complete link from original decision to refined early warning. The four-dimensional early warning vector defines the abnormal scene through the event type, quantifies the risk intensity through the probability value, divides the response level through the emergency degree, and guides the intervention measures through the disposal suggestion, providing a structured control signal for system feedback.
[0119] Specifically, the intelligent question and answer system based on multi-modal feature fusion of the application further comprises:
[0120] When the synthetic data injection unit of the neural radiation field simulation module outputs the adversarial sample data, the memory enhancement network unit of the meta-learning feature coding module processes the historical scene parameters through the neural Turing machine architecture, and outputs the enhanced features to the meta-learning strategy unit of the meta-learning feature coding module.
[0121] The trigger mechanism is activated when the neural radiation field simulation module generates adversarial samples, and enhances the anti-interference ability of the feature coding module. The synthetic data injection unit of the neural radiation field simulation module outputs adversarial sample data, which includes virtual scene features simulating real occlusion and crowd interference. The adversarial sample data acts as a trigger signal to activate the memory enhancement network unit of the meta-learning feature coding module.
[0122] The memory enhancement network unit processes historical scene parameters through the neural Turing machine architecture. The read-write controller of the neural Turing machine analyzes the adversarial sample data to generate a content addressing signal, retrieves similar historical scene parameters in the external memory matrix, and outputs enhanced features that fuse historical adversarial sample knowledge based on cosine similarity matching of stored radiation field parameters and space-time mapping rules. The enhanced features are transmitted to the meta-learning strategy unit of the meta-learning feature coding module, and are spliced with real-time features to form anti-interference input.
[0123] After receiving the enhanced features, the meta-learning strategy unit implements a model-agnostic meta-learning algorithm to update the feature extractor. Adversarial samples are injected in the support set tasks, and the network weights are optimized through gradient descent. The feature extractor learns the invariance of adversarial perturbations, enhancing the robustness of the capsule network to occlusion scenarios. This process forms a closed loop of "adversarial sample generation, historical knowledge retrieval, feature enhancement, network optimization", enabling the system to maintain feature discrimination under virtual scene interference.
[0124] The trigger mechanism logic is clear: the adversarial sample data is the activation signal, the neural Turing machine implements knowledge retrieval, and the enhanced features deliver anti-interference information. The data flow extends from the simulation module to the feature encoding module, fusing historical experience to cope with virtual scene disturbances and ensuring the accuracy of subsequent group behavior analysis.
[0125] Specifically, the intelligent question and answer system based on multi-modal feature fusion of the application further comprises:
[0126] When the Monte Carlo tree search unit of the game reasoning engine module outputs the group behavior probability matrix, the risk pattern reconstruction unit of the chaotic rule evolution module processes the fractal dimension through multi-fractal detrended fluctuation analysis, and outputs the optimized decision rule to the early warning threshold function optimization unit of the chaotic rule evolution module.
[0127] The four-dimensional early warning vector output by the cross-modal consensus module includes a credibility weight, which is output to the quantized behavior acquisition module.
[0128] When the Monte Carlo tree search unit of the game reasoning engine module outputs the group behavior probability matrix, the event serves as a system trigger signal to activate the risk pattern reconstruction unit of the chaotic rule evolution module. The group behavior probability matrix includes a spatiotemporal behavior distribution probability, and its abnormal pattern may indicate a potential risk scenario, triggering the risk pattern reconstruction unit to start a dynamic optimization mechanism.
[0129] The risk pattern reconstruction unit processes the fractal dimension through multi-fractal detrended fluctuation analysis. The multi-fractal detrended fluctuation analysis performs multi-scale segmentation on the group behavior probability matrix, calculates the fluctuation function within different time windows, extracts the singular spectral features of the fractal dimension through the scaling exponent, identifies the non-stationary pattern in the probability distribution, and reconstructs the risk feature space based on the fractal dimension, outputting the optimized decision rule to the early warning threshold function optimization unit of the chaotic rule evolution module. The optimized decision rule enhances the adaptability of the risk threshold, improving the accuracy of anomaly detection.
[0130] After receiving the optimized decision rule, the early warning threshold function optimization unit adjusts the decision boundary parameters based on the fractal boundary characteristics. The optimized decision rule dynamically updates the topology of the threshold function, enabling the early warning mechanism to adapt to group behavior evolution and forming a closed loop decision optimization.
[0131] The four-dimensional early warning vector output by the cross-modal consensus module includes a credibility weight, which is transmitted to the quantum behavior acquisition module. The credibility weight quantizes the confidence level of the early warning signal as a feedback control parameter to adjust the sampling frequency of the quantum sensor; the Josephson junction array sensor adjusts the microseismic tremor capture accuracy according to the weight value, and the single-photon avalanche diode matrix optimizes the motion vector analysis sensitivity, realizing the cooperative optimization of data acquisition and risk response.
[0132] The overall trigger logic is clear: the group behavior probability matrix output activates risk reconstruction, the reconstruction result optimizes the decision rule, the rule update drives the threshold adjustment, and the credibility weight feedback adjusts the acquisition module. The data flow from behavior analysis to rule optimization to acquisition closed loop strengthens the system's adaptive ability.
[0133] In a second aspect, referring to the drawings, the present application provides an intelligent question and answer method based on multi-modal feature fusion, applied to the intelligent question and answer system based on multi-modal feature fusion, comprising:
[0134] Step 1: Deploy a quantum sensor array to capture human micro-behavior characteristics and generate a non-Euclidean space-time behavior feature tensor. The quantum sensor array includes a Josephson junction array sensor, a single-photon avalanche diode matrix, and a superconducting quantum interference device;
[0135] Step 2: Receive the behavior feature tensor output by step 1, convert the behavior feature tensor into differentiable rendering parameters, generate initial rendering parameters, introduce a generative adversarial network architecture to create adversarial sample data for the initial rendering parameters to create a blocked crowded scene, design a hyperbolic geometry embedding model to align entity virtual coordinates, process the adversarial sample data and output an enhanced space-time sequence;
[0136] Step 3: Receive the enhanced space-time sequence output by step 2, extract hierarchical features of the behavior characteristics through a variational autoencoder, integrate a neural Turing machine architecture to store scene migration knowledge, fuse the hierarchical features to generate memory-enhanced features, implement a model-agnostic meta-learning algorithm to optimize the feature extractor parameters, process the memory-enhanced features and output a 128-dimensional hyper-vector;
[0137] Step 4: Receive the hyper-vector output by step 3, simulate group interaction strategies, and generate a group behavior probability matrix;
[0138] Step 5: Receive the group behavior probability matrix output by step 4, analyze the stability of the decision trajectory, and output a dynamic decision rule;
[0139] Step 6: Receive the dynamic decision rule output by step 5, fuse the behavior feature tensor of step 1, the enhanced space-time sequence output by step 2, and the group behavior probability matrix output by step 4, output a hierarchical early warning signal, and the hierarchical early warning signal is fed back to step 1 to adjust the sampling frequency.
[0140] Step 1: Deploy Josephson junction array sensors to capture quantum tunneling currents generated by human muscle microtremor, map current oscillation frequency to microtremor frequency data; single-photon avalanche diode matrix emits near-infrared laser pulses, modulates pulse phase according to microtremor frequency, resolves gait motion vectors through reflected photon time difference; superconducting quantum interference device detects electromagnetic field gradient changes induced by motion using a gradiometer structure, inverts electromagnetic disturbance patterns; manifold learning algorithm fuses microtremor frequency data, motion vector data, and electromagnetic disturbance data, performs nonlinear dimensionality reduction in Riemannian manifold space, generating non-Euclidean spatiotemporal behavior feature tensors. This step realizes quantum-level collection of microscopic behavior features under privacy compliance.
[0141] Step 2: Receive behavior feature tensors output from Step 1, decompose tensors into radiation field density parameters and color parameters using differentiable rendering technology, generate initial rendering parameters; generate adversarial network architecture to inject occluders and crowd disturbance factors into initial rendering parameters, synthesize adversarial sample data; hyperbolic geometry embedding model maps adversarial sample data to Poincare disk space, calculates geodesic distance to align real entities and virtual coordinates, outputs enhanced spatiotemporal sequences. This step addresses the problem of missing biometric features and improves data diversity.
[0142] Step 3: Receive enhanced spatiotemporal sequences output from Step 2, capsule network structure extracts hierarchical relationships of spatiotemporal features through dynamic routing mechanism; neural Turing machine architecture reads historical scene parameters from external memory matrix, fuses with hierarchical features to generate memory-enhanced features; model-agnostic meta-learning algorithm updates feature extractor network weights on support set tasks, outputs 128-dimensional hyper vectors. This step strengthens the scene generalization ability of features.
[0143] Step 4: Receive 128-dimensional hyper vectors output from Step 3, graph attention mechanism constructs an intelligent agent interaction graph, calculates node influence weights to output interaction strategy data; deep virtual game framework converges to a Nash equilibrium point through policy gradient descent, outputs equilibrium strategy data; temporal difference learning combined with Monte Carlo tree search predicts behavior evolution paths, generates group behavior probability matrix. This step realizes quantitative prediction of group dynamics.
[0144] Step 5: Receive group behavior probability matrix output from Step 4, recursive neural network calculates Lyapunov indicator to quantify decision trajectory stability; multi-fractal detrended fluctuation analysis extracts singular spectral features of stability index, reconstructs risk patterns; based on fractal dimension optimization, determine boundary parameters, output dynamic decision rules. This step realizes adaptive adjustment of risk thresholds.
[0145] Step 6: receive the dynamic decision rule output in step 5, Dempster-Shafer evidence theory fusion behavior feature tensor, confidence of enhanced spatiotemporal sequence and group behavior probability matrix; extended Kalman filter corrects the system deviation of the fusion result; fuzzy logic controller distributes the weight coefficients of event type, probability, urgency and disposal suggestion, and outputs graded early warning signal; the signal feedback adjusts the sampling frequency of the quantum sensor in step 1, forming a closed loop of self-optimizing collection accuracy.
[0146] The present application solves the problem of insufficient input dimensions under the restriction of privacy protection through a multi-level technical architecture, and the specific technical path is as follows:
[0147] The quantum behavior acquisition module deploys a Josephson junction array sensor to capture muscle micro tremor frequency, and converts micron-level muscle vibration into frequency signals based on quantum tunneling effect; a single-photon avalanche diode matrix analyzes limb motion vectors, and a superconducting quantum interference device inverts electromagnetic disturbance patterns. Three types of non-biological features are fused to generate a behavior feature tensor through manifold learning algorithm, and a spatiotemporal representation is constructed in non-Euclidean space. This process avoids biological feature acquisition while expanding the input dimension through quantum-level microscopic behavior data, such as muscle micro tremor frequency representing neural activity intensity, motion vector reflecting limb displacement trend, and electromagnetic disturbance mapping group movement pattern.
[0148] The neural radiation field simulation module decomposes the behavior feature tensor into differentiable rendering parameters, generates an adversarial network architecture to inject disturbance in crowded scenes, and synthesizes adversarial sample data; a hyperbolic geometry embedding model aligns real and virtual coordinates in a Poincare disk space, and outputs an enhanced spatiotemporal sequence. This step simulates the complexity of real scenes through synthetic data, making up for the single data defect caused by the lack of biological features, so that the model can still recognize abnormal postures in crowded and obstructed environments.
[0149] The meta-learning feature encoding module extracts hierarchical spatiotemporal features through a capsule network, and a neural Turing machine generates memory-enhanced features by fusing historical scene knowledge; a model-agnostic meta-learning algorithm optimizes the feature extractor on the support set task, and outputs a 128-dimensional hyper vector. This design improves the model's adaptability to unknown scenes, and when the input data lacks biological identification features, the generalized feature representation can still maintain the accuracy of behavior recognition.
[0150] The cross-modal consensus module adopts evidence theory to fuse quantum behavior feature tensors, enhanced spatiotemporal sequences and group behavior probability matrices, eliminating single-modal bias; a fuzzy logic controller generates a four-dimensional early warning vector, and the credibility weight thereof feeds back to adjust the sampling frequency of the quantum sensor. The closed-loop mechanism realizes dynamic coordination of data acquisition and risk response: when the early warning credibility is high, the sampling accuracy of the Josephson junction array is improved, and when the credibility is low, the power consumption is reduced. Under the privacy constraint, the system improves the reliability of decision-making through multi-level technology linkage, for example, in a commercial complex scene, abnormal gathering behavior is identified by fusing muscle micro-tremor and motion vector, and the false positive rate is reduced.
[0151] In the real-time passenger flow analysis scene of a commercial complex, the existing system cannot collect biological features such as faces due to privacy regulations, and only relies on human posture and movement trajectory data, resulting in insufficient input dimensions and rising abnormal behavior false positive rates. The present application solves this problem through a multi-level technology architecture, and the specific implementation process is as follows:
[0152] Deploy a quantum sensor array at the entrance of the mall:
[0153] The Josephson junction array sensor captures the quantum tunneling current generated by the muscle micro-tremor of the tourists, converts the gastrocnemius muscle micro-vibration into micro-tremor frequency data of 50-200Hz, and represents the intensity of neural activity;
[0154] The single-photon avalanche diode matrix emits 905nm wavelength laser pulses, modulates the pulse phase according to the micro-tremor frequency, analyzes the lower limb motion vector through the time difference of reflected photons, and outputs three-dimensional motion data including speed and direction;
[0155] The gradient meter structure of the superconducting quantum interference device detects the change in the bioelectromagnetic field gradient caused by the movement of the tourists, and inversely outputs the electromagnetic disturbance pattern.
[0156] The manifold learning dimension reduction unit uses the Isomap algorithm to project the micro-tremor frequency data, motion vector data and electromagnetic disturbance data into the Riemannian manifold space, generating a 128-dimensional non-Euclidean spatiotemporal behavior feature tensor. This process expands the input dimension through three-level fusion of muscle tremor (quantum feature), limb movement (classical feature) and electromagnetic field (group feature), and avoids the restriction of biological feature collection.
[0157] The dynamic scene reconstruction unit decomposes the behavior feature tensor into a density field and a color field parameter to generate initial rendering parameters of the three-dimensional structure of the virtual mall; the synthetic data injection unit uses a conditional generative adversarial network (CGAN) to inject column obstructions and crowd gathering disturbance factors into the initial rendering parameters, and outputs virtual scene data including adversarial noise; the space-time mapping function unit constructs a Poincare disk model, maps the entity coordinates to hyperbolic space, aligns the virtual and real coordinates through geodesic distance, and outputs an enhanced space-time sequence. This step synthesizes crowded scene data in the column obstruction area of the mall, making up for the insufficient model generalization caused by real data missing.
[0158] The capsule network of the variational autoencoder unit identifies the hierarchical relationship of behaviors: the primary capsule extracts local motion patterns (such as waving hands), and the senior capsule integrates into complete behavior representation (such as arguing), outputting 256-dimensional hierarchical features; the neural Turing machine of the memory enhancement network unit retrieves the fast traversal parameters of the delivery man in the historical scene, and fuses with the current features to output memory-enhanced features; the meta-learning strategy unit updates the feature extractor weights on the support set tasks (different floor layouts), outputting 128-dimensional hyper vectors. This design makes the model adapt to different mall structures, solving the problem of recognition rate fluctuation caused by scene changes.
[0159] The agent strategy network unit maps the 128-dimensional hyper vector to the agent node, and the graph attention mechanism calculates the interaction weight of tourists (such as following and avoiding), outputting interaction strategy data; the Nash equilibrium solver unit simulates the avoidance game through policy gradient descent, and converges to the equilibrium strategy; the Monte Carlo tree search unit uses time difference learning to predict the crowd gathering path, and outputs the crowd gathering probability matrix in the next 5 minutes in the A area rest area (probability value 0.75). This process quantifies the evolution trend of abnormal behavior, supporting precise evacuation decisions.
[0160] When the Monte Carlo tree search unit outputs the gathering probability matrix, the risk pattern reconstruction unit of the chaotic rule evolution module is triggered: the multifractal detrended fluctuation analysis extracts the singular spectrum of the probability matrix, and reconstructs the risk pattern of the fractal dimension; the early warning threshold function optimization unit dynamically adjusts the decision threshold according to the fractal boundary (such as triggering an early warning when the probability is greater than 0.7). The cross-modal consensus module fuses the behavior feature tensor, the enhanced space-time sequence, and the probability matrix, calculates the joint confidence (0.92) using evidence theory, and outputs a four-dimensional early warning vector: event type (gathering), probability (0.75), urgency (high), and disposal suggestion (dispersion in area A). The credibility weight in the four-dimensional early warning vector is fed back to the quantum sensor: when the probability is greater than 0.7, the sampling frequency of the Josephson junction array is increased to enhance the data capture accuracy.
[0161] The quantum sensor array of the application includes a Josephson junction array sensor, a single-photon avalanche diode matrix and a superconducting quantum interference device, a muscle microtremor frequency capturing unit measures the unit through a quantum tunneling effect, a gait optical flow field phase difference detection unit analyzes a motion vector, a superconducting quantum interference device unit inverts an electromagnetic disturbance mode, a manifold learning dimension reduction algorithm fuses multi-source data to generate a non-Euclidean space-time behavior feature tensor, expands the input dimension under the condition of privacy compliance, and avoids biological feature collection restrictions.
[0162] The differentiable rendering technology decomposes the behavior feature tensor into a radiance field density parameter and a color parameter, generates initial rendering parameters of a virtual scene, generates an adversarial network architecture to inject an occluder and a crowd disturbance factor to synthesize adversarial sample data, and a hyperbolic geometry embedding model maps and aligns the real and virtual coordinates through a Poincare disk space, outputs an enhanced space-time sequence, and improves data diversity to make up for feature loss.
[0163] The capsule network structure establishes a hierarchical relationship between primary capsules and high-level capsules through a dynamic routing mechanism, extracts local and global features of the space-time sequence, the neural Turing machine architecture retrieves historical scene parameters in an external memory matrix through a read-write controller, fuses real-time features to output memory-enhanced features, and the model-agnostic meta-learning algorithm updates the feature extractor network weights on the support set task, and outputs a 128-dimensional hyper vector with scene generalization capability.
[0164] The graph attention mechanism constructs an agent relationship graph, calculates the attention weight between nodes to generate interaction strategy data, the deep virtual game framework uses a policy gradient algorithm to simulate multi-agent strategy games, converges to a Nash equilibrium point to output an equilibrium strategy, and the time difference learning combined with the Monte Carlo tree search evaluates the long-term benefits of the decision path and predicts the evolution probability of group behavior.
[0165] The recurrent neural network reconstructs the phase space of the decision system, calculates the trajectory divergence rate to output the Lyapunov stability index, extracts the scaling exponent at different scales through multi-fractal detrended fluctuation analysis, reconstructs the risk feature dimension, adaptively determines the boundary to set the warning threshold function based on the fractal dimension, and realizes dynamic decision rule generation.
[0166] The Dempster-Shafer evidence theory constructs a basic probability assignment function, fuses the confidence of multi-modal input data, eliminates evidence conflict through an orthogonal sum rule to output a fusion result, extends a Kalman filter to model a nonlinear system state equation, corrects the observation trajectory deviation, and a fuzzy logic controller defines a Gaussian membership function to quantify the rules and outputs a four-dimensional warning vector weight coefficient.
[0167] The neural Turing machine architecture matches the historical scene parameters through the content addressing mechanism, outputs the enhanced feature promotion model anti-interference ability, multi-fractal de-trending fluctuation analysis reconstructs the fractal dimension optimization decision rule, the credibility weight feedback adjusts the quantum sensor sampling frequency, and forms the closed-loop optimization mechanism of data acquisition and risk response.
Claims
1. An intelligent question-answering system based on multi-modal feature fusion, characterized in that, The application relates to a quantum behavior acquisition module, a neural radiation field simulation module, a meta-learning feature coding module, a game reasoning engine module, a chaotic rule evolution module and a cross-modal consensus module. The quantum behavior acquisition module is configured to capture human microscopic behavior characteristics by deploying a quantum sensor array, generate a non-Euclidean space-time behavior characteristic tensor, and the quantum sensor array comprises a Josephson junction array sensor, a single-photon avalanche diode matrix and a superconducting quantum interference device. The neural radiation field simulation module is configured to receive the behavior characteristic tensor output by the quantum behavior acquisition module, convert the behavior characteristic tensor into differentiable rendering parameters, generate initial rendering parameters, introduce a generative adversarial network architecture to create adversarial sample data of the initial rendering parameters to create a blocked crowded scene, design a hyperbolic geometry embedding model to align entity virtual coordinates, process the adversarial sample data and output an enhanced space-time sequence. The meta-learning feature coding module is configured to receive the enhanced space-time sequence output by the neural radiation field simulation module, extract hierarchical features of the behavior characteristics by a variational autoencoder, integrate a neural Turing machine architecture to store scene migration knowledge, fuse the hierarchical features to generate memory-enhanced features, implement a model-agnostic meta-learning algorithm to optimize feature extractor parameters, process the memory-enhanced features and output a 128-dimensional hyper vector. The game reasoning engine module is configured to receive the 128-dimensional hyper vector output by the meta-learning feature coding module, simulate group interaction strategies and generate a group behavior probability matrix. The chaotic rule evolution module is configured to receive the group behavior probability matrix output by the game reasoning engine module, analyze decision trajectory stability and output dynamic decision rules. The cross-modal consensus module is configured to receive the dynamic decision rules output by the chaotic rule evolution module, fuse the behavior characteristic tensor of the quantum behavior acquisition module, the enhanced space-time sequence output by the neural radiation field simulation module and the group behavior probability matrix output by the game reasoning engine module, output a hierarchical early warning signal and feed back the hierarchical early warning signal to the quantum behavior acquisition module to adjust a sampling frequency.
2. The intelligent question answering system based on multi-modal feature fusion according to claim 1, characterized in that, The quantum behavior acquisition module comprises a quantum tunneling effect measurement unit, a gait optical flow field phase difference detection unit and a superconducting quantum interference device unit. The quantum tunneling effect measurement unit is configured to process human muscle microtremor by the Josephson junction array sensor and output microtremor frequency data. The gait optical flow field phase difference detection unit is configured to receive the microtremor frequency data output by the quantum tunneling effect measurement unit, process the microtremor frequency data by the single-photon avalanche diode matrix and output motion vector data. The superconducting quantum interference device unit is configured to receive the motion vector data output by the gait optical flow field phase difference detection unit, process the motion vector data by the superconducting quantum interference device configured with a gradient meter structure and output electromagnetic disturbance data. The manifold learning dimension reduction unit is configured to receive the microtremor frequency data output by the quantum tunneling effect measurement unit, the motion vector data output by the gait optical flow field phase difference detection unit and the electromagnetic disturbance data output by the superconducting quantum interference device unit, process the microtremor frequency data, the motion vector data and the electromagnetic disturbance data by a manifold learning algorithm, generate a behavior characteristic tensor and output the behavior characteristic tensor to the neural radiation field simulation module.
3. The intelligent question answering system based on multi-modal feature fusion according to claim 2, characterized in that, The neural radiation field simulation module comprises a dynamic scene reconstruction unit and a neural radiance field simulation unit. The dynamic scene reconstruction unit is configured to receive the behavior characteristic tensor output by the quantum behavior acquisition module, process the behavior characteristic tensor by differentiable rendering and output initial rendering parameters. The neural radiance field simulation unit is configured to receive the initial rendering parameters output by the dynamic scene reconstruction unit, generate a neural radiance field by a neural radiance field simulation algorithm, output a neural radiance field model, introduce a generative adversarial network architecture to create adversarial sample data of the neural radiance field model to create a blocked crowded scene, design a hyperbolic geometry embedding model to align entity virtual coordinates, process the adversarial sample data and output an enhanced space-time sequence. The synthetic data injection unit receives the initial rendering parameters output by the dynamic scene reconstruction unit, processes the initial rendering parameters through a generative adversarial network architecture, and outputs adversarial sample data. The spatiotemporal mapping function unit receives the adversarial sample data output by the synthetic data injection unit, processes the adversarial sample data through a hyperbolic geometry embedding model, and outputs enhanced spatiotemporal sequences to the meta-learning feature encoding module.
4. The intelligent question answering system based on multi-modal feature fusion according to claim 3, characterized in that, The meta-learning feature encoding module includes: The variational autoencoder unit receives the enhanced spatiotemporal sequences output by the neural radiance field simulation module, processes the enhanced spatiotemporal sequences through a capsule network structure, and outputs hierarchical features. The memory enhancement network unit receives the hierarchical features output by the variational autoencoder unit, processes the hierarchical features through a neural Turing machine architecture, and outputs memory-enhanced features. The meta-learning strategy unit receives the memory-enhanced features output by the memory enhancement network unit, processes the memory-enhanced features through a model-agnostic meta-learning algorithm, and outputs a 128-dimensional hyper-vector to the game reasoning engine module.
5. The intelligent question answering system based on multi-modal feature fusion according to claim 4, characterized in that, The game reasoning engine module includes: The agent policy network unit receives the 128-dimensional hyper-vector output by the meta-learning feature encoding module, processes the 128-dimensional hyper-vector through a graph attention mechanism, and outputs interaction policy data. The Nash equilibrium solver unit receives the interaction policy data output by the agent policy network unit, processes the interaction policy data through a deep fictitious play framework, and outputs equilibrium policy data. The Monte Carlo tree search unit receives the equilibrium policy data output by the Nash equilibrium solver unit, processes the equilibrium policy data through a temporal difference learning, and outputs a group behavior probability matrix to the chaotic rule evolution module.
6. The intelligent question answering system based on multi-modal feature fusion according to claim 5, characterized in that, The chaotic rule evolution module includes: The Lyapunov exponent analysis unit receives the group behavior probability matrix output by the game reasoning engine module, processes the group behavior probability matrix through a recurrent neural network, and outputs a stability index. The risk pattern reconstruction unit receives the stability index output by the Lyapunov exponent analysis unit, processes the stability index through a multifractal detrended fluctuation analysis, and outputs a reconstructed risk pattern. The early warning threshold function optimization unit receives the reconstructed risk pattern output by the risk pattern reconstruction unit, processes the reconstructed risk pattern through an adaptive decision boundary setting, and outputs a dynamic decision rule to the cross-modal consensus module.
7. The intelligent question answering system based on multi-modal feature fusion according to claim 6, characterized in that, The cross-modal consensus module includes: The cross-validation mechanism unit receives the dynamic decision rule output by the chaotic rule evolution module, processes the dynamic decision rule, the behavior feature tensor output by the quantumized behavior acquisition module, the enhanced spatiotemporal sequences output by the neural radiance field simulation module, and the group behavior probability matrix output by the game reasoning engine module through evidence theory fusion, and outputs a fusion result. The chaotic rule correction unit receives the fusion result output by the cross-validation mechanism unit, processes the fusion result through nonlinear filtering, and outputs correction data. The credibility weight allocation unit receives the correction data output by the chaotic rule correction unit, processes the correction data through a fuzzy logic controller, and outputs a four-dimensional early warning vector. The four-dimensional early warning vector includes event type, probability, urgency, and disposal suggestion.
8. The intelligent question answering system based on multi-modal feature fusion according to claim 7, characterized in that, Further includes: When the synthetic data injection unit of the neural radiation field simulation module outputs the adversarial sample data, the memory enhancement network unit of the meta-learning feature encoding module processes the historical scene parameters through the neural Turing machine architecture and outputs the enhanced features to the meta-learning strategy unit of the meta-learning feature encoding module.
9. The intelligent question answering system based on multi-modal feature fusion according to claim 8, characterized in that, Also includes: When the Monte Carlo tree search unit of the game reasoning engine module outputs the group behavior probability matrix, the risk pattern reconstruction unit of the chaotic rule evolution module processes the fractal dimension through the multiple fractal detrended fluctuation analysis and outputs the optimized decision rule to the early warning threshold function optimization unit of the chaotic rule evolution module. The four-dimensional early warning vector output by the cross-modal consensus module includes the credibility weight and is output to the quantumized behavior collection module.
10. The intelligent question-answering method based on multi-modal feature fusion, applied to the intelligent question-answering system based on multi-modal feature fusion according to any one of claims 1-9, characterized in that, Includes: Step 1, deploy a quantum sensor array to capture human microscopic behavior characteristics and generate a non-Euclidean space-time behavior feature tensor, the quantum sensor array includes a Josephson junction array sensor, a single-photon avalanche diode matrix, and a superconducting quantum interference device; Step 2, receive the behavior feature tensor output by step 1, convert the behavior feature tensor into differentiable rendering parameters, generate initial rendering parameters, introduce a generative adversarial network architecture to create adversarial sample data for the initial rendering parameters to create a crowded scene, design a hyperbolic geometry embedding model to align entity virtual coordinates, process the adversarial sample data and output enhanced space-time sequences; Step 3, receive the enhanced space-time sequences output by step 2, extract hierarchical features of behavior characteristics through a variational autoencoder, integrate a neural Turing machine architecture to store scene migration knowledge, fuse hierarchical features to generate memory-enhanced features, implement a model-agnostic meta-learning algorithm to optimize feature extractor parameters, process memory-enhanced features and output 128-dimensional hyper vectors; Step 4, receive the hyper vector output by step 3, simulate group interaction strategies, and generate a group behavior probability matrix; Step 5, receive the group behavior probability matrix output by step 4, analyze the stability of the decision trajectory, and output a dynamic decision rule; Step 6, receive the dynamic decision rule output by step 5, fuse the behavior feature tensor of step 1, the enhanced space-time sequences output by step 2, and the group behavior probability matrix output by step 4, and output a hierarchical early warning signal, the hierarchical early warning signal is fed back to step 1 to adjust the sampling frequency.
Citation Information
Patent Citations
Intelligent multi-mode virtual digital human interaction system based on AI language large model, interaction method and application
CN120259499A
Digital twinborn simulation training system for whole process of administrative law enforcement
CN120542252A