Multimodal perception data fusion method and system based on silicon photonic neural network

By using dynamic topology adjustment and differentiated resource configuration based on silicon photonic neural networks, the problem of inflexible resource allocation in multimodal sensing data fusion is solved, achieving efficient computing resource management and steady-state optimization, and improving computing accuracy and latency performance.

CN121145153BActive Publication Date: 2026-02-03JILIN YUNTOU LAISENGOU DIGITAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511668256.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-14
Publication Date
2026-02-03
Estimated Expiration
2045-11-14

AI Technical Summary

Technical Problem

Existing multimodal sensing data fusion methods, due to their use of fixed computing models, are unable to dynamically adjust the allocation of computing resources according to the dynamic changes in the external environment and task requirements. This makes it difficult to achieve an efficient balance between computing accuracy, processing latency and computing power consumption in complex and ever-changing real-time sensing scenarios.

Method used

A multimodal sensing data fusion method based on silicon photonic neural networks is adopted. By configuring the dynamic interconnection of optical switch arrays and evolutionary algorithms, the network topology is dynamically adjusted. According to the characteristics of multimodal sensing data and the requirements of fusion tasks, computing resources are configured differently to achieve flexible adjustment and real-time optimization of inter-layer connections.

Benefits of technology

It achieves efficient allocation of computing resources in dynamic environments, improves computing accuracy and processing latency, reduces energy consumption, enhances system adaptability and resource utilization efficiency, and ensures steady-state optimal performance under different task scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121145153B_ABST
    Figure CN121145153B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of intelligent computing, in particular to a multi-modal perception data fusion method and system based on a silicon photon neural network. The method comprises the following steps: acquiring multi-modal perception data and extracting features; configuring a reconfigurable topology, dynamically adjusting interlayer connection by using on-off combination of an optical switch array; combining task and data characteristics, automatically searching for an optimal topology and generating connection parameters by using an evolutionary algorithm; according to complexity and real-time performance of each mode, selecting deep and shallow paths on demand to realize differentiated computing power allocation; online monitoring precision, time delay and computing power indexes, triggering topology reconfiguration based on evaluation results to adapt to environmental changes; based on topology structure parameters of the silicon photon neural network in multiple scenes, constructing a knowledge base, quickly adapting to new fusion tasks through similarity retrieval and transfer learning, and realizing high-precision, low-latency and high-energy-efficiency adaptive multi-modal fusion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent computing technology, specifically to a multimodal sensing data fusion method and system based on silicon photonic neural networks. Background Technology

[0002] Multimodal perception fusion technology plays a crucial role in cutting-edge fields such as autonomous driving, intelligent robots, and augmented reality. Its core lies in the comprehensive processing of data from various heterogeneous sensors such as cameras, LiDAR, and millimeter-wave radar to obtain a more comprehensive and reliable environmental understanding than that of a single sensor.

[0003] In actual operation, the external environment and task requirements are dynamically changing, which directly leads to continuous changes in the real-time characteristics of the input data from each sensor (such as information complexity and criticality) and the priority of the fusion processing task. Existing data fusion methods typically employ computational models with fixed topologies, such as fixed deep neural network architectures.

[0004] In such models, the allocation of internal computing resources across different data processing paths is static and fixed. This fixed structure makes it difficult to dynamically adjust the processing depth and information flow based on the real-time complexity of different modal perception data and task priorities. This results in a difficulty in achieving an efficient dynamic balance between key performance dimensions such as computational accuracy, processing latency, and computing power consumption when facing complex and ever-changing real-time perception scenarios. Consequently, it presents challenges in meeting the extreme adaptability and resource utilization efficiency requirements of perception systems in scenarios such as high-level autonomous driving.

[0005] In view of this, this application proposes a multimodal sensing data fusion method and system based on silicon photonic neural networks. Summary of the Invention

[0006] To achieve the above objectives, this application provides a multimodal sensing data fusion method and system based on silicon photonic neural networks, the specific technical solution of which is as follows:

[0007] Multimodal sensing data fusion methods based on silicon photonic neural networks include:

[0008] Acquire multimodal sensing data and extract features from the multimodal sensing data to obtain multimodal sensing data features;

[0009] Configure the topology of a silicon photonic neural network, create a dynamic interconnection method based on an optical switch array, and adjust the connection relationship between layers of the silicon photonic neural network by controlling the on / off state combination of the optical switch array;

[0010] The network topology is adjusted based on the characteristics of multimodal sensing data and the requirements of fusion tasks. An evolutionary algorithm is used to automatically search for the optimal network topology and generate connection configuration parameters that are suitable for the current task.

[0011] Based on the complexity score and real-time requirements of each modal sensing data, processing paths of different depths are dynamically selected. Shallow network paths are allocated to simple modalities, and deep network paths are allocated to complex modalities. Computational resources for the multimodal sensing data to be fused are configured differently.

[0012] Monitor the performance indicators of multimodal sensing data fusion under the current silicon photonic neural network topology, including computational accuracy, processing latency and computing power overhead. Determine whether to trigger the topology reconstruction process based on the performance evaluation results, and automatically adjust the network connection relationship to adapt to the dynamic changes in the fusion task requirements.

[0013] Record the topology parameters of silicon photonic neural networks under different application scenarios, establish a topology configuration knowledge base, retrieve the similarity of application scenarios to select topology parameters when facing new fusion task requirements, and adapt to new multimodal perception data fusion processing requirements through transfer learning methods.

[0014] Preferably, acquiring multimodal perception data includes: acquiring visual images, 3D point clouds, and motion state data through a high-definition digital camera, LiDAR, and inertial measurement unit, respectively;

[0015] Feature extraction of multimodal sensing data includes: aligning the timestamps of sensors using a network time protocol or a precise time protocol; extracting feature vectors from visual images using a convolutional neural network; extracting geometric features from 3D point clouds using voxelization and a 3D convolutional network; and extracting dynamic features from motion state data using a recurrent neural network, thereby obtaining multimodal sensing data features.

[0016] Preferably, the optical switch array is composed of two-dimensionally arranged Mach-Zehnder switch units, and each switch unit achieves optical path switching by changing the waveguide refractive index through electro-optic modulation or thermo-optic modulation.

[0017] The connection configuration array controls the connection relationships between layers by defining the connection configuration array. The array elements are binary values, representing the connection status of the corresponding input and output ports. The external digital control unit converts the connection configuration array into control signals to drive the optical switch array.

[0018] Preferably, the constructed silicon photonic neural network includes multiple cascaded optical processing layers, each comprising an array of optical interference units and a photoelectric conversion activation unit;

[0019] A reconfigurable optical switch array is integrated between adjacent optical processing layers, with multiple input and output ports.

[0020] Preferably, an evolutionary algorithm is created, in which a network topology containing multiple reconfigurable layers is defined as an individual of the evolutionary algorithm. The chromosome of the individual is formed by unfolding and splicing all the inter-layer connection configuration arrays in a predetermined order of layer by layer, row by row, and column by column to form a one-dimensional binary vector. Each gene position in the chromosome corresponds to a specific switching state in the optical switch array.

[0021] The fitness function of the evolutionary algorithm is constructed based on the performance metrics and network structure complexity of the silicon photonic neural network. The performance metrics include classification accuracy and target detection accuracy, and the network structure complexity is obtained by calculating the ratio of the total number of enabled connections to the maximum number of connections.

[0022] Preferably, the evolutionary algorithm includes: randomly generating an initial population and calculating the fitness score of each individual; selecting parent individuals using a tournament selection strategy; generating offspring by exchanging chromosome segments through crossover operations; performing mutation operations on the offspring chromosomes with a preset probability; iteratively evolving until the maximum number of generations or fitness convergence conditions are met; and decoding the chromosome of the optimal individual to obtain the connection configuration parameters of the silicon photonic neural network.

[0023] Preferably, the complexity score is quantified by calculating the information entropy of the feature vectors of each modality of sensing data, the feature values ​​are discretized into multiple intervals, the probability of occurrence of each interval is counted and the entropy value is calculated;

[0024] Set a maximum allowable processing latency for each modality as a real-time requirement, and establish a mapping relationship from complexity score to requirement depth and from real-time requirement to allowable depth.

[0025] Preferably, the final processing depth of each modality is determined based on the smaller of the required depth and the allowed depth; the final processing depth is converted into optical switch array control commands to configure cross-layer connections or bypasses for shallow paths so that the signal skips the intermediate computing layer, and to configure deep paths to pass through multiple computing layers sequentially; and the processing of multimodal sensing data with different complexity scores is achieved through differentiated configuration.

[0026] Preferably, the performance indicators for multimodal sensing data fusion include computational accuracy, processing latency, and computational cost.

[0027] The computational accuracy is obtained by evaluating the output confidence level or the matching degree with the real label; the processing delay is calculated by recording the time difference between signal input and output; the computing power overhead is calculated by statistically analyzing the number of activated optical switches and the power consumption of a single switch; a comprehensive performance index is constructed to normalize and weight the performance indicators of multimodal sensing data fusion.

[0028] Preferably, a topology refactoring process is established by setting a performance threshold and a time window. The topology refactoring process is triggered when the overall performance index is continuously lower than the threshold within the time window.

[0029] The topology reconstruction process includes: adding the current low-performance topology as feedback information to the initial population of the evolutionary algorithm, guiding the evolutionary algorithm to explore new structural spaces in the current solution neighborhood, generating new connection configuration parameters to update the optical switch array and complete the online reconstruction.

[0030] Preferably, the constructed knowledge entries include three parts: application scenario descriptor, optimal topology parameters, and silicon photonic neural network performance indicators;

[0031] The scene descriptor includes multimodal perception data types, complexity scores, fusion task requirements, and performance benchmark values; the most relevant historical configurations are retrieved by calculating cosine similarity as the initial configuration for the new task; and the retrieved topological structure parameters are used as the initial genes for the evolutionary algorithm for fine-tuning and optimization.

[0032] Preferably, the transfer learning process includes: generating an initial population around the retrieved historical best topology parameters; generating new candidate topologies by applying random perturbations to some connection states; and iteratively optimizing the evolutionary algorithm within a high-starting-point concentrated solution space to reduce the number of iterations required for the search.

[0033] Continuously store successfully configured knowledge items into the knowledge base to establish a mapping between application requirements and network configurations.

[0034] A multimodal sensing data fusion system based on silicon photonic neural networks is used to implement the aforementioned multimodal sensing data fusion method based on silicon photonic neural networks. It includes: a data acquisition module, a topology configuration module, a topology optimization module, a performance evaluation module, an optimization and reconstruction module, and a migration and extension module.

[0035] The data acquisition module is used to acquire multimodal sensing data and extract features from the multimodal sensing data to obtain multimodal sensing data features.

[0036] The topology configuration module is used to configure the topology of the silicon photonic neural network, create a dynamic interconnection method based on an optical switch array, and adjust the connection relationship between layers of the silicon photonic neural network by controlling the on / off state combination of the optical switch array.

[0037] The topology optimization module is used to adjust the network topology based on the characteristics of multimodal sensing data and the requirements of fusion tasks, and to automatically search for the optimal network topology using an evolutionary algorithm to generate connection configuration parameters that are suitable for the current task.

[0038] The performance evaluation module is used to dynamically select processing paths of different depths based on the complexity score and real-time requirements of each modality of sensing data, allocate shallow network paths for simple modalities and deep network paths for complex modalities, and perform differentiated configuration of computing resources for the multimodal sensing data to be fused.

[0039] The optimization and reconstruction module is used to monitor the performance indicators of multimodal sensing data fusion under the current silicon photonic neural network topology, including computational accuracy, processing latency and computing power overhead. Based on the performance evaluation results, it determines whether to trigger the topology reconstruction process and automatically adjusts the network connection relationship to adapt to the dynamic changes in the fusion task requirements.

[0040] The transfer extension module is used to record the topological parameters of silicon photonic neural networks under different application scenarios, establish a topological configuration knowledge base, retrieve the similarity of application scenarios to select topological parameters when facing new fusion task requirements, and adapt to new multimodal perception data fusion processing requirements through transfer learning methods.

[0041] The beneficial effects of this application are as follows: This application standardizes, aligns and reduces the dimensionality of multimodal raw data, improves the signal-to-noise ratio and input consistency, reduces the bandwidth and dynamic range pressure of optoelectronic interface, and provides compact features that can be directly mapped for subsequent optical domain computing.

[0042] This application implements programmable inter-layer routing on-chip through hardware-level dynamic interconnection, reducing unnecessary computation and data transfer; it can quickly adapt to different tasks without changing the hardware, and balance high bandwidth and low latency.

[0043] This application uses multiple objectives such as accuracy, latency and energy consumption for automatic optimization, and combines physical constraints to obtain feasible and robust connection configurations; compared with manual design, it reduces trial and error costs and improves global optimality and adaptation speed.

[0044] This application schedules computing resources differently based on complexity and real-time requirements, enabling "easy tasks to take short paths and difficult tasks to take deep paths," which significantly reduces overall energy consumption and congestion while ensuring the time limits of critical modes.

[0045] This application continuously assesses accuracy, latency, and computing power consumption to promptly detect performance degradation and trigger local or global refactoring, thereby maintaining the system's steady-state optimality and service quality in dynamic environments.

[0046] This application enables rapid invocation and few-sample adaptation of similar tasks based on effective configurations in different scenarios, shortens deployment and convergence time, and improves cross-scenario generalization and maintenance efficiency. Attached Figure Description

[0047] Figure 1 A flowchart of the multimodal sensing data fusion method based on silicon photonic neural networks provided in this application;

[0048] Figure 2 A flowchart illustrating the construction process of the dynamic topology reconfiguration mechanism for silicon photonic networks provided in this application;

[0049] Figure 3Flowchart of optimal topology search for silicon photonic neural networks driven by evolutionary algorithm provided in this application;

[0050] Figure 4 A flowchart illustrating the dynamic allocation of differential processing paths for multimodal sensing data provided in this application;

[0051] Figure 5 The flowchart for closed-loop monitoring and online adaptive adjustment of the performance of the silicon photonic neural network provided in this application;

[0052] Figure 6 The flowchart for the construction of the topology configuration knowledge base and rapid initialization of transfer learning provided in this application;

[0053] Figure 7 The structural diagram of the multimodal sensing data fusion system based on silicon photonic neural network provided in this application. Detailed Implementation

[0054] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the specific embodiments of this application will be described in detail below with reference to the accompanying drawings.

[0055] Many specific details are set forth in the following description in order to provide a full understanding of this application. However, this application may also be implemented in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this application. Therefore, this application is not limited to the specific embodiments disclosed below.

[0056] Secondly, the term "an embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of this application. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single embodiment or an embodiment selectively excluded from other embodiments.

[0057] Example 1

[0058] Reference Figures 1 to 6 This is the first embodiment of the present application, such as Figure 1 As shown, a multimodal sensing data fusion method based on silicon photonic neural networks is provided.

[0059] Step 1: Acquire multimodal sensing data and extract features from it. Feature extraction refers to extracting features from the multimodal sensing data.

[0060] Multimodal sensing data is synchronously acquired and preprocessed. External environmental information is acquired in real time through multiple heterogeneous sensors deployed on the sensing terminal. These heterogeneous sensors include at least a high-definition digital camera for capturing visual image data, a LiDAR for acquiring three-dimensional spatial point cloud data, and an inertial measurement unit (IMU) for measuring the carrier's attitude and motion state. To ensure the accuracy of subsequent data fusion, a high-precision time synchronization mechanism is employed during the data acquisition stage. For example, Network Time Protocol (NTP) or Precise Time Protocol (PTP) is used to timestamp align all sensors, ensuring a strict correspondence between the acquired modal sensing data in the time dimension.

[0061] Since the raw data acquired for multimodal sensing are in various formats—for example, visual images can be two-dimensional pixel arrays, LiDAR data can be unordered and sparse three-dimensional coordinate point sets, and IMU data can be time series containing acceleration and angular velocity—independent feature extraction operations are performed on each of the acquired modal raw data to map them from the raw data space to a high-dimensional feature space, facilitating subsequent processing by silicon photonic neural networks.

[0062] For example, for visual image data, a feature extraction method based on convolutional neural networks (CNNs) can be used. Each acquired image frame... Perform normalization processing, scaling its pixel values ​​to... Within a given range, the image is then input into a pre-defined convolutional neural network (CNN). The CNN processes the image through multiple convolutional layers, activation function layers, and pooling layers. The convolution operation uses a learnable convolutional kernel to perform a sliding window operation on the input image feature map to extract local spatial features. Through layer-by-layer stacking convolution and pooling operations, a flattened one-dimensional feature vector that represents high-level semantic information of the image (such as edges, textures, and object contours) is ultimately generated. This process transforms high-dimensional pixel data into a compact and semantically rich feature representation, effectively reducing data redundancy.

[0063] For LiDAR point cloud data, a voxel-based 3D feature extraction method can be used. Given the sparsity and unordered nature of the original point cloud data, the 3D space containing the point cloud is first divided into a regular voxel grid. Features are encoded for all points within each non-empty voxel, such as calculating the average coordinates, density, or reflection intensity of the point cloud, forming the initial features of that voxel. Then, a 3D convolutional network is applied to process this voxelized data to capture the geometric structure and neighborhood relationships in 3D space. After multiple layers of 3D convolution and pooling, a feature vector representing the 3D spatial layout and object geometry of the scene is generated. This processing method transforms unstructured point cloud data into a regular format suitable for neural network processing, thereby effectively learning the geometric structural features of the environment from a set of discrete points.

[0064] For IMU time series data, a feature extraction method based on recurrent neural networks (RNNs) can be used. A sequence of continuous IMU readings (including triaxial acceleration and triaxial angular velocity) within a short time window (e.g., 100 milliseconds) is formed and input into a Long Short-Term Memory (LSTM) network or a gated recurrent unit (GRU). The RNN receives an input vector at each time step and updates its internal hidden state, combining the current input and the hidden state from the previous time step. After processing the entire sequence, the hidden state vector of the last time step is taken as the dynamic feature representation of the vehicle's motion state within that time window. This feature processing method can effectively extract dynamic information such as the vehicle's motion trend and attitude changes from time series data.

[0065] This step completes the acquisition, synchronization, and preprocessing of heterogeneous, multi-source raw sensing data, and extracts structured feature vectors that characterize the core information of different modal sensing data based on their inherent characteristics. This process not only unifies data from different sources and formats into a standardized feature space, but also extracts deep semantic and structural information from the data through deep learning models, greatly improving the data's distinguishability and information density. This provides high-quality input for subsequent efficient and accurate multimodal sensing data fusion using silicon photonic neural networks, and is a solid foundation for ensuring the final performance of the entire fusion method.

[0066] Step 2: Configure the topology of the silicon photonic neural network and create a dynamic interconnection method based on an optical switch array. By controlling the on / off states of the optical switch array, the connection relationships between the layers of the silicon photonic neural network are adjusted. See also... Figure 2 This is a flowchart illustrating the construction process of the dynamic topology reconfiguration mechanism for silicon photonic networks in this step.

[0067] The basic hierarchical structure of a silicon photonic neural network is constructed, which consists of multiple cascaded optical processing layers, each implemented on a silicon-based photonic integrated chip (PIC). A typical optical processing layer includes an array of optical interference units (e.g., a mesh structure composed of Mach-Zehnder interferometers, MZI) for weighted operations and photoelectric conversion activation units for introducing nonlinearity.

[0068] This step innovatively integrates a reconfigurable optical switch array between any two adjacent optical processing layers to establish a dynamic inter-layer interconnection mechanism. This optical switch array is physically implemented as a [missing information - likely a specific type of array]. One input port and A crossbar switch array with output ports, wherein This represents the number of output neurons in the previous layer. This represents the number of input neurons in the next layer. With this design, the signal transmission path from the previous layer to the next is no longer fixed, but can be changed in real time according to control commands. This structure gives the topology of the neural network, i.e., the connections between neurons, hardware-level flexibility, providing a physical basis for adaptively adjusting the network structure according to task requirements.

[0069] Specifically, the dynamic interconnection method based on an optical switch array includes: the optical switch array is composed of a two-dimensionally arranged, independently controllable optical switch unit, each optical switch unit preferably employing an electro-optic or thermo-optic modulated 2×2 Mach-Zehnder switch (MZS). A single MZS introduces a controllable phase shift by applying a control voltage or current to one arm of the waveguide, thereby changing the refractive index of that arm. When the input optical signal enters the MZS, the optical power is directed to one of the two output ports based on the phase difference between the two arms. By precisely controlling the phase shift, the MZS can operate in either a cross or bar state, thus achieving precise selection of the optical path. The entire... Optical switch arrays, through the combination of these MZS units, can establish any kind of connection mapping relationship between the input port set and the output port set. In this way, a mechanism is established that enables high-speed manipulation of optical path connections via external electrical signals, achieving dynamic adjustment of inter-layer connectivity. The response speed of this method can reach nanoseconds (for electro-optic modulation) or microseconds (for thermo-optic modulation), far exceeding the efficiency of software-defined connections in traditional electronic computing architectures, thus ensuring real-time topology reconstruction.

[0070] To enable precise and programmatic control of this dynamic interconnection, a binary connection configuration array is defined. The array has dimensions of Its elements The value of defines whether inter-layer connections exist:

[0071] ;

[0072] in, , , For the number of input ports, Number of output ports; a specific connection configuration array. It uniquely defines the interlayer topology of a neural network.

[0073] In practice, an external digital control unit configures the logical connections to the array. This is translated into a specific set of analog voltage or current signals, which are applied to each MZS unit in the optical switch array, thereby driving the optical switch array to physically realize the switching action. The described connection topology. If the amplitude vector of the optical signal output from the previous layer is... Then, after passing through After configuring the optical switch array, the optical signal amplitude vector reaching the input of the next layer. The Each component can be represented as: ;in, Represents the next layer (the first) (Layer) The amplitude of the optical signal received at each input port; Represents the previous layer (the first) (Layer) The amplitude of the optical signal emitted from each output port; These are the elements in the configuration array, used as logic variables to control the on / off state of switches; this formula describes how a programmable array... This allows for precise control of signal flow between layers. This programmed control method makes network topology modifications simple, fast, and repeatable, providing a direct and efficient hardware interface for upper-layer algorithms to optimize and adjust the network structure.

[0074] This step configures a dynamically reconfigurable topology for the silicon photonic neural network by introducing an optical switch array composed of MZS arrays between the optical processing layers and establishing a connection configuration array. The mapping control method to physical switch states enables rapid, accurate, and programmable adjustment of the interlayer connections in neural networks. It not only breaks through the limitations of the traditional fixed hardware topology of neural networks, but also provides core hardware enabling technology for building intelligent computing that can adaptively evolve its internal structure according to data characteristics and task requirements, greatly enhancing the flexibility and applicability of the entire multimodal perception data fusion method.

[0075] Step 3: Adjust the network topology based on the characteristics of the multimodal sensing data and the requirements of the fusion task. Use an evolutionary algorithm to automatically search for the optimal network topology and generate connection configuration parameters suitable for the current task. See also... Figure 3 This is a flowchart of the optimal topology search process for the silicon photonic neural network driven by the evolutionary algorithm in this step.

[0076] To enable automatic search of network topologies using evolutionary algorithms, a scheme is established to bidirectionally map the topology of silicon photonic neural networks to the genetic code in the evolutionary algorithm.

[0077] Specifically, in the evolutionary algorithm, a silicon photonic neural network topology containing multiple reconfigurable layers is defined as an individual, and the chromosome of this individual is an array of connections between all reconfigurable layers. It is composed of linked elements. If the network contains A reconfigurable inter-layer connection, each connection consisting of a Connection configuration array Definition (where) For layer index, If the chromosome of that individual is a one-dimensional binary vector, then the entire array will be represented by the chromosome. The sequences are unfolded and assembled in a predetermined order (e.g., layer by layer, row by row, column by column). This encoding method ensures that each gene locus in the chromosome corresponds precisely to a specific switching state in the optical switch array, thereby establishing a direct and unambiguous link from algorithmic operations to physical hardware configuration, providing a solid foundation for subsequent evolutionary operations.

[0078] A fitness function is constructed to evaluate the merits of each candidate topology. The design of this function is closely related to the characteristics of multimodal sensing data and the requirements of fusion tasks, and aims to evaluate computational accuracy and resource consumption.

[0079] fitness function Defined as a multi-objective optimization problem, it combines two dimensions: model performance and network complexity. The specific expression is as follows: ;in, The final fitness score represents the candidate topology; a higher score indicates a better topology. This represents the performance metrics of the topology under a specific fusion task, such as the classification accuracy or the average precision of object detection obtained after inference on a given multimodal validation dataset. It represents the structural complexity of the network. It is calculated by normalizing the total number of elements with a value of 1 in all inter-layer connection configuration arrays, which is the ratio of the total number of connections actually used in the network to the maximum possible number of connections. and These are preset weighting coefficients, both of which are positive real numbers, used to adjust the relative importance of performance and complexity in the final evaluation; for example, for tasks requiring high precision, the weighting coefficient can be increased. For tasks deployed on resource-constrained devices, this can increase... Through this fitness function, the evolutionary algorithm is guided to find network topologies that can efficiently complete the fusion task and also have sparse connectivity and low power consumption characteristics, thus enabling the generated solution to have both high performance and high efficiency.

[0080] The optimal silicon photonic neural network topology is searched in the iterative process of the evolutionary algorithm. In the iterative process, an initial population is randomly generated, containing several candidate topology individuals with different initial chromosome codes. In each generation of evolutionary iteration, the optimal topology is first determined according to the aforementioned fitness function. The fitness score of each individual in the population is calculated. Then, strategies such as roulette wheel selection or tournament selection are used to select superior individuals as parents with a higher probability based on their fitness scores. The selected parents are paired up and crossover operations are performed to exchange partial chromosome segments of each other (e.g., single-point crossover or uniform crossover) to produce offspring that inherit the superior traits of their parents.

[0081] To maintain population diversity and explore new topological possibilities, a bitwise mutation operation is performed on the chromosomes of offspring individuals with a small preset probability (e.g., 0.01). This involves randomly flipping certain bits in the binary code (changing 0 to 1 or 1 to 0), which physically corresponds to randomly breaking or establishing a neuronal connection. The newly generated offspring population replaces the old population and enters the next generation cycle. This iterative process continues until a preset termination condition is met, such as reaching the maximum number of generations or the population's optimal fitness score no longer significantly improving over multiple generations.

[0082] After the evolutionary algorithm terminates, the individual with the highest fitness score is selected from the final population, and its chromosome is decoded to obtain an optimal connectivity array. ; Indicates the first The optimal connection configuration matrix for the layer; this array represents the network connection configuration parameters most suited to the current multimodal sensing data characteristics and fusion task requirements, as searched by the evolutionary algorithm. Loading these parameters into the digital control unit described in step 2 drives the optical switch array to achieve this optimal topology.

[0083] This step constructs an automated and adaptive method for optimizing neural network topologies, overcoming the limitations of traditional methods that rely on human experience and extensive trial and error in network design. It can efficiently search for the optimal solution that balances performance and efficiency within a vast topology space for specific tasks. The connection configuration parameters generated in this process enable silicon photonic neural networks to adaptively process specific multimodal sensing data, thereby significantly improving the final effect of data fusion and resource utilization efficiency.

[0084] Step 4: Based on the complexity score and real-time requirements of each modal sensing data, dynamically select processing paths of different depths. Assign shallow network paths to simple modalities and deep network paths to complex modalities, thus differentiating the computational resources for the multimodal sensing data to be fused. See also... Figure 4 This is a flowchart illustrating the dynamic allocation of the multimodal sensing data differentiation processing path for this step.

[0085] The complexity and real-time requirements of each modal sensing data entering the silicon photonic neural network are quantitatively evaluated to provide a basis for decision-making on subsequent differentiated resource allocation.

[0086] Specifically, for any mode The complexity score of the data is quantified by calculating the information entropy of the feature vector of that modality, thus obtaining the complexity score. This calculation is performed on the feature data extracted in step 1. First, the continuous feature values ​​are discretized to... Within each interval, the statistics for each discrete interval within the most recent time window are then calculated. probability of occurrence Final complexity score It is given by the following formula:

[0087] ;

[0088] in, Indicates the first Data stream of various modes It is the first discretized feature of the modal sensing data. One possible value, It means The frequency of occurrence in recent data This represents the logarithmic function with base 2. A higher value indicates a greater amount of information and higher uncertainty in the modality-sensing data, meaning the data is more complex. This applies to each modality... Set a maximum allowable processing latency determined by the specific application scenario. This serves as a quantitative indicator for its real-time requirements. In this way, abstract data characteristics and task requirements are transformed into precise numerical values, laying the foundation for dynamic and precise allocation of computing resources.

[0089] Based on the above quantitative evaluation results, a suitable processing path depth is dynamically determined for each modality. This process includes two mapping steps: the first mapping step establishes a mapping from complexity score to "required depth"; the second mapping step establishes a mapping from real-time requirements to "allowed depth".

[0090] Specifically, the total number of layers in the entire silicon photonic neural network Divide into several depth levels, such as {shallow, medium, deep}; score based on the complexity of all modes to be processed. The distribution is set with several thresholds to determine different... The score range is mapped to different levels of demand depth. For example, modalities with scores below a certain threshold are considered simple, and their demand depth is "shallow." Similarly, based on the maximum permissible delay for each modality... The fundamental delay required for photonic networks to process a unit layer depth (Based on pre-calibrated hardware parameters), calculate the maximum number of processing layers that the modality can handle and map it to the allowed depth levels. For example, if A larger value allows for a depth of "deep". Through these two mappings, each modality simultaneously acquires two depth attributes: one determined by its own complexity and the other limited by the latency of external tasks.

[0091] Based on the required depth and allowed depth, a specific processing path is ultimately selected for each modality. The selection principle is to provide a specific processing path for each modality. Final processing depth of allocation It is necessary to simultaneously meet the processing requirements of its data complexity and the real-time constraints of the task. Therefore, the final depth is the smaller of the required depth and the allowed depth, i.e.: ,in, To obtain the minimum value, For modality The depth of demand For modality The allowed depth. This decision-making mechanism ensures that the processing of complex data does not exceed time limits in pursuit of depth, and also avoids the waste of allocating unnecessary depth computing resources to simple data.

[0092] The final processing depth will be determined for each mode. This translates into specific control commands for the optical switch array in the silicon photonic neural network. For modes assigned shallow network paths, the control commands configure the optical switch array to generate cross-layer connections or bypasses, allowing the optical signal of this mode to skip certain intermediate computational layers and directly transmit to later layers or fusion layers. For modes assigned deep network paths, the control commands configure the optical switches to guide their optical signals sequentially through more computational layers for sufficient feature transformation and extraction.

[0093] For example, suppose there is currently a video modality and radar modes ,like Complexity score Maximum allowable delay ms; Complexity score Maximum allowable delay ms; complexity score based on preset rules. Corresponding Depth of Needs For depth, complexity score Corresponding to shallow layers; simultaneously, maximum allowable delay ms corresponds to the allowed depth For the middle layer, the maximum allowable latency ms corresponds to the deeper level; therefore, Depth of demand =Deep layer, Allowable depth =Middle layer, final selection Final processing depth ;for , Depth of demand =Shallow layer, Allowable depth =Deep, final choice Final processing depth Therefore, Configure a processing path of medium depth for Configure a shallow processing path.

[0094] This step enables dynamic and differentiated allocation of computing resources based on the inherent complexity of different modal sensing data and external task requirements. This approach of tailoring processing paths to different data not only ensures that complex data can be fully processed to extract deep features, but also allows simple data to be processed quickly to meet real-time requirements. Thus, while ensuring the overall performance of multimodal sensing data fusion, it greatly improves the computational and energy efficiency of silicon photonic neural networks, achieving intelligent and refined management of valuable computing power.

[0095] Step 5: Monitor the performance metrics of multimodal sensing data fusion under the current silicon photonic neural network topology, including computational accuracy, processing latency, and computational overhead. Based on the performance evaluation results, determine whether to trigger the topology reconstruction process and automatically adjust the network connections to adapt to the dynamic changes in the fusion task requirements. See also... Figure 5 This is a flowchart of the closed-loop monitoring and online adaptive adjustment of the silicon photonic neural network performance in this step.

[0096] Real-time multi-dimensional and multi-modal sensing data fusion performance indicators are performed on the silicon photonic neural network operating under the current topology to obtain quantitative data on its dynamic working state.

[0097] This monitoring process collects three core performance metrics in parallel: computational accuracy, processing latency, and computational overhead. Computational accuracy is indirectly measured by evaluating the confidence or determinism of the fusion result at the output. For scenarios with labeled reference data, the accuracy metric is obtained by directly calculating the matching degree between the current model output and the true labels. The processing delay is achieved by recording the time when the optical signal enters the network input terminal. The corresponding fusion result is generated at the output time. The time difference is used to calculate the delay precisely. The computing power overhead is calculated by counting the number of optical switches in the current network topology that are in an active state, and combining this with the rated power consumption of each individual switch, to determine the total power consumption of the entire network under the current configuration. This value directly reflects the scale of computing resources utilized. By continuously collecting these three indicators, time-series data on network performance is generated, providing a precise basis for subsequent performance evaluation and decision-making.

[0098] Based on the collected real-time performance metrics, a comprehensive performance evaluation function is constructed to comprehensively assess the effectiveness of the current network topology. This comprehensive performance evaluation function unifies multiple, and even conflicting, performance metrics into a single evaluation score.

[0099] Specifically, a comprehensive performance index is defined. The calculation method is as follows:

[0100] ;

[0101] in, , and These are the real-time performance values ​​that were previously monitored; , and These are performance target benchmarks set based on the current requirements of the fusion task, for example, a target accuracy of 95%, a target latency of 10ms, and a target power consumption of 400mW; , and These are the weighting coefficients of each performance indicator, and satisfy... These weighting coefficients can be adjusted according to the focus of the task. For example, for tasks with high real-time requirements, the weighting coefficients can be increased. The value of this comprehensive performance index. The design transforms the performance metrics of accuracy, latency, and computational overhead—three different dimensions—into a standardized, easily comparable scalar through normalization and weighted summation; a higher... The value represents the overall performance of the current network topology while satisfying all constraints.

[0102] Based on the calculated comprehensive performance index A topology reconfiguration trigger discrimination mechanism is established, which aims to identify a continuous decline in network performance and avoid misjudgments caused by instantaneous data fluctuations or noise.

[0103] Specifically, setting performance thresholds and time window During operation, continuous calculations are performed. Value, and check in the past During the time period, Does the value remain below the threshold? ,when The conditions are met continuously for more than If the duration is determined, the current topology is deemed unsuitable for the current data characteristics or task requirements, and the topology reconfiguration process is formally triggered. This time-window-based approach enhances the robustness of the decision-making process, ensuring that the resource-intensive reconfiguration process is only initiated when network performance experiences a stable degradation.

[0104] Once the topology reconstruction process is triggered, the topological connections of the neural network are adjusted. During this adjustment process, the current, poorly performing topology and its corresponding low-performance scores are removed. As feedback, the evolutionary algorithm defined in step 3 is invoked again.

[0105] Specifically, the connection configuration parameters of the current topology are added as a low-fitness individual to the initial population of the evolutionary algorithm. This guides the algorithm to focus on exploring and mutating the neighborhood of the current solution in the next iteration, or to tend to explore new structural spaces that differ significantly from the current solution, in order to discover more quickly those that can improve the overall performance index. The novel topology generates a new set of connection configuration parameters after the algorithm converges, and immediately uses them to update the on / off state of the optical switch array, thus completing the online reconstruction of the network topology.

[0106] This step establishes a complete closed-loop adaptive control process of monitoring, evaluation, decision-making, and adjustment, making the silicon photonic neural network no longer a static computational structure, but an intelligent network that can actively perceive its own performance, diagnose problems, and perform self-optimization. This dynamic adjustment capability ensures that the fusion method can always maintain a near-optimal working state when facing environmental changes, data distribution drift, or changes in task requirements, thereby significantly improving the robustness, timeliness, and energy efficiency of the entire multimodal sensing data fusion method.

[0107] Step 6: Record the topology parameters of the silicon photonic neural network under different application scenarios, establish a topology configuration knowledge base, and when facing new fusion task requirements, retrieve the similarity of application scenarios to select topology parameters, and adapt to new multimodal sensing data fusion processing requirements through transfer learning methods. See also Figure 6 This is a flowchart for the topology configuration knowledge base construction and transfer learning quick initialization process in this step.

[0108] The configurations of silicon photonic neural networks that have been verified through the aforementioned steps and have reached a stable and high-performance state are recorded in a structured manner to build a topology configuration knowledge base that can be reused for future tasks.

[0109] After each fusion task is successfully executed and optimized, a set of key information is extracted and stored to form a complete configuration-performance knowledge entry. The configuration-performance knowledge entry specifically includes three parts: the first part is the application scenario descriptor, which is specifically a vector quantifying the characteristics of the current task. Its dimensions include the types of multimodal perception data processed (such as vision, radar, and lidar), the complexity score of each modality, the specific fusion task requirements (such as target recognition and scene segmentation), and the performance requirement benchmark (such as...). The second part is the corresponding optimal topology parameters. The first part is the final determined optical switch array connection array composed of 0s and 1s; the second part is the actual stable performance indicators achieved under this topology. That is, the comprehensive performance index calculated in step 5.

[0110] By continuously storing these successfully configured "scenario-topology-performance" triples into the database, a rich knowledge base is gradually built. This knowledge base effectively links abstract application requirements with specific photonic network physical configurations, laying a data foundation for subsequent intelligent decision-making.

[0111] When faced with a new multimodal sensing data fusion task, the system will first use the established topology configuration knowledge base for rapid initialization.

[0112] Specifically, the requirements and data characteristics of the new task are first transformed into application scenario descriptors with a format consistent with the knowledge base, denoted as... Then, through calculation With each historical scene descriptor stored in the knowledge base The similarity between them is used to retrieve the most relevant historical experiences.

[0113] The similarity is calculated using the cosine similarity metric, and the formula is as follows:

[0114] ;

[0115] in, Indicates new mission scenarios and historical scenarios Similarity score; and It is a scene descriptor vector; and These two vectors are respectively at the th... Components in each dimension; It is the total dimension of the scene descriptor vector.

[0116] After the calculation is complete, select the historical entry with the highest similarity score. Then, directly extract the topology parameters corresponding to the optimal historical entry. This serves as the initial network configuration for new tasks. This process enables rapid knowledge retrieval and reuse, providing a high-quality, proven "warm-start" configuration for new tasks and significantly reducing the time required to search for the optimal topology from scratch.

[0117] Based on the retrieved initial topology parameters, a transfer learning approach is used to fine-tune them to precisely adapt to the specific requirements of the new task. Since the new task is highly similar to, but not entirely identical to, the retrieved historical tasks, directly reusing the historical topology may not achieve optimal performance. Therefore, this method uses the retrieved topology parameters... Instead of starting the search from a completely random population, the genes are used as the initial superior genes in the evolutionary algorithm in step 3.

[0118] Specifically, the initial population of the evolutionary algorithm will revolve around To generate, for example, by... A small-probability random perturbation (mutation) is applied to some of the connection states (0 or 1) in the array to generate a set of... Candidate topologies with similar but slightly different structures are selected. The evolutionary algorithm then iteratively optimizes within this high-starting, more concentrated solution space. Since the starting point of the search is already a fairly good solution, the algorithm can converge to the optimal topology for the new task much faster, requiring far fewer iterations and less computational resources than a completely recursive search. This approach constitutes an efficient transfer from old knowledge to new applications, ensuring rapid adaptation of network configuration and high performance.

[0119] This step elevates the entire method from a single-task optimization process to a self-evolving framework with learning, memory, and reasoning capabilities. It can solidify past successful experiences into reusable knowledge, and when facing new challenges, it can quickly find a high-quality solution starting point through analogical reasoning, and then perform fine-tuning through transfer learning, which greatly improves the efficiency and response speed in dealing with new integrated tasks.

[0120] Example 2

[0121] Reference Figure 7 The second embodiment of this application provides a multimodal sensing data fusion system based on silicon photonic neural networks. The system includes: a data acquisition module, a topology configuration module, a topology optimization module, a performance evaluation module, an optimization and reconstruction module, and a migration and extension module.

[0122] The data acquisition module is used to acquire multimodal sensing data and extract features from the multimodal sensing data.

[0123] The topology configuration module is used to configure the topology of the silicon photonic neural network, create a dynamic interconnection method based on an optical switch array, and adjust the connection relationship between layers of the silicon photonic neural network by controlling the on / off state combination of the optical switch array.

[0124] The topology optimization module is used to adjust the network topology based on the characteristics of multimodal sensing data and the requirements of fusion tasks. It uses an evolutionary algorithm to automatically search for the optimal network topology and generate connection configuration parameters that are suitable for the current task.

[0125] The performance evaluation module is used to dynamically select processing paths of different depths based on the complexity score and real-time requirements of each modality of sensing data, allocate shallow network paths for simple modalities and deep network paths for complex modalities, and perform differentiated configuration of computing resources for the multimodal sensing data to be fused.

[0126] The optimization and reconstruction module is used to monitor the performance indicators of multimodal sensing data fusion under the current silicon photonic neural network topology, including computational accuracy, processing latency and computing power overhead. Based on the performance evaluation results, it determines whether to trigger the topology reconstruction process and automatically adjusts the network connection relationship to adapt to the dynamic changes in the fusion task requirements.

[0127] The transfer extension module is used to record the topological parameters of silicon photonic neural networks under different application scenarios, establish a topological configuration knowledge base, retrieve the similarity of application scenarios to select topological parameters when facing new fusion task requirements, and adapt to new multimodal perception data fusion processing requirements through transfer learning methods.

[0128] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the displayed or discussed mutual couplings or direct couplings or communication connections may be through some communication interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.

[0129] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments under the guidance of this application without departing from the spirit and scope of protection of the claims. All of these variations are within the protection scope of this application.

Claims

1. A multimodal sensing data fusion method based on silicon photonic neural networks, characterized in that, include: Acquire multimodal sensing data and extract features from the multimodal sensing data to obtain multimodal sensing data features; Configure the topology of a silicon photonic neural network, create a dynamic interconnection method based on an optical switch array, and adjust the connection relationship between layers of the silicon photonic neural network by controlling the on / off state combination of the optical switch array; The network topology is adjusted based on the characteristics of multimodal sensing data and the requirements of fusion tasks. An evolutionary algorithm is used to automatically search for the optimal network topology and generate connection configuration parameters that are suitable for the current task. Based on the complexity score and real-time requirements of each modal sensing data, processing paths of different depths are dynamically selected. Shallow network paths are allocated to simple modalities, and deep network paths are allocated to complex modalities. Computational resources for the multimodal sensing data to be fused are configured differently. Monitor the performance indicators of multimodal sensing data fusion under the current silicon photonic neural network topology, including computational accuracy, processing latency and computing power overhead. Determine whether to trigger the topology reconstruction process based on the performance evaluation results, and automatically adjust the network connection relationship to adapt to the dynamic changes in the fusion task requirements. Record the topology parameters of silicon photonic neural networks under different application scenarios, establish a topology configuration knowledge base, retrieve the similarity of application scenarios to select topology parameters when facing new fusion task requirements, and adapt to new multimodal perception data fusion processing requirements through transfer learning methods.

2. The multimodal sensing data fusion method based on silicon photonic neural networks according to claim 1, characterized in that, Acquiring multimodal perception data includes: acquiring visual images, 3D point clouds, and motion state data through high-definition digital cameras, LiDAR, and inertial measurement units, respectively; Feature extraction of multimodal sensing data includes: aligning the timestamps of sensors using a network time protocol or a precise time protocol; extracting feature vectors from visual images using a convolutional neural network; extracting geometric features from 3D point clouds using voxelization and a 3D convolutional network; and extracting dynamic features from motion state data using a recurrent neural network, thereby obtaining multimodal sensing data features.

3. The multimodal sensing data fusion method based on silicon photonic neural networks according to claim 2, characterized in that, The optical switch array is composed of two-dimensionally arranged Mach-Zehnder switch units. Each switch unit changes the waveguide refractive index through electro-optic modulation or thermo-optic modulation to achieve optical path switching. The connection configuration array controls the connection relationships between layers by defining the connection configuration array. The array elements are binary values, representing the connection status of the corresponding input and output ports. The external digital control unit converts the connection configuration array into control signals to drive the optical switch array.

4. The multimodal sensing data fusion method based on silicon photonic neural networks according to claim 3, characterized in that, The constructed silicon photonic neural network includes multiple cascaded optical processing layers, each containing an array of optical interference units and a photoelectric conversion activation unit. A reconfigurable optical switch array is integrated between adjacent optical processing layers, with multiple input and output ports.

5. The multimodal sensing data fusion method based on silicon photonic neural networks according to claim 4, characterized in that, An evolutionary algorithm is created by defining a network topology containing multiple reconfigurable layers as an individual of the evolutionary algorithm. The chromosome of an individual is formed by unfolding and splicing all the inter-layer connection configuration arrays in a predetermined order of layer-by-layer, row-by-row, and column-by-column to form a one-dimensional binary vector. Each gene position in the chromosome corresponds to a specific switching state in the optical switch array. The fitness function of the evolutionary algorithm is constructed based on the performance metrics and network structure complexity of the silicon photonic neural network. The performance metrics include classification accuracy and target detection accuracy, and the network structure complexity is obtained by calculating the ratio of the total number of enabled connections to the maximum number of connections.

6. The multimodal sensing data fusion method based on silicon photonic neural networks according to claim 1, characterized in that, The evolutionary algorithm includes: randomly generating an initial population and calculating the fitness score of each individual; selecting parent individuals using a tournament selection strategy; generating offspring by exchanging chromosome segments through crossover operations; performing mutation operations on the offspring chromosomes with a preset probability; iteratively evolving until the maximum number of generations or fitness convergence conditions are met; and decoding the chromosome of the optimal individual to obtain the connection configuration parameters of the silicon photonic neural network.

7. The multimodal sensing data fusion method based on silicon photonic neural networks according to claim 6, characterized in that, The complexity score is quantified by calculating the information entropy of the feature vectors of each modality of sensing data. The feature values ​​are discretized into multiple intervals, and the probability of occurrence of each interval is statistically analyzed and the entropy value is calculated. Set a maximum allowable processing latency for each modality as a real-time requirement, and establish a mapping relationship from complexity score to requirement depth and from real-time requirement to allowable depth.

8. The multimodal sensing data fusion method based on silicon photonic neural networks according to claim 7, characterized in that, The final processing depth of each modality is determined based on the smaller of the required depth and the allowed depth. The final processing depth is then converted into optical switch array control commands to configure cross-layer connections or bypasses for shallow paths, allowing signals to skip intermediate computation layers, and to configure deep paths to pass through multiple computation layers sequentially. Differentiated configurations are used to process multimodal sensing data with different complexity scores.

9. The multimodal sensing data fusion method based on silicon photonic neural networks according to claim 8, characterized in that, The performance metrics for multimodal sensing data fusion include computational accuracy, processing latency, and computational cost. The calculation accuracy is obtained by evaluating the output confidence level or the degree of matching with the real label; the processing delay is calculated by recording the time difference between the signal input and output. The computing power overhead is calculated by statistically analyzing the number of activated optical switches and the power consumption of a single switch; a comprehensive performance index is constructed to normalize and weight the performance indicators of multimodal sensing data fusion.

10. The multimodal sensing data fusion method based on silicon photonic neural networks according to claim 9, characterized in that, Establish a topology refactoring process by setting performance thresholds and time windows. The topology refactoring process is triggered when the overall performance index is consistently below the threshold within the time window. The topology reconstruction process includes: adding the current low-performance topology as feedback information to the initial population of the evolutionary algorithm, guiding the evolutionary algorithm to explore new structural spaces in the current solution neighborhood, generating new connection configuration parameters to update the optical switch array and complete the online reconstruction.

11. The multimodal sensing data fusion method based on silicon photonic neural networks according to claim 10, characterized in that, The constructed knowledge entries include three parts: application scenario descriptors, optimal topology parameters, and silicon photonic neural network performance metrics. The scene descriptor includes multimodal perception data types, complexity scores, fusion task requirements, and performance benchmark values; the most relevant historical configurations are retrieved by calculating cosine similarity as the initial configuration for the new task; and the retrieved topological structure parameters are used as the initial genes for the evolutionary algorithm for fine-tuning and optimization.

12. The multimodal sensing data fusion method based on silicon photonic neural networks according to claim 11, characterized in that, The transfer learning process includes: generating an initial population based on the retrieved historical best topology parameters; generating new candidate topologies by applying random perturbations to some connection states; and iteratively optimizing the evolutionary algorithm within a high-starting-point concentrated solution space to reduce the number of iterations required for the search. Continuously store successfully configured knowledge items into the knowledge base to establish a mapping between application requirements and network configurations.

13. A multimodal sensing data fusion system based on silicon photonic neural networks, used to implement the multimodal sensing data fusion method based on silicon photonic neural networks as described in any one of claims 1 to 12, characterized in that, include: The module includes a data acquisition module, a topology configuration module, a topology optimization module, a performance evaluation module, an optimization and reconstruction module, and a migration and expansion module. The data acquisition module is used to acquire multimodal sensing data and extract features from the multimodal sensing data to obtain multimodal sensing data features. The topology configuration module is used to configure the topology of the silicon photonic neural network, create a dynamic interconnection method based on an optical switch array, and adjust the connection relationship between layers of the silicon photonic neural network by controlling the on / off state combination of the optical switch array. The topology optimization module is used to adjust the network topology based on the characteristics of multimodal sensing data and the requirements of fusion tasks, and to automatically search for the optimal network topology using an evolutionary algorithm to generate connection configuration parameters that are suitable for the current task. The performance evaluation module is used to dynamically select processing paths of different depths based on the complexity score and real-time requirements of each modality of sensing data, allocate shallow network paths for simple modalities and deep network paths for complex modalities, and perform differentiated configuration of computing resources for the multimodal sensing data to be fused. The optimization and reconstruction module is used to monitor the performance indicators of multimodal sensing data fusion under the current silicon photonic neural network topology, including computational accuracy, processing latency and computing power overhead. Based on the performance evaluation results, it determines whether to trigger the topology reconstruction process and automatically adjusts the network connection relationship to adapt to the dynamic changes in the fusion task requirements. The transfer extension module is used to record the topological parameters of silicon photonic neural networks under different application scenarios, establish a topological configuration knowledge base, retrieve the similarity of application scenarios to select topological parameters when facing new fusion task requirements, and adapt to new multimodal perception data fusion processing requirements through transfer learning methods.

Citation Information

Patent Citations

  • Optical neural network, data processing method and device based on optical neural network, and storage medium

    CN113408720A

  • Artificial neural network-based silicon photonic interlayer coupling structure design method

    CN119989898A