A neural semantic perception imaging processing chip system based on electron trap array charge accumulation
Patent Information
- Application Number
- CN202610874604.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-17
- Publication Date
- 2026-08-21
AI Technical Summary
这种串行、中断驱动的处理模式难以模拟生物神经系统中神经元群体的全域并行激活与突触强度的时间累积机制,导致语义特征提取的延迟高、能耗大
1、本发明中,将X光平板探测器的物理成像机制抽象并泛化到语义空间,利用电子阱阵列实现全域并行的电荷累积与时间积分。这意味着语义特征一旦被映射为空间坐标上的闪烁信号,即可在硬件层面直接完成多次激活的叠加与扩散,无需在处理器和存储器之间反复搬移中间数据。这种一次性成像式的处理方式,不仅大幅降低了语义感知任务的延迟,还从根本上减少了能耗。
Smart Images

Figure CN122616624A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of artificial intelligence image processing technology, specifically a neuro-semantic perception imaging processing chip system based on charge accumulation of an electron trap array. Background Technology
[0002] Currently, artificial intelligence systems commonly employ processors or graphics processing units (GPUs) based on the von Neumann architecture to process multimodal semantic information such as text, images, and tactile data. In these architectures, the computing and storage units are separated, requiring frequent data transfer between the processor and memory when processing semantic tasks. This serial, interrupt-driven processing mode struggles to simulate the global parallel activation of neuronal populations and the time-accumulation mechanism of synaptic strength in biological neural systems, resulting in high latency and high energy consumption in semantic feature extraction.
[0003] Furthermore, existing neural network acceleration chips typically employ digital multiplication-accumulation units to implement convolution or fully connected operations, which are essentially instruction-driven data flows, failing to achieve spatial topological mapping and diffusion correlation of semantic features at the hardware level. For multimodal semantic fusion, existing solutions mostly use post-processing feature alignment and splicing, lacking the ability to directly superimpose different modal stimulus signals at the physical layer, easily introducing alignment errors and increasing computational overhead. Simultaneously, traditional chips lack native hardware time integration and spatial diffusion mechanisms when handling fuzzy semantics, long-distance semantic dependencies, and hierarchical semantic parsing, often requiring additional software algorithms for compensation, making it difficult to balance processing efficiency and parsing accuracy. Additionally, existing technologies have not yet established a complete mathematical and physical model from semantic input to charge accumulation and feature parsing, leaving the rationality, interpretability, and generalization ability of semantic perception without rigorous theoretical guarantees.
[0004] In summary, there is an urgent need for a novel processor architecture that can simulate biological neural semantic processing mechanisms and realize semantic feature imaging and hierarchical parsing at the hardware level, in order to overcome the shortcomings of existing technologies such as frequent data transfer, poor parallelism, difficulty in multimodal fusion, and lack of theoretical models. Summary of the Invention
[0005] The purpose of this invention is to provide a neuro-semantic perception imaging processing chip system based on charge accumulation of an electron trap array in order to solve the problems mentioned above.
[0006] The technical solution adopted in this invention is as follows: A neuro-semantic perception imaging processing chip system based on charge accumulation of an electron trap array, comprising: a semantic triggering layer, an electron trap array layer (core imaging layer), a sampling readout layer, and a feature parsing layer, wherein: Semantic triggering layer: Establishes a mapping from abstract semantics to physical space coordinates and generates a flashing trigger signal that simulates X-rays.
[0007] Electron trap array layer: The accumulation, spatial diffusion and time integration of semantic charge are achieved through a two-dimensional capacitor array to form a two-dimensional spatial distribution map of semantic features.
[0008] Sampling and readout layer: The capacitor voltage in the electron trap array is quantized by ADC, the data is packaged and the capacitor is automatically cleared, and a semantic charge distribution map is output.
[0009] Feature parsing layer: Extracts first-order and second-order statistical features from the semantic charge distribution map, completes hierarchical semantic recognition of major and minor categories, and outputs the final semantic analysis results.
[0010] In a preferred embodiment, the output terminals of the flicker triggering control module and the acquisition structure configuration module in the semantic triggering layer are respectively connected to the trigger signal input terminal and the mode selection terminal of the charge injection control network in the electron trap array layer. The output of the semantic blink coordinate mapping module is connected to the coordinate input of the row and column addressing network; The voltage output terminal of the two-dimensional capacitor array in the electron trap array layer (via the scan lines) is connected to the analog input terminal of the global sampling control module in the sampling readout layer. The sampling output terminal of the electron trap array layer is connected to the analog input port of the parallel ADC conversion array. The digital output terminal after ADC conversion is connected to the input terminal of the data packaging and transmission module. The semantic charge distribution map data output terminal of the sampling readout layer is connected to the data input terminal of the semantic feature extraction module in the feature parsing layer. The feature output terminal of the sampling readout layer is connected to the input terminal of the hierarchical classification module. Finally, the result output terminal of the hierarchical classification module is connected to the input terminal of the result output and interaction module.
[0011] In a preferred embodiment, the semantic triggering layer includes: a semantic blink coordinate mapping module (maintaining the correspondence between semantic features and array coordinates), a blink triggering control module (generating trigger pulses at specified positions based on semantic input), an acquisition structure configuration module (selecting single-acquisition / multi-acquisition working mode), and an intensity control circuit (adjusting the blink intensity according to semantic importance, corresponding to the charge injection amount). This layer also includes a dynamically adjustable semantic-coordinate mapping table, supporting subsequent feature parsing layers to perform flexible remapping through feedback.
[0012] In a preferred embodiment, the mathematical model of the semantic triggering layer is described by a semantic-coordinate mapping model, and the semantic space mapping is defined by the following formula: Mapping from semantic space S to physical coordinates :
[0013] The relationship between semantic similarity and coordinate distance satisfies a monotonically decreasing constraint. semantic similarity function Relationship with Euclidean distance of coordinates:
[0014] in .
[0015] Multidimensional scaling (MDS) embedding is achieved by minimizing the stress function:
[0016] Semantic distance; Euclidean distance in coordinate space; Weighting coefficient (usually taken as 1).
[0017] In a preferred embodiment, the electron trap array layer is the core imaging medium of the chip, consisting of a large-scale two-dimensional capacitor array, a row and column addressing network, a charge injection circuit, a spatial diffusion network (multi-acquisition mode), and a charge retention mechanism. Each capacitor acts as an electron trap, storing the semantic charge at the corresponding location.
[0018] In a preferred embodiment, the electron trap array layer operates in two acquisition structures: A single flash point corresponds to a single acquisition (direct conversion type): only the capacitor directly opposite the coordinate gains charge through RC charging; A single flash point corresponds to multiple acquisitions (indirect conversion type): the central and surrounding capacitors acquire charges according to a Gaussian distribution, realizing the spatial diffusion and correlation accumulation of semantics. Multiple flashes generate charge superposition, forming a time integration effect.
[0019] In a preferred embodiment, the physical model of the electron trap array layer includes RC charging dynamics and a Gaussian diffusion model, and the parameter definitions and formulas are summarized as follows: The charging differential equation for single-acquisition mode is:
[0020] : The amount of charge stored in the capacitor at any given time; Trigger voltage; Voltage across the capacitor; Charging resistor; Capacitance value; The closed-loop expression for the charging process is:
[0021] Charging time constant; : Natural constant; The formula for calculating the spatial distribution of charge in a single flash under multi-sampling mode is:
[0022] :center The maximum charge at that location; Spatial diffusion coefficient, which controls the range of correlation; Exponential function; The formula for calculating radial attenuation is:
[0023] The formula for calculating the crosstalk ratio when the spacing is d is:
[0024] The formula for calculating the sum of the contributions from each flash during the total charge accumulation process of multiple flashes is as follows:
[0025] Total number of blinks; : No. Weight of each blink (semantic importance); : Acquire the structural response function; Single data collection: (Dirac function); Multiple data collections: (Gaussian kernel); The nonlinear function to prevent overflow during charge saturation nonlinear compression is:
[0026] The maximum amount of charge a capacitor can store.
[0027] In a preferred embodiment, the sampling readout layer is responsible for converting the charge accumulated in the electron trap array into a digital quantity and outputting it. Its internal modules include: The system includes a global sampling control module (which triggers full-frame sampling and freezes the charge state of all capacitors), a parallel ADC conversion array (multi-channel analog-to-digital conversion), a line-by-line scanning circuit (reading capacitor voltages sequentially by line), and an automatic zeroing circuit (discharging capacitors to ground after sampling to prepare for the next round of acquisition). This layer supports point-by-point serial or multi-channel parallel sampling and transmits the packaged semantic charge distribution map to the feature parsing layer via a high-speed bus.
[0028] The formula for calculating the voltage reading of a single pixel is:
[0029] : No. The selection matrix for the next sampling (only) (1) : Capacitor voltage matrix; :Location The voltage; The formula for converting analog voltage to digital code during ADC quantization is as follows:
[0030] Digital output code; ADC reference voltage; : Integer function; The formula for calculating the statistical characteristics of quantization error during the quantization of noise power is as follows:
[0031] Quantization noise variance; The formula for calculating the relationship between signal-to-noise ratio (SNR) and ADC bit depth is as follows:
[0032] : Effective value of signal The formula for calculating the exponential decay law of the zeroing process is:
[0033] Initial charge amount during discharge; Discharge resistor; The percentage of residual charge after zeroing, after zeroing time The formula for calculating the remaining amount is:
[0034] when When the residual charge is < It can be considered as completely zeroed out.
[0035] In a preferred embodiment, the feature parsing layer performs high-level semantic analysis on the sampled semantic charge distribution map and outputs the final understanding result. Its internal modules include: The system comprises a semantic feature extraction module (extracting first- and second-order statistical features from a two-dimensional charge distribution map), a hierarchical classification module (fast identification of broad semantic categories and in-depth analysis of finer classifications), a region selection and scheduling module (identifying semantic regions requiring detailed analysis), and a result output and interaction module. This layer also adjusts the mapping table of the semantic triggering layer based on the analysis results (plastic remapping) and can collaborate with sub-chips to achieve local high-resolution analysis.
[0036] In a preferred embodiment, the mathematical model of the feature parsing layer is based on the statistical features of the semantic charge distribution map and hierarchical perception decision-making. The parameter definitions and formulas are summarized as follows: First-order statistical characteristics (two-dimensional distribution plot Q∈R) M×N ) The formula for calculating the mean is:
[0037] The formula for calculating variance is:
[0038] The skewness expression is:
[0039] The expression for kurtosis is:
[0040]
[0041] The formula for calculating the detection probability of Level 1 acquisition (threshold detection) in the hierarchical classification module is as follows:
[0042] Stimulation intensity Level 1 data collection threshold; Average cumulative charge; : Charge noise standard deviation; Standard normal cumulative distribution function; The formula for calculating the signal-to-noise ratio (SNR) of the secondary acquisition stage in the hierarchical classification module is as follows:
[0043] Second-level integration time; Electron trap time constant; Read out noise; Dark current noise; Average frame rate; The decision variables of the hierarchical switching decision function in the hierarchical classification module are defined as follows:
[0044] First-order cumulative charge; Level 1 threshold; With hysteresis (hysteresis factor) The decision-making rules are as follows:
[0045] The statistical features of the semantic charge distribution map have a one-to-one correspondence with the semantic tendency of the whole text, and the misclassification probability of the Bayesian optimal classifier decreases exponentially. Charge accumulation in multi-collection mode is equivalent to kernel density estimation of semantic frequency. Charge time integral is equivalent to time weighting of semantic importance. Spatial diffusion effect is equivalent to semantic smoothing regularization, which improves generalization ability.
[0046] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are: 1. In this invention, the physical imaging mechanism of an X-ray flat panel detector is abstracted and generalized to a semantic space, utilizing an electron trap array to achieve parallel charge accumulation and time integration across the entire domain. This means that once semantic features are mapped to scintillation signals on spatial coordinates, multiple activations can be superimposed and diffused directly at the hardware level, eliminating the need to repeatedly move intermediate data between the processor and memory. This one-time imaging processing method not only significantly reduces the latency of semantic perception tasks but also fundamentally reduces energy consumption.
[0047] 2. This invention constructs a two-level hierarchical processing architecture, which first extracts the general semantic categories of the whole system and then performs fine-grained analysis on key regions. This coarse-to-fine strategy balances processing efficiency and analysis depth. A complete mathematical and physical model is established, which not only provides quantifiable basis for chip design but also rigorously proves statistically the one-to-one correspondence between charge distribution and semantic categories, and that time integral and spatial diffusion are equivalent to semantic weighting and regularized smoothing, respectively, ensuring the rationality and generalization ability of its output results. Attached Figure Description
[0048] Figure 1 This is a simplified schematic diagram of the structure of the present invention; Figure 2 This is a schematic diagram of the system interface in this invention. Detailed Implementation
[0049] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0050] Example: Reference Figure 1-2 A neuro-semantic perception imaging processing chip system based on charge accumulation in an electron trap array includes: a semantic triggering layer, an electron trap array layer (core imaging layer), a sampling readout layer, and a feature parsing layer, wherein: Semantic triggering layer: Establishes a mapping from abstract semantics to physical space coordinates and generates a flashing trigger signal that simulates X-rays.
[0051] Electron trap array layer: The accumulation, spatial diffusion and time integration of semantic charge are achieved through a two-dimensional capacitor array to form a two-dimensional spatial distribution map of semantic features.
[0052] Sampling and readout layer: The capacitor voltage in the electron trap array is quantized by ADC, the data is packaged and the capacitor is automatically cleared, and a semantic charge distribution map is output.
[0053] Feature parsing layer: Extracts first-order and second-order statistical features from the semantic charge distribution map, completes hierarchical semantic recognition of major and minor categories, and outputs the final semantic analysis results.
[0054] The output terminals of the flash triggering control module and the acquisition structure configuration module in the semantic triggering layer are respectively connected to the trigger signal input terminal and the mode selection terminal of the charge injection control network in the electron trap array layer. The output of the semantic blink coordinate mapping module is connected to the coordinate input of the row and column addressing network; The voltage output terminal of the two-dimensional capacitor array in the electron trap array layer (via the scan lines) is connected to the analog input terminal of the global sampling control module in the sampling readout layer. The sampling output terminal of the electron trap array layer is connected to the analog input port of the parallel ADC conversion array. The digital output terminal after ADC conversion is connected to the input terminal of the data packaging and transmission module. The semantic charge distribution map data output terminal of the sampling readout layer is connected to the data input terminal of the semantic feature extraction module in the feature parsing layer. The feature output terminal of the sampling readout layer is connected to the input terminal of the hierarchical classification module. Finally, the result output terminal of the hierarchical classification module is connected to the input terminal of the result output and interaction module.
[0055] The semantic triggering layer includes: a semantic blink coordinate mapping module (maintaining the correspondence between semantic features and array coordinates), a blink trigger control module (generating trigger pulses at specified positions based on semantic input), an acquisition structure configuration module (selecting single-acquisition / multi-acquisition working mode), and an intensity control circuit (adjusting the blink intensity according to semantic importance, corresponding to the charge injection amount). This layer also contains a dynamically adjustable semantic-coordinate mapping table, supporting subsequent feature parsing layers to perform plastic remapping through feedback.
[0056] The mathematical model of the semantic triggering layer is described by a semantic-coordinate mapping model, and the semantic space mapping is defined by the following formula: Mapping from semantic space S to physical coordinates :
[0057] The relationship between semantic similarity and coordinate distance satisfies a monotonically decreasing constraint. semantic similarity function Relationship with Euclidean distance of coordinates:
[0058] in .
[0059] Multidimensional scaling (MDS) embedding is achieved by minimizing the stress function:
[0060] Semantic distance; Euclidean distance in coordinate space; Weighting coefficient (usually taken as 1).
[0061] The electron trap array layer is the core imaging medium of the chip, consisting of a large-scale two-dimensional capacitor array, a row and column addressing network, a charge injection circuit, a spatial diffusion network (multi-acquisition mode), and a charge retention mechanism. Each capacitor acts as an electron trap, storing the semantic charge at the corresponding location.
[0062] The electron trap array layer operates with two acquisition structures: A single flash point corresponds to a single acquisition (direct conversion type): only the capacitor directly opposite the coordinate gains charge through RC charging; A single flash point corresponds to multiple acquisitions (indirect conversion type): the central and surrounding capacitors acquire charges according to a Gaussian distribution, realizing the spatial diffusion and correlation accumulation of semantics. Multiple flashes generate charge superposition, forming a time integration effect.
[0063] The physical model of the electron trap array layer includes RC charging dynamics and a Gaussian diffusion model. The parameter definitions and formulas are summarized as follows: The charging differential equation for single-acquisition mode is:
[0064] : The amount of charge stored in the capacitor at any given time; Trigger voltage; Voltage across the capacitor; Charging resistor; Capacitance value; The closed-loop expression for the charging process is:
[0065] Charging time constant; : Natural constant; The formula for calculating the spatial distribution of charge in a single flash under multi-sampling mode is:
[0066] :center The maximum charge at that location; Spatial diffusion coefficient, which controls the range of correlation; Exponential function; The formula for calculating radial attenuation is:
[0067] The formula for calculating the crosstalk ratio when the spacing is d is:
[0068] The formula for calculating the sum of the contributions from each flash during the total charge accumulation process of multiple flashes is as follows:
[0069] Total number of blinks; : No. Weight of each blink (semantic importance); : Acquire the structural response function; Single data collection: (Dirac function); Multiple data collections: (Gaussian kernel); The nonlinear function to prevent overflow during charge saturation nonlinear compression is:
[0070] The maximum amount of charge a capacitor can store.
[0071] The sampling readout layer is responsible for converting the accumulated charge in the electron trap array into digital quantities and outputting them. Its internal modules include: The system includes a global sampling control module (which triggers full-frame sampling and freezes the charge state of all capacitors), a parallel ADC conversion array (multi-channel analog-to-digital conversion), a line-by-line scanning circuit (reading capacitor voltages sequentially by line), and an automatic zeroing circuit (discharging capacitors to ground after sampling to prepare for the next round of acquisition). This layer supports point-by-point serial or multi-channel parallel sampling and transmits the packaged semantic charge distribution map to the feature parsing layer via a high-speed bus.
[0072] The formula for calculating the voltage reading of a single pixel is:
[0073] : No. The selection matrix for the next sampling (only) (1) : Capacitor voltage matrix; :Location The voltage; The formula for converting analog voltage to digital code during ADC quantization is as follows:
[0074] Digital output code; ADC reference voltage; : Integer function; The formula for calculating the statistical characteristics of quantization error during the quantization of noise power is as follows:
[0075] Quantization noise variance; The formula for calculating the relationship between signal-to-noise ratio (SNR) and ADC bit depth is as follows:
[0076] : Effective value of signal The formula for calculating the exponential decay law of the zeroing process is:
[0077] Initial charge amount during discharge; Discharge resistor; The percentage of residual charge after zeroing, after zeroing time The formula for calculating the remaining amount is:
[0078] when When the residual charge is < It can be considered as completely zeroed out. The feature parsing layer performs high-level semantic analysis on the sampled semantic charge distribution map and outputs the final understanding result. Its internal modules include: The system comprises a semantic feature extraction module (extracting first- and second-order statistical features from a two-dimensional charge distribution map), a hierarchical classification module (fast identification of broad semantic categories and in-depth analysis of finer classifications), a region selection and scheduling module (identifying semantic regions requiring detailed analysis), and a result output and interaction module. This layer also adjusts the mapping table of the semantic triggering layer based on the analysis results (plastic remapping) and can collaborate with sub-chips to achieve local high-resolution analysis.
[0079] The mathematical model of the feature parsing layer is based on the statistical features of the semantic charge distribution map and hierarchical perception decision-making. The parameter definitions and formulas are summarized as follows: 1. First-order statistical characteristics (two-dimensional distribution plot Q∈R) M×N ) The formula for calculating the mean is:
[0080] The formula for calculating variance is:
[0081] The skewness expression is:
[0082] The expression for kurtosis is:
[0083]
[0084] The formula for calculating the detection probability of Level 1 acquisition (threshold detection) in the hierarchical classification module is as follows:
[0085] Stimulation intensity Level 1 data collection threshold; Average cumulative charge; : Charge noise standard deviation; Standard normal cumulative distribution function; The formula for calculating the signal-to-noise ratio (SNR) of the secondary acquisition stage in the hierarchical classification module is as follows:
[0086] Second-level integration time; Electron trap time constant; Read out noise; Dark current noise; Average frame rate; The decision variables of the hierarchical switching decision function in the hierarchical classification module are defined as follows:
[0087] First-order cumulative charge; Level 1 threshold; With hysteresis (hysteresis factor) The decision-making rules are as follows:
[0088] The statistical features of the semantic charge distribution map have a one-to-one correspondence with the semantic tendency of the whole text, and the misclassification probability of the Bayesian optimal classifier decreases exponentially. Charge accumulation in multi-collection mode is equivalent to kernel density estimation of semantic frequency. Charge time integral is equivalent to time weighting of semantic importance. Spatial diffusion effect is equivalent to semantic smoothing regularization, which improves generalization ability.
[0089] Example 1: Large-scale text semantic perception and sentiment analysis: Application scenario: Real-time semantic understanding of long texts (such as news comments and social media streams) to extract the sentiment and theme distribution of the entire text.
[0090] Implementation: The text input stream undergoes semantic preprocessing, with each keyword mapped to fixed coordinates on an electron trap array (coordinates of semantically similar words are adjacent). A multi-sampling mode is employed (Gaussian diffusion coefficient σ = 2 pixels), with each flash corresponding to the semantic activation of a word, and word frequency used as the number of flashes. After full-text reading, the sampling readout layer outputs a semantic charge distribution map, and the feature parsing layer calculates the mean, skewness, and kurtosis, and determines the overall sentiment tendency (positive / negative / neutral) based on the semantic centroid.
[0091] Experimental data (compared to traditional von Neumann architecture CPUs):
[0092] Beneficial effects: Processing latency is reduced by 80% to 90%, thanks to global parallel charge accumulation and single full-frame readout.
[0093] Energy consumption is significantly reduced (only 12% to 18% of that of traditional CPUs), eliminating the need for frequent data transfer.
[0094] The semantic classification accuracy improved by 0.4 to 1.1 percentage points, which is due to the natural weighting of important semantics by charge time accumulation.
[0095] Example 2: Multimodal (text + visual) semantic fusion perception: Application scenario: Simultaneously process images (object recognition) and associated text descriptions (scene understanding) to achieve cross-modal semantic association analysis.
[0096] Implementation: Visual modality: The image input is processed by a CNN to extract feature maps. The activation intensity of each channel in the feature map is mapped to the flicker intensity of the corresponding coordinate. Text modality: Same as in Example 1. The flicker events of both modalities act on the same electron trap array, satisfying the principle of linear superposition of multimodal stimulus fields (Equation 7-4). The final semantic charge distribution map shows cross-modal semantic overlap regions (e.g., the coordinates of the "dog" image and the "dog" text are superimposed in proximity). The feature parsing layer extracts the covariance matrix of the charge accumulation region to quantify the cross-modal correlation.
[0097] Experimental data (multimodal fusion task, compared with the method of processing single modalities independently before fusion):
[0098] Beneficial effects: Cross-modal association detection accuracy is higher than that of later fusion (because it is directly superimposed at the physical level, avoiding feature alignment error).
[0099] The processing latency is only 30% to 45% of that of traditional discrete fusion.
[0100] No separate storage space needs to be maintained for different modes, resulting in bandwidth savings of over 55%.
[0101] Example 3: Hierarchical semantic processing and sub-chip collaboration (high-complexity semantic parsing): Application scenario: Processing long documents (such as technical papers and legal contracts) containing multiple topics, numerous entities, and complex logical relationships. It requires quickly locating the core topic areas first, followed by detailed semantic analysis of key paragraphs.
[0102] Implementation method: Level 1 (Major Semantic Layer): A single Class A main chip uses low-resolution (downsampling) global imaging to quickly identify three main semantic clusters (such as "hardware architecture", "mathematical model", and "experimental results"). Based on the charge density and area of each cluster, the main chip allocates three Class A sub-chips to process the ROI of each cluster.
[0103] Level 2 (Fine-grained semantic layer): Each sub-chip employs a high-resolution multi-acquisition mode (σ=0.5 pixels) and undergoes Wobble semantic dynamic calibration from the Class B chip (adjustment coefficient k=0.7) to correct local coordinate offsets. Finally, the main chip fuses the sub-chip results to output complete fine-grained semantic parsing.
[0104] Experimental data (compared to the same document for a single chip without hierarchical processing):
[0105] Beneficial effects: Hierarchical processing reduces the overall processing time by 72% to 80%, and avoids serial resource contention of a single chip by using parallel sub-chips.
[0106] wobble semantic calibration improves fine-grained recognition accuracy by 7%–8%, validating the effectiveness of dynamic alignment error correction.
[0107] The dynamic resource scheduling mechanism enables highly complex documents to maintain real-time performance (7 clusters in just 136 ms).
[0108] Based on the above embodiments and experimental data, this system (Class A neural CCD chip) has the following significant advantages compared to the traditional von Neumann architecture and single-level semantic processor: Parallelism and low latency: The global electron trap array accumulates in parallel, and global semantic extraction can be completed in a single frame readout, reducing latency by more than 80%.
[0109] High energy efficiency: No frequent data transfer is required, and the power consumption of charge integration and ADC conversion is much lower than that of CPU / GPU instruction execution, reducing energy consumption to less than 1 / 5.
[0110] Multimodal fusion is natural: the stimulus fields of different modalities are linearly superimposed at the physical layer, avoiding the complex calculations and accuracy loss of later feature alignment.
[0111] Hierarchical processing and scalability: The main and sub-chips work together, along with Wobble dynamic calibration, to support highly complex semantic scenarios. The recognition accuracy and processing time are both superior to single-chip solutions.
[0112] It conforms to the biological semantic processing mechanism: characteristics such as group coding, spatial topological mapping, temporal accumulation and spatial diffusion are naturally embedded in hardware, and it is robust to fuzzy semantics.
[0113] As described above, this invention abstracts and generalizes the physical imaging mechanism of X-ray flat panel detectors to the semantic space, utilizing an electron trap array to achieve parallel charge accumulation and time integration across the entire domain. This means that once semantic features are mapped to scintillation signals on spatial coordinates, multiple activations can be superimposed and diffused directly at the hardware level, eliminating the need to repeatedly move intermediate data between the processor and memory. This one-time imaging processing method not only significantly reduces the latency of semantic perception tasks but also fundamentally reduces energy consumption.
[0114] This invention constructs a two-level hierarchical processing architecture, which first extracts the general semantic categories of the whole system and then performs fine-grained analysis on key regions. This coarse-to-fine strategy balances processing efficiency and analysis depth. A complete mathematical and physical model is established, which not only provides quantifiable basis for chip design but also rigorously proves statistically the one-to-one correspondence between charge distribution and semantic categories, and that time integral and spatial diffusion are equivalent to semantic weighting and regularized smoothing, respectively, ensuring the rationality and generalization ability of the output results.
[0115] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the term "comprising" or any other variations thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0116] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A neuro-semantic perception imaging processing chip system based on charge accumulation in an electron trap array, characterized in that: include: The semantic triggering layer, the electronic trap array layer, the sampling readout layer, and the feature parsing layer are as follows: Semantic triggering layer: Establishes a mapping from abstract semantics to physical space coordinates and generates a flashing trigger signal that simulates X-rays; Electron trap array layer: The accumulation, spatial diffusion and time integration of semantic charge are achieved through a two-dimensional capacitor array to form a two-dimensional spatial distribution map of semantic features; Sampling and readout layer: The capacitor voltage in the electron trap array is quantized by ADC, the data is packaged and the capacitor is automatically cleared to zero, and a semantic charge distribution map is output. Feature parsing layer: Extracts first-order and second-order statistical features from the semantic charge distribution map, completes hierarchical semantic recognition of major and minor categories, and outputs the final semantic analysis results.
2. The neuro-semantic perception imaging processing chip system based on charge accumulation of an electron trap array as described in claim 1, characterized in that: The outputs of the flicker trigger control module and the acquisition structure configuration module in the semantic trigger layer are respectively connected to the trigger signal input and mode selection terminal of the charge injection control network in the electron trap array layer. The output of the semantic blink coordinate mapping module is connected to the coordinate input of the row and column addressing network; The voltage output terminal of the two-dimensional capacitor array in the electron trap array layer is connected to the analog input terminal of the global sampling control module in the sampling readout layer. The sampling output terminal of the electron trap array layer is connected to the analog input port of the parallel ADC conversion array. The digital output terminal after ADC conversion is connected to the input terminal of the data packaging and transmission module. The semantic charge distribution map data output terminal of the sampling readout layer is connected to the data input terminal of the semantic feature extraction module in the feature parsing layer. The feature output terminal of the sampling readout layer is connected to the input terminal of the hierarchical classification module. Finally, the result output terminal of the hierarchical classification module is connected to the input terminal of the result output and interaction module.
3. The neuro-semantic perception imaging processing chip system based on charge accumulation of an electron trap array as described in claim 1, characterized in that: The semantic triggering layer includes: a semantic blink coordinate mapping module, a blink triggering control module, an acquisition structure configuration module, and an intensity control circuit; this layer also includes a dynamically adjustable semantic-coordinate mapping table, which supports subsequent feature parsing layers to perform plastic remapping through feedback.
4. The neuro-semantic perception imaging processing chip system based on charge accumulation of an electron trap array as described in claim 1, characterized in that: The mathematical model of the semantic triggering layer is described by a semantic-coordinate mapping model, and the semantic space mapping is defined by the following formula: Mapping from semantic space S to physical coordinates : The relationship between semantic similarity and coordinate distance satisfies a monotonically decreasing constraint. semantic similarity function Relationship with Euclidean distance of coordinates: in ; Multidimensional scaling embedding is achieved by minimizing the stress function: Semantic distance; Euclidean distance in coordinate space; Weighting coefficient.
5. The neuro-semantic perception imaging processing chip system based on charge accumulation of an electron trap array as described in claim 1, characterized in that: The electron trap array layer is the core imaging medium of the chip, consisting of a large-scale two-dimensional capacitor array, a row and column addressing network, a charge injection circuit, a spatial diffusion network, and a charge retention mechanism; each capacitor acts as an electron trap, storing the semantic charge at the corresponding position.
6. The neuro-semantic perception imaging processing chip system based on charge accumulation of an electron trap array as described in claim 1, characterized in that: The electron trap array layer operates with two acquisition structures; A single flash point corresponds to a single acquisition: only the capacitor directly opposite the coordinate gains charge through RC charging; A single flash point corresponds to multiple acquisitions: the central and surrounding capacitors obtain charges according to a Gaussian distribution, realizing the spatial diffusion and association accumulation of semantics; multiple flashes generate charge superposition, forming a time integration effect.
7. The neuro-semantic perception imaging processing chip system based on charge accumulation of an electron trap array as described in claim 1, characterized in that: The physical model of the electron trap array layer includes RC charging dynamics and a Gaussian diffusion model. The parameter definitions and formulas are summarized as follows: The charging differential equation for single-acquisition mode is: : The amount of charge stored in the capacitor at any given time; Trigger voltage; Voltage across the capacitor; Charging resistor; Capacitance value; The closed-loop expression for the charging process is: Charging time constant; : Natural constant; The formula for calculating the spatial distribution of charge in a single flash under multi-sampling mode is: :center The maximum charge at that location; Spatial diffusion coefficient, which controls the range of correlation; Exponential function; The formula for calculating radial attenuation is: The formula for calculating the crosstalk ratio when the spacing is d is: The formula for calculating the sum of the contributions from each flash during the total charge accumulation process of multiple flashes is as follows: Total number of blinks; : No. The weight of each flash; : Acquire the structural response function; Single data collection: ; Multiple data collections: ; The nonlinear function to prevent overflow during charge saturation nonlinear compression is: The maximum amount of charge a capacitor can store.
8. The neuro-semantic perception imaging processing chip system based on charge accumulation of an electron trap array as described in claim 1, characterized in that: The sampling readout layer is responsible for converting the charge accumulated in the electron trap array into digital quantities and outputting them; Its internal modules include: The layer includes a global sampling control module, a parallel ADC conversion array, a line-by-line scanning circuit, and an automatic zeroing circuit. This layer supports point-by-point serial or multi-channel parallel sampling and transmits the packaged semantic charge distribution map to the feature parsing layer via a high-speed bus. The formula for calculating the voltage reading of a single pixel is: : No. The selection matrix for the next sampling; : Capacitor voltage matrix; :Location The voltage; The formula for converting analog voltage to digital code during ADC quantization is as follows: Digital output code; ADC reference voltage; : Integer function; The formula for calculating the statistical characteristics of quantization error during the quantization of noise power is as follows: Quantization noise variance; The formula for calculating the relationship between signal-to-noise ratio and ADC bit depth is: : Effective value of signal The formula for calculating the exponential decay law of the zeroing process is: Initial charge amount during discharge; Discharge resistor; The percentage of residual charge after zeroing, after zeroing time The formula for calculating the remaining amount is: when When the residual charge is < They believed it was completely reset to zero.
9. The neuro-semantic perception imaging processing chip system based on charge accumulation of an electron trap array as described in claim 1, characterized in that: The feature parsing layer performs high-level semantic analysis on the sampled semantic charge distribution map and outputs the final understanding result. Its internal modules include: The system includes a semantic feature extraction module, a hierarchical classification module, a region selection and scheduling module, and a result output and interaction module. This layer is also responsible for adjusting the mapping table of the semantic triggering layer based on the feedback of the parsing results, and can work with the sub-chip to achieve local high-resolution parsing.
10. The neuro-semantic perception imaging processing chip system based on charge accumulation of an electron trap array as described in claim 1, characterized in that: The mathematical model of the feature parsing layer is based on the statistical features of the semantic charge distribution map and hierarchical perception decision-making. The parameter definitions and formulas are summarized as follows: First-order statistical characteristics The formula for calculating the mean is: The formula for calculating variance is: The skewness expression is: The expression for kurtosis is: The formula for calculating the probability of primary acquisition and detection in the hierarchical classification module is as follows: Stimulation intensity Level 1 data collection threshold; Average cumulative charge; : Charge noise standard deviation; Standard normal cumulative distribution function; The formula for calculating the signal-to-noise ratio (SNR) of the secondary acquisition stage in the hierarchical classification module is as follows: Second-level integration time; Electron trap time constant; Read out noise; Dark current noise; Average frame rate; The decision variables of the hierarchical switching decision function in the hierarchical classification module are defined as follows: First-order cumulative charge; Level 1 threshold; The decision-making rules with hysteresis are as follows: The statistical features of the semantic charge distribution map have a one-to-one correspondence with the semantic tendency of the whole text, and the misclassification probability of the Bayesian optimal classifier decays exponentially; the charge accumulation of the multi-collection mode is equivalent to the kernel density estimation of semantic frequency.