Spatial information extractor based on graphene / silicon heterojunction and motion sensing method
Through a spatial information extractor based on graphene/silicon heterojunction, the carrier spatial distribution detection method is used to solve the problem of insufficient spatial correlation and information acquisition capabilities of traditional image sensors when processing space-time features, and efficient target recognition and action perception are achieved, with an identification accuracy of over 97%.
Patent Information
- Application Number
- CN202510221843.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-05-30
AI Technical Summary
Traditional image sensors have problems with insufficient spatial correlation and information acquisition capabilities when processing space-time features, and cannot effectively utilize carrier dynamic characteristics for target recognition and action perception.
A spatial information extractor based on graphene/silicon heterojunction is adopted, through carrier spatial distribution detection, the lateral transmission path formed by graphene film and P-type surface layer is used, and the carrier concentration distribution is detected in combination with an annular electrode to achieve target recognition and action perception.
It significantly improves the accuracy and efficiency of target recognition, reduces data processing requirements, and has an identification accuracy of more than 97%, significantly reducing the burden of back-end data processing.
Smart Images

Figure CN120076472A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of optoelectronic devices, and particularly relates to a spatial information extractor based on a graphene / silicon heterojunction and a method for motion perception. Background Art
[0002] With the wide application of machine vision in fields such as sports science, robotics, biomechanics, virtual reality, and surveillance, the modeling, tracking, and feature extraction of target actions have become particularly important. However, traditional Complementary Metal Oxide Semiconductor (CMOS) image sensors have significant limitations in processing spatio-temporal features. These sensors typically output data in the form of pixelated frames or continuous optical flow, requiring complex algorithms to interpret spatial relationships. This method relies on high-resolution images or videos, inevitably involving multiple conversions, transmissions, and post-processings of redundant data, which not only consumes a large amount of bandwidth and computing resources but also limits the realization of real-time perception. In a material system, the movements of photo-carriers such as diffusion, drift, and recombination occur within the plane perpendicular to the incident light and are affected by non-uniform illumination and material properties. These dynamic behaviors form spatially correlated distribution features that can be used to directly process and perceive geometric and position information inside the sensor. However, traditional image sensors lack inter-pixel correlation and cannot utilize this carrier dynamic characteristic. Summary of the Invention
[0003] Object of the Invention: To solve the problems of insufficient spatial correlation and information acquisition ability of traditional image sensors, the present invention proposes a spatial information extractor based on a graphene / silicon heterojunction and a method for motion perception, which uses the spatial characteristics of carriers and a Spatial Information Extractor (SIE) to identify a target light source, and improves the accuracy and efficiency of target recognition through a carrier spatial distribution detection method, while significantly reducing the data processing requirements.
[0004] Technical Solution: A spatial information extractor based on a graphene / silicon heterojunction, comprising: a photosensitive region, a ring electrode, and a semiconductor substrate; the photosensitive region is composed of a P-type surface layer and a graphene thin film, the graphene thin film is disposed on the semiconductor substrate, and the semiconductor substrate is an N-type silicon substrate; the ring electrodes are uniformly distributed around the photosensitive region; the P-type surface layer and the graphene thin film form a lateral transmission path for photo-carriers.
[0005] Further, the graphene thin film is a film grown by chemical vapor deposition and transferred to the semiconductor substrate by a wet transfer method.
[0006] The present invention provides a method for identifying and motion sensing of a single-frame target, comprising the following steps:
[0007] Step 1: Construct a spatial information extractor, which is a spatial information extractor based on a graphene / silicon heterojunction.
[0008] Step 2: Arrange an LED light source at multiple key points of the target to ensure that each LED light source can be detected individually.
[0009] Step 3: Deploy the spatial information extractor on one side of the target, project the optical image of the target onto the photosensitive area of the spatial information extractor, generate excited light-generated carriers associated with the optical image of the target in the photosensitive area of the spatial information extractor, and then form superimposed carriers.
[0010] Step 4: Detect the concentration distribution of the superimposed carriers through a ring electrode to obtain the voltage associated with the target and its distribution.
[0011] Step 5: Map the voltage associated with the target into polar coordinates, form a closed polygon in the polar coordinates according to the voltage distribution associated with the target; quantitatively analyze the geometric features of the closed polygon, extract trend-related parameters and shape-related parameters from it, normalize each parameter to obtain a trend factor and a shape factor; realize the identification of the single-frame target and the sensing of the target motion according to the trend factor and the shape factor.
[0012] Further, the trend-related parameters include an angle parameter and a length parameter; the shape-related parameters include the average value of the light response intensity.
[0013] The present invention provides a method for identifying and motion sensing of a continuous target, comprising the following steps:
[0014] Step 1: Construct a spatial information extractor, which is a spatial information extractor based on a graphene / silicon heterojunction disclosed above.
[0015] Step 2: Arrange an LED light source at multiple key points of the target to ensure that each LED light source can be detected individually.
[0016] Step 3: Deploy the spatial information extractor on one side of the target, project the optical image of the target onto the photosensitive area of the spatial information extractor, generate excited light-generated carriers associated with the optical image of the target in the photosensitive area of the spatial information extractor, and then form superimposed carriers.
[0017] Step 4: Detect the concentration distribution of the superimposed carriers through a ring electrode to obtain the voltage associated with the target and its distribution.
[0018] Step 5: Map the voltage associated with the target into polar coordinates. According to the voltage distribution associated with the target, form a closed polygon in polar coordinates; quantitatively analyze the geometric features of the closed polygon, extract the parameters related to the trend and the parameters related to the shape therefrom, and perform normalization processing on each parameter to obtain the trend factor and the shape factor of the optical image of the single-frame target.
[0019] Step 6: Take the trend factors and shape factors of the optical images of multiple consecutive frames of the target connected end to end as an action. For any action, use each of its trend factors and shape factors as the input of the trained recurrent neural network to achieve target recognition and action perception.
[0020] Further, the parameters related to the trend include an angle parameter and a length parameter; the parameters related to the shape include the average value of the light response intensity.
[0021] Beneficial effects: Compared with the prior art, the present invention has the following advantages:
[0022] (1) Based on the dynamics of photo-generated carriers in the material, the spatial information extractor based on the graphene / silicon heterojunction proposed by the present invention adopts a large-area graphene / silicon heterostructure, combines the lateral photovoltaic effect, and utilizes the carrier dynamics characteristics to directly generate optoelectronic signals dependent on the object shape, realizing efficient spatio-temporal feature extraction.
[0023] (2) Compared with traditional pixelated sensors, the spatial information extractor based on the graphene / silicon heterojunction proposed by the present invention can not only significantly reduce the data volume by more than three orders of magnitude and the recognition accuracy exceeds 97%, but also significantly reduce the burden of backend data processing, laying a foundation for the efficient processing of high-dimensional spatio-temporal data and future energy-saving real-time intelligent vision technology. Description of the Drawings
[0024] Figure 1 It is a schematic diagram of the spatial information extractor SIE proposed by the present invention;
[0025] Figure 2 It is a schematic flowchart of a non-pixelated target recognition method based on the spatial distribution of carriers proposed in Embodiment 1;
[0026] Figure 3 It is a schematic diagram of the time-resolved carrier distribution characteristics of a single light spot by the spatial information extractor SIE in Embodiment 1;
[0027] Figure 4 It is a schematic diagram of the light-responsive carrier distribution of the spatial information extractor SIE for the target action in Embodiment 1;
[0028] Figure 5Schematic diagram of data processing of the spatial information extractor SIE in Embodiment 1;
[0029] Figure 6 Schematic diagram of the recognition accuracy rate in Embodiment 1 and comparison with traditional detectors;
[0030] Figure 7 Schematic diagram of training a model using a recurrent neural network in Embodiment 2. Detailed implementation manners
[0031] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will further describe a non-pixelated target recognition method based on the spatial distribution of carriers proposed by the present invention in conjunction with the accompanying drawings and embodiments.
[0032] Embodiment 1:
[0033] As Figure 1 shown, this embodiment proposes a spatial information extractor SIE based on a graphene / silicon heterojunction, which is composed of a photosensitive region 1, a ring electrode 2, and a semiconductor substrate 3; wherein, the photosensitive region 1 includes a graphene thin film and a lightly doped P-type surface layer, and the lightly doped P-type surface layer and the graphene thin film cover the surface of the semiconductor substrate 3 to form a lateral transmission path for photo-generated carriers. The graphene thin film used in this embodiment is a high-mobility thin film grown by chemical vapor deposition and transferred to the substrate by a wet transfer method. The semiconductor substrate 3 used in this embodiment is a lightly doped N-type silicon substrate. The ring electrode 2 is composed of 12 electrodes evenly distributed in a circle along the photosensitive region 1, and the angular interval of each electrode is 30 degrees, having the ability to detect the distribution difference of photo-generated carriers.
[0034] The spatial information extractor SIE proposed in this embodiment is based on a graphene / silicon heterostructure, uses a non-pixelated lateral photovoltaic effect to extract object geometric features, generates position and shape information through position-related photovoltaic signals, realizes spatio-temporal sequence perception of actions, and has a high response speed, with a photoelectric response time less than 10 microseconds.
[0035] Embodiment 2:
[0036] This embodiment proposes a method for recognizing and perceiving actions of a single-frame target, which mainly includes the following steps:
[0037] Step 1: Construct a spatial information extractor SIE based on a graphene / silicon heterojunction proposed in Embodiment 1 and use it as the core sensing device;
[0038] Step 2: Prepare or select multiple high-frequency controllable LED light sources and arrange multiple high-frequency controllable LED light sources at the key points of the target action to ensure that each LED light source can be detected separately;
[0039] Step 3: As shown in Figure 2 , project the optical image 4 of the target to be measured (i.e., the real human action) into the photosensitive area 1 of the spatial information extractor SIE 5, and generate excited photo-generated carriers 6 associated with the optical image 4 of the target to be measured within the photosensitive area 1; after a period of time, the excited photo-generated carriers 6 gradually reach the equilibrium state through the carrier transport process, forming a superimposed carrier concentration distribution 7; Figure 3 shows the change process of the carrier distribution with time evolution of the spatial information extractor SIE for a single light spot. As time extends, the carriers are transported to the surrounding areas, that is, the distribution area of the photo-generated carriers expands, so it can be used as the basis for detecting the spatial distribution of carriers.
[0040] The specific operation is as follows: project a multi-point light source onto the photosensitive area of the spatial information extractor SIE to generate a superimposed photo-generated carrier distribution. The spatial distribution of carriers under steady-state illumination of a single-point light source shows an atypical Gaussian characteristic, affected by the combined effects of diffusion and drift. The diffusion effect is driven by the carrier concentration gradient at the center of the light spot, and the drift effect is induced by the transverse electric field of the graphene layer. The central region of the carrier distribution shows a higher concentration, while the edge region gradually decays due to the combined action of diffusion and drift. In the global exposure mode, affected by the target geometry, the illumination of the multi-point light source will generate a superimposed photo-generated carrier distribution, and this superimposed distribution effectively correlates the optical information in different regions, providing a basis for geometric feature extraction and enhancing the sensor's ability to process and integrate spatial information.
[0041] Therefore, the above carrier spatial distribution satisfies the following relationship:
[0042]
[0043] x c = μEt
[0044] w 2 = 2Dt
[0045]
[0046] where E is the transverse electric field strength, D is the carrier diffusion coefficient, N is the original carrier concentration in the substrate under non-illumination conditions, and τ is the carrier lifetime.
[0047] Taking the jumping action as an example, Figure 4 shows the detection result of the carrier distribution of the SIE for a specific action. When the target makes a jumping action, the LED attached to the key point generates a target light intensity distribution (as shown in the left figure of Figure 4 ), and at this time, a corresponding target carrier distribution space is generated in the photosensitive area of the SIE (as shown in Figure 4as shown in the right figure in). It can be seen that the optical image features of the human motion are significantly consistent with the carrier spatial distribution in the SIE photosensitive region, especially in terms of the motion trend and the shape and size.
[0048] Step 4: As Figure 5 shown, analyzing and processing the voltage distribution associated with the human motion characteristics can achieve the recognition and perception of different types of motions. The specific analysis method is as follows:
[0049] The electrical signal detected by the circular electrode 2 is positively correlated with the carrier concentration near it. Representing the electrical signals collected by the 12 electrodes in polar coordinates can form a closed shape information polygon (Spatial information polygon, SIP) (as Figure 5 shown in the left figure in), to represent the features of the target's direction (upper and lower limbs), centroid (upper left, lower left, upper right, and lower right), and geometric shape (distance from each vertex to the origin). Quantitative analysis of the geometric features of the SIP is performed through the following two sets of eigenvalues (as Figure 5 shown in the right figure in):
[0050] Extract the eigenvalues related to the trend, including the angle parameter α and the length parameter D:
[0051] α 1 = ∠upper left - origin - upper right
[0052] α 2 = ∠lower left - origin - lower right
[0053] D 1 = L 1 - L 2
[0054] D 2 = L 3 - L 4
[0055] Extract the eigenvalues related to the shape, including:
[0056] Average value of the upper limb vertices:
[0057] Average value of the lower limb vertices:
[0058] Adopt a generalization algorithm such as t-SNE for dimensionality reduction processing, and transform the multi-dimensional trend-related parameters (α 1 , α 2 , D 1 , D 2 ) and shape parameters (Avg 1 , Avg 2Normalize to a trend factor (a, D) and a shape factor V ph There are two numerical values. Therefore, according to the trend factor and the shape factor, different actions are displayed and distinguished in a two-dimensional plane, as shown in Figure 5 the right figure in
[0059] According to the above process, the object recognition and action perception of single-frame images of 6 different actions (standing, jumping, running, walking, kicking, throwing) can be realized.
[0060] As shown in Figure 6 the figure, SIE has a significant advantage in the amount of information processing data compared with the traditional image recognition method, and the recognition accuracy is quite the same.
[0061] Embodiment 3:
[0062] This embodiment proposes a method for recognizing and perceiving actions of continuous objects. Different from Embodiment 2, the object to be measured changes from a single-frame image to a continuous action sequence. The action recognition process of the object to be measured is similar to that in Embodiment 2, but in this embodiment, a recurrent neural network is introduced to analyze the temporal sequence of the action sequence.
[0063] As shown in Figure 7 the figure, each action consists of 5 consecutive key-frame images connected end to end. The recurrent neural network (as shown in the left figure in Figure 7 the figure) supplements the sequential relationship between frames in the action sequence as the basis for training and judgment, and a training model for continuous object recognition and action perception can be obtained. The specific steps are as follows:
[0064] Map the voltage distribution associated with the characteristics of the human action to a shape information polygon SIP in the polar coordinate system for analyzing the direction, centroid, and geometric shape characteristics of the object.
[0065] Extract trend-related parameters, including the angle parameter α and the length parameter L.
[0066] Extract shape-related parameters, including the average value Avg of the light response intensity.
[0067] Adopt a generalization algorithm such as t-SNE for dimensionality reduction processing to obtain a trend factor and a shape factor as the input parameter f(t) of the first frame of the recurrent neural network.
[0068] Perform the above steps on the next frame, input the second frame f(t + 1), until 5 frames f(t + 4). During the training process, RNN will perform weight iteration in the hidden layers h(t) to h(t + 4) based on the geometric features (trend, shape parameters) of the object and the temporal features (t, t + 1,..., t + 4) of the input parameter sequence, and finally output the final recognition result m(t + 4) at the 5th frame.
[0069] Use the trained recurrent neural network for object recognition and action perception, taking the trend factor and shape factor of each frame as the input of the RNN. Finally, generate the final action sequence classification based on the input signal and recognition result.
[0070] According to the above process, the object recognition and action perception of continuous action sequence images of 6 different actions (standing, jumping, running, walking, kicking, throwing) can be realized.
Claims
1. A spatial information extractor based on graphene / silicon heterojunction, characterized in that: include: A photosensitive region, a ring electrode and a semiconductor substrate; The photosensitive area is composed of a P-type surface layer and a graphene film, wherein the graphene film is arranged on a semiconductor substrate, and the semiconductor substrate is an N-type silicon substrate; the annular electrodes are evenly distributed around the photosensitive area; the P-type surface layer and the graphene film form a lateral transmission path for light carriers.
2. A spatial information extractor based on graphene / silicon heterojunction according to claim 1, characterized in that: The graphene film is grown by chemical vapor deposition and is transferred to a semiconductor substrate by a wet method.
3. A method for identifying and sensing a single-frame target, characterized in that: The following steps are involved: Step 1: construct a spatial information extractor, wherein the spatial information extractor is a spatial information extractor based on a graphene / silicon heterojunction as claimed in claim 1 or 2; Step 2: Arrange an LED light source at multiple key points of the target to ensure that each LED light source can be detected individually; Step 3: deploying a spatial information extractor on one side of the target, projecting the optical image of the target onto the photosensitive area of the spatial information extractor, forming stimulated photogenerated carriers associated with the optical image of the target in the photosensitive area of the spatial information extractor, and then forming superimposed carriers; Step 4: Detect the concentration distribution of the superimposed carriers through the annular electrode to obtain the voltage and its distribution associated with the target; Step 5: Map the voltage associated with the target into polar coordinates, and form a closed polygon in the polar coordinates according to the voltage distribution associated with the target; quantitatively analyze the geometric features of the closed polygon, extract trend-related parameters and shape-related parameters, normalize each parameter to obtain trend factors and shape factors; based on the trend factors and shape factors, realize the recognition of single-frame targets and the perception of target actions.
4. The method for identifying and sensing a single-frame target according to claim 3, characterized in that: The trend-related parameters include angle parameters and length parameters; the shape-related parameters include the average value of light response intensity.
5. A method for identifying and sensing the motion of a continuous target, characterized in that: The following steps are involved: Step 1: construct a spatial information extractor, wherein the spatial information extractor is a spatial information extractor based on a graphene / silicon heterojunction as claimed in claim 1 or 2; Step 2: Arrange an LED light source at multiple key points of the target to ensure that each LED light source can be detected individually; Step 3: deploying a spatial information extractor on one side of the target, projecting the optical image of the target onto the photosensitive area of the spatial information extractor, forming stimulated photogenerated carriers associated with the optical image of the target in the photosensitive area of the spatial information extractor, and then forming superimposed carriers; Step 4: Detect the concentration distribution of the superimposed carriers through the annular electrode to obtain the voltage and its distribution associated with the target; Step 5: Map the voltage associated with the target into polar coordinates, and form a closed polygon in the polar coordinates according to the voltage distribution associated with the target; quantitatively analyze the geometric features of the closed polygon, extract trend-related parameters and shape-related parameters from them, normalize each parameter, and obtain the trend factor and shape factor of the optical image of the single-frame target; Step 6: Take the trend factor and shape factor of the optical images of multiple frames of the target connected end to end as an action. For any action, each trend factor and shape factor is used as the input of the trained recurrent neural network to achieve target recognition and action perception.
6. The method for identifying and sensing continuous targets according to claim 5, characterized in that: The trend-related parameters include angle parameters and length parameters; the shape-related parameters include the average value of light response intensity.