Visual task processing method, neural network model and neuron model
Patent Information
- Application Number
- CN202211474841.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-23
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2042-11-23
AI Technical Summary
这种处理方式只保留最强的前神经元的信息,丢失了其他前神经元的信息
[0012]本公开所提供的实施例,将待处理视觉特征划分为多个视觉切片特征,视觉切片特征包括多个子特征单元;针对各个视觉切片特征,确定视觉切片特征中各个子特征单元的激励值,并根据激励值确定各个子特征单元的权重,基于各个子特征单元的权重和对应的特征值得到加权视觉特征,并将加权视觉特征与视觉切片特征进行融合,得到目标视觉切片特征;根据多个目标视觉切片特征,确定目标视觉特征;其中,目标视觉特征用于获取视觉任务处理结果。换言之,在进行视觉特征处理时,通过各个子特征单元及其权重进行特征加权融合,不仅能够保留当前层级中信号最强的特征信息,同时融合了其他信号强度的视觉特征,因此,得到的下一层级的目标视觉特征中,充分保留了当前层级中各部分视觉特征的信息,提高了特征的处理效果,进而可以提高视觉任务处理结果的准确性。
Smart Images

Figure CN115731449B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to a visual task processing method, a neural network model, and a neuron model. Background Technology
[0002] In visual signal perception, higher visual cortexes often need to compress information perceived by primary visual cortexes. Typically, this compression is based on the interconnection structure and competition mechanism between multiple neurons. When multiple preneurons connect to a target neuron, they employ a winner-takes-all competition mechanism. When one preneuron is excited, it inhibits the excitation of other preneurons, allowing the subsequent neuron to receive the information fired by that neuron. This processing method retains only the information from the strongest preneuron, losing information from other preneurons. Summary of the Invention
[0003] This disclosure provides a visual task processing method, a neural network model and a neuron model, an electronic device, and a computer-readable storage medium.
[0004] In a first aspect, this disclosure provides a visual task processing method, which includes: dividing the visual features to be processed into multiple visual slice features, each visual slice feature including multiple sub-feature units; for each visual slice feature, determining the excitation value of each sub-feature unit in the visual slice feature, and determining the weight of each sub-feature unit according to the excitation value; obtaining a weighted visual feature based on the weight of each sub-feature unit and the corresponding feature value; and fusing the weighted visual feature with the visual slice feature to obtain a target visual slice feature; and determining a target visual feature based on the multiple target visual slice features; wherein the target visual feature is used to obtain the visual task processing result.
[0005] Secondly, this disclosure provides a visual task processing method based on a spiking neural network. This method includes: pulse encoding of original visual features to obtain a pulse sequence of visual features to be processed; inputting the pulse sequence of visual features to be processed into the spiking neural network for processing to obtain a target visual feature pulse sequence output by each network layer; pulse decoding of the target visual feature pulse sequence to obtain target visual features; and obtaining a visual task processing result based on at least one target visual feature. The multi-compartment dendritic model is used to execute the visual task processing method according to any one of the embodiments of this disclosure.
[0006] Thirdly, this disclosure provides a neural network model, which is composed of a spiking neural network. The spiking neural network includes multiple network layers, and at least some neurons of at least one network layer constitute a multi-compartmental dendritic model. The spiking neural network is the spiking neural network described in any one of the embodiments of this disclosure.
[0007] Fourthly, this disclosure provides a neuron model, which includes multiple dendritic chambers and at least one neuronal chamber. The output pulse signal of the neuron model is determined based on the current weighting value of each dendritic chamber and the historical current information of the neuronal chamber. The current weighting value is determined based on the excitation current of each dendritic chamber and its corresponding weight. The excitation current is determined based on the input pulse signal of the dendritic chamber and the historical current information of the dendritic chamber. Furthermore, the weight of each dendritic chamber is determined based on the excitation current of the multiple dendritic chambers.
[0008] Fifthly, this disclosure provides a visual task processing apparatus, comprising: a segmentation module for dividing a visual feature to be processed into multiple visual slice features, each visual slice feature including multiple sub-feature units; a processing module for determining, for each visual slice feature, an excitation value of each sub-feature unit in the visual slice feature, and determining a weight of each sub-feature unit based on the excitation value, obtaining a weighted visual feature based on the weight of each sub-feature unit and its corresponding feature value, and fusing the weighted visual feature with the visual slice feature to obtain a target visual slice feature; and a determination module for determining a target visual feature based on the multiple target visual slice features; wherein the target visual feature is used to obtain a visual task processing result.
[0009] In a sixth aspect, this disclosure provides an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores one or more computer programs executable by the at least one processor, the one or more computer programs being executed by the at least one processor to enable the at least one processor to perform the above-described visual task processing method.
[0010] In a seventh aspect, this disclosure provides an electronic device comprising: a plurality of processing cores; and an on-chip network configured to interact with data between the plurality of processing cores and external data; wherein one or more of the processing cores store one or more instructions, and the one or more instructions are executed by the one or more processing cores to enable the one or more processing cores to perform the above-described visual task processing method.
[0011] Eighthly, this disclosure provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor / processing core, implements the above-described visual task processing method.
[0012] The embodiments provided in this disclosure divide the visual features to be processed into multiple visual slice features, each visual slice feature including multiple sub-feature units. For each visual slice feature, the excitation value of each sub-feature unit in the visual slice feature is determined, and the weight of each sub-feature unit is determined according to the excitation value. A weighted visual feature is obtained based on the weight of each sub-feature unit and its corresponding feature value, and the weighted visual feature is fused with the visual slice feature to obtain the target visual slice feature. The target visual feature is determined based on multiple target visual slice features. The target visual feature is used to obtain the visual task processing result. In other words, when performing visual feature processing, feature weighting and fusion through each sub-feature unit and its weight not only retains the strongest feature information in the current level, but also fuses visual features with other signal strengths. Therefore, the target visual feature obtained in the next level fully retains the information of each part of the visual features in the current level, improving the feature processing effect and thus improving the accuracy of the visual task processing result.
[0013] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0014] The accompanying drawings are provided to further illustrate the present disclosure and form part of the specification. They are used together with the embodiments of the present disclosure to explain the disclosure and do not constitute a limitation thereof. The above and other features and advantages will become more apparent to those skilled in the art from the detailed description of exemplary embodiments with reference to the accompanying drawings, in which:
[0015] Figure 1 A flowchart of a visual task processing method provided in an embodiment of this disclosure;
[0016] Figure 2 A schematic diagram of a visual slice feature provided in an embodiment of this disclosure;
[0017] Figure 3 This is a schematic diagram illustrating the working process of a visual task processing method provided in an embodiment of the present disclosure;
[0018] Figure 4 This is a schematic diagram illustrating the working process of a visual task processing method provided in an embodiment of the present disclosure;
[0019] Figure 5A schematic diagram illustrating a visual task processing method based on a multi-compartment dendritic model provided in this embodiment of the present disclosure;
[0020] Figure 6 A schematic diagram illustrating a visual task processing method based on a multi-compartment dendritic model provided in this embodiment of the present disclosure;
[0021] Figure 7 A flowchart illustrating a visual task processing method based on a spiking neural network, provided in this embodiment of the disclosure;
[0022] Figure 8 A schematic diagram of a spiking neural network provided in an embodiment of this disclosure;
[0023] Figure 9 A schematic diagram of a spiking neural network provided in an embodiment of this disclosure;
[0024] Figure 10 A schematic diagram of a spiking neural network provided in an embodiment of this disclosure;
[0025] Figure 11 A schematic diagram of a spiking neural network provided in an embodiment of this disclosure;
[0026] Figure 12 A schematic diagram of a neuron model provided in an embodiment of this disclosure;
[0027] Figure 13 A block diagram of a vision task processing device provided in an embodiment of this disclosure;
[0028] Figure 14 A block diagram of an electronic device provided in an embodiment of this disclosure;
[0029] Figure 15 This is a block diagram of an electronic device provided in an embodiment of the present disclosure. Detailed Implementation
[0030] To enable those skilled in the art to better understand the technical solutions of this disclosure, exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments of this disclosure to aid understanding. These should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0031] Where there is no conflict, the various embodiments of this disclosure and the features thereof in the embodiments may be combined with each other.
[0032] As used herein, the term “and / or” includes any and all combinations of one or more related enumerated entries.
[0033] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. As used herein, the singular forms “a” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that when the terms “comprising” and / or “made of” are used in this specification, the presence of the stated feature, integral, step, operation, element, and / or component is specified, but the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof is not excluded. Words such as “connected” or “linked” are not limited to physical or mechanical connections but can include electrical connections, whether direct or indirect.
[0034] Unless otherwise specified, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and this disclosure, and will not be interpreted as having an idealized or overly formal meaning, unless expressly so defined herein.
[0035] In related technologies, during visual signal perception, the higher visual cortex typically compresses the information perceived by the primary visual cortex. This compression can be achieved based on the interconnection structure and competition mechanism between multiple neurons. That is, when one preneuron is excited, it inhibits the excitation of other preneurons, allowing the subsequent neurons to receive the information emitted by this neuron. This neural circuit based on a winner-takes-all competition mechanism retains only the information from the strongest preneuron, which may result in the loss of some background information.
[0036] In view of this, embodiments of this disclosure provide a visual task processing method. The visual features to be processed include multiple visual slice features. Based on the feature values and other information of each sub-feature unit in the visual feature slice, a certain weight is assigned to each sub-feature unit. The information of each sub-feature unit is weighted based on the weights, so that the target visual features obtained by compressing (or extracting high-dimensional semantic features) the visual features to be processed can fully retain the information of each sub-feature unit, improving the compression quality (or the quality of high-dimensional semantic features). Thus, when obtaining visual task processing results based on target visual features, the accuracy of visual task processing results can be improved.
[0037] According to the visual task processing method of this disclosure, the information of each part of the visual features in the current level is fully preserved in the target visual features obtained in the next level, which improves the feature processing effect and thus improves the accuracy of the visual task processing results.
[0038] The visual task processing method according to embodiments of this disclosure can be executed by electronic devices such as terminal devices or servers. The terminal device can be a user equipment (UE), mobile device, user terminal, terminal, cellular phone, cordless phone, personal digital assistant (PDA), handheld device, computing device, in-vehicle device, wearable device, etc. The method can be implemented by a processor calling computer-readable program instructions stored in memory. The server can be a standalone physical server, a server cluster consisting of multiple servers, or a cloud server capable of cloud computing.
[0039] In a first aspect, embodiments of this disclosure provide a visual task processing method.
[0040] Figure 1 A flowchart illustrating a visual task processing method provided in an embodiment of this disclosure. (Refer to...) Figure 1 The method includes:
[0041] In step S11, the visual features to be processed are divided into multiple visual slice features, and each visual slice feature includes multiple sub-feature units.
[0042] In step S12, for each visual slice feature, the excitation value of each sub-feature unit in the visual slice feature is determined, and the weight of each sub-feature unit is determined according to the excitation value. Based on the weight of each sub-feature unit and the corresponding feature value, a weighted visual feature is obtained, and the weighted visual feature is fused with the visual slice feature to obtain the target visual slice feature.
[0043] In step S13, target visual features are determined based on multiple target visual slice features; wherein, the target visual features are used to obtain the visual task processing results.
[0044] According to embodiments of this disclosure, the visual features to be processed are divided into multiple visual slice features, each visual slice feature including multiple sub-feature units. For each visual slice feature, the excitation value of each sub-feature unit in the visual slice feature is determined, and the weight of each sub-feature unit is determined based on the excitation value. A weighted visual feature is obtained based on the weight of each sub-feature unit and its corresponding feature value, and the weighted visual feature is fused with the visual slice feature to obtain the target visual slice feature. The target visual feature is determined based on multiple target visual slice features. The target visual feature is used to obtain the visual task processing result. In other words, when performing visual feature processing, feature weighting and fusion through each sub-feature unit and its weight not only retains the strongest feature information in the current level but also fuses visual features with other signal strengths. Therefore, the target visual feature obtained in the next level fully retains the information of each part of the visual features in the current level, improving the feature processing effect and thus improving the accuracy of the visual task processing result.
[0045] In some optional implementations, the visual processing task can be an object recognition task, an object classification task, an object tracking task, etc. The visual features to be processed belong to the data to be processed in the task; they can be image data of various formats or data obtained by encoding or other processing of images. Visual slice features are partial features of the visual features to be processed; the original visual features to be processed can be recovered through multiple visual slice features. Sub-feature units are the constituent units of visual slice features; by concatenating multiple sub-feature units, the corresponding visual slice feature can be obtained.
[0046] In some optional implementations, multiple visual slice features can be obtained by sliding windows on the visual features to be processed, or the visual features to be processed can be divided into multiple visual slice features according to preset requirements, etc. The embodiments of this disclosure do not limit this.
[0047] For example, step S11 includes: dividing the visual features to be processed into multiple visual slice features based on preset sliding window and sliding parameters. The sliding parameters include information such as sliding step size.
[0048] For example, setting the sliding window size to (k w ,k h ), sliding step size (s) w ,s h Based on this sliding window, multiple visual slice features are obtained by sequentially sliding across the visual features to be processed according to the sliding step size. Among them, k w k is the width of the sliding window. h s is the height of the sliding window w s is the sliding step size along the width direction. h This represents the sliding step size along the height direction.
[0049] For example, step S11 includes: dividing the visual features to be processed into multiple visual slice features according to preset slicing parameters. The slicing parameters include information such as slice size and slice shape.
[0050] For example, set the slice shape to rectangle and the slice size to (t). w ,t h ), where t w t is the width of the slice. h Given the height of the slice, based on the above slice parameters, the visual features to be processed can be divided into multiple visual slice features. Since visual slice features are obtained by segmenting the visual features to be processed from the visual features through slices, the visual slice features are usually the same as the shape and size of the slice. For example, a slice is (t... w ,t h If a rectangle is t, then the corresponding visual slice feature is also of size (t). w ,t h For example, if the slice is a rectangle of size (t,t), then the corresponding visual slice feature is also a square of size (t,t).
[0051] It should be noted that the slice shape can be a regular rectangle, square, or an irregular shape, and this disclosure does not impose any limitations on this. Moreover, for the multiple visual slice features obtained from the segmentation of the visual features to be processed, their size and shape can be the same or different. As long as it is ensured that each feature point is segmented into at least one visual slice feature, the feature point can be subsequently processed according to the method of this disclosure to ensure that the information of the feature point can be retained in the next level of visual features. This disclosure also does not impose any limitations on this.
[0052] Figure 2 This is a schematic diagram illustrating a visual slice feature provided in an embodiment of this disclosure. Wherein, Figure 2 (a) shows multiple visual slice features obtained based on the sliding window method. Figure 2 (b) illustrates multiple visual slice features obtained based on rule-based slice shapes. Figure 2 (c) shows several visual slice features obtained based on irregular slice shapes.
[0053] like Figure 2 As shown in (a), the size of the visual feature 200 to be processed is (W, H), based on a sliding window (k w ,k h ) with sliding step size (s) w ,s hThe process involves sliding across the visual features to be processed, obtaining one visual slice feature with each slide. Multiple visual slice features can be obtained through multiple slides.
[0054] like Figure 2 As shown in (b), the size of the visual feature 200 to be processed is (W, H), and the slice is set as a rectangle with a size of (t). w ,t h ), where t w =W / 2,t h =H / 2, thus we can obtain four visual slice features 221, 222, 223 and 224.
[0055] like Figure 2 As shown in (c), the size of the visual feature 200 to be processed is (W, H), and the slice is set to an "L" shape with a size of ((p) w1 ,p w2 ),(p h1 ,p h2 Thus, four visual slice features 231, 232, 233 and 234 can be obtained.
[0056] Furthermore, for each of the above visual slice features, one pixel can be regarded as a sub-feature unit, or multiple adjacent pixels can be regarded as a sub-feature unit. This disclosure does not limit this.
[0057] In step S12, for each visual slice feature, firstly, the activation value of each sub-feature unit in the visual slice is determined. Then, based on the activation value of each sub-feature unit, the weight of each sub-feature unit is obtained. Furthermore, the weight of each sub-feature unit and its corresponding feature value are weighted and calculated to obtain a weighted visual feature. Finally, the weighted visual feature and the visual slice feature are fused to obtain the target visual slice feature corresponding to each visual slice feature.
[0058] Therefore, in the above processing, the features of the current level are taken as the visual features to be processed, and features are extracted or compressed in units of visual slice features to obtain higher-level visual features. Correspondingly, in the visual cortex, the visual features to be processed are equivalent to the information perceived by the lower-level visual cortex. The above processing is equivalent to the higher-level visual cortex further compressing or extracting the information perceived by the lower-level visual cortex, thereby obtaining high-dimensional features with richer semantic information.
[0059] In the above processing, based on weights, the information of each sub-feature unit is retained to obtain weighted visual features. These weighted visual features are then further fused with visual slice features to obtain target visual slice features. Therefore, compared to visual slice features, target visual slice features acquire richer semantic information, becoming higher-level visual features. Moreover, in the process of extracting semantic information, the information of low-level features and the distribution of features (including the distribution of feature intensity at each feature location) are fully preserved, reducing the omission of feature information.
[0060] In some alternative implementations, the activation value can reflect the degree of contribution of a sub-feature unit to the feature value of the visual slice feature. Generally, the larger the activation value of a sub-feature unit, the higher its contribution to the feature value of the visual slice feature; conversely, the smaller the activation value, the lower its contribution.
[0061] In this embodiment, the excitation value of the sub-feature unit is analogous to the concept of a neuron, similar to the excitation current signal generated by a neuron after receiving an external pulse signal. Generally, the stronger the external pulse signal, the stronger the excitation current signal; conversely, the weaker the external pulse signal, the weaker the excitation current signal, or even the inability to generate a current signal. The same applies to the excitation value of the sub-feature unit, except that the feature value of the sub-feature unit can be regarded as the external pulse signal. That is, the larger the feature value of the sub-feature unit, the larger the excitation value; and the smaller the feature value of the sub-feature unit, the smaller the excitation value.
[0062] For example, determining the excitation value of each sub-feature unit in a visual slice feature includes: determining the average feature value of the visual slice feature based on the feature values of multiple sub-feature units; and determining the excitation value of each sub-feature unit based on the average feature value and the feature values of each sub-feature unit.
[0063] In this implementation, the average feature value of the visual slice features is used as the basis for judging the size of the feature value of the sub-feature unit, and the excitation value of each sub-feature unit is determined based on this.
[0064] In some optional implementations, the excitation values of each sub-feature unit can be calculated using a preset excitation function.
[0065] For example, the size of the visual slice feature is (k w ,k h First, calculate the average feature value of the visual slice feature x.
[0066]
[0067] in, x is the average feature value of the visual slice features. ij The sub-feature unit is represented by i and j, which are the width and height identifiers of the sub-feature unit.
[0068] Secondly, an exponential excitation function f() is used to calculate the excitation value of each sub-feature unit.
[0069]
[0070] Where, f(x) ij ) represents the sub-feature unit x ij The incentive value.
[0071] It should be noted that the above activation function f() is an extension of the LIF (Leaky-Integrated and Fire) neuron model, that is, when the feature value is greater than the threshold, a pulse signal with a higher value is emitted, and when the feature value is less than the threshold, current leakage or attenuation occurs. The average feature value of the visual slice features is used as the above threshold.
[0072] In some alternative implementations, a simpler activation function can be used to calculate the activation value of each sub-feature unit.
[0073] For example:
[0074] In this implementation, the difference between the eigenvalue and the average eigenvalue of a sub-feature unit is directly used as the activation value for that sub-feature unit. Compared to exponential activation functions, this reduces computational complexity and cost while still preserving the distribution characteristics of the features.
[0075] It should be noted that the above examples of excitation functions are merely illustrative and are not intended to limit the scope of this disclosure.
[0076] In some alternative implementations, after obtaining the activation values of each sub-feature unit, the weight of each sub-feature unit can be determined based on the activation values in order to perform weighted calculation of the features.
[0077] In some optional implementations, the weights of each sub-feature unit are determined based on the excitation values, including: performing positive correlation nonlinear processing on the excitation values of each sub-feature unit based on a preset function to obtain the initial weights of each sub-feature unit; and normalizing multiple initial weights to obtain the weights of each sub-feature unit.
[0078] For example, the preset function includes various positively correlated nonlinear functions, such as the tanh function, the sigmoid function, etc., and this disclosure does not limit this.
[0079] For example, the initial weights w of each sub-feature unit ij ′ can be calculated using the following formula:
[0080]
[0081] Where σ() represents a preset function.
[0082] After obtaining the above initial weight w ij Afterwards, it is quantized to the range (0,1) while maintaining positive correlation, thus obtaining the final weight w of each sub-feature unit. ij .
[0083] Therefore, it can be seen that the larger the eigenvalue of a sub-feature unit, the greater its weight, and the greater the proportion it retains in the next layer of visual features. In this way, while highlighting saliency information, neighborhood-related semantic information can be better integrated.
[0084] Furthermore, after obtaining the weights of each sub-feature unit, weighted visual features can be obtained based on the weights of each sub-feature unit and their corresponding feature values.
[0085] For example, the weight w of each sub-feature unit ij This can form a weight vector. The eigenvalues x of each sub-feature unit ij It can form an eigenvalue vector Weight vector and eigenvalue vectors By performing dot product calculation, the weighted visual feature y′ can be obtained.
[0086]
[0087] Here, "·" represents the dot product operation. Therefore, y′ obtained through the above dot product operation is essentially the result of multiplying the feature values and corresponding weights of each sub-feature unit to obtain the product of each sub-feature unit, and then accumulating the product results of multiple sub-feature units.
[0088] In some optional implementations, weighted visual features are fused with visual slice features to obtain target visual slice features, including: superimposing weighted visual features with the average feature value of visual feature slices to obtain target visual slice features.
[0089] For example, the target visual slice feature y can be calculated using the following formula.
[0090]
[0091] in, The average feature value of a visual slice can be considered as an approximation of the cumulative feature value of the previous time step; w ij *x ij Corresponding to different sub-feature units in the visual slice features, since different sub-feature units correspond to different locations in the visual slice features, and different locations correspond to different dendritic compartments in a multi-compartment dendrite, therefore, w ij *x ij It can be viewed as a weighted feature value about different locations in the visual slice features, or as a weighted feature value of multi-compartment dendrites. Let y′ and By superimposing the features, not only are the accumulated feature values from the previous time step preserved, but the feature values generated at the current time step due to the influence of the excitation signal at each position are also weighted. This can highlight the feature values with strong excitation and also integrate the feature values with relatively weak excitation, thus achieving the effect of highlighting salient information while integrating neighborhood-related semantic information.
[0092] In summary, through the above processing, the size (k) is... w ,k h The visual slice features of y are compressed into target visual slice features of size (1,1). In this process, not only is the information of each sub-feature unit and the intensity distribution information of the sub-feature unit preserved as much as possible, but also higher-dimensional semantic information is extracted from these sub-feature units, so that the semantic information representation ability of the target visual slice feature y is improved compared with the video slice feature x.
[0093] In some optional implementations, after obtaining the target visual slice features corresponding to each visual slice feature, multiple target visual slice features are concatenated to obtain the next-level visual features of the visual features to be processed, i.e., the target visual features. Furthermore, for the target visual features, one or more feature processing steps can still be performed based on the above method to obtain multiple levels of target visual features. Finally, by performing corresponding processing based on one or more of the above target visual features, the final visual task processing result can be obtained.
[0094] The following is combined Figures 3-6 The visual task processing method of the present disclosure will be described in detail below.
[0095] Figure 3 This is a schematic diagram illustrating the working process of a visual task processing method provided in an embodiment of this disclosure. (Refer to...) Figure 3 The visual feature X to be processed is a 4*4 feature, which is divided into four visual slice features. The top-left visual slice feature consists of four sub-feature units, with corresponding feature values x... 11 x 12 x 13 and x14 The visual slice feature in the upper right corner consists of four sub-feature units, with corresponding feature values x. 21 x 22 x 23 and x 24 The visual slice feature in the lower left corner consists of four sub-feature units, with corresponding feature values x. 31 x 32 x 33 and x 34 The visual slice feature in the lower right corner consists of four sub-feature units, with corresponding feature values x. 41 x 42 x 43 and x 44 .
[0096] Let's take the visual slice feature in the top left corner as an example for further explanation. First, calculate the average feature value of this visual slice feature.
[0097]
[0098] Secondly, the excitation value of each sub-feature unit is calculated using the excitation function f(), thereby obtaining x. 11 The incentive value is f(x) 11 ), x 12 The incentive value is f(x) 12 ), x 13 The incentive value is f(x) 13 ), x 14 The incentive value is f(x) 14 ), where f() can be any type of activation function.
[0099] Next, calculate the initial weights for each sub-feature unit.
[0100]
[0101]
[0102]
[0103]
[0104] Among them, w 11 ′ is the sub-feature unit x 11 The corresponding initial weight, w 12 ′ is the sub-feature unit x 12 The corresponding initial weight, w 13 ′ is the sub-feature unit x 13 The corresponding initial weight, w 14 ′ is the sub-feature unit x14 The corresponding initial weights.
[0105] Furthermore, by normalizing the above weights, we can obtain the sub-feature unit x. 11 weight w 11 Sub-feature unit x 12 weight w 12 Sub-feature unit x 13 weight w 13 and sub-feature unit x 14 weight w 14 .
[0106] After obtaining the weights of each sub-feature unit, the weighted visual feature y1′ can be obtained by performing a dot product operation based on the weights and the corresponding feature values.
[0107] y1′=w 11 *x 11 +w 12 *x 12 +w 13 *x 13 +w 14 *x 14
[0108] Finally, the weighted visual features and the average feature value of the visual slice feature are superimposed to obtain the target visual slice feature.
[0109]
[0110] The remaining three visual slice features are processed in a similar manner to obtain the target visual slice features y2, y3, and y4. Concatenating y1, y2, y3, and y4 yields the next-level visual feature Y of the visual features to be processed.
[0111] Therefore, in the above processing, the size of the visual feature is compressed from 4*4 to 2*2, thus achieving the compression of the visual feature; at the same time, it is equivalent to extracting features from visual feature X, obtaining richer semantic information, so that visual feature Y has a stronger semantic representation ability than visual feature X.
[0112] It should be noted that, Figure 3 The processing procedure for a visual feature to be processed is only shown in one channel. When the visual feature to be processed exists in multiple channels, the above method can be used to process the feature for each channel.
[0113] It should also be noted that in some optional implementations, for the visual feature X to be processed, other slicing methods can be used to obtain visual slice features of other sizes and / or shapes, and processed in a similar manner to obtain another set or more sets of target visual features. Additionally, the target visual feature Y can be regarded as a new visual feature to be processed and processed in the above manner to obtain higher-level target visual features. When obtaining the visual task processing results, any one or more of the above multiple / multi-layer target visual features can be selected for further processing as needed; this disclosure does not impose any limitations on this.
[0114] Figure 4 This is a schematic diagram illustrating the working process of a visual task processing method provided in an embodiment of this disclosure. (Refer to...) Figure 4 The visual task processing method includes:
[0115] In step S41, the sliding window size and sliding step size are set.
[0116] In step S42, the i-th visual slice feature is obtained by sliding the sliding window on the visual feature to be processed.
[0117] Where i is an integer greater than or equal to 1, and the visual slice feature includes multiple sub-feature units.
[0118] In step S43, the average feature value of each sub-feature unit in the i-th visual slice feature is calculated to obtain the average feature value of the i-th visual slice feature.
[0119] In step S44, the initial weights of each sub-feature unit in the i-th visual slice feature are obtained based on the average feature value of the i-th visual slice feature and the feature value of each sub-feature unit, combined with the preset activation function.
[0120] In step S45, the initial weights are normalized to obtain the weights of each sub-feature unit in the i-th visual slice feature.
[0121] In step S46, a dot product operation is performed on the weight vector and the feature value vector formed by each weight and the feature value of the corresponding sub-feature unit to obtain the i-th weighted visual feature.
[0122] In step S47, the average feature values of the i-th weighted visual feature and the i-th visual slice feature are superimposed to obtain the i-th target visual slice feature.
[0123] In step S48, the sliding window is moved according to the sliding step size, and then the process jumps to step S42 to perform the (i+1)th sliding until all areas of the visual features to be processed have been slid.
[0124] In step S49, multiple target visual slice features are stitched together to obtain higher-level visual features.
[0125] In some optional implementations, the aforementioned visual slice features are obtained by dividing the visual features to be processed according to a preset multi-compartment dendritic model. Each sub-feature unit in the visual slice feature corresponds to a dendritic compartment, and the excitation value of the sub-feature unit corresponds to the excitation current of the dendritic compartment. The input current of the multi-compartment dendritic model is a current weighted value determined based on the excitation current of each dendritic compartment and its corresponding weight. The multi-compartment dendritic model refers to a neuron comprising multiple dendritic compartments, and the neuron's output being jointly determined by these multiple dendritic compartments.
[0126] Figure 5 This is a schematic diagram illustrating a visual task processing method based on a multi-compartment dendritic model, provided as an embodiment of this disclosure. (Refer to...) Figure 5 , its Figure 3 Taking the visual slice features in the upper left corner as an example, we will elaborate on the processing procedure of the multi-compartment dendritic model.
[0127] like Figure 5 As shown, the dendritic chambers include 1-1, 1-2, 1-3, and 1-4, and the neuronal chamber includes 2-1. The input signal of dendritic chamber 1-1 corresponds to the sub-feature unit x. 11 The input signal of dendritic chamber 1-2 corresponds to the sub-feature unit x. 12 The input signals of dendritic chambers 1-3 correspond to sub-feature units x. 13 The input signals of dendritic chambers 1-4 correspond to sub-feature units x. 14 The four dendritic chambers mentioned above together serve as the input signals for neuron chamber 2-1.
[0128] First, construct the compartment dynamics engineering of this multi-compartment dendritic model.
[0129] The kinetic equations for dendritic chambers 1-1 to 1-4 are as follows:
[0130]
[0131] Where t represents time; τ is the time constant of the dendritic chamber; i is used to identify the dendritic chamber, and its value is 1, 2, 3, or 4; S i (t represents the input pulse information of the preneuron, I) i (t represents the current in the i-th dendritic chamber.)
[0132] The dynamic equation of neuron atrioventricular 2-1 is:
[0133]
[0134] Where τ′ is the time constant of the neuronal compartment; I5(t) represents the current in the neuronal compartment, W i (t) represents the weight of each dendritic compartment.
[0135] Considering that neural network computational models are typically in discrete form, that is, the time dimension is decomposed into several discrete moments Δt, the Euler method can be used to solve the above dynamic equations to obtain the corresponding difference iterative equations.
[0136] I i (n)=ρI i (n-1)+S i (n-1)
[0137]
[0138] Where n is the current time step; S i (n-1) represents the input pulse information of the i-th dendritic chamber at the (n-1)-th time step; I i (n-1) represents the current in the i-th dendritic chamber at the (n-1)-th time step; I i Ii(n) represents the current in the i-th dendritic chamber at the n-th time step; I5(n-1) represents the current in the neuronal chamber at the (n-1)-th time step; I5(n) represents the current in the neuronal chamber at the n-th time step; ρ and ρ′ are the current leakage coefficients, and t n = n*Δt.
[0139] Combining the above-described multi-compartment dendritic model theory with the visual task processing method of this embodiment, it can be seen that, in this embodiment, through the excitation of the input signal, each dendritic compartment generates an excitation current, namely I(x 11 ), I(x 12 ), I(x 13 ) and I(x 14 Then, based on the above excitation current, the weights of each dendritic compartment are determined, respectively w 11 w 12 w 13 and w 14 Based on the above weights, the excitation currents of each dendritic chamber are weighted and calculated to obtain the input current I1 = w in neuron chamber 2-1. 11 *I(x 11 )+w 12 *I(x 12 )+w 13 *I(x 13 )+w 14 *I(x 14Neuron chamber 2-1 combines its historical current information with the input current to jointly determine the outwardly emitted signal y1. In other words, when merging signals from multiple dendritic chambers, the pulse signals from multiple dendritic chambers are not simply superimposed or the strongest signal is selected. Instead, the currents of each dendritic chamber are weighted. Based on this method, the characteristics of multiple input signals can be preserved, improving the accuracy of the task processing results.
[0140] Figure 6 This is a schematic diagram of a visual task processing method based on a multi-compartment dendritic model provided in an embodiment of this disclosure.
[0141] in, Figure 6 (a) illustrates the processing procedure of a multi-compartment dendritic model based on an embodiment of the present disclosure. Figure 6 (b) illustrates the processing procedure of the relevant technology.
[0142] Reference Figure 6 (a) A visual feature slice in the visual feature X to be processed includes four sub-feature units (corresponding to the colored parts in the figure, different coloring methods represent one sub-feature unit), which are processed using a multi-compartment dendritic model. Each sub-feature unit corresponds to a dendritic compartment (including dendritic compartments 1, 2, 3, and 4). The excitation currents of each dendritic compartment are weighted and summed, and the weighted current is input into neuron compartment 5, outputting the target visual slice features. Finally, the target visual feature Y can be obtained through multiple target visual slice features. Based on this method, not only is the information of each sub-feature unit preserved, but the proportion of the preserved features in the new layer is also basically the same as the proportion of each sub-feature unit in the original layer.
[0143] Reference Figure 6 (b) In the visual feature X to be processed, each of the four sub-feature units of the visual feature slice corresponds to a neuron, and these four neurons (including neurons 1, 2, 3, and 4) are input into neuron 5. For example... Figure 6 As shown in (b), neuron 5 can only receive information from the neuron with the highest signal strength among the preceding neurons. Assuming that neuron 4 has the highest signal strength, neuron 5 can only retain the feature information of one sub-feature unit (the corresponding line is a solid line), while the feature information of the other three sub-feature units is not retained (the corresponding line is a dashed line).
[0144] Secondly, embodiments of this disclosure provide a visual task processing method based on a spiking neural network.
[0145] Figure 7 A flowchart illustrating a visual task processing method based on a spiking neural network, provided as an embodiment of this disclosure. (Refer to...) Figure 7 The method includes:
[0146] In step S71, the original visual features are pulse-coded to obtain a pulse sequence of visual features to be processed.
[0147] In step S72, the visual feature pulse sequence to be processed is input into the spiking neural network for processing to obtain the target visual features output by each network layer.
[0148] In step S73, pulse decoding is performed on the target visual feature pulse sequence to obtain the target visual features.
[0149] In step S74, the visual task processing result is obtained based on at least one target visual feature.
[0150] The multi-compartment dendritic model is used to perform the visual task processing method of any of the embodiments of this disclosure.
[0151] In some alternative implementations, the original visual features may be in formats such as images, which cannot be directly processed by spiking neural networks. They need to be pulse encoded to obtain a pulse sequence that can be processed by the spiking neural network.
[0152] For example, in step S71, frequency coding, time coding, or other coding methods can be used to encode the original visual features to obtain a pulse sequence of visual features to be processed that can be processed by a spiking neural network.
[0153] In some optional implementations, the spiking neural network includes multiple network layers, and at least some neurons in one network layer constitute a multi-compartment dendritic model. In step S72, after the visual feature pulse sequence to be processed is input into the spiking neural network, each network layer processes the pulse sequence sequentially, thereby obtaining the target visual features output by each network layer.
[0154] The following is combined Figure 8 The spiking neural network of the present disclosure will be described in detail.
[0155] Figure 8 This is a schematic diagram of a spiking neural network provided in an embodiment of this disclosure. (Refer to...) Figure 8 This spiking neural network comprises multiple network layers, wherein at least a portion of the neurons in one network layer constitute a multi-compartmental dendritic model. For example... Figure 8As shown, in the first network layer, neurons 1-1, 1-2, 1-3, 1-4, and 2-1 constitute a multi-compartment dendritic model. In this model, neurons 1-1, 1-2, 1-3, and 1-4 respond to the input pulse sequence by firing out a pulse sequence, generating an excitation current. For neuron 2-1, the received excitation current is the weighted sum of the individual excitation currents, and the weights of each excitation current are determined based on the excitation currents from neurons 1-1 to 1-4.
[0156] In some optional implementations, the original visual feature X is in image form. It is pulse-coded using frequency coding or time coding to obtain a pulse sequence of visual features to be processed. This pulse sequence is then input into a spiking neural network. First, the first network layer processes this pulse sequence to obtain the target visual feature pulse sequence corresponding to the first network layer. This target visual feature pulse sequence is then used as the input visual feature pulse sequence for the second network layer. After processing by the second network layer, the target visual feature pulse sequence corresponding to the second network layer is obtained. This process continues until the last network layer processes the data to obtain the target visual feature pulse sequence for the last network layer. Since the output data of each network layer is in pulse sequence form, it can be pulse-decoded to obtain the target visual features of each network layer.
[0157] Figure 8 The diagram only shows pulse decoding of the target visual feature pulse sequence from the last network layer to obtain the corresponding target visual feature Y. In practical applications, pulse decoding can be performed on the target visual feature pulse sequences output from any one or more network layers to obtain the corresponding target visual features, and this disclosure does not impose any limitations on this.
[0158] In some alternative implementations, after obtaining one or more of the target visual features mentioned above, several target visual features can be selected for further processing as needed to obtain the corresponding visual task processing results.
[0159] Furthermore, when performing the above processing procedure with a time step Δt, at each time step t n Inside, the corresponding visual slice features are obtained through a sliding window. Each sub-feature unit in the visual slice feature corresponds to a dendritic chamber. Through feature sampling and pulse coding, a pulse sequence can be obtained. Based on the aforementioned processing method, a compressed visual feature pulse sequence can be obtained.
[0160] by Figure 3Taking the visual feature X shown as an example, within time step t1 (i.e., the first time step), we obtain the result from x. 11 x 12 x 13 and x 14 The visual slice features are composed of sub-feature units x. 11 Corresponding to neuron 1-1 in the first network layer, sub-feature unit x 12 Corresponding to neurons 1-2 in the first network layer, sub-feature unit x 13 Corresponding to neurons 1-3 in the first network layer, sub-feature unit x 14 This corresponds to neurons 1-4 in the first network layer. By performing feature sampling and pulse coding on each of the above sub-feature units, the corresponding pulse sequences can be obtained, which are then input into each neuron. Figure 3 The description of the relevant content allows us to obtain the weights of each neuron. By performing a dot product operation on the weights and eigenvalues, and combining this with the accumulated eigenvalues from the previous time step, we can obtain the output characteristic pulse sequence of neuron 2-1. Other multi-compartment dendritic models are similar and will not be described in detail here.
[0161] It should be noted that pulse coding of the original visual features and pulse decoding of the target visual feature pulse sequence can be implemented using dedicated encoders and decoders, or by corresponding spiking neural network structures. For the implementation using spiking neural network structures, the corresponding neural network layers (e.g., pulse coding neural network layers, pulse decoding neural network layers) can be embedded into... Figure 8 The first network layer can be mentioned before and the last network layer, but this embodiment does not limit this.
[0162] In this embodiment of the disclosure, a multi-compartment dendritic model is constructed using multiple neurons, and a spiking neural network is built based on the multi-compartment dendritic model. This allows the spiking neural network to be used to perform the aforementioned visual task processing method, ensuring that the target visual features obtained at the next level fully retain the information of each part of the visual features in the previous level, thereby improving the feature processing effect and thus improving the accuracy of the visual task processing results.
[0163] Thirdly, embodiments of this disclosure provide a spiking neural network that can be applied to processing visual tasks.
[0164] Figure 9 This is a schematic diagram of a spiking neural network provided in an embodiment of this disclosure. (Refer to...) Figure 9 The spiking neural network includes multiple network layers, wherein at least some neurons in at least one network layer constitute a multi-compartmental dendritic model.
[0165] like Figure 9 As shown, the spiking neural network includes a first network layer, a second network layer, ..., an Lth network layer (L≥1). The first network layer includes a multi-compartmental dendritic model consisting of 5 neurons.
[0166] It should be noted that other multi-compartment dendritic models may exist in the first network layer, and other multi-compartment dendritic models may also exist in the second to Lth network layers. This disclosure does not limit these aspects.
[0167] It should also be noted that, within the first network layer, other multi-compartment dendritic models can have different specifications (the specifications are related to the number of neurons). For example, the first network layer can also include a multi-compartment dendritic model consisting of 4 neurons. Furthermore, the specifications of the multi-compartment dendritic models can differ between different network layers, and this disclosure does not impose any limitations on this.
[0168] Figure 10 This is a schematic diagram of a spiking neural network provided in an embodiment of this disclosure. (Refer to...) Figure 10 The spiking neural network consists of a first network layer and a second network layer. The first network layer includes three identical multi-compartment dendritic models, each composed of five neurons. The second network layer includes one multi-compartment dendritic model, which consists of four neurons.
[0169] Figure 11 This is a schematic diagram of a spiking neural network provided in an embodiment of this disclosure. (Refer to...) Figure 11 The spiking neural network comprises a first network layer and a second network layer. The first network layer includes four multi-compartment dendritic models, two of which consist of three neurons each, and the remaining two consist of five neurons each. The second network layer includes one multi-compartment dendritic model, which consists of five neurons.
[0170] Fourthly, embodiments of this disclosure provide a neuron model that can be applied to processing visual tasks.
[0171] Figure 12 This is a schematic diagram of a neuron model provided in an embodiment of this disclosure. (Refer to...) Figure 12 The neuron model includes multiple dendritic chambers and at least one neuronal chamber. The output pulse signal of the neuron model is determined based on the current weighting value of each dendritic chamber and the historical current information of the neuronal chamber.
[0172] The current weighting value is determined based on the excitation current and corresponding weight of each dendritic chamber. The excitation current is determined based on the input pulse signal of the dendritic chamber and the historical current information of the dendritic chamber. The weight of each dendritic chamber is determined based on the excitation current of multiple dendritic chambers.
[0173] It should be noted that this neuron model is based on the theory of the multi-compartment dendritic model. Its working process can be found in the relevant content of the multi-compartment dendritic model, and will not be described in detail here.
[0174] Fifthly, embodiments of this disclosure provide a visual task processing apparatus.
[0175] Figure 13 This is a block diagram of a vision task processing device provided in an embodiment of the present disclosure.
[0176] Reference Figure 13 This disclosure provides a visual task processing device 1300, which includes:
[0177] The segmentation module 1301 is used to divide the visual features to be processed into multiple visual slice features, each of the visual slice features including multiple sub-feature units.
[0178] The processing module 1302 is configured to determine the excitation value of each sub-feature unit in the visual slice feature for each visual slice feature, determine the weight of each sub-feature unit according to the excitation value, obtain a weighted visual feature based on the weight of each sub-feature unit and the corresponding feature value, and fuse the weighted visual feature with the visual slice feature to obtain the target visual slice feature.
[0179] The determining module 1303 is used to determine target visual features based on multiple target visual slice features; wherein the target visual features are used to obtain visual task processing results.
[0180] In this embodiment, a segmentation module divides the visual features to be processed into multiple visual slice features, each visual slice feature comprising multiple sub-feature units. A processing module determines the excitation value of each sub-feature unit within each visual slice feature and determines the weight of each sub-feature unit based on the excitation value. A weighted visual feature is obtained based on the weights and corresponding feature values of each sub-feature unit, and this weighted visual feature is fused with the visual slice features to obtain the target visual slice feature. Finally, a determination module determines the target visual feature based on the multiple target visual slice features. The target visual feature is used to obtain the visual task processing result. In other words, during visual feature processing, feature weighting and fusion using each sub-feature unit and its weight not only retains the strongest signal feature information in the current level but also fuses visual features with other signal intensities. Therefore, the target visual feature obtained in the next level fully retains the information of each part of the visual features in the current level, improving the feature processing effect and thus enhancing the accuracy of the visual task processing result.
[0181] It is understood that the various method embodiments mentioned above in this disclosure can be combined with each other to form combined embodiments without violating the principle and logic. Due to space limitations, this disclosure will not elaborate further. Those skilled in the art will understand that in the above methods of specific implementation, the specific execution order of each step should be determined by its function and possible internal logic.
[0182] In addition, this disclosure also provides electronic devices and computer-readable storage media, all of which can be used to implement any of the visual task processing methods provided in this disclosure. The corresponding technical solutions and descriptions are described in the corresponding descriptions in the method section, and will not be repeated here.
[0183] Figure 14 This is a block diagram of an electronic device provided in an embodiment of the present disclosure.
[0184] Reference Figure 14 This disclosure provides an electronic device, which includes: at least one processor 1401; at least one memory 1402; and one or more I / O interfaces 1403 connected between the processor 1401 and the memory 1402; wherein the memory 1402 stores one or more computer programs that can be executed by at least one processor 1401, and the one or more computer programs are executed by at least one processor 1401 to enable at least one processor 1401 to perform the above-described visual task processing method.
[0185] Figure 15 This is a block diagram of an electronic device provided in an embodiment of the present disclosure.
[0186] Reference Figure 15This disclosure provides an electronic device that includes multiple processing cores 15401 and an on-chip network 1502. The multiple processing cores 1501 are all connected to the on-chip network 1502, and the on-chip network 1502 is used to exchange data between the multiple processing cores and external data.
[0187] One or more processing cores 1501 store one or more instructions, and the one or more instructions are executed by one or more processing cores 1501 to enable one or more processing cores 1401 to perform the above-described visual task processing method.
[0188] In some embodiments, the electronic device may be a neuromorphic chip. Since neuromorphic chips can employ vectorized computation and require external memory, such as Double Data Rate (DDR) synchronous dynamic random access memory, to load parameters such as weights of the neural network model, the batch processing method used in this embodiment offers higher computational efficiency.
[0189] This disclosure also provides a computer-readable storage medium storing a computer program thereon, wherein the computer program, when executed by a processor / processor core, implements the aforementioned visual task processing method. The computer-readable storage medium may be volatile or non-volatile.
[0190] This disclosure also provides a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code, wherein when the computer-readable code is run in a processor of an electronic device, the processor in the electronic device performs the above-described visual task processing method.
[0191] Those skilled in the art will understand that all or some of the steps, systems, and apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware implementations, the division between functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit (ASIC). Such software can be distributed on a computer-readable storage medium, which may include computer storage media (or non-transitory media) and communication media (or transient media).
[0192] As is known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable program instructions, data structures, program modules, or other data). Computer storage media includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), static random access memory (SRAM), flash memory or other memory technologies, portable compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer. Furthermore, it is known to those skilled in the art that communication media typically contain computer-readable program instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.
[0193] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.
[0194] Computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing the status information of the computer-readable program instructions to implement various aspects of this disclosure.
[0195] The computer program product described herein can be implemented specifically through hardware, software, or a combination thereof. In one alternative embodiment, the computer program product is specifically embodied in a computer storage medium; in another alternative embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.
[0196] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0197] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0198] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0199] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0200] Example embodiments have been disclosed herein, and while specific terminology has been used, it is for illustrative purposes only and should be construed as such, and is not intended to be limiting. In some instances, it will be apparent to those skilled in the art that features, characteristics, and / or elements described in connection with particular embodiments may be used alone, or in combination with features, characteristics, and / or elements described in connection with other embodiments, unless otherwise expressly indicated. Therefore, those skilled in the art will understand that various changes in form and detail may be made without departing from the scope of this disclosure as set forth by the appended claims.
Claims
1. A visual task processing method, characterized in that, include: The visual features to be processed are divided into multiple visual slice features, and each visual slice feature includes multiple sub-feature units; For each visual slice feature, the activation value of each sub-feature unit in the visual slice feature is determined, and the weight of each sub-feature unit is determined according to the activation value. A weighted visual feature is obtained based on the weight of each sub-feature unit and the corresponding feature value. The weighted visual feature is then fused with the visual slice feature to obtain the target visual slice feature. The activation value is used to characterize the degree of contribution of the sub-feature unit to the feature value of the visual slice feature. Based on multiple target visual slice features, target visual features are determined; wherein, the target visual features are used to obtain the visual task processing results; The step of fusing the weighted visual features with the visual slice features to obtain the target visual slice features includes: The target visual slice feature is obtained by superimposing the average feature value of the weighted visual feature and the visual slice feature; The process of dividing the visual features to be processed into multiple visual slice features includes: Based on preset sliding window and sliding parameters, the visual features to be processed are divided into multiple visual slice features.
2. The method according to claim 1, characterized in that, Determining the activation value of each sub-feature unit in the visual slice feature includes: The average feature value of the visual slice feature is determined based on the feature values of multiple sub-feature units; The excitation value of each sub-feature unit is determined based on the average eigenvalue and the eigenvalue of each sub-feature unit.
3. The method according to claim 1, characterized in that, Determining the weight of each sub-feature unit based on the excitation value includes: The initial weights of each sub-feature unit are obtained by performing positive correlation nonlinear processing on the excitation values of each sub-feature unit based on a preset function. The initial weights are normalized to obtain the weights of each sub-feature unit.
4. The method according to claim 1, characterized in that, The weighted visual features obtained based on the weights and corresponding feature values of each sub-feature unit include: The weighted visual features are obtained by performing dot product calculation based on the feature values and corresponding weights of each sub-feature unit.
5. The method according to claim 1, characterized in that, The visual slice features are obtained by dividing the visual features to be processed according to a preset multi-compartment dendritic model; In this context, each sub-feature unit in the visual slice feature corresponds to a dendritic chamber, and the excitation value of the sub-feature unit corresponds to the excitation current of the dendritic chamber. The input current of the multi-chamber dendritic model is a current weighted value determined based on the excitation current of each dendritic chamber and its corresponding weight.
6. A visual task processing method based on a spiking neural network, characterized in that, The spiking neural network comprises multiple network layers, with some neurons in at least one network layer forming a multi-compartmental dendritic model. The method includes: The original visual features are pulse-coded to obtain the pulse sequence of visual features to be processed; The visual feature pulse sequence to be processed is input into the spiking neural network for processing to obtain the target visual feature pulse sequence output by each network layer; The target visual feature pulse sequence is pulse-decoded to obtain the target visual features; Based on at least one of the target visual features, the visual task processing result is obtained; The multi-compartment dendritic model is used to perform the visual task processing method as described in any one of claims 1-5.
7. A neural network model, characterized in that, For processing visual tasks, the neural network model is composed of a spiking neural network, which includes multiple network layers, and at least some neurons of the network layer constitute a multi-compartment dendritic model. The spiking neural network described in claim 6 is used.
8. A neuron model, characterized in that, The method is applied to visual tasks to implement the visual task processing method as described in any one of claims 1-5. The neuron model includes multiple dendritic chambers and at least one neuronal chamber. The output pulse signal of the neuron model is determined based on the current weighting value of each dendritic chamber and the historical current information of the neuronal chamber. The current weighting value is determined based on the excitation current and corresponding weight of each dendritic chamber. The excitation current is determined based on the input pulse signal of the dendritic chamber and the historical current information of the dendritic chamber. The weight of each dendritic chamber is determined based on the excitation current of multiple dendritic chambers.
9. A visual task processing device, characterized in that, A method for implementing the visual task processing method as described in any one of claims 1-5, comprising: A segmentation module is used to divide the visual features to be processed into multiple visual slice features, each of which includes multiple sub-feature units. The processing module is used to determine the excitation value of each sub-feature unit in the visual slice feature for each visual slice feature, determine the weight of each sub-feature unit according to the excitation value, obtain a weighted visual feature based on the weight of each sub-feature unit and the corresponding feature value, and fuse the weighted visual feature with the visual slice feature to obtain the target visual slice feature. The determination module is used to determine target visual features based on multiple target visual slice features; wherein the target visual features are used to obtain the visual task processing results.
10. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores one or more computer programs that can be executed by the at least one processor, the one or more computer programs being executed by the at least one processor to enable the at least one processor to perform the visual task processing method as described in any one of claims 1-5, or the visual task processing method based on a spiking neural network as described in claim 6.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the visual task processing method as described in any one of claims 1-5, or the visual task processing method based on a spiking neural network as described in claim 6.
Citation Information
Patent Citations
Image processing method and device, computer equipment and storage medium
CN111476806A
Asynchronous pulse modulation for threshold-based signal coding
US20150372805A1
Correlative time coding method for spiking neural networks
US20210357725A1