Computer-implemented or hardware-implemented entity recognition method, computer program product and apparatus for entity recognition

CN115699018BActive Publication Date: 2026-08-21INTUISEL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180042523.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-11-25
Filing Date
2021-06-16
Publication Date
2026-08-21
Estimated Expiration
2041-06-16

AI Technical Summary

Technical Problem

[0003]然而,现有的神经网络解决方案具有差的性能和/或低的可靠性

Benefits of technology

[0022]又一优点在于,该装置能够进行自我训练,即,有限的初始训练数据量在该装置中被渗透,使得其表示以新的组合通过网络被馈送,这实现了某种“数据增强”,但在这种情况下,向该装置重放的是传感信息的内部表示,而不是调整后的传感器数据,从而提供了比由数据增强提供的自我训练更有效的自我训练。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115699018B_ABST
    Figure CN115699018B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a computer implemented or hardware implemented entity recognition method (100) comprising: a) providing (110) input from a plurality of sensors to a network of nodes; b) generating (120) an activity level by each node of the network based on the input from the plurality of sensors; c) comparing (130) the activity level of each node to a threshold level; d) based on the comparison, setting (140) the activity level for each node to a preset value or keeping the generated activity level; e) calculating (150) a total activity level as a sum of all activity levels of the nodes of the network; f) iterating (160) a) to e) until a local minimum of the total activity level is reached; and g) when the local minimum of the total activity level is reached, utilizing (170) a distribution of activity levels at the local minimum to identify a measurable property of an entity. The present disclosure also relates to a computer program product and an apparatus (300) for entity recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to computer-implemented or hardware-implemented entity recognition methods, computer program products, and apparatuses for entity recognition. More specifically, this disclosure relates to computer-implemented or hardware-implemented entity recognition methods, computer program products, and apparatuses for entity recognition provided according to the first, second, and third aspects described below. Background Technology

[0002] Entity recognition is known from existing technologies. One technique used to perform entity recognition is a neural network. One type of neural network that can be used for entity recognition is the Hopfield network. A Hopfield network is a form of recurrent artificial neural network. Hopfield networks are used as content-addressable (“associative”) memory systems with binary threshold nodes.

[0003] However, existing neural network solutions suffer from poor performance and / or low reliability. Furthermore, existing solutions require significant time to train and therefore may require substantial computing power and / or energy, especially for training. Additionally, existing neural network solutions may require large amounts of storage space.

[0004] Therefore, alternative entity recognition methods are needed. Preferably, such methods provide or achieve one or more of the following: improved performance, higher reliability, increased efficiency, faster training, less computing power, less training data, less storage space, and / or less energy. Summary of the Invention

[0005] The purpose of this disclosure is to mitigate, alleviate, or eliminate one or more of the defects and disadvantages specified above in the prior art, and to solve at least the problems mentioned above. According to a first aspect, a computer-implemented or hardware-implemented entity recognition method is provided, comprising: a) providing inputs from multiple sensors to a network of nodes; b) generating an activity level by each node of the network based on the inputs from the multiple sensors; c) comparing the activity level of each node with a threshold level; d) based on the comparison, setting the activity level to a preset value or maintaining the generated activity level for each node; e) calculating the total activity level as the sum of all activity levels of the nodes in the network; f) iterating through a) to e) until a local minimum of the total activity level is reached; and g) when the local minimum of the total activity level is reached, using the distribution of activity levels at the local minimum to identify measurable characteristics of the entity. The first aspect has the advantage that the effective structure of the network can be dynamically changed, which enables, for example, each unit / node to identify a large number of entities.

[0006] In some implementations, the input changes dynamically over time and follows the sensor input trajectory. One advantage is that this method is less sensitive to noise. Another advantage is faster recognition. Yet another advantage is that it enables more accurate recognition.

[0007] In some implementations, multiple sensors monitor the correlation between the sensors. This has the advantage that the method is less sensitive to noise.

[0008] According to some implementations, when following the sensor input trajectory with a deviation smaller than the user-defined deviation threshold over a time period longer than the user-defined time threshold, a local minimum of the total activity level has been reached.

[0009] According to some implementations, the activity level of each node is used as input to all other nodes, each input is weighted using weights, and wherein at least one weighted input is negative, and / or wherein at least one weighted input is positive, and / or wherein all maintained generated activity levels are positive scalars.

[0010] According to some implementations, the network is activated by an activation energy X, which influences the location of a local minimum in the overall activity level.

[0011] According to some implementations, the input from multiple sensors is pixel values, such as intensity, of images captured by a camera, and wherein the distribution of activity levels across all nodes is further utilized to control the camera's positioning via rotational and / or translational movement, thereby controlling the sensor input trajectory, and wherein the identified entity is an object or a feature of an object present in at least one of the captured images. One advantage of this method is its low sensitivity to noise. Another advantage is that identification is independent of absolute time quantities, such as the absolute time spent on each still camera image and the absolute time spent between different such still camera images as the camera's positioning changes.

[0012] According to some implementations, the multiple sensors are touch sensors, and the input from each of the multiple sensors is a touch event signal with a force-related value, wherein the distribution of activity levels across all nodes is used to identify the sensor input trajectory as a new contact event, the end of a contact event, a gesture, or as applied pressure.

[0013] According to some implementations, each of a plurality of sensors is associated with a different frequency band of an audio signal, wherein each sensor reports the energy present in the associated frequency band, and wherein a combined input from a plurality of such sensors follows a sensor input trajectory, and wherein the distribution of activity levels across all nodes is used to identify the speaker and / or spoken letters, syllables, words, phrases or phonemes present in the audio signal.

[0014] According to the second aspect, a computer program product including a non-transitory computer-readable medium is provided, wherein the non-transitory computer-readable medium has a computer program including program instructions, the computer program being loadable into a data processing unit and configured to cause any of the embodiments or methods mentioned above to be performed when the computer program is run by the data processing unit.

[0015] According to a third aspect, an apparatus for entity recognition is provided, the apparatus comprising a control circuit system configured such that: a) inputs from multiple sensors are provided to a network of nodes; b) an activity level is generated by each node of the network based on the inputs from the multiple sensors; c) the activity level of each node is compared with a threshold level; d) based on the comparison, for each node, the activity level is set to a preset value or the generated activity level is maintained; e) the total activity level is calculated as the sum of all activity levels of the nodes of the network; f) a) through e) are iterated until a local minimum of the total activity level is reached; and g) when the local minimum of the total activity level is reached, the distribution of activity levels at the local minimum is used to identify measurable characteristics of the entity.

[0016] The effects and features of the second and third aspects are largely similar to those described above in conjunction with the first aspect, and the effects and features of the first aspect are largely similar to those described above in conjunction with the second and third aspects. The implementation methods mentioned in the first aspect are largely compatible with the second and third aspects, and the implementation methods mentioned in the second and third aspects are largely compatible with the first aspect.

[0017] Some implementations have the advantage of providing alternative methods for entity recognition.

[0018] Some implementations offer the advantage of improved performance in entity recognition.

[0019] Another advantage of some implementations is that they provide more reliable entity recognition.

[0020] Another advantage of some implementations is that the device is trained faster, for example because the device is more versatile or generalized due to, for example, improved dynamic performance.

[0021] Another advantage of some implementations is that the processing elements are trained faster, for example, because only a small set of training data is required.

[0022] Another advantage is that the device is capable of self-training, that is, a limited amount of initial training data is permeated into the device so that its representation is fed through the network in a new combination, which achieves a kind of "data augmentation". However, in this case, what is replayed to the device is the internal representation of the sensing information, rather than the adjusted sensor data, thus providing a more effective self-training than the self-training provided by data augmentation.

[0023] Another advantage of some implementations is that they provide an effective or more efficient method for identifying entities.

[0024] Another advantage of some implementations is that they provide an energy-efficient method for identifying entities, for example, because the method saves computing power and / or storage space.

[0025] Another advantage of some implementations is that they provide a bandwidth-efficient method for identifying information fragments, for example, because the method saves the bandwidth required to transmit data.

[0026] This disclosure will become apparent from the detailed description given below. The detailed description and specific examples disclose preferred embodiments of this disclosure by way of illustration only. Those skilled in the art will understand from the guidance of the detailed description that changes and modifications can be made within the scope of this disclosure.

[0027] Therefore, it should be understood that the disclosure herein is not limited to the specific components of the described apparatus or the steps of the described method, as such apparatus and methods can vary. It should also be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. It should be noted that, when used in the specification and appended claims, the articles “a,” “an,” “the,” and “described” are intended to indicate the presence of one or more elements, unless expressly stated in the context. Thus, for example, a reference to “unit” or “the unit” can include several devices, etc. Furthermore, the words “comprising,” “including,” “containing,” and similar terms do not exclude other elements or steps.

[0028] Terminology—The term “measurable” should be interpreted as something that can be measured or detected, i.e., detectable. The terms “measurement” and “sensing” should be interpreted as synonyms. The term “entity” should be interpreted as an entity such as a physical entity, or a more abstract entity such as a financial entity, such as one or more financial datasets. The term “physical entity” should be interpreted as an entity with a physical existence, such as an object, a feature (of an object), a posture, applied pressure, a speaker, spoken letters, syllables, phonemes, words, or phrases. The term “node” can be a neuron (in a neural network) or another processing element. Attached Figure Description

[0029] The above-described objectives, as well as other objectives, features, and advantages of this disclosure, will be more fully understood by referring to the following illustrative and non-limiting detailed description of exemplary embodiments of this disclosure taken in conjunction with the accompanying drawings.

[0030] Figure 1 This is a flowchart illustrating example method steps according to some embodiments of this disclosure;

[0031] Figure 2 This is a schematic diagram illustrating an example computer-readable medium according to some embodiments;

[0032] Figure 3 This is a schematic block diagram illustrating an example apparatus according to some embodiments;

[0033] Figure 4 This is a schematic diagram illustrating the operating principle of the device based on examples of using multiple sensors and multiple processing elements in some implementations;

[0034] Figure 5 This is a schematic diagram illustrating the operating principle of the device based on examples of using multiple sensors and multiple processing elements in some implementations;

[0035] Figure 6 This is a schematic diagram illustrating the operating principle of the device based on examples of using multiple processing elements in some implementation methods;

[0036] Figure 7 This is a schematic diagram illustrating the operating principle of the device based on examples of using multiple sensors and multiple processing elements in some implementations;

[0037] Figure 8 A to Figure 8 H is a schematic diagram illustrating the operating principle of the device based on an example of using multiple tactile sensors in some implementations;

[0038] Figure 9This is a schematic diagram illustrating the operating principle of the device based on examples of using a camera in some implementation methods;

[0039] Figures 10A to 10C It is a graph showing the frequency and power of different sensors sensing audio signals; and

[0040] Figure 11 It is a graph of the sensor trajectory. Detailed Implementation

[0041] This disclosure will now be described with reference to the accompanying drawings, in which preferred exemplary embodiments of the disclosure are illustrated. However, this disclosure may be implemented in other forms and should not be construed as limited to the embodiments disclosed herein. The disclosed embodiments are provided to fully communicate the scope of this disclosure to those skilled in the art.

[0042] The implementation methods will be described below, wherein Figure 1 This is a flowchart illustrating example method steps according to an embodiment of this disclosure. Figure 1 An entity recognition method 100, implemented in computer or hardware, is illustrated. Therefore, the method can be implemented in hardware, software, or any combination of both. The method includes directing data to a network 520 (e.g., nodes 522, 524, 526, 528) of nodes 522, 524, 526, 528. Figure 5(As shown) 110 inputs 502, 504, 506, 508 from multiple sensors are provided. Network 520 can be a recurrent network, such as a recurrent neural network. Sensors can be any suitable sensor; for example, an image sensor (e.g., pixels), an audio sensor (e.g., a microphone), or a tactile sensor (e.g., a pressure sensor array or a bio-inspired tactile sensor). Furthermore, the inputs can be generated in-situ, i.e., the sensors are directly connected to network 520. Alternatively, the inputs are first generated and recorded / stored, and then fed to network 520. Additionally, the inputs may have been preprocessed. One way to preprocess the inputs is by combining (e.g., by averaging or adding the signals) or recombining multiple sensor signals, and using such combinations or recombinations as one or more inputs to network 520. In either case, the input data may change and evolve over time. Furthermore, the method includes generating 120 activity levels by each node 522, 524, 526, 528 of network 520 based on the inputs from multiple sensors. Therefore, each node 522, 524, 526, 528 has an associated activity level. The activity level of a node represents its state, such as its internal state. Furthermore, the set of activity levels, i.e., the activity levels of each node, represents the internal state of network 520. The internal state of network 520 can be used as a feedback signal to one or more actuators, which directly or indirectly control the input, thereby controlling the sensor input trajectory. Additionally, the method includes comparing the activity level of each node 522, 524, 526, 528 with a threshold level (i.e., an activity level threshold) 130. The threshold level can be any suitable value. The method includes, based on comparison 130, setting the activity level 140 to a preset value or maintaining the generated activity level for each node 522, 524, 526, 528. In some embodiments, if the generated activity level is higher than the threshold level, the generated activity level is maintained, while if the generated activity level is equal to or lower than the threshold level, the activity level is set to a preset value. The activity level can be set to a preset value of zero or to any other suitable value. Furthermore, the method includes calculating the total activity level 150 as the sum of all activity levels of nodes 522, 524, 526, and 528 of network 520. For the calculation, a set value can be used together with any retained generated values. If the preset value is zero, the total activity level can be calculated as the sum of only all retained generated values, thus ignoring all zero values ​​and providing faster computation. Additionally, the method includes iterating 160 on the previously described steps 110, 120, 130, 140, and 150 until a local minimum of the total activity level is reached. Iteration is performed using continuously evolving input signals, for example, each input being a continuous signal from a sensor. A local minimum is reached when the lowest possible total activity level is reached.The lowest possible total activity level can be considered reached when the total activity level falls below the total activity threshold over multiple iterations. The number of iterations can be any suitable number, such as two. Once a local minimum of the total activity level has been reached, the distribution of the activity level at the local minimum is used to identify measurable characteristics (or multiple measurable characteristics) of the entity. Measurable characteristics can be features of an object, a portion of a feature, a trajectory of position, a trajectory of applied pressure, or a frequency signature of a speaker when uttering a letter, syllable, phoneme, word, or phrase. Such measurable characteristics can then be mapped to entities. For example, a feature of an object can be mapped to an object, a portion of a feature can be mapped to a feature (of the object), a trajectory of position can be mapped to a posture, a trajectory of applied pressure can be mapped to (maximum) applied pressure, a speaker's frequency signature can be mapped to that speaker, and uttered letters, syllables, phonemes, words, or phrases can be mapped to actual letters, syllables, phonemes, words, or phrases. Such mapping can be simply a lookup in memory, a lookup table, or a database. This lookup can be based on finding the entity among multiple physical entities that has the characteristic closest to the identified measurable characteristic. From this search, the actual entity can be identified. When a local minimum is reached, the network 520 follows the local minimum trajectory for a user-defined amount of time. The accuracy of the identification depends on the total amount of time spent following the local minimum trajectory.

[0043] In some implementations, the input changes dynamically over time and follows a (temporally) sensor input trajectory or sensing trajectory. In some implementations, multiple sensors monitor the correlation between the sensors. Such correlation may be due to the fact that the sensors are located in different parts of the same underlying substrate, for example, the sensors are positioned close to each other and thus measure different aspects of signals that are correlated or dependent on each other. Alternatively, the quantities measured by the sensors may be correlated due to the laws governing the world monitored by the sensors. For example, when the visual world (via a camera) is mapped onto a set of sensor pixels in a first image, two adjacent sensors / pixels may have high contrast in their intensity, but not all adjacent sensors / pixels will have high intensity contrast between them, because the visual world is not structured in this way. Therefore, there may be a certain degree of predictability or correlation in the visual world. This means that if there is a high intensity contrast between the center pixel and the first pixel (e.g., the pixel to the left of the center pixel), but a much lower intensity contrast between the center pixel and other adjacent pixels (e.g., the pixel to the right of the center pixel, the pixel above the center pixel, and the pixel below the center pixel), then this relationship can also be reflected in other images, such as the second image. That is, in the second image, there may be a high intensity contrast between the center pixel and the first pixel, while there may be a much lower intensity contrast between the center pixel and other adjacent pixels.

[0044] In some implementations, a local minimum of the total activity level is reached when the sensor input trajectory is followed with a deviation smaller than a user-defined deviation threshold over a time period longer than a user-defined time threshold. In other words, a local minimum of the total activity level is reached when the internal trajectory, which is the trajectory followed by nodes 522, 524, 526, and 528, follows the sensor input trajectory well enough for a sufficient amount of time. Therefore, in some implementations, the sensor input trajectory is copied or represented by the internal trajectory. Furthermore, a local minimum is reached when the lowest possible total activity level is achieved, i.e., when the total activity level (of network 520) is as low as possible relative to the sum of the activities provided to network 520 by inputs from multiple sensors. Since the deviation threshold and time threshold are user-defined, the user can select an appropriate level of precision. Moreover, since the deviation threshold and time threshold are user-defined, it is not necessary to reach or find an actual local minimum; rather, it is sufficient for the total activity level to be near or close to a local minimum, depending on the set deviation threshold and time threshold, thereby allowing deviations from the internal trajectory. In addition, if the total activity level is not within the set deviation threshold and the set time threshold, the method may optionally include a step of reporting that no entity can be identified with appropriate accuracy / determinism.

[0045] According to some implementations, the computer program product includes a non-transitory computer-readable medium 200, such as, for example, a Universal Serial Bus (USB) memory, a plug-in card, an embedded driver, a Digital Universal Optical Disc (DVD), or a Read-Only Memory (ROM). Figure 2 An example computer-readable medium in the form of an optical disc (CD) ROM 200 is shown. The computer-readable medium stores thereon a computer program including program instructions. The computer program can be loaded into a data processor (PROC) 220, which may, for example, be included in a computer or computing device 210. When the computer program is loaded into a data processing unit, it can be stored in a memory (MEM) 230 associated with or included in the data processing unit. According to some embodiments, the computer program, when loaded into and executed by the data processing unit, can cause, for example, events as described herein. Figure 1 The execution of the method steps shown.

[0046] Figure 3 This is a schematic block diagram illustrating an example device according to some embodiments. Figure 3 A device 300 for entity recognition is shown. The device 300 can be configured to cause, for example... Figure 1 The execution of one or more of the method steps shown or otherwise described herein (e.g., apparatus 300 may be configured to perform such...). Figure 1 (One or more of the method steps shown or otherwise described herein). Apparatus 300 includes a control circuitry system 310. Control circuitry 310 is configured to provide inputs from multiple sensors (and...) to a network 520 of nodes 522, 524, 526, 528. Figure 1 Compared to step 110); this allows each node 522, 524, 526, 528 of network 520 to generate an activity level based on inputs from multiple sensors (compared to step 110). Figure 1 Compared to step 120); this allows comparing the activity levels of each node 522, 524, 526, and 528 with threshold levels (compared to...). Figure 1 Compared to step 130); based on this comparison, for each node 522, 524, 526, 528, the activity level is set to a preset value or the generated activity level is maintained (compared to step 130). Figure 1 Compared to step 140); this makes the total activity level calculated as the sum of the activity levels of all nodes 522, 524, 526, and 528 of network 520 (compared to step 140). Figure 1 Compared to step 150); this allows for iterative processes of providing, generating, comparing, setting / holding, and calculating until a local minimum of the total activity level is reached (compared to step 150). Figure 1Compared to step 160); and enabling the identification of measurable properties of an entity by utilizing the distribution of activity levels at the local minimum when a local minimum of the total activity level is reached (compared ... Figure 1 Compared to step 170).

[0047] The control circuitry 310 may include or otherwise associate with the following: a provider (e.g., a providing circuitry or providing module) 312, which may be configured to provide inputs from multiple sensors to a network of nodes; a generator (e.g., a generating circuitry or generating module) 314, which may be configured to generate an activity level for each node in the network based on the inputs from multiple sensors; a comparator (e.g., a comparing circuitry or a comparing module) 316, which may be configured to compare the activity level of each node with a threshold level; and a setter / holder (e.g., a setter / holder circuitry or a setter / holder module) 318, which may be configured to: Based on this comparison, for each node, the activity level is set to a preset value or the generated activity level is maintained; a calculator (e.g., a computing circuit system or computing module) 320, which can be configured to calculate the total activity level as the sum of the activity levels of all nodes in the network; an iterator (e.g., an iterating circuit system or iterating module) 322, which can be configured to iterate over providing, generating, comparing, setting / maintaining, and calculating until a local minimum of the total activity level is reached; and an exploiter (e.g., an exploiting circuit system or exploiting module) 324, which can be configured to: when a local minimum of the total activity level is reached, use the distribution of activity levels at the local minimum to identify measurable characteristics of the entity.

[0048] Figure 4 The operating principle of the device is illustrated using an example with multiple sensors and multiple nodes / processing elements according to some implementations. Figure 4 At point 1, network 420, including nodes i1, i2, i3, and i4, is activated by energy X. In some examples, such as... Figure 4 The network of nodes i1, i2, i3, and i4, schematically illustrated, exhibits nonlinear attractor dynamics. More specifically, the network of nodes is an attractor network. Furthermore, nonlinearity is introduced by setting a preset value for the activity level based on a comparison 130 between the activity level and a threshold level for each node. The internal state of network 420 is equivalent to the distribution of activity across nodes i1 to i4. The internal state evolves over time according to the structure of network 420. Figure 4 In two instances, the internal state of network 420 is used to trigger the generation of sensor-activated movements or simply to perform matrix operations on sensor data. In some examples, such as... Figure 4As shown, network 420 generates asynchronous data that leads to sensor activation. Figure 4 At three locations, sensors j1 to j4 in sensor network 430 measure the external state of the surrounding world, which is in Figure 4 The image is schematically illustrated as an object sensed by multiple sensors (e.g., bio-touch sensors). In some implementations, for example, if the internal state of network 420 is used to trigger movement that generates sensor activation (actuator activation through changes in sensor activation caused by the internal state), the relationship between the sensor and the external world can be altered. Figure 4 At point 4, sensor network 430 is shown generating, for example, asynchronous data, which is fed to network 420. Sensors j1 to j4 of sensor network 430 are always mechanically correlated with visual or audio signals and / or due to the physical properties of the external world, and their correlation can be compared to a network with state-related weights. Activation energy X will affect the location of local minima of the total activity level. Furthermore, activation energy X can also drive the output from i1 to i4 to j1 to j4 (or actuators) (the distribution of activity levels across all nodes). Therefore, activation energy X is useful for making the sensor input trajectory a function of the internal state (e.g., a function of the internal trajectory). Activation energy X can be an initial guess, i.e., an expectation, for a particular entity, or activation energy X can be a request for a specific piece of information / entity under a given sensing condition, known to consist of a combination of many entities.

[0049] Figure 5 The operating principle of the device is illustrated using multiple sensors and multiple processing elements or nodes according to some embodiments. More specifically, Figure 5 The network 520, with nodes 522, 524, 526, and 528, shows that each node has a corresponding input 502, 504, 506, and 508 from a corresponding sensor (not shown).

[0050] Figure 6 This is a schematic diagram illustrating the operating principle of a device using multiple interconnected nodes or processing elements according to some implementation methods. Figure 6A network 520 is shown with nodes 522, 524, 526, and 528, each node connected to all other nodes 522, 524, 526, and 528. If all nodes 522, 524, 526, and 528 are connected to all other nodes 522, 524, 526, and 528, a system / method with the greatest potential difference can be obtained. Therefore, the potential for maximum performance richness is achieved. In this scheme, each added node can increase the richness of the performance. To achieve this, the precise distribution of connections / weights between nodes becomes an important licensing factor.

[0051] Figure 7 This is a schematic diagram illustrating the operating principle of the device based on examples of some implementations using multiple sensors and multiple nodes / processing elements. Figure 7A network 520 is shown with nodes 522, 524, 526, and 528, each node connected to all other nodes 522, 524, 526, and 528 via connections. Furthermore, each node 522, 524, 526, and 528 is provided with at least one input from multiple sensors, such as inputs 512 and 514. In some embodiments, the activity level of each node 522, 524, 526, and 528 is used as an input to all other nodes 522, 524, 526, and 528 via connections, and each input is weighted using weights (input weights, such as synaptic weights). At least one weighted input is negative. One way to achieve this is by utilizing at least one negative weight. Alternatively or additionally, at least one weighted input is positive. One way to achieve this is by utilizing at least one positive weight. In one implementation, some nodes 522 and 524 influence all other nodes with weights ranging from 0 to +1, while other nodes 526 and 528 influence all other nodes with weights ranging from -1 to 0. Alternatively or additionally, all maintained generated activity levels are positive scalars. By combining the use of negative weights with the case where all maintained generated activity levels are positive scalars, and combining the use of negative weights with the case where a preset value for setting all other activity levels is, for example, zero, at any given time point, some nodes (which were not below a threshold level at the previous time point) can fall below the threshold level used to generate the output. This means that the effective structure of the network can change dynamically during the recognition process. In this implementation, the method differs from that using Hopfield networks not only in that a threshold is applied to the activity level of each node, but also in that the nodes in this implementation only have positive scalar outputs but can generate negative inputs. Furthermore, maximum performance richness can be achieved if all nodes / neurons are interconnected and all inputs from the sensor target all nodes / neurons of the network. In this scenario, inputs from sensors will trigger different states in the network, depending on the exact spatiotemporal pattern of the sensor input (and the temporal evolution of that pattern).

[0052] In addition, as mentioned above Figure 1As explained, a local minimum is reached when the lowest possible total activity level is achieved, i.e., when the total activity level (of network 520) is as low as possible relative to the sum of activities provided to network 520 by inputs from multiple sensors. However, if the inputs from multiple sensors are the product of the activities of nodes 522, 524, 526, and 528 of network 520 (and nodes 522, 524, 526, and 528 of the network are driven by activation energy X), the most efficient solution of simply setting all input weights to zero will not work. In fact, each of nodes 522, 524, 526, and 528 of network 520 also has a driving mechanism to actively avoid all input weights becoming zero. More specifically, in some embodiments, each of nodes 522, 524, 526, and 528 of network 520 has means / mechanisms to prevent all of its input weights from becoming zero. Furthermore, in some embodiments, network 520 has additional means / mechanisms to prevent all sensor input weights from being zero.

[0053] Figure 8 This is a schematic diagram illustrating the operating principle of the device using multiple tactile sensors. More specifically, Figure 8 This illustrates how the same type of sensor correlation, defining a dynamic characteristic, can occur under two different sensing conditions—contact or extension against a rigid surface and contact or extension against a compliant surface. In some embodiments, the multiple sensors are touch sensors or tactile sensors. Tactile sensors can be, for example, an array of pressure sensors or bio-inspired tactile sensors. For example, each of the tactile sensors in an array senses whether the surface is touched or extended, for example, by a finger, pen, or other object, and is activated by touch / extension. The tactile sensor is located at the finger, pen, or other object. If the tactile sensor is activated, it outputs a touch event signal, such as +1. If the tactile sensor is not activated, it outputs a no-touch event signal, such as 0. Alternatively, if the tactile sensor senses that the surface is touched, it outputs a touch event signal with a force-related value, such as a value between 0 and +1, and if the tactile sensor does not sense that the surface is touched, it outputs a no-touch event signal, such as 0. The outputs of the tactile sensors are provided as inputs to network 520, nodes 522, 524, 526, and 528. Therefore, the input from each of the multiple sensors is a touch event (e.g., a force-related value) or a non-touch event. In some embodiments, the surface being touched / stretched is a rigid (non-compliant) surface. Figure 8 As shown in Figure A, when, for example, a finger begins to touch the surface, only one or a few tactile sensors in the array can sense this as a touch event. Subsequently, as... Figure 8 B to Figure 8As shown in D, when a compliant finger is pushed against the surface with a constant force, this results in contact involving a larger surface area, thus allowing more sensors to detect the touch event. Figure 8 A to Figure 8 As shown in D, if the threshold level (using the combination) Figure 1 The described method is selected such that only nodes 522, 524, 526, 528 with activity levels from the tactile sensors receiving the highest shear force are considered active. Only these sensors, which maintain the generated activity levels, are considered active. Figure 8 A to Figure 8 The edge / perimeter of the circle in D represents the sensor in the central region. Therefore, when a finger is pushed against the surface with a constant force, the contact involves a gradually increasing surface area, and the resulting level of activity over time can be described as a radially outward-traveling wave; that is, the following sensor input trajectory is a radially outward-traveling wave involving a predictable sequence of sensor activations. As the finger is lifted and thus contacts the surface less, the following sensor input trajectory is a radially inward-traveling wave across the skin's sensor array. The trajectory can be used to identify new contact events and / or the end of contact events. The trajectory can also be used, for example, to distinguish between different types of contact events by comparing the trajectory with the trajectories of known contact events. Alternatively, both can follow the same overall trajectory, but with the addition of some adjacent trajectory paths / components that may or may not be detected by the system, depending on the threshold used for identification. By utilizing the trajectory used for identification, new contact events (or the end of contact events) can be identified as spatiotemporal sequences or qualitative events of the same type, regardless of the amount of finger force applied and the absolute level of the resulting shear force, i.e., regardless of how quickly the finger is applied to the surface (where the speed of finger movement can also depend, for example, on the aforementioned activation energy X). Furthermore, identification is independent of the presence or absence of faults / noise in one or more sensors, thus leading to robust identification.

[0054] In some implementations, the tactile sensor or array of tactile sensors contacts / extends against a compliant surface. In the case of a compliant surface, the intermediate region is instead... Figure 8 E to Figure 8H shows an increase, and the intermediate region will widen. However, the overall sensor activation relationship (sensor input trajectory) remains unchanged, and if the threshold level is set to a sufficiently permissible / low level, the method will end with the same local minimum of total system activity, and contact opening features (new contact events and / or the end of contact events) are being identified. Alternatively, the distribution of activity levels across all nodes can be used to identify the sensor input trajectory as a posture. Rigid surfaces can be used to identify, for example, posture, while compliant surfaces can be used to identify, for example, the maximum shear force applied within a certain interval.

[0055] Figure 9 This is a schematic diagram illustrating the operating principle of the device together with a camera. In some embodiments, the input from multiple sensors is pixel values. Pixel values ​​can be intensity values. Alternatively or additionally, pixel values ​​can be one or more component intensities representing a color—for example, red, green, and blue; or cyan, magenta, yellow, and black. Pixel values ​​can be generated by a camera 910, such as a digital camera (e.g., a digital camera). Figure 9 (As shown) A portion of the captured image. Furthermore, the image can be an image captured sequentially. Alternatively, the image can be a subset of the captured images, such as every other image in a sequence. Utilizing combination Figure 1 The described method utilizes the distribution of activity levels across all nodes 522, 524, 526, and 528 to control the positioning of camera 910 via rotational and / or translational movement, thereby controlling the sensor input trajectory. Rotational and / or translational movement can be performed by actuators 912 (such as one or more motors) configured to rotate the camera / by an angle and / or move the camera forward / backward or left / right. Therefore, the distribution of activity levels across all nodes 522, 524, 526, and 528 is used as a feedback signal to actuator 912. Figure 9In this process, camera 912 brings object 920 into its field of focus. Therefore, object 920 will appear in one or more of the captured images. Object 920 can be a person, a tree, a house, or any other suitable object. By controlling the camera's angle or position, the input from multiple pixels is influenced / altered. The sensor signal, i.e., the pixel, then becomes a function of the distribution of activity levels across all nodes 522, 524, 526, 528, and thus a function of its own internal state. The active movement of camera 912 generates a flow of sensor input that evolves over time. Therefore, the input changes dynamically over time and follows the sensor input trajectory. The sensor input trajectory is controlled by the movement of camera 912, and when a local minimum of the total activity level has been reached, the distribution of activity levels at that local minimum is used to identify measurable characteristics of an entity, such as features of an object or a portion of a feature; if the measurable characteristic is a feature of an object, then the entity is that object, and if the measurable characteristic is a portion of a feature, then the entity is that feature. Features can be biometrics, such as the distance between two biometric points, like the distance between a person's eyes. Alternatively, the feature can be the width or height of an object, such as the width or height of a tree. In some implementations, the number of pixels used as input can decrease as the distance between object 920 and camera 912 increases, and the number of pixels used as input can increase as the distance between object 920 and camera 912 decreases, thereby ensuring that the object or features of the same object are identified as the same entity. For example, if the distance between object 920 and camera 912 is doubled, only a quarter of the pixels used as input at the shorter distance—that is, the pixels covering the object at the longer distance—are used as input at the longer distance. Therefore, regardless of the distance, the object / entity can be identified as the same object / entity because even though the number of sensors / pixels occupied will be less when the object is further away, sensor correlation will be able to be identified as being of the same nature when the camera scans the object. In another example, the feature to be identified is a vertical contrast line between a white area and a black area. If camera 912 scans across a vertical contrast line (e.g., from left to right), the four sensors will detect the same features (i.e., the vertical contrast line) as the 16 or 64 sensors. Therefore, there exists a central feature element that becomes relevant to the activation of a particular type of sensor, traveling across the sensors as the camera moves. Furthermore, the speed of camera movement can be controlled. When camera 912 moves at an increasing speed, the pixels used as input can come from fewer images, such as every other image, while when camera 912 moves at a decreasing speed, the pixels used as input can come from more images, thus ensuring that objects or features of objects are identified as the same entity regardless of speed.Furthermore, entity recognition and / or the recognition of measurable characteristics of entities can be associated with an acceptance threshold. That is, when comparing an entity and / or its measurable characteristics with known physical entities or their characteristics, or matching a physical entity or its characteristics with actual physical entities or their characteristics in a memory, lookup table, or database (by means of the distribution of activity levels at a found local minimum), an acceptance threshold can be used to determine if a match exists. The acceptance threshold can be user-set, i.e., user-definable. The use of the acceptance threshold ensures that features or objects can be identical even if the activity distribution across sensors is not entirely identical. Because the activity distribution across sensors does not need to be identical to determine if a match exists, correct recognition can be achieved even if one or more sensors / pixels are faulty or noisy; that is, it is relatively robust to noise. Therefore, features or objects can be identified at near / short distances, but the same features can also be identified at greater distances; that is, features or objects can be identified independently of distance. The same reasoning applies to two objects of different sizes but with the same features at the same distance. In both cases, the total sensor activity changes, but the overall spatiotemporal relationship they activate can still be identified.

[0056] Figures 10A to 10C The frequency and power of different sensors used to sense audio signals are shown. (As shown from...) Figure 10A As can be seen, the spectrum of an audio signal can be divided into different frequency bands. The power or energy in different frequency bands can be sensed (and reported) by sensors. Figure 10B The figure shows the power in the frequency band sensed by sensor 1 and sensor 2. As can be seen from the figure, each sensor senses a different frequency band. Figure 10C The diagram illustrates a sensing trajectory based on power sensed by sensors 1 and 2 over time. The audio signal can include sounds from speech (and possibly other sounds) containing dynamic variations in power or energy across several frequency bands. Therefore, for each spoken syllable, speech can be identified as belonging to a specific individual (within a group of individuals) based on a specific dynamic signature. The dynamic signature includes the specific variation in power or energy levels within each frequency band over a given time period. Thus, by dividing the spectrum into frequency bands and having sensors sense the power or energy in each of multiple frequency bands (using one or more sensors for each band), a specific sensor input trajectory is created. Therefore, a combined input from multiple such sensors follows the sensor input trajectory. In some embodiments, each of the multiple sensors is associated with a frequency band of the audio signal. Preferably, each sensor is associated with a different frequency band. Each sensor senses (and reports) the power or energy present in the frequency band associated with that sensor over a given time period. Utilizing a combination... Figure 1 The described method reaches a local minimum, where the distribution of activity levels across all nodes 522, 524, 526, and 528 at the local minimum is used to identify the speaker. Alternatively or additionally, the activity levels across all nodes 522, 524, 526, and 528 can be used to identify spoken letters, syllables, phonemes, words, or phrases present in the audio signal. For example, syllables are identified by comparing the distribution of activity levels across all nodes 522, 524, 526, and 528 at the found local minimum with a stored distribution of activity levels associated with known syllables. Similarly, speakers are identified by comparing the distribution of activity levels across all nodes 522, 524, 526, and 528 at the found local minimum with a stored distribution of activity levels associated with known speakers. Figure 9 The described acceptance threshold can also be used to determine whether a match exists (for syllables or speakers). Furthermore, when the trajectory is followed for recognition and reaches a local minimum, the recognition is independent of the speed and volume of the sound signal.

[0057] Figure 11 It is a graph of the input trajectory of another sensor. Figure 11 The sensor input trajectories based on three sensors (sensor 1, sensor 2, and sensor 3) over time are shown. Figure 11 As seen, the measurement value of sensor 1 over time is used as the X-coordinate in a Cartesian coordinate system of three-dimensional space, the measurement value of sensor 2 over time is used as the Y-coordinate, and the measurement value of sensor 3 over time is used as the Z-coordinate. In one embodiment, the sensors measure different frequency bands of the audio signal. In another embodiment, the sensors measure the presence of a touch event. In yet another embodiment, the sensors measure the intensity value of a pixel. Figure 11 The coordinates drawn in the image together constitute the sensing trajectory.

[0058] Those skilled in the art will recognize that this disclosure is not limited to the preferred embodiments described above. They will also recognize that modifications and variations can be made within the scope of the appended claims. For example, other entities such as fragrance or taste can be identified. Furthermore, based on a study of the drawings, the disclosure, and the appended claims, those skilled in the art can understand and implement variations of the disclosed embodiments in practicing the claimed disclosure.

Claims

1. A computer-implemented or hardware-implemented method (100) for recognizing measurable characteristics of entities used for entity recognition, the method comprising: a) Provide (110) inputs from multiple sensors to a recurrent neural network comprising multiple data processing nodes; b) Each data processing node of the recurrent neural network generates a (120) node activity level based on the input from the plurality of sensors to the recurrent neural network; c) Compare the node activity level of each data processing node of the recurrent neural network with a threshold level (130). d) Based on the comparison, for each data processing node of the recurrent neural network, the node activity level is set (140) to a preset value or the generated node activity level is maintained, wherein if the generated node activity level is higher than the threshold level, the generated activity level is maintained, and if the generated node activity level is equal to or lower than the threshold level, the node activity level is set to the preset value. e) The total network activity level is calculated (150) as the sum of the activity levels of all nodes in the data processing node of the recurrent neural network; f) Iterate (160) from a) to e) until a local minimum of the total network activity level is reached; and g) When a local minimum of the total network activity level is reached, the distribution of node activity levels at the local minimum is used (170) to identify the measurable characteristics of the entity. The inputs from the plurality of sensors change dynamically over time and follow the sensor input trajectory. A local minimum of the total network activity level is reached when the internal trajectory followed by all data processing nodes of the recurrent neural network follows the sensor input trajectory with a deviation less than the user-defined deviation threshold over a time period longer than the user-defined time threshold.

2. The method implemented in a computer or hardware according to claim 1, wherein, The user-definable deviation threshold is user-defined.

3. The computer-implemented or hardware-implemented method according to claim 1 or 2, wherein, The user-definable time threshold is user-defined.

4. The method implemented in a computer or in hardware according to any one of claims 1 to 2, wherein, The multiple sensors monitor the correlation between the sensors.

5. The method implemented in a computer or in hardware according to any one of claims 1 to 2, wherein, The node activity level of each data processing node is used as input to all other data processing nodes, each input is weighted using weights, and wherein at least one weighted input is negative, and / or wherein at least one weighted input is positive, and / or wherein all maintained generated node activity levels are positive scalars.

6. The method implemented in a computer or in hardware according to any one of claims 1 to 2, wherein, The recurrent neural network is activated by network activation energy X, which affects the location of the local minimum of the total network activity level.

7. The method implemented in a computer or in hardware according to any one of claims 1 to 2, wherein, The input from the plurality of sensors is the pixel value of the image captured by the camera, and wherein the distribution of node activity levels across all data processing nodes is further utilized to control the positioning of the camera by rotational and / or translational movement of the camera, thereby controlling the trajectory of the sensor input, and wherein the identified entity is an object or a feature of an object present in at least one of the captured images.

8. The method implemented in a computer or hardware according to claim 7, wherein, The pixel value includes intensity.

9. The method implemented in a computer or in hardware according to any one of claims 1 to 2, wherein, The plurality of sensors are touch sensors, and the input from each of the plurality of sensors is a touch event signal with a force-related value, wherein the distribution of node activity levels across all data processing nodes is used to identify the sensor input trajectory as a new contact event, the end of a contact event, a gesture, or as applied pressure.

10. The method implemented in a computer or in hardware according to any one of claims 1 to 2, wherein, Each of the plurality of sensors is associated with a different frequency band of the audio signal, wherein each sensor reports the energy present in the associated frequency band, and wherein combined inputs from a plurality of such sensors follow the sensor input trajectory, and wherein the distribution of activity levels across all data processing nodes is used to identify the speaker and / or spoken letters, syllables, phonemes, words or phrases present in the audio signal.

11. The method implemented in a computer or in hardware according to any one of claims 1 to 2, wherein, The method also includes identifying entities based on the measurable characteristics.

12. A computer program product comprising a non-transitory computer-readable medium (200), having on the non-transitory computer-readable medium a computer program including program instructions, the computer program being loadable into a data processing unit (220) and configured to cause the execution of the method according to any one of claims 1 to 11 when the computer program is run by the data processing unit (220).

13. An apparatus (300) for identifying measurable characteristics of an entity for entity recognition, the apparatus comprising a control circuitry (310) configured such that: a) Provide inputs from multiple sensors to a recurrent neural network that includes multiple data processing nodes; b) Each data processing node of the recurrent neural network generates a node activity level based on the input from the plurality of sensors to the recurrent neural network; c) Compare the node activity level of each data processing node with the threshold level; d) Based on the comparison, for each data processing node, the node activity level is set to a preset value or the generated activity level is maintained, wherein if the generated node activity level is higher than the threshold level, the generated activity level is maintained, and if the generated node activity level is equal to or lower than the threshold level, the node activity level is set to the preset value. e) Calculate the total network activity level as the sum of the activity levels of all nodes in the data processing node of the recurrent neural network; f) Iterate through a) to e) until a local minimum of the total network activity level is reached; and g) When a local minimum of the total network activity level is reached, the distribution of node activity levels at the local minimum is used to identify the measurable characteristics of the entity. The inputs from the plurality of sensors change dynamically over time and follow the sensor input trajectory. A local minimum of the total network activity level is reached when the internal trajectory followed by all data processing nodes of the recurrent neural network follows the sensor input trajectory with a deviation less than the user-defined deviation threshold over a time period longer than the user-defined time threshold.

Citation Information

Patent Citations

  • Method for computer-assisted processing of measured values detected in a sensor network

    CN101276435A

  • Model generation method and device, entity identification method and device and electronic equipment

    CN111027325A