System and method for training a recurrent neural network for device location identification in an environment

The double-regressor neural network with a three-stage training procedure addresses the challenges of domain adaptation in regression tasks by extracting domain-invariant features, enabling accurate device localization across different environments.

JP2025517820AActive Publication Date: 2025-06-10MITSUBISHI ELECTRIC CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025515037
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-10-20
Filing Date
2023-06-22
Publication Date
2025-06-10
Estimated Expiration
2043-06-22

AI Technical Summary

Technical Problem

Existing domain adaptation methods for regression tasks, such as indoor Wi-Fi localization, face challenges in adapting features across different domains due to differences in loss functions and the lack of adaptability in regression models compared to classifiers.

Method used

A double-regressor neural network with a three-stage training procedure is introduced, where the network includes a feature extractor and two regressors with the same architecture. The three-stage training process involves initial training with labeled data, followed by training with both labeled and unlabeled data to differentiate between domains, and finally, using an adversarial discriminator to extract domain-invariant features.

Benefits of technology

The proposed method effectively adapts regression models to new environments by extracting domain-invariant features, allowing for accurate localization of devices even in environments with different statistical distributions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025517820000001_ABST
    Figure 2025517820000001_ABST
Patent Text Reader

Abstract

The present disclosure provides a method and a system for training a neural network suitable for identifying the position of a device in an environment based on signals received by the device. The method includes training a double-regressor neural network to identify a position from labeled data, the double-regressor neural network including a feature extractor and a double-regressor including two regressors, the method further including training the parameters of the double-regressor using labeled data and unlabeled data such that each of the two regressors identifies the same labeled position while processing the labeled data and different positions while processing the unlabeled data, and training the parameters of the feature extractor using an adversarial discriminator to extract domain-invariant features having statistical characteristics of the labeled data from the unlabeled data according to the adversarial discriminator such that each of the two regressors identifies the same position while processing the domain-invariant features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to device localization, and more specifically, to systems and methods for training a regression neural network for device localization in an environment.

Background Art

[0002] Domain adaptation (DA) aims to transfer knowledge from a sufficiently labeled source domain to facilitate label-free target learning. When looking at specific tasks such as indoor (Wi-Fi) localization, it is essential to learn a cross-domain regressor to mitigate domain shift. However, in contrast to classification tasks, indoor (Wi-Fi) localization, for example, relies on regression because the trained neural network needs to identify positions within a continuous space, but the training data can actually only be obtained from samples within discrete positions.

Summary of the Invention

Problems to be Solved by the Invention

[0003] Regression, as a counterpart to classification, is a major paradigm with a wide range of applications. Domain adaptation regression (DAR) extends it by generalizing a regressor from a labeled source domain to an unlabeled target domain. Deep learning has brought about remarkable changes in various regression applications across many fields. Nevertheless, training high-quality deep models relies on large-scale labeled datasets. Furthermore, in many real-world regression applications, accurately annotating rich training instances is time-consuming and labor-intensive. A solution to such problems is to utilize off-the-shelf labeled data from related domains and apply domain adaptation techniques to overcome domain shift or dataset bias.

[0004] Deep domain adaptation methods have achieved remarkable progress in the domain adaptation classification (DAC) problem. However, despite these developments in DAC, the learning of invariant representations in the deep regime of DAR remains undeveloped. One reason for such a deficiency is the difference in loss functions during the training of classifiers and regressors. The loss function commonly used in regression is the squared loss (L2), while in classification, it is the cross-entropy loss (CE) with the softmax activation function. Softmax enables the activation values of different categories to compete with each other. By introducing this competition mechanism, the classifier can be quickly adapted to changes in the feature scale. However, in the regression task, the regressor may not have such adaptability.

[0005] Therefore, a domain adaptation method for the regressor is needed.

Means for Solving the Problem

[0006] Some embodiments are based on the recognition that the problems in the domain adaptation of regression neural networks can be addressed using an adversarial discriminator applied at the feature extraction level. The purpose of such a discriminator is to train the feature extractor to extract features from the measurements of the target domain that have a similar statistical distribution to the features extracted from the measurements of the source domain. However, reaching such a discriminator has the problem that an additional judgment unit is required to determine the quality of the extracted features because the feature extractor is neither a classifier nor a regressor. Such a judgment unit is not present in typical regression architectures used for device localization.

[0007] To that end, some embodiments disclose a double-regressor neural network and a three-stage training procedure for domain adaptation regression. The double-regressor neural network includes a feature extractor and a double-regressor including two regressors. The two regressors have the same architecture with potentially different weights and biases. The three-stage training procedure is used to train the double-regressor neural network for localizing devices in an environment based on signals received by devices located within the environment. The signals can include different types of radio frequency (RF) signals, such as received signal strength indicator (RSSI) signals, channel state information (CSI) signals, at various signal levels. The different types of RF signals include ultra-wideband (UWB) signals, Wi-Fi signals, and inertial measurement unit (IMU) signals. The three-stage training procedure includes training the double-regressor neural network in the environment using labeled data and / or in a modified or new environment using unlabeled data. In one embodiment, the labeled data includes measurements of signals, and the measurements are labeled with the coordinates of the location of the measurements. In one embodiment, the unlabeled data includes measurements of signals at unidentified locations within the environment, and the measurements of the labeled data and the unlabeled data are sampled at the same or different discrete locations of the environment.

[0008] The following describes each training stage of the three-stage training procedure. During the first training stage, the double-regressor neural network is trained using labeled data to identify positions within the continuous space of the environment from the labeled data. In particular, the feature extractor is trained to extract features from the labeled data, and each regressor is trained to determine a position within the continuous space of the environment from the features extracted by the feature extractor. In other words, the goal of each regressor is to learn a regression that can map the extracted features to positions within the continuous space of the environment. The need for two regressors stems from the need to have a determination unit for the features extracted by the feature extractor and the ability to determine whether the input sensor signal is collected from an environment similar to the training data or from an unrecognized environment.

[0009] During the first stage of training, the feature extractor and the two regressors are trained with the labeled data based on the loss function L r As a result, the parameters of the two regressors are most likely to be the same or very similar. Thus, when both regressors process the features of the labeled data, both regressors produce similar outputs that approximate the ground truth data. In one embodiment, the loss function L r is based on the mean squared error (MSE) criterion between the labeled data and the coordinate estimates.

[0010] During the second stage of training, the two regressors are trained using labeled data and unlabeled data, generating the same correct output for the labeled data but different outputs for the unlabeled data from different domains. In this way, the two regressors are trained to generate different outputs for unlabeled data when the unlabeled data has a statistical distribution different from that of the labeled data. As a result, the two regressors are trained to identify whether the statistical distribution of the unlabeled data has a distribution similar to or different from that of the labeled data. As a result, these two regressors can serve as a determination unit regarding the similarity between the input statistical distribution and the statistical distribution of the labeled data.

[0011] Furthermore, the difference in the outputs of both regressors for unlabeled data can be maximized by minimizing the overlap / similarity between the outputs of both regressors. To minimize the overlap between the outputs of both regressors, it is necessary to quantify the overlap between the dual-regressor outputs. To quantify such overlap, some embodiments use the Jaccard similarity coefficient (IoU score) from the field of object detection using visual sensors. However, the non-differentiable nature of the IoU score poses a problem for the optimization of the dual-regressor neural network weights. To mitigate such problems, the soft similarity function L s is realized. Minimization of the soft similarity function L s reduces the overlap between the outputs of the two regressors and further increases the difference in the outputs of the two regressors for unlabeled data.

[0012] However, it is still necessary that both regressors should generate the same output for the labeled data. For this purpose, in the second stage of training, not only the soft similarity function L s is minimized, but also the objective function based on the loss function L r (used in the first stage of training) and the soft similarity function L s is minimized.

[0013] Therefore, the combination of the two regressors can be envisioned as an implicit discriminator for directly distinguishing labeled data from unlabeled data (i.e., detecting source samples / target samples outside the support).

[0014] Given these new capabilities of these two regressors, during the third training stage, the feature extractor is trained using an adversarial discriminator to use these two regressors as discriminators to extract domain-invariant features of the unlabeled data using the statistical distribution of the labeled data. The extracted domain-invariant features are such that when processed by the two regressors, each of the two regressors identifies the same location while processing the domain-invariant features.

[0015] Some embodiments are based on the recognition that domain adaptation becomes difficult when the distribution differences are large. As used herein, the distribution difference refers to the difference in the outputs of the two regressors. For example, in the case of multimodal indoor positioning, the distribution difference becomes relatively prominent due to the fact that the change in one multipath component can contribute to the overall sensor signal in either a constructive or destructive manner. Some embodiments are based on the recognition that such problems can be solved by constructing two intermediate domains between the source domain and the target domain and gradually eliminating their mismatches to achieve statistical distribution alignment. The intermediate domains are constructed based on data augmentation of the labeled data and the unlabeled data.

[0016]

Number

[0017] After executing the above three-stage training procedure, a trained double-regressor neural network is obtained. The trained double-regressor neural network includes a trained feature extractor and two trained regressors. The trained double-regressor neural network can be used to determine the position for the input unlabeled data. Here, the two trained regressors of the trained double-regressor neural network generate similar positions for the input unlabeled data. For this purpose, some embodiments are based on the recognition that the architecture of the trained double-regressor neural network can include only a single trained regressor instead of two trained regressors for determining the position. Alternatively, in some embodiments, the trained double-regressor neural network may include a trained feature extractor and the original regressor (the regressor obtained from the first stage of training) that can provide additional robustness of the regression output.

[0018] Accordingly, one embodiment discloses a computer-implemented method for training a neural network suitable for determining the location of a device in an environment based on signals received by a device located in the environment.The method uses a processor together with stored instructions that implement the method, and when executed by the processor, the instructions perform the steps of the method, the steps including collecting labeled data including a measured value of the signal, the labeled data being labeled with coordinates of the position of the measured value, the steps further including collecting unlabeled data including a measured value of the signal at an unidentified position within the environment, the measured value of the labeled data and the measured value of the unlabeled data being sampled at the same or different discrete positions of the environment, the steps further including training a bi-regressor neural network to identify a position within the continuous space of the environment from the labeled data during a first training stage, the bi-regressor neural network including a feature extractor configured to extract features from the labeled data and a bi-regressor including two regressors having the same architecture, each regressor being trained to determine a position within the continuous space of the environment from the features received from the feature extractor, the steps further including training parameters of the bi-regressor using the labeled data and the unlabeled data during a second training stage such that each of the two regressors identifies the same labeled position while processing the labeled data and different positions while processing the unlabeled data, using fixed parameters of the feature extractor trained during the first training stage, and training parameters of the feature extractor using the labeled data and the unlabeled data during a third training stage using a fixed parameter of the bi-regressor trained during the second training stage such that each of the two regressors identifies the same position while processing domain invariant features, and extracting domain invariant features from the unlabeled data using statistical characteristics of the labeled data according to an adversarial discriminator.

[0019] Accordingly, another embodiment discloses a system for training a neural network suitable for identifying the location of a device within an environment based on signals received by a device located within the environment.When the system includes a processor and a memory storing instructions, which, when executed by the processor, cause the system to collect labeled data including the measured values of the signal, the labeled data being labeled with the coordinates of the positions of the measured values, and the instructions further cause the system to collect, when executed by the processor, unlabeled data including the measured values of the signal at unidentified positions within the environment, the measured values of the labeled data and the measured values of the unlabeled data being sampled at the same or different discrete positions of the environment, and the instructions further cause the system to train a double-regressor neural network to identify positions within the continuous space of the environment from the labeled data during a first training stage using the labeled data, the double-regressor neural network including a feature extractor configured to extract features from the labeled data and a double-regressor including two regressors having the same architecture, each regressor being trained to determine a position within the continuous space of the environment from the features received from the feature extractor, and the instructions further cause the system to train the parameters of the double-regressor using the labeled data and the unlabeled data during a second training stage such that each of the two regressors identifies the same labeled position while processing the labeled data and different positions while processing the unlabeled data, using the fixed parameters of the feature extractor trained during the first training stage, and to extract domain-invariant features from the unlabeled data using the statistical characteristics of the labeled data according to an adversarial discriminator such that each of the two regressors identifies the same position while processing the domain-invariant features, and to train the parameters of the feature extractor using the adversarial discriminator during a third training stage using the labeled data and the unlabeled data and the fixed parameters of the double-regressor trained during the second training stage.

[0020] Accordingly, yet another embodiment discloses a non-transitory computer-readable storage medium embodying a program executable by a processor for performing a method for training a neural network suitable for identifying the position of a device within an environment based on signals received by a device located within the environment.The method includes the step of collecting labeled data including measured values of the signal, where the labeled data is labeled with the coordinates of the position of the measured value. The method further includes the step of collecting unlabeled data including measured values of the signal at unidentifiable positions within the environment. The measured values of the labeled data and the measured values of the unlabeled data are sampled at the same or different discrete positions in the environment. The method further includes the step of training a double-regressor neural network to identify positions within the continuous space of the environment from the labeled data during a first training stage. The double-regressor neural network includes a feature extractor configured to extract features from the labeled data and a double-regressor including two regressors having the same architecture. Each regressor is trained to determine the position within the continuous space of the environment from the features received from the feature extractor. The method further includes the step of training the parameters of the double-regressor using the labeled data and the unlabeled data during a second training stage such that each of the two regressors identifies the same labeled position while processing the labeled data and different positions while processing the unlabeled data, using the fixed parameters of the feature extractor trained during the first training stage. The method further includes the step of training the parameters of the feature extractor using the labeled data and the unlabeled data during a third training stage such that each of the two regressors identifies the same position while processing domain-invariant features, by extracting domain-invariant features from the unlabeled data using the statistical characteristics of the labeled data according to an adversarial discriminator, using the fixed parameters of the double-regressor trained during the second training stage.

[0021] The present invention will be described in detail below with reference to the accompanying drawings. The drawings shown are not necessarily to scale; instead, emphasis is generally placed on explaining the principles of the embodiments of the present disclosure.

Brief Description of the Drawings

[0022]

Figure 1A

Figure 1B

Figure 1C

Figure 2A

Figure 2B

Figure 2C

Figure 3

Figure 4

Modes for Carrying Out the Invention

[0023] In the following description, for the purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. It will be apparent to one of ordinary skill in the art, however, that the present disclosure may be practiced without these specific details. In other instances, well-known devices and methods are shown only in block diagram form in order to avoid obscuring the present disclosure.

[0024] As used in this specification and the appended claims, the phrases “for example,” “by way of example,” “such as,” and the verbs “comprising,” “having,” “including,” and other verb forms, when used in conjunction with a list of one or more components or other items, are each to be construed as open-ended, meaning that the list is not to be considered as excluding other, additional components or items. The phrase “based on” means at least in part based on. Further, it should be understood that the expressions and terms used in this specification are for the purpose of description and should not be regarded as limiting. Any headings utilized within this description are for convenience only and have no legal or limiting effect.

[0025] Domain adaptation (DA) aims to transfer knowledge from a sufficiently labeled source domain in order to facilitate label-free target learning. When focusing on specific tasks such as indoor (Wi-Fi) localization, it is essential to learn a cross-domain regressor to mitigate domain shift. However, in contrast to classification tasks, indoor (Wi-Fi) localization, for example, relies on regression because the trained neural network needs to identify positions within a continuous space, but the training data can actually only be obtained from samples within discrete positions.

[0026] Regression is a technique for finding the relationship between independent variables or features and dependent variables or outcomes. Regression may be used as a method for predictive modeling in machine learning, where algorithms are used to predict continuous outcomes. Domain adaptation regression extends regression by generalizing a regressor from a labeled source domain to an unlabeled target domain.

[0027] For example, a regressor can be trained with labeled data in one environment. However, during the online testing phase, the regressor can receive unlabeled data in a new environment. As a result, a regressor trained with labeled data may suffer from performance degradation when the regressor processes unlabeled data. Therefore, domain adaptation techniques for regressors are needed.

[0028] Some embodiments are based on the recognition that challenges in domain adaptation of regression neural networks can be addressed using an adversarial discriminator applied at the feature extraction level. The purpose of such a discriminator is to train the feature extractor to extract features from the measurements of the target domain that have a similar statistical distribution to the features extracted from the measurements of the source domain. However, reaching such a discriminator has the problem that an additional decision part is needed to judge the quality of the extracted features because the feature extractor is neither a classifier nor a regressor. Such a decision part is not part of the typical regression architectures used for device localization.

[0029] Therefore, some embodiments disclose a double-regressor neural network and a three-stage training procedure for domain adaptation regression. FIG. 1A shows a schematic diagram of the architecture of a double-regressor neural network 100 according to some embodiments of the present disclosure. The double-regressor neural network 100 includes a feature extractor F(·) 103 and a double-regressor including two regressors 105 and 107.

Number

[0030] The following describes each training stage of the three-stage training procedure.

[0031] During the first training stage, the double-regressor neural network 100 is trained using the labeled data 101 to identify locations within the continuous space of the environment. Specifically, the feature extractor 103 is trained to extract features from the labeled data 101, and each regressor, namely 105 and 107, is trained to determine the location within the continuous space of the environment from the features extracted by the feature extractor 103. In other words, the goal of each regressor is to learn a regression that can map the extracted features to locations within the continuous space of the environment. The need for the two regressors 105 and 107 stems from the need to have a decision-making unit for the features extracted by the feature extractor 103.

[0032] In some embodiments, during the first training stage, the feature extractor 101 and the two regressors 105 and 107 are trained using the labeled data 101 based on a loss function L r As a result, the parameters of the two regressors 105 and 107 are mostly similar. Therefore, when processing the features extracted by the two regressors 105 and 107, the two regressors 105 and 107 generate similar labeled positions that approximate the ground truth data. In one embodiment, the loss function L r is based on the mean squared error (MSE) criterion. For example, the first training stage is performed in a supervised manner as follows,

Equation

[0033] FIG. 1B shows a schematic diagram of a second training stage according to some embodiments of the present disclosure. During the second stage of training, the parameters of the feature extractor 101 trained during the first training stage are fixed, and the two regressors 105 and 107 are trained using the labeled data 101 and the unlabeled data 109 to generate the same correct labeled positions for the labeled data 101, but different positions for the unlabeled data 109. Thus, the two regressors 105 and 107 are trained to generate different outputs (i.e., positions) for the unlabeled data 109 when the unlabeled data 109 has a statistical distribution different from the statistical distribution of the labeled data 101. As a result, the two regressors 105 and 107 are trained to distinguish whether the statistical distribution of the unlabeled data 109 has a distribution similar to or different from the statistical distribution of the labeled data 101. As a result, the two regressors 105, 107 can serve as a statistical distribution determination unit.

[0034] Furthermore, the difference between the outputs of the two regressors 105 and 107 for the unlabeled data can be maximized by minimizing the overlap / similarity between the two regressors 105 and 107. In order to minimize the overlap between the two regressors 105 and 107, the overlap between the outputs of the two regressors needs to be quantified. To quantify the overlap between the two regressors 105 and 107, some embodiments use the Jaccard similarity coefficient (IoU score) from the field of object detection using visual sensors. However, the non-differentiable nature of the IoU score poses a problem for the optimization of the weights of the double-regressor neural network 100. To mitigate such problems, a soft similarity function Ls111 between the concatenated outputs of the two components G and R of the two regressors is realized. Let the output of the first regression component be G(f), and the output of the second regression component be R(G(f)), which is also the output of the regression. The concatenated output h = [G, R(G(f))].

Number

[0035] However, it is still necessary that the two regressors 105 and 107 should generate the same output for the labeled data 101. For this purpose, in the second stage of training, not only the soft similarity function Ls111 is minimized, but also the loss function L r (used in the first stage of training) and the objective function based on the soft similarity function Ls111 are minimized. The objective function minimization is given as follows.

Number

[0036] Therefore, the combination of the two regressors 105 and 107 can be envisioned as an implicit discriminator for directly distinguishing between the labeled data 101 and the unlabeled data 109.

[0037] Figure 1C shows a schematic diagram of a third training stage according to some embodiments of the present disclosure. In the third training stage, the parameters of the feature extractor 103 are trained using the fixed parameters of the two regressors 105 and 107 trained during the second training stage using labeled data and unlabeled data, using the adversarial discriminator 113. During the third training stage, the parameters of the feature extractor 103 are trained by the adversarial discriminator 113 to extract domain-invariant features of the unlabeled data 109 using the statistical distribution of the labeled data 101, using the two regressors 105 and 107 as the determination unit. The extracted domain-invariant features are such that when processed by the two regressors 105 and 107, each of the two regressors 105 and 107 identifies the same position while processing the domain-invariant features.

[0038] In addition, some embodiments are based on the recognition that domain adaptation becomes difficult when the difference in distributions is large. As used herein, the difference in distributions refers to the difference in the outputs of the two regressors 105 and 107. For example, in the case of multimodal indoor positioning, the difference in distributions becomes relatively prominent. Some embodiments are based on the recognition that such problems can be solved by constructing two intermediate domains between the source domain and the target domain and gradually eliminating their mismatches to achieve statistical distribution alignment. The intermediate domains are constructed based on data augmentation of the labeled data 101 and the unlabeled data 109. First, a fixed ratio λ is selected. For example, λ is set to 0.7.

Number

[0039]

Number

[0040] After executing the above three - stage training procedure, a trained double - regressor neural network is obtained. During online, i.e., in the real - time stage, the trained double - regressor neural network can be used to determine the position of the device in the environment for the input unlabeled data.

[0041] FIG. 2A shows a schematic diagram of the architecture of a trained double - regressor neural network 200 for determining the position of a device for input unlabeled data 201 according to some embodiments of the present disclosure. The trained double - regressor neural network 200 includes a trained feature extractor 203 obtained from the third training stage and two trained regressors 205 and 207 obtained from the second training stage. The unlabeled data 201 is input into the trained feature extraction unit 203. The trained feature extractor 203 extracts domain - invariant features from the unlabeled data 201. Further, the extracted domain - invariant features are processed by the two trained regressors 205 and 207. Each of the two trained regressors 205 and 207 outputs the same position of the device in the environment.

[0042] Since the two trained regressors 205 and 207 of the trained double - regressor neural network 200 output the same position for the input unlabeled data 201, some embodiments are based on the recognition that the architecture of the trained double - regressor neural network 200 can include only one trained regressor instead of two trained regressors for determining the position. Such an architecture of the double - regressor neural network is shown in FIG. 2B.

[0043] Figure 2B shows the architecture of a double-regressor neural network 209 for determining a position, according to some other embodiments of the present disclosure. The double-regressor neural network 209 includes a trained feature extractor 203 and a trained regressor 207. The unlabeled data 201 is input into the trained feature extraction unit 203. The trained feature extractor 203 extracts domain-invariant features from the unlabeled data 201. Further, the extracted domain-invariant features are processed by the trained regressor 207 to output the position of the device within the environment.

[0044] Alternatively, in some embodiments, the double-regressor neural network 209 may include a trained regressor 205 instead of the trained regressor 207.

[0045] Some embodiments are based on the recognition that a trained double-regressor neural network for determining a position may include a trained feature extractor 203 and the original regressor (the regressor obtained from the first stage of training). Such a double-regressor neural network is shown in Figure 2C. Figure 2C shows the architecture of a double-regressor neural network 211 for determining a position, according to some other alternative embodiments of the present disclosure. The double-regressor neural network 211 includes a trained feature extractor 203 obtained from the third training stage and two regressors 105 and 107 obtained from the first training stage. Alternatively, in some embodiments, the double-regressor neural network 211 may include a trained feature extractor 203 obtained from the third training stage and one of the two regressors 105 and 107 obtained from the first training stage.

[0046] According to some embodiments, the position of a device in an environment output by a trained double-recurrent neural network (e.g., trained double-recurrent neural network 200 or 209) can be used to control the device in the environment to perform a task. For example, the device may be a robot and the environment may be an autonomous factory. Based on the position of the robot in the autonomous factory, the robot may be controlled to perform tasks such as reaching a target position, as described later with reference to FIG. 3.

[0047] FIG. 3 shows the control of a robot 301 in an autonomous factory 300 according to some embodiments of the present disclosure. A trained double-recurrent neural network 200 is embedded in the robot 301. The robot 301 also includes a controller 303. The robot 301 receives unlabeled data including measurement values of signals at an unidentified position within the autonomous factory 300. The unlabeled data is input to the trained double-recurrent neural network 200 to determine the position 305 of the robot 301 within the autonomous factory 300. Based on the position 305 of the robot 301, the controller 303 determines an operation path 307 that connects the position 303 and a target position 309 within the autonomous factory 300. Further, the controller 303 controls the robot 301 according to the operation path 307. In some alternative embodiments, the controller 303 may make a determination regarding obstacle avoidance based on the position 305 of the robot 301.

[0048] In addition, in some embodiments, the robot 301 may also be tracked within the autonomous factory 300 based on the position of the robot 301. In addition, or alternatively, the trained double-recurrent neural network 200 can be used in any environment such as an autonomous factory 300, a warehouse, or a smart home where asset (such as a robot) tracking is required or important.

[0049] FIG. 4 is a schematic diagram showing a system 400 for implementing the method of the present disclosure. The computing device 400 includes a power supply 401, a processor 403, a memory 405, and a storage device 407, all connected to a bus 409. Further, a high-speed interface 411, a low-speed interface 413, a high-speed expansion port 415, and a low-speed connection port 417 can be connected to the bus 409. In addition, a low-speed expansion port 419 is connected to the bus 409. Further, an input interface 421 can be connected to an external receiver 423 and an output interface 425 via the bus 409. A receiver 427 can be connected to an external transmitter 429 and a transmitter 431 via the bus 409. Also, an external memory 433, an external sensor 435, a machine 437, and an environment 439 can be connected to the bus 409. Further, one or more external input / output devices 441 can be connected to the bus 409. A network interface controller (NIC) 443 can be adapted to connect to a network 445 through the bus 409, and data or other data can be rendered, inter alia, on a third-party display device, a third-party imaging device, and / or a third-party printing device external to the computing device 400.

[0050] The memory 405 can store instructions executable by the computing device 400, as well as any data that can be utilized by the methods and systems of the present disclosure. The memory 405 can include a random access memory (RAM), a read-only memory (ROM), a flash memory, or any other suitable memory system. The memory 405 can be a volatile memory unit and / or a non-volatile memory unit. The memory 405 can also be another form of computer-readable medium, such as a magnetic disk or an optical disk.

[0051] The memory device 407 can be adapted to store supplementary data and / or software modules used by the computer device 400. The memory device 407 can include a hard drive, an optical drive, a thumb drive, an array of drives, or any combination thereof. Further, the memory device 407 can include a computer-readable medium such as a floppy (registered trademark) disk device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid-state memory device, or an array of devices including a device in a storage area network or other configuration. Instructions can be stored on an information carrier. When the instructions are executed by one or more processing devices (e.g., the processor 403), one or more methods such as those described above are executed.

[0052] The computing device 400 can be optionally linked via the bus 409 to a display interface or user interface (HMI) 447 adapted to connect the computing device 400 to a display device 449 and a keyboard 451. The display device 449 can include, among other things, a computer monitor, a camera, a television, a projector, or a mobile device. In some implementations, the computer device 400 may include a printer interface for connecting to a printing device, and the printing device can include, among other things, a liquid inkjet printer, a solid ink printer, a large-scale commercial printer, a thermal printer, a UV printer, or a dye sublimation printer.

[0053] The high-speed interface 411 manages bandwidth-intensive operations for the computing device 400, and the low-speed interface 413 manages lower bandwidth-intensive operations. Such a function assignment is merely an example. In some implementations, the high-speed interface 411 can be coupled to the memory 405, the user interface (HMI) 444, the keyboard 451 and the display 449 (e.g., via a graphics processor or accelerator), and a high-speed expansion port 415 that can receive various expansion cards via the bus 409.

[0054] In one example, the low-speed interface 413 is coupled to the storage device 407 and the low-speed expansion port 417 via the bus 409. The low-speed expansion port 417, which may include various communication ports (e.g., USB, Bluetooth®, Ethernet®, wireless Ethernet®), may be coupled to one or more input / output devices 441. The computing device 400 may be connected to the server 453 and the rack server 455. The computing device 400 may be implemented in several different forms. For example, the computing device 400 may be implemented as part of the rack server 455.

[0055] This description provides only exemplary embodiments and is not intended to limit the scope, applicability, or configuration of the present disclosure. Rather, the following description of the exemplary embodiments provides an enabling description for those skilled in the art to implement one or more exemplary embodiments. Various changes can be made in the functions and configurations of the elements without departing from the spirit and scope of the disclosed subject matter as set forth in the claims.

[0056] In the following description, specific details are given for a thorough understanding of the embodiments. However, it will be understood by those skilled in the art that the embodiments may be practiced without these specific details. For example, systems, processes, and other elements in the disclosed subject matter may be shown as components in block diagram form so as not to obscure the embodiments with unnecessary detail. In other instances, well-known processes, structures, and techniques may be shown without unnecessary detail to avoid obscuring the embodiments. Further, like reference numbers and designations in the various drawings indicate like elements.

[0057] Also, individual embodiments may be described as a process shown as a flowchart, a flow diagram, a data flow diagram, a structure diagram, or a block diagram. A flowchart may describe the operations as a sequential process, but many of the operations can be performed in parallel or simultaneously. In addition, the order of the operations may be rearranged. The process may end when its operations are completed, but may have additional steps not discussed or included in the figure. Further, not all operations in any particular process described will occur in all embodiments. The process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When the process corresponds to a function, the end of the function can correspond to the return of the function to the calling function or the main function.

[0058] Furthermore, embodiments of the disclosed subject matter may be realized, at least in part, either manually or automatically. Examples of manual or automatic realization may be executed or at least assisted by a machine, hardware, software, firmware, middleware, microcode, a hardware description language, or any combination thereof. When realized in software, firmware, middleware, or microcode, the program code or code segments for performing the necessary tasks may be stored on a machine-readable medium. The necessary tasks may be executed by a processor.

[0059] The various methods or processes outlined in this specification may be encoded as software executable on one or more processors using any one of a variety of operating systems or platforms. Additionally, such software may be written using any of several suitable programming languages and / or programming or scripting tools, and may also be compiled as executable machine language code or intermediate code to be executed on a framework or virtual machine. Typically, the functionality of program modules may be combined or distributed as desired in various embodiments.

[0060] Embodiments of the present disclosure may be embodied as a method for which an example is provided. The acts performed as part of the method may be ordered in any suitable manner. Accordingly, embodiments may be constructed in which acts are performed in a different order than illustrated, including performing some acts shown as consecutive acts in an exemplary embodiment simultaneously.

[0061] Furthermore, the embodiments of the present disclosure and the functional operations described herein can be implemented in digital electronic circuitry, in tangibly embodied computer software or firmware, in computer hardware including the structures disclosed herein and their structural equivalents, or in one or more combinations thereof. Additionally, some embodiments of the present disclosure can be realized as one or more computer programs, i.e., as one or more modules of computer program instructions encoded on a tangible non-transitory program carrier for execution by, or to control the operation of, a data processing apparatus. Further, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal generated to encode information for transmission to a suitable receiver apparatus for execution by a data processing apparatus. The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them.

[0062] According to embodiments of the present disclosure, the term "data processing apparatus" can include, by way of example, all kinds of apparatus, devices, and machines for processing data, including programmable processors, computers, or multiple processors or computers. The apparatus can include dedicated logic circuitry, such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit). The apparatus can also include, in addition to hardware, code for generating an execution environment for the computer program, e.g., processor firmware, protocol stack, database management system, operating system, or code constituting one or more combinations thereof.

[0063] Computer programs (which may also be referred to as or described as programs, software, software applications, modules, software modules, scripts, or code) can be written in a compiled or interpreted language, or any form of programming language including declarative or procedural languages, and can be deployed in any form as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program may or may not correspond to a file in a file system. The program can be stored as part of a file that holds other programs or data, such as one or more scripts stored in a markup language document, a single file dedicated to the program in question, or multiple cooperating files, such as files that hold one or more modules, subprograms, or portions of code.

[0064] A computer program can be deployed to be executed on one computer or located on one site, or distributed across multiple sites and executed on multiple computers interconnected by a communication network. Computers suitable for the execution of a computer program include, by way of example, general-purpose or special-purpose microprocessors or both, or any other kind of central processing unit, and may be based thereon. Generally, the central processing unit receives instructions and data from read-only memory or random access memory or both. Essential elements of a computer are a central processing unit for executing instructions and one or more memory devices for storing instructions and data.

[0065] Generally, a computer will also be operatively coupled to include, receive data from, transfer data to, or both, one or more mass storage devices for storing data, such as magnetic disks, magneto - optical disks, or optical disks. However, a computer need not have such devices. Further, a computer can be incorporated into other devices, such as a cellular phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable memory device, such as a universal serial bus (USB) flash drive, to name a few.

[0066] To provide for interaction with a user, embodiments of the subject matter described herein can be implemented on a computer having a display device, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user, and a keyboard and a pointing device, such as a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; input received from the user can be received in any form including acoustic input, speech (voice) input, or tactile input. Further, a computer can interact with a user by sending documents to and receiving documents from devices used by the user, such as by sending a web page to a web browser on a user's client device in response to a request received from the web browser.

[0067] Embodiments of the subject matter described in this specification may be implemented in a computing system that includes, for example, backend components as a data server, or includes middleware components such as, for example, an application server, or a frontend component, such as a client computer having a graphical user interface or a web browser through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such backend components, middleware components, or frontend components. The components of the system can be interconnected by digital data communication in any form or medium, such as by a communication network. Examples of communication networks include local area networks (“LANs”) and wide area networks (“WANs”), such as the Internet.

[0068] A computing system can include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The relationship between a client and a server arises by computer programs that run on respective computers and have a client-server relationship to each other.

[0069] Although the present disclosure has been described with reference to particular preferred embodiments, it should be understood that various other adaptations and modifications can be made within the spirit and scope of the present disclosure. Accordingly, it is an aspect of the claims to embrace all such variations and modifications that fall within the true spirit and scope of the present disclosure.

Claims

Claim 1 A computer-implemented method for training a neural network suitable for identifying the position of a device in an environment based on signals received by a device located in the environment, the method using a processor together with stored instructions for implementing the method, the instructions, when executed by the processor, performing the steps of the method, the steps comprising: collecting labeled data including measured values of the signals, the labeled data being labeled with coordinates of the position of the measured values, the step further comprising: collecting unlabeled data including measured values of the signals at unidentifiable positions in the environment, the measured values of the labeled data and the measured values of the unlabeled data being sampled at the same or different discrete positions in the environment, the step further comprising: during a first training stage using the labeled data, training a double-regressor neural network to identify positions in the continuous space of the environment from the labeled data, the double-regressor neural network including a feature extractor configured to extract features from the labeled data and a double-regressor including two regressors having the same architecture, each regressor being trained to determine a position in the continuous space of the environment from the features received from the feature extractor, the step further comprising: during a second training stage, using the labeled data and the unlabeled data to train the parameters of the double-regressor using the fixed parameters of the feature extractor trained during the first training stage such that each of the two regressors identifies the same labeled position while processing the labeled data and different positions while processing the unlabeled data; During each of the two regressors processing domain invariant features to identify the same position, according to the adversarial discriminator, from the unlabeled data, using the statistical characteristics of the labeled data to extract domain invariant features, during a third training stage, using the labeled data and the unlabeled data, with the fixed parameters of the double regressor trained during the second training stage, training the parameters of the feature extractor using the adversarial discriminator, a method implemented by a computer, including the step of.

2. During the first training stage, the double regressor neural network is trained using a loss function such that each of the two regressors identifies the same position while processing the labeled data, and the loss function is based on the mean squared error (MSE) criterion, the method implemented by a computer according to claim 1.

3. During the second training stage with the fixed parameters of the feature extractor trained during the first training stage, each of the two regressors, while processing the unlabeled data, to identify different positions, the parameters of the double regressor are trained based on a soft similarity function, the method implemented by a computer according to claim 1.

4. During the second training stage with the fixed parameters of the feature extractor trained during the first training stage, each of the two regressors, while processing the labeled data, identifies the same labeled position, and while processing the unlabeled data, to identify different positions, the parameters of the double regressor are trained based on an objective function, and the objective function is based on a loss function and the soft similarity function, the method implemented by a computer according to claim 3.

5. During the third training stage with the fixed parameters of the double regressor trained during the second training stage, the parameters of the feature extractor and the adversarial discriminator are trained based on an adversarial loss function and the soft similarity function, and the adversarial loss function is configured to mitigate domain distribution shift, the method implemented by a computer according to claim 3.

6. During the third training stage with the fixed parameters of the double regressor trained during the second training stage, the parameters of the feature extractor and the adversarial discriminator are trained based on the maximization of the adversarial loss function, the method realized by the computer according to claim 1.

7. The method further includes determining extended labeled data and extended unlabeled data that form a plurality of intermediate domains, the method realized by the computer according to claim 1.

8. The method further includes reducing the domain shift between the extended labeled data and the extended unlabeled data based on the adversarial relationship between the adversarial discriminator and the feature extractor to align the statistical characteristics of the labeled data and the unlabeled data, the method realized by the computer according to claim 7.

9. The extended labeled data is determined by linearly combining the labeled data and the unlabeled data at a fixed ratio such that the extended labeled data is similar to the labeled data, and the extended unlabeled data is determined by linearly combining the unlabeled data and the labeled data at the fixed ratio such that the extended unlabeled data is similar to the unlabeled data, the method realized by the computer according to claim 7.

10. The method further includes inputting the unlabeled data collected by the device into a trained double regressor neural network, the trained double regressor neural network including the feature extractor trained during the third training stage and at least one of the two regressors trained during the second training stage, and the method further includes processing the unlabeled data using the trained double regressor neural network to output the position of the device in the environment, the method realized by the computer according to claim 1.

11. The method further includes including the step of inputting the unlabeled data collected by the device into a trained double-regressor neural network, the trained double-regressor neural network including the feature extractor trained during the third training stage and at least one of the two regressors trained during the first training stage using the labeled data, the method further comprising the method realized by a computer according to claim 1, including the step of processing the unlabeled data using the trained double-regressor neural network to output the position of the device in the environment.

12. The method further comprises determining an operation path for the device to execute a task in the environment based on the position of the device in the environment, and controlling the device based on the operation path, the method realized by a computer according to claim 11.

13. The method realized by a computer according to claim 11, further comprising the step of tracking the device in the environment based on the output position of the device.

14. The method realized by a computer according to claim 11, wherein the device is a robot and the environment is any one of an autonomous factory, a warehouse, or a smart home.

15. A system for training a neural network suitable for identifying the position of a device in an environment based on signals received by the device located in the environment, comprising a processor and a memory storing instructions, which when executed by the processor cause the system to collect labeled data including measurement values of the signals, the labeled data being labeled with coordinates of the position of the measurement values, and the instructions further cause the system to, when executed by the processor, collect unlabeled data including measurement values of the signals at unidentified positions in the environment, the measurement values of the labeled data and the measurement values of the unlabeled data being sampled at the same or different discrete positions in the environment, and the instructions further cause the system to, when executed by the processor, During a first training stage of using the labeled data, train a double-regressor neural network to identify a position in the continuous space of the environment from the labeled data, the double-regressor neural network including a feature extractor configured to extract features from the labeled data and a double-regressor including two regressors having the same architecture, each regressor being trained to determine a position in the continuous space of the environment from the features received from the feature extractor, and the instructions, when further executed by the processor, cause the system to During a second training stage, use the labeled data and the unlabeled data to train the parameters of the double-regressor using the fixed parameters of the feature extractor trained during the first training stage such that each of the two regressors identifies the same labeled position while processing the labeled data and different positions while processing the unlabeled data. During a third training stage, use the labeled data and the unlabeled data to train the parameters of the feature extractor using the adversarial discriminator such that each of the two regressors identifies the same position while processing domain-invariant features, by extracting domain-invariant features from the unlabeled data using the statistical characteristics of the labeled data according to the adversarial discriminator, using the fixed parameters of the double-regressor trained during the second training stage. A system. **Claim 16** The system according to claim 15, wherein during the first training stage, the double-regressor neural network is trained using a loss function such that each of the two regressors identifies the same position while processing the labeled data, and the loss function is based on a mean squared error (MSE) criterion. **Claim 17** The system according to claim 15, wherein during the second training stage with the fixed parameters of the feature extractor trained during the first training stage, the parameters of the double-regressor are trained based on a soft similarity function such that each of the two regressors identifies different positions while processing the unlabeled data.

18. During the second training stage with the fixed parameters of the feature extractor trained during the first training stage, the parameters of the double regressor are trained based on an objective function such that each of the two regressors identifies the same labeled position while processing the labeled data and different positions while processing the unlabeled data, and the objective function is based on a loss function and the soft similarity function. The system according to claim 17.

19. The processor further is configured to input unlabeled data collected by the device into a trained double regressor neural network, the trained double regressor neural network including the feature extractor trained during the third training stage and at least one of the two regressors trained during the second training stage, and the processor further is configured to process the unlabeled data using the trained double regressor neural network to output the position of the device in the environment. The system according to claim 15.

20. A non-transitory computer-readable storage medium embodying a processor-executable program for performing a method for training a neural network suitable for identifying the position of a device in an environment based on signals received by the device in the environment, the method including collecting labeled data including measurements of the signals, the labeled data being labeled with the coordinates of the position of the measurements, and the method further including collecting unlabeled data including measurements of the signals at unidentified positions in the environment, the measurements of the labeled data and the measurements of the unlabeled data being sampled at the same or different discrete positions in the environment, and the method further including During a first training phase of using the labeled data, the method includes training a double-regressor neural network to identify a position in a continuous space of the environment from the labeled data, the double-regressor neural network including a feature extractor configured to extract features from the labeled data and a double-regressor including two regressors having the same architecture, each regressor being trained to determine a position in the continuous space of the environment from the features received from the feature extractor, the method further comprising during a second training phase, using the labeled data and the unlabeled data to train the parameters of the double-regressor using the fixed parameters of the feature extractor trained during the first training phase such that each of the two regressors identifies the same labeled position while processing the labeled data and different positions while processing the unlabeled data; during a third training phase, using the labeled data and the unlabeled data to train the parameters of the feature extractor using the adversarial discriminator such that each of the two regressors identifies the same position while processing domain-invariant features, the method including extracting domain-invariant features from the unlabeled data using the statistical characteristics of the labeled data according to the adversarial discriminator, using the fixed parameters of the double-regressor trained during the second training phase. A non-transitory computer-readable storage medium.