Circuitry, method and system
The system addresses finetuning challenges by generating synthetic training data with reduced noise and privacy-sensitive information, enhancing robot performance and security in specific environments.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SONY SEMICON SOLUTIONS CORP
- Filing Date
- 2025-12-01
- Publication Date
- 2026-06-11
AI Technical Summary
Existing robot control systems face challenges in finetuning for specific environments due to high information noise in environmental data, resource limitations, and privacy concerns when uploading data for training, leading to suboptimal performance and potential privacy breaches.
The system generates synthetic training data with a predetermined similarity to real environmental data, replacing sensitive information with virtual objects to enhance finetuning while protecting privacy and optimizing performance.
This approach enhances finetuning efficiency, reduces computational effort, and safeguards privacy by generating training data with reduced noise and increased environmental representation, resulting in a more robust and secure robot control system.
Smart Images

Figure EP2025084825_11062026_PF_FP_ABST
Abstract
Description
[0001] Sony Semiconductor Solutions Corporation et al.
[0002] CIRCUITRY, METHOD AND SYSTEM
[0003] TECHNICAL FIELD
[0004] The present disclosure generally pertains to circuitry, a method and a system.
[0005] TECHNICAL BACKGROUND
[0006] Known robots may interact with their environment. For example, a robot may acquire environmental data that indicate an environment of the robot, and may perform an action based on the environmental data.
[0007] Although there exist techniques for controlling robots, it is generally desirable to provide improved circuitry, an improved method and an improved system.
[0008] SUMMARY
[0009] According to a first aspect, the disclosure provides circuitry that is configured to obtain environmental data that indicate an environment of a robot; generate training data based on the environmental data, wherein the training data include a synthetic portion that has a predetermined similarity to a corresponding source portion of the environmental data, wherein the synthetic portion represents a virtual object in the environment, and wherein the source portion represents a real object in the environment; and provide the training data to training circuitry for finetuning a control network of the robot based on the training data.
[0010] According to a second aspect, the disclosure provides a method that includes obtaining environmental data that indicate an environment of a robot; generating training data based on the environmental data, wherein the training data include a synthetic portion that has a predetermined similarity to a corresponding source portion of the environmental data, wherein the synthetic portion represents a virtual object in the environment, and wherein the source portion represents a real object in the environment; and providing the training data to training circuitry for finetuning a control network of the robot based on the training data.
[0011] According to a third aspect, the disclosure provides a system that includes the circuitry according to the first aspect; a sensor that is configured to acquire environmental data that indicate the environment of the robot and to provide the environmental data to the circuitry; the training circuitry, wherein the training circuitry is configured to determine updated weights of the control network based on the training data and to provide the updated weights to the control network; and the control network, wherein the control network is configured to control the robot based on Sony Semiconductor Solutions Corporation et al. environmental data that are acquired by the sensor and based on the updated weights that are determined by the training circuitry.
[0012] Further aspects are set forth in the dependent claims, the drawings and the following description.
[0013] BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Embodiments are explained by way of example with respect to the accompanying drawings, in which:
[0015] Fig. 1 illustrates an embodiment of circuitry;
[0016] Fig. 2 illustrates an embodiment of a method;
[0017] Fig. 3 illustrates a first embodiment of a system;
[0018] Fig. 4 illustrates a second embodiment of a system; and
[0019] Fig. 5 illustrates an embodiment of a general-purpose computer.
[0020] DETAILED DESCRIPTION OF EMBODIMENTS
[0021] Before a detailed description of the embodiments under reference of Fig. 1 is given, general explanations are made.
[0022] As mentioned in the outset, robots may interact with their environment. For example, a robot may obtain data from a sensor that indicate an environment of the robot (which may also be referred to as environmental data), and may determine an action based on the obtained environmental data.
[0023] Finetuning of robots for specific area of operation may be necessary to achieve a high performance. For example, a robot may be controlled by a control network. The control network may include an artificial neural network and may be configured to receive the environmental data as input, and to determine an action to be performed by the robot based on the environmental data. Finetuning of the robot may include determining updated weights of the artificial neural network (control network) according to the environment of the robot.
[0024] It has been recognized that, in some embodiments, it is desirable to perform finetuning of a robot based on training data that have been modified with respect to environmental data that represent the environment of the robot as it was acquired by a sensor.
[0025] For example, information noise may be minimized in the training data such that, e.g., the finetuning may converge faster. Sony Semiconductor Solutions Corporation et al.
[0026] For example, a representation of the environment of the robot may be made more complicated in the training data, e.g., to make the control network of the robot more robust.
[0027] For example, the training data may be modified to represent another, similarly structured environment in order to increase an amount of training data for finetuning.
[0028] For example, training (e.g., finetuning) may be performed in a server and not on the edge (e.g., in the robot), and uploading environmental data to a server may involve issues in terms of privacy. Therefore, the training data may be modified to not include personal or confidential information before uploading to the server for finetuning.
[0029] Consequently, some embodiments of the disclosure pertain to circuitry that is configured to: obtain environmental data that indicate an environment of a robot; generate training data based on the environmental data, wherein the training data include a synthetic portion that has a predetermined similarity to a corresponding source portion of the environmental data, wherein the synthetic portion represents a virtual object in the environment, and wherein the source portion represents a real object in the environment; and provide the training data to training circuitry for finetuning a control network of the robot based on the training data.
[0030] The circuitry may include a processing portion, e.g., an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a programmed microprocessor, a central processing unit (CPU), or the like. The circuitry may include a hardware configuration and / or software (and / or firmware) that may cause the circuitry to perform the processing according to the present technology.
[0031] The circuitry may include a storage portion with volatile and / or non-volatile memory (e.g., a hard disk, a solid-state drive (SSD), a flash memory, a dynamic random-access memory (DRAM), a static random-access memory (SRAM), a double data rate synchronous dynamic random-access memory (DDR), or the like) that may store software / firmware to be executed by the circuitry and / or data that are used or generated by the circuitry (e.g., the environmental data, the training data, temporary data etc.).
[0032] The circuitry (e.g., in its processing portion) may also include a portion (e.g., a graphics processing unit (GPU) and / or a tensor processing unit (TPU)) that may be specialized for executing a machine-learning model (e.g., an artificial neural network). Sony Semiconductor Solutions Corporation et al.
[0033] The circuitry may further include a communication portion with one or more interfaces (e.g., a camera serial interface (CSI), an audio interface, an analog input with analog-to-digital converter (ADC), a peripheral component interconnect (PCI) interface, a high-definition multimedia interface (HDMI), a universal serial bus (USB) interface, an ethernet interface, an InfiniBand interface, an optical interface, a Wi-Fi interface (e.g., IEEE 802.11), a mobile telecommunications interface (e.g., New Radio (NR), Long Term Evolution (LTE), High-Speed Packet Access (HSPA), etc.), a Bluetooth interface, a ZigBee interface, or the like) for communicating (e.g., exchanging, sending and / or receiving data) with a sensor that may acquire the environmental data and / or with the training circuitry.
[0034] The circuitry may include a general-purpose computer, which is described below with respect to Fig. 5.
[0035] The circuitry may be included in the robot, for example, in a central control portion of the robot, in a camera of the robot and / or on a separate chip. The circuitry may be provided separately from the robot, e.g., in a remote control, smartphone, laptop, server etc. that may be used for controlling the robot.
[0036] The robot may include a machine that may be configured to autonomously perform an action in its environment, e.g., to move within the environment and / or to interact with the environment. For example, the robot may be configured to help in the household, to assist a user with daily tasks, to support in nursing, to repair / maintain a product, to transport persons or goods, to explore its environment, to rescue people or animals, or the like. The robot may be configured as an android, a robot pet (e.g., Aibo by Sony), a robotic vehicle, an autonomous car, a drone, or the like.
[0037] The environment of the robot may include one or more objects in the surroundings of the robot, such as a wall, a door, stairs, furniture, a household / office article, a sign, a document (e.g., book, letter, etc.), a person, an animal, a vehicle (e.g., car, motorcycle, bicycle, etc.), a building, a tree, a fence, a street, a light source (e.g., sun, window, artificial light source, etc.), etc.
[0038] The environmental data may indicate the environment of the robot, e.g., may include a representation of one or more real objects that may be present in the environment. Based on the environmental data, the robot (e.g., a control network executed by control circuitry of the robot) may determine an environment of the robot and / or may determine an action to be performed by the robot (e.g., for moving in the environment, for interacting with the environment (e.g., with an object in the environment), etc.). Sony Semiconductor Solutions Corporation et al.
[0039] The robot may include a sensor (environmental sensor) for obtaining the environmental data.
[0040] The environmental sensor may also be provided separate from the robot, e.g., in a remote control, smartphone, laptop etc. used for controlling the robot, or as a separate sensor device that may communicate with the circuitry via Wi-Fi, Bluetooth, ZigBee, Ethernet, HDMI, USB, or the like. The environmental sensor may include a camera (e.g., red-green-blue (RGB) camera, grayscale camera, infrared camera, event-based vision sensor, etc.) for acquiring image data (e.g., single image, series of image frames, video, etc.) as the environmental data, a microphone for acquiring sound data as the environmental data, a depth sensor (e.g., time-of-flight (ToF) sensor, lidar sensor, radar sensor, ultrasound sensor, etc.) for acquiring depth data (e.g., a three- dimensional (3D) map) as the environmental data, or the like.
[0041] It may be desirable to perform finetuning of the robot for improving a performance of the robot in the environment, for example, for increasing an ability (e.g., robustness, probability, etc.) of the robot to identify its environment, for increasing a relevance of actions performed by the robot in the environment, for increasing a robustness of the robot in the environment, for increasing a safety of the robot in the environment, for reducing a computational effort for determining an action of the robot, etc. The finetuning of the robot may be performed based on data (training data) that may indicate the environment of the robot.
[0042] The finetuning may include training an artificial neural network (control network) that may be configured to determine an operation / action of the robot (e.g., movement, interaction with the environment, etc.). In the training, the training circuitry may update one or more parameters (network weights) of the control network to improve an output of the control network based on the training data.
[0043] However, as mentioned, it may be desirable that the training data used for finetuning include a synthetic portion instead of unmodified environmental data that represent the environment of the robot as it was acquired by the environmental sensor.
[0044] For example, the environmental data as acquired by the environmental sensor may include high information noise. Therefore, the circuitry may replace a source portion of the environmental data that may include high information noise with the synthetic portion, which may include less information noise. Thus, for example, the finetuning may converge faster, a performance of the finetuned robot may be optimized, or the like.
[0045] For example, the environment may have a simple structure (e.g., may include few features). Therefore, the circuitry may generate a synthetic portion that may include more features, such Sony Semiconductor Solutions Corporation et al. that the representation of the environment in the training data may be made more complicated, e.g., to make the robot more robust in the environment and / or to avoid overfitting. Likewise, the circuitry may generate, based on the environmental data, training data with several representations of an environment that may be structured similar to the representation of the environment in the acquired environmental data, but may include different features. Thus, an amount of training data for finetuning (e.g., a number of different representations of an environment of the robot) may be increased, e.g., to make the robot more robust in the environment and / or to avoid overfitting.
[0046] The finetuning may require dedicated hardware and / or may consume a high amount of energy. The finetuning may also be based on training data from a plurality of robots (federated learning). Therefore, in some embodiments, the finetuning is performed in a server and not on the edge (e.g., in the robot). In such a case, the training circuitry may be provided in the server and may include dedicated hardware (e.g., a GPU and / or a TPU) for finetuning the robot.
[0047] In some embodiments, the robot is instead finetuned on the edge, e.g., in its onboard processor. In such a case, the robot may include the training circuitry (e.g., its onboard processor, CPU, GPU, TPU or the like). However, due to resource limitations in a robot, a neural network (e.g., control network) used by the robot may not be as big as in a server if the neural network can be finetuned on the edge, hence a performance of a neural network (control network) that is finetuned on the edge may be lower as compared to a control network that is finetuned on a dedicated server.
[0048] The task of finetuning robots in the real world is, in some embodiments, difficult enough in itself but may in addition present issues based on privacy. The robot may need to collect data (e.g., the environmental data) for finetuning, but the collected data may contain personal information of people near a robot operation area (e.g., in the environment of the robot). If the robot transmits (e.g., uploads) the collected data to the server for finetuning, a secure connection may be required and consent of persons represented in the considered training data (e.g., in images represented by the training data) may be needed, even if the server deletes the data after finetuning. Further, it may be possible to (at least partially) reconstruct, based on the finetuned control network, the training data that have been used for finetuning. Therefore, the circuitry may generate the synthetic portion such that it includes less or no personal information than the source portion of the environmental data acquired by the environmental sensor. The circuitry may then provide the training data with the synthetic portion to the training circuitry. Sony Semiconductor Solutions Corporation et al.
[0049] In summary, by providing the synthetic portion in the training data, the training (finetuning) may be enhanced and / or privacy may be protected.
[0050] The circuitry may input the environmental data (or at least the source portion of the environmental data) to a synthesizing model. The synthesizing model may generate and output the synthetic portion based on the inputted environmental data (or source portion). The circuitry may execute the synthesizing model.
[0051] The synthesizing model may include a machine learning model, e.g., an artificial neural network, which may be configured as a semantic segmentation network, as an (image) variation network, as an inpainting model, as a foundation model, or as any other model that may be suitable for generating the training data (e.g., image data, sound data, depth data, or the like).
[0052] The synthesizing model may generate the synthetic portion such that the synthetic portion has a predetermined similarity to the source portion. The similarity between the synthetic portion and the source portion may be based on a defined metric. The predetermined similarity may correspond to a predetermined threshold in the defined metric.
[0053] The synthesizing model may generate, in the synthetic portion, a representation of the virtual object. The virtual object may be based on the real object that may be present in the environment and represented by the environmental data. The similarity between the synthetic portion and the source portion may correspond to a similarity between the representation of the virtual object in the synthetic portion and the representation of the real object in the source portion. For example, the synthesizing model may generate the representation of the virtual object based on interpolating between features of training data of the synthesizing model. The interpolating may be controlled based on random values (e.g., true random numbers or pseudo-random numbers). The synthesizing model may interpolate between the features such that the similarity between the representations of the virtual object and the real object satisfies the predetermined threshold in the defined metric.
[0054] For example, the synthesizing model may generate the representation of the virtual object (e.g., obstacle, person, animal, item, etc.) such that a position and / or size of the virtual object in the environment may correspond to a position and / or size, respectively, of the real object as indicated in the environmental data, wherein an appearance of the virtual object may differ from an appearance of the real object.
[0055] The representation of the virtual object may replace the representation of the real object in the training data. For example, the circuitry (e.g., the synthesizing model and / or a postprocessing Sony Semiconductor Solutions Corporation et al. function for postprocessing an output of the synthesizing model) may remove the representation of the real object and insert the representation of the virtual object in the synthetic portion, may cover the representation of the real object with the representation of the virtual object (e.g., in the case of image data) in the synthetic portion, and / or may generate the training data (or the synthetic portion) with the representation of the virtual object but without inserting the representation of the real object into the training data.
[0056] The synthetic portion may correspond to the entire environmental data or to a sub-portion (e.g., the source portion) of the environmental data. In the latter case, a portion of the training data that is not included in the synthetic portion may correspond to a corresponding portion of the environmental data.
[0057] The circuitry may then transmit the generated training data to the training circuitry. The training circuitry, as described, may include a GPU, a TPU or any other configuration suitable for performing the finetuning of the control network. The training circuitry may be provided in a server, to which the circuitry may provide the training data via a communication network (e.g., the internet, an intranet, or the like), or the training circuitry may be provided in the robot and the circuitry may provide the training data to the training circuitry, e.g., via a data bus, via PCI, or the like.
[0058] The training circuitry may update one or more parameters (e.g., network weights) of the control network based on the training data, e.g., using stochastic gradient descent (SGD), mini-batch gradient descent, adaptive gradient algorithm (Adagrad), momentum, adaptive moment estimation (Adam), root mean square propagation (RMSprop), or the like. The training circuitry may then provide the updated param eter(s) to the robot, and the robot (e.g., control circuitry of the robot) may execute the control network with the updated parameter(s).
[0059] In some embodiments, the synthetic portion corresponds to a variation of the source portion.
[0060] The variation may include a modification of the source portion according to the predetermined similarity. For example, the variation may include providing the representation of the virtual object in the synthetic portion instead of the representation of the real object.
[0061] The variation may include generating the representation of the virtual object with a difference in one or more features with respect to the representation of the real object.
[0062] In some embodiments, the variation of the source portion is based on a random value in a latent space. Sony Semiconductor Solutions Corporation et al.
[0063] The latent space may be a multidimensional space in which objects and / or other concepts may be represented as vectors (embeddings), wherein a similarity between objects / concepts may correspond to a distance between their corresponding embeddings. A metric for the distance between embeddings (vectors in the latent space) may include or be based on the cosine distance, the Euclidean distance, the Manhattan distance, the Hamming distance, or the like.
[0064] The similarity between the synthetic portion and the source portion may correspond to the distance between embeddings that represent the synthetic portion and the source portion, or that represent the virtual object and the real object, respectively, in the latent space.
[0065] The random value may be obtained from a true random number generator (e.g., based on thermal noise, optical noise, electronic noise, radioactive decay, quantum phenomena, or the like) or from a pseudo-random number generator (e.g., Mersenne-Twister, linear congruential generator (LCG), inversive congruential generator, multiply-with-carry (MWC), xorshift, cryptographical functions, or the like).
[0066] For generating the synthetic portion, the circuitry (e.g., the synthesizing model) may determine an embedding of the source portion (e.g., of the real object). The circuitry may generate an embedding of the synthetic portion (e.g., of the virtual object) that may have a distance to the embedding of the source portion according to the predetermined similarity, for example, by transforming the embedding of the source portion according to the random value or by selecting a random embedding according to the random value. The circuitry may then determine the synthetic portion (e.g., the representation of the virtual object in the synthetic portion) according to the generated embedding of the virtual object.
[0067] In some embodiments, the generating of the training data includes replacing the source portion with the synthetic portion.
[0068] The circuitry (e.g., the synthesizing model) may remove the source portion (e.g., the representation of the real object) from the training data and insert the synthetic portion (e.g., the representation of the virtual object) in the training data, or the circuitry may cover the source portion (e.g., the representation of the real object) with the synthetic portion (e.g., the representation of the virtual object).
[0069] For example, if the environmental data include image data, the circuitry may generate the training data based on one or more of the following options for image variation.
[0070] The circuitry may input an original image indicated by the source portion into a semantic segmentation network to obtain a semantic map of the original image, and may then input the Sony Semiconductor Solutions Corporation et al. semantic map into a semantic map-to-image generative model that may generate the synthetic portion.
[0071] The circuitry may input the original image into an image variation network to obtain an image with a similar content but changed objects, e.g., changed faces, as the synthetic portion.
[0072] The circuitry may obtain the original image, identify, using object detection in the original image, features like faces, number plates, etc. that may cause privacy issues, and replace such features with representations of synthetic objects (e.g., face by synthetic face, etc.) using inpainting (e.g. based on stable diffusion, DALL-E, Midjourney, or the like).
[0073] For example, if the environmental data include image data that represent a real person as the real object, the synthesizing model may generate, as the representation of the virtual object, a representation of a virtual person at a similar position in the environment and with a similar size and pose as the real person, but the synthesizing model may generate an artificial face for the virtual person such that the real person may not be identified based on the representation of the virtual person, and / or the synthesizing model may increase or decrease a degree of information noise of the virtual person, e.g., by inserting or removing features of the virtual person such as glasses, headdress jewelry, wrinkles, birthmarks, pimples, or the like.
[0074] For example, if the environmental data include depth data, the circuitry may generate the training data based on similar processing as described with respect to image data.
[0075] For example, the environmental data may include audio data that may represent, as the real object, a real voice of a person that may be speaking in the environment of the robot.
[0076] In such a case, the synthesizing model may generate, as the representation of the virtual object, an artificial virtual voice that may differ from the real voice according to the predetermined similarity. The source portion may correspond to a frequency spectrum of the real voice, and the synthetic portion may correspond to a frequency spectrum of the virtual voice. The synthesizing model may generate the synthetic portion such that the frequency spectrum of the virtual voice may differ from the frequency spectrum of the real voice according to the predetermined similarity. Thus, a person who spoke with the real voice in the environment of the robot may not be identifiable based on the training data, and / or the synthesizing model may generate the virtual voice with a higher or lower information noise (e.g., with a more or less complicated frequency spectrum) than the real voice.
[0077] Additionally or alternatively, the synthesizing model may generate, as the representation of the virtual object, an artificial virtual utterance. The source portion may correspond to a time interval Sony Semiconductor Solutions Corporation et al. in which the real voice is speaking, and the synthetic portion may correspond to a time interval in which the virtual utterance is generated. The virtual utterance may include different words or expressions than words / expressions in a real utterance spoken by the real voice. For example, the virtual utterance may include words or expressions with a different information noise than the real utterance, or the virtual utterance may fictive data instead of personal data (e.g., name, address, phone number, bank account, insurance number, etc.) spoken by the real voice. The synthesizing model may generate the virtual utterance according to a frequency spectrum of the real voice or of an artificial voice.
[0078] In some embodiments, the predetermined similarity corresponds to a degree of deviation between the source portion and the synthetic portion.
[0079] As mentioned, the degree of deviation may be based on a defined metric, e.g., a distance between embeddings in a latent space.
[0080] In some embodiments, the predetermined similarity corresponds to a difference in a degree of information noise between the synthetic portion and the source portion.
[0081] The information noise may correspond to a number of features indicated in the environmental data, to a degree of deviation of the features indicated in the environmental data from an average, or the like. For example, the information noise may correspond to a number of embeddings that the circuitry may generate for representing the environmental data (or at least the source portion) in the latent space, and / or the information noise may correspond to a variance, standard deviation or the like of the embeddings that the circuitry may generate for representing the environmental data (or at least the source portion) in the latent space.
[0082] In some embodiments, the generating of the training data includes causing an artificial neural network to generate the synthetic portion based on the source portion.
[0083] As mentioned, the circuitry may execute the synthesizing model. The synthesizing model may include the artificial neural network. For generating the training data, the circuitry may provide the environmental data (or at least the source portion) to the artificial neural network of the synthesizing model and may cause the artificial neural network of the synthesizing model to generate the synthetic portion based on the source portion (e.g., to generate a representation of the virtual object with the predetermined similarity to the representation of the real object), as described herein.
[0084] In some embodiments, the environmental data include image data. Sony Semiconductor Solutions Corporation et al.
[0085] The robot may include a camera and may acquire the image data with the camera. The image data may represent a single image, a sequence (e.g., stream) of image frames, and / or a video.
[0086] For example, the circuitry may be included in an image variation camera module that may be included in the camera of the robot. The camera may be operable in a first mode, in which the camera may output acquired image data, and in a second mode, in which the camera may provide acquired image data to the image variation camera module and may output training data generated by the circuitry in the image variation camera module. Thus, the robot may switch between the first mode for controlling the robot and the second mode for generating training data for finetuning the robot.
[0087] Alternatively, the circuitry may be provided separately from the camera, e.g., in control circuitry of the robot.
[0088] When finetuning of the robot is performed based on image data acquired by the camera (e.g., for every image of the image data that may be uploaded to a server for finetuning of the robot and / or inputted to the training circuitry), the circuitry may generate training data (e.g., a synthetic image that may include the synthetic portion) may be generated based on the original image data as acquired by the camera to replace the original image data, such that privacy of people around the robot may be protected and / or the finetuning may be enhanced. The training data with the synthetic portion may then be uploaded and the original image data may be deleted. The training data with the synthetic portion may be stored and used for finetuning without any concern about privacy.
[0089] Further examples of image generation techniques that may be used in the synthesizing model may include a variational autoencoder, an adversarial model, and latent diffusion.
[0090] The synthetic portion may represent a same kind of information as the corresponding source portion (e.g., an embedding of the synthetic portion may have a small distance to an embedding of the source portion). For example, proportions in the synthetic portion (e.g., proportions of the virtual object) may be equal to proportions in the source portion (e.g., proportions of the real object), or the proportions may differ within the predetermined similarity. The synthesizing model may be trained to keep and / or vary proportions according to the predetermined similarity.
[0091] In some embodiments, the environmental data include sound data.
[0092] The robot may include a microphone for acquiring sound data from its environment. The sound data may represent voices from people in the environment, acoustic signals in the environment (e.g., a horn, an alarm, a doorbell, a notification sound, etc.), environmental sounds (e.g., engine Sony Semiconductor Solutions Corporation et al. and / or tire sound of a car, steps of people, etc.). The robot may detect a category of a sound represented in the sound data, e.g., whether the sound corresponds to a voice of a speaking person, to an acoustic signal (e.g., horn, alarm, doorbell, notification sound, etc.), to an environmental sound (e.g., car, steps, etc.).
[0093] The robot may receive a speech command via a voice represented by the sound data. The robot may determine an action based on an acoustic signal represented by the sound data. The robot may identify its environment based on the sound data. The robot may determine an action based on the sound data, e.g., based on a recognized sound in the sound data.
[0094] If the robot acquires the sound data with two or more microphones, the robot may determine a direction from which a sound is coming, and may locate a corresponding sound source (e.g., the robot may locate a person that is uttering a speech command, and the robot may turn to the person, or the robot may detect an approaching car based on the sound data and may wait before crossing a road until the car has passed).
[0095] The generating of the synthetic portion may include generating, as the virtual object, an artificial sound. The generating of the training data may include cancelling, from the acquired sound data, a frequency spectrum of a sound that corresponds to the real object and that should be excluded from the training data. The circuitry may generate the training data in a frequency domain and / or in a time domain. The circuitry may execute the synthesizing model, which may include an artificial neural network, for generating the training data, as described herein.
[0096] For processing audio data, the circuitry may execute, as the synthesizing model, a foundation model that may be configured to process audio data appropriately, or the circuitry may execute, as the synthesizing model, an artificial neural network that may be specialized for processing audio data.
[0097] In some embodiments, the circuitry is further configured to delete the environmental data after generating the training data.
[0098] Thus, a smaller memory may be sufficient for the robot, and leakage of the original environmental data may be avoided, e.g., in a case when the environmental data include personal or other confidential information.
[0099] Some embodiments pertain to a method that includes: obtaining environmental data that indicate an environment of a robot; generating training data based on the environmental data, wherein the training data include a synthetic portion that has a predetermined Sony Semiconductor Solutions Corporation et al. similarity to a corresponding source portion of the environmental data, wherein the synthetic portion represents a virtual object in the environment, and wherein the source portion represents a real object in the environment; and providing the training data to training circuitry for finetuning a control network of the robot based on the training data.
[0100] The method may be performed by the circuitry described herein. Features described herein with respect to the circuitry may accordingly correspond to features of the method.
[0101] Some embodiments pertain to a system that includes: the circuitry according to any embodiment described above; a sensor that is configured to acquire environmental data that indicate the environment of the robot and to provide the environmental data to the circuitry; the training circuitry, wherein the training circuitry is configured to determine updated weights of the control network based on the training data and to provide the updated weights to the control network; and the control network, wherein the control network is configured to control the robot based on environmental data acquired by the sensor and based on the updated weights determined by the training circuitry.
[0102] The system may be included in the robot or may include the robot, a server, a control device (e.g., remote control, smartphone, laptop, server etc.) and / or a sensor device.
[0103] The sensor may be configured as an environmental sensor as described herein.
[0104] The control network may include an artificial neural network that may include the weights. The weights may be parameters of the artificial neural network and may determine how the artificial neural network processes its input.
[0105] The training circuitry may perform finetuning of the robot by updating weights for the control network, e.g., by determining values for the weights such that an output of the control network based on the training data is optimized.
[0106] In some embodiments, the training circuitry is configured to: receive, via a communication network, the training data from the circuitry and further training data that indicate an environment of a further robot; determine the updated weights based on the training data and on the further training data; and provide the updated weights to the control network via the communication network. Sony Semiconductor Solutions Corporation et al.
[0107] The communication network may include the internet, an intranet, a local area network (LAN), a wireless LAN (WLAN), a controller area network (CAN), or the like.
[0108] The communication network may connect the circuitry, the further robot, and the training circuitry such that communication (e.g., data exchange) between the circuitry, the further robot and the training circuitry may be possible.
[0109] The training circuitry may receive training data that may indicate respective environments of a plurality of robots, and may perform the finetuning based on the training data from the plurality of robots (federated learning). The training circuitry may then provide the updated weights to the plurality of robots (e.g., to their respective control circuitries / control networks). The robot and / or the control circuitry of the robot may receive or download the updated weights from the training circuitry, may update the control network to use the updated weights, and may determine an action of the robot by executing the control network with the updated weights.
[0110] The methods as described herein are also implemented in some embodiments as a computer program causing a computer and / or a processor to perform the method, when being carried out on the computer and / or processor. In some embodiments, also a non-transitory computer- readable recording medium is provided that stores therein a computer program product, which, when executed by a processor, such as the processor described above, causes the methods described herein to be performed.
[0111] Returning to Fig. 1, Fig. 1 illustrates an embodiment of circuitry 1. The circuitry 1 includes a processing portion 2, a storage portion 3 and a communication portion 4.
[0112] The processing portion 2 is configured to perform data processing. The storage portion 3 is configured to store data which the processing portion 2 uses for data processing or generates by data processing. The communication portion 4 is configured to exchange data with a device external to the circuitry 1.
[0113] In some embodiments, the circuitry 1 is provided in a robot. In some embodiments, the circuitry 1 is provided in a device configured to control a robot, e.g., in a remote control, smartphone, laptop, server, etc.
[0114] The circuitry 1 is configured to perform the method 10 of Fig. 2.
[0115] Fig. 2 illustrates an embodiment of a method 10. The method 10 is an example of a method performed by the circuitry 1 of Fig. 1. Sony Semiconductor Solutions Corporation et al.
[0116] At 11, the circuitry 1 obtains environmental data that indicate an environment of a robot. The environmental data are acquired by an environmental sensor, and the communication portion 4 receives the environmental data from the environmental sensor.
[0117] At 12, the processing portion 2 generates training data based on the environmental data. The processing portion 2 generates the training data such that the training data include a synthetic portion that has a predetermined similarity to a corresponding source portion of the environmental data.
[0118] The source portion represents a real object in the environment. The synthetic portion represents a virtual object in the environment and corresponds to a variation of the source portion. In some embodiments, the variation of the source portion is based on a random value in a latent space.
[0119] The predetermined similarity corresponds to a degree of deviation between the source portion and the synthetic portion. In some embodiments, the predetermined similarity corresponds to a difference in a degree of information noise between the synthetic portion and the source portion.
[0120] The generating of the training data at 12 includes causing, at 13, a synthesizing model to generate the synthetic portion based on the source portion. The synthesizing model is an example of an artificial neural network and is trained to generate a synthetic portion based on a source portion.
[0121] The generating of the training data at 12 further includes replacing, at 14, the source portion with the synthetic portion, such that the training data include the synthetic portion instead of the source portion.
[0122] At 15, the storage portion 3 deletes the environmental data after the generating of the training data at 12.
[0123] At 16, the communication portion 4 provides the training data to training circuitry for finetuning a control network of the robot based on the training data.
[0124] It is noted that, in some embodiments, the environmental data include image data. In some embodiments, the environmental data include audio data. In some embodiments, the environmental data include both image data and audio data. In some embodiments, the environmental data additionally or alternatively include depth data.
[0125] It is further noted that, in some embodiments, the replacing of the source portion with the synthetic portion at 14 is omitted and, e.g., the circuitry 1 generates the training data without including a portion of the environmental data in the training data. Sony Semiconductor Solutions Corporation et al.
[0126] It is also noted that, in some embodiments, the deleting of the environmental data at 15 is omitted, and the storage portion 3 stores the environmental data, e.g., for logging or for further processing.
[0127] Fig. 3 illustrates a first embodiment of a system 20. The system 20 is configured as a robot.
[0128] The system 20 includes circuitry 21, which is an example of the circuitry 1 of Fig. 1 and is configured to perform the method 10 of Fig. 2.
[0129] The system 20 includes a camera 22 and a microphone 23. The camera 22 and the microphone 23 are examples of environmental sensors and are configured to acquire environmental data that indicate the environment of the robot 20, and to provide the environmental data to the circuitry 21. The camera 22 acquires image data as environmental data, and the microphone 23 acquires audio data as environmental data.
[0130] The system 20 further includes control circuitry 24 and training circuitry 25.
[0131] The control circuitry 24 is configured to execute a control network. The control network includes an artificial neural network and is configured to control the robot 20 based on environmental data acquired by the environmental sensors 22 and 23. The control network includes weights that determine an output of the control network.
[0132] When the circuitry 21 provides, at 16 of the method 10, training data to the training circuitry 25, the training circuitry 25 performs finetuning of the robot by determining updated weights for the control network based on the training data. The training circuitry 25 then provides the updated weights to the control network of the control circuitry 24.
[0133] After receiving the updated weights from the training circuitry 25, the control network controls the robot 20 based on the updated weights.
[0134] It is noted that, in some embodiments, the camera 22 or the microphone 23 is omitted, and / or the system 20 further includes a depth sensor.
[0135] Fig. 4 illustrates a second embodiment of a system 30.
[0136] The system 30 includes a first robot 31. The first robot 31 differs from the robot 20 of Fig. 3 in that the first robot 31 does not include training circuitry for finetuning the first robot 31.
[0137] The first robot 31 includes circuitry 32, which is an example of the circuitry 1 of Fig. 1 and is configured to perform the method 10 of Fig. 2. Sony Semiconductor Solutions Corporation et al.
[0138] The first robot 31 further includes a camera 33, a microphone 34 and control circuitry 35, which are configured similar to the camera 22, the microphone 23 and the control circuitry 24 of Fig 3, respectively.
[0139] The system 30 further includes a server 36. The server 36 includes training circuitry 37 and is connected to a communication network 38. At 16 of the method 10, the circuitry 32 provides the generated training data to the training circuitry 37, and the training circuitry 37 receives the training data from the circuitry 31 via the communication network 38.
[0140] The system 30 further includes a second robot 39. The second robot 39 is configured similar to the first robot 31. The second robot 39 generates further training data that indicate an environment of the second robot 39, and provides the further training data to the training circuitry 37. The training circuitry 37 also receives the further training data via the communication network 38.
[0141] The training circuitry 37 then performs finetuning of the first robot 31 and of the second robot 39 by determining, based on the training data and on the further training data, updated weights for the control network of the control circuitry 35, as described with respect to Fig. 3, as well as for a control network of the second robot 39.
[0142] The training circuitry 37 provides the updated weights to the control network of the control circuitry 35 in the first robot 31 and to the second robot 39 via the communication network 38.
[0143] As described with respect to Fig. 3, after receiving the updated weights from the training circuitry 37, the control network controls the robot 31 based on the updated weights.
[0144] Fig. 5 illustrates an embodiment of a general -purpose computer 150. The general -purpose computer 150 can be implemented such that it can basically function as any type of circuitry, for example, the circuitry 1 of Fig. 1, the circuitry 21 of Fig. 3, the circuitry 32 of Fig. 4, the control circuitry 24 of Fig. 3, the control circuitry 35 of Fig. 4, the training circuitry 25 of Fig. 3, the training circuitry 37 of Fig. 4, the server 36 of Fig. 4, or the like. The general-purpose computer 150 can be configured as a smartphone, smart glasses, a head-mounted display, a smartwatch, a mobile phone, a mobile tablet, a laptop, a server, a terminal device, or the like. The general -purpose computer 150 is an example of a data processing apparatus that includes circuitry that is configured to perform the method according to the present technology (e.g., the method 10 of Fig. 2). The computer has components 151 to 161, which can form a circuitry, such as any one of the portion 2, the unit 3, the portion 4, or the like, as described herein. Sony Semiconductor Solutions Corporation et al.
[0145] Embodiments which use software, firmware, programs or the like for performing the methods as described herein can be installed on computer 150, which is then configured to be suitable for the concrete embodiment.
[0146] The computer 150 has a CPU 151 (Central Processing Unit), which can execute various types of procedures and methods as described herein, for example, in accordance with programs stored in a read-only memory (ROM) 152, stored in a storage 157 and loaded into a random-access memory (RAM) 153, stored on a medium 160 which can be inserted in a respective drive 159, etc.
[0147] Furthermore, the computer 150 includes an artificial intelligence (Al) processor 151a. The Al processor 151a may include a graphics processing unit (GPU) and / or a tensor processing unit (TPU). The Al processor 151a may be configured to execute an Al model (e.g., an artificial neural network), for example, the synthesizing model executed by the circuitry 1, 2 lor 32, or the control network executed by the control circuitry 24 or 35.
[0148] The CPU 151, the ROM 152 and the RAM 153 are connected with a bus 161, which in turn is connected to an input / output interface 154. The number of CPUs, memories and storages is only exemplary, and the skilled person will appreciate that the computer 150 can be adapted and configured accordingly for meeting specific requirements which arise when it functions as an information processing apparatus according to the present technology.
[0149] At the input / output interface 154, several components are connected: an input 155, an output 156, the storage 157, a communication interface 158 and the drive 159, into which a medium 160 (compact disc (CD), digital video disc (DVD), universal serial bus (USB) flash drive, secure digital (SD) card, CompactFlash (CF) memory, or the like) can be inserted.
[0150] The input 155 can be a pointer device (mouse, graphic table, or the like), a keyboard, a microphone, a camera, a touchscreen, an eye-tracking unit etc.
[0151] The output 156 can have a display (liquid crystal display (LCD), cathode ray tube (CRT) display, light-emitting diode (LED) display, electronic paper, etc.; e.g., included in a touchscreen), loudspeakers, etc.
[0152] The storage 157 can have a hard disk drive (HDD), a solid-state drive (SSD), a flash drive and the like.
[0153] The communication interface 158 can be adapted to communicate, for example, via universal serial bus (USB), a serial port (RS-232), parallel port (IEEE 1284), a local area network (LAN; Sony Semiconductor Solutions Corporation et al. e.g., ethemet), wireless local area network (WLAN; e.g., Wi-Fi, IEEE 802.11), mobile telecommunications system (GSM, UMTS, LTE, NR etc.), Bluetooth, near-fteld communication (NFC), ZigBee, infrared, etc.
[0154] It should be noted that the description above only pertains to an example configuration of computer 150. Alternative configurations may be implemented with additional or other sensors, storage devices, interfaces or the like. For example, the communication interface 158 may support other radio access technologies than the mentioned UMTS, LTE and NR.
[0155] Accordingly, the disclosure provides finetuning a robot based on training data that include a synthetic portion.
[0156] In some embodiments, a robot generates, based on environmental data, training data to finetune its neural networks. The training data with the synthetic portion (e.g., image data, audio data, depth data, etc.) may be generated on the robot’s onboard processor(s). The original environmental data may be deleted and the training data (which may indicate, e.g., synthesized image variations) may be kept for training.
[0157] The robot may synthesize, as the synthetic portion in the training data, synthetic replacement images for original images of the environmental data to preserve privacy. The robot may provide the replacement images to a finetuning system (e.g., training circuitry) and may delete the original images (e.g., the environmental data). The finetuning system may train on the synthetic replacement images without restrictions due to data protection requirements, and / or the training may be enhanced by using the synthetic replacement images. Further, no consent of a person in an environment of the robot may be required when storing the training data with the synthetic images (e.g., for later use) because the training data may not include personal information of the person. Weights of the neural networks of the robot may then be updated by the finetuning system.
[0158] Further, a camera device is proposed that may output image variations where privacy is not an issue in some embodiments. The camera device may include the circuitry described herein and may generate training data with a synthetic portion that may represent a virtual object.
[0159] It should be recognized that the embodiments describe methods with an exemplary ordering of method steps. The specific ordering of method steps is however given for illustrative purposes only and should not be construed as binding. For example, the ordering of 15 and 16 in the embodiment of Fig. 2 may be exchanged. Other changes of the ordering of method steps may be apparent to the skilled person. Sony Semiconductor Solutions Corporation et al.
[0160] Please note that the division of the circuitry 1 into portions 2, 3 and 4 is only made for illustration purposes and that the present disclosure is not limited to any specific division of functions in specific units. For instance, the circuitry 1 could be implemented by a respective programmed processor, field programmable gate array (FPGA) and the like.
[0161] All units and entities described in this specification and claimed in the appended claims can, if not stated otherwise, be implemented as integrated circuit logic, for example on a chip, and functionality provided by such units and entities can, if not stated otherwise, be implemented by software.
[0162] In so far as the embodiments of the disclosure described above are implemented, at least in part, using software-controlled data processing apparatus, it will be appreciated that a computer program providing such software control and a transmission, storage or other medium by which such a computer program is provided are envisaged as aspects of the present disclosure.
[0163] Note that the present technology can also be configured as described below.
[0164] (1) Circuitry, configured to: obtain environmental data indicating an environment of a robot; generate training data based on the environmental data, wherein the training data include a synthetic portion that has a predetermined similarity to a corresponding source portion of the environmental data, wherein the synthetic portion represents a virtual object in the environment, and wherein the source portion represents a real object in the environment; and provide the training data to training circuitry for finetuning a control network of the robot based on the training data.
[0165] (2) The circuitry of (1), wherein the synthetic portion corresponds to a variation of the source portion.
[0166] (3) The circuitry of (2), wherein the variation of the source portion is based on a random value in a latent space.
[0167] (4) The circuitry of any one of (1) to (3), wherein the generating of the training data includes replacing the source portion with the synthetic portion. Sony Semiconductor Solutions Corporation et al.
[0168] (5) The circuitry of any one of (1) to (4), wherein the predetermined similarity corresponds to a degree of deviation between the source portion and the synthetic portion.
[0169] (6) The circuitry of any one of (1) to (5), wherein the predetermined similarity corresponds to a difference in a degree of information noise between the synthetic portion and the source portion.
[0170] (7) The circuitry of any one of (1) to (6), wherein the generating of the training data includes causing an artificial neural network to generate the synthetic portion based on the source portion.
[0171] (8) The circuitry of any one of (1) to (7), wherein the environmental data include image data.
[0172] (9) The circuitry of any one of (1) to (8), wherein the environmental data include sound data.
[0173] (10) The circuitry of any one of (1) to (9), further configured to: delete the environmental data after generating the training data.
[0174] (11) A method, compri sing : obtaining environmental data indicating an environment of a robot; generating training data based on the environmental data, wherein the training data include a synthetic portion that has a predetermined similarity to a corresponding source portion of the environmental data, wherein the synthetic portion represents a virtual object in the environment, and wherein the source portion represents a real object in the environment; and providing the training data to training circuitry for finetuning a control network of the robot based on the training data.
[0175] (12) The method of ( 11 ), wherein the synthetic portion corresponds to a variation of the source portion.
[0176] (13) The method of ( 12), wherein the variation of the source portion is based on a random value in a latent space.
[0177] (14) The method of any one of (11) to (13), wherein the generating of the training data includes replacing the source portion with the synthetic portion. Sony Semiconductor Solutions Corporation et al.
[0178] (15) The method of any one of (11) to (14), wherein the predetermined similarity corresponds to a degree of deviation between the source portion and the synthetic portion.
[0179] (16) The method of any one of (11) to (15), wherein the predetermined similarity corresponds to a difference in a degree of information noise between the synthetic portion and the source portion.
[0180] (17) The method of any one of (11) to (16), wherein the generating of the training data includes causing an artificial neural network to generate the synthetic portion based on the source portion.
[0181] (18) The method of any one of (11) to (17), wherein the environmental data include image data.
[0182] (19) The method of any one of (11) to (18), wherein the environmental data include sound data.
[0183] (20) The method of any one of (11) to (19), further comprising: deleting the environmental data after generating the training data.
[0184] (21) A system, comprising: the circuitry according to any one of (1) to (10); a sensor configured to acquire environmental data indicating the environment of the robot and to provide the environmental data to the circuitry; the training circuitry, wherein the training circuitry is configured to determine updated weights of the control network based on the training data and to provide the updated weights to the control network; and the control network, wherein the control network is configured to control the robot based on environmental data acquired by the sensor and based on the updated weights determined by the training circuitry.
[0185] (22) The system of (21), wherein the training circuitry is configured to: receive, via a communication network, the training data from the circuitry and further training data that indicate an environment of a further robot; determine the updated weights based on the training data and on the further training data; and provide the updated weights to the control network via the communication network. Sony Semiconductor Solutions Corporation et al.
[0186] (23) A computer program comprising program code causing a computer to perform the method according to anyone of (11) to (20), when being carried out on a computer.
[0187] (24) A non-transitory computer-readable recording medium that stores therein a computer program product, which, when executed by a processor, causes the method according to anyone of (11) to (20) to be performed.
Claims
Sony Semiconductor Solutions Corporation et al.CLAIMS1. Circuitry, configured to: obtain environmental data indicating an environment of a robot; generate training data based on the environmental data, wherein the training data include a synthetic portion that has a predetermined similarity to a corresponding source portion of the environmental data, wherein the synthetic portion represents a virtual object in the environment, and wherein the source portion represents a real object in the environment; and provide the training data to training circuitry for finetuning a control network of the robot based on the training data.
2. The circuitry of claim 1, wherein the synthetic portion corresponds to a variation of the source portion.
3. The circuitry of claim 2, wherein the variation of the source portion is based on a random value in a latent space.
4. The circuitry of claim 1, wherein the generating of the training data includes replacing the source portion with the synthetic portion.
5. The circuitry of claim 1, wherein the predetermined similarity corresponds to a degree of deviation between the source portion and the synthetic portion.
6. The circuitry of claim 1, wherein the predetermined similarity corresponds to a difference in a degree of information noise between the synthetic portion and the source portion.
7. The circuitry of claim 1, wherein the generating of the training data includes causing an artificial neural network to generate the synthetic portion based on the source portion.
8. The circuitry of claim 1, wherein the environmental data include at least one of image data and sound data.
9. The circuitry of claim 1, further configured to: delete the environmental data after generating the training data.Sony Semiconductor Solutions Corporation et al.
10. A method, comprising: obtaining environmental data indicating an environment of a robot; generating training data based on the environmental data, wherein the training data include a synthetic portion that has a predetermined similarity to a corresponding source portion of the environmental data, wherein the synthetic portion represents a virtual object in the environment, and wherein the source portion represents a real object in the environment; and providing the training data to training circuitry for finetuning a control network of the robot based on the training data.
11. The method of claim 10, wherein the synthetic portion corresponds to a variation of the source portion.
12. The method of claim 11, wherein the variation of the source portion is based on a random value in a latent space.
13. The method of claim 10, wherein the generating of the training data includes replacing the source portion with the synthetic portion.
14. The method of claim 10, wherein the predetermined similarity corresponds to a degree of deviation between the source portion and the synthetic portion.
15. The method of claim 10, wherein the predetermined similarity corresponds to a difference in a degree of information noise between the synthetic portion and the source portion.
16. The method of claim 10, wherein the generating of the training data includes causing an artificial neural network to generate the synthetic portion based on the source portion.
17. The method of claim 10, wherein the environmental data include at least one of image data and sound data.
18. The method of claim 10, further comprising: deleting the environmental data after generating the training data.
19. A system, comprising: circuitry that is configured to:Sony Semiconductor Solutions Corporation et al. obtain environmental data indicating an environment of a robot; generate training data based on the environmental data, wherein the training data include a synthetic portion that has a predetermined similarity to a corresponding source portion of the environmental data, wherein the synthetic portion represents a virtual object in the environment, and wherein the source portion represents a real object in the environment; and provide the training data to training circuitry for finetuning a control network of the robot based on the training data; a sensor configured to acquire environmental data indicating the environment of the robot and to provide the environmental data to the circuitry; the training circuitry, wherein the training circuitry is configured to determine updated weights of the control network based on the training data and to provide the updated weights to the control network; and the control network, wherein the control network is configured to control the robot based on environmental data acquired by the sensor and based on the updated weights determined by the training circuitry.
20. The system of claim 19, wherein the training circuitry is configured to: receive, via a communication network, the training data from the circuitry and further training data that indicate an environment of a further robot; determine the updated weights based on the training data and on the further training data; and provide the updated weights to the control network via the communication network.
Citation Information
Patent Citations
Training machine learning models using simulation for robotics systems and applications
US20240095527A1
Hybrid audio synthesis using neural networks
WO2020010338A1
Control of an industrial robot for a gripping task
WO2023078884A1
System and training system for computer implemented generating synthetic images representing compressed versions of original images
WO2023175199A2
Training policy neural networks in simulation using scene synthesis machine learning models
WO2024056892A1