Data acquisition method, device and system
By combining a multi-antenna receiving matrix with mobile information as a data acquisition strategy, and utilizing reinforcement learning to optimize robot path and communication resource allocation, the problem of real-time and efficient sensor data acquisition in the Internet of Things is solved, and the robustness and signal quality of the system are improved.
Patent Information
- Application Number
- PCT/CN2024/088314
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-04-17
- Publication Date
- 2025-10-23
AI Technical Summary
Existing data acquisition methods face challenges such as high sensor power consumption, signal attenuation, obstacle interference, network congestion, and insufficient robot flexibility when collecting large amounts of sensor data in the Internet of Things in real time and efficiently, thus failing to fully realize the potential of robot mobility and communication technology.
A data acquisition strategy combining a multi-antenna receiving matrix and motion information is adopted. By optimizing robot path and communication resource allocation through reinforcement learning, efficient data acquisition in dynamic environments is achieved.
It improves the real-time performance and efficiency of data acquisition, reduces sensor power consumption, enhances network robustness and signal quality, adapts to dynamic and complex environments, and improves system reliability and stability.
Smart Images

Figure CN2024088314_23102025_PF_FP_ABST
Abstract
Description
Data collection method, device and system TECHNICAL FIELD
[0001] The present application relates to the technical field of wireless communication, and more particularly, to a data collection method, device and system. BACKGROUND
[0002] With the development of communication technology, higher requirements are put forward for the rate of data collection. For example, in the Internet of Things, the wide deployment of sensors not only leads to the surge of big data, but also some important data (e.g., sensing data) in it needs to be collected in real time and efficiently.
[0003] In order to avoid the problems such as large delay in remote data transmission and complex multi-hop transmission, deploying mobile intelligent devices such as robots for data collection becomes a solution. However, how to fully exert the flexibility and mobility of robots becomes a problem to be solved.
[0004] SUMMARY
[0005] The present application provides a data collection method, device and system, and the following introduces each aspect related to the embodiments of the present application.
[0006] In a first aspect, a data collection method is provided, comprising: a first device receiving data sent by at least one second device according to first information in a first time unit, to perform data collection; and the first device sending relevant information of performing the data collection to a third device; wherein the first information comprises one or more of the following information: a multi-antenna receiving matrix of the first device in the first time unit; movement information of the first device in the first time unit; and an allocation strategy of a first communication resource used by the first device to perform the data collection.
[0007] In a second aspect, a data collection method is provided, comprising: a second device sending data to a first device in a first time unit; wherein the data is data collected by the first device according to first information, and the first information comprises one or more of the following information: a multi-antenna receiving matrix of the first device in the first time unit; movement information of the first device in the first time unit; and an allocation strategy of a first communication resource used by the first device to perform the data collection.
[0008] In a third aspect, a data collection method is provided, comprising: receiving, by a third device, information related to data collection sent by a first device; wherein the information related to data collection is used to indicate data collection performed by the first device on at least one second device based on first information in a first time unit, and the first information comprises one or more of the following: a multiple-antenna receiving matrix of the first device in the first time unit; a moving track of the first device in the first time unit; and an allocation strategy of a first communication resource used by the first device to perform the data collection.
[0009] In a fourth aspect, a data collection device is provided, which is a first device, comprising: a receiving unit configured to receive data sent by at least one second device in a first time unit according to first information to perform data collection; and a sending unit configured to send information related to the data collection to a third device; wherein the first information comprises one or more of the following: a multiple-antenna receiving matrix of the first device in the first time unit; moving information of the first device in the first time unit; and an allocation strategy of a first communication resource used by the first device to perform the data collection.
[0010] In a fifth aspect, a data collection device is provided, which is a second device, comprising: a sending unit configured to send data to a first device in a first time unit; wherein the data is data collected by the first device according to first information, and the first information comprises one or more of the following: a multiple-antenna receiving matrix of the first device in the first time unit; moving information of the first device in the first time unit; and an allocation strategy of a first communication resource used by the first device to perform the data collection.
[0011] In a sixth aspect, a data collection device is provided, which is a third device, comprising: a receiving unit configured to receive information related to data collection sent by a first device; wherein the information related to data collection is used to indicate data collection performed by the first device on at least one second device based on first information in a first time unit, and the first information comprises one or more of the following: a multiple-antenna receiving matrix of the first device in the first time unit; a moving track of the first device in the first time unit; and an allocation strategy of a first communication resource used by the first device to perform the data collection.
[0012] In a seventh aspect, a data collection system is provided, which comprises a control device configured to control a plurality of first devices in the data collection system to perform the method according to the first aspect, and / or the control device is configured to control a third device in the data collection system to perform the method according to the third aspect.
[0013] In an eighth aspect, a communication device is provided, which comprises a memory configured to store a program, and a processor configured to invoke the program in the memory to perform the method according to any one of the first aspect to the third aspect.
[0014] In a ninth aspect, a device is provided, which comprises a processor configured to invoke a program from a memory to perform the method according to any one of the first aspect to the third aspect.
[0015] In a tenth aspect, a chip is provided, which comprises a processor configured to invoke a program from a memory to cause a device in which the chip is installed to perform the method according to any one of the first aspect to the third aspect.
[0016] In an eleventh aspect, a computer readable storage medium is provided, which has stored thereon a program, the program causing a computer to perform the method according to any one of the first aspect to the third aspect.
[0017] In a twelfth aspect, a computer program product is provided, which comprises a program, the program causing a computer to perform the method according to any one of the first aspect to the third aspect.
[0018] In a thirteenth aspect, a computer program is provided, which causes a computer to perform the method according to any one of the first aspect to the third aspect.
[0019] In the first device in the embodiments of the present application, the first information can comprise a multi-antenna receiving matrix of the first device in the first time unit, movement information, and / or resource allocation information of the first device for data collection. Therefore, the first device can flexibly perform data collection based on multi-antenna and mobility in the first time unit, thereby improving real-time performance and efficiency of data collection. BRIEF DESCRIPTION OF DRAWINGS
[0020] FIG. 1 is a wireless communication system to which the embodiments of the present application are applied.
[0021] FIG. 2 is a flow diagram of a data collection method according to an embodiment of the present application.
[0022] FIG. 3 is a flow diagram of a possible implementation of the method shown in FIG. 2.
[0023] FIG. 4 is a structural schematic diagram of a multi-robot assisted data acquisition system according to an embodiment of the present application.
[0024] FIG. 5 is a structural schematic diagram of a first device according to an embodiment of the present application.
[0025] FIG. 6 is a structural schematic diagram of a second device according to an embodiment of the present application.
[0026] FIG. 7 is a structural schematic diagram of a third device according to an embodiment of the present application.
[0027] FIG. 8 is a structural schematic diagram of a control device of a data acquisition system according to an embodiment of the present application.
[0028] FIG. 9 is a structural schematic diagram of an electronic device according to an embodiment of the present application.
[0029] FIG. 10 is a schematic block diagram of a communication device according to an embodiment of the present application. DETAILED DESCRIPTION
[0030] The technical solutions in the embodiments of the present application will be described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments of the present application. Any other embodiments obtained by those of ordinary skill in the art without creative effort based on the embodiments in the present application are within the scope of protection of the present application.
[0031] The embodiments of the present application can be applied to various communication systems. For example, the embodiments of the present application can be applied to a global system of mobile communication (GSM) system, a code division multiple access (CDMA) system, a wideband code division multiple access (WCDMA) system, a general packet radio service (GPRS), a long term evolution (LTE) system, an advanced long term evolution (LTE-A) system, a new radio (NR) system, an evolved system of the NR system, an LTE-based access to unlicensed spectrum (LTE-U) system, an NR-based access to unlicensed spectrum (NR-U) system, a universal mobile telecommunication system (UMTS), a wireless local area networks (WLAN), a wireless fidelity (WiFi), a 5th-generation (5G) system. The embodiments of the present application can also be applied to other communication systems, for example, a future communication system. The future communication system can be, for example, a 6th-generation (6G) mobile communication system, or a satellite communication system, etc.
[0032] The conventional communication system supports a limited number of connections, which is easy to implement. However, with the development of communication technology, the communication system can not only support traditional cellular communication, but also support one or more types of other communications. For example, the communication system can support one or more of the following communications: device to device (D2D) communication, machine to machine (M2M) communication, machine type communication (MTC), enhanced MTC (eMTC), vehicle to vehicle (V2V) communication, and vehicle to everything (V2X) communication, and the like. The embodiments of the present application can also be applied to a communication system supporting the above communication modes.
[0033] The communication system in the embodiments of the present application can be applied to a carrier aggregation (CA) scenario, a dual connectivity (DC) scenario, and a standalone (SA) network deployment scenario.
[0034] The communication system in the embodiments of the present application can be applied to unlicensed spectrum. The unlicensed spectrum can also be considered as shared spectrum. Alternatively, the communication system in the embodiments of the present application can also be applied to licensed spectrum. The licensed spectrum can also be considered as dedicated spectrum.
[0035] The embodiments of the present application can be applied to a non-terrestrial network (NTN) system. As an example, the NTN system can include a 4G-based NTN system, an NR-based NTN system, an internet of things (IoT)-based NTN system, and a narrow band internet of things (NB-IoT)-based NTN system.
[0036] The communication system can include one or more terminal devices. The terminal device mentioned in the embodiments of the present application can also be referred to as user equipment (UE), access terminal, subscriber unit, subscriber station, mobile station, mobile station (MS), mobile terminal (MT), remote station, remote terminal, mobile device, user terminal, terminal, wireless communication device, user agent, or user equipment, etc.
[0037] In some embodiments, the terminal device can be a station (STATION, ST) in a WLAN. In some embodiments, the terminal device can be a cellular phone, a cordless phone, a session initiation protocol (SIP) phone, a wireless local loop (WLL) station, a personal digital assistant (PDA) device, a handheld device having wireless communication function, a computing device, or other processing device connected to a wireless modem, an in-vehicle device, a wearable device, a terminal device in a next-generation communication system (e.g., an NR system), or a terminal device in a future evolved public land mobile network (PLMN) network, etc.
[0038] In some embodiments, the terminal device can be a device that provides voice and / or data connectivity to a user. For example, the terminal device can be a handheld device having wireless connection function, an in-vehicle device, etc. As some specific examples, the terminal device can be a mobile phone, a Pad, a notebook computer, a palmtop computer, a mobile internet device (MID), a wearable device, a virtual reality (VR) device, an augmented reality (AR) device, a wireless terminal in industrial control, a wireless terminal in self driving, a wireless terminal in remote medical surgery, a wireless terminal in smart grid, a wireless terminal in transportation safety, a wireless terminal in smart city, a wireless terminal in smart home, etc.
[0039] In some embodiments, the terminal device can be deployed on land. For example, the terminal device can be deployed indoors or outdoors. In some embodiments, the terminal device can be deployed on water surface, such as on a ship. In some embodiments, the terminal device can be deployed in air, such as on an airplane, a balloon, and a satellite.
[0040] In addition to the terminal device, the communication system can also include one or more network devices. The network device in the embodiments of the present application can be a device for communicating with the terminal device, which can also be referred to as an access network device or a radio access network device. The network device can be, for example, a base station. The network device in the embodiments of the present application can refer to a radio access network (RAN) node (or device) that accesses the terminal device to a wireless network. The base station can broadly cover various names in the following or be replaced by the following names, such as: Node B (NodeB), evolved Node B (eNB), next generation Node B (gNB), relay station, access point, transmitting and receiving point (TRP), transmitting point (TP), master station MeNB, auxiliary station SeNB, multi-standard radio (MSR) node, home base station, network controller, access node, wireless node, access point (AP), transmission node, transceiver node, baseband unit (BBU), remote radio unit (RRU), active antenna unit (AAU), remote radio head (RRH), central unit (CU), distributed unit (DU), positioning node, etc. The base station can be a macro base station, a micro base station, a relay node, a donor node, or the like, or a combination thereof. The base station can also refer to a communication module, modem, or chip used in the aforementioned devices or apparatuses. The base station can also be a mobile switching center and a device that performs the function of a base station in D2D, V2X, M2M communication, a network side device in a 6G network, a device that performs the function of a base station in a future communication system, etc. The base station can support networks of the same or different access technologies. The embodiments of the present application do not limit the specific technology and specific device form adopted by the network device.
[0041] The base station can be fixed or mobile. For example, a helicopter or a drone can be configured to act as a mobile base station, and one or more cells can move according to the location of the mobile base station. In other examples, the helicopter or the drone can be configured to serve as a device that communicates with another base station.
[0042] In some deployments, the network device in the embodiments of the present application can refer to a CU or a DU, or the network device includes a CU and a DU. The gNB can also include an AAU.
[0043] By way of example and not limitation, in embodiments of the application, a network device can have mobile characteristics, e.g., the network device can be a mobile device. In some embodiments of the application, a network device can be a satellite, a balloon station. In some embodiments of the application, a network device can also be a base station disposed at a location on land, water, etc.
[0044] In embodiments of the application, a network device can serve a cell, and a terminal device communicates with the network device through a transmission resource (e.g., a frequency domain resource, or a spectrum resource) used by the cell. The cell can be a cell corresponding to the network device (e.g., a base station), and the cell can belong to a macro base station or a base station corresponding to a small cell. The small cell can include a metro cell, a micro cell, a pico cell, a femto cell, etc., and these small cells have small coverage and low transmit power, and are suitable for providing high-rate data transmission services.
[0045] In some embodiments, the application can also be applied to an artificial intelligence communication system. One of the goals of the 3rd Generation Partnership Project (3GPP) Release 18 is to enhance the functionality of 5G and extend its application to new devices, deployments, and industries. As network design becomes more and more complex, which can include a wide range of deployment and usage options, traditional methods will not be able to provide a quick solution. Since it is costly and inefficient to manually reconfigure a cellular communication system, it is necessary to use artificial intelligence (AI) and machine learning (ML) to automate the operational process to reduce costs by automating functions that require human interaction. By way of example, AI and ML can solve complex and unstructured network problems by using a large amount of data collected from wireless networks.
[0046] As an example, AI can be used for intelligent network operations in the core network and RAN. For example, AI can be used to enhance quality of service (QoS), improve efficiency, simplify deployment, and improve security.
[0047] As an example, on-device artificial intelligence will benefit the overall communication system. AI's potential support capabilities are radio awareness. AI can provide valuable knowledge through environmental and contextual awareness, thus reducing overhead and latency. With radio awareness, the communication system can support enhanced device experiences, such as smart beamforming and power management. In addition, AI helps improve system performance, such as reducing interference, achieving better spectrum utilization, and improving radio security. For example, AI helps better detect and prevent malicious attacks.
[0048] As an example, in the face of unknown environments without prior knowledge, deep reinforcement learning has strong adaptability to dynamic environments, so deep reinforcement learning is considered an effective method to achieve network intelligence. At the same time, network intelligence will gradually be pushed to the edge end to run, to further provide distributed autonomous intelligence.
[0049] By way of example, FIG. 1 is a schematic diagram of an architecture of a communication system provided by embodiments of the present application. As shown in FIG. 1, the communication system 100 can include a network device 110, which can be a device that communicates with a terminal device 120 (or called a communication terminal, a terminal). The network device 110 can provide communication coverage for a specific geographic area, and can communicate with terminal devices located within the coverage area.
[0050] By way of example, FIG. 1 shows one network device and two terminal devices. In some embodiments of the present application, the communication system 100 can include multiple network devices and each network device can include other numbers of terminal devices within its coverage, which is not limited.
[0051] In embodiments of the present application, the communication system shown in FIG. 1 further includes a mobility management entity (MME), an access and mobility management function (AMF), and other network entities, which are not limited in embodiments of the present application.
[0052] It should be understood that the devices with communication functions in the network / system in embodiments of the present application can be referred to as communication devices. Taking the communication system 100 shown in FIG. 1 as an example, the communication devices can include the network device 110 and the terminal device 120 with communication functions, which can be the specific devices described above, and will not be described here. The communication devices can also include other devices in the communication system 100, such as network controllers, mobility management entities, and other network entities, which are not limited in embodiments of the present application.
[0053] To facilitate the detailed description of the innovation points of the technical solutions, some related technical knowledge involved in the embodiments of the present application is introduced first. The following related technologies can be combined with the technical solutions of the embodiments of the present application as optional solutions, which all belong to the protection scope of the embodiments of the present application. The embodiments of the present application include at least part of the following contents.
[0054] With the development of technology, future intelligent production / life requires seamless integration of multiple technologies such as communication, control and AI computing. In some embodiments, with the continuous progress of wireless communication technology and the reduction of manufacturing cost, the data collected by various sensors contains a large amount of information. For data-driven tasks, efficient data collection is particularly important. For example, with the development of the Internet of Things, the widespread deployment of sensors will lead to an explosion of big data. Important perception data in these big data needs to be collected in real time and efficiently.
[0055] However, real-time transmission of massive data faces challenges such as small sensor transmission power, limited battery energy, short transmission distance and non-line-of-sight channel. Therefore, how to efficiently collect real-time or dynamic data is a challenge that needs to be solved urgently.
[0056] However, the data collection method in the related art may not meet the needs of real-time and efficient collection. Specifically, when sensors are distributed in a large range, the related sensor network mainly uses three methods of long-distance transmission, multi-hop transmission and unmanned aerial vehicle for data collection.
[0057] Long-distance transmission usually requires sensor nodes to increase transmission power to overcome signal attenuation and transmission distance limitations, which will increase the energy consumption of sensor nodes. Further, long-distance transmission is also susceptible to signal attenuation, interference and obstacles, thereby increasing the risk of data packet loss or damage.
[0058] Multi-hop transmission is relatively complex and can cause network congestion. In particular, when the traffic at the intermediate nodes is too large, congestion is likely to occur. In addition, multi-hop transmission is highly dependent on multiple intermediate nodes. Performance problems of any of the multiple intermediate nodes can affect the performance of the entire transmission link, so multi-hop transmission has the problems of unstable performance and complex maintenance.
[0059] Unmanned aerial vehicle data collection is limited by the site and the battery, and often ignores the impact of non-line-of-sight links caused by obstacles on the system. In addition, due to the limited computing power of the unmanned aerial vehicle, the deployment and development of edge intelligence are greatly limited.
[0060] To solve these problems, seamless integration of communication technology, control technology, and AI computing technology has become a major research direction. The application of these technologies in data collection is exemplarily illustrated below.
[0061] As an example, movable intelligent devices such as robots are used for data collection. In recent years, intelligent manufacturing and robot automation technology have developed rapidly. In order to avoid data remote transmission delay, multi-hop transmission complexity, etc., to achieve data collection in toxic and other harmful human environments, high pressure and other human inaccessible scenes, deploying multiple robots for data collection has become a promising solution. This solution can meet the real-time needs of data-driven and effectively deal with the difficulties that dynamic time-varying channels may bring. Multiple robots can use their mobility to collect data from nearby sensors at a shorter communication distance, thereby reducing the communication energy consumption of sensors and prolonging the life cycle of sensor networks. At the same time, the flexibility of robots allows them to quickly adapt to dynamic and complex environments, ensuring high signal quality links. In addition, robots often have certain computing resources. Robots can quickly process sensor information at the local end to meet the low latency requirements of time-sensitive services. Therefore, deploying multiple robots in actual production and life can speed up data collection and make accurate and timely decisions.
[0062] As an example, for a multi-robot system, it is instructive to make full use of limited communication resources and consider appropriate communication technologies to provide reliable links. In future networks, multi-antenna beamforming technology is considered an indispensable component. Multi-antenna technology uses multiple antennas to enhance the performance of communication systems. Exemplarily, multi-antenna technology can achieve directional transmission of signals through beamforming technology, improving signal transmission efficiency and coverage. Under scarce spectrum resources, multi-antenna technology can provide higher spectral efficiency and larger-scale connectivity, thereby playing an important role in multi-robot systems. Therefore, through multi-antenna technology, better signal coverage and transmission quality can be achieved, improving the reliability and stability of the system. However, designing a multi-antenna non-orthogonal multiple access technology network is still challenging. The beams of the multi-antenna system need to be carefully designed to ensure clustering without inter-cluster interference. Considering the characteristics and advantages of multi-antenna technology, effective design and optimization of the communication network of the multi-robot system will be an important direction for future research.
[0063] As an example, in a multi-robot assisted system, in the face of multi-device collaborative training problems, existing research on distributed machine learning has shown great interest in federated learning. This research enables devices to collaboratively train models while maintaining original data. Compared with centralized methods, federated learning based on deep reinforcement learning supports scalable network size and can accelerate convergence speed, making it suitable for multi-robot assisted systems.
[0064] Based on this, a multi-robot assisted data collection system integrating multiple technologies is applied to data collection. However, in actual application, there are various problems to be solved in the above-mentioned technologies, which will affect the efficiency and performance of the data collection system.
[0065] For example, the related multi-antenna sensor network often uses zero-forcing and other algorithms that require global information for data collection. The traditional algorithm such as zero-forcing usually needs global information for data collection decision, which means a large amount of calculation and communication overhead is needed to collect and process the information. Therefore, this technology will lead to high computational complexity of the algorithm, especially in large-scale sensor networks, which may consume a large amount of time and resources, affecting the real-time performance of data collection.
[0066] For another example, the related sensor data collection often uses fixed communication strategy, which cannot cope with time-varying channel and dynamic complex environment. In addition, the fixed strategy will greatly affect the flexibility and mobility of multi-robot, and cannot adapt to the change of sensor deployment or number, so the network robustness is poor.
[0067] For another example, the related multi-robot system often considers the communication strategy and the moving strategy separately. However, communication and robot position are inseparable. Specifically, the robot can enhance communication by position movement, while the communication strategy will also be affected by the movement of multi-robot, such as beamforming with strong directivity. Therefore, considering communication and movement separately will make the network strategy performance poor, and cannot fully exert the potential of robot mobility and mobile communication technology.
[0068] In view of the above problems, an embodiment of the present application proposes a data collection method. Through this method, the first device (for example, robot) can receive the data of at least one second device (for example, sensor) in a first time unit according to the first information, and send the related information of data collection to the third device (for example, cloud device). Wherein, the first information can indicate the multi-antenna receiving matrix, the moving information and / or the resource allocation strategy of the first device, so as to ensure the real-time and high efficiency of the first device for data collection. Further, when there are multiple first devices in the data collection system, multiple first devices can collect the data of a part of second devices respectively, so as to improve the collection efficiency of massive data. In order to facilitate understanding, the data collection method is described in detail below in combination with FIG. 2. FIG. 2 is introduced from the perspective of interaction among the first device, the second device and the third device.
[0069] Referring to FIG. 2, in step S210, the first device receives the data sent by at least one second device. As shown in FIG. 2, the first device can receive the data sent by the at least one second device in a first time unit according to the first information, in order to perform data collection.
[0070] The first device can be a device with communication capability to facilitate data transmission and reception. In some embodiments, the first device can be any of the terminal devices described above. For example, a UE. In some embodiments, the first device can be any of the communication terminals used for data collection in a data collection system. In some embodiments, the first device can be a dedicated collection device in a data collection system to facilitate data collection. For example, the first device can be a robot assisting data collection as described above.
[0071] As a possible implementation, the first device can better implement its communication capability by configuring multiple antennas. That is, the first device can deploy multiple antenna technology to improve spectral efficiency. As an example, for a robot equipped with multiple antennas, a suitable multiple antenna reception matrix can be designed to efficiently receive information from multiple second devices.
[0072] As an example, the first device can be a robot with L antennas, where L is a positive integer.
[0073] As a possible implementation, the first device can receive data transmitted by the second device through a data line to facilitate data collection.
[0074] In some embodiments, the first device can be one of multiple first devices. That is, the data collection method can be used in a data collection system including multiple first devices. For example, the data collection system can include K robots, where K is a positive integer. In the data collection system, the set of robots can be represented as where 1≤k≤K when the first device is a robot k.
[0075] As a possible implementation, the multiple first devices in the data collection system can be multiple devices of different types, multiple devices of the same type, or multiple devices of different types, without limitation.
[0076] As a possible implementation, the number of first devices in the data collection system can vary.
[0077] In some embodiments, the first device is a movable data collection device. Taking a robot as an example, to improve the average rate of data transmission, multiple robots can avoid communication dead zones by moving, and can also avoid collisions with obstacles and / or other robots by cooperating.
[0078] As an example, the first device can be self-moving, such as a robot. As an example, the first device can be a device placed on a movable platform, and the position is changed based on the movement of the platform.
[0079] As an example, the position information of the first device at a certain time can be represented in a three-dimensional Cartesian coordinate system or in various other ways. For example, the first device is a robot k in a data collection system, and at time t (t is an integer), the position coordinates of the robot k can be represented as q k (t) = (x k (t), y k (t), H0), where H0may represent the height of the antenna of the robot. The antenna heights of different first devices can be different.
[0080] Optionally, the position of the first device can be a relative position with respect to a certain reference object, or an absolute position calculated by radar or indoor positioning, etc. Different determination methods will have different measurement errors. The embodiments of the present application do not limit this.
[0081] As a possible implementation manner, the movement trajectory of the first device can be designed or pre-configured. The movement trajectory of the first device can also be referred to as the path of the first device. Designing the movement trajectory is to perform path planning.
[0082] As an example, the first device or other device for controlling the first device can determine the movement trajectory of the first device according to the position information of the at least one second device, so as to perform efficient collection and reduce energy consumption. As another example, when the first device is a robot, the first device or the control device of the robot can design the movement trajectory according to the remaining power of the first device and other information.
[0083] As an example, the movement trajectory of the first device can include a starting point, an ending point, and a path of the first device from the starting point to the ending point.
[0084] As a possible implementation manner, the movement speed of the first device can also be designed or pre-configured. For example, the maximum movement speed v max of the first device at a certain time can be configured to constrain the first device.
[0085] In some embodiments, the first device can be a device with certain computing capability. As an example, the computing capability of the first device can be used to process the collected data in time to improve timeliness. As an example, the computing capability of the first device can also be used to design or pre-configure the movement trajectory of the first device to improve data collection efficiency.
[0086] As a possible implementation manner, when the first device is a robot, the robot can be an intelligent robot.
[0087] As a possible implementation, the first device can train an AI model, and can also receive data based on the AI model.
[0088] In some embodiments, the first device can also complete tasks other than data collection. For example, complete a given item delivery task.
[0089] In some embodiments, the first device can include a control device to facilitate control of the first device for data collection. That is, the execution subject of the data collection method can be a control device, or include a control device. As an example, the control device can be deployed on the first device to facilitate direct control of the first device. In this scenario, the first device can also be referred to as a control device. As an example, the control device can be deployed on a separate device outside the first device. The separate device can be referred to as a control device. In this scenario, the control device can control the operation of the first device through wireless signals.
[0090] As an example, the control device can be a microcomputer, a processor, a mobile phone, or the like. The embodiments of the present application do not limit the installation location and device type of the control device.
[0091] In some embodiments, the first device can be any kind of movable device for data collection. The data collection performed by the first device can include one or more of the following data collection-related operations: receiving data, processing data, sending data, training data collection-related models, and sending data collection-related information, without limitation.
[0092] In some embodiments, the movable first device needs to consider the remaining energy of the device when performing data collection. Energy can also be referred to as power, energy consumption, etc. When the energy of the first device is low, timely charging is needed to avoid being unable to move due to insufficient energy.
[0093] As an example, when the first device is a robot k, the remaining energy at time t can be represented as e k (t), and e k (t) ∈ [0, E max ]. Where E max is the maximum battery capacity of the robot k.
[0094] As an example, the data collection system can include multiple charging stations or charging piles to facilitate charging of multiple first devices. For example, the data collection system can include C charging stations, where C is a positive integer. The location coordinates of the charging station c (1 ≤ c ≤ C) in the C charging stations can be represented as q charging = (x charging , y charging , H charging ). Hcharging represents the height of the charging station antenna.
[0095] As an example, when the robot k stays at the charging station location, i.e. x charging = x k (t), y charging = y k (t), the robot will charge at a rate of E charging per second, i.e. e k (t+1) = min(e k (t) + E charging , E max ).
[0096] When the first device is a robot, the robot position and energy level can correspond to the mobility and energy saving requirement of the robot. If a robot does not have mobility, i.e. the robot position is constant. If the energy source of the robot is sufficient, the remaining energy can also be assumed to be a constant value, and this state information is strongly related to the system requirement. For this, embodiments of the present application do not make any limitation.
[0097] The second device can be a device with certain communication capability and data collection capability. The data sent by the second device can be the data collected by the second device. The second device can directly send the data to the first device after collecting the data, or can send the data to the first device after converting the data. As an example, the second device can be any sensing device that can directly determine the collected data, such as a sensor.
[0098] Optionally, the second device can be any sensor in a multi-sensor system. The multi-sensor can integrate multiple different types of sensors to collect and monitor various data and information in the environment. These sensors can detect different physical quantities or environmental parameters, such as temperature, humidity, pressure, light intensity, sound, etc. Therefore, the second device can be any of various sensors such as a temperature sensor, a humidity sensor, a pressure sensor, etc., or a sensor that can sense multiple types of data, which is not limited herein.
[0099] In some embodiments, the second device can also have certain computing capability to facilitate preliminary processing of the collected data.
[0100] The position of the second device can be fixed or variable. For example, when the sensor is deployed on a wall, the position of the sensor can remain unchanged to collect environmental parameters. For another example, when the sensor is deployed on a mobile workstation, the position of the sensor is time-varying. In this scenario, the communication channel and communication environment between the sensor and the robot are also time-varying.
[0101] As an example, the position information of the second device at a certain time can also be represented in a three-dimensional Cartesian coordinate system or the like. Taking a sensor m (m is a positive integer) in the system as an example, when the position of the sensor m is fixed, the position coordinates can be represented as q m = (x m ,y m ,H m ); when the position of the sensor m is variable, at time t, the position coordinates of the sensor m can be represented as q m (t) = (c m (t), y m (t), H m ), where H m may represent the height of the antenna of the sensor m.
[0102] Alternatively, the position of the second device can be a relative position with respect to a certain reference, or an absolute position calculated by radar or indoor positioning or the like. Different determination methods will also have different measurement errors. The embodiments of the present application do not limit this.
[0103] In some embodiments, the data acquisition system can include a plurality of second devices. For example, the data acquisition system can deploy M sensors, where M is a positive integer. The M sensors can include sensor m, that is, 1≤m≤M.
[0104] In some embodiments, the actual number of sensors in the data acquisition system is time-varying. As an example, the number of sensors deployed in the data acquisition system can change due to device movement. As an example, in some scenarios, part of the sensors in the data acquisition system are effective, that is, it can not be necessary to receive data of all deployed sensors.
[0105] As an example, at time t, the data acquisition system can deploy M(t) sensors, where M(t) is a positive integer. The set of sensors effective at the current time is
[0106] Alternatively, the plurality of second devices in the data acquisition system can be multiple second devices of different types, or multiple devices of the same type, or multiple devices of different types, which are not limited here.
[0107] In some embodiments, at least one second device corresponding to the first device can be determined according to the indication of the third device, or can be determined by the first device according to its own position information and the distribution information of the second device.
[0108] In some embodiments, the plurality of second devices can be divided into a plurality of device subgroups, respectively corresponding to the plurality of first devices. The division is also referred to as the division or matching of the plurality of first devices and the plurality of second devices. Taking robots and sensors as an example, for a sensor network with a change in position or distribution, a suitable multi-robot and multi-sensor matching strategy (i.e., the division manner of the plurality of sensors) needs to be designed. Based on the actual deployment environment of the sensors of the system, each robot can be responsible for the sensor data collection in a specific area and make decisions at each time unit.
[0109] As an example, a first device subgroup in the plurality of device subgroups corresponds to a first device, that is, the first device subgroup includes at least one second device in step S210. For example, the first device subgroup includes all second devices in the plurality of second devices corresponding to the first device.
[0110] As an example, the division manner of the plurality of device subgroups can be determined according to the position information of the plurality of first devices, so that each first device is responsible for data collection in a specific area and makes decisions at each time unit.
[0111] As an example, the division manner of the plurality of device subgroups can be determined according to the position information of the plurality of first devices and the plurality of second devices.
[0112] As an example, the division manner of the plurality of device subgroups can be determined by the third device based on a first global model, which will be described in detail below.
[0113] In some embodiments, the division manner of the plurality of device subgroups can be fixed or time-varying, which is not limited herein.
[0114] As an example, when the first device is a robot k, at time t, the at least one second device can be M k (t) sensors. Further, the service sensor set of the robot k at time t can be and
[0115] The first device receiving data sent by the at least one second device can include that the first device receives data sent by one or more second devices through a wireless channel. That is, the first device and the second device can communicate wirelessly.
[0116] Optionally, the communication resource for the first device to receive data can be preconfigured or dynamically changed.
[0117] In some embodiments, the channel quality of the communication between the first device and the second device can be determined according to the channel coefficients. The channel coefficients can include large-scale fading and small-scale fading, and the specific values can be measured by a channel estimation method.
[0118] As an example, at time t, the channel coefficient h k,m (t) can be represented as Based on this, the system channel coefficient matrix can be represented as:
[0119] In some embodiments, the data rate of the data transmission between the first device and the second device can be related to the resource size of the communication between the devices, the communication environment, the channel coefficient, etc. Alternatively, in an actual system, the data rate can be obtained by a communication speed test or the like. In this regard, the embodiments of the present application do not limit the specific measurement method.
[0120] In some embodiments, the data sent by the second device, also referred to as data samples, includes but is not limited to pictures, audio, signals, etc. For example, the data sent by the sensor includes but is not limited to pictures, audio, signals, etc. In this regard, the embodiments of the present application do not limit.
[0121] In some embodiments, the number of second devices for which the first device collects data is variable at different times, which is not limited herein.
[0122] The first device receives data sent by at least one second device in a first time unit. In some embodiments, the first time unit can represent each time period in which the first device receives data based on a time discretization method. The first time unit can also be referred to as a time slot, a time.
[0123] Alternatively, the setting of the first time unit needs to take into account multiple aspects such as the frequency of collecting data, energy consumption management, real-time requirements, etc. For example, when the data collection frequency is high or the real-time requirement is high, the length of the first time unit can be relatively short.
[0124] The first time unit can be any time period in the data collection process, such as time unit t. In some embodiments, the first time unit is the length of each time period in which data is collected in a data collection method. The entire cycle of the data collection method can include a data collection phase and a data processing phase. Alternatively, the data collection phase can mainly collect data and can also process data. Alternatively, the data processing phase can mainly process data and can also collect data.
[0125] In some embodiments, the first time unit can be determined based on a time unit of the communication for facilitating the resource allocation. For example, the first time unit can be one or more of: one or more subframes, one or more slots, one or more symbols.
[0126] As an example, the first time unit can be determined according to a period of data collection by the first device. The period can also be a moving time of the first device. For example, the time for robot k to move from a start point to an end point is T k , the first time unit can be any time period within T k .
[0127] The first device receiving data transmitted by the at least one second device according to the first information means that the first device can determine a manner of data collection according to the first information. That is, the first information can also indicate how the first device collects data.
[0128] The first information can include one or more of: a multi-antenna receiving matrix of the first device in the first time unit; movement information of the first device in the first time unit; and an allocation strategy of a first communication resource used by the first device for data collection.
[0129] In some embodiments, the multi-antenna receiving matrix of the first device is used by the first device to determine a beam for data reception. Exemplarily, the first device can select a suitable beam from a plurality of beam manners.
[0130] As an example, the multi-antenna receiving matrix of the first device can be different or the same in different time units. For example, the first device is robot k, and the multi-antenna receiving matrix of the first device in time unit t can be represented as W k (t).
[0131] In some embodiments, the movement information of the first device in the first time unit can include the movement trajectory and the movement speed described above, and can also include other movement parameters related to data collection by the first device.
[0132] As an example, when robot k moves from a start point to an end point, a maximum speed v max may be assumed. T k may represent the number of time units included in the moving time, and thus the first time unit can be any time unit within the time set . In this scenario, robot k can start from the start point at an initial time based on the maximum speed v to collect data from associated sensors and / or complete a given related task.
[0133] In some embodiments, the first communication resource for data collection of the first device can refer to a resource for the first device to communicate with at least one second device or third device. Optionally, the first communication resource can be a time domain resource and / or a frequency domain resource, which is not limited herein. For example, the first communication resource can include a first frequency band.
[0134] As an example, the allocation strategy of the first communication resource can be used to determine a resource for the first device to communicate with each of the at least one second device respectively. That is, the first device can allocate a corresponding communication resource for each of the corresponding plurality of second devices, so as to simultaneously receive data transmitted by the plurality of second devices.
[0135] As an example, the allocation strategy of the first communication resource can be determined based on a plurality of multiple access manners or decoding manners, so as to reduce mutual interference. Optionally, the multiple access manner used to determine the allocation strategy includes, but is not limited to, a frequency division multiple access technology, a time division multiple access technology, a code division multiple access technology, an orthogonal frequency division multiple access technology, or a non-orthogonal multiple access technology, which is not limited herein.
[0136] Optionally, when the first communication resource includes a first frequency band, the allocation strategy of the first communication resource can include one of the following: a strategy of equally dividing the first frequency band based on frequency division orthogonal multiple access; a strategy based on non-orthogonal multiple access and successive interference cancellation.
[0137] For the convenience of understanding, the following takes an example of a robot k receiving data of a sensor m to exemplarily illustrate the two allocation strategies.
[0138] In an embodiment, considering a frequency division orthogonal multiple access technology, a plurality of sensors equally divide a first frequency band B, so there is no mutual interference. When the robot k receives data / signals of the sensor m, the uplink data transmission rate at a first time unit t is:
[0139] where σ is the noise power in the received signal; is a module of a set , i.e., the number of sensors collected by the robot k; is a channel gain, p m is a signal power.
[0140] In another embodiment, considering a power domain non-orthogonal multiple access technology and a successive interference cancellation strategy. In order to alleviate intra-zone interference, an interference successive cancellation technology is used on each robot. Without loss of generality, the effective channel gain of the sensors served by the robot k is sorted as wherein represents any sensor. According to this ordering, it can help the weak signal sensor to eliminate the intra-group interference from the strong signal sensor in the same group. When the robot k receives the data / signal of the sensor m, the uplink data transmission rate at time unit t is:
[0141] wherein, represents the sensor other than the sensor served by the robot k.
[0142] The first device can determine the first information in multiple ways. In some embodiments, the first information can be sent to the first device by the cloud / third device. In some embodiments, the first information can be determined by the first device according to the position information and the environmental information of the second device. In some embodiments, the first device can determine the first information by a first local model, which will be described later.
[0143] Referring back to FIG. 2, at step S220, the first device sends the third device the relevant information for data collection.
[0144] The third device can be used to provide services, implement business and / or provide server functions, etc. in the data collection system. The third device can be any device for implementing cloud functions, and thus the third device can also be referred to as a cloud, a cloud device or a cloud-related device.
[0145] In some embodiments, the third device can be any network device that can provide cloud services as described above, such as a base station.
[0146] In some embodiments, the third device can be any cloud device participating in data collection, which is not limited herein.
[0147] In some embodiments, the third device can send information to the first device and the second device by broadcasting, so as to facilitate data collection. For example, the third device can broadcast the period of data collection to all first devices.
[0148] The relevant information of the first device for data collection can include but is not limited to whether the first device collides when collecting data, whether the energy of the first device when collecting data is higher than a threshold value, and the local data collection rate of the first device. The third device can re-match the first device and the second device according to the relevant information of the multiple first devices respectively collecting data, i.e. determine the division relationship.
[0149] The above describes a data collection method with a movable first device in combination with FIG. 2. The data collection method can be used in a data collection system in various data collection scenarios. As an example, the first device can be used for data collection in the Internet of Things. As an example, the first device can be deployed in a high-temperature, high-pressure, or other harmful environment to the human body.
[0150] As an example, the data collection method can be used in a multi-robot assisted sensor data collection system. In the data collection system, multiple robots can assist a large number of sensors in data collection, and both the design of the multi-antenna receiving end beamforming matrix and the multi-robot path planning problem are considered to maximize the total throughput of the system for long-term operation under the premise that the battery level of the robot is always higher than the warning value.
[0151] It should be understood that the multi-robot assisted sensor data collection method in the embodiments of the present application can be applied to any data collection system in which a robot can be deployed. Further, the data collection system includes at least one robot and at least one sensor.
[0152] When a large number of first devices and second devices are included in the data collection system, the matching of the first devices and the second devices, the multi-antenna receiving matrix of the first device, the path planning, and other issues are related to the application scenario. In particular, under the premise of no prior knowledge, how to perform data collection to meet the needs of data-driven tasks in the Internet of Things or other scenarios is a problem to be solved.
[0153] To solve this problem, the embodiments of the present application also propose a data collection system based on reinforcement learning to improve the adaptability of the algorithm to dynamic unknown environments. The reinforcement learning can refer to multi-agent reinforcement learning. For example, when the first device is an intelligent robot and the third device is deployed with an intelligent computing unit, the data collection system can be a multi-robot data collection system based on multi-agent reinforcement learning.
[0154] In some embodiments, the first device and the third device can both determine a data collection strategy that adapts to the time-varying environment based on an artificial intelligence model. The local model of the first device and the global model of the third device are described below.
[0155] In the data collection method based on reinforcement learning, the first device can perform data collection according to the first local model. Optionally, the first local model is an artificial intelligence model based on reinforcement learning, which can also be referred to as a local deep reinforcement learning network model, a local network model, etc. For example, when the first device has a certain computing power, the first device can perform data collection based on the local network model.
[0156] As an example, the first local model is an artificial intelligence model. As an example, the first local model can be: a Q-learning model, a deep Q-learning model, an actor-critic network model, a deep deterministic policy gradient model, etc. The embodiments of the present application are not limited in this regard.
[0157] Still taking the robot k as an example, at the time unit t, the first local model run by the robot k can include a real-time state-action function (also referred to as a real-time Q function) Q k,t (s k,t ,a k,t ;w k,t ) and a corresponding target Q function wherein s k,t is the local state observed by the robot k, a k,t is the local action, w k,t is the neural network parameter of the real-time Q function Q k,t (s k,t ,a k,t ;w k,t ) and is the neural network parameter of the target Q function .
[0158] In some embodiments, the first information can be determined according to the first local model. That is, the first device can design the multi-antenna receiving matrix, the own trajectory, the communication resource allocation, etc. based on the first local model.
[0159] As an example, at the initial stage of the time unit t, the first device can run the first local model to obtain the first information. Optionally, the first device determines the local state before data collection; and then inputs the local state into the first local model to determine the first information.
[0160] For example, after the robot k collects the local state s k,t of the first time unit t, the robot k can input the state s k,t as input and perform action selection based on the real-time Q function Q k,t (s k,t ,a k,t ;w k,t ): Further, the action selection is also to determine the first information. Specifically, wherein Δq k (t) is the movement of the robot k at the first time unit t, the movement of the robot k has a speed constraint |Δq k (t)|<v max ; W k (t) is the multi-antenna receiving matrix of the robot k.
[0161] Optionally, the local state can comprise one or more of: energy information of the first device; location information of the first device; and location information of the at least one second device.
[0162] As an example, the energy information of the first device can be used to indicate a remaining energy of the first device.
[0163] As an example, the location information of the first device and / or the location information of the at least one second device can be used to determine a movement trajectory.
[0164] Taking a multi-robot assisted data collection system as an example, at each time unit, after the cloud determines the partition relationship between the multi-robot and the multi-sensor, the local end of each robot needs to obtain the current self-position information and the remaining energy, so as to determine the first information. For example, the robot k collects the self-position information q k (t) and the remaining energy e k (t) at the time unit t, and takes them as the current local state s k,t ={q k (t), e k (t)}.
[0165] In some embodiments, the first local model needs to be trained to improve the adaptability to the dynamic environment. Therefore, the data collection process is usually divided into two stages: a training stage and an inference (or test) stage. In the training stage, the first local model learns the strategy and the value function by interacting with the environment, in which stage the system can continuously update the strategy, but there is a certain training time delay and energy consumption. In the inference stage, the first local model has been trained, so that the system decision delay is low, and the decision can be made in actual application, but the strategy will not be updated.
[0166] The length of time for which the first local model is trained and / or inferred can be determined by a first period. The first period can comprise a plurality of time units. The first time unit is any time unit within the first period.
[0167] In some embodiments, within the first period, the first device can train the first local model to update the first local model. Within the first period, the first device can also test the first local model, that is, collect data based on the first local model.
[0168] Illustratively, the first period can comprise a training stage and an inference stage of the first local model. That is, within the first period, the first device can train and test the first local model. Illustratively, the first period can comprise only a period of the training stage or a period of the inference stage. That is, within the first period, the first device only trains or infers the first local model.
[0169] It should be noted that, in the corresponding time of the training stage, the first device can continuously update the first local model based on the training situation. However, in the inference stage, the first local model is not updated.
[0170] Optionally, the data acquisition system can determine the duration of the training stage and / or the inference stage of the first local model by using a pre-defined training round number, designing a training end indicator according to task requirements, or remote control, without limitation. By this method, unnecessary resource waste and performance degradation can be avoided in actual application, while ensuring that the system can normally run and work in different stages.
[0171] As an example, the training stage of the first local model is determined by a pre-defined manner for G cycles, each cycle containing T max time units (T max is a positive integer). When the first device determines that the training stage of the first local model is over, it can enter the inference stage.
[0172] As an example, the first cycle is set to T max time units, and when the running time of the continuous multiple first cycles reaches T max time units, the first local model converges. For example, when all non-faulty robots reach T max time units in continuous multiple cycles, it can be considered that the multiple local models in the system converge, and the multiple local models change from the training stage to the inference stage.
[0173] As an example, whether the running time of the first cycle reaches T max time units can be determined according to the energy of the first device. For example, in the first cycle, the energy of any of the multiple first devices is lower than a first threshold, and the first cycle ends. Taking a multi-robot data acquisition task as an example, as long as the energy consumption of one robot is lower than E min , the current cycle is immediately ended and reset to enter the next cycle.
[0174] As an example, the first threshold can be E min . Taking the cloud and robots as an example, if the energy of a robot is lower than E min , the robot will be in a fault state and no longer have the functions of data acquisition, movement, and parameter updating. Once the cloud ends the current cycle, all robots immediately end the current cycle and reset to enter the next cycle.
[0175] In the training stage of each cycle, the first device can update the first local model based on deep reinforcement learning. For example, the first device can set a suitable reward function and / or return function to facilitate reasonable updating of the first local model.
[0176] In some embodiments, the first device can set a parameter related to the actions of the third device and the first device as a reward function. For example, the first device sets the local data collection rate as the reward function. Taking the cloud and robots as an example, the local data collection rate is related to the actions of the cloud and each robot. In this reward setting, positive rewards are designed to encourage good behavior that improves system throughput, while negative rewards are designed to punish bad behavior that causes performance loss. In order to maximize long-term rewards through immediate rewards, the reward function can take into account the total upload rate of the sensors responsible by the robots and other necessary constraints such as collision avoidance and low energy, which helps the robots to maintain a high total system throughput during movement.
[0177] As an example, the first device can determine a local reward function based on data collection to update the first local model. The local reward function can include three parts: local data collection rate, collision penalty, and downtime penalty.
[0178] Taking robot k as an example, at time unit t, after robot k performs movement and beam selection, robot k can obtain a local reward function r k,t , r k,t ∈[0,1] through interaction with the environment, as follows:
[0179] where κ is a parameter for balancing the data rate and energy consumption of the robot; 1≤m≤M k (t); R k,m (t) is the data upload rate of sensor m to robot k at time unit t, i.e., the local data collection rate; R colli is the collision penalty of robot k, i.e., the collision penalty of robot k with other obstacles or robots; R down is the low energy penalty of robot k, i.e., the downtime penalty.
[0180] Optionally, the parameter κ can be used to balance the importance of the data rate and energy consumption of the robot.
[0181] Optionally, the first part can improve the total data collection rate of the sensors collected by robot k at time unit t.
[0182] Optionally, the second part R colli is the collision penalty. If the distance between robots or between a robot and an obstacle is less than a safe distance, R colli is a negative penalty, and if there is no collision, it is 0. This part is used to prevent the robot from colliding with obstacles and other robots.
[0183] Optionally, the third part R downPenalty for low robot energy. If the robot energy is e k (t)<E min , R down It is a negative penalty if the robot energy is greater than E min , then it is 0. This part is used to prevent the robot from having too low energy.
[0184] In some embodiments, during a first period, the first device may update the first local model based on a small batch of experiences. For example, the first device may store the local experiences corresponding to the first time unit into a local experience pool; then, the first device may sample a small batch of experiences from the local experience pool to update the first local model.
[0185] As an example, the first device updating the first local model may include the first device updating the real-time Q function. Taking robot k as an example, robot k updates the Q function Q k,t (s k,t ,a k,t ;w k,t ) will convert the experience gained per time unit into k,t =(s k,t ,a k,t ,s k,t+1 ,r k,t ) into the local experience pool of robot k, and take small batches of experience to update the real-time Q function Q during the training cycle k,t (s k,t ,a k,t ;w k,t ).
[0186] As an example, the first device may determine a loss function based on the local reward function, and determine the neural network parameters in the real-time Q function based on the loss function. In other words, the loss function may be used to update the real-time Q function.
[0187] As an example, the loss function can be obtained by using the mean square error method. Taking robot k as an example, the loss function can be:
[0188] Among them, y k,t For the goal, μ is the future impact on time unit t.
[0189] As an example, a gradient descent step is performed on the loss function to update the Q function Q k,t (s k,t ,a k,t ;w k,t ) to minimize the loss function. For example, the neural network parameter w at time unit t+1 is k,t+1 for:
[0190] in, represents the gradient operator; α represents the learning rate, which is used to control the update amplitude of the first local model.
[0191] In the data collection method based on reinforcement learning, a third device (or cloud) can participate in data collection through the first global model.
[0192] Optionally, the first global model is an artificial intelligence model based on reinforcement learning, which may also be referred to as a global deep reinforcement learning network model or a cloud-based learning model. As an example, the first global model is an artificial intelligence model. As an example, the first global model may be a model such as Q-learning, deep Q-learning, an actor-critic network, or a deep deterministic policy gradient. This is not limited in the present embodiment.
[0193] For example, at time unit t, the first global model may include the real-time state-action (Q) function and the corresponding target Q function in, is the global state observed by a third device (or cloud), For global actions, is the real-time Q function The neural network parameters, is the target Q function The neural network parameters.
[0194] Taking robots and sensors as an example, at time unit t, the third device collects the position information of all valid sensors and the location information of all robots And use this location information as the global state of the current system.
[0195] In some embodiments, the division relationship between the plurality of first devices and the plurality of second devices can be determined according to the first global model. That is, the third device can match the plurality of first devices with at least one corresponding second device based on the first global model.
[0196] As an example, at the initial stage of time unit t, the third device may determine location information of multiple first devices and multiple second devices. Then, based on the first global model, the third device may determine multiple device subgroups corresponding to the multiple first devices, wherein the multiple device subgroups include a first device subgroup corresponding to the first device, and the first device subgroup includes the at least one second device in step S210.
[0197] As an example, without any prior knowledge, the cloud can train a first global model based on global observation information to design a partitioning relationship between multiple robots and multiple sensors.
[0198] For example, when the cloud collects global state at time unit t, the cloud can take state as input, and perform action selection based on real-time Q function . Further, the action selection can be used to determine the partitioning relationship. Specifically, can represent the partitioning relationship between multiple robots and multiple sensors. Wherein, τ m (t) can represent the robot corresponding to sensor m at time unit t, For example, when τ m (t) = k, it means that sensor m is collecting data for robot k.
[0199] In some embodiments, the first global model needs to be trained to improve its adaptability to dynamic environments. Similarly, the data collection process of the cloud is divided into two stages: training stage and inference (or testing) stage. In the training stage, the first global model learns the policy and value function by interacting with the environment, in which the system can continuously update the policy, but there is a certain training delay and energy consumption. In the inference stage, the first global model has been trained, so the system decision delay is low, and it can make decisions in actual applications, but the policy will no longer be updated.
[0200] The first period can also be used for the third device to train and / or test the first global model, also known as the cloud period, denoted as N c1oud time units. The length of time for the first global model to train and / or infer can also be determined based on the first period.
[0201] In some embodiments, within the first period, the third device can train the first global model to update the first global model. Within the first period, the third device can also test the first global model, that is, participate in data collection based on the first global model.
[0202] Illustratively, the first period can include the training stage and the inference stage of the first global model. That is, within the first period, the third device can train and test the first global model. Illustratively, the first period can only include the training stage or the inference stage. That is, within the first period, the third device only trains or infers the first global model. It should be noted that in the training stage, the third device can continuously update the first global model based on the training. However, in the inference stage, the first global model is not updated.
[0203] Optionally, the data acquisition system can determine the length of the training phase and / or the inference phase of the first global model by using a variety of methods such as a predefined number of training rounds, designing a training end indicator according to task requirements, or remote control, without limitation. By this method, unnecessary waste of resources and performance degradation in actual application can be avoided, while ensuring that the system can normally run and work in different stages.
[0204] As an example, the training of the first global model by the third device or the cloud can be periodic.
[0205] As an example, the training phase of the first global model is determined by a predefined manner for G cycles, each cycle containing T max time units. When the first device determines that the training phase of the first global model ends, it can enter the inference phase.
[0206] As an example, the first cycle is set to T max time units, and when the running time of the continuous multiple first cycles reaches T max time units, the first global model converges. The T max time units can include the first time unit.
[0207] As an example, whether the running time of the first cycle reaches T max time units can be determined according to the energy of the first device. For example, in the first cycle, the energy of any of the plurality of first devices is lower than the first threshold, and the first cycle ends.
[0208] For example, the first threshold can be E min . Taking the cloud and the robot as an example, in a multi-robot data acquisition task, as long as the energy consumption of one robot is lower than E min , the current cycle ends immediately and resets to the next cycle. When the continuous multiple cycles reach T max time units, the cloud first global model converges, records the current optimal robot operation quantity, and changes the first global model from the training mode to the inference phase.
[0209] Optionally, the determination of whether the first device is in the training phase and the determination of whether the third device is in the training phase are independent of each other, and can be determined according to their respective requirements. The embodiments of the present application do not limit this.
[0210] As an example, the time when the first local model enters the inference phase can be different from the time when the first global model enters the inference phase in the first cycle.
[0211] In the training phase of each cycle, the third device can update the first global model based on deep reinforcement learning. For example, the third device can set a suitable reward function and / or return function to facilitate reasonable updates to the first global model.
[0212] In some embodiments, the third device can set the system total communication rate as the reward function. Taking the cloud and robots as an example, in order to maximize the long-term reward through feedback instant return, the cloud reward function can simultaneously consider the system total communication rate and other necessary constraints such as collision avoidance and energy shortage, which helps the cloud to make globally optimal choices, so that the multi-robot assisted data acquisition system has high total throughput.
[0213] As an example, the third device can determine a global return function based on the relevant information of the plurality of first devices respectively performing data acquisition to update the first global model. The global return function can include three parts: system total throughput, collision penalty, and downtime penalty.
[0214] Taking the cloud and robots k as an example, at time unit t, when the cloud determines the matching relationship between the multi-sensor and the robots, and all robots perform their respective movements and beam selection, the cloud interacts with the environment and obtains the global return function as follows:
[0215] wherein, is a parameter for balancing the total throughput and the number of surviving robots; is the collision penalty of all robots, i.e., the collision penalty of all robots with other obstacles or robots; is the energy shortage penalty of all robots, i.e., the downtime penalty.
[0216] Optionally, the parameter can be used to balance the importance ratio of the system total throughput and the number of surviving robots.
[0217] Optionally, the first part can improve the total data acquisition rate of the sensors collected by all robots.
[0218] Optionally, the second part is the collision penalty, which is a negative value if the distance between robots or between a robot and an obstacle is less than a safe distance, R colli , and is 0 if there is no collision. This part is used to prevent robots from colliding with obstacles and other robots.
[0219] Optionally, the third part is the energy shortage penalty of the robots, which is a negative value if the energy of all robots is less than E min , and is 0 if there is no collision. This part is used to prevent robots from colliding with obstacles and other robots. downis a negative penalty, if the robot energy is all greater than E min is 0. This part is used to avoid the robot energy being too low.
[0220] In some embodiments, in the first cycle, the third device can update the first global model based on the mini-batch experience. Illustratively, the third device can store the global experience corresponding to the first time unit into a global experience pool; then sample a mini-batch experience from the global experience pool to update the first global model.
[0221] As an example, the third device updating the first global model can include the third device updating a real-time Q function. Taking the cloud as an example, the cloud can put the experience obtained by each time unit into a global experience pool of the cloud, and take a mini-batch experience in a training cycle to update a real-time Q function
[0222] As an example, the third device can determine a loss function according to the global return function, and determine the neural network parameters in the real-time Q function according to the loss function. That is, the loss function can be used to update the real-time Q function.
[0223] As an example, the loss function can be obtained in the form of mean square error method. For example, the loss function can be:
[0224] wherein, is a target, can represent the influence of the future on the time unit t.
[0225] As an example, a gradient descent step is performed on the loss function to update the Q function to minimize the loss function. For example, the neural network parameters of the time unit t+1 are:
[0226] wherein, denotes a gradient operator; denotes a learning rate, used to control the update amplitude of the first local model.
[0227] The training and inference of the first local model and the first global model in the first period are introduced above. To avoid the limitation of local information, the data collection method in the embodiments of the present application can adopt a federated learning manner to periodically aggregate the local models (also referred to as federation). After the local models of the plurality of first devices are aggregated at the third device, the third device can send the aggregated model to each first device to update the local model of each first device. When applied to a multi-robot assisted sensor data collection system, this method can effectively cover a large number of changing multi-sensors, has good robustness to changes in the number of robots and the size of the sensor network, can fully utilize the flexibility of robots and effectively utilize edge computing power, and is suitable for scenarios where sensors are distributed and time-varying.
[0228] The period of model aggregation is the second period relative to the first period. The second period can refer to the time interval before the local model updates of each first device are aggregated to the central server or the global model after a certain number of local updates (such as model training or parameter update) are completed, and can also be referred to as the aggregation period. The setting of the second period directly affects the performance and efficiency of the federated learning system.
[0229] In some embodiments, the second period can be an integer multiple of the first time unit, for example, N aggregate time units.
[0230] In some embodiments, the setting of the second period can consider one or more of the following factors: communication overhead, computing resources, model convergence speed, privacy and security, and system real-time requirements. The embodiments of the present application do not limit this. Exemplarily, according to specific application scenarios and requirements, different second periods can be used to balance various considerations and ensure the performance and efficiency of the federated learning system.
[0231] In some embodiments, the second period can be decided by the third device first and then broadcast to all first devices by the third device. The embodiments of the present application do not limit this.
[0232] In some embodiments, the third device and / or the first device can determine whether the second period is reached. For example, after the system determines that the second period is reached, the plurality of first devices can upload the models, and the third device can receive the uploaded plurality of models and aggregate the models.
[0233] As an example, the third device can determine whether the current time unit reaches the second period. If the second period is reached, the third device receives the plurality of local models sent by the plurality of first devices; or if the second period is not reached, the third device receives the information related to the data collection of the plurality of first devices.
[0234] As an example, the system can determine whether the second period is reached according to the current time unit and the second period. For example, the current time unit is t, and the second period is T aggre If t mod (T aggre ) = 0, the second period is reached.
[0235] In some embodiments, the first devices and the third device perform model or model parameter interaction in the second period. On the first device side, a plurality of first devices respectively upload a plurality of local models updated in the first period in the second period. The plurality of local models include a first local model updated by the first device in the first period. On the third device side, the third device can receive the plurality of local models in the second period and perform model aggregation on the plurality of local models.
[0236] As an example, the federated model obtained after model aggregation can be used to update the corresponding local model of the plurality of first devices, thereby avoiding the limitation of local information.
[0237] As an example, the related parameters uploaded by the plurality of first devices can also be used to update the first global model by the third device. For example, the model aggregation method of the federation includes aggregating the model parameters of the local training of the robots participating in uploading to the cloud to update the global model.
[0238] As an example, the model aggregation method can be a technique such as federated averaging, weighted federated averaging, federated optimization with penalty, federated stochastic gradient descent, or hierarchical aggregation. The embodiments of the present application do not limit this.
[0239] As an example, the plurality of first devices can perform model uploading in the initial stage of the second period. It should be understood that the plurality of first devices can be all first devices in the data acquisition system, or can be part of the first devices. The embodiments of the present application do not limit this.
[0240] For example, when the first device is a robot, not all robot ends need to perform model uploading, and only robots with similar environments or located in the same environment can perform model uploading.
[0241] As an example, the first device uploading the local model can be uploading the model update gradient, the local training model parameter, the model evaluation index, or other meta information, for updating the global model and evaluating the contribution of the participants. The embodiments of the present application do not limit this.
[0242] In some embodiments, the third device can distribute the aggregated model to the plurality of first devices after aggregating the model. For example, after completing the model aggregation in the cloud, the cloud will distribute the aggregated model to all robot ends participating in the model aggregation.
[0243] As an example, the distributed aggregated model should mainly depend on the uploaded model parameters, which can be gradients, model parameters, model evaluation indicators or other meta information, and is related to the model aggregation method. The embodiments of the present application do not limit this.
[0244] As an example, the third device can distribute the aggregated model to all first devices, or only to part of the first devices. The embodiments of the present application do not limit this. For example, the cloud does not distribute the aggregated model to all robot ends, but to robots with similar environments or located in the same environment.
[0245] In some embodiments, after receiving the aggregated model, the plurality of first devices participating in model aggregation can update the local model.
[0246] As an example, the method of updating the local model depends on the distributed aggregated model and the model aggregation method, which can be a soft update method. The embodiments of the present application do not limit this.
[0247] As an example, it should be that the robots participating in model aggregation update the parameters according to the aggregated model, rather than requiring all robots to update, mainly for robots with similar environments or located in the same environment. The embodiments of the present application do not limit this.
[0248] The following takes an example of a robot k uploading a local model to the cloud. At time unit t, the first local model of robot k includes real-time Q function Q k,t (s k,t ,a k,t ;w k,t ) and target Q function The meanings of the parameters in the Q function are not repeated.
[0249] When the system judges that the second cycle is reached, it is assumed that all robots are located in the same environment, and each robot k uploads the local model parameters w k,t and to the cloud, and other robots upload corresponding neural network parameters.
[0250] When the cloud receives the local model parameters and of all robots participating in the federation, the cloud can perform federated averaging and to obtain the aggregated model parameters and
[0251] After federated averaging is completed in the cloud, the aggregated model parameters and will be distributed to all robots participating in the federation
[0252] When robot k receives the aggregated model parameters and , the local real-time state-action function Q k,t (s k,t ,a k,t ;w k,t ) and the corresponding target Q function are soft-updated as follows:
[0253] where γ is the soft-update amplitude, γ ∈ [0, 1).
[0254] The above describes various method embodiments of training, inference, and model aggregation of local models and global models, respectively. As a possible implementation, for a multi-robot assisted sensor data acquisition system, the actions of the cloud and the multi-robot can be as follows:
[0255] On the one hand, the cloud trains the global model based on global observation information to design the partition relationship of multi-robot and multi-sensor. Specifically, every N cloud time units, the cloud collects the position information of the global effective sensor and the robot position as the global state. The cloud takes the global state as the input of the global neural network, and designs the multi-robot and multi-sensor matching according to the current action selection output by the network. Further, the cloud can also obtain the current system throughput by interacting with the environment as the reward function. Every N cloud time units, the cloud stores the global experience into the experience pool, and updates the global network model parameters from the experience pool.
[0256] On the other hand, the multi-robot trains the local model on the local end based on local observation, and accelerates model convergence and global information acquisition through periodic aggregation. Specifically, every time unit, each robot obtains its own geographic position coordinates and remaining power as the local state, takes the local state of the current time unit as the input of the local neural network, and designs the communication resource allocation and robot movement according to the current action selection output by the network. Further, each robot can obtain the current robot total data acquisition rate by interacting with the environment as the reward function. In each time unit, each robot stores the local experience into the local experience pool, and updates the local network model parameters from the local experience pool.
[0257] Further, in each aggregation cycle, each robot uploads the local network model to the cloud, and the cloud performs a weighted average of all collected local models to obtain an aggregated network model. The cloud distributes the aggregated network model to all robots, and each robot updates the local network model according to the aggregated network model.
[0258] It can be seen that when the data collection method provided by the embodiments of the present application is applied, the multi-robot assisted data collection system based on reinforcement learning can be deployed in an unknown environment without prior knowledge. The data collection system can achieve dynamic and efficient data collection of multiple robots and multiple sensors, and improve the long-term throughput of the data collection system. Further, the data collection system can also effectively cover a large number of changing multiple sensors, and has good robustness to changes in the number of robots and the size of the sensor network, and fully utilizes the flexibility of robots and edge computing power. Therefore, the data collection method can be applied to a data collection system with time-varying sensor distribution.
[0259] The embodiments of the present application will be described in more detail below with reference to specific examples of FIGS. 3 and 4. It should be noted that the example of FIG. 2 is only to help those skilled in the art to understand the embodiments of the present application, and is not intended to limit the embodiments of the present application to the specific values or specific scenarios illustrated. Those skilled in the art can obviously make various equivalent modifications or changes based on the example of FIG. 2, and such modifications or changes also fall within the scope of the embodiments of the present application.
[0260] For ease of understanding, the data collection method in the embodiments of the present application will be exemplarily described below with reference to a multi-robot assisted data collection system based on reinforcement learning, in combination with the flowchart of FIG. 3. FIG. 3 considers the operations of the cloud and the robots respectively.
[0261] Referring to FIG. 3, at step S302, at the beginning of each time unit, the cloud collects the position information of all valid sensors and the robot position as the global state.
[0262] At step S304, the cloud inputs the global state into the first global model to design the division relationship of multiple robots and multiple sensors.
[0263] At step S306, it is determined whether the second cycle is reached. If yes, step S308 is performed, and if no, step S316 is performed.
[0264] At step S308, each robot end uploads the local model to the cloud.
[0265] At step S310, the cloud performs model aggregation based on the received local model.
[0266] At step S312, the cloud distributes the aggregated model to all robot ends.
[0267] At step S314, the robot end updates the local model according to the received aggregated model.
[0268] As can be seen from FIG. 3, steps S308, S310, S312 and S314 are in a serial relationship, and are all model aggregation parts of the robot and the cloud end.
[0269] At step S316, each robot end confirms its own position coordinates and remaining power.
[0270] At step S318, each robot end designs its own trajectory and communication resource allocation based on the local model.
[0271] At step S320, the robot end determines whether it is in a training phase. If yes, step S322 is performed; if no, step S328 is performed.
[0272] At step S322, each robot end obtains a local data collection rate as a reward function, and updates the local model based on an experience mini-batch.
[0273] As can be seen from FIG. 3, steps S316, S318, S320 and S322 are in a serial relationship, and are all decision-making parts of the robot end.
[0274] At step S324, the cloud end determines whether it is in a training phase. If yes, step S326 is performed; if no, step S328 is performed.
[0275] At step S326, the cloud end obtains a total communication rate of the system as a reward function, and updates the first global model based on an experience mini-batch.
[0276] As can be seen from FIG. 3, steps S302, S304, S324 and S326 are in a serial relationship, and are all decision-making parts of the cloud end.
[0277] At step S328, the current time unit ends.
[0278] Embodiments of the present application also propose a data collection system. The data collection system comprises a plurality of first devices, a plurality of second devices and a third device, any first device in the plurality of first devices performs the method for the first device to perform in the method described above, any second device in the plurality of second devices performs the method for the second device to perform in the method described above, and the third device performs the method for the third device to perform in the method described above.
[0279] For ease of understanding, the data collection system in the embodiments of the present application will be exemplarily described below in combination with the structural schematic diagram of FIG. 4.
[0280] FIG. 4 is a structural schematic diagram of a multi-robot assisted data collection system based on reinforcement learning. As shown in FIG. 4, the data collection system can be deployed in a place with a production line and operating equipment. FIG. 4 exemplarily shows 5 sensors, 3 robots, 1 charging station 430, and 1 cloud 440. Among them, the cloud 440 can provide services 441, businesses 442, and server 443 functions for multiple devices in the data collection system. The sensors 411-414 among the 5 sensors are deployed on the wall, and the sensor 415 is deployed on the operating equipment. The 3 robots are respectively the robots 421-423.
[0281] FIG. 4 takes the sensor 423 as an example. The cloud 440 can interact with the robot 423. The sensors served by the robot 423 are the sensors 412, 414, and 415. Based on the data collection method of the embodiments of the present application, the robot 423 can receive the data of the sensors 412, 414, and 415 to perform data collection. The robot 423 can send relevant information of data collection to the cloud 440.
[0282] When the energy of the 3 robots is lower than the first threshold value, the robots can be charged at the charging station 430.
[0283] The above describes the method embodiments of the present application in detail in combination with FIGS. 1-4. The device embodiments of the present application are described in detail below in combination with FIGS. 5-10. It should be understood that the description of the device embodiments corresponds to the description of the method embodiments, and therefore, the parts not described in detail can be referred to the foregoing method embodiments.
[0284] FIG. 5 is a schematic block diagram of a first device according to an embodiment of the present application. The first device 500 can be any movable device described above participating in data collection. The first device 500 shown in FIG. 5 includes a receiving unit 510 and a sending unit 520.
[0285] The receiving unit 510 can be configured to receive data sent by at least one second device at a first time unit according to first information to perform data collection; wherein the first information includes one or more of the following information: a multi-antenna receiving matrix of the first device at the first time unit; movement information of the first device at the first time unit; and an allocation strategy of a first communication resource used by the first device to perform data collection.
[0286] The sending unit 520 can be configured to send relevant information of data collection to a third device.
[0287] Optionally, the first information is determined according to a first local model.
[0288] Optionally, the first device 500 further comprises a first determining unit, configured to determine a local state before data collection; and a second determining unit, configured to input the local state into the first local model to determine the first information.
[0289] Optionally, the local state comprises one or more of the following: energy information of the first device; position information of the first device; and position information of at least one second device.
[0290] Optionally, the first device 500 further comprises a third determining unit, configured to determine a local reward function based on the data collection; wherein the local reward function is used by the first device to update the first local model.
[0291] Optionally, the first device is a robot k of K robots, and the at least one second device is a sensor m of M k (t) of the robot k, the first time unit is t, and the local reward function r k,t is:
[0292] wherein K is a positive integer, 1≤k≤K, t is an integer, κ is a parameter balancing data rate and energy consumption of the robot, M k (t) is a positive integer, 1≤m≤M k (t), R k,m (t) is a data upload rate of the sensor m to the robot k in the time unit t, R colli is a collision penalty of the robot k, R down is an energy deficiency penalty of the robot k.
[0293] Optionally, the first device 500 further comprises a first processing unit, configured to store local experiences corresponding to the first time unit into a local experience pool; and a second processing unit, configured to sample a small batch of experiences from the local experience pool to update the first local model.
[0294] Optionally, the first time unit is any time unit in a first period, and the first period comprises a training phase and an inference phase of the first local model, and the first local model is not updated in the inference phase.
[0295] Optionally, the first device is one of a plurality of first devices, and the sending unit is further configured to upload the first local model updated in the first period in a second period; wherein a plurality of local models uploaded by the plurality of first devices respectively comprise the first local model updated, and the plurality of local models are collectively used by a third device to perform model aggregation.
[0296] Optionally, the first global model is updated in the first period according to information related to data collection of the plurality of first devices respectively.
[0297] Optionally, the energy of any first device in the plurality of first devices is lower than the first threshold value in a first period, and the first period ends.
[0298] Optionally, the first period is set as T max time units, T max is a positive integer, and the T max time units include a first time unit, and the first local model converges when the running time of the continuous plurality of first periods reaches T max time units.
[0299] Optionally, the first local model is an artificial intelligence model based on reinforcement learning.
[0300] Optionally, the allocation strategy is used to determine the resource of the first device for communication with each of the at least one second device, the first communication resource includes a first frequency band, and the allocation strategy includes one of the following: a strategy of equally dividing the first frequency band based on frequency division orthogonal multiple access; a strategy based on non-orthogonal multiple access and serial interference cancellation.
[0301] Optionally, the at least one second device belongs to a first device sub-group, the first device sub-group includes all second devices corresponding to the first device in the plurality of second devices, and the first device sub-group is determined according to the first global model.
[0302] Optionally, the first time unit is one or more of the following: one or more subframes, one or more time slots, and one or more symbols.
[0303] FIG. 6 is a schematic block diagram of a second device according to an embodiment of the present application. The second device 600 can be any sensing device described above that can collect data. The second device 600 shown in FIG. 6 includes a sending unit 610.
[0304] The sending unit 610 can be configured to send data to the first device at the first time unit, wherein the data is data collected by the first device according to the first information, and the first information includes one or more of the following information: a multiple antenna receiving matrix of the first device at the first time unit; movement information of the first device within the first time unit; and an allocation strategy of the first communication resource used by the first device to collect data.
[0305] Optionally, the second device is any device in a first device sub-group, the first device sub-group corresponds to the first device, and the first device sub-group is determined according to the first global model.
[0306] FIG. 7 is a schematic block diagram of a third device according to an embodiment of the present application. The third device 700 can be any cloud device described above that participates in data collection. The third device 700 shown in FIG. 7 includes a receiving unit 710.
[0307] The receiving unit 710 can be configured to receive relevant information for data collection sent by the first device; wherein the relevant information is used to indicate data collection performed by the first device on the at least one second device within a first time unit based on first information, and the first information includes one or more of the following information: a multi-antenna receiving matrix of the first device in the first time unit; movement information of the first device within the first time unit; and an allocation strategy of a first communication resource used by the first device for data collection.
[0308] Optionally, the first device is one of a plurality of first devices, and the at least one second device is one or more of a plurality of second devices, and the third device 700 further includes a first determining unit configured to determine position information of the plurality of first devices and the plurality of second devices; and a second determining unit configured to determine a plurality of device subgroups corresponding to the plurality of first devices respectively according to the first global model, wherein the plurality of device subgroups includes a first device subgroup corresponding to the first device, and the first device subgroup includes the at least one second device.
[0309] Optionally, the third device 700 further includes a third determining unit configured to determine a global return function based on the relevant information of data collection performed by the plurality of first devices respectively; wherein the global return function is used for the third device to update the first global model.
[0310] Optionally, the first device is a robot k of K robots, the at least one second device is M k (t) sensors, the first time unit is t, and the global return function
[0311] wherein K is a positive integer, 1≤k≤K, t is an integer, is a parameter for balancing the total throughput and the number of surviving robots, M k (t) is a positive integer, 1≤m≤M k (t), R k,m (t) is a data upload rate of the sensor m to the robot k within the time unit t, is a collision penalty of all robots, is an energy deficiency penalty of all robots.
[0312] Optionally, the third device 700 further includes a first processing unit configured to store global experience corresponding to the first time unit into a global experience pool; and a second processing unit configured to sample a small batch of experience from the global experience pool to update the first global model.
[0313] Optionally, the first time unit is any time unit in a first period, and the first period includes a training phase and an inference phase of the first global model, and in the inference phase, the first global model is not updated.
[0314] Optionally, the receiving unit 710 is further configured to receive a plurality of local models in a second period; and the third device 700 further includes a third processing unit configured to perform model aggregation on the plurality of local models; wherein the plurality of local models are respectively the local models updated by the plurality of first devices in the first period, and the plurality of local models include the first local model updated by the first device in the first period.
[0315] Optionally, the second period includes a plurality of time units, and the third device 700 further includes a third determining unit configured to determine whether a current time unit reaches the second period; if the current time unit reaches the second period, the receiving unit 710 is further configured to receive the plurality of local models sent by the plurality of first devices; or if the current time unit does not reach the second period, the receiving unit 710 is further configured to receive information related to data collection performed by the plurality of first devices respectively.
[0316] Optionally, in the first period, the energy of any first device in the plurality of first devices is lower than a first threshold, and the first period ends.
[0317] Optionally, the first period is set to T max time units, T max is a positive integer, and the T max time units include the first time unit, and when the running time of the continuous plurality of first periods reaches the T max time units, the first global model converges.
[0318] Optionally, the first global model is an artificial intelligence model based on reinforcement learning.
[0319] Optionally, the first time unit is one or more of the following: one or more subframes, one or more time slots, and one or more symbols.
[0320] The embodiments of the present application also provide a data collection system including a control device. The control device is configured to control a plurality of first devices in the data collection system to perform the method performed by the first device in the embodiments of the present application, and / or the control device is configured to control a third device of the data collection system to perform the method performed by the third device in the embodiments of the present application.
[0321] Optionally, if the control device is deployed in the cloud, any third device in the cloud instructs the first device to collect data according to the control device.
[0322] Optionally, if the control device is deployed at the robot end as the first device, the robot performs data collection based on the indication of the control device.
[0323] FIG. 8 is a structural diagram of a control device in the data collection system. When the first device is a robot, the control device 1000 can be a control device in a multi-robot assisted data collection system based on reinforcement learning. As shown in FIG. 8, the control device 800 can include an information acquisition module 810, a scheme determination module 820, and a resource allocation module 830.
[0324] The information acquisition module 710 can be used to acquire the position information of the active sensors in the system at each moment, the geographical position of each robot, and the total communication rate of the currently served sensors in the multi-robot assisted data collection device, if the information acquisition module 710 is deployed in the cloud. If the information acquisition module 710 is deployed at the robot end, it can be used to acquire the geographical position of each robot and the remaining power in the multi-robot assisted data collection device at each moment.
[0325] The scheme determination module 720 determines the target resource allocation scheme of the current model based on a global deep reinforcement learning method. The target resource allocation scheme can include the matching relationship between multiple sensors and multiple robots, the movement scheme of multiple robots, and the communication resource allocation scheme of multiple robots.
[0326] Optionally, the scheme determination module 720 can include a sensor and robot matching unit, a robot direction control unit, and a robot antenna control unit.
[0327] Exemplarily, the sensor and robot matching unit can be used to make each sensor upload data to a specific robot according to the strategy of global deep reinforcement learning.
[0328] Exemplarily, the robot direction control unit can be used to make each robot move according to the action output by the local deep reinforcement learning.
[0329] Exemplarily, the robot antenna control unit can be used to make each robot deploy a multi-antenna receiving matrix according to the action output by the local deep reinforcement learning.
[0330] The resource allocation module 730 can be used to control the matching of robots and sensors according to the target scheme, control the movement of robots according to the set direction, and control the multi-antenna receiving beam of the robot, so as to maximize the long-term throughput of the data collection system.
[0331] FIG. 9 is a structural schematic diagram of an electronic device provided in an embodiment of the present application. The electronic device is used to implement any step in the data collection method described above. The following takes the multi-robot assisted data collection process based on reinforcement learning as an example for illustration. As shown in FIG. 9, the electronic device structure includes a processor 910, a memory 920, a communication interface 930, and a communication bus 940.
[0332] The processor 910 can be used to execute the program stored in the memory 920, and implement any step in the multi-robot assisted data collection process based on reinforcement learning provided in the embodiments of the present application.
[0333] The memory 920 can be used to store the program related to the multi-robot assisted data collection based on reinforcement learning.
[0334] The communication interface 930 is used for an external entity to modify the program stored in the memory 920. The external entity includes but is not limited to the multi-robot assisted data collection system based on reinforcement learning maintenance personnel, system management device, etc., and the embodiments of the present application do not make specific limitations thereto.
[0335] The communication bus 940 can be used to complete the communication among the processor 910, the memory 920, and the communication interface 930.
[0336] FIG. 10 is a structural schematic diagram of a communication device according to an embodiment of the present application. The dashed line in FIG. 10 indicates that the unit or module is optional. The device 1000 can be used to implement the method described in the above method embodiments. The device 1000 can be a chip, a first device, a second device, or a third device.
[0337] The device 1000 can include one or more processors 1010. The processor 1010 can support the device 1000 to implement the method described in the above method embodiments. The processor 1010 can also be a general-purpose processor or a special-purpose processor, which will not be described again. As described above, the processor 1010 can be a central processing unit (CPU). Alternatively, the processor can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.
[0338] The apparatus 1000 can further include one or more memories 1020. The memories 1020 store a program that can be executed by the processor 1010, so that the processor 1010 performs the methods described in the foregoing method embodiments. The memories 1020 can be independent of the processor 1010 or integrated in the processor 1010.
[0339] The apparatus 1000 can further include a transceiver 1030. The processor 1010 can communicate with other devices or chips through the transceiver 1030. For example, the processor 1010 can perform data transceiving with other devices or chips through the transceiver 1030.
[0340] The embodiments of the present application also provide a computer readable storage medium for storing a program. The computer readable storage medium can be applied in the terminal device or the network device provided by the embodiments of the present application, and the program causes the computer to execute the method performed by the first device, the second device or the third device in the embodiments of the present application.
[0341] The computer readable storage medium can be any available medium or a data storage device such as a server, data center, etc. integrated with one or more available media. The available medium can be a magnetic medium, an optical medium or a semiconductor medium, etc. Examples of computer storage media include, but are not limited to: phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read only memory (ROM), electrically erasable programmable read only memory (EEPROM), flash memory or other memory technology, compact disc-read only memory (CD-ROM), digital video disc (DVD) or other optical storage, magnetic cassette, tape / magnetic disk storage or other magnetic storage devices, or any other non-transmission medium. According to the definition herein, the computer readable medium does not include transitory media such as modulated data signals and carriers.
[0342] Optionally, a computer-readable medium can be used to store information that can be accessed by a computing device. A computer-readable medium can include permanent and non-permanent, movable and non-movable media that can be implemented by any method or technology to store information. Information can be computer-readable instructions, data structures, program modules or other data.
[0343] The embodiments of the present application further provide a readable storage medium, which stores a program or instructions, and the program or instructions are executed by a processor to implement the processes of the above method embodiments or achieve the same technical effects. To avoid repetition, details are not described here.
[0344] The embodiments of the present application further provide a computer program product. The computer program product includes a program. The computer program product can be applied to the terminal device or the network device provided by the embodiments of the present application, and the program causes the computer to execute the method performed by the first device, the second device or the third device in the embodiments of the present application.
[0345] Each of the embodiments in the specification is described in a related manner, and the same or similar parts of each of the embodiments can be referred to each other. Each of the embodiments focuses on the difference from other embodiments. In particular, for the device embodiments, the electronic device embodiments, the computer-readable storage medium embodiments and the computer program product embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the related parts can be referred to the part of the method embodiments.
[0346] In the above embodiments, all or part of the embodiments can be realized by software, hardware, firmware or any combination thereof. When realized by software, all or part of the embodiments can be realized in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center through wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) mode.
[0347] The embodiments of the present application further provide a computer program. The computer program can be applied to the terminal device or the network device provided by the embodiments of the present application, and the computer program enables a computer to execute the method performed by the first device, the second device or the third device in the embodiments of the present application.
[0348] The terms "system" and "network" can be used interchangeably in the present application. In addition, the terms used in the present application are only used to explain the specific embodiments of the present application, and are not intended to limit the present application.
[0349] The relationship terms such as "first", "second" and "third" and the like in the present application are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between the entities or operations.
[0350] It should be understood that the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. Without more limitations, the element defined by the sentence "including a" does not exclude the presence of other identical elements in the process, method, article or device including the element.
[0351] In the embodiments of the present application, the "indication" mentioned can be direct indication, indirect indication, or can be an indication of an associated relationship. For example, A indicates B, which can mean that B can be obtained directly through A; or it can mean that A indirectly indicates B, for example, A indicates C, and B can be obtained through C; or it can mean that A and B have an associated relationship.
[0352] In the embodiments of the present application, the term "corresponding" can mean a direct or indirect corresponding relationship between the two, or can mean an associated relationship between the two, or can mean an indication and being indicated, configuration and being configured, etc.
[0353] In the embodiments of the present application, the "protocol" can refer to a standard protocol in the communication field, which can include, for example, an LTE protocol, an NR protocol and a related protocol applied to a future communication system, and the present application does not limit this.
[0354] In the embodiments of the present application, according to A to determine B does not mean that B is determined only according to A, but B can also be determined according to A and / or other information.
[0355] The term "and / or" in the embodiments of the present application is only used to describe the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can represent the following three cases: A exists alone, A and B exist together, and B exists alone. In addition, the character " / " in this paper generally represents an "or" relationship between the front and rear associated objects.
[0356] In the embodiments of the present application, the size of the serial number of each process does not mean the order of execution, and the execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0357] In several embodiments provided by the present application, it should be understood that the disclosed system, device and method can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There can be another division way in actual implementation. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections between the units can be indirect couplings or communication connections through some interfaces, devices or units, and can be electrical, mechanical or in other forms.
[0358] The units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, that is, they can be located in one place, or can be distributed on a plurality of network units. According to actual needs, some or all of the units can be selected to achieve the purpose of the embodiments of the present application.
[0359] In addition, each functional unit in each embodiment of the present application can be integrated into a processing unit, or each unit can exist physically independently, or two or more units can be integrated into one unit.
[0360] The above embodiments of the present application can be preferred embodiments. Among them, the embodiment serial number is only for description, not representing the advantages and disadvantages of the embodiments.
[0361] Those skilled in the art can clearly understand the method embodiments described above can be realized by means of software and necessary universal hardware platforms, of course, can also be realized by hardware, but in many cases, the former is a better implementation. Based on such understanding, the technical solutions of the present application can be embodied in the form of software products, and the computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), including a plurality of instructions to make a service classification device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) execute the method described in each embodiment of the present application.
[0362] The above merely describes the specific implementation of the present application, but the protection scope of the present application is not limited to this. Any modification, equivalent replacement or improvement within the technical scope disclosed in the present application can be easily thought by any person skilled in the art, which should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A data acquisition method, characterized by, The method comprises: a first device receiving data transmitted by at least one second device according to first information in a first time unit to perform data collection; the first device transmitting relevant information of the data collection to a third device; wherein the first information comprises one or more of the following information: a multi-antenna receiving matrix of the first device in the first time unit; movement information of the first device in the first time unit; and an allocation strategy of a first communication resource of the first device for performing the data collection.
2. The data collection method of claim 1, wherein, The first information is determined according to a first local model.
3. The data collection method of claim 2, wherein, Before the first device receives data transmitted by at least one second device according to first information in a first time unit, the method further comprises: the first device determining a local state before the data collection; the first device inputting the local state into the first local model to determine the first information.
4. The data collection method of claim 3, wherein, The local state comprises one or more of the following: energy information of the first device; position information of the first device; and position information of the at least one second device.
5. The data acquisition method according to any one of claims 2-4, characterized in that, After the first device receives data transmitted by at least one second device according to first information in a first time unit, the method further comprises: the first device determining a local reward function based on the data collection; wherein the local reward function is used by the first device to update the first local model.
6. The data collection method of claim 5, wherein, The first device is a robot k of K robots, the at least one second device is a robot m of M robots k (t) sensors, the first time unit is t, and the local return function r k,t is: wherein K is a positive integer, 1≤k≤K, t is an integer, κ is a parameter balancing the data rate and energy consumption of the robot, M k (t) is a positive integer, 1≤m≤M k (t), R k,m (t) is the data upload rate of sensor m to robot k in time unit t, R colli is the collision penalty for robot k, R down is the energy too low penalty for robot k.
7. The data acquisition method according to any one of claims 2-6, characterized in that, The method further comprises: the first device storing local experience corresponding to the first time unit into a local experience pool; the first device sampling a small batch of experience from the local experience pool to update the first local model.
8. The data acquisition method according to any one of claims 2-7, characterized in that, The first time unit is any time unit in a first period, and the first period comprises a training phase and an inference phase of the first local model, and in the inference phase, the first local model is not updated.
9. The data collection method of claim 8, wherein, The first device is one of a plurality of first devices, and the method further comprises: the first device uploading the first local model updated in a first period to a second period; wherein a plurality of local models uploaded by the plurality of first devices respectively include the first local model updated, and the plurality of local models are collectively used by a third device to perform model aggregation.
10. The data collection method of claim 9, wherein, The first global model is updated according to relevant information of data collection performed by the plurality of first devices respectively in the first period.
11. The data acquisition method according to any one of claims 8-10, characterized in that, In the first period, the energy of any first device of the plurality of first devices is lower than a first threshold, and the first period ends.
12. The data acquisition method according to any one of claims 8-11, characterized in that, The first period is set as T max time units, T max is a positive integer, the T max time units include the first time unit, when the running time of a plurality of consecutive first periods reaches T max time units, the first local model converges.
13. The data acquisition method of any one of claims 2-12, wherein, The first local model is an artificial intelligence model based on reinforcement learning.
14. The data acquisition method of any one of claims 1-13, wherein, The allocation strategy is used to determine resources for communication between the first device and each second device of the at least one second device, and the first communication resource includes a first frequency band, and the allocation strategy comprises one of the following: a strategy of equally dividing the first frequency band based on frequency division orthogonal multiple access; a strategy based on non-orthogonal multiple access and serial interference cancellation.
15. The data acquisition method of any one of claims 1-14, wherein, The at least one second device belongs to a first device sub-group, the first device sub-group includes all second devices corresponding to the first device among the plurality of second devices, and the first device sub-group is determined according to a first global model.
16. The data acquisition method of any one of claims 1-15, wherein, The first time unit is one or more of the following: one or more subframes, one or more time slots, and one or more symbols.
17. A method of data acquisition, characterized by, The method comprises the following steps: The second device sends data to the first device at the first time unit; The data is data collected by the first device according to first information, and the first information includes one or more of the following information: The multi-antenna receiving matrix of the first device at the first time unit; The movement information of the first device within the first time unit; and The allocation strategy of the first communication resource used by the first device to collect the data.
18. The data collection method of claim 17, wherein, The second device is any device in a first device sub-group, the first device sub-group corresponds to the first device, and the first device sub-group is determined according to a first global model.
19. A data acquisition method, characterized by, The method comprises the following steps: The third device receives relevant information sent by the first device for data collection; The relevant information is used to indicate that the first device collects data from at least one second device within a first time unit based on first information, and the first information includes one or more of the following information: The multi-antenna receiving matrix of the first device at the first time unit; The movement trajectory of the first device within the first time unit; and The allocation strategy of the first communication resource used by the first device to collect the data.
20. The data collection method of claim 19, wherein, The first device is one of a plurality of first devices, and the at least one second device is one or more of a plurality of second devices, and the method further comprises the following steps: The third device determines position information of the plurality of first devices and the plurality of second devices; The third device determines a plurality of device sub-groups corresponding to the plurality of first devices respectively according to a first global model, the plurality of device sub-groups include a first device sub-group corresponding to the first device, and the first device sub-group includes the at least one second device.
21. The data collection method of claim 20, wherein, The method further comprises the following steps: The third device determines a global return function based on relevant information of data collection by the plurality of first devices respectively; The global return function is used by the third device to update the first global model.
22. The data collection method of claim 21, wherein, The first device is a robot k of K robots, the at least one second device is M k (t) sensors, the first time unit is t, the global return function is: wherein K is a positive integer, 1 < k < K, and t is an integer, To balance the total throughput and the parameters of surviving robots, M k (t) is a positive integer, 1 < m < M k (t), R k,m (t) is the data upload rate of sensor m to robot k for time unit t, Penalize all robots for collisions, All robots are penalized for low energy.
23. The data acquisition method of any of claims 20-22, wherein, The method further comprises the following steps: The third device stores global experience corresponding to the first time unit into a global experience pool; The third device samples a small batch of experience from the global experience pool to update the first global model.
24. The data acquisition method of any of claims 20-23, wherein, The first time unit is any time unit within a first period, and the first period includes a training phase and an inference phase of the first global model, and in the inference phase, the first global model is not updated.
25. The data acquisition method of claim 24, wherein, The method further comprises the following steps: The third device receives a plurality of local models within a second period; The third device aggregates the plurality of local models; The multiple local models are respectively local models of the multiple first devices after the multiple first devices complete updating in the first period.
26. The data acquisition method of claim 25, wherein, The second period includes multiple time units, and the method further includes: The third device determines whether a current time unit reaches the second period; If the second period is reached, the third device receives the multiple local models sent by the multiple first devices; or If the second period is not reached, the third device receives related information of data collection performed by the multiple first devices respectively.
27. The data acquisition method of any of claims 24-26, wherein, In the first period, energy of any first device in the multiple first devices is lower than a first threshold, and the first period ends.
28. The data acquisition method of any of claims 24-27, wherein, The first period is set as T max time units, T max is a positive integer, the T max time units include the first time unit, when the running time of a plurality of consecutive first periods reaches T max time units, the first global model converges.
29. The data collection method of any of claims 20-28, wherein, The first global model is an artificial intelligence model based on reinforcement learning.
30. The data collection method of any of claims 19-29, wherein, The first time unit is one or more of the following: one or more subframes, one or more time slots, and one or more symbols.
31. A data acquisition device, comprising: The data collection device is a first device, which includes: a receiving unit configured to receive data sent by at least one second device in a first time unit according to first information, so as to perform data collection; a sending unit configured to send, to a third device, related information of the data collection; The first information includes one or more of the following information: a multiple-antenna receiving matrix of the first device in the first time unit; movement information of the first device in the first time unit; and an allocation strategy of a first communication resource of the first device for performing the data collection.
32. The data acquisition device of claim 31, wherein, The first information is determined according to a first local model.
33. The data acquisition device of claim 32, wherein, The first device further includes: a first determining unit configured to determine a local state before the data collection; a second determining unit configured to input the local state into the first local model, so as to determine the first information.
34. The data acquisition device of claim 33, wherein, The local state includes one or more of the following: energy information of the first device; position information of the first device; and position information of the at least one second device.
35. The data acquisition device of any of claims 32-34, wherein, The first device further includes: a third determining unit configured to determine a local return function based on the data collection; The local return function is used by the first device to update the first local model.
36. The data acquisition device of claim 35, wherein, The first device is a robot k of K robots, the at least one second device is a robot m of M robots k (t) sensors, the first time unit is t, and the local return function r k,t is: where K is a positive integer, 1≤k≤K, t is an integer, κ is a parameter balancing the data rate and energy consumption of the robot, M k (t) is a positive integer, 1≤m≤M k (t), R k,m (t) is the data upload rate of sensor m to robot k in time unit t, R colli is the collision penalty for robot k, R down is the energy too low penalty for robot k.
37. The data acquisition device of any of claims 32-36, wherein, The first device further includes: a first processing unit configured to store local experience corresponding to the first time unit into a local experience pool; a second processing unit configured to sample a small batch of experience from the local experience pool, so as to update the first local model.
38. The data acquisition device of any of claims 32-37, wherein, The first time unit is any time unit in a first period, and the first period includes a training phase and an inference phase of the first local model. In the inference phase, the first local model is not updated.
39. The data acquisition device of claim 38, wherein, The first device is one of multiple first devices, and the sending unit is further configured to upload a first local model after the first local model is updated in a first period in a second period; the multiple local models uploaded by the multiple first devices respectively include the first local model after the first local model is updated, and the multiple local models are collectively used by a third device to perform model aggregation.
40. The data acquisition device of claim 39, wherein, The first global model is updated according to the related information of the data collection of the plurality of first devices respectively in the first period.
41. The data acquisition device of any of claims 38-40, wherein, In the first period, the energy of any first device in the plurality of first devices is lower than a first threshold value, and the first period ends.
42. The data acquisition device of any of claims 38-41, wherein, The first period is set as T max time units, T max is a positive integer, the T max time units include the first time unit, when the running time of a plurality of consecutive first periods reaches T max time units, the first local model converges.
43. The data acquisition device of any of claims 32-42, wherein, The first local model is an artificial intelligence model based on reinforcement learning.
44. The data acquisition device of any of claims 31-43, wherein, The allocation strategy is used to determine the resource of the first device for communication with each second device in the at least one second device, the first communication resource includes a first frequency band, and the allocation strategy includes one of the following: A strategy of equally dividing the first frequency band based on frequency division orthogonal multiple access; A strategy based on non-orthogonal multiple access and serial interference cancellation.
45. The data acquisition device of any of claims 31-44, wherein, The at least one second device belongs to a first device subgroup, the first device subgroup includes all second devices corresponding to the first device in the plurality of second devices, and the first device subgroup is determined according to a first global model.
46. The data acquisition device of any of claims 31-45, wherein, The first time unit is one or more of the following: one or more subframes, one or more time slots, and one or more symbols.
47. A data acquisition device, comprising: The data collection device is a second device, and the second device includes: A sending unit configured to send data to a first device at a first time unit; The data is data collected by the first device according to first information, and the first information includes one or more of the following information: A multiple antenna receiving matrix of the first device at the first time unit; Movement information of the first device in the first time unit; and An allocation strategy of a first communication resource used by the first device for data collection.
48. The data acquisition device of claim 47, wherein, The second device is any device in a first device subgroup, the first device subgroup corresponds to the first device, and the first device subgroup is determined according to a first global model.
49. A data acquisition device, comprising: The data collection device is a third device, and the third device includes: A receiving unit configured to receive related information for data collection sent by a first device; The related information is used to indicate data collection of at least one second device by the first device based on first information in a first time unit, and the first information includes one or more of the following information: A multiple antenna receiving matrix of the first device at the first time unit; Movement trajectory of the first device in the first time unit; and An allocation strategy of a first communication resource used by the first device for data collection.
50. The data acquisition device of claim 49, wherein, The first device is one of a plurality of first devices, the at least one second device is one or more of a plurality of second devices, and the third device further includes: A first determination unit configured to determine position information of the plurality of first devices and the plurality of second devices; A second determination unit configured to determine a plurality of device subgroups corresponding to the plurality of first devices respectively according to a first global model, the plurality of device subgroups include a first device subgroup corresponding to the first device, and the first device subgroup includes the at least one second device.
51. The data acquisition device of claim 50, wherein, The third device further includes: A third determination unit configured to determine a global return function based on related information of data collection of the plurality of first devices respectively. The global return function is used by the third device to update the first global model.
52. The data acquisition device of claim 51, wherein, The first device is a robot k of K robots, the at least one second device is a robot m of M robots k (t) sensors, the first time unit is t, the global return function is: wherein K is a positive integer, 1 < k < K, and t is an integer, To balance the total throughput and the parameters of surviving robots, M k (t) is a positive integer, 1≤m≤M k (t), R k,m (t) is the data upload rate of sensor m to robot k for time unit t, Penalize all robots for collisions, All robots are penalized for low energy.
53. The data acquisition device of any of claims 50-52, wherein, The third device further comprises: A first processing unit configured to store global experience corresponding to the first time unit into a global experience pool. A second processing unit configured to sample a small batch of experience from the global experience pool to update the first global model.
54. The data acquisition device of any of claims 50-53, wherein, The first time unit is any time unit within a first period, and the first period comprises a training phase and an inference phase of the first global model, and during the inference phase, the first global model is not updated.
55. The data acquisition device of claim 54, wherein, The receiving unit is further configured to receive a plurality of local models during a second period, and the third device further comprises: A third processing unit configured to aggregate the plurality of local models. The plurality of local models are respectively local models of the plurality of first devices after being updated during the first period, and the plurality of local models comprise the first local model of the first device after being updated during the first period.
56. The data acquisition device of claim 55, wherein, The second period comprises a plurality of time units, and the third device further comprises: A third determining unit configured to determine whether a current time unit reaches the second period. If the second period is reached, the receiving unit is further configured to receive a plurality of local models sent by the plurality of first devices; or If the second period is not reached, the receiving unit is further configured to receive information related to data collection by the plurality of first devices respectively.
57. The data acquisition device of any of claims 51-56, wherein, During the first period, the energy of any first device in the plurality of first devices is lower than a first threshold, and the first period ends.
58. The data acquisition device of any of claims 51-57, wherein, The first period is set to T max , T max is a positive integer, the T max time units include the first time unit, when the running time of a plurality of consecutive first periods reaches T max time units, the first global model converges.
59. The data acquisition device of any of claims 50-58, wherein, The first global model is an artificial intelligence model based on reinforcement learning.
60. The data acquisition device of any of claims 49-59, wherein, The first time unit is one or more of the following: one or more subframes, one or more time slots, and one or more symbols.
61. A data acquisition system, comprising: The control device is configured to control a plurality of first devices in the data collection system to perform the method of any one of claims 1-17, and / or the control device is configured to control a third device in the data collection system to perform the method of any one of claims 19-30.
62. A communications device, characterized by The memory is configured to store a program, and the processor is configured to invoke the program in the memory to perform the method of any one of claims 1-30.
63. An apparatus, comprising: The processor is configured to invoke a program from the memory to perform the method of any one of claims 1-30.
64. A chip, comprising: The processor is configured to invoke a program from the memory to cause a device in which the chip is installed to perform the method of any one of claims 1-30.
65. A computer-readable storage medium, comprising: The program causes a computer to perform the method of any one of claims 1-30.
66. A computer program product, characterised in that, The program causes a computer to perform the method of any one of claims 1-30.
67. A computer program characterised in that, The computer program causes a computer to perform the method of any one of claims 1-30.
Citation Information
Patent Citations
Robot communication control method, system and equipment based on federal reinforcement learning
CN113392539A
Model aggregation method for reinforcement learning based on federated learning and related equipment
CN114692893A
Unmanned aerial vehicle federal learning method fusing clustering and selection optimization
CN117835280A
Multi-rotor unmanned aerial vehicle cluster training neural network model by applying federated learning framework
CN117873171A
Unmanned aerial vehicle assisted edge computing method for random inspection of power grid line
WO2023160012A1