Apparatus and method for semantic communication

By allocating child targets in distributed semantic communication, the child device performs semantic extraction and processing, and the parent device comprehensive verification solves the problem that edge devices cannot resolve global targets, and achieves low-cost and efficient data fusion and target verification.

CN120266129APending Publication Date: 2025-07-04HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380081354.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-01-10
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In distributed semantic communication, edge devices and intermediate nodes cannot effectively parse global targets because they only observe part of the data, resulting in object misdetection or duplication problems, and existing communication methods lead to high communication costs and delays.

Method used

The child target is allocated by the parent device, and the child device performs semantic extraction and processing, only sends decision bits or compressed data, and the parent device performs comprehensive verification, using semantic processing and logical operations to reduce communication overhead, improve robustness and efficiency.

Benefits of technology

Significantly reduces communication costs and latency, improves robustness and flexibility, and achieves more efficient data fusion and target verification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120266129A_ABST
    Figure CN120266129A_ABST
Patent Text Reader

Abstract

The invention relates to semantic communication. The parent device determines one or more child targets based on the target. The parent device allocates a child target to one or more child devices. Each sub-device obtains an input, performs semantic extraction on the input to obtain an intermediate feature, and performs semantic processing on the intermediate feature to validate the assigned sub-target. If the child target is verified, the child device sends a decision flag to its parent device. If the child target cannot be validated, the child device compresses the input using the neural network and provides an activation vector of the neural network output to the parent device. If the decision flag is received, the parent device may directly validate its own target based on the decision flag. If the activation vector is received, the parent device performs semantic extraction and semantic processing on the activation vector to verify its own target.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to the field of communication technologies. For example, the present disclosure relates to devices and methods for semantic communication. Background Art

[0002] An increasing number of applications and services (e.g., robots, autonomous driving, traffic management, and smart factories) rely on artificial intelligence (AI) technologies such as object recognition and computer vision. In these applications and services, multiple distributed sensors collect information about the environment for some complex decision-making in a control center. However, due to the increasing quantity and / or complexity of the sensor data to be transmitted by the sensors and processed by the control center, efficient decision-making has become a very challenging task.

[0003] To process the distributed data obtained by multiple distributed sensors, distributed machine learning models such as neural networks (NNs) can be used. These distributed models require appropriate training, which can be referred to as the learning phase or the training phase. In many cases, learning needs to be performed in a distributed manner. For example, in the case where a part of the relevant data is obtained or measured at multiple distributed local sites. Sometimes, due to limited bandwidth and / or privacy issues, the data / measurements cannot be directly transmitted to a remote central center. Therefore, in the training and inference phases, a part of the potentially associated data needs to be locally processed by each agent / node deployed at each local site. Then, the processed compressed measurements can be transmitted over the network to the remote central center.

[0004] A possible and efficient technical solution for training a distributed machine learning model is "In-Network Learning (INL)", which is disclosed in "In-Network Learning: Distributed Training and Inference in Networks", M. Moldoveanu and A. Zaidi, 2021. INL provides a new distributed learning and inference architecture, in which any number of agents / nodes are involved in both the training phase and the inference phase. All agents / nodes involved in the training phase are also active in the inference phase. The nodes run simultaneously, rather than sequentially. Specifically, in the training phase, each node performs a forward pass on its data using its own NN, possibly also using all the received information from the previous agents / nodes in the network as part of the NN. If a node has no data, it only uses the incoming information as the input to its NN. If a node has no parent node in the network, it only uses its available data as the input. If a node has no available data, it only uses the incoming information as the input to its NN. In these cases, the available / acquired information is vertically concatenated in the input vector before being used as the input to the node's NN.

[0005] Then, the agent / node sends the vector output of the last layer of its NN (referred to as the activation vector) to the next node connected to it in the graph network. The propagation of the forward pass continues until it reaches the end agent / node that needs to make a decision. This node continues the forward pass. Then, it computes a backward pass on its local NN. In the backward step, the output of the first layer of its NN is first vertically split and then sent back to the agent / node of the parent node. Each of these nodes first computes the sum of all the vectors it receives and then continues the backward pass. This process continues until convergence.

[0006] In the application scenarios of distributed semantic communication, multiple edge devices (or nodes), some possible intermediate nodes (such as base stations), and a fusion center (FC) may be involved. Edge nodes include sensors for observing the environment and collecting possible raw data. The raw data can be multimodal (e.g., video, audio, etc.). Therefore, edge nodes can also be referred to as sensing devices. Bidirectional communication can occur between nodes. However, due to communication / privacy restrictions, it is impossible or not allowed for sensing nodes to share raw data with the FC. The FC is used to solve complex objectives (or called global objectives). The global objective can be expressed using some suitable combination languages (e.g., logic-based languages or graph-based languages). Examples of such global objectives can be detecting events that pose a safety risk to pedestrians, determining the root cause of a road accident, counting specific objects present in the observed scene, and so on. Each element for semantic communication may include a machine learning model (e.g., neural network) that needs to be trained and then separately perform inference for the semantic network. Summary of the Invention

[0007] For distributed semantic training and / or inference, the FC facing the global objective not only needs to correctly interpret the observed data but also must perform semantic reasoning using logical rules and some context / background knowledge (BK). Therefore, solving this global objective goes beyond solving simple tasks such as traditional classification / regression tasks and scene graph generation.

[0008] Therefore, finding suitable signaling, encoding, and decoding mechanisms is crucial, which helps the FC efficiently solve its global objective, preferably without directly accessing the raw data collected by sensors.

[0009] There are several challenges. First, edge devices and any intermediate nodes that only observe part of the data can only correctly extract partial semantic information from the raw data that is semantically meaningful for the global objective of the FC. In addition, the extracted semantic information needs to be appropriately processed using context / background knowledge. However, each node may only have access to incomplete context / background knowledge. Moreover, some semantic facts of the global objective can only be deduced by jointly considering data from multiple sensors, which means that in this case, no edge device can solve the global objective alone. These lead to typical problems such as object misdetection or duplication that occur when each edge device only observes a part of the entire environment.

[0010] In some traditional distributed inference / training methods, edge devices may send the activation values output by their machine learning models. However, the magnitudes of these activation values are still relatively large, and relatively high communication costs are still incurred. As a result, scarce communication resources may be wasted and unnecessary latency may be introduced.

[0011] In view of the above disadvantages and problems, the present disclosure aims to provide a technical solution for distributed semantic processing and communication. Other objectives may be to improve the performance of learning / inference for performing distributed semantic processing and communication (e.g., lower latency, robustness, flexibility, and energy consumption). These objectives and other objectives are achieved by the descriptions in the independent claims, etc. Advantageous implementations are further described in the dependent claims.

[0012] A first aspect of the present disclosure provides a parent device. The parent device is configured to:

[0013] - obtain a target, where the target includes one or more computing tasks;

[0014] - determine one or more sub-targets based on the target, where each sub-target is at least a subset of the target;

[0015] - send the one or more sub-targets to one or more sub-devices;

[0016] - in response to receiving a decision for a corresponding sub-target from a corresponding sub-device, perform semantic processing on the corresponding decision to verify the target;

[0017] - in response to not receiving a decision for the corresponding sub-target from the corresponding sub-device, obtain processed data from the corresponding sub-device, perform semantic extraction on the processed data to obtain one or more intermediate features of the processed data, and perform semantic processing on the one or more intermediate features to verify the target.

[0018] Optionally, for performing semantic processing, when only one intermediate feature is received, the parent device may be configured to directly verify the target based on the corresponding intermediate feature. When two or more intermediate features are received, the parent device may be configured to combine two or more corresponding semantic features and verify the target based on the combination. Optionally, logical operations (such as "AND", "OR", etc.) may be used to combine two or more corresponding intermediate features. Optionally, the background knowledge of the parent device may be used to verify the target.

[0019] Optionally, corresponding intermediate features may be used to represent characteristics associated with objects or events related to the target.

[0020] Optionally, the processed data may be an activation vector (or activation value). The activation vector may be the output of the last layer of the neural network of the corresponding sub-device.

[0021] By processing the input data at the semantic level, the sub-device can send only the information relevant to or useful for its sub-goal. Therefore, the data rate can be significantly reduced. In addition, performance improvements such as lower latency, stronger robustness, flexibility, and lower energy consumption can also be achieved.

[0022] In one implementation of the first aspect, the decision may include an indication bit indicating that the sub-goal is verified on the corresponding sub-device.

[0023] Optionally, the indication bit may be a 1-bit value or a flag. The indication bit can be used by the sub-device to notify the parent device that the sub-goal has been successfully verified. That is, the corresponding computing task is successfully completed on the corresponding sub-device.

[0024] In this way, by sending only one bit instead of the original data, the communication overhead can be significantly reduced.

[0025] In another implementation of the first aspect, the decision may further include the confidence of the decision.

[0026] Optionally, the confidence may be a numerical value, such as a percentage value. Alternatively, the confidence can be represented by various levels, such as high confidence level, medium confidence level, low confidence level. These different levels can be represented by different values, such as bit values.

[0027] In this way, when the parent device receives multiple decisions from multiple sub-devices, the parent device can use the confidence of each decision when combining the multiple decisions through semantic processing. For example, corresponding weights can be assigned to each decision according to its confidence.

[0028] In another implementation of the first aspect, the decision may further include one or more extracted features.

[0029] Optionally, the extracted features may be intermediate features extracted through semantic extraction. Alternatively or additionally, the decision may include additional information, such as the geographical location of the corresponding sub-device, and / or a timestamp. For example, the timestamp can be used to indicate the time point when the input data is captured by the sub-device or the time point when the decision is made.

[0030] By additionally providing one or more of the extracted features, the additional information, and / or a portion of the processed data, the parent device can use the additional useful information to verify the target. In this way, the accuracy of verifying the target can be improved.

[0031] In another implementation of the first aspect, for performing semantic extraction on the processed data, the parent device can be used to detect one or more objects. Optionally, the parent device can also be used to detect one or more attributes of the detected object from the processed data.

[0032] Optionally, the one or more objects and their associated optional attributes can be represented using a scene graph. However, it should be noted that the scene graph is optional and not necessary, and can be selectively generated according to the application scenario.

[0033] In this way, the accuracy of verifying the target can also be improved.

[0034] In another implementation of the first aspect, the parent device can include a first neural network model (or simply referred to as a neural network (NN)) for performing semantic extraction.

[0035] In another implementation of the first aspect, for performing semantic processing, the parent device can be used to verify semantic facts related to the target.

[0036] Optionally, the semantic facts can be determined by the parent device based on the one or more intermediate features obtained through semantic extraction.

[0037] In another implementation of the first aspect, in response to receiving multiple decisions from multiple child devices, the parent device can be used to combine the multiple decisions and perform semantic processing on the combined decisions.

[0038] Optionally, two or more of the multiple child devices can be assigned the same sub-goal. In this way, by combining multiple decisions from two or more child devices, due to the wisdom of the crowd, the accuracy of verifying the target can be further improved. Or, different sub-goals can be assigned to the two or more of the multiple child devices. In this way, the accuracy of verifying the target can also be further improved due to inputs from various angles.

[0039] In another implementation of the first aspect, in response to obtaining two or more pieces of processed data from two or more child devices, the parent device can be used to concatenate, sum, or splice the two or more pieces of processed data.

[0040] Optionally, the joining, summing, or splicing may be performed before semantic extraction to obtain the data after joining, summing, or splicing. The parent device may be used to perform semantic extraction based on the data after joining, summing, or splicing.

[0041] In this way, by combining different processed data from different sub-devices, the accuracy of verifying the target can be further improved.

[0042] In another implementation manner of the first aspect, in response to determining that the target is not verified, the parent device may further be used to:

[0043] - Determine one or more updated sub-goals based on the target;

[0044] - Send the one or more updated sub-goals to the one or more sub-devices.

[0045] In this way, the one or more sub-goals can be dynamically updated according to the performance of the verification result of the target. Therefore, the overall performance of semantic communication can be improved.

[0046] In another implementation manner of the first aspect, the parent device may be used to communicate with the one or more sub-devices in one or more predetermined time slots.

[0047] Optionally, the parent device may be used to communicate with the one or more sub-devices in a synchronous manner.

[0048] In this way, synchronization between all devices can be achieved, improving communication efficiency.

[0049] In another implementation manner of the first aspect, the parent device may further be used to select the one or more sub-devices based on side information. The side information may include one or more of the following:

[0050] - The usefulness of the input data obtained by each sub-device;

[0051] - The processing capacity of each sub-device;

[0052] - The acquired knowledge of each sub-device;

[0053] - The acquired knowledge of the parent device.

[0054] In this way, the assignment of sub-goals can be more targeted based on the side information. Therefore, since one or more suitable sub-goals can be assigned to each device, the communication efficiency can be improved.

[0055] A second aspect of the present disclosure provides a sub-device. The sub-device is used for:

[0056] Obtain a sub-goal from a parent device, where the sub-goal is at least a subset of the goals of the parent device, and the goals include one or more computing tasks;

[0057] Obtain input data, and perform semantic extraction on the input data to obtain one or more intermediate features of the input data;

[0058] Perform semantic processing on the one or more intermediate features to verify the sub-goal.

[0059] In response to determining that the sub-goal is verified, send a decision on the sub-goal to the parent device;

[0060] In response to determining that the sub-goal is not verified, compress the input data to obtain processed data, and send the processed data to the parent device.

[0061] Optionally, the sub-device can be used to receive two or more sub-goals from the parent device.

[0062] Optionally, for performing semantic processing, when only one intermediate feature is received, the sub-device can be used to directly verify the sub-goal based on the corresponding intermediate feature. When two or more intermediate features are received, the sub-device can be used to combine the two or more corresponding semantic features and verify the sub-goal based on the combination. Optionally, logical operations (such as "AND", "OR", etc.) can be used to combine two or more corresponding intermediate features. Optionally, the background knowledge of the sub-device can be used to verify the sub-goal.

[0063] Optionally, the corresponding intermediate features can be used to represent the characteristics associated with the objects or events related to the sub-goal.

[0064] By processing the input data at the semantic level, the sub-device can only send information related to or useful for its sub-goal. Therefore, the data rate can be significantly reduced. In addition, performance improvements such as lower latency, stronger robustness, flexibility, and lower energy consumption can also be achieved.

[0065] In one implementation of the second aspect, the decision can include an indication bit indicating that the sub-goal is verified.

[0066] Optionally, the indication bit can be a 1-bit value or a flag. The indication bit can be used by the sub-device to notify the parent device that the sub-goal is successfully verified. That is, the corresponding computing task is successfully completed on the corresponding sub-device.

[0067] In this way, by sending only one bit instead of the original data, the communication overhead can be significantly reduced.

[0068] In another implementation of the second aspect, the decision may further include the confidence level of the decision.

[0069] Optionally, the confidence level may be a numerical value, such as a percentage value. Alternatively, the confidence level may be represented by various levels, such as a high confidence level, a medium confidence level, and a low confidence level. These different levels may be represented by different values, such as bit values.

[0070] Optionally, the decision may further include one or more extracted features. The extracted features may be intermediate features extracted through semantic extraction. Alternatively or additionally, the decision may include additional information, such as the geographical location of the sub-device, and / or a timestamp. For example, the timestamp may be used to indicate the time point when the input data is captured by the sub-device or the time point when the decision is made.

[0071] In another implementation of the second aspect, for performing semantic extraction on the input data, the sub-device may be used to detect one or more objects. Optionally, the sub-device may also be used to detect one or more attributes of the detected object from the input data.

[0072] In another implementation of the second aspect, the sub-device may include a second neural network model for performing semantic extraction.

[0073] In another implementation of the second aspect, for performing semantic processing, the sub-device may be used to verify semantic facts related to the sub-goal.

[0074] Optionally, the semantic facts may be determined based on one or more intermediate features obtained by the sub-device through semantic extraction.

[0075] In another implementation of the second aspect, for compressing the input data, the sub-device may include a third neural network model for inferring the input data to obtain a feature map of the input data as the compressed data.

[0076] Optionally, the feature map of the input data may be an activation vector (or referred to as an activation value). The activation vector may be the output of the last layer of the third neural network.

[0077] In another implementation of the second aspect, the sub-device may be used to communicate with the parent device in one or more predetermined time slots.

[0078] A third aspect of the present disclosure provides a system, including one or more parent devices according to the first aspect or any of its implementations, and one or more sub-devices according to the second aspect or any of its implementations.

[0079] The fourth aspect of the present disclosure provides a method, including the following steps:

[0080] - The parent device obtains a target, where the target includes one or more computing tasks;

[0081] - The parent device determines one or more sub-goals based on the target, where each sub-goal is at least a subset of the target;

[0082] - The parent device sends the one or more sub-goals to one or more sub-devices;

[0083] - In response to receiving a decision on the corresponding sub-goal from the corresponding sub-device: The parent device performs semantic processing on the decision to verify the target;

[0084] - In response to not receiving a decision on the sub-goal from the corresponding sub-device: The parent device obtains the processed data from the corresponding sub-device; performs semantic extraction on the processed data to obtain one or more intermediate features of the processed data; and performs semantic processing on the one or more intermediate features to verify the target.

[0085] In one implementation of the fourth aspect, the decision may include an indication bit indicating that the sub-goal is verified on the corresponding sub-device.

[0086] In another implementation of the fourth aspect, the decision may further include the confidence level of the decision.

[0087] In another implementation of the fourth aspect, the decision may further include one or more extracted features.

[0088] In another implementation of the fourth aspect, the step of performing semantic extraction on the processed data may include the parent device detecting one or more objects. Optionally, the parent device may detect one or more attributes of the detected object from the processed data.

[0089] In another implementation of the fourth aspect, the parent device may include a first neural network model for performing semantic extraction.

[0090] In another implementation of the fourth aspect, the step of performing semantic processing may include: The parent device verifies semantic facts related to the target.

[0091] In another implementation of the fourth aspect, in response to receiving multiple decisions from multiple sub-devices, the method may include: The parent device combines the multiple decisions; and the parent device performs semantic processing on the combined decisions.

[0092] In another implementation of the fourth aspect, in response to obtaining two or more pieces of processed data from two or more sub-devices, the step of performing semantic processing on the one or more intermediate features may include: the parent device joins, sums, or splices the two or more pieces of processed data.

[0093] In another implementation of the fourth aspect, in response to determining that the target is not verified, the method may further include the following steps:

[0094] - The parent device determines one or more updated sub-goals based on the target;

[0095] - The parent device sends the one or more updated sub-goals to the one or more sub-devices.

[0096] In another implementation of the fourth aspect, the method may include: the parent device communicates with the one or more sub-devices in one or more predetermined time slots.

[0097] In another implementation of the fourth aspect, the method may include: the parent device selects the one or more sub-devices based on side information. The side information may include one or more of the following:

[0098] - The usefulness of the input data obtained by each sub-device;

[0099] - The processing capacity of each sub-device;

[0100] - The acquired knowledge of each sub-device;

[0101] - The acquired knowledge of the parent device.

[0102] The method of the fourth aspect may share the same optional features with the parent device of the first aspect and has the same technical advantages and benefits.

[0103] The fifth aspect of the present disclosure provides a method, including the following steps:

[0104] - The sub-device obtains a sub-goal from the parent device, where the sub-goal is at least a subset of the goal of the parent device, where

[0105] the goal includes one or more computing tasks;

[0106] - The sub-device obtains input data and performs semantic extraction on the input data to obtain one or more intermediate features of the input data;

[0107] - The sub-device performs semantic processing on the one or more intermediate features to verify the sub-goal;

[0108] - In response to determining that the sub-goal is verified: the sub-device sends the decision of the sub-goal to the parent device;

[0109] - In response to determining that the sub-goal is not verified, the sub-device compresses the input data to obtain processed data, and sends the processed data to the parent device.

[0110] In one implementation of the fifth aspect, the decision may include an indication bit indicating that the sub-goal is verified.

[0111] In another implementation of the fifth aspect, the decision may further include the confidence level of the decision.

[0112] In another implementation of the fifth aspect, the step of performing semantic extraction on the input data may include: the sub-device detecting one or more objects. Optionally, the sub-device may detect one or more attributes of the detected object from the input data.

[0113] In another implementation of the fifth aspect, the sub-device may include a second neural network model for performing semantic extraction.

[0114] In another implementation of the fifth aspect, the step of performing semantic processing may include: the sub-device verifying semantic facts related to the sub-goal.

[0115] In another implementation of the fifth aspect, for compressing the input data, the sub-device may include a third neural network model for inferring the input data to obtain a feature map of the input data as the compressed data.

[0116] In another implementation of the fifth aspect, the method may include: the sub-device communicating with the parent device in one or more predetermined time slots.

[0117] The method of the fifth aspect may share the same optional features with the sub-device of the second aspect, and has the same technical advantages and benefits.

[0118] The sixth aspect of the present disclosure provides a computer program, which includes program code for executing the method according to the fourth aspect or any implementation of the fourth aspect or according to the fifth aspect or any implementation of the fifth aspect.

[0119] The seventh aspect of the present disclosure provides a non-transitory storage medium storing executable program code, which, when executed by a processor, executes the method according to the fourth aspect or any of its implementations or according to the fifth aspect or any of its implementations.

[0120] The eighth aspect of the present disclosure provides a chipset, which includes a memory and a processor. The memory and the processor are used to store and execute program codes to perform the method according to the fourth aspect or any of its implementation manners or according to the fifth aspect or any of its implementation manners.

[0121] It should be noted that all devices, elements, units, and modules described in this application can be implemented in software or hardware elements or any combination thereof. The steps performed by various entities described in this application and the functions to be performed by various entities described are intended to mean that each entity is used to perform each step and function. Even in the following description, if the specific functions or steps to be performed by an external entity are not reflected in the description of the specific detailed elements of the entity performing the specific step or function, those skilled in the art should understand that these methods and functions can be implemented in the corresponding software or hardware elements or any combination thereof. Description of the Drawings

[0122] In combination with the accompanying drawings, the descriptions of the following specific embodiments will elaborate on the above aspects and their implementation manners, where:

[0123] Figure 1 Shows the parent device and the child device provided by the present disclosure;

[0124] Figure 2 Shows an example of the target transformation provided by the present disclosure;

[0125] Figure 3 Shows the system provided by the present disclosure;

[0126] Figure 4 Shows the application scenario provided by the present disclosure;

[0127] Figure 5 Is an example for detecting a vehicle with potential safety hazards and the corresponding signaling between nodes;

[0128] Figure 6 Shows the synchronization mechanism provided by the present disclosure;

[0129] Figure 7 Shows the distributed joint training of the neural network provided by the present disclosure;

[0130] Figure 8 Shows the method provided by the present disclosure for constructing a set of potential targets that each node can solve;

[0131] Figure 9 Shows the method provided by the present disclosure;

[0132] Figure 10 Shows another method 1000 provided by the present disclosure;

[0133] Figure 11 Shows the effect of intermittent communication provided by the present disclosure. Detailed implementation

[0134] The present disclosure provides a technical solution for distributed semantic processing and communication. Compared with the situation where the data collected by the sensing nodes is directly sent to the fusion center, this technical solution greatly reduces the communication cost and enhances the privacy. Compared with the general distributed training / learning technology, this technical solution also provides significantly improved performance (e.g., lower latency, robustness, flexibility, and energy consumption), while supporting efficient data fusion from multiple sensors.

[0135] In Figures 1 to 11 it, the corresponding elements may share the same features and may function in the same way.

[0136] Figure 1 Shows the parent device 110 and the child device 120 provided by the present disclosure. The parent device 110 and the child device 120 are communication nodes in the semantic communication network. The child device 120 and the parent device 110 can be connected in the network through wired or wireless connections. In the present disclosure, the terms "device" and "node" can be used interchangeably.

[0137] The parent device 110 may include a semantic extraction (SE) component 111 for performing semantic extraction and a semantic processing (SP) component 112 for performing semantic processing. When the parent device is an intermediate node, the parent device may include an NN processing component (or referred to as the fourth NN) 113. When the parent device is an end server such as an FC, the parent device does not include the NN processing component 113.

[0138] The child device 120 may include its own SE component 121 for performing semantic extraction, an SP component 122 for performing semantic processing, and an NN processing component (or referred to as the third NN) 123.

[0139] Optionally, the connection may be cascaded. That is, when the parent device 110 is an intermediate node, the parent device 110 may be connected to its own parent device (or referred to as "other parent device", Figure 1 not shown in ). In this case, the parent device 110 facing its own parent device can be used to act as a child device similar to the child device 120. Therefore, in Figure 1 it, the elements 111, 112, and 113 may correspond to the elements 121, 122, and 123, and thus, they may share the same features accordingly under appropriate circumstances.

[0140] The parent device 110 is used to obtain a target 101. The target 101 can be sent from other devices at a higher level (e.g., other parent devices of the parent device 110). The target 101 includes one or more computing tasks. Each computing task can be verified.

[0141] The parent device 110 is used to determine one or more sub-goals based on the target. Each sub-goal can be at least one subset of the target 101. That is, each sub-goal can include at least one computing task among the one or more computing tasks included in the target 101. The parent device 110 is used to send one or more sub-goals to one or more sub-devices. There is no need to have a one-to-one correspondence between the one or more sub-goals and the one or more sub-devices. That is, the sub-devices can be used to receive one or more sub-goals from the parent device. Two sub-devices can be used to receive the same target. For example, the first sub-device can receive a sub-goal #1, the second sub-device can receive two sub-goals #2 and #3, and the third sub-device can receive the sub-goal #1. In Figure 1 For simplicity, one sub-device 120 that receives the sub-goal 102 is shown.

[0142] It should be noted that each target (or each sub-goal) of the corresponding device can be called a local target (or a local sub-goal). For example, the parent device 110 can have its local target 101, and the sub-device 120 can have its local sub-goal 102.

[0143] To verify the received sub-goal 102, the sub-device 120 is used to obtain input data 103 and perform semantic extraction on the input data 103 to obtain one or more intermediate features 104 of the input data 103. The sub-device 120 is also used to perform semantic processing on the one or more intermediate features 104 to verify the sub-goal.

[0144] Optionally, each intermediate feature can be used to represent the characteristics of a described object or event. For example, the symbol of a "car" may have intermediate features such as "moving object", "wheel", "windshield", "license plate", "carrying people", "various colors", etc. During semantic processing, one or more such semantic features can be combined to verify whether it is a car.

[0145] For performing semantic processing, when only one intermediate feature is received, the parent device can be used to directly verify the target based on the corresponding semantic feature. When two or more intermediate features are received, the parent device can be used to combine two or more corresponding semantic features and verify the target based on the combination.

[0146] When the sub-goal 102 is verified by the sub-device 120 (or can be verified with a high confidence), the sub-device 120 is used to send a decision 105 of the sub-goal 102 to the parent device 110. The decision 105 may include an indication bit (e.g., a 1-bit value) or a flag indicating that the sub-goal is verified. The decision may also include the confidence of the result. The confidence may be a percentage value indicating the confidence level of the decision (or result). Optionally, the decision may also include one or more extracted features of the input data. The decision may be included in a response sent from the sub-device 120 to the parent device 110.

[0147] For example, consider a sub-goal of "there is a car". The sub-device 120 may include a camera that takes a video / photo as input. The optional SE component 121 may be used to perform semantic extraction on the video / photo to detect the presence of any characteristics associated with a car (i.e., extract any semantic features of the car). For example, if the SE component 121 only extracts a semantic feature of "moving object" and provides it to the SP component 122, the SP component 122 cannot verify whether it is a car, or at least cannot verify it with a high confidence. In this case, the video / photo is input to the NN processing component 123, and one or more intermediate features 106 are provided to the parent device 110. If two semantic features of "windshield" and "license plate" are extracted, the SP component 122 can verify that a car is detected, or at least can verify that a car is detected with a high confidence. In this case, the sub-device 120 signals to the parent device 110 a flag with a 1-bit value such as "1". If no intermediate features are extracted at all, or if there is no valid input for performing semantic extraction: then the sub-device 120 may not send anything, or may send a synchronization signal to the parent device 110, such as "SYN" predefined using a certain bit pattern.

[0148] Optionally, each sub-goal may be associated with an identity (ID). The decision and each intermediate feature may be associated with the sub-goal ID and / or the ID of the sub-device 120.

[0149] After receiving the decision 102 from the sub-device 120, the parent device 110 is used to perform semantic processing on the decision to verify the goal 101.

[0150] When the sub-goal 102 is not verified by the sub-device 120 (or cannot be verified with high confidence), the sub-device 120 is used to compress the input data 103 to obtain the processed data 106 and send the processed data 106 to the parent device. The compression of the input data 103 can be regarded as preprocessing of the input data 103 locally captured by the sub-device 120. The third NN 123 can be used to extract the feature map of the input data 103 as the processed data 106. The feature map can be the activation vector (or activation value) of the last layer of the third NN 123. The processed data 106 can be included in the response sent by the sub-device 120 to the parent device 110.

[0151] Since the sub-goal 102 is not verified, the sub-device 120 is used not to send any decision to the parent device 110. When the parent device 110 cannot receive the decision of the sub-goal 102 from the sub-device 120, the parent device 110 is used to obtain the processed data 106 from the sub-device 120. The parent device 110 is also used to perform semantic extraction on the processed data 106 to obtain one or more intermediate features 107 of the processed data, and perform semantic processing on the one or more intermediate features 107 to verify the goal 101. This is similar to the sub-device 120 verifying its sub-goal 102. For example, if the goal is verified, the decision 108 of the goal 101 is output by the parent device 110. If the goal 110 is not verified, the processed data compressed by the sub-device 120 is further compressed by the parent device 110 to obtain other processed data 109 that can be provided to other parent devices.

[0152] For performing semantic extraction on the processed data, the parent device 110 can be used to detect one or more objects (or referred to as "symbolic facts") from the processed data 106. Optionally, the parent device 110 can be used to determine one or more attributes of the detected objects. To this end, the parent device can include an NN for performing semantic extraction (referred to as the "first NN" here). The first NN can be hosted in the SE component 111 of the parent device 110. Similarly, the second NN can be hosted in the SE component 121 of the sub-device 120.

[0153] For performing semantic processing on the decision, the parent device 110 can be used to verify the decision. Optionally, when the parent device 110 receives multiple decisions from multiple sub-devices, the parent device 110 can be used to combine the multiple decisions and perform semantic processing on the combined decisions.

[0154] Optionally and additionally, one or more other intermediate features can be provided by other sub-devices ( Figure 1 not shown in the figure). In this case, the one or more intermediate features and the one or more other intermediate features can be combined to perform semantic processing.

[0155] In the case where the target cannot be verified, the parent device may need to rethink target allocation. To this end, the parent device 110 can update the sub-goal 102 by determining one or more other (or updated) sub-goals based on the target, and send one or more updated sub-goals to one or more sub-devices. For Figure 1 the sub-device 120 in

[0156] its sub-goal 102 is updated.

[0157] Optionally, the parent device 110 can also be used to communicate with one or more sub-devices simultaneously in one or more predetermined time slots.

[0158] - The usefulness of the input data obtained by each sub-device;

[0159] - The processing capacity of each sub-device;

[0160] - The acquired knowledge of each sub-device;

[0161] - The acquired knowledge of the parent device.

[0162] Figure 2 Shows an example of target transformation provided by the present disclosure. The task of target transformation is to transform a target into one or more sub-goals. Figure 2 Shows based on Figure 1 the sub-device 120 and the parent device 110 constructed system. The system includes FC207, two intermediate nodes 205, 206 and four edge nodes 201 to 204. When FC 207 is regarded as the parent device, the intermediate nodes 205, 206 can be regarded as sub-devices. When the intermediate node 205 is regarded as the parent device, the corresponding two sensing nodes 201, 202 can be regarded as one or more sub-devices. In Figure 1 and Figure 2 the corresponding elements may have similar features and functions.

[0163] FC 207 can be used to obtain a global target. This global target can be formed according to an input instruction or given by an operator. For example, a human operator can input an instruction as a natural language query, and then the FC is used to parse the instruction into a formula in a suitable machine-oriented semantic language. Alternatively, a formula in a machine-oriented semantic language can be directly input into FC 207. Then, FC 207 is used to decompose the global target into one or more sub-goals. Any suitable algorithm for decomposing the global target into sub-goals can be used. One or more sub-goals will be assigned to appropriate sub-devices (or nodes) of the FC. One or more sub-goals assigned to a sub-node can be regarded as local goals. When assigning sub-goals to sub-nodes, one or more of the following information can be considered:

[0164] - The usefulness of the data that the sub-node can obtain;

[0165] - The processing capacity of the sub-node (e.g., the ability to extract specific types of facts from its input data);

[0166] - The available context / background knowledge of the sub-node (e.g., some domain-specific knowledge or implicitly learned knowledge about the usefulness of the sub-node for solving a specific task).

[0167] For example, as Figure 2 shown, the edge node 201 can have background knowledge BK1, while the intermediate node 205 can have background knowledge BK2.

[0168] The generated local goals are assigned by FC 207 as the parent node to appropriate intermediate nodes 205, 206 as sub-nodes. These intermediate nodes 205, 206 have their own sub-nodes, such as sensing nodes 201 to 204. The intermediate nodes 205, 206 can further transform the obtained local goals and use the same method as above to assign new sub-goals to the sensing nodes 201 to 204. The method of sub-goal generation and assignment can be executed iteratively layer by layer until all nodes of the network have local goals assigned.

[0169] It should be noted that there is no strict one-to-one correspondence between one or more local goals and one or more sub-nodes. For example, the same local goal can be assigned to multiple sub-nodes. One sub-node can be assigned multiple local sub-goals. Local sub-goals can be expressed in the semantic language understood by the sub-node. Local sub-goals can also include multiple other sub-goals, which can be assigned by the sub-node to other sub-nodes (i.e., at the next level). The assignment of goals (or sub-goals) may require each pair of directly connected parent and sub-nodes to reach an agreement on a common semantic language for communication. Optionally, the parent device and its sub-devices can reach an agreement on a shared semantic language. For example, as Figure 2As shown, the intermediate node 205 and the edge nodes 201 and 202 agree on the first shared semantic language L1, while the FC 207 and the intermediate nodes 205 and 206 agree on the second semantic language L2.

[0170] Optionally, multiple rounds of goal refinement can be allowed. For example, if the sub-goal assigned by the parent device to the child device is not applicable to the child device, the parent device can be used to update the sub-goal and assign the updated sub-goal to the child device. This process can be called goal refinement and can be executed for multiple rounds.

[0171] Figure 3 The system provided by the present disclosure is shown.

[0172] Similar to Figure 2 the system includes an FC 307, several intermediate nodes 305 and 306, and several edge nodes 301 to 304. The components introduced below can be included in the nodes of the system as functional / logical units.

[0173] The goal formation component can be included at each intermediate node and the FC 307. The goal formation component can be used to form goals, which represent or include one or more goals that the corresponding node wants to verify (or invalidate). Each goal may contribute to solving the main goal, which is finally calculated at the FC 307.

[0174] Optionally, formulas expressed in a suitable combined semantic language can be used to represent each goal. Each goal can include multiple sub-goals.

[0175] The goal assignment component can be included at each intermediate node and the FC 307. The goal assignment component can be used to convert the goals of the parent node (or global goals) into sub-goals (or local goals) to be assigned to each child node. This relationship can be recursive. For example, each sub-goal can be further transformed into other sub-goals. The sub-goals assigned to different child nodes can be expressed in different semantic languages. The same sub-goal can be assigned to multiple child nodes. A child node can be assigned one or more sub-goals.

[0176] The NN (i.e., the third NN of the child device, or the fourth NN of the parent device) processing component can be included at each sensing and intermediate node. The NN processing component can be used to process input data (e.g., sensing data of the sensing node or activation values of the intermediate node) and output a compressed representation of the input data. The compressed representation can be low-dimensional activation values. In the present disclosure, the NN processing component can also be referred to as the "NN-based compression component".

[0177] An SE component may be included at each node. The SE component can be used to extract one or more intermediate features from the input data. For a parent node, the input data can be the processed data from one or more of its child devices. Optionally and additionally, the parent node may further include a sensing unit for capturing its own sensing data as the input data. As an example regarding the function, the SE component can be used to detect one or more objects from the input data. Optionally, the attributes and / or relationships of one or more objects can be determined by the SE component. The SE component may include a neural network (the first NN of the parent device or the second NN of the child device) for performing the above tasks. Generally, the SE component can be used to perform any one or more of the following semantic tasks: classification task, object detection task, object recognition task, object relationship recognition task, action recognition task, and scene graph generation task. Various existing machine learning algorithms (such as neural networks) well-known in the art can be used to perform each of the above tasks.

[0178] An SP component may be included at each node. The SP component can be used to process the one or more intermediate features extracted by the corresponding SE component to verify the (local) goal of the node (i.e., the goal of the parent device or the sub-goal of the child device). The SP component can be used to better interpret symbolic facts using pre-coded (or learned) context or background knowledge. The SP component can be used to perform logical operations on one or more intermediate features and output a decision (if the local goal is verified). Optionally, the decision may be accompanied by a confidence level and / or additional information. The additional information may include one or more of the following: the geographical location and / or timestamp of the corresponding node. The timestamp can be used to indicate when the decision is made or when the input data is acquired.

[0179] After Figure 2 assigning local goals to each node in the network, the inference phase begins. In the inference phase, each edge device can be used to collect input data. For example, each edge device may include a sensor or may be connected to a sensor. The sensor is used to collect input data, such as multimodal data, and forward the input data to the corresponding edge device. On each edge device, the SE component is used to process the input data and output one or more semantic facts. One or more semantic facts can be a semantic interpretation of the observed scene. Then the one or more extracted semantic facts are provided to the corresponding SP component. Each SP component is used to verify the assigned local goal using semantic reasoning and local background knowledge.

[0180] If the local goal of the edge node can be verified, the edge node is used to send a decision and possibly some relevant data or features to the parent node. If the local goal of the edge cannot be verified, the edge node is used to compress the input data and send the compressed data to the parent device.

[0181] Optionally, a confidence value can be used to define whether a local target is verified. For example, a confidence level of at least 70% can indicate a high confidence, while a confidence level below 70% can indicate a low confidence. It should be noted that the value of 70% for the confidence level in this example is only given as an example, and different thresholds can be used in different scenarios, such as 50%, 80%, etc. Again, the confidence level can be fine-tuned according to various requirements such as accuracy and sensitivity.

[0182] If the local target of the edge node can be verified with a high confidence, the edge node is used to send a decision (e.g., a binary value or a flag) and possibly some related data / features to the parent node. As Figure 3 shown, if the condition of "Is the target solved?" is yes, the data path of the input sensed data is cut off. The NN processing component has no valid input, and thus, the NN processing component does not output anything. If the local target cannot be verified (e.g., the result of the SP component has a relatively low confidence), the data path of the input sensed data is opened. In this case, the edge node can be used to process its input data using its NN processing component. The NN processing component outputs a compressed version of the input data, such as an activation vector. These compressed data, possibly together with some related symbolic facts / data output by the SP component, are forwarded to the parent node. If the local target cannot be verified, the edge node does not send any decision. Alternatively, the edge node can also send only a synchronization flag. The synchronization flag can be used to notify the parent node that the local processing has been completed and the child node has no valid result, so that the parent node does not need to wait for a valid response from the child node.

[0183] The parent (e.g., intermediate) node is used to collect all the data (or responses) received from its child nodes and attempts to solve its own local target. When the parent node is FC 307, then FC 307 is used to solve the global target. The response can include a decision or compressed data. The compressed data can be an activation vector.

[0184] The activation values from multiple nodes are combined (e.g., joined, summed, concatenated, etc.) and processed by the SE component of the parent node. The SP component uses the extracted symbolic facts, as well as the facts indicated by any received decisions, to verify the local target. According to the verification result, the parent node can be used to continue processing as described above for the child node and send a decision or compressed data (e.g., an activation vector). Optionally and additionally, the output of the SP component can be used to refine the local target assigned to the child node. If necessary, the parent node can be used to update the local target and assign the updated local target to the child node.

[0185] Processing and communication proceed until FC 307 is reached. FC 307 is used to collect all the data received from its direct children and attempt to verify its global objective. Finally, FC 307 is used to put the final result. Optionally, FC 307 can be used to generate a new query (i.e., a new global objective), or send a termination signal as a result of verifying the global objective, or request additional data, etc.

[0186] Optionally, verifying a local objective at a parent node may require multiple rounds of communication ( "multiple-round communication") between the parent node and the child node. For example, this may be helpful when the parent node needs more evidence to verify its objective, or when the parent node decides to refine the assigned local objective based on the information already provided by the child node.

[0187] Optionally, the semantic communication network can adopt appropriate synchronization techniques to keep all nodes synchronized.

[0188] Figure 4 An application scenario provided by the present disclosure is shown. On Figure 4 the left side, a distributed network suitable for monitoring suburban security is shown. The distributed network exemplarily includes four sensing nodes (denoted as nodes 401 to 404), each node having a camera; two base stations (denoted as nodes 405, 406) as intermediate nodes; and a remote FC (denoted as node 407). The sensing nodes with cameras are used to monitor their respective sensing areas. Sensing nodes 401 and 402 (as sub-devices) are connected to base station 405 (as the parent device). Sensing nodes 403 and 404 (as sub-devices) are connected to base station 406 (as the parent device). Base stations 405 and 406 (as sub-devices) are connected to FC 407 (as the parent device).

[0189] For example, a simple global objective obtained or generated at FC 407 can be expressed as G0 = gg1 ∨ gg2. G0 represents the global objective, gg1 and gg2 represent two computing tasks, and the logical operator "∨" represents logical disjunction as logical "or (OR)". Each computing task can be a potential sub-objective. A sub-objective can also be a logical combination of multiple computing tasks, which results in the sub-objective being further divisible into multiple other sub-objectives. For example, G0 can represent the task of "detecting security risks", gg1 can represent the statement "there is a vehicle performing a security task", and gg2 can represent the statement "there is a person with a gun". As Figure 4 another example not shown in, gg1 can be further divided into two other sub-objectives: "there is a moving car" and "there is a person walking on the path of the car". These two other sub-objectives can be connected using the logical "and (AND)".

[0190] Figure 4An example of target allocation according to the present disclosure is shown on the right side. FC 407 obtains the global target G0 and converts this target into local targets for base stations 5 and 6. In this example, it is assumed that FC 407, based on its background knowledge and the obtained context, knows that base stations 405 and 406 can solve sub-goals gg1 and gg2. Therefore, the allocated local targets G5 and G6 are the same as the global target G0. Or, if base station 405 can only solve sub-goal gg1, the allocated local target G5 can be only gg1 ( Figure 4 not shown in

[0191] ). Base stations 405 and 406, as intermediate nodes, receive their local targets and further convert them into sub-goals for the sensing nodes. In this example, similar to FC 407, node 405 assigns exactly the same local target to nodes 401 and 402. On the other hand, node 406 only assigns gg1 to node 403 and assigns gg2 to node 404. This decision is made because node 406 realizes that node 403 does not know the concept of "gun" (therefore, it cannot verify gg2), and node 404 does not obtain useful input data for solving gg1 (for example, node 404 has never observed any cars in a pure pedestrian area).

[0192] Taking gg1 "a vehicle is performing a security task" as an example. The output (i.e., the one or more intermediate features) of the SE component at a sub-node with sensors (e.g., node 401) based on the input data (e.g., captured video / photo) can be:

[0193] - "The input data represents a car", with a confidence of 95%,

[0194] - "The car is red", with a confidence of 90%,

[0195] - "The car is moving fast (or exceeding a certain speed)", with a confidence of 80%,

[0196] - "The car is driven by a person", with a confidence of 90%,

[0197] - "The car is in a pedestrian area", with a confidence of 85%,

[0198] - "The input data also represents a tree", with a confidence of 95% etc.

[0199] Then, the one or more intermediate features are provided to the SP component for semantic processing (e.g., applying logical rules and semantic reasoning). For example, the SP component can perform logical operations on multiple intermediate features and verify the target based on the one or more intermediate features using its background knowledge.

[0200] For example, the SP component can consider and process the following intermediate features related to the target gg1:

[0201] - "The input data represents a car", with a confidence of 95%, "AND",

[0202] - "The car is moving fast (or exceeding a certain speed)", with a confidence of 80%, "AND",

[0203] - "The car is in a pedestrian area", with a confidence of 85%.

[0204] Use the following background knowledge of node 401 to verify these three intermediate features:

[0205] - Background knowledge: "A car is a vehicle";

[0206] - Background knowledge: "Moving fast in a pedestrian area poses a safety risk."

[0207] Then, the SP component outputs a decision flag. Optionally, a confidence value can be determined for the decision, such as a confidence of 80%.

[0208] It should be noted that when the child node detects that there is no car (for example, the confidence that the data represents a car is close to 0%), the child node can be configured to remain silent (for example, not send anything). Alternatively, only a synchronization flag can be sent. That is, in this case, the child node is not used to send a decision to its parent node. This is because the child node has no relevant information to send. However, if the local sub-goal is "detect that there is no car in the pedestrian area", then detecting no car is useful information and should be passed to the parent node.

[0209] Optionally, the decision can be accompanied by one or more extracted features and / or additional information. One or more extracted features can include a part of one or more intermediate features derived by the SE component. The additional information can include one or more of the following: the geographical location of the node, and / or a timestamp indicating when the decision was made (or when the input was captured).

[0210] Figure 5 An example for detecting a car with a safety hazard and the corresponding signaling between nodes.

[0211] Based on Figure 4 In the example, the cameras at nodes 401 and 402 can share a partially overlapping area. It should be noted that Figure 5 The description is given for the sub-goal gg1 introduced in Figure 4 Assume that in Figure 5 In the example, there is no valid input regarding the sub-goal gg2 (for example, no gun is detected at all). Figure 5The top shows frames of video (or photos) recorded by nodes 401 and 2 at three different frames (or timestamps) t1, t2, and t3. These frames show a moving car that may pose a safety risk to pedestrians. The communication between devices is as shown Figure 5 at the bottom.

[0212] At timestamp t1, only node 401 detects (or observes) the car. Node 401, which can solve sub-goal gg1, evaluates the data it has collected and determines with high confidence that the observed car poses a safety risk to the walking area. Therefore, it sends a decision flag to node 405 to report the event. At the same time, node 402 does not observe any cars, so it remains silent (i.e., does not send any decisions or processed data). Alternatively, node 402 can only send a synchronization flag (e.g., in each synchronization slot).

[0213] It should be noted that in some configurations, optionally, if the local goal can be verified as positive, the corresponding node can be used to send a decision or flag (also known as a "decision flag") indicating that the result of the computational task of the local goal is positive. If the local goal can be verified as negative, the corresponding node can remain silent to save communication costs (e.g., node 401 at frame t3). When the corresponding node is unsure whether it can verify the local goal, the corresponding node can compress its input data and send the compressed data to its parent node (e.g., node 401 at frame t2).

[0214] At timestamp t2, the dangerous car moves forward and is now partially observed by nodes 1 and 2. Nodes 401 and 402 detect the car but cannot evaluate with high confidence whether the car poses a safety risk. Therefore, both nodes 401 and 402 use their NN-based processing components to compress the corresponding images and send the activation vectors to their parent node: the intermediate node 405, instead of sending a decision flag. The intermediate node 405 receives the activation values and processes them jointly. The joint evaluation of the input data can help node 405 determine the presence of the unauthorized car with relatively high confidence. Therefore, node 5 is able to send a decision to FC407. In an alternative scenario where node 405 cannot verify its local goal, node 5 can be used to combine the data received from its child nodes 401 and 402, compress the combined data through its own NN processing component, and send the activation values to its parent node: FC 407.

[0215] At timestamp t3, the car posing the risk moves forward again, and it is only visible to node 402. Then, node 402 sends a decision to node 405 while node 401 remains silent. It can be noted that node 401 observes another car which is outside the pedestrian area and does not pose a danger to the pedestrians. Therefore, the presence of this safe car is not reported.

[0216] Figure 6 illustrates the synchronization mechanism provided by the present disclosure. Similar to Figures 2 to 4 Similar, Figure 6 illustrates a system which includes FC 607, several intermediate nodes 605, 606 and several edge nodes 601 to 604.

[0217] Solving the global objective in a distributed manner in the network may sometimes require a certain degree of synchronization among all nodes in the network. The benefit of synchronization may be to avoid devices reporting events that occur at different timestamps, while the corresponding parent devices do not know whether these events are related or isolated cases. This may lead to incorrect interpretations on the parent nodes and ultimately result in errors on the FC. Before reporting to FC 607, the intermediate nodes may perform multiple rounds of communication with their child nodes, which may exacerbate this problem.

[0218] To solve this problem, Figure 6 a synchronization mechanism is given. Nodes at the same level are given predefined time slots for communicating with their parent nodes. For example, any of the nodes 601 to 604 can be used to communicate with their parent nodes using any one or more of the three allocated time slots. The time slots can be allocated in such a way that the intermediate nodes can perform multiple rounds of data exchange with the edge nodes before sending their own decisions to their parent nodes or FC 607.

[0219] It should be noted that in the present disclosure, the synchronization mechanism is not essential but an optional feature.

[0220] Figure 7 illustrates the distributed joint training of the neural network provided by the present disclosure.

[0221] Similar to Figure 2 、 Figure 3 、 Figure 4 、 Figure 6 Similar, Figure 7 illustrates a system which includes FC 707, several intermediate nodes 705, 706 and several edge nodes 701 to 704.

[0222] In this disclosure, all neural networks (the first, second, third, and fourth NNs) for inference at nodes are trained and suitable for performing corresponding tasks. For obtaining these trained neural networks, a distributed joint training method using a multi-task loss function is disclosed. This training method is based on traditional in-network learning. In this training method, the network topology and the training data set are known. Different training data sets can be obtained according to different application scenarios. In addition, each parent node can efficiently combine the activation vectors separately sent by their own child nodes. This can be achieved by setting the hyperparameters of the neural network (e.g., the dimension of the activation vector), or by performing an initial calibration before starting the training.

[0223] For training the neural networks in the semantic communication network, appropriate training samples are loaded into the corresponding edge devices. For example, the training samples can include images, videos, and audio signals corresponding to various application scenarios. For example, if a global objective involves public safety, then images / videos of humans (e.g., criminals, etc.) and weapons can be provided as training samples. If the global objective is related to traffic control, then images / videos of cars and license plate numbers can be provided as training samples.

[0224] In addition, appropriate training labels are provided for each node in the network. The training labels at each node can help train the corresponding SE component to extract relevant semantic facts from the input data.

[0225] Then, the distributed joint training can be started using the multi-task loss function. The loss function is a metric used to describe how close the output of a neural network is to the ground truth. The goal of training is to minimize the value of the loss function. Usually, a neural network is designed to solve only one specific task, so usually a single-task loss function is required. In this disclosure, each neural network can be used to perform multiple tasks (e.g., object detection, object recognition, scene graph generation, etc.). Therefore, in this disclosure, a multi-task loss function that takes multiple tasks into account is used to train the corresponding neural networks.

[0226] In the forward pass, the edge device passes its input data through its neural network. The activation values output by the corresponding SE component are used to calculate the value of the local loss function, and the activation values output by the corresponding NN-based processing component are forwarded to the parent node.

[0227] The intermediate node, which is the parent node, combines the received activation values using predefined techniques (e.g., joining, concatenating, summing, etc.). The combined activation values are copied and then passed through the neural network at the intermediate node. Similar to before, the activation values output by the SE component are used to calculate the local loss, and the activation values output by the NN-based processing component are sent to the subsequent parent node at the next level. This process will continue until the FC is reached.

[0228] Backpropagation typically reverses the operations done in the forward pass. At each intermediate node, the error vectors at the inputs of the SE component and the NN-based processing component are fused (e.g., added). Then, the fused error vector is split by reversing the "combination" operation done in the forward pass. Then, the obtained split vectors are distributed to the respective child nodes. This process continues until each edge device processes the error vector and updates the weights of its respective neural network.

[0229] The forward pass and backpropagation are repeated until all neural networks converge.

[0230] It should be noted that in this disclosure, Figure 7 the distributed joint training is not essential but an optional feature / step. For example, when the neural network models of the devices (parent device and child devices) have already been trained, the distributed joint training does not need to be performed.

[0231] Figure 8 shows a method provided by this disclosure for constructing a set of potential goals that each node can solve. Similar to Figure 2 , Figure 3 , Figure 4 , Figure 6 and Figure 7 similar, Figure 8 shows a system that includes an FC 807, several intermediate nodes 805, 806, and several edge nodes 801 to 804.

[0232] One feature of this technical solution is to gradually assign local goals (or sub-goals) to each node in the network so that all nodes can contribute to solving the global goal. For each parent node, it would be beneficial to know what types of queries their child nodes can answer.

[0233] One possible technical solution is for the operator to manually pre-code the set of potential (or possible) local goals for each node. However, in a large network with hundreds of devices, this method may become cumbersome. Therefore, a method using the training dataset is introduced to gradually construct and propagate the set of possible goals from the edge devices to the FC. This method is shown in Figure 8 . As a first step, each edge device uses its local training labels and local background knowledge to construct a set of possible goals that the edge device can verify. The background knowledge can be pre-coded or learned from its own input data. Then, the obtained set of possible goals is passed to the parent node. Then, the parent node constructs its own set of possible goals and forwards it to the next node. This process is repeated until the FC is reached.

[0234] The target discovery process described above provides a technical solution for target allocation. During target allocation, the parent node checks whether any of its local targets (or any subset of the local targets) match any of the possible targets in the child node's set of possible targets. If so, the parent node assigns the corresponding sub-goal to the child node.

[0235] Optionally, to enrich the local context or background knowledge of the node, a mechanism of implicit learning (e.g., learning from examples) is introduced. During the inference phase, a processing node observing many running examples may derive some new facts and relationships that do not appear in the training dataset. If the updated background knowledge helps the processing node solve any new goals, this information should be passed to its parent node and then propagated until it reaches the FC.

[0236] Figure 9 Method 900 provided by the present disclosure is shown. Method 900 includes the following steps:

[0237] - Step 901: The parent device obtains a target, where the target includes one or more computing tasks;

[0238] - Step 902: The parent device determines one or more sub-goals based on the target, where each sub-goal is at least a subset of the target;

[0239] - Step 903: The parent device sends one or more sub-goals to one or more child devices;

[0240] - Step 904: In response to receiving a decision on the corresponding sub-goal from the corresponding child device: The parent device performs semantic processing on the decision to verify the target.

[0241] In response to not receiving a decision on the corresponding sub-goal from the corresponding child device, the following steps 905 to 907 are performed:

[0242] - Step 905: The parent device obtains the processed data from the corresponding child device;

[0243] - Step 906: The parent device performs semantic extraction on the processed data to obtain one or more intermediate features of the processed data;

[0244] - Step 907: The parent device performs semantic processing on one or more intermediate features to verify the target.

[0245] Figure 10 Another method 1000 provided by the present disclosure is shown. Method 1000 includes the following steps:

[0246] - Step 1001: The child device obtains a sub-goal from the parent device;

[0247] - Step 1001: The sub-device obtains input data, performs semantic extraction on the input data to obtain one or more intermediate features of the input data;

[0248] - Step 1003: The sub-device performs semantic processing on one or more intermediate features to verify the sub-goal;

[0249] - Step 1004: In response to the sub-device determining that the sub-goal is verified: The sub-device sends a decision of the sub-goal to the parent device.

[0250] In response to the sub-device determining that the sub-goal is not verified, the following steps 1005 and 1006 are performed:

[0251] - Step 1005: The sub-device compresses the input data to obtain processed data;

[0252] - Step 1006: The sub-device sends the processed data to the parent device.

[0253] Optionally, before determining whether the sub-goal is verified, the sub-device can be used to determine whether there is valid input. This can be determined based on any of the following conditions:

[0254] - Whether there is any valid input for semantic extraction;

[0255] - Whether there are any extracted intermediate features.

[0256] If there is no valid input or no extracted intermediate features, the sub-device can be used to remain silent (i.e., not send any decisions or processed data). Alternatively, the sub-device can be used to send a synchronization flag to the parent device.

[0257] It should be noted that, from the above Figures 1 to 8 perspective, the steps of methods 900 and 1000 can have the same functions and details. Therefore, the corresponding method implementation manners are not described in detail at this time.

[0258] Figure 11 illustrates the effect of the intermittent communication provided by the present disclosure.

[0259] In a distributed communication network, a device only sends data when it detects relevant information. In addition, a decision flag or a compressed activation value is sent. Compared with traditional distributed inference technologies, the present disclosure can achieve intermittent communication, thereby significantly reducing the average communication rate.

[0260] Figure 11 The idea of intermittent communication is shown in Figure 4The multi-hop topology illustrated as an example in the present disclosure. In the centralized solution, the edge device 401 does not process its input data, but continuously sends the uncompressed photo / video to the parent node 405 until the FC 407. In this case, the required coding rate is relatively high. When the edge device uses a trained neural network to compress the video (e.g., frame by frame) and continuously transmits the generated activation values to the parent node, the required coding rate can be slightly reduced compared to the uncompressed data. In the technical solution according to the present disclosure, the edge device 401 sends the decision flag 1101 only when the sub-goal is verified, or sends the processed data (e.g., activation values) 1102 only when the sub-goal cannot be (reliably) verified. In this way, the required coding rate can be significantly reduced.

[0261] Generally speaking, compared with traditional methods, the present disclosure provides significantly improved performance (e.g., lower latency, robustness, flexibility, and energy consumption).

[0262] The present disclosure can be applied to any type of communication network. For example, the present disclosure can be applied to a V2X network, where traffic monitoring is required for various purposes (or goals), such as vehicle detection, obstacle detection, speed control, collision detection, etc. For another example, the present disclosure can be applied to an IoT network, such as Industry 4.0, where various types of sensing nodes are distributed to verify one or more goals.

[0263] The devices in the present disclosure (e.g., the parent device and the child device) can include a processing circuit (not shown) for respectively performing, conducting, or initiating the various operations of the devices described herein. The processing circuit can include hardware and software. The hardware can include analog circuits or digital circuits, or both analog circuits and digital circuits. The digital circuit can include components such as an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), or a general-purpose processor. The processing circuit includes one or more processors and a non-transitory memory connected to the one or more processors. The non-transitory memory can carry executable program code, which when executed by the one or more processors, causes the device to respectively perform, conduct, or initiate the operations or methods described herein.

[0264] Optionally, the device in the present disclosure can be a single electronic device capable of performing calculations, or can include a set of connected electronic devices capable of performing calculations using a shared system memory. As is well known in the art, such computing capabilities can be incorporated into many different devices, so the term "device" can include a PC, server, mobile terminal, tablet computer, wearable device, graphics processing unit, graphics card, etc.

[0265] The present disclosure has been described in connection with various examples and implementations. However, upon study of the drawings, the present disclosure, and the independent claims, other variations can be understood and implemented by those skilled in the art when practicing the claimed subject matter. In the claims as well as the specification, the word "comprising" does not exclude other elements or steps, and "a" does not exclude a plurality. A single element or other unit can fulfill the functions of several entities or items described in the claims. The fact that certain measures are recited in mutually different dependent claims does not mean that a combination of these measures cannot be used in an advantageous implementation.

Claims

1. A parent device (110), characterized in that, For: Obtain a target (101), where the target (101) includes one or more computing tasks; Determine one or more sub - targets based on the target (101), where each sub - target (102) is at least a subset of the target (101); Send the one or more sub - targets to one or more sub - devices; In response to receiving a decision (105) of a corresponding sub - target (102) from a corresponding sub - device (120), perform semantic processing on the decision (105) to verify the target; In response to not receiving the decision (105) of the corresponding sub - target (102) from the corresponding sub - device (120), obtain processed data (106) from the corresponding sub - device (120), perform semantic extraction on the processed data (106) to obtain one or more intermediate features (107) of the processed data (106), and perform semantic processing on the one or more intermediate features (107) to verify the target (101).

2. The parent device (110) according to claim 1, characterized in that, The decision (105) includes an indication bit indicating that the sub - target (102) is verified.

3. The parent device (110) according to claim 1 or 2, characterized in that, The decision (105) includes the confidence of the decision (105).

4. The parent device (110) according to any one of claims 1 to 3, characterized in that The decision (105) further includes one or more extracted features.

5. The parent device (110) according to any one of claims 1 to 4, characterized in that, For performing semantic extraction on the processed data (106), the parent device (110) is used to detect one or more objects, and optionally, detect one or more attributes of the detected objects from the processed data (106).

6. The parent device (110) according to any one of claims 1 to 5, characterized in that, The parent device (110) includes a first neural network model for performing semantic extraction.

7. The parent device (110) according to any one of claims 1 to 6, characterized in that, For performing semantic processing, the parent device (110) is used to verify semantic facts related to the target (101).

8. The parent device (110) according to any one of claims 1 to 7, characterized in that, In response to receiving multiple decisions from multiple sub - devices, the parent device (110) is used to combine the multiple decisions and perform semantic processing on the combined decisions.

9. The parent device (110) according to any one of claims 1 to 8, characterized in that, In response to obtaining multiple processed data from multiple sub - devices, the parent device (110) is used to concatenate, sum, or splice the multiple processed data.

10. The parent device (110) according to any one of claims 1 to 9, characterized in that, In response to determining that the target (101) is not verified, the parent device (110) is further used to: - Determine the one or more updated sub - targets based on the target (101); - Send the one or more updated sub - targets to the one or more sub - devices.

11. The parent device (110) according to any one of claims 1 to 10, characterized in that, The parent device (110) is used to communicate with the one or more sub - devices in one or more predetermined time slots.

12. The parent device (110) according to any one of claims 1 to 11, characterized in that, The parent device (110) is further used to select the one or more sub - devices based on side information, where the side information includes one or more of the following: - The usefulness of the input data obtained by each sub - device; - The processing capacity of each sub - device; - The acquired knowledge of each sub - device; - The acquired knowledge of the parent device (110).

13. A sub-device (120), characterized in that, For: Obtain a sub - target (102) from a parent device (110), where the sub - target (102) is at least a subset of the target (101) of the parent device (110), and where the target (101) includes one or more computing tasks; Obtain input data (103) and perform semantic extraction on the input data (103) to obtain one or more intermediate features (104) of the input data (103); Perform semantic processing on the one or more intermediate features (104) to verify the sub-goal (102); In response to determining that the sub-goal (102) is verified, send a decision (105) of the sub-goal (102) to the parent device (110); In response to determining that the sub-goal (102) is not verified, compress the input data (103) to obtain processed data (106), and send the processed data (106) to the parent device (110).

14. The sub-device (120) according to claim 13, characterized in that, The decision (105) includes an indication bit indicating that the sub-goal (102) is verified.

15. The sub-device (120) according to claim 13 or 14, characterized in that, The decision (105) includes the confidence level of the decision (105).

16. The sub-device (120) according to any one of claims 13 to 15, characterized in that, For performing semantic extraction on the input data (103), the sub-device (120) is used to detect one or more objects, and optionally, one or more attributes of the detected object are detected from the input data (103).

17. The sub-device (120) according to any one of claims 13 to 16, characterized in that, The sub-device (120) includes a second neural network model for performing semantic extraction.

18. The sub-device (120) according to any one of claims 13 to 17, characterized in that, For performing semantic processing, the sub-device (120) is used to verify semantic facts related to the sub-goal (102).

19. The sub-device (120) according to any one of claims 13 to 18, characterized in that, For compressing the input data (103), the sub-device (120) includes a third neural network model (123) for inferring the input data (103) to obtain a feature map of the input data (103) as the processed data (106).

20. The sub-device (120) according to any one of claims 13 to 19, characterized in that The sub-device (120) is used to communicate with the parent device (110) in one or more predetermined time slots.

21. A system, characterized in that, Includes one or more parent devices (110) according to any one of claims 1 to 12 and one or more sub-devices (120) according to any one of claims 13 to 20.

22. A method (900), characterized in that, Includes: The parent device obtains (901) a goal, where the goal includes one or more computing tasks; The parent device determines (902) one or more sub-goals based on the goal, where each sub-goal is at least a subset of the goal; The parent device sends (903) the one or more sub-goals to one or more sub-devices; In response to receiving a decision of a corresponding sub-goal from a corresponding sub-device: the parent device performs (904) semantic processing on the decision to verify the goal; In response to not receiving the decision of the corresponding sub-goal from the corresponding sub-device: obtain (905) processed data from the corresponding sub-device; perform (906) semantic extraction on the processed data to obtain one or more intermediate features of the processed data, and perform (907) semantic processing on the one or more intermediate features to verify the goal.

23. A method (1000), characterized in that, Includes: The sub-device obtains (1001) a sub-goal from the parent device, where the sub-goal is at least a subset of the goal of the parent device, and the goal includes one or more computing tasks; The sub-device obtains (1002) input data and performs semantic extraction on the input data to obtain one or more intermediate features of the input data; The sub-device performs (1003) semantic processing on the one or more intermediate features to verify the sub-goal; In response to determining that the sub-goal is verified: The sub-device sends (1004) the decision of the sub-goal to the parent device; In response to determining that the sub-goal is not verified: The sub-device compresses (1005) the input data to obtain processed data; The sub-device sends (1006) the processed data to the parent device.

24. A computer program comprising instructions, characterized in that, When the program is executed by a computer, the instructions cause the computer to execute the method according to claim 22 or 23.