Task determination method and apparatus, communication device, storage medium and program product

By deploying miniature machine learning models and encoder-decoder systems on drones to process multimodal data, the problem of limited storage and computing resources for drones in an integrated air-ground network is solved, thereby improving the accuracy and efficiency of mission execution.

WO2026085879A1PCT designated stage Publication Date: 2026-04-30BEIJING XIAOMI MOBILE SOFTWARE CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
BEIJING XIAOMI MOBILE SOFTWARE CO LTD
Filing Date
2024-10-25
Publication Date
2026-04-30

Smart Images

  • Figure CN2024127479_30042026_PF_FP_ABST
    Figure CN2024127479_30042026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to a task determination method and apparatus, a communication device, a storage medium and a program product. The method is executed by a first device, and comprises: acquiring multi-modal data, wherein the multi-modal data comprises a plurality of pieces of data, and the plurality of pieces of data belong to at least two types; on the basis of the multi-modal data, determining a semantic intention by means of a coder and a decoder; and, on the basis of the semantic intention, determining an utterance intention corresponding to the multi-modal data, wherein the utterance intention is used for indicating a first task. The solution of the present disclosure can improve the task execution capability of communication devices deployed in the air.
Need to check novelty before this filing date? Find Prior Art

Description

Task determination methods and apparatus, communication equipment, storage media and program products Technical Field

[0001] This disclosure relates to the field of wireless communication, and more particularly to a task determination method and apparatus, communication equipment, storage medium, and program product. Background Technology

[0002] The concept of space-air-ground integrated communication (SAGIN) helps to build a communication network with wide coverage, high security and reliability, and high speed and intelligence. As an important component of SAGIN, communication equipment deployed in the air, such as drones, provides communication services with high flexibility, wide coverage, and high communication efficiency.

[0003] Summary of the Invention

[0004] This disclosure provides a task determination method and apparatus, communication equipment, storage medium, and program product.

[0005] According to a first aspect of the present disclosure, a task determination method is provided. The method is performed by a first device. The method includes: acquiring multimodal data, wherein the multimodal data includes multiple data sets belonging to at least two types; determining semantic intent based on the multimodal data using an encoder and a decoder; and determining terminological intent corresponding to the multimodal data based on the semantic intent, wherein the terminological intent is used to indicate a first task.

[0006] According to a second aspect of the present disclosure, a task determination apparatus is provided. The apparatus is disposed in a first device. The apparatus includes a processing module. The processing module is configured to: acquire multimodal data, wherein the multimodal data includes multiple data items belonging to at least two types; determine semantic intent based on the multimodal data using an encoder and a decoder; and determine terminological intent corresponding to the multimodal data based on the semantic intent, wherein the terminological intent is used to indicate a first task.

[0007] According to a third aspect of the present disclosure, a communication device is provided. The communication device includes: one or more processors and a memory storing instructions. When executed by the communication device, the instructions cause the communication device to implement the task determination method as described in the first aspect.

[0008] According to a fourth aspect of the present disclosure, a storage medium is provided. The storage medium stores instructions. When executed on a communication device, the instructions cause the communication device to perform the task determination method as described in the first aspect.

[0009] According to a fifth aspect of the present disclosure, a program product is provided. When executed by a communication device, the program product causes the communication device to perform the task determination method as described in the first aspect.

[0010] According to a sixth aspect of the present disclosure, a computer program is provided. When this computer program is run on a computer, it causes the computer to perform the task determination method as described in the first aspect.

[0011] According to a seventh aspect of this disclosure, a chip or chip system is provided. The chip or chip system includes processing circuitry. The processing circuitry is configured to perform the task determination method as described in the first aspect.

[0012] According to embodiments of this disclosure, the task execution capability of communication devices deployed in the air can be improved.

[0013] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not constitute a limitation on the embodiments of this disclosure. Attached Figure Description

[0014] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments of the invention and, together with the description, serve to explain the principles of the embodiments of the invention.

[0015] Figure 1 is a schematic diagram of the architecture of a communication system provided according to an embodiment of the present disclosure.

[0016] Figure 2 is a schematic diagram of the architecture of the integrated air-space-ground network.

[0017] Figure 3 is a schematic diagram of the semantic communication system model.

[0018] Figure 4 is a flowchart illustrating the task determination method provided according to an embodiment of the present disclosure.

[0019] Figure 5 is a schematic diagram of acquiring multimodal data in an embodiment of this disclosure.

[0020] Figure 6 is a schematic diagram of the framework of virtual semantic communication provided according to an embodiment of the present disclosure.

[0021] Figure 7A is an interactive schematic diagram illustrating the task execution implemented according to the task determination method provided in the embodiments of this disclosure.

[0022] Figure 7B is an interactive schematic diagram of model training implemented according to the task determination method provided in the embodiments of this disclosure.

[0023] Figure 8A is a schematic diagram of the task execution process provided according to an embodiment of the present disclosure.

[0024] Figure 8B is a schematic diagram of the conversion of multimodal input to data vector according to an embodiment of the present disclosure.

[0025] Figure 8C is a schematic diagram of the correspondence between the task execution system and virtual semantic communication provided according to an embodiment of the present disclosure.

[0026] Figure 8D is a schematic diagram of the logical layering of a task execution system provided according to an embodiment of the present disclosure.

[0027] Figure 8E is a schematic diagram of the framework of a virtual semantic communication system provided according to an embodiment of the present disclosure.

[0028] Figure 9 is a schematic diagram of the structure of a task determination device provided according to an embodiment of the present disclosure.

[0029] Figure 10A is a schematic diagram of the structure of a communication device provided according to an embodiment of the present disclosure.

[0030] Figure 10B is a schematic diagram of the structure of a chip provided according to an embodiment of the present disclosure. Detailed Implementation

[0031] This disclosure provides a task determination method and apparatus, a communication device, a storage medium, and a program product.

[0032] In a first aspect, embodiments of this disclosure provide a task determination method. The method is performed by a first device. The method includes: acquiring multimodal data, wherein the multimodal data includes multiple data sets belonging to at least two types; determining semantic intent based on the multimodal data using an encoder and a decoder; and determining terminological intent corresponding to the multimodal data based on the semantic intent, wherein the terminological intent is used to indicate a first task.

[0033] In this embodiment, multimodal data is processed by an encoder and a decoder to obtain the corresponding semantic intent. This semantic intent indicates a first task. Thus, the first device can determine and execute the first task based on the multimodal data using the encoder and decoder in semantic communication. This improves the accuracy of the first device in determining the first task, thereby enhancing its task execution capability.

[0034] In conjunction with some embodiments of the first aspect, in some embodiments, the operation of determining semantic intent based on multimodal data by an encoder and a decoder may include: determining a first symbol based on multimodal data by an encoder, wherein the first symbol contains semantic information of the multimodal data; and determining semantic intent based on the first symbol by a decoder.

[0035] In conjunction with some embodiments of the first aspect, in some embodiments, the operation of determining a first symbol by an encoder based on multimodal data may include: extracting features from each of a plurality of data in the multimodal data to obtain a first feature; transforming the first feature of each data to obtain a second feature, wherein the second features corresponding to the plurality of data have the same dimension; and channel coding the second feature to obtain a first symbol corresponding to the second feature.

[0036] In conjunction with some embodiments of the first aspect, in some embodiments, the operation of determining semantic intent through a decoder based on a first symbol may include: performing channel decoding on the first symbol to obtain multiple third features, wherein each of the multiple third features corresponds to one of the multiple data; performing semantic decoding on each third feature to obtain semantic information; and fusing the multiple semantic information of the multimodal data to determine the semantic intent.

[0037] In conjunction with some embodiments of the first aspect, in some embodiments, the operation of determining the terminology intent corresponding to multimodal data based on semantic intent may include: determining the terminology intent corresponding to the semantic intent based on a terminology intent database.

[0038] In conjunction with some embodiments of the first aspect, in some embodiments, the multimodal data is obtained based on at least one of the following: communication signals; sensor signals; power supply signals; task calculation results.

[0039] In conjunction with some embodiments of the first aspect, in some embodiments, the operation of acquiring multimodal data may include at least one of the following: receiving multimodal data; acquiring multimodal data from a local source.

[0040] In conjunction with some embodiments of the first aspect, in some embodiments, the encoder and decoder are implemented based on a first machine learning model; wherein, the above method may further include: performing model training on the first machine learning model according to the intended meaning.

[0041] In this embodiment, an encoder and decoder are implemented using a first machine learning model, and the encoder and decoder are used to determine the first task corresponding to the multimodal data. Thus, after training the first machine learning model with a large amount of multimodal data, the first machine learning model can extract and process features from heterogeneous data across multiple modalities to determine the first task. This improves the accuracy of the first device in determining the first task.

[0042] In conjunction with some embodiments of the first aspect, in some embodiments, the first machine learning model includes a micro machine learning model.

[0043] In this embodiment, the first machine learning model can be a micro machine learning model. By deploying the micro machine learning model on the first device and implementing task determination, the requirements for hardware resources such as storage space and computing resources of the first device are greatly reduced.

[0044] In conjunction with some embodiments of the first aspect, in some embodiments, the training of the first machine learning model is implemented based on federated learning or segmentation learning.

[0045] In conjunction with some embodiments of the first aspect, in some embodiments, the first device is a drone.

[0046] In a second aspect, embodiments of this disclosure provide that the aforementioned apparatus includes a processing module. The processing module is configured to: acquire multimodal data, wherein the multimodal data includes multiple data, the multiple data belonging to at least two types; determine semantic intent based on the multimodal data by an encoder and a decoder; and determine terminology intent corresponding to the multimodal data according to the semantic intent, wherein the terminology intent is used to indicate a first task.

[0047] In conjunction with some embodiments of the second aspect, in some embodiments, the processing module may be configured to: determine a first symbol by an encoder based on multimodal data, wherein the first symbol contains semantic information of the multimodal data; and determine a semantic intent by a decoder based on the first symbol.

[0048] In conjunction with some embodiments of the second aspect, in some embodiments, the processing module may be configured to: extract features from each of the multiple data in the multimodal data to obtain a first feature; transform the first feature of each data to obtain a second feature, wherein the second features corresponding to the multiple data have the same dimension; and perform channel coding on the second feature to obtain a first symbol corresponding to the second feature.

[0049] In conjunction with some embodiments of the second aspect, in some embodiments, the processing module may be configured to: perform channel decoding on the first symbol to obtain a plurality of third features, wherein each of the plurality of third features corresponds to one of the plurality of data; perform semantic decoding on each third feature to obtain semantic information; and perform fusion processing on the plurality of semantic information of the multimodal data to determine the semantic intent.

[0050] In conjunction with some embodiments of the second aspect, in some embodiments, the processing module may be configured to: determine the terminology intent corresponding to the semantic intent based on a terminology intent database.

[0051] In conjunction with some embodiments of the second aspect, in some embodiments, the multimodal data is obtained based on at least one of the following: communication signals; sensor signals; power supply signals; task calculation results.

[0052] In conjunction with some embodiments of the second aspect, in some embodiments, the above-described apparatus may further include a transceiver module configured to receive multimodal data; and a processing module configured to acquire multimodal data from a local source.

[0053] In conjunction with some embodiments of the second aspect, in some embodiments, the encoder and decoder are implemented based on a first machine learning model; wherein the processing module may further include: performing model training on the first machine learning model according to the intended meaning.

[0054] In conjunction with some embodiments of the second aspect, in some embodiments, the first machine learning model includes a micro machine learning model.

[0055] In conjunction with some embodiments of the second aspect, in some embodiments, the training of the first machine learning model is implemented based on federated learning or segmentation learning.

[0056] In conjunction with some embodiments of the second aspect, in some embodiments, the first device is a drone.

[0057] In a third aspect, embodiments of this disclosure provide a communication device. The communication device includes one or more processors and a memory storing instructions. When executed by the communication device, the instructions cause the communication device to implement the task determination method as described in the first aspect and its possible embodiments.

[0058] In a fourth aspect, embodiments of this disclosure provide a storage medium storing instructions. When executed on a communication device, the instructions cause the communication device to perform the task determination method as described in the first aspect and any of its possible embodiments.

[0059] In a fifth aspect, embodiments of this disclosure provide a program product. When executed by a communication device, this program product causes the communication device to perform the task determination method as described in the first aspect and any of its possible embodiments.

[0060] In a sixth aspect, embodiments of this disclosure provide a computer program. When this computer program is run on a computer, it causes the computer to perform the task determination method as described in any of the first aspect and its possible implementations.

[0061] In a seventh aspect, embodiments of this disclosure provide a chip or chip system. The chip or chip system includes processing circuitry. The processing circuitry is configured to perform the task determination method as described in any of the first aspects and their possible implementations.

[0062] Understandably, the aforementioned task determining apparatus, communication device, storage medium, program product, computer program, chip, and chip system are all used to execute the methods provided in the embodiments of this disclosure. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects in the corresponding methods, and will not be repeated here.

[0063] This disclosure provides a task determination method and apparatus, a communication device, a storage medium, and a program product. In some embodiments, terms such as task determination method, information processing method, and information transmission method can be used interchangeably; terms such as task determination apparatus, communication device, network device, network function, and network entity can be used interchangeably; and terms such as communication system and information processing system can be used interchangeably.

[0064] This disclosure is not exhaustive, but merely illustrative of some embodiments, and is not intended to limit the scope of protection of this disclosure. Unless otherwise specified, each step in a particular embodiment can be implemented as an independent embodiment, and the steps can be arbitrarily combined. For example, a solution after removing some steps in a particular embodiment can also be implemented as an independent embodiment, and the order of the steps in a particular embodiment can be arbitrarily interchanged. Furthermore, the optional implementation methods in a particular embodiment can be arbitrarily combined; moreover, the embodiments can be arbitrarily combined, for example, some or all steps of different embodiments can be arbitrarily combined, and a particular embodiment can be arbitrarily combined with the optional implementation methods of other embodiments.

[0065] In the embodiments disclosed herein, unless otherwise specified or in case of logical conflict, the terminology and / or descriptions of the various embodiments are consistent and can be referenced by each other. Technical features in different embodiments can be combined to form new embodiments based on their inherent logical relationships.

[0066] The terminology used in the embodiments of this disclosure is for the purpose of describing particular embodiments only and is not intended to limit the scope of this disclosure.

[0067] In the embodiments of this disclosure, unless otherwise stated, elements expressed in the singular form, such as “a,” “one,” “a kind,” “the,” “the,” “the,” “the,” “the,” “the,” “the,” “the,” “this,” etc., can mean “one and only one,” or “one or more,” “at least one,” etc. For example, when articles such as “a,” “an,” and “the” are used in translation, the noun following the article can be understood as either a singular or a plural expression.

[0068] In the embodiments of this disclosure, "a plurality of" means two or more.

[0069] In some embodiments, terms such as “at least one (at least one, at least one item, at least one)” and “one or more” may be used interchangeably.

[0070] In some embodiments, the notation "at least one of A and B", "A and / or B", "A in one case, B in another", "in response to one case A, in response to another case B", etc., may include the following technical solutions depending on the situation: in some embodiments, A (execute A regardless of B); in some embodiments, B (execute B regardless of A); in some embodiments, execution is selected from A and B (A and B are selectively executed); in some embodiments, A and B (both A and B are executed). The same applies when there are more branches such as A, B, C, etc.

[0071] In some embodiments, the notation "A or B" may include the following technical solutions, depending on the situation: in some embodiments, A (execution of A regardless of B); in some embodiments, B (execution of B regardless of A); in some embodiments, execution is selected from A and B (A and B are selectively executed). The same applies when there are more branches such as A, B, C, etc.

[0072] The prefixes "first," "second," etc., used in the embodiments of this disclosure are merely for distinguishing different descriptive objects and do not impose restrictions on the position, order, priority, quantity, or content of the descriptive objects. The description of the descriptive objects is found in the claims or the context of the embodiments, and the use of prefixes should not constitute unnecessary restrictions. For example, if the descriptive object is a "field," the ordinal numbers preceding "field" in "first field" and "second field" do not restrict the position or order of the "fields." "First" and "second" do not restrict whether the "fields" they modify are in the same message, nor do they restrict the order of "first field" and "second field." Similarly, if the descriptive object is a "level," the ordinal numbers preceding "level" in "first level" and "second level" do not restrict the priority between "levels." Furthermore, the number of descriptive objects is not limited by ordinal numbers and can be one or more. For example, in "first device," the number of "devices" can be one or more. Furthermore, the objects modified by different prefixes can be the same or different. For example, if the object being described is "device", then "first device" and "second device" can be the same device or different devices, and their types can be the same or different. As another example, if the object being described is "information", then "second information" and "first information" can be the same information or different information, and their content can be the same or different.

[0073] In some embodiments, “including A,” “containing A,” “for indicating A,” and “carrying A” can be interpreted as directly carrying A or indirectly indicating A.

[0074] In some embodiments, the terms “in response to…”, “in response to determining…”, “in the case of…”, “when…”, “if…”, “if…”, etc., can be used interchangeably.

[0075] In some embodiments, devices, etc., can be interpreted as physical or virtual, and their names are not limited to the names recorded in the embodiments. Terms such as “device”, “equipment”, “circuit”, “network element”, “node”, “function”, “unit”, “section”, “system”, “network”, “chip”, “chip system”, “entity”, and “subject” can be used interchangeably.

[0076] In some embodiments, "network" can be interpreted as devices included in a network (e.g., access network devices, core network devices, etc.).

[0077] In some embodiments, the terms "access network device (AN device)," "radio access network device (RAN device)," "base station (BS)," "radio base station," "fixed station," "node," "access point," "transmission point (TP)," "reception point (RP)," "transmission / reception point (TRP)," "panel," "antenna panel," "antenna array," "cell," "macro cell," "small cell," "femto cell," "pico cell," "sector," "cell group," "serving cell," "carrier," "component carrier," and "bandwidth part (BWP)" can be used interchangeably.

[0078] In some embodiments, the terms "terminal", "terminal device", "user equipment (UE)", "user terminal", "mobile station (MS)", "mobile terminal (MT)", "subscriber station", "mobile unit", "subscriber unit", "wireless unit", "remote unit", "mobile device", "wireless device", "wireless communication device", "remote device", "mobile subscriber station", "access terminal", "mobile terminal", "wireless terminal", "remote terminal", "handset", "user agent", "mobile client", and "client" can be used interchangeably.

[0079] In some embodiments, access network devices, core network devices, or network devices can be replaced by terminals. For example, embodiments of this disclosure can also be applied to structures where communication between access network devices, core network devices, or network devices and terminals is replaced by communication between multiple terminals (e.g., device-to-device (D2D), vehicle-to-everything (V2X), etc.). In this case, the structure can also be configured such that the terminal has all or part of the functions of the access network device. Furthermore, terms such as "uplink" and "downlink" can be replaced with terms corresponding to communication between terminals (e.g., "sidelink"). For example, uplink channel, downlink channel, etc., can be replaced with sidelink channel, and uplink link, downlink, etc., can be replaced with sidelink link.

[0080] In some embodiments, the terminal may be replaced by an access network device, a core network device, or a network device. In this case, the access network device, core network device, or network device may also be configured to have all or some of the functions of the terminal.

[0081] In some embodiments, the acquisition of data, information, etc., may comply with the laws and regulations of the country where the location is situated.

[0082] In some embodiments, data, information, etc., may be obtained with the user's consent.

[0083] Furthermore, each element, each row, or each column in the table of this disclosure can be implemented as an independent embodiment, and any combination of any element, any row, or any column can also be implemented as an independent embodiment.

[0084] Figure 1 is a schematic diagram of the architecture of a communication system provided according to an embodiment of the present disclosure. As shown in Figure 1, the communication system 100 includes a first device 101, a second device 102, and a third device 103.

[0085] In some embodiments, the first device 101 may be used to perform services such as terminal communication and sensing.

[0086] In some embodiments, the first device 101 may be an access network device. In one example, the first device 101 may be a drone. In another example, the first device 101 may be a high-altitude platform.

[0087] In some embodiments, the second device 102 may be used to wirelessly communicate with the first device 101 and be responsible for transmitting signaling and / or data to the first device 101.

[0088] In some embodiments, the second device 102 may be a gateway (GW) device. In one example, the second device 102 may be a signaling gateway. In another example, the second device 102 may be a service gateway.

[0089] In some embodiments, the second device 102 may be a core network (CN) device.

[0090] In some embodiments, the third device 103 may be responsible for training artificial intelligence (AI) and / or machine learning (ML) models.

[0091] In some embodiments, the third device 103 may be a core network device.

[0092] In some embodiments, the third device 103 may be an application function (AF) and / or an application server (AS).

[0093] In some embodiments, the terminal includes, but is not limited to, at least one of the following: mobile phone, wearable device, Internet of Things device, car with communication function, smart car, tablet computer, computer with wireless transceiver function, virtual reality (VR) terminal device, augmented reality (AR) terminal device, wireless terminal device in industrial control, wireless terminal device in self-driving, wireless terminal device in remote medical surgery, wireless terminal device in smart grid, wireless terminal device in transportation safety, wireless terminal device in smart city, and wireless terminal device in smart home.

[0094] In some embodiments, the access network device is, for example, a node or device that connects a terminal to a wireless network. In some embodiments, the access network device may include, but is not limited to, at least one of the following in a 5G communication system: evolved Node B (eNB), next-generation eNB (ng-eNB), next-generation Node B (gNB), node B (NB), home node B (HNB), home evolved node B (HeNB), radio backhaul device, radio network controller (RNC), base station controller (BSC), base transceiver station (BTS), base band unit (BBU), mobile switching center, base station in a 6G communication system, open RAN, cloud RAN, base station in other communication systems, and access node in a Wi-Fi system.

[0095] In some embodiments, the technical solutions of this disclosure can be applied to Open Radio Access Network (Open RAN) architectures. In this case, the interfaces between or within access network devices involved in the embodiments of this disclosure can be transformed into internal interfaces of Open RAN. The processes and information interactions between these internal interfaces can be implemented by software or programs.

[0096] In some embodiments, the access network device may be composed of a central unit (CU) and a distributed unit (DU). The CU may also be called a control unit. The CU-DU structure can separate the protocol layer of the access network device. Some of the protocol layer functions are centrally controlled by the CU, while the remaining part or all of the protocol layer functions are distributed in the DU and centrally controlled by the CU. However, this is not the only possibility.

[0097] In some embodiments, the core network equipment can be a single device, multiple devices, or a group of devices. The device can be virtual or physical. The core network includes, for example, at least one of the Evolved Packet Core (EPC), 5G Core Network (5GCN), and Next Generation Core (NGC).

[0098] In some embodiments, the communication system 100 described above may be a 4G communication system, a 5G communication system, or a 6G communication system. It should be noted that the communication system 100 may also be other communication systems, and this disclosure does not specifically limit them.

[0099] The integrated space-air-ground network for 6G (6th generation) refers to using satellite communication networks as an important supplement and extension to terrestrial communication networks. By deeply integrating space-based, air-based, and terrestrial network resources, it constructs a globally covering, ubiquitous, high-speed, intelligent, secure, and reliable communication network. This network features wide coverage, flexible deployment, ultra-low power consumption, ultra-high precision, and resilience to ground-based disasters, providing users with high-quality, efficient, and intelligent communication services. The integrated space-air-ground network will fully utilize aerial resources such as satellites in different orbits, drones, and high-altitude platforms, as well as terrestrial cellular mobile networks, the Internet of Things (IoT), and cloud computing technologies to achieve a new converged architecture with multiple layers, connections, and access. This network will support the communication needs of high-speed mobility, high capacity, low latency, and high reliability, meeting the communication network requirements of fields such as the IoT, industrial internet, smart manufacturing, and intelligent transportation.

[0100] Figure 2 is a schematic diagram of the architecture of an integrated air-space-ground network. As shown in Figure 2, the network architecture of an integrated air-space-ground network can include three parts: a space-based network, a space-based network, and a ground-based network. The space-based network can consist of multiple satellites, including geostationary satellites, medium-Earth orbit / low-Earth orbit satellites, and relay satellites. These satellites will form a multi-layered, multi-connected, multi-source data transmission and processing system, providing high-speed, reliable, and continuous communication services to global users. The space-based network can consist of equipment such as drones and high-altitude platforms, which can be deployed in the air to provide flexible and efficient communication services to ground users. At the same time, the space-based network can also cooperate with the space-based network to achieve wider coverage and higher communication performance. The ground-based network can consist of ground base stations, IoT devices, etc., which can achieve direct connection and data transmission with users. The ground-based network will work in conjunction with the space-based and space-based networks to form a seamless global communication network.

[0101] In some embodiments, as an important component of the air-based network in an integrated air-space-ground network, unmanned aerial vehicles (UAVs) may need to perform complex tasks based on multimodal data. In this case, UAVs may face the following challenges:

[0102] (1) UAVs have limited storage space and computing power, making it difficult to process large amounts of multimodal data.

[0103] (2) Data in different modalities of multimodal data exhibits strong heterogeneity. For example, some modal data correspond to physical signals, and these physical signals are continuous signals. Similarly, data processed or received by a communication system typically corresponds to discrete signals. The subspaces inhabited by the feature vectors of data from different modalities may not overlap. In some cases, data with the same or similar semantics in different modalities may have completely different feature vectors.

[0104] (3) UAVs can perform tasks based on received multimodal data. During the transmission of multimodal data, the signals carrying the multimodal data may be interfered with. This may affect the completion of the task.

[0105] Therefore, how to accurately determine the task to be executed based on multimodal data using limited storage and computing resources is an urgent problem to be solved.

[0106] In the following text, important terms used in the embodiments of this disclosure will be explained.

[0107] Semantic communication:

[0108] Communication can be divided into three levels: syntactic, semantic, and pragmatic. Correspondingly, the information transmitted during communication can be divided into three levels: syntactic information, semantic information, and pragmatic information. Syntactic information transmission focuses on the accurate transmission of communication symbols. Semantic information transmission focuses on the precise expression of meaning by the transmitted symbols. Pragmatic information transmission focuses on the effective impact of the received information on behavior.

[0109] Semantic communication is a task-oriented communication method that follows a "understand first, then transmit" approach. It involves selectively extracting, compressing, and transmitting features from the original signal, and then using semantic information for communication.

[0110] Figure 3 is a schematic diagram of the semantic communication system model. As shown in Figure 3, semantic communication can be implemented between the sender 301 and the receiver 302. The sender 301 includes a semantic information extraction module 3011, a semantic encoding module 3012, and a channel encoding module 3013. The receiver 302 may include a semantic information recovery module 3021, a semantic decoding module 3022, and a channel decoding module 3023.

[0111] At the sender 301, the semantic information extraction module 3011 can extract features from the source information to obtain semantic features; the semantic coding module 3012 and the channel coding module 3013 can sequentially perform semantic coding and channel coding on the semantic features. The encoded symbols carrying semantic information can be sent to the receiver 302. At the receiver 302, the channel decoding module 3023 and the semantic decoding module 3022 can sequentially perform channel decoding and channel decoding on the received symbols to obtain semantic features; then, the semantic information recovery module 3021 can perform semantic recovery based on the semantic features to obtain the source information.

[0112] In some embodiments, semantic communication between sender 301 and receiver 302 can be implemented based on knowledge base 303. Knowledge base 303 can be deployed in both sender 301 and receiver 302. In some embodiments, the processing of at least one of semantic information extraction module 3011, semantic encoding module 3012, and channel encoding module 3013 can be implemented based on knowledge base 303. In some embodiments, the processing of at least one of semantic information recovery module 3021, semantic decoding module 3022, and channel decoding module 3023 can be implemented based on knowledge base 303. Knowledge base 303 in sender 301 and receiver 302 can be synchronized to maintain consistency. Based on knowledge base 303, accurate transmission and reconstruction of semantic information can be achieved between sender 301 and receiver 302.

[0113] TinyML (Micro Machine Learning) model:

[0114] TinyML refers to the technology of running machine learning models on microcontroller units (MCUs) or other resource-constrained hardware platforms. Traditional machine learning models typically require powerful computing resources and storage space, while TinyML aims to compress and optimize these models to fit into small devices with limited memory, computing power, and energy consumption. TinyML's core goal is to extend AI technology to Internet of Things (IoT) devices, enabling them to possess intelligent sensing and processing capabilities, thereby achieving broader applications.

[0115] TinyML's technical architecture mainly includes three stages: model training, model compression, and model deployment.

[0116] Model training: Model training is performed using powerful computing resources (such as cloud servers or high-performance computers). Common training algorithms include Convolutional Neural Networks (CNNs), Recurrent Neural Networks (RNNs), and deep learning. After training, the resulting model is usually large in size and not suitable for direct deployment on embedded devices.

[0117] Model compression: To adapt models to resource-constrained devices, they need to be compressed and optimized. Common model compression techniques include model pruning, quantization, and distillation. These techniques can significantly reduce the number of model parameters and computational complexity, thereby reducing memory usage and energy consumption.

[0118] Model Deployment: The compressed model is deployed onto embedded devices. This stage requires consideration of the hardware platform's characteristics, such as computing power, memory size, and power consumption. Commonly used embedded hardware platforms include the ARM Cortex-M series, RISC-V, and custom AI acceleration chips. To improve model execution efficiency, hardware acceleration technologies such as Graphics Processing Units (GPUs) and Field Programmable Gate Arrays (FPGAs) can also be utilized.

[0119] Figure 4 is a flowchart illustrating a task determination method according to an embodiment of the present disclosure. The task determination method described in this embodiment can be applied to a first device 101. As shown in Figure 4, the task determination method of this embodiment includes steps S401 to S404.

[0120] In step S401, the first device 101 acquires multimodal data.

[0121] In some embodiments, the first device 101 may receive multimodal data. For example, the first device 101 may receive multimodal data sent by the second device 102. For example, the first device 101 may receive multimodal data sent by the third device 103. For example, the first device 101 may receive multimodal data sent by one or more other first devices 101. In some embodiments, the first device 101 may obtain multimodal data locally.

[0122] In some embodiments, multimodal data may include multiple data.

[0123] In some embodiments, multiple data points in multimodal data may belong to two or more types.

[0124] In some embodiments, multimodal data may be obtained based on at least one of the following: communication signals, sensor signals, power supply signals, and task calculation results.

[0125] In some embodiments, multimodal data may include data obtained based on communication signals.

[0126] In some embodiments, the data type in multimodal data may include at least one of the following: image data, video data, text data, and audio data.

[0127] In some embodiments, the communication signal may be a received signal. For example, the communication signal may include a signal transmitted by the first device 101 and / or the second device 102.

[0128] In some embodiments, multimodal data may include data obtained based on sensor signals.

[0129] In some embodiments, the first device 101 may be equipped with one or more sensors.

[0130] In some embodiments, the sensors of the first device 101 may include at least one of the following: a temperature sensor, a barometer, radar, a camera, a microphone, and an accelerometer. In some embodiments, the sensing signals may include at least one of the following: temperature values, air pressure values, radio signals, visual signals, sound signals, and acceleration values. It is understood that the first device 101 may also include other types of sensors, and this disclosure does not specifically limit them.

[0131] In some embodiments, the data types in the multimodal data may include at least one of the following: temperature data, air pressure data, distance measurement data, video data, audio data, and acceleration data.

[0132] In some embodiments, multimodal data may include data obtained based on power supply signals.

[0133] In some embodiments, the power supply signal may be a signal for charging the first device 101. In some embodiments, the power supply signal may be a communication signal. In some embodiments, the power supply signal may be a sensing signal.

[0134] In some embodiments, multimodal data may include data obtained based on task computation results.

[0135] In some embodiments, the task calculation result may include task support information. For example, the first device 101 may generate task support information through calculation.

[0136] In some embodiments, the task calculation result may include communication mode, sensing mode, and computational consumption. It is understood that the task calculation result may also include other data, and this disclosure does not impose specific limitations on the embodiments described.

[0137] In some embodiments, the communication mode and the sensing mode can be determined based on computational operations such as feature extraction, data compression, data pattern recognition, and simplified source coding.

[0138] In some embodiments, the first device 101 in communication mode can perform data communication.

[0139] In some embodiments, the first device 101 in sensing mode can realize environmental sensing.

[0140] In some embodiments, the first device 101 may be in a communication mode, a sensing mode, or both a communication mode and a sensing mode.

[0141] In some embodiments, the first device 101 may determine the computational cost of the first task based on its own computing power.

[0142] In some embodiments, the computational consumption of the first task may include at least one of the following: computing power consumption, storage space consumption, communication resource consumption, and energy consumption. Computing power consumption may refer to the computing power occupied by the first task when executed in the first device 101. Storage space consumption may refer to the storage space occupied by the first task when executed in the first device 101. Communication resource consumption may refer to the communication resources occupied by the first task when executed in the first device 101. Energy consumption may refer to the energy consumed by the first device 101 when the first task is executed in the first device 101. In one example, the first device 101 may rely on electrical energy to operate, then energy consumption may be electrical power consumption.

[0143] Figure 5 is a schematic diagram of acquiring multimodal data according to an embodiment of this disclosure. As shown in Figure 5, the source of multimodal data may include at least one of the following: communication signal Scomm, sensing signal Ssen, power supply signal Spow, and calculation result Scomp.

[0144] In some embodiments, after the communication signal Scomm is received by the receiver of the first device 101, it can be input into a filter amplifier circuit to perform filtering and amplification processing; then, the filtered and amplified communication signal Scomm can be input into an analog-to-digital converter to generate a corresponding digital signal.

[0145] In some embodiments, after the sensing signal Ssen is received by the receiver of the first device 101, it can be input into a filtering and amplification circuit to perform filtering and amplification processing; then, the filtered and amplified sensing signal Ssen can be input into an analog-to-digital converter to generate a corresponding digital signal.

[0146] In some embodiments, the power supply signal Spow can be implemented by the communication signal Scomm and / or the sensing signal Ssen, or it can be independent of the communication signal Scomm and the sensing signal Ssen. After receiving the power supply signal Spow, the first device 101 can determine whether the power supply signal Spow is an information signal. If so, the power supply signal Spow can be input to a filter amplifier circuit for filtering and amplification; then, the filtered and amplified power supply signal Spow can be input to an analog-to-digital converter to generate a corresponding digital signal. If not, the power supply signal Spow can be used to charge the battery of the first device 101. It is understood that if the power supply signal Spow is determined to be an information signal, then the power supply signal Spow can be the communication signal Scomm or the sensing signal Ssen.

[0147] In some embodiments, the communication signal Scomm, the sensing signal Ssen, and the power supply signal Spow can be encoded digital signals. The digital signal obtained based on the communication signal Scomm, the sensing signal Ssen, and the power supply signal Spow can then be input to a decoder for decoding. In some embodiments, the sensing signal Ssen can be directly input to a demodulator for demodulation. In some embodiments, the demodulated sensing signal Ssen can also be input to a decoder for decoding.

[0148] In some embodiments, based on at least one of the communication signal Scomm, the sensing signal Ssen, the power supply signal Spow, and local data of the first device 101, the first device 101 can perform task calculations for a first task to obtain a calculation result Scomp. The calculation result Scomp obtained through this task calculation can be task support information.

[0149] In some embodiments, the digital signals obtained from the communication signal Scomm and / or the sensing signal Ssen and / or the power supply signal Spow, and / or the calculation result Scomp, can be used to obtain multimodal data corresponding to the first task of the first device 101. This multimodal data includes heterogeneous multimodal data components, namely S... k ∈{Scomm, Ssen, Spow, Scomp}. In other words, k∈M, M={comm, sen, pow, comp}.

[0150] In some embodiments, an information vector can be obtained based on each component of the multimodal data. In some embodiments, information vectors It can be obtained through a conversion function, i.e.

[0151] In some embodiments, multimodal data may include information vectors obtained from each component. In some embodiments, the information vector is obtained from at least one of the communication signal Scomm, the sensing signal Ssen, the power supply signal Spow, and the calculation result Scomp. It can form multimodal data.

[0152] In some embodiments, after obtaining the multimodal data, the multimodal data can be used by the first device 101 to determine the language intent. The determined language intent may correspond to a first task.

[0153] In step S402, the first symbol is determined by the encoder based on the multimodal data.

[0154] In some embodiments, the first device 101 can input the obtained multimodal data into the encoder to obtain a first symbol.

[0155] In some embodiments, the first symbol may contain semantic information about the multimodal data.

[0156] In some embodiments, the multimodal data obtained by the first device 101 in step S401 may include multiple data sets. These data sets may belong to at least two types. In other words, these data sets may belong to at least two modalities.

[0157] In some embodiments, the multimodal data input to the encoder may include multiple information vectors, each of which can be used to characterize the semantic information of one data point in the multimodal data. It is understood that in the multimodal data, the number of data types can be greater than one, and the number of data points of each type can be greater than or equal to one.

[0158] Figure 6 is a schematic diagram of the framework of a virtual semantic communication system provided according to an embodiment of the present disclosure. As shown in Figure 6, the virtual semantic communication system 600 may include an encoder 610, a decoder 620, and a semantic matching module 630. In some embodiments, the encoder 610 may be used to implement encoding processing in semantic communication of multimodal data. In some embodiments, the encoder 610 may include a feature extraction module 611, a tensor generation module 612, and a channel coding module 613. In some embodiments, the decoder 620 may be used to implement decoding processing in semantic communication of multimodal data. In some embodiments, the decoder 620 may include a channel decoding module 621, a semantic decoding module 622, and a semantic fusion module 623. In some embodiments, the semantic matching module 630 may be used to match the intended meaning of the language.

[0159] In some embodiments, the virtual semantic communication system 600 may further include an estimation module 640. The estimation module 640 may be used to perform channel estimation of the semantic communication channel.

[0160] In some embodiments, at least one of the encoder 610, decoder 620, semantic matching module 630, and estimation module 640 may be implemented based on a first model. In some embodiments, the first model may include at least one of the following: an AI model and an ML model. In one example, the first model may be a first AI model. In one example, the first model may be a first ML model. In some embodiments, the first ML model may include a TinyML model.

[0161] In some embodiments, the first model may be implemented based on a transformer model. In some embodiments, the first model may be implemented based on a recurrent neural network (RNN).

[0162] In some embodiments, step S402 may include: extracting features from each of the multiple data in the multimodal data to obtain a first feature; performing a transformation process on the first feature of each data to obtain a second feature; and performing channel coding on the second feature to obtain a first symbol corresponding to the second feature.

[0163] In some embodiments, the first device 101 may input each data point in the multimodal data into the feature extraction module 611 for feature extraction to obtain a first feature. In some embodiments, the first device 101 may perform feature extraction on the information vector of each data point in the multimodal data to obtain the first feature. In some embodiments, the first feature may be used to characterize the corresponding data. For example, the first feature may include the eigenvector of the data.

[0164] In some embodiments, the number of first features obtained by the first device 101 based on the multimodal data may be equal to the number of data contained in the multimodal data. It is understood that the first features and the data in the multimodal data may have a one-to-one correspondence.

[0165] In some embodiments, the first feature may be a one-dimensional vector or a multi-dimensional array. In some embodiments, the dimensions of the first features of different data in multimodal data may be the same or different. In some embodiments, the dimensions of the first features of data in different modalities may be the same or different.

[0166] In some embodiments, after obtaining the first feature, the first device 101 can input the first feature into the tensor generation module 612 for transformation processing to obtain the second feature. In some embodiments, the first device 101 can perform transformation processing on the first feature to obtain the second feature.

[0167] In some embodiments, each second feature may be determined for a corresponding first feature. It is understood that there may be a one-to-one correspondence between the second features and the first features.

[0168] In some embodiments, the second features obtained from multiple data points of the multimodal data may have the same dimension. In some embodiments, the second features of multiple data points of the multimodal data may all be one-dimensional vectors. For example, the second features obtained by the tensor generation module 612 may be one-dimensional semantic vectors. In some embodiments, the second features of multiple data points of the multimodal data may all be two-dimensional vectors.

[0169] In some embodiments, the first device 101 may perform channel coding on the second feature to obtain a first symbol corresponding to the second feature. In some embodiments, the first device 101 may use the channel coding module 613 to map the second feature onto a symbol in the semantic channel to obtain the first symbol. In some embodiments, the number of first symbols corresponding to the second feature may be one or more.

[0170] It is understandable that the first device 101 can obtain the first symbol based on multimodal data through the feature extraction module 611, the tensor generation module 612, and the channel coding module 613.

[0171] In some embodiments, the first symbol of the multimodal data obtained by the first device 101 can be transmitted in a semantic communication channel. In some embodiments, the task determination method can be executed within the first device 101, in which case the semantic communication channel for transmitting the first symbol can be a "virtual" semantic communication channel. In some embodiments, the task determination method can be executed between two or more first devices 101, in which case the semantic channel for transmitting the first symbol can be a real semantic communication channel, i.e., a semantic communication channel between two first devices 101.

[0172] In step S403, the semantic intent is determined by the decoder based on the first symbol.

[0173] In some embodiments, the first symbol generated in step S402, after being transmitted through the semantic communication channel, can be input to the decoder 620 for decoding.

[0174] In some embodiments, during the process of receiving multimodal data, the signal carrying the multimodal data may be interfered with during transmission. In this case, the multimodal data obtained by the first device 101 may contain interference information and / or noise information. For example, the multimodal data may include communication interference, noise, and environmental interference. In some embodiments, the multimodal data may include redundancy information related to the first task. For example, the multimodal data may include residual information related to the first task. This interference information, noise information, and redundancy information can all be considered as interference information in the multimodal data.

[0175] In some embodiments, the interference information present in the multimodal data can be regarded as channel noise introduced into the semantic communication channel. In some embodiments, the first symbol obtained by the decoder 620 can be regarded as containing semantic information (i.e., the true content part) and interference information (i.e., the channel noise part).

[0176] In some embodiments, step S403 may include: performing channel decoding on the first symbol to obtain multiple third features; performing semantic decoding on each third feature to obtain semantic information; and performing fusion processing on the multiple semantic information of the multimodal data to determine the semantic intent.

[0177] In some embodiments, the first symbol can be channel-decoded by the channel decoding module 621 to obtain the third feature. In some embodiments, the channel decoding module 621 can perform channel decoding on one or more first symbols corresponding to each second feature to obtain the third feature corresponding to that second feature. In some embodiments, each third feature can correspond to a piece of data in the multimodal data.

[0178] In some embodiments, the third feature can be input into the semantic decoding module 622 for semantic decoding to obtain semantic information. In some embodiments, the semantic decoding module 622 can perform semantic decoding on each third feature to obtain the semantic information of the data corresponding to that third feature.

[0179] In some embodiments, multiple semantic information pieces can be input into the semantic fusion module 623 for fusion processing to determine the semantic intent. In some embodiments, the semantic fusion module 623 can perform fusion processing on the semantic information corresponding to all third features to determine the semantic intent.

[0180] It is understood that, through the channel decoding module 621, the semantic decoding module 622, and the semantic fusion module 623, the first device 101 can obtain the semantic intent based on the first symbol. This semantic intent can correspond to multimodal data.

[0181] In some embodiments, the task determination method can be executed among two or more first devices 101, in which case the estimation module 640 can perform channel estimation on the semantic communication channel for receiving a first symbol from another first device 101. After the channel estimation is completed, the first symbol can be input to the channel decoding module 621.

[0182] In step S404, the linguistic intent is determined based on the semantic intent.

[0183] In some embodiments, when the semantic intent corresponding to the multimodal data is obtained, the first device 101 can determine the terminology intent based on the semantic intent.

[0184] In some embodiments, the first device 101 may determine the phrasing intent that matches the semantic intent through the semantic matching module 630.

[0185] In some embodiments, the determination of language intent can be achieved through at least one of the following methods: a language intent database, a first model.

[0186] In some embodiments, the first device 101 can determine the terminology intent corresponding to the semantic intent based on a terminology intent database. In some embodiments, the terminology intent database may include one or more terminology intents. In some embodiments, the terminology intent database may also include one or more semantic intents corresponding to each terminology intent. It is understood that the terminology intent database may include the association between terminology intents and semantic intents. Thus, the terminology intent database allows for matching between semantic intents and terminology intents. In this case, the first device 101 can search for a matching terminology intent from the terminology intent database based on the decoded semantic intent.

[0187] In some embodiments, the language intent database may be pre-configured. For example, the language intent database may be pre-configured in the first device 101 by the manufacturer or operator of the first device 101.

[0188] In some embodiments, the first device 101 may determine the terminology intent corresponding to the semantic intent based on the first model. In some embodiments, the semantic matching module 630 may include the first model. In some embodiments, the first device 101 may input the semantic intent into the first model of the semantic matching module 630 to obtain the corresponding terminology intent.

[0189] In some embodiments, the terminology intent may be used to indicate a first task. In some embodiments, the terminology intent may be used by the first device 101 to determine the first task.

[0190] The task determination method of this disclosure embodiment can be implemented through steps S401 to S404.

[0191] It should be noted that steps S401 to S404 can be executed by a single first device 101 or by different first devices 101 respectively. In one example, one first device 101 can execute steps S401 to S404. In this case, the first symbol can be transmitted in a virtual semantic transmission channel. In another example, two first devices 101 can execute steps S401 to S404 separately. For example, one first device 101 can execute steps S401 and S402, and the other first device 101 can execute steps S403 and S404. In this case, the first symbol can be transmitted in a semantic transmission channel between the two first devices 101.

[0192] The task determination method involved in the embodiments of this disclosure may include at least one of steps S401 to S404. For example, step S402 may be implemented as a separate embodiment. For example, step S403 may be implemented as a separate embodiment. For example, a combination of steps S401 and S402 may be implemented as a separate embodiment. For example, a combination of steps S402 and S403 may be implemented as a separate embodiment. For example, a combination of steps S403 and S404 may be implemented as a separate embodiment. For example, a combination of steps S401, S402, and S403 may be implemented as a separate embodiment. For example, a combination of steps S402, S403, and S404 may be implemented as a separate embodiment. For example, a combination of steps S401, S402, S403, and S404 may be implemented as a separate embodiment. It should be noted that the possible separate embodiments consisting of one or more steps S401 to S404 are not limited thereto.

[0193] In some embodiments, steps S401, S403, and S404 are optional, and one or more of these steps may be omitted or substituted in different embodiments.

[0194] In some embodiments, other optional implementations may be described before or after the specification corresponding to Figure 4.

[0195] In some embodiments, the names of information, etc., are not limited to the names described in the embodiments. Terms such as "information", "message", "signal", "signaling", "report", "configuration", "indication", "instruction", "command", "channel", "parameter", "domain", "field", "symbol", "symbol", "codebook", "codeword", "codepoint", "bit", "data", "program", and "chip" can be used interchangeably.

[0196] In some embodiments, the terms “radio”, “wireless”, “radio access network (RAN)”, “access network (AN)”, and “RAN-based” can be used interchangeably.

[0197] In some embodiments, terms such as “moment,” “point in time,” “time,” and “time location” can be used interchangeably, as can terms such as “duration,” “segment,” “time window,” “window,” and “time.”

[0198] In some embodiments, “get,” “obtain,” “receive,” “transmit,” “bidirectional transmission,” and “send and / or receive” can be used interchangeably and can be interpreted as receiving from other entities, obtaining from protocols, obtaining from higher layers, obtaining through self-processing, or autonomous implementation, among other meanings.

[0199] In some embodiments, terms such as “send,” “transmit,” “report,” “distribute,” “transfer,” “bidirectional transmission,” “send and / or receive” can be used interchangeably.

[0200] In some embodiments, terms such as "certain", "preset", "default", "set", "indicated", "a certain", "any", and "first" can be used interchangeably. "Certain A", "preset A", "default A", "set A", "indicated A", "a certain A", "any A", and "first A" can be interpreted as A pre-defined in a protocol or the like, or as A obtained through setting, configuration, or instruction, or as specific A, a certain A, any A, or first A, but are not limited thereto.

[0201] In some embodiments, the determination or judgment can be made by a value represented by 1 bit (0 or 1), or by a true or false value (boolean), or by a comparison of numerical values ​​(e.g., a comparison with a predetermined value), but is not limited thereto.

[0202] In some embodiments, terms such as "decoding" and "decode" can be used interchangeably.

[0203] In some embodiments, the terms "vector", "tensor", etc., can be used interchangeably.

[0204] In some embodiments, the terms "encoding module", "encoder", "encoding circuit" and other terms can be used interchangeably, as can the terms "decoding module", "decoder", "decoding circuit" and other terms.

[0205] Figure 7A is an interactive schematic diagram illustrating the task execution implemented according to the task determination method provided in this embodiment of the present disclosure. The task execution process involved in this embodiment of the present disclosure can be applied to the communication system 100. As shown in Figure 7A, the task execution process of this embodiment of the present disclosure may include steps S7101 to S7103.

[0206] In step S7101, the second device 102 sends the first information to the first device 101.

[0207] In some embodiments, the second device 102 may send the first information. In some embodiments, the first information may be sent by the second device 102, but is not limited thereto, and may also be sent by other entities.

[0208] In some embodiments, the first device 101 may receive first information. In some embodiments, the first information may be received by the first device 101, but is not limited thereto, and may also be received by other entities.

[0209] In some embodiments, the first information may be used by the first device 101 to perform a first task. In some embodiments, the first information may be used to instruct the first device 101 to perform a first task. In some embodiments, the first information may be used to trigger the first device 101 to perform a first task.

[0210] In some embodiments, the first task may include at least one of the following: a communication task, a sensing task, and a charging task. It is understood that the first task may also include other tasks, and this disclosure does not impose specific limitations on the embodiments described.

[0211] In some embodiments, the communication task may be that the first device 101 provides communication services. For example, the first device 101 may provide communication services to one or more terminals.

[0212] In some embodiments, the sensing task may be environmental sensing performed by the first device 101. For example, the first device 101 may use one or more sensors for sensing.

[0213] In some embodiments, the charging task may be performed by the first device 101. For example, the first device 101 may charge itself using a power supply signal.

[0214] In step S7102, the first device 101 determines the first task.

[0215] The optional implementation of step S7102 can be found in the optional implementation of steps S401 to S404 in Figure 4, as well as other related parts in the embodiments involved in Figure 4, which will not be repeated here.

[0216] In some embodiments, the first device 101 may determine a first task in response to first information.

[0217] In some embodiments, the first task may be determined by the steps in the task determination method shown in FIG4.

[0218] In some embodiments, multimodal data may include at least one of the following: data obtained from the second device 102, data obtained from the first device 101 itself, and data obtained from one or more other first devices 101.

[0219] In step S7103, the first device 101 performs the first task.

[0220] In some embodiments, the first device 101 may perform a first task determined based on multimodal data.

[0221] Through steps S7101 to S7103, the execution of the first task can be achieved based on the task determination method of this disclosure embodiment.

[0222] It should be noted that the number of first devices 101 can be one or more. In some embodiments, when there are multiple first devices 101, each first device 101 can execute steps S7101 to S7103. In some embodiments, multiple first devices 101 can share the multimodal data they have obtained. For example, in step S7102, each first device 101 can send its obtained multimodal data to other first devices 101. In some embodiments, in step S7101, the second device 102 can send the first information to each first device 101 by means of broadcasting, grouping, or individual transmission; this disclosure does not specifically limit this.

[0223] The task determination method involved in the embodiments of this disclosure may include at least one of steps S7101 to S7103. For example, step S7102 may be implemented as a standalone embodiment. For example, a combination of steps S7102 and S7103 may be implemented as a standalone embodiment. For example, a combination of steps S7101, S7102, and S7103 may be implemented as a standalone embodiment. It should be noted that the possible standalone embodiments consisting of one or more steps S7101 to S7103 are not limited thereto.

[0224] In some embodiments, steps S7101 and S7103 are optional, and one or more of these steps may be omitted or substituted in different embodiments.

[0225] Figure 7B is an interactive schematic diagram illustrating model training implemented according to the task determination method provided in this embodiment of the present disclosure. The model training process involved in this embodiment of the present disclosure can be applied to the communication system 100. As shown in Figure 7B, the model training process in this embodiment of the present disclosure may include steps S7201 to S7207.

[0226] In step S7201, the third device 103 sends the second information to the first device 101.

[0227] In some embodiments, the third device 103 may send the second information. In some embodiments, the second information may be sent by the third device 103, but is not limited thereto, and may also be sent by other entities.

[0228] In some embodiments, the first device 101 may receive the second information. In some embodiments, the second information may be received by the first device 101, but is not limited thereto, and may also be received by other entities.

[0229] In some embodiments, the second information may be used by the first device 101 to perform training on the first model. In some embodiments, the second information may be used to instruct the first device 101 to train the first model. In some embodiments, the first information may be used to trigger the first device 101 to train the first model.

[0230] In some embodiments, the second information may include multimodal data. In some embodiments, the multimodal data included in the second information may be used for training the first model. In some embodiments, the multimodal data in the second information may be used to implement training of the first model on the first device 101. In some embodiments, the multimodal data included in the second information may be training sample data of the first model.

[0231] In some embodiments, the number of first devices 101 can be one or more. A third device 103 can send second information to one or more first devices 101.

[0232] In some embodiments, the second information may include the first model to be trained. For example, the second information may include the model parameters of the initial first model.

[0233] It is understood that step S7201 is optional. In some embodiments, step S7201 may not be performed, and the first device 101 may autonomously perform training on the first model.

[0234] In step S7202, the first device 101 determines the first task.

[0235] The optional implementation of step S7202 can be found in the optional implementation of steps S401 to S404 in Figure 4, as well as other related parts in the embodiments involved in Figure 4, which will not be repeated here.

[0236] In some embodiments, the first device 101 may determine the first task in response to the second information.

[0237] In some embodiments, the first task may be determined by the steps in the task determination method shown in FIG4.

[0238] In some embodiments, multimodal data may include at least one of the following: data obtained from a third device 103, data obtained from the first device 101 itself, and data obtained from one or more other first devices 101.

[0239] In step S7203, the first device 101 determines the loss value.

[0240] In some embodiments, based on the first task determined in step S7202, the first device 101 may determine the loss value of the loss function.

[0241] In some embodiments, the loss function used to determine the loss value can be used in classification scenarios.

[0242] In some embodiments, the loss function may include a cross-entropy (CE) loss function, a log loss function, or other types of loss functions. In one example, the cross-entropy loss function may include a binary cross-entropy loss function or a multi-class cross-entropy loss function.

[0243] In step S7204, the first device 101 adjusts the parameters of the first model.

[0244] In some embodiments, the first device 101 may determine whether to adjust the parameters of the first model based on the loss value.

[0245] In some embodiments, the first device 101 may determine whether a first condition is met. It is understood that the first condition can be used to determine whether the training of the first model is complete. In other words, the first condition can be used to determine whether the performance of the first model meets the requirements. In some embodiments, if the first condition is met, it indicates that the performance of the first model does not meet the requirements; conversely, if the first condition is not met, it indicates that the performance of the first model meets the requirements. In some embodiments, if the first condition is met, the first device 101 may adjust the parameters of the first model; if the first condition is not met, the first device 101 may not adjust the parameters of the first model.

[0246] In some embodiments, the first condition may include: the loss value is greater than a threshold. In one example, the first condition is met if the loss value is greater than the threshold, or if the loss value is greater than or equal to the threshold. In another example, the first condition is not met if the loss value is less than the threshold, or if the loss value is less than or equal to the threshold.

[0247] In some embodiments, the threshold used to determine whether the first condition is met may be preset. In some embodiments, the threshold used to determine whether the first condition is met may be sent to the first device 101 via second information.

[0248] It is understood that the first model can be trained in the first device 101 through steps S7202 to S7204. The first device 101 can continuously adjust the parameters of the first model by repeatedly executing steps S7202 to S7204 based on the obtained multimodal data until the performance of the first model meets the requirements.

[0249] In step S7205, the first device 101 sends third information to the third device 103.

[0250] In some embodiments, the first device 101 may send third information. In some embodiments, the third information may be sent by the first device 101, but is not limited thereto, and may also be sent by other entities.

[0251] In some embodiments, the third device 103 may receive third information. In some embodiments, the third information may be received by the third device 103, but is not limited thereto, and may also be received by other entities.

[0252] In some embodiments, the third information may be used to indicate that the training of the first model has been completed.

[0253] In some embodiments, the third information may be used to return the trained first model to the third device 103.

[0254] In some embodiments, the third information may include a trained first model. In one example, the third information may include model parameters of the trained first model in the first device 101. For example, the third information may include all model parameters. For example, the third information may include parameters that have changed among the model parameters.

[0255] In some embodiments, when the number of first devices 101 is greater than one, a third device 103 may receive third information from multiple first devices 101. In some embodiments, the multiple pieces of third information received by the third device 103 from the multiple first devices 101 may be the same or different. For example, the parameters and parameter values ​​of the first model contained in the third information sent by different first devices 101 may be the same. For example, the parameters and / or parameter values ​​of the first model contained in the third information sent by different first devices 101 may be at least partially different.

[0256] In step S7206, the third device 103 determines the parameters of the first model.

[0257] In some embodiments, the third device 103 may determine the parameters of the first model based on third information from one or more first devices 101.

[0258] In some embodiments, the third device 103 may aggregate model parameters indicated by third information from multiple first devices 101 to determine model parameters of the first model. In one example, for the same parameter in the first model, the third device 103 may perform a weighted average of the parameter values ​​sent by the multiple first devices 101 to determine the parameter value. For example, the weights of the parameter values ​​sent by the multiple first devices 101 in the weighted average may be the same or different.

[0259] It should be noted that after the third device 103 determines the parameters of the first model based on the third information, it can determine whether the first model meets the predetermined conditions. If the predetermined conditions are met, the training of the first model is stopped; otherwise, steps S7201 to S7206 are continued.

[0260] In some embodiments, the predetermined conditions may include at least one of the following: the number of iterations in steps S7201 to S7206 reaches a preset threshold, and the first model meets the performance requirements.

[0261] In some embodiments, the preset threshold for the number of iterations can be set based on experience. In some embodiments, the preset threshold for the number of iterations can be, for example, 10 times, 100 times, 1000 times, etc.

[0262] In some embodiments, the performance requirement of the first model may include achieving a task recognition accuracy that reaches a preset threshold. For example, the accuracy of task recognition using the first model may reach 99%, 99.99%, etc.

[0263] In step S7207, the third device 103 sends the fourth information to the first device 101.

[0264] In some embodiments, the third device 103 may send fourth information. In some embodiments, the fourth information may be sent by the third device 103, but is not limited thereto, and may also be sent by other entities.

[0265] In some embodiments, the first device 101 may receive fourth information. In some embodiments, the fourth information may be received by the first device 101, but is not limited thereto, and may also be received by other entities.

[0266] In some embodiments, after determining that the training of the first model has ended, the third device 103 may send fourth information to the first device 101.

[0267] In some embodiments, the fourth information may be used to indicate that the training of the first model has been completed.

[0268] In some embodiments, the fourth information may include the trained first model. For example, the fourth information may include the model parameters of the trained first model.

[0269] In some embodiments, the training of the first model can be performed independently and locally on the first device 101. In this case, the first device 101 can complete the local training of the first model by executing steps S7202 to S7204. In some embodiments, the first model can be trained by each first device 101. In some embodiments, the first model can be trained by one first device 101. After the first device 101 completes the training of the first model, it can send the trained first model to other first devices 101.

[0270] In some embodiments, the training of the first model can be based on federated learning. In this case, through steps S7201 to S7206, the training of the first model can be completed through federated learning between multiple first devices 101 and third device 103. After the training of the first model is completed, the third device 103 can deploy the trained first model to each of the first devices 101.

[0271] It should be noted that the training of the first model can also be based on split learning. Partial training of the first model can be performed in the first device 101, and partial training of the first model can be performed in the third device 103.

[0272] Through steps S7201 to S7207, the execution of the first task can be achieved based on the task determination method of this embodiment.

[0273] The task determination method involved in the embodiments of this disclosure may include at least one of steps S7201 to S7207. For example, step S7202 may be implemented as a standalone embodiment. For example, a combination of steps S7202, S7203, and S7204 may be implemented as a standalone embodiment. It should be noted that the possible standalone embodiments consisting of one or more steps S7201 to S7207 are not limited thereto.

[0274] In some embodiments, steps S7201, S7203, S7204, S7205, S7206, and S7207 are optional, and one or more of these steps may be omitted or substituted in different embodiments.

[0275] In the following, the technical solutions of the embodiments of this disclosure will be described by way of specific implementation.

[0276] This disclosure proposes viewing the successful execution of advanced UAV missions as a process of semantic recognition and pragmatic execution. TinyML provides advanced UAV algorithms and models that can run on low-power and resource-constrained platforms. From a semantic communication perspective, leveraging the applicability of TinyML (i.e., the first model) to UAVs, heterogeneous multimodal communication and UAV mission execution processes are mapped, aiming to better utilize the capabilities of machine learning and semantic communication to enhance the practical mission execution capabilities of UAVs.

[0277] In some embodiments, multimodal virtual semantic communication can provide task-related auxiliary information, enabling complementary integration of multiple independent modalities within the task domain. The proposed scheme and model can achieve deep fusion of communication, perception, and computation, ultimately improving the practical mission execution capabilities of unmanned aerial vehicle (UAV) systems.

[0278] In this embodiment, starting from the concept of semantic communication, the input and processing of multimodal information are regarded as a virtual semantic source of communication. Simultaneously, an intent-driven network with features such as flexible reconfigurability and adaptive policy optimization is proposed to improve network operating efficiency.

[0279] In some embodiments, within a virtual semantic source, the original multimodal input is converted into multimedia data suitable for TinyML processing to obtain a comprehensive semantic vector of planar information, thereby achieving multimodal interoperability. Furthermore, biases in multimodal perception, transmission, and analysis are unified as semantic noise at the receiver of semantic communication. Additionally, the subsequent processes of UAV mission execution are treated as a virtual semantic receiver, enabling methods for decoding, decision-making, and error correction of semantic communication to address biases in semantic understanding and generate appropriate mission execution reference information based on contextual relationships.

[0280] Figure 8A is a schematic diagram of a task execution process provided according to an embodiment of the present disclosure. As shown in Figure 8A, a drone (i.e., a first device) with a task to be performed is capable of receiving communication signals, performing multimedia sensing, and receiving wireless charging signals or extracting energy recovery information from wireless signals during flight. Simultaneously, the drone possesses limited computing power, enabling it to perform basic deep learning. These multimodal methods can provide basic reference information for the execution of drone tasks.

[0281] Figure 8B is a schematic diagram of a multimodal input to data vector conversion interface provided according to an embodiment of the present disclosure. As shown in Figure 8B, the multimodal input is converted into an aligned data vector.

[0282] In some embodiments, the input for the UAV to perform the mission is composed of heterogeneous multimodal components S. k The basic information vector consists of (k∈M={comm, sen, comp, pow, ...}). It can be represented as a transformation function before information fusion.

[0283] In some embodiments, the wireless communication signal Scomm is actively transmitted from the signal source, and the wirelessly transmitted communication information is then transformed into a semantic vector. After TinyML feature extraction, the data is transmitted to the virtual semantic receiver.

[0284] In some embodiments, the drone senses mission-related external signals Ssen, such as radar signals, camera signals, and audio feedback, and then converts them into semantic vectors. To the virtual semantic receiver.

[0285] In some embodiments, the power mode allows for wireless charging via communication and sensor signals to compensate for the drone's energy, ensuring that the drone has the necessary conditions to complete its mission.

[0286] In some embodiments, computational modes typically support communication and sensing modes by providing features such as feature extraction, compression, data pattern recognition, and simplified source coding. The computational cost of the task is generated based on the computational capabilities available.

[0287] Figure 8C is a schematic diagram illustrating the correspondence between the task execution system and virtual semantic communication provided according to embodiments of this disclosure. As shown in Figure 8C, if the UAV uses semantic methods to process task-related multimodal information, the entire process of "acquiring, understanding, digesting, and executing the task" in the figure can be simplified into a simple process of integrating perception, communication, computing, and wireless power supply into "virtual semantic information generation and pragmatic intent matching." This is suitable for practical use with TinyML processing capabilities and can complete the multimodal communication of the UAV. The negative impacts generated in the entire process, such as communication interference, task-irrelevant information, and noise, can be unified into a bias (such as additive noise) mapped in the virtual semantic channel transmission.

[0288] Figure 8D is a schematic diagram of the logical layering of a task execution system provided according to an embodiment of the present disclosure. As shown in Figure 8D, through a virtual semantic information source, based on the essence of semantic communication, the entire UAV task execution system is logically decomposed into three layers: a syntax representation layer, a semantic merging layer, and a pragmatic implementation layer.

[0289] In some embodiments, in the syntax representation layer, multimodal signals such as sensory, communication, computation, and dynamic signals are collected and converted into information that is easy for TinyML to process through a multimodal communication interface.

[0290] In some embodiments, in the semantic merging layer, TinyML can directly process the multimodal alignment information obtained from the syntax representation layer to generate corresponding multimodal planar semantic vectors for communication, sensation, computation, force, etc.

[0291] In some embodiments, at the pragmatic implementation layer, task execution is achieved by matching semantic intent and pragmatic intent. TinyML is used to handle semantic subjective and objective biases, thereby using semantic communication methods to ensure accurate task execution.

[0292] Figure 8E is a schematic diagram of the framework of a virtual semantic communication system provided according to an embodiment of the present disclosure.

[0293] The virtual semantic information source consists of three parts: a feature extraction module, a semantic tensor module, and an encoding module.

[0294] In some embodiments, the main function of the feature extraction module is to obtain semantic vectors of multimodal information.

[0295] In some embodiments, since the feature vectors have different dimensions, the processing is relatively complex. Therefore, the semantic tensor module converts the feature vectors into one-dimensional vectors.

[0296] In some embodiments, semantic communication encoding and decoding can be trained and deployed on the same UAV, allowing them to complete tasks independently. Alternatively, they can be trained and deployed on multiple UAVs, enabling collaborative work between them. Multimodal data collected and fused on one UAV can be sent to other UAVs via a virtual semantic channel, and the other UAVs can then interpret the intent and work together to complete the task. If multi-UAV training is used, split learning or federated learning methods can be employed, followed by deployment as needed after training.

[0297] In some embodiments, for symbol mapping in virtual semantic communication, the purpose of this module is to convert the transformed aligned semantic vector into bit codes. In some embodiments, this module treats semantic information as the content to be transmitted and treats any objective deviations in the multimodal information collection process (such as communication interference, noise, environmental noise, and residual task-irrelevant information) as channel noise. The one-dimensional semantic vector is compressed and mapped to transmission symbols by a channel encoder consisting of multiple dense layers.

[0298] In some embodiments, these symbols are actually decomposed into the correct content to be transmitted, X, and the noise portion, n (pure noise in a fading-free semantic channel). After virtual channel transmission, the "received signal" can be represented as y = X + n, where X represents the correct bits and n represents the deviation (noise).

[0299] In some embodiments, semantic encoding and decoding may employ a transformer model to maximize attention to and understanding of its semantics.

[0300] In some embodiments, objective bias is an inherent and unavoidable characteristic of communication systems. It is widespread and can amplify subjective bias. Although a virtual channel is designed in this system, any biases during the multimodal information acquisition process (such as communication interference, noise, environmental noise, task-related information residues, etc.) are treated as interference information.

[0301] In some embodiments, intent is defined as the smallest unit of semantic information, and the understanding of the semantics of the knowledge base task is considered at the semantic layer. Because the sender and receiver may have different knowledge bases, receivers may interpret the correct semantic intent differently, leading to subjective bias.

[0302] In some embodiments, the TinyML method can effectively reduce model size, making it suitable for drone deployment. By selectively removing unnecessary weights, it can significantly reduce computational load without affecting model performance.

[0303] In some embodiments, the TinyML-based virtual semantic receiver has two main functions: (1) converting the received signal y into semantic information; and (2) determining the semantic intent.

[0304] In some embodiments, the virtual semantic receiver consists of a decoding module, a multimodal semantic fusion module, and an intent recognition module to realize the semantic matching process.

[0305] This disclosure proposes viewing the successful execution of advanced UAV missions as a process of semantic recognition and pragmatic execution. TinyML provides advanced UAV algorithms and models that can run on low-power and resource-constrained platforms. From a semantic communication perspective, leveraging TinyML's applicability to UAVs, heterogeneous multimodal communication and UAV mission execution processes are mapped, aiming to better utilize the capabilities of machine learning and semantic communication to enhance the practical mission execution capabilities of UAVs.

[0306] In the embodiments disclosed herein, some or all of the steps and their optional implementations may be arbitrarily combined with some or all of the steps in other embodiments, or may be arbitrarily combined with the optional implementations in other embodiments.

[0307] This disclosure also provides a task determination apparatus for implementing any of the above methods. For example, this disclosure provides a task determination apparatus including units or modules for implementing the steps performed by a first device in any of the above methods. For example, this disclosure provides a task determination apparatus including units or modules for implementing the steps performed by a second device in any of the above methods. For example, this disclosure provides a task determination apparatus including units or modules for implementing the steps performed by a third device in any of the above methods.

[0308] It should be understood that the division of units or modules in the above device is only a logical functional division. In actual implementation, they can be fully or partially integrated into a single physical entity, or they can be physically separated. Furthermore, the units or modules in the device can be implemented by a processor calling software: for example, the device includes a processor connected to a memory containing instructions. The processor calls the instructions stored in the memory to implement any of the above methods or to implement the functions of the units or modules in the above device. The processor can be, for example, a general-purpose processor, such as a Central Processing Unit (CPU) or a microprocessor, and the memory can be internal or external to the device. Alternatively, the units or modules in the device can be implemented in the form of hardware circuits. The functionality of some or all of the units or modules can be achieved through the design of these hardware circuits, which can be understood as one or more processors. For example, in one implementation, the hardware circuit is an application-specific integrated circuit (ASIC). The functionality of some or all of the units or modules is achieved through the design of the logical relationships between the components within the circuit. In another implementation, the hardware circuit can be implemented using a programmable logic device (PLD). Taking a field-programmable gate array (FPGA) as an example, it can include a large number of logic gates. The connection relationships between the logic gates are configured through configuration files, thereby achieving the functionality of some or all of the units or modules. All units or modules of the above device can be implemented entirely through processor-called software, entirely through hardware circuits, or partially through processor-called software with the remaining parts implemented through hardware circuits.

[0309] In this embodiment, the processor is a circuit with signal processing capabilities. In one implementation, the processor can be a circuit with instruction read and execute capabilities, such as a central processing unit, microprocessor, graphics processing unit (GPU) (which can be understood as a type of microprocessor), or digital signal processor (DSP). In another implementation, the processor can implement certain functions through the logical relationships of hardware circuits. The logical relationships of the aforementioned hardware circuits are fixed or reconfigurable. For example, the processor is a hardware circuit implemented by an application-specific integrated circuit (ASIC) or a programmable logic device, such as an FPGA. In a reconfigurable hardware circuit, the process of the processor loading a configuration document and configuring the hardware circuit can be understood as the process of the processor loading instructions to implement the functions of some or all of the above units or modules. Furthermore, it can also be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC, such as a neural network processing unit (NPU), tensor processing unit (TPU), deep learning processing unit (DPU), etc.

[0310] Figure 9 is a schematic diagram of the structure of a task determination device provided according to an embodiment of the present disclosure. As shown in Figure 9, the task determination device 900 may include at least one of the following: a transceiver module 901 and a processing module 902.

[0311] In some embodiments, the communication device 900 may be the first device 101. In some embodiments, the processing module 901 is configured to: acquire multimodal data, wherein the multimodal data includes multiple data, the multiple data belonging to at least two types; determine semantic intent based on the multimodal data by an encoder and a decoder; and determine the terminology intent corresponding to the multimodal data according to the semantic intent, wherein the terminology intent is used to indicate a first task. Optionally, the transceiver module 901 may be configured to perform at least one of the communication steps such as sending and / or receiving performed by the first device 101 in any of the above methods (e.g., steps S7101, S7201, S7205, S7207, but not limited thereto), which will not be elaborated here. Optionally, the processing module 902 may be configured to perform at least one of the steps performed by the first device 101 in any of the above methods, other than communication steps such as sending and receiving (e.g., steps S401, S402, S403, S404, S7102, S7103, S7202, S7203, S7204, but not limited thereto).

[0312] In some embodiments, the transceiver module may include a transmitting module and / or a receiving module. The transmitting and receiving modules may be separate or integrated. Optionally, the transceiver module may be interchangeable with a transceiver.

[0313] In some embodiments, the processing module may be a single module or may include multiple sub-modules. Optionally, the multiple sub-modules may each perform all or part of the steps required by the processing module. Optionally, the processing module may be interchangeable with a processor.

[0314] Figure 10A is a schematic diagram of the structure of a communication device provided according to an embodiment of the present disclosure. The communication device 10100 can be a network device (e.g., access network device, core network device, etc.), a terminal (e.g., user equipment, etc.), a chip, chip system, or processor that supports the network device in implementing any of the above methods, or a chip, chip system, or processor that supports the terminal in implementing any of the above methods. The communication device 10100 can be used to implement the methods described in the above method embodiments; for details, please refer to the descriptions in the above method embodiments.

[0315] As shown in Figure 10A, the communication device 10100 includes one or more processors 10101. The processor 10101 can be a general-purpose processor or a dedicated processor, such as a baseband processor or a central processing unit (CPU). The baseband processor can be used to process communication protocols and communication data, while the CPU can be used to control the communication device (e.g., base station, baseband chip, terminal device, terminal device chip, DU or CU, etc.), execute programs, and process program data. Optionally, the communication device 10100 can be used to execute any of the above methods. Optionally, one or more processors 10101 can be used to invoke instructions to cause the communication device 10100 to execute any of the above methods.

[0316] In some embodiments, the communication device 10100 further includes one or more transceivers 10102. When the communication device 10100 includes one or more transceivers 10102, the transceiver 10102 performs at least one of the communication steps such as sending and / or receiving in the above-described method (e.g., steps S7101, S7201, S7205, S7207, but not limited thereto), and the processor 10101 performs at least one of other steps (e.g., steps S401, S402, S403, S404, S7102, S7103, S7104, S7202, S7203, S7204, S7206, but not limited thereto). In optional embodiments, the transceiver may include a receiver and / or a transmitter, which may be separate or integrated together. Optionally, terms such as transceiver, transceiver unit, transceiver, transceiver circuit, interface circuit, and interface can be used interchangeably; terms such as transmitter, transmitting unit, transmitter, and transmitting circuit can be used interchangeably; and terms such as receiver, receiving unit, receiver, and receiving circuit can be used interchangeably.

[0317] In some embodiments, the communication device 10100 further includes one or more memories 10103 for storing data. Optionally, all or part of the memories 10103 may be located outside the communication device 10100. In optional embodiments, the communication device 10100 may include one or more interface circuits 10104. Optionally, the interface circuits 10104 are connected to the memories 10103 and can be used to receive data from the memories 10103 or other devices, and to send data to the memories 10103 or other devices. For example, the interface circuits 10104 can read data stored in the memories 10103 and send the data to the processor 10101.

[0318] The communication device 10100 described in the above embodiments may be a network device or a terminal, but the scope of the communication device 10100 described in this disclosure is not limited thereto, and the structure of the communication device 10100 may not be limited by FIG10A. The communication device may be a standalone device or may be part of a larger device. For example, the communication device may be: (1) a standalone integrated circuit IC, or chip, or chip system or subsystem; (2) a collection of one or more ICs, optionally, the IC collection may also include storage components for storing data and programs; (3) an ASIC, such as a modem; (4) a module that can be embedded in other devices; (5) a receiver, terminal device, smart terminal device, cellular phone, wireless device, handheld device, mobile unit, vehicle device, network device, cloud device, artificial intelligence device, etc.; (6) others, etc.

[0319] Figure 10B is a schematic diagram of the structure of a chip provided according to an embodiment of the present disclosure. For cases where the communication device 10100 can be a chip or a chip system, please refer to the schematic diagram of the structure of the chip 10200 shown in Figure 10B, but it is not limited thereto.

[0320] Chip 10200 includes one or more processors 10201. Chip 10200 is used to perform any of the above methods.

[0321] In some embodiments, chip 10200 further includes one or more interface circuits 10202. Optionally, terms such as interface circuit, interface, and transceiver pin can be used interchangeably. In some embodiments, chip 10200 further includes one or more memories 10203 for storing data. Optionally, all or part of the memories 10203 may be located outside of chip 10200. Optionally, interface circuit 10202 is connected to memory 10203, and interface circuit 10202 can be used to receive data from memory 10203 or other devices, and interface circuit 10202 can be used to send data to memory 10203 or other devices. For example, interface circuit 10202 can read data stored in memory 10203 and send the data to processor 10201.

[0322] In some embodiments, the interface circuit 10202 performs at least one of the communication steps such as sending and / or receiving in the above-described method (e.g., steps S7101, S7201, S7205, S7207, but not limited thereto). The interface circuit 10202 performing the communication steps such as sending and / or receiving in the above-described method refers, for example, to the interface circuit 10202 performing data interaction between the processor 10201, the chip 10200, the memory 10203, or the transceiver device. In some embodiments, the processor 10201 performs at least one of other steps (e.g., steps S401, S402, S403, S404, S7102, S7103, S7104, S7202, S7203, S7204, S7206, but not limited thereto).

[0323] The modules and / or devices described in the various embodiments, such as virtual devices, physical devices, and chips, can be combined or separated arbitrarily as needed. Optionally, some or all steps can also be performed collaboratively by multiple modules and / or devices, which is not limited here.

[0324] This disclosure also proposes a storage medium storing instructions that, when executed on a communication device 10100, cause the communication device 10100 to perform any of the methods described above. Optionally, the storage medium is an electronic storage medium. Optionally, the storage medium is a computer-readable storage medium, but is not limited thereto; it may also be a storage medium readable by other devices. Optionally, the storage medium may be a non-transitory storage medium, but is not limited thereto; it may also be a temporary storage medium.

[0325] This disclosure also proposes a program product that, when executed by the communication device 10100, causes the communication device 10100 to perform any of the above methods. Optionally, the program product is a computer program product.

[0326] This disclosure also proposes a computer program that, when run on a computer, causes the computer to perform any of the above methods.

[0327] Other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of the invention are indicated by the following claims.

[0328] It should be understood that the present invention is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.

Claims

1. A task determination method, executed by a first device, wherein, The method includes: Acquire multimodal data, wherein the multimodal data includes multiple data, and the multiple data belong to at least two types; Based on the multimodal data, semantic intent is determined through an encoder and a decoder; Based on the semantic intent, the terminology intent corresponding to the multimodal data is determined, wherein the terminology intent is used to indicate the first task.

2. The method according to claim 1, wherein, The determination of semantic intent based on the multimodal data through an encoder and decoder includes: Based on the multimodal data, a first symbol is determined by the encoder, wherein the first symbol contains the semantic information of the multimodal data; The semantic intent is determined by the decoder based on the first symbol.

3. The method according to claim 2, wherein, The step of determining the first symbol based on the multimodal data using the encoder includes: Feature extraction is performed on each of the multiple data points in the multimodal data to obtain a first feature; The first feature of each data is transformed to obtain a second feature, wherein the second features corresponding to the multiple data have the same dimension; The second feature is channel-coded to obtain a first symbol corresponding to the second feature.

4. The method according to claim 2 or 3, wherein, The semantic intent based on the first symbol, as determined by the decoder, includes: The first symbol is channel decoded to obtain a plurality of third features, wherein each of the plurality of third features corresponds to one of the plurality of data; Semantic decoding is performed on each of the third features to obtain semantic information; The semantic information of the multimodal data is fused to determine the semantic intent.

5. The method according to any one of claims 1 to 4, wherein, Determining the terminology intent corresponding to the multimodal data based on the semantic intent includes: Based on the terminology intent database, the terminology intent corresponding to the semantic intent is determined.

6. The method according to any one of claims 1 to 5, wherein, The multimodal data is obtained based on at least one of the following: communication signals; Sensing signals; Power supply signal; Task calculation results.

7. The method according to any one of claims 1 to 6, wherein, The acquisition of multimodal data includes at least one of the following: Receive the multimodal data; The multimodal data is obtained locally.

8. The method according to any one of claims 1 to 7, wherein, The encoder and the decoder are implemented based on a first machine learning model; The method further includes: Based on the stated terminology, model training is performed on the first machine learning model.

9. The method according to claim 8, wherein, The first machine learning model includes a micro machine learning model.

10. The method according to claim 8 or 9, wherein, The training of the first machine learning model is based on federated learning or segmentation learning.

11. The method according to any one of claims 1 to 10, wherein, The first device is a drone.

12. A task determination device, disposed in a first device, wherein, The device includes: The processing module is configured as follows: Acquire multimodal data, wherein the multimodal data includes multiple data, and the multiple data belong to at least two types; Based on the multimodal data, semantic intent is determined through an encoder and a decoder; Based on the semantic intent, the terminology intent corresponding to the multimodal data is determined, wherein the terminology intent is used to indicate the first task.

13. A communication device, comprising: One or more processors; A memory that stores instructions; When the instruction is executed by the communication device, it causes the communication device to implement the task determination method as described in any one of claims 1 to 11.

14. A storage medium storing instructions, wherein, When the instruction is executed on the communication device, the communication device implements the task determination method as described in any one of claims 1 to 11.

15. A computer program product comprising instructions, wherein, When the instruction is executed on the communication device, the communication device implements the task determination method as described in any one of claims 1 to 11.

Citation Information

Patent Citations

  • Abnormality detection method and device and electronic equipment

    CN115964628A

  • Image cognition semantic communication system and method based on multi-modal knowledge graph

    CN118260432A

  • Power plant scene multi-modal data collaborative calibration method based on attention mechanism

    CN118587640A

  • Systems and methods for shared cross-modal trajectory prediction

    US20210286371A1