Semantic communication methods, communication devices, communication system and storage medium
By segmenting and encoding data based on a semantic knowledge base, using a large-scale AI model to identify semantic targets, and filtering and compressing semantic fragments, the problem of redundant data transmission in semantic communication is solved, achieving efficient semantic data transmission and accurate semantic recovery.
Patent Information
- Application Number
- PCT/CN2024/124542
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-10-12
- Publication Date
- 2026-04-16
Smart Images

Figure CN2024124542_16042026_PF_FP_ABST
Abstract
Description
Semantic communication methods, communication devices, communication systems and storage media Technical Field
[0001] This disclosure relates to the field of communications, and in particular to a semantic communication method, communication device, communication system, and storage medium. Background Technology
[0002] With the rapid development of internet and artificial intelligence (AI) technologies, semantic communication (SC), as a new intelligent paradigm, is expected to become one of the core technologies of future communication. Through semantic communication technology, selective feature extraction, compression, and transmission of raw signals can be achieved, thereby utilizing semantic-level information to realize communication. This not only reduces the transmission of redundant data and improves bandwidth utilization but also ensures the effective delivery of critical information.
[0003] Summary of the Invention
[0004] To improve semantic communication performance, embodiments of this disclosure provide a semantic communication method, communication device, communication system, and storage medium.
[0005] According to a first aspect of the present disclosure, a semantic communication method is provided, applied to a first communication device. The method includes: semantically segmenting first data based on a semantic knowledge base to obtain a plurality of first semantic segments; encoding the plurality of first semantic segments to obtain second data; and sending the second data to a second communication device, wherein the second data is used by the second communication device for data decoding and semantic recovery.
[0006] According to a second aspect of the present disclosure, a semantic communication method is provided, applied to a second communication device, the method comprising: receiving second data sent by a first communication device; decoding the second data to obtain a plurality of second semantic segments; and performing semantic recovery on the plurality of second semantic segments based on a semantic knowledge base to obtain third data.
[0007] According to a third aspect of the present disclosure, a first communication device is provided, comprising: a processing module configured to perform semantic segmentation on first data based on a semantic knowledge base to obtain a plurality of first semantic segments; the processing module is further configured to encode the plurality of first semantic segments to obtain second data; and a transceiver module configured to send the second data to a second communication device, wherein the second data is used by the second communication device for data decoding and semantic recovery.
[0008] According to a fourth aspect of the present disclosure, a second communication device is provided, comprising: a transceiver module configured to receive second data sent by a first communication device; a processing module configured to decode the second data to obtain a plurality of second semantic segments; the processing module is further configured to perform semantic recovery on the plurality of second semantic segments based on a semantic knowledge base to obtain third data.
[0009] According to a fifth aspect of the present disclosure, a first communication device is provided, comprising: one or more processors; wherein the first communication device is configured to perform the semantic communication method as described in the first aspect above.
[0010] According to a sixth aspect of the present disclosure, a second communication device is provided, comprising: one or more processors; wherein the second communication device is configured to perform the semantic communication method as described in the second aspect above.
[0011] According to a seventh aspect of the present disclosure, a communication system is provided, including a first communication device and a second communication device, wherein the first communication device is configured to implement the semantic communication method as described in the first aspect above, and the second communication device is configured to implement the communication method as described in the second aspect above.
[0012] According to an eighth aspect of the present disclosure, a storage medium is provided that stores instructions that, when executed on a communication device, cause the communication device to perform the semantic communication method as described in the first or second aspect above.
[0013] In this embodiment, a first communication device segments first data based on a semantic knowledge base to obtain multiple first semantic segments. These segments are then encoded to obtain second data, which is subsequently sent to a second communication device. The second communication device receives the second data from the first device, decodes it to obtain multiple second semantic segments, and then performs semantic recovery on these segments based on the semantic knowledge base to obtain third data. This embodiment enables semantic communication-based data transmission, improving communication efficiency. Furthermore, by using a semantic knowledge base for semantic segment acquisition and recovery, the accuracy of semantic data processing at both the sending and receiving ends is ensured, thus improving semantic communication performance.
[0014] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0015] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0016] Figure 1 is a schematic diagram of the architecture of a communication system according to an embodiment of the present disclosure.
[0017] Figure 2 is an interactive schematic diagram of a semantic communication method according to an embodiment of the present disclosure.
[0018] Figure 3 is a flowchart illustrating a semantic communication process according to an embodiment of the present disclosure.
[0019] Figure 4 is a schematic diagram of the workflow of a semantic communication framework based on a large AI model according to an embodiment of the present disclosure.
[0020] Figure 5 is a schematic diagram illustrating the principle of a training process according to an embodiment of the present disclosure.
[0021] Figure 6A is a schematic flowchart illustrating a semantic communication method according to an embodiment of the present disclosure.
[0022] Figure 6B is a schematic flowchart illustrating a semantic communication method according to an embodiment of the present disclosure.
[0023] Figure 7A is a flowchart illustrating a semantic communication method according to an embodiment of the present disclosure.
[0024] Figure 7B is a schematic flowchart illustrating a semantic communication method according to an embodiment of the present disclosure.
[0025] Figure 8A is a schematic diagram of the structure of the first communication device proposed in an embodiment of this disclosure.
[0026] Figure 8B is a schematic diagram of the structure of the second communication device proposed in an embodiment of this disclosure.
[0027] Figure 9A is a schematic diagram of the structure of the communication device 9100 proposed in an embodiment of this disclosure.
[0028] Figure 9B is a schematic diagram of the structure of the chip 9200 proposed in an embodiment of this disclosure. Detailed Implementation
[0029] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0030] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure. The singular forms “a,” “the,” and “the” as used in this disclosure and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of at least one associated listed item.
[0031] It should be understood that although the terms first, second, third, etc., may be used in this disclosure to describe various messages, these messages should not be limited to these terms. These terms are used only to distinguish messages of the same type from one another. For example, without departing from the scope of this disclosure, a first message may also be referred to as a second message, and similarly, a second message may also be referred to as a first message. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."
[0032] This disclosure presents a semantic communication method, communication device, communication system, and storage medium.
[0033] In a first aspect, embodiments of this disclosure propose a semantic communication method applied to a first communication device. The method includes: semantically segmenting first data based on a semantic knowledge base to obtain multiple first semantic segments; encoding the multiple first semantic segments to obtain second data; and sending the second data to a second communication device, wherein the second data is used by the second communication device for data decoding and semantic recovery.
[0034] In the above embodiments, the first communication device segments the first data based on a semantic knowledge base to obtain multiple first semantic segments, and then encodes the multiple first semantic segments to obtain second data, which is then sent to the second communication device. This allows the second communication device to perform data decoding and semantic recovery based on the second data, thereby realizing data transmission based on semantic communication, improving communication efficiency. Furthermore, by processing the first data to be transmitted based on the semantic knowledge base, the accuracy of semantic data processing at the sending end can be guaranteed, improving semantic communication performance.
[0035] In conjunction with some embodiments of the first aspect, in some embodiments, the semantic knowledge base is a large-scale artificial intelligence (AI) model; the step of semantically segmenting the first data based on the semantic knowledge base to obtain multiple first semantic segments includes: identifying semantic targets contained in the first data through a large-scale AI model; and segmenting the first data based on the identified semantic targets through a large-scale AI model to obtain the multiple first semantic segments.
[0036] In the above embodiments, an optional implementation method for semantic segmentation of the first data based on a semantic knowledge base is provided to ensure the smooth progress of the data processing process based on the semantic knowledge base, thereby ensuring the smooth progress of the semantic communication process.
[0037] In conjunction with some embodiments of the first aspect, in some embodiments, the step of encoding based on the plurality of first semantic segments to obtain second data includes: performing semantic encoding based on the plurality of first semantic segments using a semantic encoder to obtain first semantic information; and performing channel encoding based on the first semantic information using a channel encoder to obtain the second data.
[0038] In the above embodiments, an optional implementation method is provided to obtain corresponding data by encoding based on semantic segments, so as to ensure the smooth progress of the encoding process based on semantic segments, thereby obtaining second data that can be transmitted in the channel, so as to ensure the smooth progress of the semantic communication process.
[0039] In some embodiments, in conjunction with the first aspect, the method further includes: compressing the first semantic information using a semantic compression network to obtain first semantic information for channel coding.
[0040] In the above embodiments, the first semantic information is compressed by using a semantic compression network to reduce redundant information in the first semantic information. The compressed first semantic information is then used as the semantic information for channel coding, thereby reducing the waste of resources caused by channel coding of redundant semantic information, reducing communication overhead, and improving communication efficiency.
[0041] In conjunction with some embodiments of the first aspect, in some embodiments, the step of compressing the first semantic information through a semantic compression network to obtain first semantic information for channel coding includes: masking a portion of the first semantic information through a semantic compression network to obtain first semantic information for channel coding.
[0042] In the above embodiments, an optional implementation method is provided to compress the first semantic information using a semantic compression network, so as to ensure that the compression of the first semantic information can be achieved, thereby reducing the waste of resources caused by channel coding of redundant semantic information, reducing communication overhead, and improving communication efficiency.
[0043] In conjunction with some embodiments of the first aspect, in some embodiments, different mask ratios correspond to different first data, and the mask ratio indicates the proportion of the first semantic information used for masking in the first semantic information obtained by semantic encoding based on the plurality of first semantic segments.
[0044] In the above embodiments, by configuring different mask ratios for different first data, the appropriate mask ratio can be used to compress the semantic information obtained from the processing of the first data based on the actual situation of the first data, thereby ensuring that the semantic compression process is more targeted and improving the semantic compression performance.
[0045] In conjunction with some embodiments of the first aspect, in some embodiments, the method further includes: determining, based on the semantic importance of the plurality of first semantic segments, a first semantic segment whose semantic importance meets the requirements, as the first semantic segment for encoding.
[0046] In the above embodiments, by filtering the first semantic segment based on semantic importance, key semantic segments that meet the semantic importance requirements are selected as the first semantic segments for encoding, thereby achieving more accurate semantic perception, reducing the semantic encoding processing pressure while ensuring the semantic communication effect, and improving the semantic communication efficiency.
[0047] In conjunction with some embodiments of the first aspect, in some embodiments, the method further includes: determining the semantic importance of the plurality of first semantic segments through an attention network.
[0048] In the above embodiments, an optional implementation method for determining the semantic importance of semantic segments is provided to ensure that the semantic importance of each semantic segment can be evaluated, thereby ensuring that key semantic segments can be screened based on the evaluated semantic segments.
[0049] In conjunction with some embodiments of the first aspect, in some embodiments, the first data is image data, and the semantic knowledge base is a large-scale AI model for implementing image recognition and / or image classification; and / or, the first data is text data, and the semantic knowledge base is a large-scale AI model for implementing text content recognition and / or human-computer dialogue; and / or, the first data is audio data, and the semantic knowledge base is a large-scale AI model for implementing speech recognition and / or speech separation.
[0050] In the above embodiments, semantic knowledge base types corresponding to different types of first data are provided to ensure that a large AI model of the corresponding type can be used as a semantic knowledge base to match the first data of the corresponding type, thereby ensuring that different types of first data can be semantically segmented based on the corresponding type of semantic knowledge base, thus ensuring the smooth progress of the semantic communication process.
[0051] In conjunction with some embodiments of the first aspect, in some embodiments, the large AI model, the attention network, the semantic compression network, the semantic encoder, and the channel encoder are all pre-trained; for any network structure among the large AI model, the attention network, the semantic compression network, the semantic encoder, and the channel encoder, when the network structure has been trained, the network parameters of the trained network structure are frozen to continue training other network structures.
[0052] In the above embodiments, by pre-training a large AI model, attention network, semantic compression network, semantic encoder, and channel encoder, the semantic data to be transmitted can be processed based on the pre-trained network structure, thereby ensuring the smooth progress of the semantic communication process. Furthermore, by freezing the network parameters of the trained network structure after its training is complete, and continuing to train other network structures, the training of other network structures is guaranteed not to affect the already trained network structure, thus ensuring the training effect of the network structure.
[0053] In conjunction with some embodiments of the first aspect, in some embodiments, the large AI model and the attention network are obtained by rapid learning or fine-tuning training based on a pre-trained model, wherein the pre-trained model is trained based on a general dataset.
[0054] In the above embodiments, an optional implementation method for training large-scale AI models and attention networks is provided to ensure that the training of large-scale AI models and attention networks can be achieved, thereby ensuring the smooth progress of the semantic communication process.
[0055] In conjunction with some embodiments of the first aspect, in some embodiments, the semantic compression network is trained with the network parameters of the large AI model and the attention network frozen, and the loss function of the semantic compression network is the difference between the first semantic information used for channel coding and the first semantic information obtained by decompression by the second communication device.
[0056] In the above embodiments, an optional implementation of training a semantic compression network is provided to ensure that the training of the semantic compression network can be achieved, thereby ensuring the smooth progress of the semantic communication process.
[0057] In conjunction with some embodiments of the first aspect, in some embodiments, the channel encoder is trained with the network parameters of the semantic encoder frozen, wherein the loss function of the semantic encoder is the difference between the first semantic information obtained by semantic encoding and the second semantic information obtained by the second communication device by semantic decoding, and the loss function of the channel encoder is the difference between the second data and the third data obtained by the second communication device by channel decoding.
[0058] In the above embodiments, an optional implementation of training the channel encoder and semantic encoder is provided to ensure that the training of the channel encoder and semantic encoder can be realized, thereby ensuring the smooth progress of the semantic communication process.
[0059] Secondly, this disclosure proposes a semantic communication method applied to a second communication device. The method includes: receiving second data sent by a first communication device; decoding the second data to obtain multiple second semantic segments; and performing semantic recovery on the multiple second semantic segments based on a semantic knowledge base to obtain third data.
[0060] In the above embodiments, by receiving second data sent by the first communication device through the second communication device, decoding the second data to obtain multiple second semantic segments, and then performing semantic recovery on the multiple second semantic segments based on the semantic knowledge base to obtain third data, the data transmission based on semantic communication is realized, improving communication efficiency. Furthermore, by realizing semantic recovery of semantic segments based on the semantic knowledge base, the accuracy of semantic data processing at the receiving end can be guaranteed, thus improving semantic communication performance.
[0061] In conjunction with some embodiments of the second aspect, in some embodiments, the step of decoding based on the second data to obtain multiple second semantic segments includes: performing channel decoding based on the second data using a channel decoder to obtain second semantic information; and performing semantic decoding based on the second semantic information using a semantic decoder to obtain the multiple second semantic segments.
[0062] In the above embodiments, an optional implementation is provided to decode the received second data to obtain the corresponding semantic fragment, so as to ensure the smooth progress of the decoding process based on the channel transmitted data, thereby ensuring the smooth progress of the semantic communication process.
[0063] In conjunction with some embodiments of the second aspect, in some embodiments, the semantic knowledge base is a large-scale AI model; the step of performing semantic recovery on the plurality of second semantic segments based on the semantic knowledge base to obtain third data includes: identifying the semantic target corresponding to each of the plurality of second semantic segments through the large-scale AI model; and performing semantic recovery on the plurality of second semantic segments based on the identified semantic target through the large-scale AI model to obtain the third data.
[0064] In the above embodiments, an optional implementation method for semantic recovery of semantic fragments based on a semantic knowledge base is provided to ensure the smooth progress of the semantic recovery process based on the semantic knowledge base, thereby ensuring the smooth progress of the semantic communication process.
[0065] In conjunction with some embodiments of the second aspect, in some embodiments, the second data is image data, and the semantic knowledge base is a large-scale AI model for implementing image recognition and / or image classification; and / or, the second data is text data, and the semantic knowledge base is a large-scale AI model for implementing text content recognition and / or human-computer dialogue; and / or, the second data is audio data, and the semantic knowledge base is a large-scale AI model for implementing speech recognition and / or speech separation.
[0066] In conjunction with some embodiments of the second aspect, in some embodiments, the large AI model, the semantic decoder, and the channel decoder are all pre-trained; for any network structure among the large AI model, the semantic decoder, and the decoding encoder, when the network structure has been trained, the network parameters of the trained network structure are frozen so as to continue training other network structures.
[0067] In conjunction with some embodiments of the second aspect, in some embodiments, the large AI model is obtained by rapid learning or fine-tuning training based on a pre-trained model, wherein the pre-trained model is trained based on a general dataset.
[0068] In conjunction with some embodiments of the second aspect, in some embodiments, the channel decoder is trained with the network parameters of the semantic decoder frozen, wherein the loss function of the semantic decoder is the difference between the second semantic information obtained by semantic decoding and the first semantic information obtained by the first communication device by semantic encoding, and the loss function of the channel decoder is the difference between the second data and the third data.
[0069] Thirdly, embodiments of this disclosure propose a first communication device, comprising: a processing module configured to perform semantic segmentation on first data based on a semantic knowledge base to obtain multiple first semantic segments; the processing module is further configured to encode the multiple first semantic segments to obtain second data; and a transceiver module configured to send the second data to a second communication device, wherein the second data is used by the second communication device for data decoding and semantic recovery.
[0070] Fourthly, embodiments of this disclosure propose a second communication device, comprising: a transceiver module configured to receive second data sent by a first communication device; a processing module configured to decode the second data to obtain a plurality of second semantic segments; the processing module is further configured to perform semantic recovery on the plurality of second semantic segments based on a semantic knowledge base to obtain third data.
[0071] Fifthly, embodiments of this disclosure provide a first communication device, comprising: one or more processors; wherein the first communication device is configured to perform the semantic communication method as described in the first aspect and any embodiment thereof.
[0072] In a sixth aspect, embodiments of this disclosure provide a second communication device, comprising: one or more processors; wherein the second communication device is configured to perform the semantic communication method as described in the second aspect and any embodiment thereof.
[0073] In a seventh aspect, embodiments of this disclosure provide a communication system including a first communication device and a second communication device, wherein the first communication device is configured to implement the semantic communication method as described in the first aspect and any embodiment thereof, and the second communication device is configured to implement the communication method as described in the second aspect and any embodiment thereof.
[0074] Eighthly, embodiments of this disclosure provide a storage medium storing instructions that, when executed on a communication device, cause the communication device to perform the semantic communication method as described in the first aspect and any embodiment of the first aspect, the second aspect and any embodiment of the second aspect.
[0075] Ninthly, embodiments of this disclosure provide a program product that, when executed by a communication device, causes the communication device to perform the semantic communication method as described in the first aspect and any embodiment of the first aspect, the second aspect and any embodiment of the second aspect.
[0076] In a tenth aspect, embodiments of this disclosure provide a computer program that, when run on a computer, causes the computer to perform the semantic communication method as described in the first aspect and any embodiment of the first aspect, the second aspect and any embodiment of the second aspect.
[0077] Eleventhly, embodiments of this disclosure provide a chip or chip system. The chip or chip system includes processing circuitry configured to perform the semantic communication methods as described in the first aspect and any embodiment of the first aspect, the second aspect and any embodiment of the second aspect.
[0078] It is understood that the aforementioned first communication device, second communication device, communication system, storage medium, program product, computer program, chip, or chip system are all used to execute the methods proposed in the embodiments of this disclosure. Therefore, the beneficial effects that can be achieved can be referred to the beneficial effects in the corresponding methods, and will not be repeated here.
[0079] This disclosure provides a semantic communication method, communication device, communication system, and storage medium. In some embodiments, the terms semantic communication method, information processing method, and communication method can be used interchangeably; the terms semantic communication device, information processing device, and communication device can be used interchangeably; and the terms information processing system, communication system, and semantic communication system can be used interchangeably.
[0080] This disclosure is not exhaustive, but merely illustrative of some embodiments, and is not intended to limit the scope of protection of this disclosure. Unless otherwise specified, each step in a particular embodiment can be implemented as an independent embodiment, and the steps can be arbitrarily combined. For example, a solution after removing some steps in a particular embodiment can also be implemented as an independent embodiment, and the order of the steps in a particular embodiment can be arbitrarily interchanged. Furthermore, the optional implementation methods in a particular embodiment can be arbitrarily combined; moreover, the embodiments can be arbitrarily combined, for example, some or all steps of different embodiments can be arbitrarily combined, and a particular embodiment can be arbitrarily combined with the optional implementation methods of other embodiments.
[0081] In each of the disclosed embodiments, unless otherwise specified or in case of logical conflict, the terminology and / or descriptions of the embodiments are consistent and can be referenced by each other. Technical features in different embodiments can be combined to form new embodiments based on their inherent logical relationships.
[0082] The terminology used in the embodiments of this disclosure is for the purpose of describing particular embodiments only and is not intended to limit the scope of this disclosure.
[0083] In this embodiment of the disclosure, unless otherwise stated, elements expressed in the singular form, such as "a," "an," "the," "the," "the," "the," "the," "the," "this," etc., can mean "one and only one," or "one or more," "at least one," etc. For example, when using articles such as "a," "an," "the," etc. in translation, the noun following the article can be understood as either a singular expression or a plural expression.
[0084] In the embodiments disclosed herein, "multiple" refers to two or more.
[0085] In some embodiments, the terms “at least one of”, “one or more”, “a plurality of”, “multiple”, etc., may be used interchangeably.
[0086] In some embodiments, the notation "at least one of A and B", "A and / or B", "A in one case, B in another", "in response to one case A, in response to another case B", etc., may include the following technical solutions depending on the situation: in some embodiments, A (execute A regardless of B); in some embodiments, B (execute B regardless of A); in some embodiments, execution is selected from A and B (A and B are selectively executed); in some embodiments, A and B (both A and B are executed). The same applies when there are more branches such as A, B, C, etc.
[0087] In some embodiments, the notation "A or B" may include the following technical solutions, depending on the situation: in some embodiments, A (execution of A regardless of B); in some embodiments, B (execution of B regardless of A); in some embodiments, execution is selected from A and B (A and B are selectively executed). The same applies when there are more branches such as A, B, C, etc.
[0088] The prefixes "first," "second," etc., used in the embodiments of this disclosure are merely for distinguishing different descriptive objects and do not impose restrictions on the position, order, priority, quantity, or content of the descriptive objects. The description of the descriptive objects is found in the claims or the context of the embodiments, and the use of prefixes should not constitute unnecessary restrictions. For example, if the descriptive object is a "field," the ordinal numbers preceding "field" in "first field" and "second field" do not restrict the position or order of the "fields." "First" and "second" do not restrict whether the "fields" they modify are in the same message, nor do they restrict the order of "first field" and "second field." Similarly, if the descriptive object is a "level," the ordinal numbers preceding "level" in "first level" and "second level" do not restrict the priority between "levels." Furthermore, the number of descriptive objects is not limited by ordinal numbers and can be one or more. For example, in "first device," the number of "devices" can be one or more. Furthermore, the objects modified by different prefixes can be the same or different. For example, if the object being described is "device", then "first device" and "second device" can be the same device or different devices, and their types can be the same or different. Similarly, if the object being described is "information", then "first information" and "second information" can be the same information or different information, and their content can be the same or different.
[0089] In some embodiments, “including A,” “containing A,” “for indicating A,” and “carrying A” can be interpreted as directly carrying A or indirectly indicating A.
[0090] In some embodiments, the terms “in response to…”, “in response to determining…”, “in the case of…”, “when…”, “if…”, “if…”, etc., can be used interchangeably.
[0091] In some embodiments, the terms “greater than,” “greater than or equal to,” “not less than,” “more than,” “more than or equal to,” “not less than,” “higher than,” “higher than or equal to,” “not lower than,” and “above” can be used interchangeably, as can the terms “less than,” “less than or equal to,” “not greater than,” “less than,” “less than or equal to,” “not more than,” “lower than,” “lower than or equal to,” “not higher than,” and “below”.
[0092] In some embodiments, the apparatus and device may be interpreted as physical or virtual, and their names are not limited to the names recorded in the embodiments. In some cases, they may also be understood as "equipment", "device", "circuit", "network element", "node", "function", "unit", "section", "system", "network", "chip", "chip system", "entity", "body", etc.
[0093] In some embodiments, "network" can be interpreted as devices included in the network, such as access network devices, core network devices, etc.
[0094] In some embodiments, "access network device (AN device)" may also be referred to as "radio access network device (RAN device)," "base station (BS)," "radio base station," or "fixed station." In some embodiments, it may also be understood as "node," "access point," "transmission point (TP)," "reception point (RP)," "transmission / reception point (TRP)," "panel," "antenna panel," "antenna array," "cell," "macro cell," "small cell," "femto cell," "pico cell," "sector," "cell group," "serving cell," "carrier," "component carrier," or "bandwidth part (BWP)."
[0095] In some embodiments, "terminal" or "terminal device" may be referred to as "user equipment (UE)," "user terminal," "mobile station (MS)," "mobile terminal (MT)," "subscriber station," "mobile unit," "subscriber unit," "wireless unit," "remote unit," "mobile device," "wireless device," "wireless communication device," "remote device," "mobile subscriber station," "access terminal," "mobile terminal," "wireless terminal," "remote terminal," "handset," "user agent," "mobile client," "client," etc.
[0096] In some embodiments, the acquisition of data, information, etc., may comply with the laws and regulations of the country where the location is situated.
[0097] In some embodiments, data, information, etc., may be obtained with the user's consent.
[0098] Furthermore, each element, each row, or each column in the table of this disclosure can be implemented as an independent embodiment, and any combination of any element, any row, or any column can also be implemented as an independent embodiment.
[0099] Figure 1 is a schematic diagram of the architecture of a communication system according to an embodiment of the present disclosure. As shown in Figure 1, the communication system 100 includes a first communication device 101 and a second communication device 102.
[0100] In some embodiments, the first communication device 101 is a semantic communication sender, and the second communication device 102 is a semantic communication receiver.
[0101] In some embodiments, the first communication device 101 is a terminal, and the second communication device 102 is a network device.
[0102] In some embodiments, the first communication device 101 is a network device, and the second communication device 102 is a terminal.
[0103] In some embodiments, the terminal includes, but is not limited to, at least one of the following: mobile phone, wearable device, Internet of Things device, car with communication function, smart car, tablet computer, computer with wireless transceiver function, virtual reality (VR) terminal device, augmented reality (AR) terminal device, wireless terminal device in industrial control, wireless terminal device in self-driving, wireless terminal device in remote medical surgery, wireless terminal device in smart grid, wireless terminal device in transportation safety, wireless terminal device in smart city, and wireless terminal device in smart home.
[0104] In some embodiments, the network device includes at least one of an access network device and a core network device.
[0105] In some embodiments, the access network device is, for example, a node or device that connects a terminal to a wireless network. The access network device may include, but is not limited to, at least one of the following in a 5G communication system: evolved Node B (eNB), next-generation eNB (ng-eNB), next-generation Node B (gNB), node B (NB), home node B (HNB), home evolved node B (HeNB), radio backhaul device, radio network controller (RNC), base station controller (BSC), base transceiver station (BTS), base band unit (BBU), mobile switching center, base station in a 6G communication system, open RAN, cloud RAN, base station in other communication systems, and access node in a Wi-Fi system.
[0106] In some embodiments, the technical solutions of this disclosure can be applied to the Open RAN architecture. In this case, the interfaces between or within access network devices involved in the embodiments of this disclosure can be transformed into internal interfaces of Open RAN. The processes and information interactions between these internal interfaces can be implemented by software or programs.
[0107] In some embodiments, the access network device may be composed of a central unit (CU) and a distributed unit (DU). The CU may also be called a control unit. The CU-DU structure can separate the protocol layer of the access network device. Some of the protocol layer functions are centrally controlled by the CU, while the remaining part or all of the protocol layer functions are distributed in the DU and centrally controlled by the CU. However, this is not the only possibility.
[0108] In some embodiments, the core network equipment can be a single device comprising multiple network elements, or it can be multiple devices or a group of devices, each comprising all or part of the multiple network elements. Network elements can be virtual or physical. The core network includes, for example, at least one of the following: Evolved Packet Core (EPC), 5G Core Network (5GCN), 6G Core Network (6GCN), and Next Generation Core (NGC).
[0109] In some embodiments, the core network equipment may include a first network element, such as an Access and Mobility Management Function (AMF) node.
[0110] In some embodiments, the first network element is used for user access management and mobility management, but is not limited thereto.
[0111] In some embodiments, the core network device may include a second network element, such as an Application Function (AF) node.
[0112] In some embodiments, the second network element is used to support applications or services running on the network edge or device, such as video streaming, voice calls, messaging, etc., but is not limited thereto.
[0113] In some embodiments, the core network device may include a third network element, such as a User Plane Function (UPF).
[0114] In some embodiments, the third network element is used for user plane data forwarding, traffic statistics, Quality of Service (QoS) management, etc., but is not limited to these.
[0115] In some embodiments, the core network device may include a fourth network element, such as a Policy Control Function (PCF).
[0116] In some embodiments, the fourth network element is used to implement user control policy management, including but not limited to QoS control, service access control, etc.
[0117] In some embodiments, the core network equipment may include a fifth network element, such as a unified data management function (UDM).
[0118] In some embodiments, the fifth network element is used to implement user subscription data management, roaming control, etc., but is not limited to these.
[0119] In some embodiments, the core network device may include a sixth network element, such as an Authentication Server Function (AUSF).
[0120] In some embodiments, the sixth network element is used to implement user authentication, but is not limited thereto.
[0121] In some embodiments, each of the above network elements can be independent of the core network equipment.
[0122] In some embodiments, each of the above network elements may be part of the core network equipment.
[0123] It is understood that the communication system described in this disclosure is for the purpose of more clearly illustrating the technical solutions of this disclosure, and does not constitute a limitation on the technical solutions proposed in this disclosure. As those skilled in the art will know, with the evolution of system architecture and the emergence of new business scenarios, the technical solutions proposed in this disclosure are also applicable to similar technical problems.
[0124] The following embodiments of this disclosure can be applied to the communication system 100 shown in FIG1, or to some of the main bodies, but are not limited thereto. The main bodies shown in FIG1 are illustrative. The communication system may include all or some of the main bodies in FIG1, or may include other main bodies outside of FIG1. The number and form of each main body are arbitrary. Each main body may be physical or virtual. The connection relationship between the main bodies is illustrative. The main bodies may not be connected or may be connected. The connection can be in any way, it can be a direct connection or an indirect connection, it can be a wired connection or a wireless connection.
[0125] The embodiments disclosed herein can be applied to Long Term Evolution (LTE), LTE-Advanced (LTE-A), LTE-Beyond (LTE-B), SUPER 3G, IMT-Advanced, 4th generation mobile communication system (4G), 5th generation mobile communication system (5G), 5G new radio (NR), Future Radio Access (FRA), New-Radio Access Technology (RAT), New Radio (NR), New radio access (NX), Future generation radio access (FX), Global System for Mobile communications (GSM), CDMA2000, Ultra Mobile Broadband (UMB), IEEE 802.11 (Wi-Fi), IEEE 802.16 (WiMAX), and IEEE 802.20, Ultra-Wideband (UWB), Bluetooth (a registered trademark), Public Land Mobile Network (PLMN) networks, Device-to-Device (D2D) systems, Machine-to-Machine (M2M) systems, Internet of Things (IoT) systems, Vehicle-to-Everything (V2X) systems, systems utilizing other communication methods, and next-generation systems built upon them, etc. Furthermore, multiple systems can be combined (e.g., a combination of LTE or LTE-A with 5G).
[0126] Semantic communication (SC), as a new intelligent paradigm, holds promise for contributing to various applications such as the virtual world, mixed reality (MR), and the Internet of Everything (IoE). Compared to traditional communication methods, semantic communication prioritizes conveying the intended meaning with minimal data, effectively improving information transmission efficiency.
[0127] In some embodiments, a semantic communication system may consist of a semantic encoder, a channel encoder (or channel encoder), a channel decoder (or channel decoder, communication decoder, channel decoder), a semantic decoder (or semantic decoder), and a semantic knowledge base (KB).
[0128] In some embodiments, a semantic encoder can extract semantic information from raw data and encode the extracted information into semantic features to understand the meaning of the data at the semantic level and reduce the scale of transmitted information.
[0129] In some embodiments, the channel encoder can encode and modulate semantic features to combat channel impairments, improve robustness, and ensure effective data transmission over the wireless physical channel.
[0130] In some embodiments, a channel decoder can be used to demodulate and decode the received signal to obtain the semantic features of the transmission before recovering the original data.
[0131] In some embodiments, a semantic decoder can be used to understand received semantic features, infer semantic information, and recover the original data from a semantic level.
[0132] In some embodiments, a semantic knowledge base can be viewed as a general knowledge model that can help semantic encoders and semantic decoders effectively understand and infer semantic information.
[0133] In some embodiments, semantic knowledge bases play a crucial role in enabling semantic communication to stand out from traditional communication systems due to their ability to understand and infer semantic information. However, there is currently no clear implementation method for how semantic knowledge bases play a role in the semantic communication process.
[0134] Furthermore, in some embodiments, a general semantic knowledge base can be constructed by learning a large amount of world knowledge. This general semantic knowledge base can consist of prior knowledge and background knowledge that users can understand and recognize, to help the semantic encoder and semantic decoder effectively understand and infer semantic information. However, semantic knowledge bases require a complex and time-consuming learning process. Therefore, the construction of semantic knowledge bases faces problems such as limited knowledge representation, frequent knowledge updates, and insecure knowledge sharing.
[0135] In view of this, the present disclosure aims to provide a semantic communication method to address the shortcomings of related technologies.
[0136] Figure 2 is an interactive schematic diagram of a semantic communication method according to an embodiment of the present disclosure. As shown in Figure 2, the embodiments of the present disclosure relate to a semantic communication method, which includes:
[0137] In step S2101, the first communication device performs semantic segmentation on the first data based on the semantic knowledge base to obtain multiple first semantic segments.
[0138] In some embodiments, the first communication device is a semantic communication sender.
[0139] In some embodiments, the first data is the data to be sent by the semantic communication sender, or in other words, the first data is the semantic communication data to be transmitted.
[0140] In some embodiments, the first data can be of various types, such as text data, image data, audio data, etc., but is not limited thereto.
[0141] In some embodiments, large artificial intelligence models (LAMs) have advantages such as accurate knowledge representation, rich prior knowledge and / or background knowledge, and low-cost knowledge updates, which can demonstrate excellent performance in fields such as natural language processing, image recognition, and speech recognition. Therefore, large AI models can be considered as the semantic knowledge base of a semantic communication system.
[0142] Optionally, for different types of primary data, different types of large-scale AI models can be used as their corresponding semantic knowledge bases.
[0143] In some embodiments, the first data is image data, and the corresponding semantic knowledge base should be able to segment various targets in the image and identify the categories of the segmentation results and the relationships between them. Therefore, a large AI model used to implement image recognition and / or image classification can be used as the semantic knowledge base of the image data.
[0144] For example, if the first data is image data, a Segment Anything Model (SAM) can be used as the semantic knowledge base for the image data. This allows zero-shot results to be generalized to unfamiliar images and targets without any additional training. The semantic communication sender can then use SAM to segment the input image and select important and meaningful segments for semantic encoding. The semantic communication receiver can generate reconstructed image data and use SAM to reduce semantic noise or interference. Furthermore, SAM excels at transmitting visual information through images, capturing complex details, spatial organization, and color, and can accurately represent facial expressions, emotions, and nonverbal cues to improve text semantic perception performance.
[0145] In some embodiments, the first data is text data, and its corresponding semantic knowledge base should be able to understand the content of the text and identify the attributes and relationships of various subjects and different topics. Therefore, a large-scale AI model used to realize text content recognition and / or human-computer dialogue can be used as the semantic knowledge base of the text data.
[0146] For example, if the first data is text data, a large language model can be used as the semantic knowledge base for the text data. Because a large language model can accurately understand the text content and provide correct answers to various questions, the semantic communication sender can use the large language model to extract key content from the input text according to user needs. Furthermore, the semantic communication receiver can use the large language model to eliminate semantic noise. Moreover, large language models excel at clearly conveying ideas and viewpoints through textual information, making text data easy to store, retrieve, and analyze.
[0147] In some embodiments, the first data is audio data. In order to ensure that the analysis of the original audio data and the effective extraction of semantic information can be achieved, the semantic knowledge base should be able to perform various audio processing tasks, including automatic speech recognition, speaker recognition and speech separation. Therefore, a large AI model used to achieve speech recognition and / or speech separation can be used as the semantic knowledge base for the audio data.
[0148] For example, if the first data is audio data, a Wav Language Model (WavLM) or Dasheng can be used as the semantic knowledge base for the audio data. This allows the semantic communication sender to separate and identify audio data from different speakers, discarding unimportant information such as background noise. Furthermore, the semantic communication receiver can perform speech denoising and recognition based on user needs, thereby achieving audio restoration. Moreover, WavLM and Dasheng are well-suited for real-time interaction and instant communication, enabling rapid and efficient information exchange.
[0149] It should be noted that regardless of the type of large-scale AI model used as the semantic knowledge base, the first communication device can achieve semantic segmentation of the first data through the large-scale AI model to obtain multiple first semantic fragments.
[0150] In some embodiments, a large AI model can be used to identify semantic targets contained in the first data, thereby segmenting the first data based on the identified semantic targets to obtain multiple first semantic segments.
[0151] In some embodiments, a large AI model can be used to comprehensively identify and segment all semantic targets in the input data to analyze the semantic information conveyed by the input data, thereby identifying each individual semantic target in the image and generating multiple semantic segments based on the identified semantic targets, each semantic segment corresponding to a semantic target.
[0152] Taking image data as an example, a large-scale AI model can segment raw or unstructured images into different semantic segments or targets, thereby better understanding the visual information conveyed by the input image.
[0153] In step S2102, the first communication device determines the semantic importance of multiple first semantic segments.
[0154] In some embodiments, the semantic importance of multiple first semantic segments can be determined using an attention network.
[0155] In some embodiments, an attention network can be used to perform max pooling and average pooling operations on multiple first semantic segments, thereby outputting the max pooling result and the average pooling result respectively through a multilayer perceptron (MLP). The max pooling result and the average pooling result are then summed to obtain the semantic importance of multiple first semantic segments.
[0156] In some embodiments, attention networks can mimic human perception to select the most attentional semantic segments through channel and spatial attention, thereby achieving attention-based semantic integration (ASI).
[0157] Furthermore, a method of manually prompting the direct acquisition of semantic segments of interest can be combined in the process of semantic importance assessment based on attention mechanism, so as to achieve the assessment of semantic importance based on the selection results of human intention.
[0158] In other words, after determining the semantic importance of multiple first semantic segments through the attention network, the determined semantic importance and the selection result of human intention can be convolved to obtain the final semantic importance of the first semantic segment.
[0159] In step S2103, the first communication device determines a first semantic segment whose semantic importance meets the requirements from the multiple first semantic segments based on the semantic importance of the multiple first semantic segments, and uses it as the first semantic segment for encoding.
[0160] In some embodiments, the semantic importance requirement may be that the semantic importance meets a first threshold. For example, the semantic importance requirement may be that the semantic importance is higher than the first threshold, or the semantic importance requirement may be that the semantic importance is higher than or equal to the first threshold, and so on, but not limited thereto.
[0161] By selecting the first semantic fragment that meets the semantic importance requirements, key semantic segments can be retained without human intervention, thereby enabling subsequent semantic perception data integration based on the retained key semantic fragments.
[0162] In some embodiments, key semantic segments can be filtered using a spatial attention network.
[0163] In some embodiments, the spatial attention network can perform max pooling and average pooling operations based on the semantic importance of multiple first semantic segments, thereby enabling the selection of key semantic segments through a convolutional neural network (CNN).
[0164] In step S2104, the first communication device performs semantic encoding based on the first semantic segment used for encoding to obtain the first semantic information.
[0165] In some embodiments, a semantic encoder can be used to perform semantic encoding based on a first semantic segment used for encoding to obtain first semantic information.
[0166] In some embodiments, a semantic encoder can be used to perform semantic encoding based on a first semantic segment used for encoding, so as to encode the first semantic segment into semantic features, and the encoded semantic features are the first semantic information.
[0167] In some embodiments, a semantic encoder can be created based on a CNN to fully utilize the excellent feature extraction capabilities of convolutional neural networks.
[0168] In some embodiments, the names of information, etc., are not limited to the names described in the embodiments. Terms such as "information", "message", "signal", "signaling", "report", "configuration", "indication", "instruction", "command", "channel", "parameter", "domain", "field", "symbol", "symbol", "codebook", "codeword", "codepoint", "bit", "data", "program", and "chip" can be used interchangeably.
[0169] In step S2105, the first communication device compresses the first semantic information to obtain the first semantic information used for channel coding.
[0170] In some embodiments, the first semantic information can be compressed using a semantic compression network to obtain the first semantic information for channel coding.
[0171] In some embodiments, a semantic compression network can be used to mask a portion of the first semantic information to obtain the first semantic information used for channel coding.
[0172] In some embodiments, different mask ratios correspond to different first data, and the mask ratio indicates the proportion of the first semantic information used for masking in the first semantic information obtained by semantic encoding based on multiple first semantic segments.
[0173] Semantic compression can further eliminate redundant semantic features, thereby significantly reducing communication overhead.
[0174] In step S2106, the first communication device performs channel coding based on the first semantic information used for channel coding to obtain the second data.
[0175] In some embodiments, the second data can be obtained by channel coding based on the first semantic information used for channel coding using a channel encoder.
[0176] In some embodiments, a channel encoder can be built based on an MLP to achieve signal encoding and modulation for the physical channel.
[0177] In step S2107, the first communication device sends the second data to the second communication device.
[0178] In some embodiments, the first communication device may send second data to the second communication device via a wireless communication channel.
[0179] In some embodiments, the second communication device may be a semantic communication receiver.
[0180] In some embodiments, the second communication device may receive second data sent by the first communication device.
[0181] In step S2108, the second communication device performs channel decoding based on the second data to obtain the second semantic information.
[0182] In some embodiments, the second semantic information can be obtained by channel decoding based on the second data using a channel decoder.
[0183] In some embodiments, a channel decoder can be built based on an MLP to demodulate and decode signals for a physical channel.
[0184] In step S2109, the second communication device performs semantic decoding based on the second semantic information to obtain multiple second semantic segments.
[0185] In some embodiments, a semantic decoder can be used to perform semantic decoding based on the second semantic information to obtain multiple second semantic segments.
[0186] In some embodiments, the second semantic information may be encoded semantic features. The second communication device may decode the second semantic information through a semantic encoder to decode the semantic features into corresponding semantic segments, thereby obtaining multiple second semantic segments.
[0187] In some embodiments, a semantic decoder can be created based on a deconvolutional neural network (DCNN) to achieve data recovery.
[0188] In step S2110, the second communication device performs semantic recovery on multiple second semantic segments based on a semantic knowledge base to obtain third data.
[0189] In some embodiments, the second communication device can use a large AI model to perform semantic recovery on multiple second semantic segments to obtain third data.
[0190] In some embodiments, the second communication device can use a large AI model to identify the semantic targets corresponding to each of the multiple second semantic segments, and then use the large AI model to perform semantic recovery on the multiple second semantic segments based on the identified semantic targets to obtain third data.
[0191] It should be noted that the solutions provided in the above embodiments can be integrated into a semantic communication framework based on a large-scale AI model. Through this semantic communication framework, semantic communication can be implemented according to the process shown in Figure 3. Referring to Figure 3, which is a flowchart illustrating a semantic communication process according to an embodiment of this disclosure, for the semantic communication sender, after obtaining the original data, semantic segmentation can be achieved through a semantic knowledge base to obtain semantic fragments. The original data can be text data, image data, or audio data. For text data, GPT can be used as the semantic knowledge base to process the text data and obtain key content. For image data, SAM can be used as the semantic knowledge base to process the image data and obtain relevant image segments. For audio data, WavLM can be used as the semantic knowledge base to process the audio data and obtain important audio. After obtaining the semantic fragments, semantic encoding can be performed using a semantic encoder to obtain the first semantic information. Optionally, text semantics can be obtained by semantically encoding key text, image semantics can be obtained by semantically encoding relevant image fragments, and audio semantics can be obtained by semantically encoding important audio. After obtaining the second semantic information, channel encoding can be performed using a channel encoder to obtain semantic data that can be used for wireless channel transmission.
[0192] For the semantic communication receiver, after receiving semantic data transmitted via a wireless channel, channel decoding can be performed using a channel decoder to obtain second semantic information (which may be text semantics, image semantics, or audio semantics). Then, semantic decoding can be performed using a semantic decoder to obtain the corresponding semantic content. Specifically, semantic decoding of text semantics yields decoded content, semantic decoding of image semantics yields decoded segments, and semantic decoding of audio semantics yields decoded audio. Finally, semantic data can be recovered using a semantic knowledge base based on the decoded content to obtain recovered data. Corresponding to the original data, the recovered data can be recovered text, recovered image, or recovered audio.
[0193] In some embodiments, taking the semantic communication process for image data as an example, the semantic communication framework based on a large AI model can implement semantic communication according to the workflow shown in Figure 4. Referring to Figure 4, which is a schematic diagram of the workflow of a semantic communication framework based on a large AI model according to an embodiment of this disclosure, after obtaining the original image, it can be encoded by the image encoder in the large AI model to obtain the image encoding result. Then, combined with the prompt information (such as the parameter matrix used for encoding) provided by the prompt encoder in the large AI model, the first semantic segment can be obtained through the mask decoder in the large AI model. Then, the first semantic segment can be subjected to max pooling and average pooling operations by an attention network, thereby outputting the max pooling result and the average pooling result by the MLP respectively. Then, the max pooling result and the average pooling result are summed to obtain the semantic importance of the first semantic segment. Then, the determined semantic importance and the selection result of human intention are convolved to obtain the final semantic importance of the first semantic segment. Then, a spatial attention network performs max pooling and average pooling operations based on the semantic importance of multiple first semantic segments. This allows for the selection of key semantic segments using a CNN. The retained key semantic segments are then used for subsequent semantic-aware data integration to obtain a semantic-aware image. After acquiring the semantic-aware image, a semantic encoder performs semantic encoding to obtain semantic features. A mask network then masks these semantic features to obtain a mask matrix. This mask matrix and semantic features are then convolved to obtain masked semantic features. A channel encoder then performs channel encoding on these masked semantic features to obtain information for wireless channel transmission. At the receiving end, a channel decoder performs channel decoding on the received information, and a semantic decoder performs semantic decoding on the channel decoding result. Finally, a large-scale AI model performs semantic recovery to obtain the reconstructed image.
[0194] The above embodiments mainly introduce the implementation process of semantic communication. For the large AI model, attention network, semantic compression network, semantic encoder and channel encoder involved in the semantic communication process, these network structures can all be pre-trained.
[0195] In some embodiments, for any network structure among a large AI model, attention network, semantic compression network, semantic encoder, and channel encoder, once the training of one network structure is completed, the network parameters of the completed network structure can be frozen to continue training other network structures.
[0196] In some embodiments, a large AI model can be obtained by rapid learning or fine-tuning training based on a pre-trained model, wherein the pre-trained model can be trained based on a general dataset.
[0197] In some embodiments, the attention network can be trained by fast learning or fine-tuning based on a pre-trained model, wherein the pre-trained model can be trained based on a general dataset.
[0198] In some embodiments, the semantic compression network can be trained with the network parameters of a large AI model and attention network frozen, and the loss function of the semantic compression network can be the difference between the first semantic information used for channel coding and the first semantic information obtained by decompression by the second communication device.
[0199] In some embodiments, the channel encoder can be trained with the network parameters of the semantic encoder frozen, wherein the loss function of the semantic encoder can be the difference between the first semantic information obtained by semantic encoding and the second semantic information obtained by the second communication device by semantic decoding, and the loss function of the channel encoder can be the difference between the second data and the third data obtained by the second communication device by channel decoding.
[0200] Referring to Figure 5, which is a schematic diagram illustrating the principle of a training process according to an embodiment of this disclosure, as shown in Figure 5, taking the training process of various network structures in a semantic communication framework for image data processing as an example, for the training process of a large AI model, the original image can be used as training data input to the large AI model, and semantic fragments can be output by the large AI model. The selection result of human prompts can be used as the relevant label of the original image, and the loss function is based on the semantic fragments output by the large AI model and the selection result of human prompts to achieve the training of the large AI model. For the encoder (including semantic encoder and channel encoder) and decoder (including semantic decoder and channel decoder), the channel encoder and channel decoder can be trained first, and then their parameters can be frozen. Then the semantic encoder and semantic decoder can be trained. Next, the semantic model parameters are frozen and the channel model is trained again. The above process is repeated until the entire SC model converges. For the training process of the attention network, the image can be recovered using the received mask semantic features. The difference between the reconstructed image and the transmitted original image (i.e., the semantically aware image) is calculated based on the loss function, and the training of the mask network is guided so that it learns how to generate the optimal mask matrix that minimizes the difference.
[0201] In some embodiments, “get,” “obtain,” “receive,” “transmit,” “bidirectional transmission,” and “send and / or receive” can be used interchangeably and can be interpreted as receiving from other entities, obtaining from protocols, obtaining from higher layers, obtaining through self-processing, or autonomous implementation, among other meanings.
[0202] In some embodiments, terms such as “send,” “transmit,” “report,” “distribute,” “transfer,” “bidirectional transmission,” “send and / or receive” can be used interchangeably.
[0203] In some embodiments, terms such as "certain," "preset," "default," "set," "indicated," "a certain," "any," and "first" can be used interchangeably. "Certain A," "preset A," "default A," "set A," "indicated A," "a certain A," "any A," and "first A" can be interpreted as A pre-defined in a protocol or the like, or as A obtained through setting, configuration, or instruction, or as specific A, a certain A, any A, or first A, but are not limited thereto.
[0204] In some embodiments, the determination or judgment can be made by a value represented by 1 bit (0 or 1), or by a true or false value (boolean), or by a comparison of numerical values (e.g., a comparison with a predetermined value), but is not limited thereto.
[0205] In some embodiments, "not expecting to receive" can be interpreted as not receiving on time domain resources and / or frequency domain resources, or as not performing subsequent processing on the data after receiving it; "not expecting to send" can be interpreted as not sending, or as sending but not expecting the receiver to respond to the sent content.
[0206] The semantic communication method disclosed in this embodiment may include at least one of steps S2101 to S2110. For example, step S2101, step S2103, step S2105, step S2110, step S2101+S2102, step S2101+S2103, step S2101+S2102+S2103, step S2101+S2103+S2104, step S2101+S2104+S2105, and step S2101+S2104+S2106 may be implemented as independent embodiments. Steps S2108+S2109+S2110 can be implemented as independent embodiments, as can steps S2101+S2102+S2103+S2104, S2101+S2104+S2106+S2107, S2101+S2104+S2106+S2107+S2108, S2101+S2104+S2106+S2107+S2108+S2109, and S2101+S2104+S2106+S2107+S2108+S2109+S2110, but are not limited thereto.
[0207] In some embodiments, steps S2102, S2103, S2104, S2105, S2106, S2107, S2108, S2109, and S2110 are optional, and one or more of these steps may be omitted or substituted in different embodiments.
[0208] In some embodiments, steps S2101, S2102, S2104, S2105, S2106, S2107, S2108, S2109, and S2110 are optional, and one or more of these steps may be omitted or substituted in different embodiments.
[0209] In some embodiments, steps S2101, S2102, S2103, S2104, S2106, S2107, S2108, S2109, and S2110 are optional, and one or more of these steps may be omitted or substituted in different embodiments.
[0210] In some embodiments, steps S2101, S2102, S2103, S2104, S2105, S2106, S2107, S2108, and S2109 are optional, and one or more of these steps may be omitted or substituted in different embodiments.
[0211] In some embodiments, other optional implementations described before or after the specification corresponding to FIG2 may be referred to.
[0212] According to the solution provided in the embodiments of this disclosure, a semantic communication framework based on a large-scale artificial intelligence model (LAM-SC) is proposed. Taking image data transmission as an example, this semantic communication framework can provide the following processing capabilities:
[0213] (1) High-precision semantic segmentation: A large semantic segmentation model based KB (SKB) is adopted. SKB can use accurate knowledge representation to segment the original or unstructured image into different semantic segments or targets. Each semantic segment or target can be selected and encoded by the sender individually, so that the sender can focus on specific semantic targets related to its communication needs.
[0214] (2) Goal-Oriented Semantic Alignment: An attention-based semantic integration (ASI) mechanism is proposed, which can accurately measure the semantic importance of fragments generated by SKB to integrate important fragments as new semantically aware data. Therefore, ASI can achieve more accurate semantic awareness, thus preserving key semantic segments without human intervention.
[0215] (3) High-quality semantic compression: A novel adaptive semantic compression (ASC) coding scheme is proposed in the semantic encoder. ASC can mask some of the transmitted semantic features, and the masking ratio can be adaptively adjusted according to the content of the transmitted features, which can further eliminate redundant semantic features and thus significantly reduce communication overhead.
[0216] In some embodiments, the design schemes of different types of semantic communication systems (i.e., text, images, audio, etc.) can all adopt the method of integrating LAM into KB creation, and the LAM inherited by different types of semantic communication systems can be different.
[0217] In some embodiments, for a text-based SC system, the knowledge base should be able to understand the content of the text and identify various topics, their attributes, and relationships. A large language model (e.g., ChatGPT) can be considered as the semantic knowledge base for the text data to extract key content from the input text according to the user's needs. At the receiving end, the received text data can be fed back to the SC decoder and fed to ChatGPT to eliminate semantic noise.
[0218] In some embodiments, for image-based SC systems, the knowledge base should be able to segment various targets in an image and identify their respective categories and relationships. SAM is a groundbreaking segmentation system that can generalize zero-shot results to unfamiliar images and targets without any additional training. Therefore, SAM can be considered as a semantic knowledge base for image data, allowing the sender to segment the input image using SAM and select important and meaningful segments for the SC encoder. Furthermore, at the receiver, the SC decoder generates reconstructed image data, which is then used with SAM to mitigate semantic noise or interference, thus effectively identifying and extracting relevant segments.
[0219] In some embodiments, to enable the SC system to support audio, the KB should be able to perform various tasks, including automatic speech recognition, speaker recognition, and speech separation, to ensure effective analysis of the raw audio data and extraction of semantic information. Optionally, WavLM or Dasheng can be considered as the semantic knowledge base for the audio data. By using WavLM or Dasheng as the KB, the sender can first separate and identify audio data from different speakers, discarding unimportant information such as background noise. The remaining audio data can then be integrated and encoded by the SC encoder. At the receiving end, the SC decoder can recover the audio, and then perform speech denoising and recognition using WavLM / Dasheng based on user needs.
[0220] In the above embodiments, SC based on WavLM or Dasheng is well-suited for real-time interaction and instant communication, enabling rapid and efficient information exchange. SC based on GPT excels at clearly conveying thoughts and opinions through textual information, making text data easy to store, retrieve, and analyze. SC systems based on SAM focus on transmitting visual information through images, capturing complex details, spatial organization, and color, as well as accurately representing facial expressions, emotions, and nonverbal cues to achieve a more intuitive communication experience.
[0221] In some embodiments, taking image data as an example, the workflow of the LAM-SC framework can be roughly divided into the following parts:
[0222] (1) Knowledge base construction and semantic segmentation: In order to achieve semantic segmentation of any raw image without KB training, SKB can be applied to comprehensively identify and segment all semantic targets in the input image. This process includes analyzing the visual information conveyed by the image to identify each individual target. Therefore, multiple segments can be generated, each containing only one semantic target.
[0223] (2) Attention-based semantic integration: ASI can simulate human perception by selecting the most noteworthy semantic segments through channel and spatial attention. In addition, it provides a method for directly obtaining semantic segments of interest through artificial prompts, which refer to the selection results based on human intent. The selected segments can be merged into a new semantic perception image.
[0224] (3) Semantic Adaptive Coding and Channel Coding: Semantically perceptual images are encoded into semantic features by a semantic encoder. Here, the semantic encoder is built on CNN, which has excellent image feature extraction capabilities and can adaptively mask unimportant features in the semantic information according to the content of the semantic information. In addition, a channel encoder based on MLP can be used to encode and modulate signals in the physical channel.
[0225] (4) Channel Decoding and Semantic Decoding: This function enables signal demodulation and decoding, acquiring the semantic features of the transmitted signal as it reaches the receiver via the wireless physical channel. The channel decoder uses an MLP structure to obtain semantic features. The semantic decoder, composed of a DCNN, decodes the semantic features of the image, then recovers the image data. Subsequently, SKB can be used again on the recovered image to accurately identify the target.
[0226] In some embodiments, the LAM-SC framework can be pre-trained, and the training method is roughly as follows:
[0227] (1) Human Experience-Based ASI Training: The purpose of ASI is to mimic human perception, identifying targets of interest in raw images and then generating semantically perceptual images that conform to human preferences. To achieve this goal, semantics of interest to humans are recorded as experiences, which constitute the foundation for training the semantic attention network, including both channel and spatial attention networks. In this experience database, semantic segments can serve as input samples for the attention network, while selection results based on human cues can be seen as relevant labels. By using supervised learning on the experience database, the attention network can effectively adapt to human behavior and make decisions very similar to human perception.
[0228] (2) Cross-training of Encoder and Decoder Based on SC: The SC encoder consists of a semantic decoder and a channel decoder, and the SC decoder consists of a channel decoder and a semantic decoder. First, the difference between the original semantic features and the recovered semantic features is used as a loss function to guide the training of the channel encoder and decoder. Then, the difference between the original image and the recovered image is used as a loss function to guide the learning and training of the semantic encoder and decoder. A cross-training strategy involving the channel encoder / decoder and the semantic encoder / decoder model is proposed. More specifically, the channel encoder / decoder is trained first, then its parameters are frozen, and then the semantic encoder / decoder is trained. Next, the semantic model parameters are frozen and the channel model is trained again. This process can be repeated until the entire SC model converges.
[0229] (3) ASC Training: To generate a mask array that accurately reflects the importance of semantic features, a joint training method for the mask network and the SC model (i.e., the SC encoder / decoder) is proposed, where the parameters of the SC model and the attention network are frozen. The process includes the following steps: transmitting the mask semantic features compressed by the mask network; reconstructing the image using the received mask semantic features; calculating the difference between the reconstructed image and the transmitted original image (i.e., the semantically aware image) based on the loss function, and guiding the training of the mask network to learn how to generate the optimal mask matrix that minimizes this difference.
[0230] (4) LAM can be pre-trained on a wide range of datasets using self-supervised learning with unlabeled data, and the pre-trained model can be applied to a variety of tasks by learning quickly or fine-tuning.
[0231] The solutions provided by the embodiments of this disclosure enable semantic knowledge bases to provide accurate knowledge representations, rich prior / background knowledge, and low-cost knowledge update capabilities. Specifically, the semantic knowledge base (LAM) possesses trillions of parameters, allowing it to learn complex knowledge representations from transformer models with multi-head attention mechanisms. These multi-head attention mechanisms develop a powerful understanding of semantics and knowledge structures, thus enabling LAMs to provide high-quality semantic representations of input data. Furthermore, LAMs are pre-trained on extensive datasets, enabling them to learn from vast amounts of information across various domains and store rich prior / background knowledge. They exhibit significant generalization capabilities, achieving high performance on various tasks, even exceeding the knowledge domain of their pre-training, thereby eliminating the need for frequent KB updates. Additionally, LAMs typically come with pre-trained weights, allowing for prompting using only a few examples or fine-tuning with a small amount of labeled data. This alleviates concerns about frequent knowledge updates and insecure knowledge sharing.
[0232] Figure 6A is a flowchart illustrating a semantic communication method according to an embodiment of the present disclosure. As shown in Figure 6A, the embodiments of the present disclosure relate to a semantic communication method, which includes:
[0233] Step S6101: Based on the semantic knowledge base, perform semantic segmentation on the first data to obtain multiple first semantic segments.
[0234] The optional implementation of step S6101 can be found in the optional implementation of step S2101 in Figure 2 and other related parts in the embodiments involved in Figure 2, which will not be repeated here.
[0235] Step S6102: Determine the semantic importance of multiple first semantic segments.
[0236] The optional implementation of step S6102 can be found in the optional implementation of step S2102 in Figure 2 and other related parts in the embodiments involved in Figure 2, which will not be repeated here.
[0237] Step S6103: Based on the semantic importance of multiple first semantic segments, determine the first semantic segment whose semantic importance meets the requirements from the multiple first semantic segments, and use it as the first semantic segment for encoding.
[0238] The optional implementation of step S6103 can be found in the optional implementation of step S2103 in Figure 2 and other related parts in the embodiments involved in Figure 2, which will not be repeated here.
[0239] Step S6104: Semantic encoding is performed based on the first semantic segment used for encoding to obtain the first semantic information.
[0240] The optional implementation of step S6104 can be found in the optional implementation of step S2104 in Figure 2 and other related parts in the embodiments involved in Figure 2, which will not be repeated here.
[0241] Step S6105: Compress the first semantic information to obtain the first semantic information used for channel coding.
[0242] The optional implementation of step S6105 can be found in the optional implementation of step S2105 in Figure 2 and other related parts in the embodiments involved in Figure 2, which will not be repeated here.
[0243] Step S6106: Channel coding is performed based on the first semantic information used for channel coding to obtain the second data.
[0244] The optional implementation of step S6106 can be found in the optional implementation of step S2106 in Figure 2 and other related parts in the embodiments involved in Figure 2, which will not be repeated here.
[0245] Step S6107: Send the second data.
[0246] The optional implementation of step S6107 can be found in the optional implementation of step S2107 in Figure 2 and other related parts in the embodiments involved in Figure 2, which will not be repeated here.
[0247] In some embodiments, the first communication device sends second data to the second communication device, but is not limited thereto; it may also send second data to other entities.
[0248] The semantic communication method involved in the embodiments of this disclosure may include at least one of steps S6101 to S6110. For example, step S6101 can be implemented as an independent embodiment, step S6103 can be implemented as an independent embodiment, step S6105 can be implemented as an independent embodiment, step S6101+S6102 can be implemented as an independent embodiment, step S6101+S6103 can be implemented as an independent embodiment, step S6101+S6102+S6103 can be implemented as an independent embodiment, step S6101+S6103+S6104 can be implemented as an independent embodiment, step S6101+S6104+S6105 can be implemented as an independent embodiment, step S6101+S6104+S6106 can be implemented as an independent embodiment, step S6101+S6102+S6103+S6104 can be implemented as an independent embodiment, and step S6101+S6104+S6106+S6107 can be implemented as an independent embodiment, but are not limited thereto.
[0249] In some embodiments, steps S6102, S6103, S6104, S6105, S6106, and S6107 are optional, and one or more of these steps may be omitted or substituted in different embodiments.
[0250] In some embodiments, steps S6101, S6102, S6104, S6105, S6106, and S6107 are optional, and one or more of these steps may be omitted or substituted in different embodiments.
[0251] In some embodiments, steps S6101, S6102, S6103, S6104, S6106, and S6107 are optional, and one or more of these steps may be omitted or substituted in different embodiments.
[0252] Figure 6B is a flowchart illustrating a semantic communication method according to an embodiment of the present disclosure. As shown in Figure 6B, the present disclosure relates to a semantic communication method, which includes:
[0253] Step S6201: Obtain the second data.
[0254] The optional implementation of step S6201 can be found in the optional implementations of steps S2101 to S2107 in Figure 2, as well as other related parts in the embodiments involved in Figure 2, which will not be repeated here.
[0255] In some embodiments, the second communication device receives second data sent by the first communication device, but is not limited thereto; it may also receive second data sent by other entities.
[0256] In some embodiments, the second communication device acquires second data as defined by the protocol.
[0257] In some embodiments, the second communication device obtains second data from the upper layer(s).
[0258] In some embodiments, the second communication device processes the data to obtain the second data.
[0259] In some embodiments, step S6201 is omitted, and the second communication device autonomously implements the function indicated by the second data, or the above function is default or default.
[0260] Step S6202: Channel decoding is performed based on the second data to obtain the second semantic information.
[0261] The optional implementation of step S6202 can be found in the optional implementation of step S2108 in Figure 2 and other related parts in the embodiments involved in Figure 2, which will not be repeated here.
[0262] Step S6203: Semantic decoding is performed based on the second semantic information to obtain multiple second semantic segments.
[0263] The optional implementation of step S6203 can be found in the optional implementation of step S2109 in Figure 2 and other related parts in the embodiments involved in Figure 2, which will not be repeated here.
[0264] Step S6204: Based on the semantic knowledge base, perform semantic recovery on multiple second semantic segments to obtain third data.
[0265] The optional implementation of step S6204 can be found in the optional implementation of step S2110 in Figure 2 and other related parts in the embodiments involved in Figure 2, which will not be repeated here.
[0266] The semantic communication method involved in the embodiments of this disclosure may include at least one of steps S6201 to S6204. For example, step S6204 may be implemented as an independent embodiment, step S6201+S6204 may be implemented as an independent embodiment, step S6202+S6204 may be implemented as an independent embodiment, step S6203+S6204 may be implemented as an independent embodiment, step S6201+S6202+S6204 may be implemented as an independent embodiment, and step S6201+S6203+S6204 may be implemented as an independent embodiment, but is not limited thereto.
[0267] In some embodiments, steps S6201, S6202, and S6203 are optional, and one or more of these steps may be omitted or substituted in different embodiments.
[0268] In this embodiment of the disclosure, step S6201 can be combined with step S6107 of FIG6A.
[0269] Figure 7A is a flowchart illustrating a semantic communication method according to an embodiment of the present disclosure. As shown in Figure 7A, the embodiments of the present disclosure relate to a semantic communication method, which includes:
[0270] Step S7101: Based on the semantic knowledge base, perform semantic segmentation on the first data to obtain multiple first semantic segments.
[0271] The optional implementation of step S7101 can be found in step S2101 of Figure 2, the optional implementation of step S6101 of Figure 6A, and other related parts in the embodiments involved in Figures 2 and 6A, which will not be repeated here.
[0272] In some embodiments, the semantic knowledge base is a large AI model.
[0273] In some embodiments, the first communication device identifies semantic targets contained in the first data through a large AI model, and then segments the first data based on the identified semantic targets using the large AI model to obtain multiple first semantic segments.
[0274] In some embodiments, the first data is image data, and the semantic knowledge base is a large AI model used to implement image recognition and / or image classification.
[0275] In some embodiments, the first data is text data, and the semantic knowledge base is a large AI model used to implement text content recognition and / or human-computer dialogue.
[0276] In some embodiments, the first data is audio data, and the semantic knowledge base is a large AI model used to implement speech recognition and / or speech separation.
[0277] Step S7102: Encode multiple first semantic segments to obtain second data.
[0278] The optional implementation of step S7102 can be found in the optional implementations of steps S2102 to S2106 in Figure 2, steps S6102 to S6106 in Figure 6A, and other related parts in the embodiments involved in Figure 2 and Figure 6A, which will not be repeated here.
[0279] In some embodiments, the first communication device performs semantic encoding based on multiple first semantic segments using a semantic encoder to obtain first semantic information; and then performs channel encoding based on the first semantic information using a channel encoder to obtain second data.
[0280] In some embodiments, the first communication device may further compress the first semantic information through a semantic compression network to obtain the first semantic information for channel coding.
[0281] In some embodiments, the first communication device may use a semantic compression network to mask a portion of the first semantic information to obtain the first semantic information for channel coding.
[0282] In some embodiments, different mask ratios correspond to different first data, and the mask ratio indicates the proportion of the first semantic information used for masking in the first semantic information obtained by semantic encoding based on multiple first semantic segments.
[0283] In some embodiments, the first communication device may further determine, based on the semantic importance of multiple first semantic segments, a first semantic segment whose semantic importance meets the requirements, and use it as the first semantic segment for encoding.
[0284] In some embodiments, the first communication device may also determine the semantic importance of multiple first semantic segments through an attention network.
[0285] Step S7103: Send the second data.
[0286] The optional implementation of step S7103 can be found in step S2107 of Figure 2, the optional implementation of step S6107 of Figure 6A, and other related parts in the embodiments involved in Figures 2 and 6A, which will not be repeated here.
[0287] In some embodiments, the first communication device sends second data to the second communication device, but is not limited thereto; it may also send second data to other entities.
[0288] In some embodiments, the second data is used by the second communication device for data decoding and semantic recovery.
[0289] In some embodiments, the large AI model, attention network, semantic compression network, semantic encoder, and channel encoder are all pre-trained; for any network structure among the large AI model, attention network, semantic compression network, semantic encoder, and channel encoder, once the network structure has been trained, the network parameters of the trained network structure are frozen to continue training other network structures.
[0290] In some embodiments, the large AI model and the attention network are obtained by rapid learning or fine-tuning training based on a pre-trained model, wherein the pre-trained model is trained based on a general dataset.
[0291] In some embodiments, the semantic compression network is trained with the network parameters of the large AI model and the attention network frozen, and the loss function of the semantic compression network is the difference between the first semantic information used for channel coding and the first semantic information obtained by decompression by the second communication device.
[0292] In some embodiments, the channel encoder is trained with the network parameters of the semantic encoder frozen, wherein the loss function of the semantic encoder is the difference between the first semantic information obtained by semantic encoding and the second semantic information obtained by the second communication device by semantic decoding, and the loss function of the channel encoder is the difference between the second data and the third data obtained by the second communication device by channel decoding.
[0293] The semantic communication method involved in the embodiments of this disclosure may include at least one of steps S7101 to S7103. For example, step S7101 may be implemented as a standalone embodiment, step S7101+S7102 may be implemented as a standalone embodiment, and step S7101+S7103 may be implemented as a standalone embodiment, but is not limited thereto.
[0294] In some embodiments, steps S7102 and S7103 are optional, and one or more of these steps may be omitted or substituted in different embodiments.
[0295] Figure 7B is a flowchart illustrating a semantic communication method according to an embodiment of the present disclosure. As shown in Figure 7B, the embodiments of the present disclosure relate to a semantic communication method, which includes:
[0296] Step S7201: Obtain the second data.
[0297] The optional implementation of step S7201 can be found in steps S2101 to S2107 in Figure 2, the optional implementation of step S6201 in Figure 6B, and other related parts in the embodiments involved in Figures 2 and 6B, which will not be repeated here.
[0298] In some embodiments, the second communication device receives second data sent by the first communication device, but is not limited thereto; it may also receive second data sent by other entities.
[0299] In some embodiments, the second communication device acquires second data as defined by the protocol.
[0300] In some embodiments, the second communication device obtains second data from the upper layer(s).
[0301] In some embodiments, the second communication device processes the data to obtain the second data.
[0302] In some embodiments, step S7201 is omitted, and the second communication device autonomously implements the function indicated by the second data, or the above function is default or set to default.
[0303] Step S7202: Decode the second data to obtain multiple second semantic segments.
[0304] The optional implementations of step S7202 can be found in steps S2108 and S2109 in Figure 2, steps S6202 and S6203 in Figure 6B, and other related parts in the embodiments involved in Figures 2 and 6B, which will not be repeated here.
[0305] In some embodiments, the second communication device performs channel decoding based on the second data using a channel decoder to obtain second semantic information; and then performs semantic decoding based on the second semantic information using a semantic decoder to obtain multiple second semantic segments.
[0306] Step S7203: Based on the semantic knowledge base, perform semantic recovery on multiple second semantic segments to obtain third data.
[0307] The optional implementation of step S7203 can be found in the optional implementation of step S2110 in Figure 2, step S6204 in Figure 6B, and other related parts in the embodiments involved in Figures 2 and 6B, which will not be repeated here.
[0308] In some embodiments, the semantic knowledge base is a large AI model.
[0309] In some embodiments, the second communication device uses a large-scale AI model to identify the semantic targets corresponding to each of the multiple second semantic segments; thereby, based on the identified semantic targets, the large-scale AI model performs semantic recovery on the multiple second semantic segments to obtain third data.
[0310] In some embodiments, the first data is image data, and the semantic knowledge base is a large AI model used to implement image recognition and / or image classification.
[0311] In some embodiments, the first data is text data, and the semantic knowledge base is a large AI model used to implement text content recognition and / or human-computer dialogue.
[0312] In some embodiments, the first data is audio data, and the semantic knowledge base is a large AI model used to implement speech recognition and / or speech separation.
[0313] In some embodiments, the large AI model, the semantic decoder, and the channel decoder are all pre-trained; for any network structure among the large AI model, the semantic decoder, and the decoding encoder, once the network structure has been trained, the network parameters of the trained network structure are frozen to continue training other network structures.
[0314] In some embodiments, the large AI model is obtained by rapid learning or fine-tuning training based on a pre-trained model, wherein the pre-trained model is trained based on a general dataset.
[0315] In some embodiments, the channel decoder is trained with the network parameters of the semantic decoder frozen, wherein the loss function of the semantic decoder is the difference between the second semantic information obtained by semantic decoding and the first semantic information obtained by the first communication device by semantic encoding, and the loss function of the channel decoder is the difference between the second data and the third data.
[0316] The semantic communication method involved in the embodiments of this disclosure may include at least one of steps S7201 to S7203. For example, step S7203 may be implemented as a standalone embodiment, step S7201+S7203 may be implemented as a standalone embodiment, and step S7202+S7203 may be implemented as a standalone embodiment, but is not limited thereto.
[0317] In some embodiments, steps S7201 and S7202 are optional, and one or more of these steps may be omitted or substituted in different embodiments.
[0318] In this embodiment of the disclosure, step S7201 can be combined with step S7103 of FIG7A.
[0319] In the embodiments disclosed herein, some or all of the steps and their optional implementations may be arbitrarily combined with some or all of the steps in other embodiments, or may be arbitrarily combined with the optional implementations in other embodiments.
[0320] This disclosure also provides embodiments of an apparatus for implementing any of the above methods. For example, an apparatus is provided that includes units or modules for implementing the steps performed by the first communication device in any of the above methods. Furthermore, another apparatus is provided that includes units or modules for implementing the steps performed by the second communication device in any of the above methods.
[0321] It should be understood that the division of units or modules in the above device is only a logical functional division. In actual implementation, they can be fully or partially integrated into a single physical entity, or they can be physically separated. Furthermore, the units or modules in the device can be implemented by a processor calling software: for example, the device includes a processor connected to a memory containing instructions. The processor calls the instructions stored in the memory to implement any of the above methods or to implement the functions of the units or modules in the above device. The processor can be, for example, a general-purpose processor, such as a Central Processing Unit (CPU) or a microprocessor, and the memory can be internal or external to the device. Alternatively, the units or modules in the device can be implemented in the form of hardware circuits. The functionality of some or all of the units or modules can be achieved through the design of these hardware circuits, which can be understood as one or more processors. For example, in one implementation, the hardware circuit is an application-specific integrated circuit (ASIC). The functionality of some or all of the units or modules is achieved through the design of the logical relationships between the components within the circuit. In another implementation, the hardware circuit can be implemented using a programmable logic device (PLD). Taking a field-programmable gate array (FPGA) as an example, it can include a large number of logic gates. The connection relationships between the logic gates are configured through configuration files, thereby achieving the functionality of some or all of the units or modules. All units or modules of the above device can be implemented entirely through processor-called software, entirely through hardware circuits, or partially through processor-called software with the remaining parts implemented through hardware circuits.
[0322] In this embodiment, the processor is a circuit with signal processing capabilities. In one implementation, the processor can be a circuit with instruction read and execute capabilities, such as a Central Processing Unit (CPU), a microprocessor, a graphics processing unit (GPU) (which can be understood as a microprocessor), or a digital signal processor (DSP). In another implementation, the processor can implement certain functions through the logical relationships of hardware circuits. The logical relationships of the aforementioned hardware circuits are fixed or reconfigurable. For example, the processor is a hardware circuit implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD), such as an FPGA. In a reconfigurable hardware circuit, the process of the processor loading a configuration document and configuring the hardware circuit can be understood as the process of the processor loading instructions to implement the functions of some or all of the above units or modules. Furthermore, it can also be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC, such as a Neural Network Processing Unit (NPU), a Tensor Processing Unit (TPU), or a Deep Learning Processing Unit (DPU).
[0323] Figure 8A is a schematic diagram of the structure of the first communication device proposed in an embodiment of this disclosure. As shown in Figure 8A, the first communication device 8100 may include at least one of a processing module 8101 and a transceiver module 8102. In some embodiments, the processing module 8101 is configured to perform semantic segmentation on the first data based on a semantic knowledge base to obtain multiple first semantic segments; the processing module 8101 is also configured to encode the multiple first semantic segments to obtain second data; the transceiver module 8102 is configured to send the second data to a second communication device, the second data being used by the second communication device for data decoding and semantic recovery. Optionally, the processing module 8101 may execute at least one of the other steps (e.g., steps S2101, S2102, S2103, S2104, S2105, S2106, but not limited thereto) executed by the first communication device in any of the above methods, which will not be elaborated here. Optionally, the transceiver module 8102 is used to perform at least one of the communication steps (such as step S2107, but not limited thereto) performed by the first communication device in any of the above methods, which will not be described in detail here.
[0324] Figure 8B is a schematic diagram of the structure of the second communication device proposed in an embodiment of this disclosure. As shown in Figure 8B, the second communication device 8200 may include at least one of a transceiver module 8201, a processing module 8202, etc. In some embodiments, the transceiver module 8201 is configured to receive second data sent by the first communication device; the processing module 8202 is configured to decode the second data to obtain a plurality of second semantic segments; the processing module 8202 is further configured to perform semantic recovery on the plurality of second semantic segments based on a semantic knowledge base to obtain third data. Optionally, the transceiver module is used to perform at least one of the communication steps (e.g., step S2107, but not limited thereto) performed by the second communication device in any of the above methods, which will not be described in detail here. Optionally, the processing module is used to perform at least one of the other steps (e.g., steps S2108, S2109, S2110, but not limited thereto) performed by the second communication device in any of the above methods, which will not be described in detail here.
[0325] In some embodiments, the transceiver module may include a transmitting module and / or a receiving module, which may be separate or integrated. Optionally, the transceiver module may be interchangeable with a transceiver.
[0326] In some embodiments, the processing module may be a single module or may include multiple sub-modules. Optionally, the multiple sub-modules may each perform all or part of the steps required by the processing module. Optionally, the processing module may be interchangeable with a processor.
[0327] Figure 9A is a schematic diagram of the structure of the communication device 9100 proposed in an embodiment of this disclosure. The communication device 9100 can be a first communication device, a second communication device, a chip, chip system, or processor that supports the first communication device in implementing any of the above methods, or a chip, chip system, or processor that supports the second communication device in implementing any of the above methods. The communication device 9100 can be used to implement the methods described in the above method embodiments; for details, please refer to the descriptions in the above method embodiments.
[0328] As shown in Figure 9A, the communication device 9100 includes one or more processors 9101. The processor 9101 can be a general-purpose processor or a dedicated processor, such as a baseband processor or a central processing unit (CPU). The baseband processor can be used to process communication protocols and communication data, while the CPU can be used to control communication devices (e.g., base stations, baseband chips, terminal devices, terminal device chips, DUs or CUs, etc.), execute programs, and process program data. The communication device 9100 is used to execute any of the above methods.
[0329] In some embodiments, the communication device 9100 further includes one or more memories 9102 for storing instructions. Optionally, all or part of the memories 9102 may also be located outside the communication device 9100.
[0330] In some embodiments, the communication device 9100 further includes one or more transceivers 9103. When the communication device 9100 includes one or more transceivers 9103, the transceivers 9103 perform at least one of the communication steps such as sending and / or receiving in the above method (e.g., step S2107, but not limited thereto), and the processor 9101 performs at least one of the other steps (e.g., steps S2101, S2102, S2103, S2104, S2105, S2106, S2108, S2109, S2110, but not limited thereto).
[0331] In some embodiments, a transceiver may include a receiver and / or a transmitter, which may be separate or integrated. Optionally, the terms transceiver, transceiver unit, transceiver, transceiver circuit, etc., may be used interchangeably; the terms transmitter, transmitting unit, transmitter, transmitting circuit, etc., may be used interchangeably; and the terms receiver, receiving unit, receiver, receiving circuit, etc., may be used interchangeably.
[0332] In some embodiments, the communication device 9100 may include one or more interface circuits 9104. Optionally, the interface circuit 9104 is connected to the memory 9102, and the interface circuit 9104 can be used to receive signals from the memory 9102 or other devices, and can be used to send signals to the memory 9102 or other devices. For example, the interface circuit 9104 can read instructions stored in the memory 9102 and send the instructions to the processor 9101.
[0333] The communication device 9100 described in the above embodiments may be a network device or a terminal, but the scope of the communication device 9100 described in this disclosure is not limited thereto, and the structure of the communication device 9100 may not be limited by FIG. 9A. The communication device may be a standalone device or a part of a larger device. For example, the communication device may be: (1) a standalone integrated circuit IC, or chip, or chip system or subsystem; (2) a collection of one or more ICs, optionally, the IC collection may also include storage components for storing data and programs; (3) an ASIC, such as a modem; (4) a module that can be embedded in other devices; (5) a receiver, terminal device, smart terminal device, cellular phone, wireless device, handheld device, mobile unit, vehicle device, network device, cloud device, artificial intelligence device, etc.; (6) others, etc.
[0334] Figure 9B is a schematic diagram of the structure of the chip 9200 proposed in an embodiment of this disclosure. For cases where the communication device 9100 can be a chip or a chip system, please refer to the schematic diagram of the chip 9200 shown in Figure 9B, but it is not limited thereto.
[0335] Chip 9200 includes one or more processors 9201, which are used to perform any of the above methods.
[0336] In some embodiments, chip 9200 further includes one or more interface circuits 9202. Optionally, the interface circuit 9202 is connected to memory 9203, and the interface circuit 9202 can be used to receive signals from memory 9203 or other devices, and the interface circuit 9202 can be used to send signals to memory 9203 or other devices. For example, the interface circuit 9202 can read instructions stored in memory 9203 and send the instructions to processor 9201.
[0337] In some embodiments, the interface circuit 9202 performs at least one of the communication steps such as sending and / or receiving in the above method (e.g., step S2107, but not limited thereto), and the processor 9201 performs at least one of the other steps (e.g., steps S2101, S2102, S2103, S2104, S2105, S2106, S2108, S2109, S2110, but not limited thereto).
[0338] In some embodiments, the terms interface circuit, interface, transceiver pin, transceiver, etc., can be used interchangeably.
[0339] In some embodiments, chip 9200 further includes one or more memories 9203 for storing instructions. Optionally, all or part of the memories 9203 may be located outside of chip 9200.
[0340] This disclosure also proposes a storage medium storing instructions that, when executed on the communication device 9100, cause the communication device 9100 to perform any of the above methods. Optionally, the storage medium is an electronic storage medium. Optionally, the storage medium is a computer-readable storage medium, but not limited thereto; it may also be a storage medium readable by other devices. Optionally, the storage medium may be a non-transitory storage medium, but not limited thereto; it may also be a temporary storage medium.
[0341] This disclosure also provides a program product that, when executed by the communication device 9100, causes the communication device 9100 to perform any of the above methods. Optionally, the program product is a computer program product.
[0342] This disclosure also proposes a computer program that, when run on a computer, causes the computer to perform any of the above methods.
[0343] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.
[0344] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. A semantic communication method, characterized in that, Applied to a first communication device, the method includes: Based on the semantic knowledge base, semantic segmentation is performed on the first data to obtain multiple first semantic segments; Encode the multiple first semantic segments to obtain second data; The second data is sent to the second communication device, and the second data is used by the second communication device to perform data decoding and semantic recovery.
2. The method of claim 1, wherein, The semantic knowledge base is a large-scale artificial intelligence (AI) model; The first data is semantically segmented based on a semantic knowledge base to obtain multiple first semantic segments, including: Identify semantic targets contained in the first data using a large AI model; The first data is segmented using a large-scale AI model based on the identified semantic targets to obtain the multiple first semantic segments.
3. The method according to claim 1 or 2, characterized in that, The process of encoding the plurality of first semantic segments to obtain second data includes: First semantic information is obtained by semantically encoding the plurality of first semantic segments using a semantic encoder. The second data is obtained by channel coding based on the first semantic information using a channel encoder.
4. The method according to claim 3, characterized in that, The method further includes: The first semantic information is compressed using a semantic compression network to obtain the first semantic information used for channel coding.
5. The method according to claim 4, characterized in that, The step of compressing the first semantic information using a semantic compression network to obtain first semantic information for channel coding includes: A portion of the first semantic information is masked using a semantic compression network to obtain the first semantic information used for channel coding.
6. The method of claim 5, wherein, Different first data correspond to different mask ratios, and the mask ratio indicates the proportion of the first semantic information used for masking in the first semantic information obtained by semantic encoding based on the multiple first semantic segments.
7. The method according to any one of claims 1 to 6, characterized in that, The method further includes: Based on the semantic importance of the plurality of first semantic segments, a first semantic segment whose semantic importance meets the requirements is determined from the plurality of first semantic segments and used as the first semantic segment for encoding.
8. The method according to claim 7, characterized in that, The method further includes: The semantic importance of the plurality of first semantic segments is determined by an attention network.
9. The method according to any one of claims 1 to 8, characterized in that, The first data is image data, and the semantic knowledge base is a large-scale AI model used to implement image recognition and / or image classification; and / or, The first data is text data, and the semantic knowledge base is a large-scale AI model used to achieve text content recognition and / or human-computer dialogue; and / or, The first data is audio data, and the semantic knowledge base is a large-scale AI model used to achieve speech recognition and / or speech separation.
10. The method according to any one of claims 2 to 9, characterized in that, The large AI model, the attention network, the semantic compression network, the semantic encoder, and the channel encoder are all pre-trained. For any network structure among the large AI model, the attention network, the semantic compression network, the semantic encoder, and the channel encoder, once the network structure has been trained, the network parameters of the trained network structure are frozen to continue training other network structures.
11. The method according to claim 10, characterized in that, The large-scale AI model and the attention network are obtained through rapid learning or fine-tuning training based on a pre-trained model, wherein the pre-trained model is trained based on a general dataset.
12. The method according to claim 10 or 11, characterized in that, The semantic compression network is trained with the network parameters of the large AI model and the attention network frozen, and the loss function of the semantic compression network is the difference between the first semantic information used for channel coding and the first semantic information obtained by decompression by the second communication device.
13. The method according to any one of claims 10 to 12, characterized in that, The channel encoder is trained with the network parameters of the semantic encoder frozen. The loss function of the semantic encoder is the difference between the first semantic information obtained by semantic encoding and the second semantic information obtained by the second communication device by semantic decoding. The loss function of the channel encoder is the difference between the second data and the third data obtained by the second communication device by channel decoding.
14. A semantic communication method, characterized in that, Applied to a second communication device, the method includes: Receive the second data sent by the first communication device; Decoding is performed based on the second data to obtain multiple second semantic segments; Based on the semantic knowledge base, semantic recovery is performed on the multiple second semantic segments to obtain the third data.
15. The method according to claim 14, characterized in that, The decoding based on the second data yields multiple second semantic segments, including: The second semantic information is obtained by channel decoding based on the second data using a channel decoder. The semantic decoder performs semantic decoding based on the second semantic information to obtain the plurality of second semantic segments.
16. The method according to claim 14 or 15, characterized in that, The semantic knowledge base is a large-scale AI model; The process of semantically restoring the multiple second semantic segments based on a semantic knowledge base to obtain third data includes: Using a large-scale AI model, the semantic targets corresponding to each of the multiple second semantic segments are identified; Using a large-scale AI model, semantic recovery is performed on the multiple second semantic segments based on the identified semantic targets to obtain the third data.
17. The method according to any one of claims 14 to 16, characterized in that, The second data is image data, and the semantic knowledge base is a large-scale AI model used to achieve image recognition and / or image classification; and / or, The second data is text data, and the semantic knowledge base is a large-scale AI model used to achieve text content recognition and / or human-computer dialogue; and / or, The second data is audio data, and the semantic knowledge base is a large AI model used to achieve speech recognition and / or speech separation.
18. The method according to any one of claims 15 to 17, characterized in that, The large AI model, the semantic decoder, and the channel decoder are all pre-trained. For any network structure among the large AI model, the semantic decoder, and the decoding encoder, once the network structure has been trained, the network parameters of the trained network structure are frozen to continue training other network structures.
19. The method of claim 18, wherein, The large-scale AI model is obtained by rapid learning or fine-tuning training based on a pre-trained model, wherein the pre-trained model is trained based on a general dataset.
20. The method of claim 18 or 19, wherein, The channel decoder is trained with the network parameters of the semantic decoder frozen. The loss function of the semantic decoder is the difference between the second semantic information obtained by semantic decoding and the first semantic information obtained by the first communication device by semantic encoding. The loss function of the channel decoder is the difference between the second data and the third data.
21. A first communication device, characterized by include: The processing module is configured to perform semantic segmentation on the first data based on a semantic knowledge base to obtain multiple first semantic fragments; The processing module is further configured to encode the plurality of first semantic segments to obtain second data; The transceiver module is configured to send the second data to the second communication device, and the second data is used by the second communication device for data decoding and semantic recovery.
22. A second communication device, characterized in that, include: The transceiver module is configured to receive second data sent by the first communication device; The processing module is configured to decode based on the second data to obtain multiple second semantic segments; The processing module is also configured to perform semantic recovery on the multiple second semantic segments based on a semantic knowledge base to obtain third data.
23. A first communication device, characterized in that, include: One or more processors; The first communication device is used to execute the semantic communication method according to any one of claims 1-13.
24. A second communication device, characterized in that, include: One or more processors; The second communication device is used to execute the semantic communication method according to any one of claims 14-20.
25. A communication system, characterized by The device includes a first communication device and a second communication device, wherein the first communication device is configured to implement the semantic communication method of any one of claims 1-13, and the second communication device is configured to implement the communication method of any one of claims 14-20.
26. A storage medium storing instructions, characterized in that, When the instruction is executed on a communication device, the communication device performs the semantic communication method as described in any one of claims 1-13 or 14-20.
Citation Information
Patent Citations
Generative multi-mode mutual benefit enhancement video semantic communication method
CN116939320A
Model training method, semantic communication transmission method and model training device
CN117524203A
Semantic communication coding and decoding method and device, equipment and storage medium
CN118540024A
Data compression with controllable semantic loss
EP4436048A1
Semantic Communication: Protocol Stack and Model Selection
US20230412709A1